Voice processing system, voice processing apparatus, and voice processing method
The audio processing system on vehicles with multiple seats addresses the issue of unspecified speaker positions by using an input and determination unit to restrict or prioritize voice commands, ensuring reliable and appropriate processing.
Patent Information
- Application Number
- JP2022550338
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-18
- Filing Date
- 2021-04-20
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2041-04-20
AI Technical Summary
Existing audio processing systems fail to execute appropriate processing when the position of the speaker cannot be specified, leading to potential unintended operations.
An audio processing system mounted on a vehicle with multiple seats, incorporating an input unit, determination unit, and audio recognition unit, which determines the seated speaker's position and restricts or prioritizes voice commands based on microphone status and speaker identification.
Ensures appropriate processing by restricting unintended operations and prioritizing commands, even when the speaker's position cannot be identified, thereby enhancing system reliability.
Smart Images

Figure 0007702212000001 
Figure 0007702212000002 
Figure 0007702212000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an audio processing system, an audio processing apparatus, and an audio processing method.
Background Art
[0002] An audio processing system that processes an audio recognition command based on audio spoken by a speaker is known.
[0003] Patent Document 1 discloses an audio processing system that processes an audio recognition command based on the position where the speaker speaks.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
[0005] However, Patent Document 1 does not disclose control in the case where the position of the speaker cannot be specified. If the case where the position of the speaker cannot be specified is not assumed, the audio processing system may execute unintended processing.
[0006] An object of the present disclosure is to execute appropriate processing even when the position of the speaker cannot be specified in an audio processing system.
[0007] The audio processing system according to the present disclosure An audio processing system mounted on a vehicle having a plurality of seats, includes an input unit, a determination unit, and an audio recognition unit. The input unit receives first audio that is audio spoken by a first speaker. The determination unit determines whether the first speaker the seated seat can be specified. The audio recognition unit is an audio recognition unit that outputs an audio command, which is a signal that is specified by audio and controls a target device, to the target device, and when the determination unit determines that the first speaker the seated seatWhen it is determined that the speaker cannot be identified, among the voice commands, the output of the the seated seat utterance position command for the speaker a command for which a process to be executed is determined according to is 、 restricted.
[0008] According to the present disclosure, in a voice processing system, appropriate processing can be executed even when the position of the speaker cannot be identified. BRIEF DESCRIPTION OF THE DRAWINGS
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings as appropriate. However, a more detailed description than necessary may be omitted. Note that the accompanying drawings and the following description are provided for those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims.
[0011] (First Embodiment) FIG. 1 is a diagram showing an example of the schematic configuration of the voice system 5 in the first embodiment. The voice system 5 is mounted on, for example, the vehicle 10. Hereinafter, an example in which the voice system 5 is mounted on the vehicle 10 will be described.
[0012] A plurality of seats are provided in the passenger compartment of the vehicle 10. The plurality of seats are, for example, four seats including the driver's seat, the front passenger seat, and the left and right rear seats. Note that the number of seats is not limited to this. Hereinafter, the person sitting in the driver's seat will be referred to as the occupant hm1, the person sitting in the front passenger seat will be referred to as the occupant hm2, the person sitting on the right side of the rear seat will be referred to as the occupant hm3, and the person sitting on the left side of the rear seat will be referred to as the occupant hm4.
[0013] The voice system 5 includes a microphone MC1, a microphone MC2, a microphone MC3, a microphone MC4, a voice processing system 20, and an electronic device 30. The voice system 5 shown in FIG. 1 has the same number of microphones as the number of seats, that is, four microphones, but the number of microphones may not be the same as the number of seats.
[0014] The microphones MC1, MC2, MC3, and MC4 output voice signals to the voice processing system 20. Then, the voice processing system 20 outputs the voice recognition result to the electronic device 30. The electronic device 30 executes the process specified by the voice recognition result based on the input voice recognition result.
[0015] The microphone MC1 is a microphone that picks up the voice spoken by the occupant hm1. In other words, the microphone MC1 acquires a voice signal including the voice component spoken by the occupant hm1. The microphone MC1 is arranged, for example, on the right side of the overhead console. The microphone MC2 picks up the voice spoken by the occupant hm2. In other words, the microphone MC2 is a microphone that acquires a voice signal including the voice component spoken by the occupant hm2. The microphone MC2 is arranged, for example, on the left side of the overhead console. That is, the microphone MC1 and the microphone MC2 are arranged at adjacent positions.
[0016] The microphone MC3 is a microphone that picks up the voice of the occupant hm3. In other words, the microphone MC3 acquires an audio signal including the voice component of the occupant hm3 speaking. The microphone MC3 is arranged, for example, at the center right side of the ceiling near the rear seat. The microphone MC4 is a microphone that picks up the voice of the occupant hm4. In other words, the microphone MC4 acquires an audio signal including the voice component of the occupant hm4 speaking. The microphone MC4 is arranged, for example, at the center left side of the ceiling near the rear seat of the ceiling. That is, the microphone MC3 and the microphone MC4 are arranged at adjacent positions.
[0017] Also, the arrangement positions of the microphones MC1, MC2, MC3, and MC4 shown in FIG. 1 are examples, and they may be arranged at other positions.
[0018] Each microphone may be a directional microphone or an omnidirectional microphone. Each microphone may be a small MEMS (Micro Electro Mechanical Systems) microphone or an ECM (Electret Condenser Microphone). Each microphone may be a microphone capable of beamforming. For example, each microphone may be a microphone array having directivity in the direction of each seat and capable of picking up the voice of the pointing method.
[0019] The audio system 5 shown in FIG. 1 includes a plurality of audio processing systems 20 corresponding to each of the microphones. Specifically, the audio system 5 includes an audio processing system 21, an audio processing system 22, an audio processing system 23, and an audio processing system 24. The audio processing system 21 corresponds to the microphone MC1. The audio processing system 22 corresponds to the microphone MC2. The audio processing system 23 corresponds to the microphone MC3. The audio processing system 24 corresponds to the microphone MC4. Hereinafter, the audio processing system 21, the audio processing system 22, the audio processing system 23, and the audio processing system 24 may be collectively referred to as the audio processing system 20.
[0020] A signal output from the audio processing system 20 is input to the electronic device 30. The electronic device 30 executes processing corresponding to the signal output from the audio processing system 20. Here, the signal output from the audio processing system 20 is, for example, an audio command which is a command input by voice. The audio command is a signal that is specified by voice and controls a target device. That is, the electronic device 30 executes processing corresponding to the audio command output from the audio processing system 20. For example, the electronic device 30 executes processing such as opening and closing a window, processing related to driving the vehicle 10, changing the temperature of an air conditioner, and changing the volume of an audio device based on the audio command. The electronic device 30 is an example of a target device.
[0021] In addition, in FIG. 1, the case where four people are riding in the vehicle 10 is shown, but the number of passengers is not limited to this. The number of passengers may be less than or equal to the maximum seating capacity of the vehicle 10. For example, when the maximum seating capacity of the vehicle 10 is six, the number of passengers may be six or may be five or less.
[0022] FIG. 2 is a diagram showing an example of the hardware configuration of the audio processing system 20 in the first embodiment. In the example shown in FIG. 2, the audio processing system 20 includes a DSP (Digital Signal Processor) 2001, a RAM (Random Access Memory) 2002, a ROM (Read Only Memory) 2003, and an I / O (Input / Output) interface 2004.
[0023] The DSP 2001 is a processor capable of executing a computer program. Note that the type of the processor included in the audio processing system 20 is not limited to the DSP 2001. For example, the audio processing system 20 may be a CPU (Central Processing Unit) or may be other hardware. Also, the audio processing system 20 may include a plurality of processors.
[0024] RAM2002 is a volatile memory used as a cache or buffer, etc. Note that the type of volatile memory included in the audio processing system 20 is not limited to RAM2002. The audio processing system 20 may include a register instead of RAM2002. Also, the audio processing system 20 may include a plurality of volatile memories.
[0025] ROM2003 is a non-volatile memory that stores various information including computer programs. The DSP2001 realizes the functions of the audio processing system 20 by reading and executing a specific computer program from the ROM2003. The functions of the audio processing system 20 will be described later. Note that the type of non-volatile memory included in the audio processing system 20 is not limited to ROM2003. For example, the audio processing system 20 may include a flash memory instead of ROM2003. Also, the audio processing system 20 may include a plurality of non-volatile memories.
[0026] The I / O interface 2004 is an interface device to which an external device is connected. Here, the external device is, for example, a device such as a microphone MC1, a microphone MC2, a microphone MC3, a microphone MC4, or an electronic device 30. Also, the audio processing system 20 may include a plurality of I / O interfaces 2004.
[0027] In this way, the audio processing system 20 includes a memory in which a computer program is stored and a processor capable of executing the computer program. That is, the audio processing system 20 can be regarded as a computer. Note that the number of computers required to realize the functions of the audio processing system 20 is not limited to 1. The functions of the audio processing system 20 may be realized by the cooperation of two or more computers.
[0028] FIG. 3 is a block diagram showing an example of the configuration of the voice processing system 20 in the first embodiment. A voice signal is input to the voice processing system 20 from the microphones MC1, MC2, MC3, and MC4. Then, the voice processing system 20 outputs the voice recognition result to the electronic device 30. The voice processing system 20 includes a voice input unit 210, a failure detection unit 220, and a voice processing device 230.
[0029] The microphone MC1 generates a voice signal by converting the collected voice into an electrical signal. Then, the microphone MC1 outputs the voice signal to the voice input unit 210. The voice signal is a signal including the voice of the occupant hm1 and noises such as the voices of persons other than the occupant hm1, music emitted from an audio device, and driving noise.
[0030] The microphone MC2 generates a voice signal by converting the collected voice into an electrical signal. Then, the microphone MC2 outputs the voice signal to the voice input unit 210. The voice signal is a signal including the voice of the occupant hm2 and noises such as the voices of persons other than the occupant hm2, music emitted from an audio device, and driving noise.
[0031] The microphone MC3 generates a voice signal by converting the collected voice into an electrical signal. Then, the microphone MC3 outputs the voice signal to the voice input unit 210. The voice signal is a signal including the voice of the occupant hm3 and noises such as the voices of persons other than the occupant hm3, music emitted from an audio device, and driving noise.
[0032] The microphone MC4 generates a voice signal by converting the collected voice into an electrical signal. Then, the microphone MC4 outputs the voice signal to the voice input unit 210. The voice signal is a signal including the voice of the occupant hm4 and noises such as the voices of persons other than the occupant hm4, music emitted from an audio device, and driving noise.
[0033] The voice input unit 210 receives voice signals from each of the microphones MC1, MC2, MC3, and MC4. That is, the voice input unit 210 receives a first voice which is the voice uttered by the first speaker. In other words, the voice input unit 210 receives the voice uttered by any one of the plurality of speakers as the first speaker. The voice input unit 210 is an example of an input unit. Then, the voice input unit 210 outputs the voice signal to the failure detection unit 220.
[0034] The failure detection unit 220 detects failures of each of the microphones MC1, MC2, MC3, and MC4. Also, the failure detection unit 220 determines whether it is possible to identify the position of the first speaker. The failure detection unit 220 is an example of a determination unit. Here, the voice processing system 20 identifies the position of the speaker who uttered the voice included in each voice signal by comparing the voice signals output from each of the microphones MC1, MC2, MC3, and MC4. When any one of the microphones MC1, MC2, MC3, or MC4 fails, the voice processing system 20 may not be able to identify the position of the speaker. Therefore, the failure detection unit 220 detects the presence or absence of failures of the plurality of microphones, and determines whether it is possible to identify the position of the first speaker based on the detection result.
[0035] Specifically describe the determination of whether the microphone is faulty. The microphone MC1 and the microphone MC2 are arranged at adjacent positions. Therefore, the sound pressure received by the microphone MC1 and the sound pressure received by the microphone MC2 are substantially the same. Thus, the levels of the audio signals output from the microphone MC1 and the microphone MC2 are substantially the same. However, when one of the microphones MC1 and MC2 fails, one of the microphones MC1 and MC2 cannot normally pick up sound. Therefore, a difference occurs in the levels of the audio signals output from the microphone MC1 and the microphone MC2. When the difference between the level of the audio signal output from the microphone MC1 and the level of the audio signal output from the microphone MC2 is equal to or greater than a threshold value, the failure detection unit 220 determines that a failure has occurred in one of the microphones MC1 and MC2. For example, the failure detection unit 220 determines that the microphone that outputs the audio signal with the lower level among the two audio signals is faulty.
[0036] For the same reason, when the difference between the level of the audio signal output from the microphone MC3 and the level of the audio signal output from the microphone MC4 is equal to or greater than a threshold value, the failure detection unit 220 determines that a failure has occurred in one of the microphones MC3 and MC4.
[0037] When the failure detection unit 220 detects a failure in at least one of the microphones MC1, MC2, MC3, and MC4, the failure detection unit 220 outputs a failure detection signal indicating that a failure has been detected. That is, the failure detection unit 220 outputs a failure detection signal indicating whether it is possible to identify the position of the speaker who spoke the audio received by the audio input unit 210. The failure detection signal is an example of the first signal. In addition, the failure detection unit 220 outputs the audio signals output from the microphones MC1, MC2, MC3, and MC4 to the audio processing device 230.
[0038] The audio processing device 230 includes a signal reception unit 231, a BF (Beam Forming) processing unit 232, an EC (Echo Canceller) processing unit 233, a CTC (Cross Talk Canceller) processing unit 234, and an audio recognition unit 235.
[0039] The signal receiving unit 231 receives a failure detection signal indicating whether it is possible to identify the position of the speaker who uttered the voice received by the voice input unit 210. The signal receiving unit 231 is an example of a receiving unit. The signal receiving unit 231 receives the failure detection signal from the failure detection unit 220. The signal receiving unit 231 transmits the failure detection signal to the BF processing unit 232, the EC processing unit 233, the CTC processing unit 234, and the voice recognition unit 235.
[0040] The BF processing unit 232 emphasizes the voice in the direction of the target seat by directivity control. Regarding the operation of the BF processing unit 232, the case where the voice in the direction of the driver's seat is emphasized among the voice signals output from the microphone MC1 will be described as an example. The microphone MC1 and the microphone MC2 are arranged at adjacent positions. Therefore, it is assumed that the voice signals output from the microphone MC1 include the voices of the driver's seat passenger hm1 and the passenger hm2 in the passenger seat. Similarly, it is assumed that the voice signals output from the microphone MC2 include the voices of the driver's seat passenger hm1 and the passenger hm2 in the passenger seat.
[0041] However, the distance from the microphone MC1 to the passenger seat is farther than that of the microphone MC2. Therefore, when the passenger hm2 in the passenger seat speaks, the voice of the passenger hm2 in the passenger seat included in the voice signal output from the microphone MC1 is delayed compared to the voice of the passenger hm2 in the passenger seat included in the voice signal output from the microphone MC2. Thus, the BF processing unit 232 emphasizes the voice in the direction of the target seat by, for example, applying time delay processing to the voice signal. Then, the BF processing unit 232 outputs the voice signal with the voice in the direction of the target seat emphasized to the EC processing unit 233. However, the method by which the BF processing unit 232 emphasizes the voice in the direction of the target seat is not limited to the above.
[0042] The EC processing unit 233 cancels the voice components other than the voice uttered by the speaker among the voice signals output from the BF processing unit 232. Here, the voice components other than the voice uttered by the speaker are, for example, music emitted by the audio device of the vehicle 10, running noise, etc. In other words, the EC processing unit 233 performs echo cancellation processing.
[0043] More specifically, the EC processing unit 233 cancels the voice components specified by the reference signal from the voice signal output from the BF processing unit 232. Thereby, the EC processing unit 233 cancels the voice components other than the voice uttered by the speaker. Here, the reference signal is a signal indicating the voice components other than the voice uttered by the speaker. For example, the reference signal is a signal indicating the voice components due to the music emitted by the audio device. Thereby, the EC processing unit 233 can cancel the voice components other than the voice uttered by the speaker by canceling the voice components specified by the reference signal.
[0044] The CTC processing unit 234 cancels the voice emitted from directions other than the target seat. In other words, the CTC processing unit 234 performs crosstalk cancellation processing. All the voice signals from the microphones are input to the CTC processing unit 234 after passing through the echo cancellation processing by the EC processing unit 233. The CTC processing unit 234 cancels the voice components picked up from directions other than the target seat by using, as the reference signal, the voice signal output from the microphones other than the microphone of the target seat among the input voice signals. That is, the CTC processing unit 234 cancels the voice components specified by the reference signal from the voice signal related to the microphone of the target seat. Then, the CTC processing unit 234 outputs the voice signal after the crosstalk cancellation processing to the voice recognition unit 235.
[0045] The voice recognition unit 235 outputs a voice command to the electronic device 30 based on the voice signal and the fault detection signal. More specifically, the voice recognition unit 235 executes voice recognition processing on the voice signal output from the CTC processing unit 234 to identify the voice command included in the voice signal. The voice command also includes a voice position command that is a command regarding the position of the speaker. The electronic device 30 executes processing corresponding to the voice position command. The electronic device 30 executes processing such as changing the temperature of the air conditioner, changing the volume of the speaker, or opening and closing the window based on the voice position command.
[0046] The voice position command is a command for which the processing to be executed is determined according to the position of the speaker. For example, when the passenger hm2 in the passenger seat says "Open the window", the voice recognition unit 235 determines that the voice from this utterance is a voice position command indicating the process of opening the window on the left side of the passenger seat. Also, when the passenger hm3 in the right seat of the rear seat says "Open the window", the voice recognition unit 235 determines that the voice from this utterance is a voice position command indicating the process of opening the window on the right side of the rear seat.
[0047] The voice position command also includes a driving command related to driving. The driving command is a command related to the driving of the vehicle 10. For example, if the control of a device related to the driving of the vehicle 10 is performed by the utterance of a passenger hm3 in the rear seat or the like where driving is not originally assumed, there is a possibility that the control will be different from the intention of the driver hm1 in the driver's seat, which may be dangerous. Therefore, the voice recognition unit 235 is configured to be able to distinguish the driving command from other voice position commands. For example, the driving command is a command for controlling the car navigation system, a command for controlling the vehicle speed by accelerator control, or a command for controlling the vehicle speed by brake control.
[0048] Also, based on the microphone position where the voice signal containing the utterance position command is input, the voice recognition unit 235 determines from which position the voice signal was uttered. The voice recognition unit 235 determines that the voice signal based on the microphone MC1 is a voice uttered from the direction of the driver's seat. The voice recognition unit 235 determines that the voice signal based on the microphone MC2 is a voice uttered from the direction of the passenger seat. The voice recognition unit 235 determines that the voice signal based on the microphone MC3 is a voice uttered from the direction of the right side of the rear seat. The voice recognition unit 235 determines that the voice signal based on the microphone MC4 is a voice uttered from the direction of the left side of the rear seat.
[0049] Also, the voice recognition unit 235 determines whether it is possible to identify the position of the speaker based on the fault detection signal output from the fault detection unit 220. Here, if any one of the microphones MC1, MC2, MC3, or MC4 is faulty, the BF processing unit 232 and the CTC processing unit 234 may not be able to execute the processing normally. For example, the microphone MC1 picks up the voice of the driver's seat occupant hm1 and the voice of the passenger seat occupant hm2. If the microphone MC2 is faulty in this case, the BF processing unit 232 and the CTC processing unit 234 may not be able to execute the processing normally. That is, the CTC processing unit 234 cannot cancel the voice component of the voice that should have been picked up by the microphone MC2 from the voice signal output from the microphone MC1. Therefore, the voice signal output from the microphone MC1, which contains both the voice of the driver's seat occupant hm1 and the voice of the passenger seat occupant hm2, is input to the voice recognition unit 35. In that case, the voice recognition unit 235 also treats the voice of the passenger seat occupant hm2 contained in the voice signal output from the microphone MC1 as the voice of the driver's seat occupant hm1. Therefore, the voice recognition unit 235 determines whether it is possible to identify the position of the speaker based on the fault detection signal.
[0050] Here, when the voice recognition unit 235 determines that the voice command included in the voice signal is a speaking position command, if it cannot identify the position of the speaker who issued the speaking position command, it cannot determine the speaking position command to be output. For example, when the voice recognition unit 235 determines that the voice signal includes a speaking position command such as "Open the window", if it cannot identify the position of the speaker, it cannot identify which speaking position command for opening the window should be output.
[0051] Therefore, the voice recognition unit 235 is a voice recognition unit that outputs a voice command, which is a signal identified by voice and controls the electronic device 30, to the electronic device 30. When the fault detection signal indicates that the position of the first speaker cannot be identified, the output of the speaking position command, which is related to the position of the speaker among the voice commands, is restricted. The voice recognition unit 235 is an example of a voice recognition unit. In other words, the voice recognition unit 235 is a voice recognition unit that outputs a voice command, which is a signal identified by voice and controls the electronic device 30, to the electronic device 30. When the fault detection unit 220 determines that the position of the first speaker cannot be identified, the output of the speaking position command, which is related to the position of the speaker among the voice commands, is restricted.
[0052] Next, a method for restricting the output of the speaking position command will be described.
[0053] For example, when the fault detection unit 220 determines that the position of the speaker cannot be identified, the voice recognition unit 235 does not output the speaking position command. As a result, the electronic device 30 does not execute the process according to the speaking position command. Therefore, the voice recognition unit 235 can suppress the execution of an unintended process by the electronic device 30.
[0054] Alternatively, when the voice recognition unit 235 determines that the speaker's position cannot be identified because the failure detection unit 220 has detected a failure of the microphone, the voice recognition unit 235 restricts the output of the utterance position command specified by the voice input from the microphone associated with the failed microphone. In other words, when any one of the microphones belonging to a group composed of a plurality of adjacent microphones fails, the voice recognition unit 235 does not output the utterance position command specified by the voice input from the other microphones belonging to the group to the electronic device 30. On the other hand, the voice recognition unit 235 does not restrict the output of the utterance position command specified by the voice input from the microphones belonging to other groups. That is, the voice recognition unit 235 outputs the utterance position command specified by the voice input from the microphones belonging to other groups.
[0055] For example, the microphone MC1 and the microphone MC2 form a group. The voice input unit 210 receives voice including the first voice output from a plurality of microphones including the first microphone and the second microphone associated with the first microphone. The first microphone is, for example, the microphone MC2. The second microphone is, for example, the microphone MC1. The first voice is, for example, the voice spoken by the passenger hm2 in the passenger seat. For example, the microphone MC1 picks up the voice spoken by the driver hm1 and the voice spoken by the passenger hm2 in the passenger seat. In this case, if the microphone MC2 fails, the BF processing unit 232 and the CTC processing unit 234 cannot execute the processing normally. Therefore, the voice signal based on the microphone MC1 is input to the voice recognition unit 235 while including the voice spoken by the driver hm1 and the voice spoken by the passenger hm2 in the passenger seat. Thus, the voice recognition unit 235 may misjudge the voice spoken by the passenger hm2 in the passenger seat as the voice spoken by the driver hm1. On the other hand, since the microphone MC3 and the microphone MC4 are away from the driver hm1 and the passenger hm2 in the passenger seat, the possibility of picking up the voice spoken by the driver hm1 and the voice spoken by the passenger hm2 in the passenger seat is low. Therefore, when the failure detection unit 220 detects a failure of the first microphone and determines that the position of the first speaker cannot be specified, the voice recognition unit 235 does not output the speech position command specified by the voice input from the second microphone among the speech position commands. The first speaker is, for example, the passenger hm2 in the passenger seat.
[0056] Alternatively, when the failure detection unit 220 determines that the position of the first speaker cannot be specified, the voice recognition unit 235 changes the priority of the output of the driving commands among the voice position commands related to driving. For example, when the voice recognition unit 235 receives a plurality of voice position commands, it assigns the voice position commands to any one of the priority levels divided into multiple levels. Then, the voice recognition unit 235 outputs the voice position commands with a priority higher than the threshold value to the electronic device 30. That is, the voice recognition unit 235 preferentially causes the electronic device 30 to execute the voice position commands. On the other hand, the voice recognition unit 235 does not output the voice position commands with a priority lower than the threshold value to the electronic device 30. In this way, when the failure detection unit 220 determines that the position of the speaker cannot be specified, the voice recognition unit 235 changes the priority of the output of the driving commands.
[0057] For example, when the failure detection unit 220 determines that the position of the first speaker cannot be specified, the voice recognition unit 235 increases the priority of the output of the driving commands. Thereby, when any one of the microphones fails, the voice recognition unit 235 prevents the operation related to driving by voice from becoming impossible.
[0058] Alternatively, when the failure detection unit 220 determines that the position of the speaker cannot be specified, the voice recognition unit 235 decreases the priority of the output of the driving commands. Thereby, when any one of the microphones fails, the voice recognition unit 235 prevents the operation related to driving from being performed by the voice of a person who is not originally related to driving, such as the passenger hm4 in the rear seat.
[0059] Next, the operation of the voice processing system 20 according to the first embodiment will be described. FIG. 4 is a flowchart showing an example of the operation of the voice processing system 20 in the first embodiment.
[0060] The voice input unit 210 receives an input of a voice signal from the microphones MC1, MC2, MC3, and MC4 (step S11).
[0061] Based on the voice signal output from the voice input unit 210, the fault detection unit 220 determines whether any one of the microphones MC1, MC2, MC3, or MC4 is faulty (step S12).
[0062] The fault detection unit 220 outputs a fault detection signal indicating whether any one of the microphones MC1, MC2, MC3, or MC4 is faulty to the signal receiving unit 231 of the voice processing device 230 (step S13).
[0063] The signal receiving unit 231 transmits a fault detection signal indicating whether any one of the microphones MC1, MC2, MC3, or MC4 is faulty to the BF processing unit 232, the EC processing unit 233, the CTC processing unit 234, and the voice recognition unit 235 (step S14).
[0064] Based on the fault detection signal output from the signal receiving unit 231, the voice recognition unit 235 determines whether it can identify the position of the speaker of the voice included in the voice signal input via the BF processing unit 232, the EC processing unit 233, and the CTC processing unit 234 (step S15).
[0065] When the position of the speaker can be identified (step S15; Yes), the voice recognition unit 235 outputs the voice command included in the voice signal to the electronic device 30 (step S16). Thereby, the voice recognition unit 235 causes the electronic device 30 to execute the process specified by the voice command.
[0066] When the position of the speaker cannot be identified (step S15; No), the voice recognition unit 235 determines whether the voice command included in the voice signal is a command other than the speaking position command (step S17). When it is a command other than the speaking position command (step S17; Yes), the voice recognition unit 235 proceeds to step S16.
[0067] When the command included in the voice signal is a speaking position command (step S17; No), the voice recognition unit 235 restricts the output of the speaking position command (step S18). That is, as shown in step S16, the voice recognition unit 235 outputs a voice command, which is a signal identified by voice and controls the target device, to the electronic device 30. However, when it is determined in step S15 that the position of the first speaker cannot be identified, the voice recognition unit 235 restricts the output of the speaking position command, which is related to the position of the speaker, among the voice commands. Thereby, the voice recognition unit 235a restricts the execution of the process specified by the voice command.
[0068] As described above, the voice processing system 20 ends the process.
[0069] As described above, according to the first embodiment, the voice input unit 210 receives the first voice spoken by the first speaker, who is any one of the plurality of speakers. The failure detection unit 220 determines whether it is possible to identify the position of the first speaker who spoke the first voice received by the voice input unit 210 by detecting failures of the microphones MC1, MC2, MC3, and MC4. Then, when it is determined that the position of the first speaker cannot be identified, the voice recognition unit 235 that outputs a voice command, which is a signal identified by voice and controls the target device, to the electronic device 30 restricts the output of the speaking position command, which is related to the position of the speaker, among the voice commands identified by voice. Therefore, since the execution of an unintended process is restricted, the voice processing system 20 can execute an appropriate process even when the position of the speaker cannot be identified.
[0070] (Second Embodiment) The voice processing system 20a in the second embodiment will be described. In the second embodiment, matters different from those in the first embodiment will be described, and matters the same as those in the first embodiment will be briefly described or the description will be omitted.
[0071] FIG. 5 is a block diagram showing an example of the configuration of the voice processing system 20a in the second embodiment. The voice processing apparatus 230a of the voice processing system 20a in the second embodiment is different from the voice processing system 20 in the first embodiment in that it includes a speaker recognition unit 236.
[0072] The speaker recognition unit 236 determines whether the first voice, which is the voice spoken by the first speaker who is one of the plurality of speakers, is the voice of a pre-registered registrant. The speaker recognition unit 236 is an example of a speaker determination unit. More specifically, the speaker recognition unit 236 compares the voice signal of the pre-registered registrant with the voice signal output from the CTC processing unit 234 to determine whether the voice included in the voice signal output from the CTC processing unit 234 is the voice of the pre-registered registrant. For example, the speaker recognition unit 236 determines whether the voice included in the voice signal is the voice of the owner of the vehicle 10. Then, the speaker recognition unit 236 outputs a recognition result signal indicating whether it has been determined that the speaker who spoke the voice included in the voice signal is the registrant to the voice recognition unit 235a.
[0073] The voice recognition unit 235a outputs a speech position command on the condition that the speaker recognition unit 236 determines that the first voice is the speech of the registrant. More specifically, when the failure detection unit 220 determines that the position of the speaker can be specified, the voice recognition unit 235a outputs a speech position command regardless of whether it is the speech of the pre-registered registrant. Also, when the failure detection unit 220 determines that the position of the speaker cannot be specified, the voice recognition unit 235a outputs a speech position command on the condition that the speaker recognition unit 236 recognizes that it is the speech of the registrant. For example, the voice recognition unit 235a executes the processing of the speech position command on the condition that it is the voice of the pre-registered owner of the vehicle 10. On the other hand, when the failure detection unit 220 determines that the position of the speaker cannot be specified, the voice recognition unit 235a restricts the output of the speech position command on the condition that the speaker recognition unit 236 recognizes that it is the speech of the registrant.
[0074] Next, the operation of the voice processing system 20a according to the second embodiment will be described. FIG. 6 is a flowchart showing an example of the operation of the voice processing system 20a in the second embodiment.
[0075] The voice input unit 210 receives an input of a voice signal from the microphones MC1, MC2, MC3, and MC4 (step S21).
[0076] The failure detection unit 220 determines whether any of the microphones MC1, MC2, MC3, or MC4 has failed based on the voice signal output from the voice input unit 210 (step S22).
[0077] The failure detection unit 220 outputs a failure detection signal indicating whether any of the microphones MC1, MC2, MC3, or MC4 has failed to the signal reception unit 231 of the voice processing apparatus 230a (step S23).
[0078] The signal reception unit 231 transmits a failure detection signal indicating whether the microphones MC1, MC2, MC3, or MC4 has failed to the BF processing unit 232, the EC processing unit 233, the CTC processing unit 234, and the voice recognition unit 235a (step S24).
[0079] Based on the failure detection signal output from the signal reception unit 231, the voice recognition unit 235a determines whether it is possible to identify the position of the speaker of the voice included in the voice signal input via the BF processing unit 232, the EC processing unit 233, and the CTC processing unit 234 (step S25).
[0080] When it is possible to identify the position of the speaker (step S25; Yes), the voice recognition unit 235a outputs the voice command included in the voice signal to the electronic device 30 (step S26). Thereby, the voice recognition unit 235a causes the electronic device 30 to execute the process specified by the voice command.
[0081] When the position of the speaker cannot be specified (step S25; No), the voice recognition unit 235a determines whether the voice included in the voice signal is due to the speech of the registered user based on the recognition result signal (step S27).
[0082] When the voice included in the voice signal is due to the speech of the registered user (step S27; Yes), the voice recognition unit 235a proceeds to step S26.
[0083] When the voice included in the voice signal is not due to the speech of the registered user (step S27; No), the voice recognition unit 235a determines whether the voice command included in the voice signal is a command other than the speaking position command (step S28). When the voice command included in the voice signal is a command other than the speaking position command (step S28; Yes), the voice recognition unit 235a proceeds to step S26.
[0084] When the voice command included in the voice signal is the speaking position command (step S28; No), the voice recognition unit 235a restricts the output of the speaking position command (step S29). Thereby, the voice recognition unit 235a restricts the execution of the process specified by the voice command.
[0085] As described above, the voice processing system 20a ends the process.
[0086] As described above, according to the second embodiment, the speaker recognition unit 236 determines whether the first voice spoken by the first speaker, who is one of the plurality of speakers, is the voice of a registered user registered in advance. Then, the voice recognition unit 235a outputs a voice position command to the electronic device 30 on the condition that the speaker recognition unit 236 determines that the first voice is the voice of a registered user. As a result, the electronic device 30 executes the process of the voice position command on the condition that the voice is spoken by a specific registered user such as the owner of the vehicle 10. On the other hand, the voice recognition unit 235a restricts the output of the voice position command in the case of a voice spoken by a person other than the registered user. Therefore, since the execution of an unintended process is restricted, the voice processing system 20a can execute an appropriate process even when the position of the speaker cannot be specified.
[0087] (Modification Example 1) A modification example 1 of the first embodiment or the second embodiment will be described.
[0088] The voice processing device 230 in the first embodiment and the voice processing device 230a in the second embodiment have a CTC processing unit 234. However, the voice processing device 230 and the voice processing device 230a may not have the CTC processing unit 234. Further, the voice processing device 230 shown in FIG. 3 and the voice processing device 230a shown in FIG. 5 include an EC processing unit 233 at the subsequent stage of the BF processing unit 232. However, the voice processing device 230 and the voice processing device 230a may include the BF processing unit 232 at the subsequent stage of the EC processing unit 233.
[0089] (Modification Example 2) A modification example 2 of the first embodiment or the second embodiment will be described.
[0090] When the microphone MC3 or the microphone MC4 installed near the rear seat shown in FIG. 1 malfunctions, the audio processing device 230 in the first embodiment and the audio processing device 230a in the second embodiment may perform partial multi-zone sound collection using the non-faulty microphones. Specifically, when the microphone MC3 malfunctions, the audio processing device 230 and the audio processing device 230a perform sound collection of the voices of the rear seats using the microphone MC4. Or, when the microphone MC4 malfunctions, the audio processing device 230 and the audio processing device 230a perform sound collection of the voices of the rear seats using the microphone MC3.
[0091] In the first embodiment, the second embodiment, and their modification examples 1 and 2, the functions of the audio processing system 20 and the audio processing system 20a have been described as being realized by the DSP2001 executing a specific computer program. The computer program for realizing the functions of the audio processing system 20 and the audio processing system 20a on a computer may be pre-stored and provided in the ROM2003. The computer program for realizing the functions of the audio processing system 20 and the audio processing system 20a on a computer may be configured to be recorded and provided on a computer-readable recording medium such as a CD (Compact Disc)-ROM (Read Only Memory), a flexible disk (FD: Flexible Disc), a CD-R (Recordable), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, an SD (Secure Digital) card, etc. in an installable or executable file format.
[0092] Furthermore, a computer program for implementing the functions of the voice processing system 20 and the voice processing system 20a on a computer may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the computer program for implementing the functions of the voice processing system 20 and the voice processing system 20a may be configured to be provided or distributed via a network such as the Internet.
[0093] Also, some or all of the functions of the voice processing system 20 and the voice processing system 20a may be realized by a logic circuit. Some or all of the functions of the voice processing system 20 and the voice processing system 20a may be realized by an analog circuit. Some or all of the functions of the voice processing system 20 and the voice processing system 20a may be realized by an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), etc.
[0094] Although some embodiments of the present disclosure have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and the equivalent scope thereof.
Description of Reference Numerals
[0095] 5 Voice system 10 Vehicle 20, 20a, 21, 22, 23, 24 Voice processing system 30 Electronic device 210 Voice input unit 220 Fault detection unit 230, 230a Voice processing device 231 Signal reception unit 232 BF (Beam Forming) processing unit 233 EC (Echo Canceller) processing unit 234 CTC (Cross Talk Canceller) processing unit 235, 235a Speech recognition unit 236 Speaker recognition unit hm1, hm2, hm3, hm4 Occupants MC1, MC2, MC3, MC4 Microphones 2001 DSP (Digital Signal Processor) 2002 RAM (Random Access Memory) 2003 ROM (Read Only Memory) 2004 I / O (Input / Output) interface
Claims
A voice processing system mounted on a vehicle having a plurality of seats, comprising: An input unit that receives a first voice that is the voice uttered by a first speaker; A determination unit that determines whether or not the seat on which the first speaker is seated can be specified; A voice recognition unit that outputs, to the target device, a voice command that is a signal for controlling the target device and is specified by voice, and when the determination unit determines that the seat on which the first speaker is seated cannot be specified, restricts the output of a voice command that is a command for which the process to be executed according to the seat on which the speaker is seated is determined among the voice commands; A voice processing system comprising the above. **Claim 2** When the determination unit determines that the seat on which the first speaker is seated cannot be specified, the voice recognition unit does not output the voice command related to the seating position. The voice processing system according to claim 1. **Claim 3** The input unit receives a voice including the first voice output from a plurality of microphones including a first microphone and a second microphone associated with the first microphone; The determination unit detects the presence or absence of a failure in the plurality of microphones, and when a failure is detected in at least one of the first microphone and the second microphone, determines that the seat on which the first speaker is seated cannot be specified; When the determination unit detects a failure in the first microphone and determines that the seat on which the first speaker is seated cannot be specified, the voice recognition unit does not output a voice command related to the seating position specified by the voice input from the second microphone among the voice commands related to the seating position. The voice processing system according to claim 1. **Claim 4** When the determination unit determines that the seat on which the first speaker is seated cannot be specified, the voice recognition unit changes the priority of output of a driver's seat command related to the driver's seat among the voice commands related to the seating position. The voice processing system according to any one of claims 1 to 3. **Claim 5** When the determination unit determines that the seat on which the first speaker is seated cannot be specified, the voice recognition unit increases the priority of the output of the driver's seat command. The voice processing system according to claim 4. **Claim 6** Further comprising a speaker determination unit that determines whether the first voice is the voice of a pre-registered speaker; The voice recognition unit outputs the voice command related to the seating position on the condition that the speaker determination unit determines that the first voice is the voice of the registered speaker. The voice processing system according to any one of claims 1 to 5.
7. Further comprising an identifying unit that identifies the seat on which the first speaker who uttered the first voice is seated by comparing the voice signals of each of the plurality of microphones. The voice processing system according to claim 3.
8. When the difference between the level of the voice signal output from the first microphone and the level of the voice signal output from the second microphone is equal to or greater than a threshold value, the determination unit determines that a failure has occurred. The voice processing system according to claim 3.
9. The voice recognition unit outputs the utterance position command for which the process to be executed is determined according to the seat on which the speaker is seated. The voice processing system according to claim 6.
10. When the determination unit determines that the seat on which the speaker is seated cannot be identified because a failure of the microphone is detected, the voice recognition unit restricts the output of the utterance position command of the voice input from the microphone associated with the failed microphone. The voice processing system according to claim 6.
11. Further comprising an enhancement processing unit that performs directivity control to enhance the voice in the target direction. The voice processing system according to any one of claims 1 to 10.
12. Further comprising a first cancellation processing unit that cancels voice components other than the voice uttered by the speaker based on a reference signal indicating a specific voice component. The voice processing system according to any one of claims 1 to 11.
13. Further comprising a second cancellation processing unit that performs cancellation processing to cancel voices from directions other than the target direction. The voice processing system according to any one of claims 1 to 10.
14. A voice processing device mounted on a vehicle having a plurality of seats, A receiving unit that receives a first signal indicating whether or not it is possible to identify the seat on which the first speaker who uttered the first voice is seated; A voice recognition unit that outputs a voice command, which is a signal for controlling a target device identified by voice, to the target device, and when the first signal indicates that the seat on which the first speaker is seated cannot be identified, among the voice commands identified by the voice, the voice recognition unit that restricts the output of the utterance position command, which is a command for which the process to be executed is determined according to the seat on which the speaker is seated. A voice processing device comprising the above. A voice processing method executed in a voice processing system mounted on a vehicle having a plurality of seats, comprising: an input step of receiving, by an input unit, a first voice uttered by a first speaker; a determination step of determining, by a determination unit, whether the seat on which the first speaker is seated can be specified; an output step of outputting, by a voice recognition unit, a voice command, which is a signal for controlling a target device and is specified by voice, to the target device; wherein in the output step, when it is determined in the determination step that the seat on which the first speaker is seated cannot be specified, output of a speaking position command, which is a command among the voice commands for which processing to be executed according to the seat on which the speaker is seated is determined, is restricted; a voice processing method.
Citation Information
Patent Citations
Voice command inputting device and recording medium recoding program for operating the device
JP2001034454A
Exhibition panel and base plate therefor
JP2007090611A
Methods and Systems for Controlling an Electronic Device in Response to Detected Social Cues
US20170161016A1
Location based voice association system
US20180047394A1