Sound processing system, sound processing device, and sound processing method

By using multiple microphones and fault detection components in the sound processing system, the speaker's location is determined and related command output is restricted, thus solving the problem of improper processing when the location cannot be determined and achieving appropriate and safe processing.

CN115917642BActive Publication Date: 2026-05-05PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PANASONIC AUTOMOTIVE SYST CO LTD
Filing Date
2021-04-20
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing sound processing systems may perform unintended processing when the speaker's location cannot be determined, leading to improper processing.

Method used

By setting up multiple microphones and fault detection components, the speaker's location is determined, and command output related to the speaker's location is restricted to ensure that the system performs appropriate processing when the location cannot be determined.

Benefits of technology

Even when the speaker's location cannot be determined, the system can effectively limit unintended processing, ensuring the appropriateness and safety of the processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115917642B_ABST
    Figure CN115917642B_ABST
Patent Text Reader

Abstract

The sound processing system disclosed herein includes an input unit, a determination unit, and a sound recognition unit. The input unit receives a first sound, which is the sound emitted by a first speaker. The determination unit determines whether the location of the first speaker can be determined. The sound recognition unit outputs a sound command, determined based on the sound, as a signal for controlling the target device. If the determination unit determines that the location of the first speaker cannot be determined, the sound recognition unit restricts the output of a speaking position command, which is a command related to the speaker's location, from the sound command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a sound processing system, a sound processing apparatus, and a sound processing method. Background Technology

[0002] A sound processing system is known that processes sound recognition commands based on the sound emitted by the speaker.

[0003] Patent document 1 discloses a sound processing system that processes sound recognition commands based on the location of the sound emitted by the speaker.

[0004] Existing technical documents

[0005] Patent documents

[0006] Patent Document 1: Japanese Patent Application Publication No. 2017-90611 Summary of the Invention

[0007] However, Patent Document 1 does not disclose control measures in cases where the speaker's location cannot be determined. Without considering the possibility of the speaker's location being undetermined, the sound processing system may perform unintended processing.

[0008] The purpose of this disclosure is to perform appropriate processing in a sound processing system even when the speaker's location cannot be determined.

[0009] The sound processing system disclosed herein includes an input unit, a determination unit, and a sound recognition unit. The input unit receives a first sound, which is the sound emitted by a first speaker. The determination unit determines whether the location of the first speaker can be determined. The sound recognition unit outputs a sound command, determined based on the sound, as a signal for controlling the target device. If the determination unit determines that the location of the first speaker cannot be determined, the sound recognition unit restricts the output of a speaking position command, which is a command related to the speaker's location, from the sound command.

[0010] According to this disclosure, in a sound processing system, appropriate processing can be performed even when the speaker's location cannot be determined. Attached Figure Description

[0011] Figure 1 This is a diagram illustrating an example of the general structure of the in-vehicle sound processing system in the first embodiment.

[0012] Figure 2 This is a diagram illustrating an example of the hardware structure of the sound processing system in the first embodiment.

[0013] Figure 3This is a block diagram illustrating an example of the structure of the sound processing system in the first embodiment.

[0014] Figure 4 This is a flowchart illustrating an example of the operation of the sound processing system in the first embodiment.

[0015] Figure 5 This is a block diagram illustrating an example of the structure of the sound processing system in the second embodiment.

[0016] Figure 6 This is a flowchart illustrating an example of the operation of the sound processing system in the second embodiment. Detailed Implementation

[0017] The embodiments of this disclosure will now be described in detail with appropriate reference to the accompanying drawings. Sometimes, overly detailed descriptions are omitted. Furthermore, the drawings and the following description are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter of the claims.

[0018] (First Embodiment)

[0019] Figure 1 This is a diagram showing an example of the general structure of the sound system 5 in the first embodiment. The sound system 5 is, for example, mounted on a vehicle 10. An example of the sound system 5 being mounted on a vehicle 10 will be described below.

[0020] The vehicle 10 has multiple seats inside its passenger compartment. These multiple seats may include, for example, the driver's seat, the front passenger seat, and the left and right rear seats (four seats in total). However, the number of seats is not limited to these. Hereinafter, the person sitting in the driver's seat will be designated as passenger hm1, the person sitting in the front passenger seat as passenger hm2, the person sitting on the right side of the rear seats as passenger hm3, and the person sitting on the left side of the rear seats as passenger hm4.

[0021] The sound system 5 includes microphones MC1, MC2, MC3, and MC4, a sound processing system 20, and electronic devices 30. Figure 1 The sound system 5 shown has the same number of microphones as the number of seats, i.e., 4, but the number of microphones may not be the same as the number of seats.

[0022] Microphones MC1, MC2, MC3, and MC4 output sound signals to the sound processing system 20. Then, the sound processing system 20 outputs a sound recognition result to the electronic device 30. The electronic device 30 performs processing specified according to the input sound recognition result.

[0023] Microphone MC1 is a microphone that collects the sound emitted by occupant hm1. In other words, microphone MC1 acquires an audio signal containing the sound components emitted by occupant hm1. Microphone MC1 is positioned, for example, on the right side of the overhead console. Microphone MC2 collects the sound emitted by occupant hm2. In other words, microphone MC2 is a microphone that acquires an audio signal containing the sound components emitted by occupant hm2. Microphone MC2 is positioned, for example, on the left side of the overhead console. That is, microphones MC1 and MC2 are positioned close to each other.

[0024] Microphone MC3 is a microphone that collects the sound emitted by occupant hm3. In other words, microphone MC3 acquires a sound signal containing the sound components emitted by occupant hm3. Microphone MC3 is, for example, positioned on the right-center side of the ceiling near the rear seats. Microphone MC4 is a microphone that collects the sound emitted by occupant hm4. In other words, microphone MC4 acquires a sound signal containing the sound components emitted by occupant hm4. Microphone MC4 is, for example, positioned on the left-center side of the ceiling near the rear seats. That is, microphones MC3 and MC4 are positioned close to each other.

[0025] in addition, Figure 1 The configuration positions of microphones MC1, MC2, MC3, and MC4 shown are one example; they can also be configured in other positions.

[0026] Each microphone can be either directional or omnidirectional. Each microphone can be a small MEMS (Micro Electro Mechanical Systems) microphone or an ECM (Electret Condenser Microphone). Each microphone can also be a beamforming microphone. For example, each microphone can also be a microphone array that is directional along the direction of each seat and can collect sound in that direction.

[0027] Figure 1 The sound system 5 shown includes multiple sound processing systems 20, each corresponding to a microphone. Specifically, the sound system 5 includes sound processing system 21, sound processing system 22, sound processing system 23, and sound processing system 24. Sound processing system 21 corresponds to microphone MC1. Sound processing system 22 corresponds to microphone MC2. Sound processing system 23 corresponds to microphone MC3. Sound processing system 24 corresponds to microphone MC4. Hereinafter, sound processing systems 21, 22, 23, and 24 will sometimes be collectively referred to as sound processing system 20.

[0028] The electronic device 30 receives a signal output from the sound processing system 20. The electronic device 30 then performs processing corresponding to the signal output from the sound processing system 20. Here, the signal output from the sound processing system 20 refers to, for example, a command input via sound, i.e., a sound command. A sound command is a signal determined based on sound for controlling the target device. That is, the electronic device 30 performs processing corresponding to the sound command output from the sound processing system 20. For example, the electronic device 30 performs processing based on the sound command to open and close windows, processing related to driving the vehicle 10, processing to change the air conditioning temperature, and processing to change the volume of the audio equipment. The electronic device 30 is an example of a target device.

[0029] In addition, Figure 1 The diagram shows a vehicle 10 with four passengers, but the number of passengers is not limited to this. The number of passengers can be below the maximum capacity of vehicle 10. For example, if the maximum capacity of vehicle 10 is six people, the number of passengers can be six or five or less.

[0030] Figure 2 This is a diagram illustrating an example of the hardware structure of the sound processing system 20 in the first embodiment. Figure 2 In the example shown, the sound processing system 20 includes a DSP (Digital Signal Processor) 2001, a RAM (Random Access Memory) 2002, a ROM (Read Only Memory) 2003, and an I / O (Input / Output) interface 2004.

[0031] DSP 2001 is a processor capable of executing computer programs. However, the type of processor included in the sound processing system 20 is not limited to DSP 2001. For example, the sound processing system 20 may be a CPU (Central Processing Unit) or other hardware. Furthermore, the sound processing system 20 may also have multiple processors.

[0032] RAM 2002 is a volatile memory used as a cache memory or buffer, etc. However, the type of volatile memory included in the sound processing system 20 is not limited to RAM 2002. The sound processing system 20 can use registers instead of RAM 2002. Furthermore, the sound processing system 20 may also include multiple volatile memories.

[0033] ROM 2003 is a non-volatile memory that stores various information containing computer programs. DSP 2001 implements the functions of sound processing system 20 by reading a specific computer program from ROM 2003 and executing that program. The functions of sound processing system 20 are described later. Furthermore, the type of non-volatile memory provided by sound processing system 20 is not limited to ROM 2003. For example, sound processing system 20 can use flash memory instead of ROM 2003. Additionally, sound processing system 20 can also have multiple non-volatile memories.

[0034] I / O interface 2004 is an interface device for connecting external devices. These external devices include, for example, microphones MC1, MC2, MC3, MC4, and electronic device 30. Additionally, the sound processing system 20 may also have multiple I / O interfaces 2004.

[0035] In this way, the sound processing system 20 includes a memory storing a computer program and a processor capable of executing the computer program. That is, the sound processing system 20 can be considered a computer. Furthermore, the number of computers required to implement the functions of the sound processing system 20 is not limited to one. The functions of the sound processing system 20 can also be implemented through the cooperation of two or more computers.

[0036] Figure 3 This is a block diagram illustrating an example of the structure of the sound processing system 20 in the first embodiment. Sound signals are input to the sound processing system 20 from microphones MC1, MC2, MC3, and MC4. Then, the sound processing system 20 outputs a sound recognition result to the electronic device 30. The sound processing system 20 includes a sound input unit 210, a fault detection unit 220, and a sound processing device 230.

[0037] Microphone MC1 generates an audio signal by converting the collected sound into an electrical signal. Then, microphone MC1 outputs the audio signal to audio input unit 210. The audio signal is a signal that includes the voice of occupant hm1, as well as the voices of people other than occupant hm1, music emitted from audio equipment, or noise such as driving noise.

[0038] Microphone MC2 generates an audio signal by converting the collected sound into an electrical signal. Then, microphone MC2 outputs the audio signal to audio input unit 210. The audio signal includes the voice of occupant hm2, as well as the voices of people other than occupant hm2, music emitted from audio equipment, or noise such as driving noise.

[0039] Microphone MC3 generates an audio signal by converting the collected sound into an electrical signal. Then, microphone MC3 outputs the audio signal to audio input unit 210. The audio signal includes the voice of occupant hm3, as well as the voices of people other than occupant hm3, and noise such as music emitted from audio equipment or driving noise.

[0040] Microphone MC4 generates an audio signal by converting the collected sound into an electrical signal. Then, microphone MC4 outputs the audio signal to audio input unit 210. The audio signal includes the voice of occupant hm4, as well as the voices of people other than occupant hm4, and noise such as music or driving noise emitted from audio equipment.

[0041] The sound input unit 210 receives sound signals from each of the microphones MC1, MC2, MC3, and MC4. That is, the sound input unit 210 receives a first sound, which is the sound emitted by a first speaker. In other words, the sound input unit 210 receives the sound emitted by the first speaker, who is one of a plurality of speakers. The sound input unit 210 is an example of an input unit. Then, the sound input unit 210 outputs the sound signal to the fault detection unit 220.

[0042] The fault detection unit 220 detects faults in each of microphones MC1, MC2, MC3, and MC4. Additionally, the fault detection unit 220 determines whether the location of the first speaker can be determined. The fault detection unit 220 is an example of a determination unit. Here, the sound processing system 20 determines the location of the speaker who emitted the sound contained in each sound signal by comparing the sound signals output from microphones MC1, MC2, MC3, and MC4 respectively. If any of microphones MC1, MC2, MC3, and MC4 malfunctions, the sound processing system 20 may be unable to determine the speaker's location. Therefore, the fault detection unit 220 detects whether multiple microphones are faulty and determines whether the location of the first speaker can be determined based on the detection results.

[0043] Specifically, the determination of whether a microphone is faulty is explained. Microphones MC1 and MC2 are positioned close to each other. Therefore, the sound pressure level received by microphone MC1 is approximately the same as that received by microphone MC2. Consequently, the levels of the sound signals output from microphones MC1 and MC2 are approximately the same. However, if either microphone MC1 or MC2 malfunctions, either microphone MC1 or MC2 cannot collect sound properly. Therefore, the levels of the sound signals output from microphones MC1 and MC2 differ. If the difference between the levels of the sound signals output from microphone MC1 and microphone MC2 is greater than or equal to a threshold, the fault detection unit 220 determines that either microphone MC1 or microphone MC2 is faulty. For example, the fault detection unit 220 determines that the microphone that outputs the lower-level sound signal of the two sound signals is faulty.

[0044] For the same reason, if the difference between the level of the sound signal output from microphone MC3 and the level of the sound signal output from microphone MC4 is above a threshold, the fault detection unit 220 determines that one of the microphones MC3 and MC4 has malfunctioned.

[0045] If the fault detection unit 220 detects a fault in at least one of the microphones MC1, MC2, MC3, and MC4, it outputs a fault detection signal indicating that a fault has been detected. That is, the fault detection unit 220 outputs a fault detection signal indicating whether the location of the speaker, who emitted the sound received by the sound input unit 210, can be determined. The fault detection signal is an example of the first signal. Furthermore, the fault detection unit 220 outputs the sound signals from microphones MC1, MC2, MC3, and MC4 to the sound processing device 230.

[0046] The sound processing device 230 includes a signal receiving unit 231, a BF (Beam Forming) processing unit 232, an EC (Echo Canceller) processing unit 233, a CTC (Cross Talk Canceller) processing unit 234, and a sound recognition unit 235.

[0047] Signal receiving unit 231 receives a fault detection signal indicating whether the speaker's location can be determined, after the speaker has emitted a sound received by voice input unit 210. Signal receiving unit 231 is an example of a receiving unit. Signal receiving unit 231 receives a fault detection signal from fault detection unit 220. Signal receiving unit 231 sends the fault detection signal to BF processing unit 232, EC processing unit 233, CTC processing unit 234, and voice recognition unit 235.

[0048] The BF processing unit 232 enhances the sound in the direction of the target seat through directional control. The operation of the BF processing unit 232 will be explained using the case of enhancing the sound in the direction of the driver's seat in the sound signal output from microphone MC1 as an example. Microphones MC1 and MC2 are positioned close to each other. Therefore, it is assumed that the sound signal output from microphone MC1 includes the sounds of the driver's seat occupant hm1 and the passenger seat occupant hm2. Similarly, it is assumed that the sound signal output from microphone MC2 includes the sounds of the driver's seat occupant hm1 and the passenger seat occupant hm2.

[0049] However, the distance from microphone MC1 to the passenger seat is greater than the distance from microphone MC2 to the passenger seat. Therefore, when the passenger occupant hm2 makes a sound, the sound of passenger occupant hm2 contained in the sound signal output from microphone MC1 is delayed compared to the sound of passenger occupant hm2 contained in the sound signal output from microphone MC2. Therefore, the BF processing unit 232 enhances the sound in the direction of the target seat, for example, by applying time delay processing to the sound signal. Then, the BF processing unit 232 outputs the sound signal with enhanced sound in the direction of the target seat to the EC processing unit 233. However, the method by which the BF processing unit 232 enhances the sound in the direction of the target seat is not limited to the above method.

[0050] The EC processing unit 233 eliminates sound components other than the speaker's voice from the sound signal output from the BF processing unit 232. Here, sound components other than the speaker's voice refer to, for example, music emitted by the audio equipment of the vehicle 10, driving noise, etc. In other words, the EC processing unit 233 performs echo cancellation processing.

[0051] More specifically, the EC processing unit 233 eliminates the sound components determined based on the reference signal from the sound signal output from the BF processing unit 232. Thus, the EC processing unit 233 eliminates sound components other than the sound emitted by the speaker. Here, the reference signal refers to a signal representing sound components other than the sound emitted by the speaker. For example, the reference signal refers to a signal representing the sound components of music emitted by an audio device. Therefore, by eliminating the sound components determined based on the reference signal, the EC processing unit 233 is able to eliminate sound components other than the sound emitted by the speaker.

[0052] The CTC processing unit 234 eliminates sound emitted from directions other than the target seat. In other words, the CTC processing unit 234 performs crosstalk cancellation processing. Sound signals from all microphones are input to the CTC processing unit 234 after echo cancellation processing by the EC processing unit 233. The CTC processing unit 234 eliminates sound components collected from directions other than the target seat by using sound signals from microphones other than the target seat's microphone in the input sound signal as a reference signal. That is, the CTC processing unit 234 eliminates sound components determined based on the reference signal from the sound signal associated with the microphone at the target seat. Then, the CTC processing unit 234 outputs the crosstalk-cancelled sound signal to the sound recognition unit 235.

[0053] The voice recognition unit 235 outputs voice commands to the electronic device 30 based on the voice signal and the fault detection signal. More specifically, the voice recognition unit 235 determines the voice commands contained in the voice signal by performing voice recognition processing on the voice signal output from the CTC processing unit 234. Furthermore, the voice commands include a speaking position command, which is a command related to the speaker's location. The electronic device 30 performs processing corresponding to the speaking position command. Based on the speaking position command, the electronic device 30 performs processes such as changing the air conditioner temperature, changing the speaker volume, and opening and closing windows.

[0054] A speech position command is a command whose processing is determined based on the speaker's position. For example, if the passenger in the front passenger seat, hm2, says "open the window," the voice recognition unit 235 determines that the sound is a speech position command indicating that the window on the left side of the front passenger seat should be opened. Similarly, if the passenger in the right rear seat, hm3, says "open the window," the voice recognition unit 235 determines that the sound is a speech position command indicating that the window on the right side of the rear seat should be opened.

[0055] Furthermore, the speech location commands include driving commands associated with driving. Driving commands are commands related to driving the vehicle 10. For example, if the control of driving-related devices of the vehicle 10 is performed based on the voice of an occupant hm3 in the rear seat, etc., who was not originally intended to drive, it is possible to perform control that differs from the intention of the occupant hm1 in the driver's seat, which may sometimes lead to danger. Therefore, the voice recognition unit 235 is configured to be able to distinguish between driving commands and other speech location commands. For example, driving commands include commands to control the vehicle navigation system, commands to control vehicle speed via throttle control, and commands to control vehicle speed via brake control.

[0056] Furthermore, the voice recognition unit 235 determines the origin of the sound signal based on the microphone position from which the sound signal containing the speaking position command was input. The voice recognition unit 235 determines the sound signal based on microphone MC1 to be from the driver's seat direction. The voice recognition unit 235 determines the sound signal based on microphone MC2 to be from the passenger seat direction. The voice recognition unit 235 determines the sound signal based on microphone MC3 to be from the right side of the rear seat direction. The voice recognition unit 235 determines the sound signal based on microphone MC4 to be from the left side of the rear seat direction.

[0057] Furthermore, the voice recognition unit 235 determines whether the speaker's location can be determined based on the fault detection signal output from the fault detection unit 220. Here, if any one of the microphones MC1, MC2, MC3, and MC4 malfunctions, the BF processing unit 232 and the CTC processing unit 234 may sometimes fail to perform processing correctly. For example, microphone MC1 collects the voice emitted by the driver's seat occupant hm1 and the passenger seat occupant hm2. In this case, when microphone MC2 malfunctions, the BF processing unit 232 and the CTC processing unit 234 cannot perform processing correctly. That is, the CTC processing unit 234 cannot remove the sound components that should be collected by microphone MC2 from the sound signal output from microphone MC1. Therefore, the sound signal output from microphone MC1 is input to the voice recognition unit 235 in a state that includes both the voice emitted by the driver's seat occupant hm1 and the voice emitted by the passenger seat occupant hm2. In this case, the voice recognition unit 235 processes the voice emitted by the passenger occupant hm2, which is included in the sound signal output from the microphone MC1, as if it were emitted by the driver occupant hm1. Therefore, the voice recognition unit 235 determines whether the speaker's location can be determined based on the fault detection signal.

[0058] Here, if the voice recognition unit 235 determines that the voice command contained in the voice signal is a speaking position command, it cannot determine the speaking position command to be output if it cannot determine the location of the speaker who issued the speaking position command. For example, if the voice recognition unit 235 determines that the voice signal contains the speaking position command "open the window" but cannot determine the location of the speaker, it cannot determine which window to output the speaking position command to open.

[0059] Therefore, the voice recognition unit 235 outputs a voice command, determined based on sound, to the electronic device 30 as a signal for controlling the electronic device 30. If the fault detection signal indicates that the location of the first speaker cannot be determined, the voice recognition unit 235 restricts the output of the speaking position command, which is a command related to the speaker's location, within the voice command. The voice recognition unit 235 is an example of a voice recognition unit. In other words, the voice recognition unit 235 outputs a voice command, determined based on sound, to the electronic device 30 as a signal for controlling the electronic device 30. If the fault detection unit 220 determines that the location of the first speaker cannot be determined, the voice recognition unit 235 restricts the output of the speaking position command, which is a command related to the speaker's location, within the voice command.

[0060] Next, the method for limiting the output of the speaking position command will be explained.

[0061] For example, if the fault detection unit 220 determines that the speaker's location cannot be determined, the voice recognition unit 235 does not output a speaking location command. Therefore, the electronic device 30 does not perform processing corresponding to the speaking location command. Thus, the voice recognition unit 235 can suppress situations where the electronic device 30 performs unintended processing.

[0062] Alternatively, if the fault detection unit 220 determines that the speaker's location cannot be determined due to a microphone malfunction, the voice recognition unit 235 restricts the output of a speaking location command determined based on sound input from the microphone associated with the malfunctioning microphone. In other words, if any microphone in a group of multiple nearby microphones malfunctions, the voice recognition unit 235 does not output a speaking location command determined based on sound input from other microphones in the same group to the electronic device 30. On the other hand, the voice recognition unit 235 does not restrict the output of a speaking location command determined based on sound input from microphones in other groups. That is, the voice recognition unit 235 outputs a speaking location command determined based on sound input from microphones in other groups.

[0063] For example, microphones MC1 and MC2 form a group. The sound input unit 210 receives sound output from multiple microphones containing a first sound, including a first microphone and a second microphone associated with the first microphone. The first microphone is, for example, microphone MC2. The second microphone is, for example, microphone MC1. The first sound is, for example, the sound emitted by the passenger occupant hm2. For example, microphone MC1 collects the sound emitted by both the driver occupant hm1 and the passenger occupant hm2. In this case, when microphone MC2 malfunctions, the BF processing unit 232 and the CTC processing unit 234 cannot perform processing normally. Therefore, the sound signal based on microphone MC1 is input to the sound recognition unit 235 in a state that includes the sound emitted by both the driver occupant hm1 and the passenger occupant hm2. Therefore, the sound recognition unit 235 may misidentify the sound emitted by the passenger occupant hm2 as the sound emitted by the driver occupant hm1. On the other hand, microphones MC3 and MC4 are located far from the driver's seat occupant hm1 and the front passenger seat occupant hm2, making it unlikely that they will collect the sounds emitted by the driver's seat occupant hm1 and the front passenger seat occupant hm2. Therefore, if the fault detection unit 220 detects a fault in the first microphone and determines that the location of the first speaker cannot be determined, the voice recognition unit 235 will not output the speaking location command determined based on the sound input from the second microphone. The first speaker is, for example, the front passenger seat occupant hm2.

[0064] Alternatively, if the fault detection unit 220 determines that the location of the first speaker cannot be determined, the voice recognition unit 235 changes the priority of the output of driving commands related to driving within the speaking location commands. For example, if the voice recognition unit 235 receives multiple speaking location commands, it assigns the speaking location commands to one of the priority levels divided into multiple levels. Then, the voice recognition unit 235 outputs speaking location commands with a priority higher than a threshold to the electronic device 30. That is, the voice recognition unit 235 causes the electronic device 30 to execute speaking location commands with priority. On the other hand, the voice recognition unit 235 does not output speaking location commands with a priority lower than a threshold to the electronic device 30. In this way, if the fault detection unit 220 determines that the location of the speaker cannot be determined, the voice recognition unit 235 changes the priority of the output of driving commands.

[0065] For example, if the fault detection unit 220 determines that the location of the first speaker cannot be determined, the voice recognition unit 235 increases the priority of the output of driving commands. Thus, the voice recognition unit 235 prevents situations where driving-related operations cannot be performed via sound if a microphone malfunctions.

[0066] Alternatively, if the fault detection unit 220 determines that the speaker's location cannot be determined, the voice recognition unit 235 lowers the priority of the driving command output. Thus, the voice recognition unit 235 prevents rear-seat occupants, such as hm4, who are not normally involved in driving, from performing driving-related operations via voice in the event of a microphone malfunction.

[0067] Next, the operation of the sound processing system 20 according to the first embodiment will be described. Figure 4 This is a flowchart illustrating an example of the operation of the sound processing system 20 in the first embodiment.

[0068] The sound input unit 210 receives sound signal input from microphones MC1, MC2, MC3 and MC4 (step S11).

[0069] The fault detection unit 220 determines whether any one of the microphones MC1, MC2, MC3 and MC4 has malfunctioned based on the sound signal output from the sound input unit 210 (step S12).

[0070] The fault detection unit 220 outputs a fault detection signal indicating whether any one of the microphones MC1, MC2, MC3 and MC4 has malfunctioned to the signal receiving unit 231 of the sound processing device 230 (step S13).

[0071] The signal receiving unit 231 sends a fault detection signal indicating whether any one of the microphones MC1, MC2, MC3 and MC4 has malfunctioned to the BF processing unit 232, EC processing unit 233, CTC processing unit 234 and voice recognition unit 235 (step S14).

[0072] The voice recognition unit 235 determines, based on the fault detection signal output from the signal receiving unit 231, whether it is possible to determine the location of the speaker of the voice contained in the voice signal input via the BF processing unit 232, the EC processing unit 233 and the CTC processing unit 234 (step S15).

[0073] If the speaker's location can be determined (step S15; "Yes"), the voice recognition unit 235 outputs the voice command contained in the voice signal to the electronic device 30 (step S16). Thus, the voice recognition unit 235 causes the electronic device 30 to perform processing determined according to the voice command.

[0074] If the speaker's location cannot be determined (step S15; "No"), the voice recognition unit 235 determines whether the voice command contained in the voice signal is a command other than a speech location command (step S17). If it is a command other than a speech location command (step S17; "Yes"), the voice recognition unit 235 proceeds to step S16.

[0075] If the command contained in the sound signal is a speaking position command (step S17; "No"), the sound recognition unit 235 restricts the output of the speaking position command (step S18). That is, as shown in step S16, the sound recognition unit 235 outputs a sound command, determined based on the sound, as a signal for controlling the target device to the electronic device 30. However, if the sound recognition unit 235 determines in step S15 that the position of the first speaker cannot be determined, it restricts the output of the speaking position command, which is a command related to the speaker's position, in the sound command. Thus, the sound recognition unit 235a restricts the execution of the processing determined based on the sound command.

[0076] Based on the above, the sound processing system 20 ends the processing.

[0077] As described above, according to the first embodiment, the sound input unit 210 receives a first sound emitted by a first speaker, who is one of a plurality of speakers. The fault detection unit 220 determines whether the location of the first speaker who emitted the first sound received by the sound input unit 210 can be determined by detecting faults in microphones MC1, MC2, MC3, and MC4. Then, if it is determined that the location of the first speaker cannot be determined, the sound recognition unit 235, which outputs a sound command determined based on the sound as a signal for controlling the target device to the electronic device 30, restricts the output of the speaking position command, which is a command related to the speaker's location, in the sound command determined based on the sound. Therefore, since the execution of unintended processing is restricted, the sound processing system 20 can perform appropriate processing even when the speaker's location cannot be determined.

[0078] (Second Implementation)

[0079] The sound processing system 20a in the second embodiment will be described. Furthermore, in the second embodiment, matters that differ from those in the first embodiment will be described, while matters that are the same as those in the first embodiment will be briefly described or omitted.

[0080] Figure 5This is a block diagram showing an example of the structure of the sound processing system 20a in the second embodiment. The sound processing device 230a of the sound processing system 20a in the second embodiment differs from the sound processing system 20 in the first embodiment in that it includes a speaker recognition unit 236.

[0081] Speaker identification unit 236 determines whether the voice emitted by the first speaker (one of a plurality of speakers) is the voice of a pre-registered registrant. Speaker identification unit 236 is an example of a speaker determination unit. More specifically, speaker identification unit 236 determines whether the voice contained in the voice signal output from CTC processing unit 234 is the voice of a pre-registered registrant by comparing the voice signal of a pre-registered registrant with the voice signal output from CTC processing unit 234. For example, speaker identification unit 236 determines whether the voice contained in the voice signal is the voice of the owner of vehicle 10. Then, speaker identification unit 236 outputs an identification result signal indicating whether the speaker who emitted the voice contained in the voice signal can be determined to be a registrant to voice recognition unit 235a.

[0082] The voice recognition unit 235a outputs a speaking position command based on the speaker recognition unit 236 determining that the first voice is from a registered user. More specifically, if the fault detection unit 220 determines that the speaker's location can be determined, the voice recognition unit 235a outputs a speaking position command regardless of whether the voice is from a pre-registered user. Conversely, if the fault detection unit 220 determines that the speaker's location cannot be determined, the voice recognition unit 235a outputs a speaking position command based on the speaker recognition unit 236 recognizing that the voice is from a registered user. For example, the voice recognition unit 235a processes the speaking position command based on the condition that the voice is from the pre-registered owner of vehicle 10. On the other hand, if the fault detection unit 220 determines that the speaker's location cannot be determined, the voice recognition unit 235a restricts the output of the speaking position command based on the speaker recognition unit 236 recognizing that the voice is from a registered user.

[0083] Next, the operation of the sound processing system 20a according to the second embodiment will be described. Figure 6 This is a flowchart illustrating an example of the operation of the sound processing system 20a in the second embodiment.

[0084] The sound input unit 210 receives sound signal input from microphones MC1, MC2, MC3 and MC4 (step S21).

[0085] The fault detection unit 220 determines whether any one of the microphones MC1, MC2, MC3 and MC4 has malfunctioned based on the sound signal output from the sound input unit 210 (step S22).

[0086] The fault detection unit 220 outputs a fault detection signal indicating whether any one of the microphones MC1, MC2, MC3 and MC4 has malfunctioned to the signal receiving unit 231 of the sound processing device 230a (step S23).

[0087] The signal receiving unit 231 sends a fault detection signal indicating whether microphone MC1, microphone MC2, microphone MC3 or microphone MC4 has malfunctioned to the BF processing unit 232, EC processing unit 233, CTC processing unit 234 and voice recognition unit 235a (step S24).

[0088] The voice recognition unit 235a determines, based on the fault detection signal output from the signal receiving unit 231, whether it is possible to determine the location of the speaker of the voice contained in the voice signal input via the BF processing unit 232, the EC processing unit 233 and the CTC processing unit 234 (step S25).

[0089] If the speaker's location can be determined (step S25; "Yes"), the voice recognition unit 235a outputs the voice command contained in the voice signal to the electronic device 30 (step S26). Thus, the voice recognition unit 235a causes the electronic device 30 to perform processing determined according to the voice command.

[0090] If the speaker's location cannot be determined (step S25; "No"), the voice recognition unit 235a determines whether the sound contained in the sound signal is the sound made by the registrant based on the recognition result signal (step S27).

[0091] If the sound contained in the sound signal is a sound emitted by the registrant (step S27; "Yes"), the sound recognition unit 235a proceeds to step S26.

[0092] If the sound contained in the sound signal is not a sound emitted by the registrant (step S27; "No"), the sound recognition unit 235a determines whether the sound command contained in the sound signal is a command other than a speech position command (step S28). If the sound command contained in the sound signal is a command other than a speech position command (step S28; "Yes"), the sound recognition unit 235a proceeds to step S26.

[0093] If the voice command contained in the sound signal is a speaking position command (step S28; "No"), the voice recognition unit 235a restricts the output of the speaking position command (step S29). Thus, the voice recognition unit 235a restricts the execution of the processing determined based on the voice command.

[0094] Based on the above, the sound processing system 20a ends the processing.

[0095] As described above, according to the second embodiment, the speaker identification unit 236 determines whether the first voice emitted by the first speaker, who is one of a plurality of speakers, is the voice of a pre-registered registrant. Then, the voice recognition unit 235a outputs a speaking position command to the electronic device 30 on the condition that the speaker identification unit 236 determines that the first voice is the voice of a registrant. Thus, the electronic device 30 performs the speaking position command processing on the condition that the voice is emitted by a specific registrant, such as the owner of the vehicle 10. On the other hand, if the voice is emitted by someone other than a registrant, the voice recognition unit 235a restricts the output of the speaking position command. Therefore, since the execution of unintended processing is restricted, the voice processing system 20 can perform appropriate processing even when the speaker's position cannot be determined.

[0096] (Variation Example 1)

[0097] A variation of the first or second embodiment will be described.

[0098] The sound processing apparatus 230 in the first embodiment and the sound processing apparatus 230a in the second embodiment both have a CTC processing unit 234. However, the sound processing apparatus 230 and the sound processing apparatus 230a may also not have a CTC processing unit 234. Furthermore, Figure 3 The sound processing device 230 shown and Figure 5 The sound processing device 230a shown includes an EC processing unit 233 after the BF processing unit 232. However, both the sound processing device 230 and the sound processing device 230a may also include a BF processing unit 232 after the EC processing unit 233.

[0099] (Variation Example 2)

[0100] A variation of the first or second embodiment will be described.

[0101] exist Figure 1In the event of a malfunction in either microphone MC3 or microphone MC4 located near the rear seats, the sound processing device 230 in the first embodiment and the sound processing device 230a in the second embodiment can also perform localized multi-regional sound collection using the microphone that is not malfunctioning. Specifically, in the event of a malfunction in microphone MC3, the sound processing devices 230 and 230a collect sound from the rear seats using microphone MC4. Alternatively, in the event of a malfunction in microphone MC4, the sound processing devices 230 and 230a collect sound from the rear seats using microphone MC3.

[0102] In the first embodiment, the second embodiment, and their variations 1 and 2, the functions of the sound processing system 20 and the sound processing system 20a are described as being implemented by the DSP 2001 executing a specific computer program. The computer program for enabling the computer to implement the functions of the sound processing system 20 and the sound processing system 20a can be provided by pre-storing it in the ROM 2003. The computer program for enabling the computer to implement the functions of the sound processing system 20 and the sound processing system 20a can also be configured as an installable or executable file recorded on a computer-readable recording medium such as a CD (Compact Disc)-ROM (Read Only Memory), a floppy disk (FD), a CD-R (Recordable), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, or an SD (Secure Digital) card.

[0103] Furthermore, it can also be configured to provide the computer program for enabling the computer to perform the functions of the sound processing system 20 and the sound processing system 20a by storing it on a computer connected to a network such as the Internet and downloading it via the network. Alternatively, it can be configured to provide or distribute the computer program for enabling the computer to perform the functions of the sound processing system 20 and the sound processing system 20a via a network such as the Internet.

[0104] Furthermore, some or all of the functions of the sound processing system 20 and the sound processing system 20a can also be implemented using logic circuits. Some or all of the functions of the sound processing system 20 and the sound processing system 20a can also be implemented using analog circuits. Some or all of the functions of the sound processing system 20 and the sound processing system 20a can also be implemented using FPGA (Field-Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), etc.

[0105] Several embodiments of this disclosure have been described, but these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other ways and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the scope of the invention as described in the claims and its equivalents.

[0106] Explanation of reference numerals in the attached figures

[0107] 5: Sound system; 10: Vehicle; 20, 20a, 21, 22, 23, 24: Sound processing system; 30: Electronic equipment; 210: Sound input unit; 220: Fault detection unit; 230, 230a: Sound processing device; 231: Signal receiving unit; 232: BF (BeamForming) processing unit; 233: EC (Echo Canceller) processing unit; 234: CTC (Cross Talk Canceller) processing unit; 235, 235a: Voice recognition unit; 236: Speaker recognition unit; hm1, hm2, hm3, hm4: Occupants; MC1, MC2, MC3, MC4: Microphones; 2001: DSP (Digital Signal Processor); 2002: RAM (Random Access Memory); 2003: ROM (Read Only Memory); 2004: I / O (Input / Output) interface.

Claims

1. A sound processing system, comprising: The input unit receives the first voice, which is the voice emitted by the first speaker; The determination unit determines whether the location of the first speaker can be determined; and The voice recognition unit outputs a voice command, determined based on sound, as a signal for controlling the target device. If the determination unit determines that the location of the first speaker cannot be determined, the voice recognition unit restricts the output of a speech position command, which is a command related to the speaker's location, from the voice command. The input unit receives sound including the first sound output from a plurality of microphones, the plurality of microphones including a first microphone and a second microphone associated with the first microphone. The determination unit detects whether the plurality of microphones are faulty, and based on the detection results, determines whether the location of the first speaker can be determined. If the determination unit detects a malfunction in the first microphone and determines that the location of the first speaker cannot be determined, the voice recognition unit will not output the speaking location command determined based on the sound input from the second microphone.

2. The sound processing system according to claim 1, wherein, If the determination unit determines that the location of the first speaker cannot be determined, the voice recognition unit will not output the speaking location command.

3. The sound processing system according to claim 1 or 2, wherein, If the determination unit determines that the location of the first speaker cannot be determined, the voice recognition unit changes the priority of the output of the driver's seat command associated with the driver's seat in the speaking location command.

4. The sound processing system according to claim 3, wherein, If the determination unit determines that the location of the first speaker cannot be determined, the voice recognition unit increases the priority of the output of the driver's seat command.

5. The sound processing system according to claim 1 or 2, wherein, It also includes a speaker determination unit, which determines whether the first voice is the voice of a pre-registered registrant. The voice recognition unit outputs the speaking position command based on the condition that the speaker determination unit determines the first voice as the voice of the registrant.

6. The sound processing system according to claim 1, wherein, It also includes a determining unit that determines the location of the first speaker who emitted the first sound by comparing the sound signals of each of the plurality of microphones.

7. The sound processing system according to claim 1, wherein, If the difference between the level of the sound signal output from the first microphone and the level of the sound signal output from the second microphone is greater than or equal to a threshold, the determination unit determines that a malfunction has occurred.

8. The sound processing system according to claim 5, wherein, The voice recognition unit outputs a speaking position command based on the speaker's position to determine the processing to be performed.

9. The sound processing system according to claim 5, wherein, If the determination unit determines that the speaker's location cannot be determined due to a microphone malfunction, the voice recognition unit restricts the output of the speaking location command from the sound input from the microphone associated with the malfunctioning microphone.

10. A sound processing system, comprising: The input unit receives the first voice, which is the voice emitted by the first speaker; The determination unit determines whether the location of the first speaker can be determined. The voice recognition unit outputs a voice command, determined based on the voice, as a signal for controlling the target device to the target device. If the determination unit determines that the position of the first speaker cannot be determined, the voice recognition unit restricts the output of the voice command, which is a command related to the position of the speaker. as well as The first elimination processing unit eliminates sound components other than those emitted by the speaker based on a reference signal representing a specific sound component.

11. A sound processing system, comprising: The input unit receives the first voice, which is the voice emitted by the first speaker; The determination unit determines whether the location of the first speaker can be determined. The voice recognition unit outputs a voice command, determined based on the voice, as a signal for controlling the target device to the target device. If the determination unit determines that the position of the first speaker cannot be determined, the voice recognition unit restricts the output of the voice command, which is a command related to the position of the speaker. as well as The second elimination processing unit performs elimination processing to eliminate sound from directions other than the target direction.

12. A sound processing device, comprising: The receiving unit receives a first signal indicating whether the location of the first speaker who emitted the first sound can be determined. A voice recognition unit outputs a voice command, determined based on sound, as a signal for controlling the target device to a target device. If the first signal indicates that the location of the first speaker cannot be determined, the voice recognition unit restricts the output of a speech position command, which is a command related to the speaker's location, from the voice commands determined based on the sound. The enhancement processing unit performs directional control to enhance the sound in the target direction.

13. A sound processing method, comprising: Input steps to receive the first sound from the first speaker; The determination step involves determining whether the location of the first speaker can be determined. as well as The output step involves outputting a sound command, determined based on the sound, to the target device as a signal for controlling the target device. In the determination step, if it is determined that the location of the first speaker cannot be determined, the priority of the output of the driver's seat command associated with the driver's seat in the voice command is changed in the output step.

Citation Information

Patent Citations

  • Voice recognition control system

    JP2017090611A

  • Location based voice association system

    US20180047394A1

  • Information processing device, information processing method, program, and moving body

    WO2019069731A1