Microphone System

The microphone system addresses the issue of spatial noise by classifying sounds based on reflection time differences, improving SNR and voice recognition rates in environments with multiple sound sources.

JP7819579B2Active Publication Date: 2026-02-25DENSO CORP +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022093838
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2026-02-25
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

Existing sound collection and processing devices do not account for spatial noise, such as reflected sound, leading to a decrease in Signal-to-Noise Ratio (SNR) and voice recognition rate, particularly in environments like vehicle cabins where multiple sound sources and noise sources are present.

Method used

A microphone system that includes a sound collection unit, a clustering unit to classify sounds based on reflection time differences and other parameters, and an output unit to enhance voice recognition by suppressing noise from reflected sounds.

Benefits of technology

The system effectively suppresses noise from reflected sounds, maintaining or improving the SNR and voice recognition rate by accurately classifying voice and noise types, thereby enhancing voice recognition in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007819579000001
    Figure 0007819579000001
  • Figure 0007819579000002
    Figure 0007819579000002
  • Figure 0007819579000003
    Figure 0007819579000003
Patent Text Reader

Abstract

To provide a microphone system for suppressing lowering of an audio recognition ratio.SOLUTION: A microphone system causes at least one microphone 45 to collect sounds. The microphone system, based on a value ΔTr relating to a sound reflected in an acoustic space SB where the microphone 45 is arranged, classifies kinds of sounds included in sound data collected by the microphone 45 to a kind of a voice of a person in the acoustic space SB and a kind of a noise other than the voice. Further, the microphone system outputs data relating to the classified voice to an audio recognition device.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a microphone system. [Background technology]

[0002] As described in Patent Document 1, a sound collection processing device is known that separates an acoustic signal from an object from acoustic signals acquired by a microphone array based on the coordinates and features of the object. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-12314 Summary of the Invention [Problem to be solved by the invention]

[0004] According to the inventors' research, the sound collection and processing device described in Patent Document 1 separates acoustic signals from an object based on the coordinates and characteristics of the object, but does not take into account spatial noise, such as reflected sound, that occurs in the space where the object is located. This results in a decrease in the SNR for the sound from the object. As a result, when the object is, for example, a person, the SNR for the voice decreases, resulting in a decrease in the voice recognition rate for the person. Note that SNR stands for Signal-to-Noise Ratio, and is the ratio of signal to noise. Furthermore, the voice recognition rate is the degree of match between the actual spoken content and the voice-to-text converted text.

[0005] An object of the present disclosure is to provide a microphone system that suppresses a decrease in speech recognition rate. [Means for solving the problem]

[0006] The invention described in claim 1 is a microphone system comprising: a sound collection unit (S402) that collects sound in at least one microphone (45); a clustering unit (S404) that classifies the types of sound contained in sound data, which is data related to the sound collected by the microphone, into types of human voices in the acoustic space and types of noise, which are sounds other than voices, based on a value (ΔTr) related to sound reflected in an acoustic space (Sb), which is a space in which the microphones are arranged; and an output unit (S412) that outputs the classified data related to the voices to a voice recognition device (20). The value related to the sound reflected in the acoustic space is the arrival time difference of the sound reflected in the acoustic space for the same microphone. 。

[0007] This allows the occupant's voice to be classified from the sound collected by the microphone, taking into account noise caused by reflected sounds in the acoustic space. Therefore, the noise contained in the classified voice is suppressed, thereby suppressing a decrease in the SNR related to the voice. Therefore, a decrease in the voice recognition rate is suppressed.

[0008] The reference symbols in parentheses attached to each component indicate an example of the correspondence between the component and the specific components described in the embodiments described below. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a configuration diagram of a vehicle in which a microphone system according to a first embodiment is used. [Figure 2] 1 is a schematic diagram showing an acoustic space inside a vehicle cabin. [Figure 3] FIG. 2 is a diagram illustrating the configuration of a computing device of the microphone system. [Figure 4] 6 is a flowchart showing the processing of an occupant estimation unit of the computing device. [Figure 5] 10 is a flowchart showing the processing of a space estimation unit of the arithmetic device. [Figure 6] 5 is a flowchart showing the processing of a vehicle state estimation unit of the computing device. [Figure 7] 10 is a flowchart showing the processing of an SNR estimation unit of the arithmetic device. [Figure 8] 1 is a diagram showing the relationship between speech and each noise and sound pressure. [Figure 9] Graph showing clustering by frequency and sound pressure. [Figure 10] FIG. 10 is a diagram showing the relationship between SNR and speech recognition rate. [Figure 11] 1 is a diagram showing the relationship between the number of microphones and SNR. [Figure 12] FIG. 10 is a diagram showing the relationship between noise sound pressure and number, number of microphones, audio sound pressure, SNR, and responsiveness. [Figure 13] 10 is a flowchart showing the processing of the space estimation unit of the arithmetic device in the microphone system according to the second and third embodiments. [Figure 14] FIG. 10 is a configuration diagram of a vehicle in which a microphone system according to a fourth embodiment is used. [Figure 15] 1 is a schematic diagram showing an acoustic space inside a vehicle cabin. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described with reference to the drawings. In the following embodiments, identical or equivalent parts will be denoted by the same reference numerals, and description thereof will be omitted.

[0011] (First embodiment) The microphone system 30 of this embodiment is used in, for example, a vehicle 5. First, the vehicle 5 will be described.

[0012] 1, the vehicle 5 includes a vehicle system 10 and a microphone system 30. The vehicle system 10 includes an audio system 12, an air conditioner 14, a vehicle speed sensor 16, a road surface sensor 18, a voice recognition device 20, and the like.

[0013] The audio device 12 reads the recorded sound source and amplifies the read signal. The audio device 12 then emits a sound corresponding to the amplified signal into the vehicle cabin. The audio device 12 then outputs a signal corresponding to the sound pressure of the sound emitted into the vehicle cabin to a microphone system 30 (described later).

[0014] The air conditioner 14 is an air conditioning device and includes face air outlets, foot air outlets, defroster air outlets, and a blower (not shown). The air conditioner 14 blows air into the vehicle cabin through face air outlets, foot air outlets, and defroster air outlets (not shown) to adjust the temperature and humidity in the vehicle cabin. The air conditioner 14 also outputs a signal indicating the air outlet mode and a signal corresponding to the air volume to a microphone system 30 (described later). The face air outlet opens toward the backrest or headrest of the seat 6 in the vehicle cabin (shown in FIG. 2) and is opened and closed by a face air outlet door (not shown). The foot air outlet opens toward the seat of the seat 6 or the underside of the seat and is opened and closed by a foot air outlet door (not shown). The defroster air outlet opens toward the inner surface of the front windshield (not shown) of the vehicle 5 and is opened and closed by a defroster air outlet door (not shown). The air outlet mode refers to the open / closed states of the face air outlets, foot air outlets, and defroster air outlets.

[0015] The vehicle speed sensor 16 detects the vehicle speed and outputs a signal corresponding to the detected vehicle speed to a microphone system 30 (described later).

[0016] The road surface sensor 18 detects the condition of the road surface on which the vehicle 5 is traveling by using an exterior camera, Lidar, or the like. For example, the road surface sensor 18 detects the condition of the road surface on which the vehicle 5 is traveling by detecting unevenness of the road surface on which the vehicle 5 is traveling using an image captured by the exterior camera and pattern matching. Furthermore, for example, the road surface sensor 18 detects the condition of the road surface on which the vehicle 5 is traveling by using Lidar to detect the surface roughness of the road surface on which the vehicle 5 is traveling. The road surface sensor 18 then outputs a signal corresponding to the detected road surface condition to a microphone system 30, which will be described later. Note that Lidar is an abbreviation for Light detection and ranging. Examples of surface roughness include the root mean square height, maximum peak height, maximum valley height, maximum height, and calculated average height.

[0017] The voice recognition device 20 converts voice data output from a microphone system 30 (described later) into character data by using a voice recognition engine or the like. The voice recognition device 20 also outputs a signal corresponding to the converted character data to a display (not shown), for example. As a result, characters corresponding to the voice of the passenger in the vehicle are displayed on a display (not shown), and various systems (not shown) in the vehicle 5 are caused to operate in accordance with the character strings.

[0018] The microphone system 30 includes a microphone array 40, a sensor group 50, and a computing device 60. As shown in FIG. 2, the microphone array 40 has a plurality of arranged microphones 45 to collect sound.

[0019] Returning to FIG. 1 , the sensor group 50 includes an occupant sensor 52 and an environmental sensor 54. The occupant sensor 52 includes a weight sensor, an in-vehicle camera, an ultrasonic sensor, and the like. For example, the occupant sensor 52 detects that an occupant is sitting on the seat 6 in the vehicle cabin using a weight sensor attached to the seat 6. The occupant sensor 52 also uses an image captured by the in-vehicle camera and image recognition, as well as ultrasonic waves transmitted and received from the ultrasonic sensor. Using these, the occupant sensor 52 detects the position and number of occupants in the vehicle cabin. The occupant sensor 52 then outputs a signal corresponding to the detected position and number of occupants in the vehicle cabin to the microphone system 30 (described later). The ultrasonic waves are sound waves with a frequency of 20 kHz or higher. The position of the occupant in the vehicle cabin is, for example, the position of the occupant's mouth in an absolute coordinate system. The reference position of the absolute coordinate system is, for example, the center of gravity of the vehicle 5.

[0020] The environmental sensor 54 includes an in-vehicle camera, a window open / close sensor, and the like. For example, the environmental sensor 54 detects the size of the acoustic space Sb in the vehicle cabin and the position, type, and size of objects other than the occupants in the vehicle cabin by using an image captured by the in-vehicle camera and image recognition. Furthermore, the environmental sensor 54 detects the window opening degree by using a window open / close sensor. The environmental sensor 54 then outputs signals corresponding to the detected size of the space in the vehicle cabin, the position, type, and size of the objects in the vehicle cabin, and the window opening degree to the microphone system 30, which will be described later. The objects other than the occupants in the vehicle cabin are, for example, seats 6. The window opening degree is the opening degree of a side window of the vehicle 5.

[0021] The arithmetic device 60 is mainly composed of a microcomputer and includes a CPU, ROM, flash memory, RAM, I / O, a drive circuit, an A / D converter, and bus lines connecting these components. As shown in Fig. 3, the arithmetic device 60 also includes an occupant estimation unit 62, a space estimation unit 64, a vehicle state estimation unit 66, and an SNR estimation unit 68 as functional blocks.

[0022] The occupant estimation unit 62 executes a program stored in the ROM to estimate the positions and number of occupants in the vehicle cabin based on signals from the occupant sensor 52. Details of the processing by the occupant estimation unit 62 will be described later.

[0023] The space estimation unit 64 executes a program stored in the ROM to estimate the spatial state inside the vehicle cabin based on the signal from the environment sensor 54. Details of the space estimation unit 64 will be described later.

[0024] The vehicle state estimation unit 66 executes a program stored in the ROM to estimate the state of the vehicle 5 based on a signal from the vehicle system 10. The vehicle state estimation unit 66 will be described in detail later.

[0025] The SNR estimation unit 68 corresponds to the sound collection unit, the clustering unit, and the output unit. By executing a program stored in the ROM, the SNR estimation unit 68 generates voice data of the occupants in the vehicle cabin and calculates the SNR based on signals from the occupant estimation unit 62, the space estimation unit 64, and the vehicle state estimation unit 66. If the calculated SNR is insufficient, the SNR estimation unit 68 reselects the microphone 45 to collect the sound. If the calculated SNR is sufficient, the SNR estimation unit 68 outputs the generated voice data to the voice recognition device 20, which will be described later. The SNR estimation unit 68 will be described in detail later.

[0026] The vehicle 5 is configured as described above. The microphone system 30 used in the vehicle 5 recognizes voices in the vehicle cabin and suppresses a decrease in the voice recognition rate. Next, to explain the voice recognition in the vehicle cabin by the microphone system 30, each process performed when the programs of the occupant estimation unit 62, the space estimation unit 64, the vehicle state estimation unit 66, and the SNR estimation unit 68 are executed will be described. First, the process performed by the occupant estimation unit 62 will be described with reference to the flowchart of FIG. 4. The program of the occupant estimation unit 62 is executed, for example, when the ignition of the vehicle 5 is turned on. The period of a series of operations from when the occupant estimation unit 62 starts processing at step S100 until it returns to processing at step S100 is defined as the control cycle of the occupant estimation unit 62.

[0027] In step S100, the occupant estimation unit 62 acquires from the occupant sensor 52 a signal corresponding to the occupant positions and the number of occupants in the vehicle compartment.

[0028] Subsequently, in step S102, the occupant estimation unit 62 estimates the positions and number of occupants in the vehicle cabin from the signals acquired in step S100. The occupant estimation unit 62 also outputs signals corresponding to the estimated positions and number of occupants in the vehicle cabin to the SNR estimation unit 68. Thereafter, the processing of the occupant estimation unit 62 returns to step S100.

[0029] The occupant estimation unit 62 performs processing as described above. Next, the processing of the space estimation unit 64 will be described with reference to the flowchart in Fig. 5. The program of the space estimation unit 64 is executed, for example, when the ignition of the vehicle 5 is turned on. The period of a series of operations from when the processing of step S200 of the space estimation unit 64 starts until when the processing of step S200 is returned to is defined as the control cycle of the space estimation unit 64.

[0030] In step S200, the space estimation unit 64 acquires from the environment sensor 54 a signal corresponding to the size of the space in the vehicle cabin, the position, type and size of objects in the vehicle cabin, and the window opening degree.

[0031] Next, in step S202, the space estimation unit 64 estimates the size of the acoustic space Sb within the vehicle cabin, the position, type, and size of objects within the vehicle cabin, and the window opening degree from the signals acquired in step S200. In this way, the space estimation unit 64 estimates the state of the acoustic space Sb within the vehicle cabin. The space estimation unit 64 also outputs a signal corresponding to this estimated state of the acoustic space Sb within the vehicle cabin to the SNR estimation unit 68. Thereafter, the processing of the space estimation unit 64 returns to step S200.

[0032] The space estimation unit 64 performs processing as described above. Next, the processing of the vehicle state estimation unit 66 will be described with reference to the flowchart in Fig. 6. The program of the vehicle state estimation unit 66 is executed, for example, when the ignition of the vehicle 5 is turned on. The period of a series of operations from when the processing of step S300 of the vehicle state estimation unit 66 starts until when the processing of step S300 is returned to is defined as the control cycle of the vehicle state estimation unit 66.

[0033] In step S300, the vehicle state estimation unit 66 acquires signals corresponding to the state of the audio 12, the state of the air conditioner 14, the speed of the vehicle 5, and the state of the road surface on which the vehicle 5 is traveling from the vehicle system 10. Specifically, the vehicle state estimation unit 66 acquires a signal corresponding to the sound pressure of the sound from the audio 12 from the audio 12. The vehicle state estimation unit 66 also acquires a signal indicating the air outlet mode and a signal corresponding to the amount of air being blown from the air conditioner 14. The vehicle state estimation unit 66 also acquires a signal corresponding to the vehicle speed from the vehicle speed sensor 16. The vehicle state estimation unit 66 also acquires a signal corresponding to the state of the road surface on which the vehicle 5 is traveling from the road surface sensor 18.

[0034] Next, in step S302, the vehicle state estimation unit 66 estimates the state of the audio 12, the state of the air conditioner 14, the speed of the vehicle 5, and the state of the road surface on which the vehicle 5 is traveling, from the signals acquired in step S300. In this way, the vehicle state estimation unit 66 estimates the state of the vehicle 5. The vehicle state estimation unit 66 also outputs signals corresponding to the estimated state of the audio 12, the state of the air conditioner 14, the speed of the vehicle 5, and the state of the road surface on which the vehicle 5 is traveling, to the SNR estimation unit 68.

[0035] The vehicle state estimation unit 66 performs the processing as described above. Next, the processing of the SNR estimation unit 68 will be described with reference to the flowchart in Fig. 7. The program of the SNR estimation unit 68 is executed, for example, when the ignition of the vehicle 5 is turned on. The period of a series of operations from when the SNR estimation unit 68 starts the processing of step S400 until it returns to the processing of step S400 is defined as the control cycle of the SNR estimation unit 68.

[0036] In step S400, the SNR estimation unit 68 acquires various information. Specifically, the SNR estimation unit 68 acquires signals corresponding to the positions and number of occupants in the vehicle cabin from the occupant estimation unit 62. The SNR estimation unit 68 also acquires signals corresponding to the size of the acoustic space Sb in the vehicle cabin, the positions, types, and sizes of objects in the vehicle cabin, and the window opening degree from the space estimation unit 64. The SNR estimation unit 68 also acquires signals corresponding to the sound pressure of the audio system 12, the air outlet mode, the air volume of the air conditioner 14, the vehicle speed, and the state of the road surface on which the vehicle 5 is traveling from the vehicle state estimation unit 66.

[0037] Next, in step S402, the SNR estimation unit 68 collects sound from within the vehicle cabin using a pre-selected microphone 45 or a microphone 45 selected in step S410, which will be described later. The SNR estimation unit 68 also generates sound data corresponding to the sound collected by the microphone 45. The sound data is amplitude data for each time period over a time interval of a predetermined length.

[0038] Next, in step S404, the SNR estimation unit 68 classifies the sound data generated in step S402 into speech and noise based on the information acquired in step S400, and also classifies the speech type and the noise type, thereby clustering the sound data generated in step S402.

[0039] Specifically, first, the SNR estimation unit 68 estimates the following parameters for each time period of the sound data from the information acquired in step S400 in order to classify the type of speech and the type of noise.

[0040] The SNR estimation unit 68 estimates the number of voices from the number of occupants acquired in step S400. Furthermore, the SNR estimation unit 68 estimates the speech time difference ΔTs from the occupant positions acquired in step S400, the preset positions of each microphone 45, and the speed of sound. Note that the speech time difference ΔTs is the arrival time difference of the voices uttered by the occupants between the microphones 45, as shown in FIG. 2.

[0041] The SNR estimation unit 68 also estimates frequency components by performing frequency analysis of the sound data generated in step S402 using FFT or the like. Furthermore, the SNR estimation unit 68 estimates the occupant's speech pitch P from the sound data generated in step S402 using time-frequency analysis or the like. Note that FFT is an abbreviation for Fast Fourier Transform. The speech pitch P is the interval between sounds spoken by the occupant.

[0042] The SNR estimation unit 68 also estimates the degree of sound absorption and shielding from the type of object and the map acquired in step S400. Furthermore, the SNR estimation unit 68 estimates a reflection time difference ΔTr from the estimated degree of absorption and shielding, the size of the acoustic space Sb in the vehicle cabin acquired in step S400, the position and size of the object, and the window opening degree, the sound data generated in step S402, and the map. The map for estimating the degree of sound absorption and shielding is set in advance through experiments, simulations, etc. As shown in FIG. 2 , the reflection time difference ΔTr is the difference in arrival times of sounds reflected within the vehicle cabin to the same microphone 45. The map for estimating the reflection time difference ΔTr is set in advance through experiments, simulations, etc.

[0043] The SNR estimation unit 68 also estimates the sound pressure due to the occupant's speech from the sound data generated in step S402 and the map. The SNR estimation unit 68 also estimates the sound pressure due to the audio 12 from the sound pressure of the audio 12 acquired in step S400. The SNR estimation unit 68 also estimates the sound pressure due to the air conditioner 14 from the air volume of the air conditioner 14 acquired in step S400 and the map. The SNR estimation unit 68 also estimates the sound pressure due to wind noise of the vehicle 5 from the vehicle speed acquired in step S400. As shown in FIG. 8, the maps for estimating the sound pressure due to the occupant's speech, the air volume of the air conditioner 14, and the wind noise of the vehicle 5 are set in advance through experiments, simulations, etc.

[0044] The SNR estimation unit 68 also estimates the generation position of the sound from the audio device 12 from the setting state of the audio device 12 acquired in step S400. The SNR estimation unit 68 then estimates an audio sound time difference ΔTa from the estimated generation position of the sound from the audio device 12, the preset positions of each microphone 45, and the speed of sound. The SNR estimation unit 68 also estimates an air conditioner sound time difference ΔTw from the air volume and air outlet mode of the air conditioner 14 acquired in step S400 and the map. The SNR estimation unit 68 also estimates a driving sound time difference ΔTc from the vehicle speed, window opening degree, and the map acquired in step S400. The audio sound time difference ΔTa is the arrival time difference of the sound from the audio device 12 to the same microphone 45. The air conditioner sound time difference ΔTw is the arrival time difference of the sound from the air conditioner 14 to the same microphone 45. The map for estimating the air conditioner sound time difference ΔTw is set in advance through experiments, simulations, etc. The traveling sound time difference ΔTc is the difference in arrival time of the wind noise of the vehicle 5 to the same microphone 45. Furthermore, a map for estimating the traveling sound time difference ΔTc is set in advance by experiment, simulation, or the like.

[0045] Furthermore, the SNR estimation unit 68 estimates the sound pressure due to vibration of the vehicle 5 from the map and the state of the road surface on which the vehicle 5 is traveling, acquired in step S400. The map for estimating the sound pressure due to vibration of the vehicle 5 is set in advance through experiments, simulations, etc.

[0046] The SNR estimation unit 68 then classifies the types of sounds included in the sound data into speech types and noise types using the estimated number of speeches, the speech time difference ΔTs, the frequency components of the sound data, and the speech pitch P. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the estimated reflection time difference ΔTr. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the estimated sound pressure due to the occupant's speech. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the estimated sound pressure due to the audio 12, the sound pressure due to the air conditioner 14, the sound pressure due to wind noise of the vehicle 5, and the sound pressure due to vibrations of the vehicle 5. Furthermore, the SNR estimation unit 68 classifies the types of sounds contained in the sound data into voice types and noise types using the audio sound time difference ΔTa, air conditioner sound time difference ΔTw, and running sound time difference ΔTc estimated above.

[0047] Here, for example, let us assume that there are two occupants. One occupant is referred to as the first occupant. The other occupant is referred to as the second occupant. The voice of the first occupant is referred to as the first voice X1. The voice of the second occupant is referred to as the second voice X2. The sounds caused by the audio 12, the air conditioner 14, the wind noise of the vehicle 5, and the vibrations of the vehicle 5 are referred to as the first noise Xn1 and the second noise Xn2.

[0048] The number of occupants corresponds to the number of types of sounds. Furthermore, the frequencies of the voices of the occupants, the sound from the audio system 12, the sound from the air conditioner 14, the wind noise of the vehicle 5, and the vibrations of the vehicle 5 are all different. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first sound X1, a second sound X2, a first noise Xn1, and a second noise Xn2 using the sound data frequency-analyzed as described above and a frequency threshold. As a result, for example, as shown in FIG. 9, sounds with frequencies equal to or greater than the frequency threshold are classified as the first sound X1 and the first noise Xn1. Furthermore, sounds with frequencies less than the frequency threshold are classified as the second sound X2 and the second noise Xn2. The frequency threshold is set by experiment, simulation, machine learning, or the like so that the first sound X1, the second sound X2, the first noise Xn1, and the second noise Xn2 are classified.

[0049] Furthermore, the sound pressures of the voices of the passengers, the sound from the audio system 12, the sound from the air conditioner 14, the sound due to wind noise of the vehicle 5, and the sound due to vibrations of the vehicle 5 are all different. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first voice X1, a second voice X2, a first noise Xn1, and a second noise Xn2 using the amplitude of the sound data generated in step S402, the various sound pressures estimated above, and the sound pressure threshold. As a result, for example, as shown in FIG. 9, sounds with sound pressures equal to or greater than the sound pressure threshold are classified as the first voice X1 and the second noise Xn2. Furthermore, sounds with sound pressures less than the sound pressure threshold are classified as the second voice X2 and the first noise Xn1. The sound pressure threshold is set by experiment, simulation, machine learning, or the like so that the first voice X1, the second voice X2, the first noise Xn1, and the second noise Xn2 can be classified.

[0050] Therefore, a sound whose frequency is equal to or greater than the frequency threshold and whose sound pressure is equal to or greater than the sound pressure threshold is classified as a first sound X1. Furthermore, a sound whose frequency is less than the frequency threshold and whose sound pressure is less than the sound pressure threshold is classified as a second sound X2. Furthermore, a sound whose frequency is equal to or greater than the frequency threshold and whose sound pressure is less than the sound pressure threshold is classified as a first noise Xn1. Furthermore, a sound whose frequency is less than the frequency threshold and whose sound pressure is equal to or greater than the sound pressure threshold is classified as a second noise Xn2. In this way, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first sound X1, a second sound X2, a first noise Xn1, and a second noise Xn2. Note that in FIG. 9, the ranges of the first sound X1 and the second sound X2 are indicated by diagonal hatching. Furthermore, the ranges of the first noise Xn1 and the second noise Xn2 are indicated by hatching.

[0051] Furthermore, the speech pitch P differs depending on the occupant. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first voice X1, a second voice X2, a first noise Xn1, and a second noise Xn2 using the speech pitch P estimated above and a pitch threshold. Furthermore, since the occupant's position differs depending on the occupant, the speech time difference ΔTs differs. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first voice X1, a second voice X2, a first noise Xn1, and a second noise Xn2 using the speech time difference ΔTs estimated above and a speech threshold. Furthermore, since the reverberation of the occupant's voice and noise sounds differs depending on the state of the acoustic space Sb in the vehicle cabin, the reflection time difference ΔTr differs for each. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first sound X1, a second sound X2, a first noise Xn1, and a second noise Xn2 using the reflection time difference ΔTr estimated above and the reflection threshold. Furthermore, the audio sound time difference ΔTa varies depending on the position where the sound is generated by the audio 12. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first sound X1, a second sound X2, a first noise Xn1, and a second noise Xn2 using the audio sound time difference ΔTa estimated above and the audio time difference threshold. Furthermore, the air conditioner sound time difference ΔTw varies depending on the air outlet mode. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into a first sound X1, a second sound X2, a first noise Xn1, and a second noise Xn2 using the air conditioner sound time difference ΔTw estimated above and the air conditioner time difference threshold. Furthermore, the running sound time difference ΔTc varies depending on the window opening degree. Therefore, the SNR estimation unit 68 classifies the types of sounds included in the sound data into the first sound X1, the second sound X2, the first noise Xn1, and the second noise Xn2 using the running sound time difference ΔTc estimated above and the running sound time difference threshold. As a result, the SNR estimation unit 68 classifies the types of sounds included in the sound data into the first sound X1, the second sound X2, the first noise Xn1, and the second noise Xn2.The pitch threshold, speech threshold, reflection threshold, audio time difference threshold, air conditioner time difference threshold, and road noise time difference threshold are set by experiments, simulations, machine learning, etc. so that the first voice X1, the second voice X2, the first noise Xn1, and the second noise Xn2 are classified.

[0052] Then, the SNR estimation unit 68 generates the sound data for each type of sound classified in this way, for example, the sound data of the first sound X1 and the second sound X2, by extracting them from the sound data generated in step S402.

[0053] Next, in step S406, the SNR estimation unit 68 calculates the SNR for each piece of audio data generated in step S404. Specifically, the SNR estimation unit 68 divides the sum of amplitudes over time for the audio data by the sum of amplitudes over time for the audio data recorded in the vehicle cabin when no occupant is speaking. In this way, the SNR estimation unit 68 calculates the SNR for each piece of audio data.

[0054] For example, let S1 be the sum of amplitudes over time for the first sound X1. Let S2 be the sum of amplitudes over time for the second sound X2. Let N1 be the sum of amplitudes over time for the first sound X1 recorded in the vehicle cabin when the first occupant is not speaking. Let N2 be the sum of amplitudes over time for the second sound X2 recorded in the vehicle cabin when the second occupant is not speaking. In this case, the SNR of the first sound X1 is expressed as S1 / N1. Also, the SNR of the second sound X2 is expressed as S2 / N2.

[0055] Subsequently, in step S408, the SNR estimation unit 68 determines whether the SNR calculated in step S406 is equal to or greater than an SN threshold SNR_th. This allows the SNR estimation unit 68 to determine whether the SNR is sufficient. As shown in FIG. 10, the speech recognition rate improves as the SNR increases. Therefore, the SN threshold SNR_th is set by experiment, simulation, or the like so that the speech recognition rate is sufficient, for example, so that the speech recognition rate is 80% or higher.

[0056] Then, when the SNR of the voice data of the occupant designated by the occupant's button operation or the like, among the SNRs calculated in step S406, is less than the SNR threshold SNR_th, the SNR estimation unit 68 determines that the SNR is insufficient. Then, the processing of the SNR estimation unit 68 proceeds to step S410. Furthermore, when the SNR of the voice data of the occupant designated by the occupant's button operation or the like, among the SNRs calculated in step S406, is equal to or greater than the SNR threshold SNR_th, the SNR estimation unit 68 determines that the SNR is sufficient. Then, the processing of the SNR estimation unit 68 proceeds to step S412. Note that the SNR estimation unit 68 determines that the SNR is sufficient when the SNR of the voice data of the designated occupant is equal to or greater than the SNR threshold SNR_th, but this is not limiting. For example, the SNR estimation unit 68 may determine that the SNR is sufficient when the voice data of multiple occupants is equal to or greater than the SNR threshold SNR_th.

[0057] In step S410 following step S408, since the SNR is insufficient, the SNR estimation unit 68 changes the microphone 45 that collects sound in order to make the SNR sufficient. As a result, the SNR estimation unit 68 makes the SNR in the next control cycle larger than that in the current control cycle, thereby making the SNR sufficient.

[0058] 11, when the sound pressure and number of noises and the sound pressure of voices are fixed, the SNR increases as the number of microphones 45 increases. Therefore, in step S410, the SNR estimation unit 68, for example, increases the number of microphones 45 that collect sound compared to the current control cycle and sets the increased number of microphones 45. As a result, the SNR in the next control cycle is greater than the SNR in the current control cycle.

[0059] 12, when the number of microphones 45 and the sound pressure of the sound are fixed, the SNR decreases as the sound pressure of the noise or the number of noises increases. Therefore, in step S410, the SNR estimator 68 changes the number of microphones 45 to be increased to collect sound, for example, depending on the number of types of noise. Furthermore, in step S410, the SNR estimator 68 changes the number of microphones 45 to be increased to collect sound, for example, depending on the sound pressure of the noise. As a result, the SNR in the next control cycle is more likely to be equal to or greater than the SN threshold SNR_th.

[0060] Furthermore, when the sound pressure and number of noise and the number of microphones 45 are fixed, the SNR decreases as the sound pressure of the occupant's speech decreases. Therefore, in step S410, the SNR estimation unit 68 changes the number of microphones 45 to collect sound, for example, depending on the sound pressure of the occupant's speech. This makes it more likely that the SNR in the next control cycle will be equal to or greater than the SN threshold SNR_th.

[0061] In this way, the SNR estimating unit 68 increases the SNR in the next control cycle compared to the current control cycle by changing the number of microphones 45. After that, the processing of the SNR estimating unit 68 returns to step S400.

[0062] In step S412 following step S408, since the SNR is sufficient, the SNR estimation unit 68 outputs the voice data of the specified occupant from the voice data generated in step S404 to the voice recognition device 20. The voice recognition device 20 converts the voice data output from the SNR estimation unit 68 into character data using a voice recognition engine or the like. Furthermore, the voice recognition device 20 outputs the converted character data to, for example, a display (not shown). As a result, characters corresponding to the voices of the occupants in the vehicle cabin are displayed on the display (not shown). Thereafter, the processing of the SNR estimation unit 68 returns to step S400.

[0063] As described above, the SNR estimation unit 68 performs processing. Therefore, in the microphone system 30, the voice in the vehicle cabin is recognized through processing by the occupant estimation unit 62, the space estimation unit 64, the vehicle state estimation unit 66, and the SNR estimation unit 68. Next, how a decrease in the voice recognition rate by the microphone system 30 is suppressed will be described.

[0064] Here, the decrease in SNR related to voice will be explained. The sound collection processing device described in Patent Document 1 separates acoustic signals from an object based on the coordinates and features of the object, but does not take into account spatial noise such as reflected sound generated in the space where the object is located. This results in a decrease in SNR related to the sound from the object. As a result, when the object is, for example, a person, the decrease in SNR related to voice results in a decrease in the voice recognition rate related to the person.

[0065] Furthermore, in the sound collection device described in JP 2021-197658 A, the sound collection direction is controlled based on the direction of the sound source from the speaker and the direction of the listener's line of sight in the captured image shown by the image data. However, even in the sound collection device described in JP 2021-197658 A, spatial noise such as reflected sound generated in the space where the sound source is located is not taken into consideration. This results in a decrease in the SNR related to the sound from the sound source. As a result, when the sound source is, for example, a human voice, the SNR related to the sound decreases, resulting in a decrease in the speech recognition rate.

[0066] In contrast to these, in this embodiment, in step S404, the SNR estimation unit 68 classifies the types of sounds contained in the sound data into the types of voices of occupants in the vehicle cabin and the types of noise based on the data of the sound collected by the microphone 45 and the reflection time difference ΔTr. Note that the reflection time difference ΔTr is the difference in arrival times of sounds reflected in the vehicle cabin to the same microphone 45, and corresponds to a value related to the sound reflected in the acoustic space Sb. Also, the occupant corresponds to a person.

[0067] As a result, noise due to reflected sounds occurring in the acoustic space Sb is taken into consideration when classifying the occupant's voice from the sound collected by the microphone 45. This prevents an increase in noise contained in the classified voice, thereby preventing a decrease in the SNR related to the voice. Therefore, a decrease in the voice recognition rate is prevented.

[0068] Furthermore, the microphone system 30 of the first embodiment also provides the following effects.

[0069] [1-1] In step S404, the SNR estimation unit 68 classifies the types of sounds included in the sound data into the type of voice of an occupant in the vehicle cabin and the type of noise based on the audio sound time difference ΔTa, the air conditioner sound time difference ΔTw, and the road sound time difference ΔTc. The audio sound time difference ΔTa is the arrival time difference of sound from the audio 12 to the same microphone 45, and corresponds to the arrival time difference of sounds other than voice generated in the acoustic space Sb to the same microphone 45. The air conditioner sound time difference ΔTw is the arrival time difference of sound from the air conditioner 14 to the same microphone 45, and corresponds to the arrival time difference of sounds other than voice generated in the acoustic space Sb to the same microphone 45. The road sound time difference ΔTc is the arrival time difference of wind noise from the vehicle 5 to the same microphone 45, and corresponds to the arrival time difference of sounds other than voice generated in the acoustic space Sb to the same microphone 45.

[0070] This allows the occupant's voice to be classified from the sound collected by the microphone 45, taking into consideration noises generated in the acoustic space Sb by the audio 12, the air conditioner 14, and wind noise. This prevents an increase in noise contained in the classified voice, thereby preventing a decrease in the SNR related to the voice. This also prevents a decrease in the voice recognition rate.

[0071] [1-2] Here, the frequency and sound pressure differ depending on the voice and noise, and the speech pitch P and speech time difference ΔTs also differ depending on the occupant. Therefore, in step S404, the SNR estimation unit 68 classifies the types of sounds contained in the sound data into the types of voices of occupants in the vehicle cabin and the types of noise based on the frequency, sound pressure, speech pitch P, and speech time difference ΔTs. This makes it easier to classify the types of sounds contained in the sound data. The speech time difference ΔTs is the difference in arrival times of the voices of the occupants' speech between the microphones 45.

[0072] [1-3] In step S404, the SNR estimation unit 68 estimates the audio sound time difference ΔTa, the air conditioner sound time difference ΔTw, and the road sound time difference ΔTc based on the state of the audio system 12, the state of the air conditioner 14, the vehicle speed, and the state of the road surface on which the vehicle 5 is traveling. Furthermore, the SNR estimation unit 68 classifies the types of sounds contained in the sound data into the types of voices of passengers in the vehicle cabin and the types of noise based on the estimated audio sound time difference ΔTa, the air conditioner sound time difference ΔTw, and the road sound time difference ΔTc. As a result, when the microphone system 30 is used in the vehicle 5, the voices of passengers are classified from the sounds collected by the microphone 45, taking into account noise, which is sound other than voice generated in the acoustic space Sb. Therefore, an increase in noise contained in the classified voices is suppressed, and a decrease in the SNR related to the voices is suppressed. Therefore, a decrease in the voice recognition rate is suppressed.

[0073] [1-4] In step S408, the SNR estimation unit 68 determines whether the SNR of each piece of voice data calculated in step S406 is equal to or greater than the SN threshold SNR_th. If the SNR is less than the SN threshold SNR_th, the SNR is insufficient. Therefore, in step S410, the SNR estimation unit 68 increases the number of microphones 45 that collect sound from the current number, thereby increasing the SNR related to the voice. This increases the SNR related to the voice, thereby suppressing a decrease in the voice recognition rate. Note that the SNR estimation unit 68 corresponds to the change unit. Also, the current time corresponds to when the SNR is less than the SN threshold SNR_th.

[0074] [1-5] As described above, the SNR varies depending on the number of noise types, the sound pressure of the noise, and the sound pressure of the voice. Furthermore, as the number of microphones 45 collecting sound increases, the SNR increases, but the responsiveness of the output of sound data to the input of sound data, which increases the computational load, decreases. Therefore, in step S410, the SNR estimation unit 68 changes the number of microphones 45 to be added to collect sound depending on the number of noise types, the sound pressure of the noise, and the sound pressure of the voice. This adjusts the number of microphones 45 to be added, thereby increasing the SNR to a sufficient value and preventing excessive deterioration in responsiveness.

[0075] (Second embodiment) In the second embodiment, the processing of the spatial estimation unit 64 and the SNR estimation unit 68 differs from that of the first embodiment. The rest is the same as in the first embodiment. First, the processing of the spatial estimation unit 64 in the second embodiment will be described with reference to the flowchart in FIG.

[0076] In step S200, the space estimation unit 64 acquires from the environment sensor 54 a signal corresponding to the size of the acoustic space Sb within the vehicle cabin, the position, type and size of an object within the vehicle cabin, and the window opening degree.

[0077] Next, in step S202, the space estimation unit 64 estimates the size of the acoustic space Sb within the vehicle cabin, the position, type, and size of objects within the vehicle cabin, and the window opening degree from the signals acquired in step S200. In this way, the space estimation unit 64 estimates the state of the acoustic space Sb within the vehicle cabin.

[0078] Next, in step S204, the space estimation unit 64 determines whether the spatial state estimated in step S202 has changed. For example, the space estimation unit 64 determines that the spatial state has changed when the absolute value of the difference between the size of the acoustic space Sb in the vehicle cabin in the current control cycle and the size of the acoustic space Sb in the vehicle cabin in the previous control cycle is equal to or greater than a threshold. The space estimation unit 64 also determines that the spatial state has changed when the absolute value of the difference between each coordinate of the position of an object in the vehicle cabin in the current control cycle and each coordinate of the position of an object in the vehicle cabin in the previous control cycle is equal to or greater than a threshold. The space estimation unit 64 also determines that the spatial state has changed when the type of object in the vehicle cabin in the current control cycle is different from the type of object in the vehicle cabin in the previous control cycle. The space estimation unit 64 also determines that the spatial state has changed when the absolute value of the difference between the size of an object in the vehicle cabin in the current control cycle and the size of an object in the vehicle cabin in the previous control cycle is equal to or greater than a threshold. The above thresholds are set by experiment, simulation, machine learning, or the like so that it is determined that the spatial state has changed.

[0079] Furthermore, it is assumed that the absolute value of the difference between the size of the acoustic space Sb in the vehicle cabin in the current control cycle and the size of the acoustic space Sb in the vehicle cabin in the previous control cycle is less than a threshold value. It is also assumed that the absolute value of the difference between each coordinate of the position of an object in the vehicle cabin in the current control cycle and each coordinate of the position of an object in the vehicle cabin in the previous control cycle is less than a threshold value. It is also assumed that the type of object in the vehicle cabin in the current control cycle is the same as the type of object in the vehicle cabin in the previous control cycle. It is also assumed that the absolute value of the difference between the size of an object in the vehicle cabin in the current control cycle and the size of an object in the vehicle cabin in the previous control cycle is less than a threshold value. In this case, the space estimation unit 64 determines that the spatial state has not changed.

[0080] In step S206 following step S204, since the spatial state has not changed, the space estimation unit 64 outputs a signal corresponding to the state of the acoustic space Sb in the vehicle cabin estimated in step S202 to the SNR estimation unit 68. Thereafter, the processing of the space estimation unit 64 returns to step S200.

[0081] In step S208 following step S204, the space estimation unit 64 calculates a transfer function G for correcting the frequency threshold, sound pressure threshold, pitch threshold, speech threshold, reflection threshold, audio time difference threshold, air conditioner time difference threshold, and road noise time difference threshold, which will be described later.

[0082] Specifically, the space estimation unit 64 generates a reference sound such as an impulse sound, white noise, or M sequence from a speaker. The space estimation unit 64 then collects the generated reference sound in the microphone array 40. The space estimation unit 64 then calculates the transfer function G by dividing the amplitude of the sound collected by the microphone array 40 by the amplitude of the reference sound. The reference sound is an ultrasonic wave with a frequency of 20 kHz or higher, such as an impulse sound, white noise, or M sequence. The reference sound may also be an audible sound with a frequency of 20 to 20 kHz.

[0083] Next, in step S210, the space estimation unit 64 outputs a signal corresponding to the transfer function G calculated in step S208 in addition to the signal corresponding to the state of the acoustic space Sb in the vehicle cabin estimated in step S202 to the SNR estimation unit 68. Thereafter, the processing of the space estimation unit 64 returns to step S200.

[0084] As described above, the spatial estimation unit 64 performs the processing. Next, the processing of the SNR estimation unit 68 will be described with reference to the flowchart of FIG.

[0085] In step S400, the SNR estimation unit 68 acquires from the space estimation unit 64 a signal corresponding to the transfer function G in addition to the size of the acoustic space Sb within the vehicle cabin, the position, type, and size of objects within the vehicle cabin, and the window opening degree. The SNR estimation unit 68 also acquires from the occupant estimation unit 62 a signal corresponding to the position and number of occupants within the vehicle cabin. The SNR estimation unit 68 also acquires from the vehicle state estimation unit 66 signals corresponding to the state of the audio 12, the state of the air conditioner 14, the speed of the vehicle 5, and the state of the road surface on which the vehicle 5 is traveling.

[0086] Subsequently, in step S402, the SNR estimation unit 68 performs the same process as in the first embodiment, so a description of the process in step S402 will be omitted.

[0087] In step S404 following step S402, the SNR estimation unit 68 performs frequency analysis on the sound data generated in step S402, as in the first embodiment, and estimates the speech time difference ΔTs, speech pitch P, and reflection time difference ΔTr. The SNR estimation unit 68 also estimates the sound pressure due to the speech of the occupants, the sound pressure due to the audio 12, the sound pressure due to the air conditioner 14, the sound pressure due to wind noise of the vehicle 5, and the sound pressure due to vibrations of the vehicle 5. The SNR estimation unit 68 also estimates the audio sound time difference ΔTa, the air conditioner sound time difference ΔTw, and the running sound time difference ΔTc.

[0088] In addition, the SNR estimation unit 68 corrects the frequency threshold, sound pressure threshold, pitch threshold, speech threshold, reflection threshold, audio time difference threshold, air conditioner time difference threshold, and driving sound time difference threshold using the transfer function G acquired in step S400 and machine learning.

[0089] Then, the SNR estimation unit 68 classifies the types of sounds included in the sound data into speech types and noise types using the frequency-analyzed sound data and the corrected frequency threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the sound data generated in step S402 and the corrected sound pressure threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the speech pitch P estimated above and the corrected pitch threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the speech time difference ΔTs estimated above and the corrected speech threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the reflection time difference ΔTr estimated above and the corrected reflection threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the audio sound time difference ΔTa estimated above and the corrected audio time difference threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the air conditioner sound time difference ΔTw estimated above and the corrected air conditioner time difference threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the road sound time difference ΔTc estimated above and the road sound time difference threshold. As a result, the SNR estimation unit 68 increases the SNR for speech compared to before each threshold was corrected.

[0090] Subsequently, in steps S408 to S412, the SNR estimation unit 68 performs the same processing as in the first embodiment. Therefore, a description of the processing in steps S408 to S412 will be omitted.

[0091] As described above, the SNR estimation unit 68 performs the processing. Even when such processing is performed, the same effects as those of the first embodiment are achieved. Furthermore, the second embodiment also achieves the effects described below.

[0092] [2-1] In step S208, the space estimation unit 64 calculates a transfer function G, which is a value related to the ratio between the amplitude of a reference sound collected by the microphone 45 and the amplitude of the sound collected by the microphone 45. Furthermore, the SNR estimation unit 68 corrects the frequency threshold, sound pressure threshold, pitch threshold, speech threshold, reflection threshold, audio time lag threshold, air conditioner time lag threshold, and road noise time lag threshold based on the transfer function G. As a result, the SNR estimation unit 68 increases the SNR for the classified sound compared to before correction. This prevents a decrease in the speech recognition rate. The space estimation unit 64 corresponds to a calculation unit.

[0093] [2-2] The reference sound is an ultrasonic wave with a frequency of 20 kHz or more. Furthermore, ultrasonic waves are inaudible sounds. Therefore, discomfort felt by passengers when the transfer function G is calculated is suppressed.

[0094] (Third embodiment) The third embodiment differs from the second embodiment in the calculation of the transfer function G by the process of step S208 in the space estimation unit 64. Other than this, the third embodiment is similar to the second embodiment.

[0095] In step S208, the space estimation unit 64 acquires signals corresponding to the positions and number of occupants in the vehicle cabin from the occupant estimation unit 62. The space estimation unit 64 then calculates a transfer function G from the acquired positions and number of occupants, the window opening degree acquired in step S200, and a map. The map for calculating the transfer function G is set by experiment, simulation, or the like.

[0096] As described above, in the third embodiment, the space estimation unit 64 calculates the transfer function G. Even if the transfer function G is calculated in this way, the same effects as those of the first embodiment can be achieved. Furthermore, the third embodiment also has the following effects.

[0097] [3-1] In step S208, the space estimation unit 64 calculates a transfer function G, which is a value based on the occupant positions, the number of occupants, and the window opening degrees. Furthermore, the SNR estimation unit 68 corrects the frequency threshold, sound pressure threshold, pitch threshold, speech threshold, reflection threshold, audio time lag threshold, air conditioner time lag threshold, and road noise time lag threshold based on the transfer function G. As a result, the SNR estimation unit 68 increases the SNR for the classified sounds compared to before the correction. This prevents a decrease in the speech recognition rate.

[0098] [3-2] In step S102, the occupant estimation unit 62 estimates the occupant positions and the number of occupants based on values ​​related to the transmission and reception of ultrasonic waves with a frequency of 20 kHz or higher. Furthermore, ultrasonic waves are inaudible sounds. Therefore, discomfort felt by occupants when estimating the occupant positions and the number of occupants, which are parameters for calculating the transfer function G, is reduced.

[0099] (Fourth embodiment) In the fourth embodiment, the sensor group 50 of the microphone system 30 has a microphone position sensor 56 in addition to the occupant sensor 52 and the environment sensor 54, as shown in Fig. 14. Also, the processing of steps S400 and S404 by the SNR estimator 68 differs from that of the first embodiment. Other than this, the fourth embodiment is the same as the first embodiment.

[0100] The microphone position sensor 56 is disposed in the acoustic space Sb as shown in Fig. 15. The microphone position sensor 56 detects the position of each microphone 45 in the absolute coordinate system using an ultrasonic sensor or the like. The microphone position sensor 56 then outputs a signal corresponding to the detected position of each microphone 45 in the absolute coordinate system to the SNR estimation unit 68. Next, the processing of the SNR estimation unit 68 will be described with reference to the flowchart of Fig. 7.

[0101] In step S400, the SNR estimation unit 68 acquires information from the occupant estimation unit 62, the space estimation unit 64, and the vehicle state estimation unit 66, and also acquires signals corresponding to the positions of each microphone 45 in the absolute coordinate system from the microphone position sensor 56.

[0102] Subsequently, in step S402, the SNR estimation unit 68 performs the same process as in the first embodiment, so a description of the process in step S402 will be omitted.

[0103] In step S404 following step S402, the SNR estimation unit 68 performs frequency analysis on the sound data generated in step S402, as in the first embodiment, and estimates the speech time difference ΔTs, speech pitch P, and reflection time difference ΔTr. The SNR estimation unit 68 also estimates the sound pressure due to the speech of the occupants, the sound pressure due to the audio 12, the sound pressure due to the air conditioner 14, the sound pressure due to wind noise of the vehicle 5, and the sound pressure due to vibrations of the vehicle 5. The SNR estimation unit 68 also estimates the audio sound time difference ΔTa, the air conditioner sound time difference ΔTw, and the running sound time difference ΔTc.

[0104] The SNR estimation unit 68 also corrects the frequency threshold, sound pressure threshold, pitch threshold, and speech threshold using the positions of each microphone 45 in the absolute coordinate system acquired in step S400 and machine learning. Furthermore, the SNR estimation unit 68 corrects the reflection threshold, audio time difference threshold, air conditioner time difference threshold, and road noise time difference threshold using the positions of each microphone 45 in the absolute coordinate system acquired in step S400 and machine learning.

[0105] Then, the SNR estimation unit 68 classifies the types of sounds included in the sound data into speech types and noise types using the frequency-analyzed sound data and the corrected frequency threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the sound data generated in step S402 and the corrected sound pressure threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the speech pitch P estimated above and the corrected pitch threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the speech time difference ΔTs estimated above and the corrected speech threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the reflection time difference ΔTr estimated above and the corrected reflection threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the audio sound time difference ΔTa estimated above and the corrected audio time difference threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the air conditioner sound time difference ΔTw estimated above and the corrected air conditioner time difference threshold. The SNR estimation unit 68 also classifies the types of sounds included in the sound data into speech types and noise types using the road sound time difference ΔTc estimated above and the road sound time difference threshold. As a result, the SNR estimation unit 68 increases the SNR for speech compared to before each threshold was corrected.

[0106] Subsequently, in steps S408 to S412, the SNR estimation unit 68 performs the same processing as in the first embodiment. Therefore, a description of the processing in steps S408 to S412 will be omitted.

[0107] As described above, the SNR estimation unit 68 performs the process. In this way, even when the process is performed by the SNR estimation unit 68 of the fourth embodiment, the same effects as those of the first embodiment are achieved. Furthermore, the fourth embodiment also achieves the following effects.

[0108] [4] When the position of the microphone 45 is changed, the sound data collected by the microphone 45 changes, and the SNR for the voice changes accordingly. Therefore, the SNR estimation unit 68 corrects the frequency threshold, sound pressure threshold, pitch threshold, speech threshold, reflection threshold, audio time lag threshold, air conditioner time lag threshold, and road noise time lag threshold based on the position of the microphone 45. As a result, the SNR estimation unit 68 increases the SNR for the classified voice compared to before the correction. This prevents a decrease in the voice recognition rate.

[0109] (Other embodiments) The present disclosure is not limited to the above-described embodiments, and appropriate modifications can be made to the above-described embodiments. Furthermore, it goes without saying that the elements constituting the embodiments in the above-described embodiments are not necessarily essential unless they are specifically stated as essential or are considered to be clearly essential in principle.

[0110] The sound collection unit, clustering unit, output unit, calculation unit, modification unit, estimation unit, and methods described herein may be implemented by a special-purpose computer configured by configuring a processor and memory programmed to perform one or more functions embodied in a computer program. Alternatively, the sound collection unit, clustering unit, output unit, calculation unit, modification unit, estimation unit, and methods described herein may be implemented by a special-purpose computer configured by configuring a processor with one or more dedicated hardware logic circuits. Alternatively, the sound collection unit, clustering unit, output unit, calculation unit, modification unit, estimation unit, and methods described herein may be implemented by one or more special-purpose computers configured by combining a processor and memory programmed to perform one or more functions with a processor configured with one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by a computer on a computer-readable non-transitory tangible recording medium.

[0111] In each of the above embodiments, the number of frequency thresholds, sound pressure thresholds, pitch thresholds, speech thresholds, reflection thresholds, audio time lag thresholds, air conditioner time lag thresholds, and road noise time lag thresholds for classifying the types of voice and noise is one each. However, the number of each threshold is not limited to one, and may be two or more.

[0112] In each of the above embodiments, the parameters for classifying the types of sounds included in the sound data into speech types and noise types are the number of voices, the speech time difference ΔTs, the frequency components of the sound data, the speech pitch P, and the reflection time difference ΔTr. Furthermore, the parameters for classifying the types of sounds included in the sound data into speech types and noise types are the sound pressure due to the occupant's speech, the sound pressure due to the audio 12, the sound pressure due to the air conditioner 14, the sound pressure due to wind noise of the vehicle 5, and the sound pressure due to vibrations of the vehicle 5. Furthermore, the parameters for classifying the types of sounds included in the sound data into speech types and noise types are the audio sound time difference ΔTa, the air conditioner sound time difference ΔTw, and the running sound time difference ΔTc. In contrast, the SNR estimation unit 68 is not limited to classifying the types of sounds included in the sound data into speech types and noise types using all of the above parameters. The SNR estimation unit 68 may classify the types of sounds included in the sound data into speech types and noise types using at least one of the above parameters.

[0113] In each of the above embodiments, the value related to the sound reflected in the acoustic space Sb is the reflection time difference ΔTr. However, the value related to the sound reflected in the acoustic space Sb is not limited to the reflection time difference ΔTr. Because the reflectance and attenuation rate of the reflected sound differ depending on the state of the acoustic space Sb inside the vehicle cabin, the value related to the sound reflected in the acoustic space Sb may be, for example, the sound pressure of the sound reflected in the acoustic space Sb.

[0114] The above embodiments may be combined as appropriate.

[0115] (Features of the present invention) [Claim 1] a sound collection unit (S402) that collects sound into at least one microphone (45); a clustering unit (S404) that classifies the types of sounds included in the sound data, which is data related to the sounds collected by the microphones, into types of voices of people in the acoustic space and types of noise, which is sounds other than the voices, based on a value (ΔTr) related to the sounds reflected in the acoustic space (Sb), which is the space in which the microphones are placed; an output unit (S412) that outputs the classified data related to the voice to a voice recognition device (20); A microphone system comprising: [Claim 2] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to the ratio between the amplitude of a reference sound when the reference sound is collected by the microphone and the amplitude of the sound collected by the microphone, 2. The microphone system of claim 1, wherein the clustering unit classifies the sound using a value related to the sound reflected in the acoustic space and a reflection threshold, and corrects the reflection threshold based on the transfer function, thereby increasing the SNR of the classified sound compared to before the correction of the reflection threshold. [Claim 3] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle, 2. The microphone system of claim 1, wherein the clustering unit classifies the sound using a value related to the sound reflected in the acoustic space and a reflection threshold, and corrects the reflection threshold based on the transfer function, thereby increasing the SNR of the classified sound compared to before the correction of the reflection threshold. [Claim 4] 2. The microphone system of claim 1, wherein the clustering unit classifies the sound using a value related to the sound reflected in the acoustic space and a reflection threshold, and corrects the reflection threshold based on the position of the microphone, thereby increasing the SNR of the classified sound compared to before the correction of the reflection threshold. [Claim 5] The microphone system of claim 1, wherein the clustering unit classifies the types of sounds contained in the sound data into the type of sound and the type of noise based on values ​​(ΔTa, ΔTw, ΔTc) related to the arrival time difference of sounds other than the sound generated in the acoustic space for the same microphone. [Claim 6] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to the ratio between the amplitude of a reference sound when the reference sound is collected by the microphone and the amplitude of the sound collected by the microphone, 6. The microphone system of claim 5, wherein the clustering unit classifies the audio using the value related to the arrival time difference and a time difference threshold, and corrects the time difference threshold based on the transfer function, thereby increasing the SNR of the classified audio compared to before the time difference threshold was corrected. [Claim 7] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle, 6. The microphone system of claim 5, wherein the clustering unit classifies the audio using the value related to the arrival time difference and a time difference threshold, and corrects the time difference threshold based on the transfer function, thereby increasing the SNR of the classified audio compared to before the time difference threshold was corrected. [Claim 8] 6. The microphone system of claim 5, wherein the clustering unit classifies the audio using the value related to the arrival time difference and a time difference threshold, and corrects the time difference threshold based on the position of the microphone, thereby increasing the SNR of the classified audio compared to before the time difference threshold was corrected. [Claim 9] 2. The microphone system according to claim 1, wherein the clustering unit classifies the types of sounds included in the sound data into the voice type and the noise type based on a value related to the sound pressure of the sound data. [Claim 10] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to the ratio between the amplitude of a reference sound when the reference sound is collected by the microphone and the amplitude of the sound collected by the microphone, 10. The microphone system of claim 9, wherein the clustering unit classifies the sound data using a value related to the sound pressure and a sound pressure threshold, and corrects the sound pressure threshold based on the transfer function, thereby increasing the SNR of the classified sound compared to before the sound pressure threshold was corrected. [Claim 11] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle, 10. The microphone system of claim 9, wherein the clustering unit classifies the sound data using a value related to the sound pressure and a sound pressure threshold, and corrects the sound pressure threshold based on the transfer function, thereby increasing the SNR of the classified sound compared to before the sound pressure threshold was corrected. [Claim 12] 10. The microphone system of claim 9, wherein the clustering unit classifies the sound data using a value related to the sound pressure of the sound data and a sound pressure threshold, and corrects the sound pressure threshold based on the position of the microphone, thereby increasing the SNR of the classified sound compared to before the sound pressure threshold was corrected. [Claim 13] 2. The microphone system of claim 1, wherein the clustering unit classifies the types of sounds contained in the sound data into the types of voices and the types of noises based on a value related to a speech pitch (P), which is the interval between sounds in the speech of the person. [Claim 14] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to the ratio between the amplitude of a reference sound when the reference sound is collected by the microphone and the amplitude of the sound collected by the microphone, 14. The microphone system according to claim 13, wherein the clustering unit classifies the speech using the value related to the speech pitch and a pitch threshold, and corrects the pitch threshold based on the transfer function, thereby increasing an SNR for the classified speech compared to an SNR before correcting the pitch threshold. [Claim 15] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle, 14. The microphone system of claim 13, wherein the clustering unit classifies the speech into types of speech and types of noise based on the value related to the speech pitch and a pitch threshold, and corrects the pitch threshold based on the transfer function, thereby increasing the SNR of the classified speech compared to before the pitch threshold was corrected. [Claim 16] 14. The microphone system of claim 13, wherein the clustering unit classifies the speech using the value related to the speech pitch and a pitch threshold, and corrects the pitch threshold based on the position of the microphone, thereby increasing the SNR of the classified speech compared to before the pitch threshold was corrected. [Claim 17] The microphones are plural, The microphone system of claim 1, wherein the clustering unit classifies the types of sounds contained in the sound data into the types of sounds and the types of noise based on a value related to the arrival time difference (ΔTs) of the sounds between the microphones. [Claim 18] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to the ratio between the amplitude of a reference sound when the reference sound is collected by the microphone and the amplitude of the sound collected by the microphone, 18. The microphone system of claim 17, wherein the clustering unit classifies the speech using the value related to the arrival time difference and a speech threshold, and corrects the speech threshold based on the transfer function, thereby increasing the SNR of the classified speech compared to before the speech threshold was corrected. [Claim 19] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle, 18. The microphone system of claim 17, wherein the clustering unit classifies the speech using the value related to the arrival time difference and a speech threshold, and corrects the speech threshold based on the transfer function, thereby increasing the SNR of the classified speech compared to before the speech threshold was corrected. [Claim 20] 18. The microphone system of claim 17, wherein the clustering unit classifies the speech using the value related to the arrival time difference and a speech threshold, and corrects the speech threshold based on the position of the microphone, thereby increasing the SNR for the classified speech compared to before the speech threshold was corrected. [Claim 21] 2. The microphone system according to claim 1, wherein the clustering unit classifies the types of sounds included in the sound data into the types of voices and the types of noises based on values ​​related to frequency components of the sound data. [Claim 22] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to the ratio between the amplitude of a reference sound when the reference sound is collected by the microphone and the amplitude of the sound collected by the microphone, 22. The microphone system of claim 21, wherein the clustering unit classifies the audio using values ​​related to the frequency components and a frequency threshold, and corrects the frequency threshold based on the transfer function, thereby increasing the SNR of the classified audio compared to before the frequency threshold was corrected. [Claim 23] The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle, 22. The microphone system of claim 21, wherein the clustering unit classifies the audio using values ​​related to the frequency components and a frequency threshold, and corrects the frequency threshold based on the transfer function, thereby increasing the SNR of the classified audio compared to before the frequency threshold was corrected. [Claim 24] 22. The microphone system of claim 21, wherein the clustering unit classifies the audio using values ​​related to the frequency components and a frequency threshold, and corrects the frequency threshold based on the position of the microphone, thereby increasing the SNR of the classified audio compared to before the frequency threshold was corrected. [Claim 25] 2. The microphone system according to claim 1, wherein the clustering unit classifies the types of sounds included in the sound data into the types of voices and the types of noises based on the state of an audio system (12) and an air conditioner (14) of the vehicle (5), the speed of the vehicle, and the state of a road surface on which the vehicle is traveling. [Claim 26] 26. A microphone system as described in any one of claims 1 to 25, further comprising a change unit (S408, S410) that, when a value related to the SNR for the audio is less than a threshold (SNR_th), increases the number of microphones collecting the audio from the current number, thereby increasing the SNR for the classified audio compared to before the number of microphones collecting the audio was increased. [Claim 27] 27. The microphone system according to claim 26, wherein the change unit changes the number of microphones to be increased to collect sound, depending on a value related to the number of types of noise. [Claim 28] 28. The microphone system according to claim 26, wherein the change unit changes the number of microphones to be increased to collect sound in accordance with a value related to the sound pressure of the noise. [Claim 29] 29. The microphone system according to claim 26, wherein the change unit changes the number of microphones to be increased to collect sound in accordance with a value related to the sound pressure of the sound. [Claim 30] 23. The microphone system according to claim 2, wherein the reference sound is an ultrasonic wave having a frequency of 20 kHz or more. [Claim 31] 24. The microphone system according to claim 3, further comprising an estimation unit (S102) that estimates the passenger positions and the number of passengers based on values ​​related to transmission and reception of ultrasonic waves having a frequency of 20 kHz or more. [Claim 32] 32. The microphone system according to claim 1, wherein the clustering unit estimates values ​​related to the sound reflected in the acoustic space based on the sound data, the size of the acoustic space, and the positions and sizes of objects in the acoustic space. [Explanation of symbols]

[0116] 10 Vehicle Systems 30 microphone system 40 microphone array 45 microphones 50 sensors 60 Arithmetic unit 62 Occupant Estimation Department 64 Spatial estimation part 66 Vehicle state estimation unit 68 SNR estimation unit

Claims

1. a sound collection unit (S402) that collects sound into at least one microphone (45); a clustering unit (S404) that classifies the types of sounds included in the sound data, which is data related to the sounds collected by the microphones, into types of voices of people in the acoustic space and types of noise, which is sounds other than the voices, based on a value (ΔTr) related to the sounds reflected in the acoustic space (Sb), which is the space in which the microphones are placed; an output unit (S412) that outputs the classified data related to the voice to a voice recognition device (20); Equipped with A microphone system in which the value related to the sound reflected in the acoustic space is the arrival time difference of the sound reflected in the acoustic space to the same microphone.

2. The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value related to a ratio between an amplitude of a reference sound when the reference sound is collected by the microphone and an amplitude of the sound collected by the microphone, 2. The microphone system according to claim 1, wherein the clustering unit classifies the sound using a value related to the sound reflected in the acoustic space and a reflection threshold, and corrects the reflection threshold based on the transfer function, thereby increasing the SNR of the classified sound compared to before the correction of the reflection threshold.

3. The microphone system further includes a calculation unit (S208) that calculates a transfer function (G) that is a value based on the occupant positions of the vehicle (5), the number of occupants, and the opening degree of the side window of the vehicle.

2. The microphone system according to claim 1, wherein the clustering unit classifies the sound using a value related to the sound reflected in the acoustic space and a reflection threshold, and corrects the reflection threshold based on the transfer function, thereby increasing the SNR of the classified sound compared to before the correction of the reflection threshold.

4. 2. The microphone system of claim 1, wherein the clustering unit classifies the sound using a value related to the sound reflected in the acoustic space and a reflection threshold, and corrects the reflection threshold based on the position of the microphone, thereby increasing the SNR of the classified sound compared to before the reflection threshold was corrected.

5. 5. The microphone system of claim 1, further comprising a change unit (S408, S410) that, when the value related to the SNR for the audio is less than a threshold (SNR_th), increases the number of microphones collecting the audio from the current number, thereby increasing the SNR for the classified audio compared to before the number of microphones collecting the audio was increased.

6. The microphone system according to claim 5 , wherein the change unit changes the number of microphones that collect sounds in accordance with a value related to the number of types of noise.

7. The microphone system according to claim 5 , wherein the change unit changes the number of microphones to be increased to collect sound in accordance with a value related to the sound pressure of the noise.

8. The microphone system according to claim 5 , wherein the change unit changes the number of microphones to be increased to collect the sound in accordance with a value related to the sound pressure of the sound.

9. 3. The microphone system according to claim 2, wherein the reference sound is an ultrasonic wave having a frequency of 20 kHz or more.

10. The microphone system according to claim 3 , further comprising an estimation unit (S102) that estimates the occupant positions and the number of occupants based on values ​​related to transmission and reception of ultrasonic waves having a frequency of 20 kHz or more.

11. The microphone system according to claim 1 , wherein the clustering unit estimates values ​​related to the sound reflected in the acoustic space based on the sound data, the size of the acoustic space, and the positions and sizes of objects in the acoustic space.

Citation Information

Patent Citations

  • Sound signal processing method, device, and program

    JP2007010897A

  • Sound source separation and localization method

    JP2008145610A

  • Signal processing device, signal processing method and program

    JP2011215357A

  • Sound collection processor, sound collection processing method and program

    JP2021012314A