Information processing system, information processing method, and program

The system uses internal vehicle microphones to generate phase and acoustic features for accurate sound source estimation, addressing placement challenges and enhancing vehicle control and entertainment management.

WO2026004622A1PCT designated stage Publication Date: 2026-01-02SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021259
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-12
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately estimating the type and direction of acoustic events using microphones installed inside vehicles due to placement restrictions and environmental factors, which affect the accuracy of sound source detection.

Method used

An information processing system that utilizes multiple microphones inside a vehicle to generate input features indicating phase differences and acoustic features, employing signal processing and machine learning to estimate the direction and type of acoustic events, such as sirens or other sounds, by distributing microphones to distinguish between internal and external sources.

Benefits of technology

The system accurately estimates the type and direction of acoustic events, enabling effective vehicle control and in-car entertainment management by distinguishing between internal and external sounds, reducing design constraints and improving accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025021259_02012026_PF_FP_ABST
    Figure JP2025021259_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing system, an information processing method, and a program that make it possible to more accurately estimate the type of an acoustic event and the direction of a sound source. In the present invention, a first input feature amount indicating a phase difference between a plurality of sound signals and a second input feature amount indicating an acoustic feature are generated on the basis of a sound signal in which sound arriving from a predetermined sound source is collected and inputted by a plurality of microphones distributed and arranged inside a vehicle, and the direction of the sound source, whether the sound source is inside or outside the vehicle, and the type of an acoustic event are estimated by signal processing using the first input feature amount and the second input feature amount. The present technology can be applied to, for example, an environmental sound monitoring system of a vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and program

[0001] The present disclosure relates to an information processing system, an information processing method, and a program, and more particularly to an information processing system, an information processing method, and a program that enable more accurate estimation of the type of acoustic event and the direction of the sound source.

[0002] Conventionally, in order to popularize in-car entertainment services in autonomous vehicles, for example, it has been considered important to monitor environmental sounds inside or outside the vehicle.

[0003] Patent Document 1 discloses a technology that uses multiple microphones installed outside the vehicle to detect siren noise corresponding to an emergency vehicle, estimates the direction of the emergency vehicle, determines how to respond to the emergency vehicle, and controls the vehicle in autonomous driving mode.

[0004] Japanese Patent Application Laid-Open No. 2023-146132

[0005] However, when installing a microphone outside the vehicle, it is preferable to install the microphone inside the vehicle due to, for example, the cost of weatherproofing and dustproofing, placement restrictions, etc. However, it has been difficult to accurately estimate the type of acoustic event and the direction of the sound source with a microphone installed inside the vehicle.

[0006] The present disclosure has been made in view of such circumstances, and aims to make it possible to more accurately estimate the type of acoustic event and the direction of the sound source.

[0007] An information processing system according to one aspect of the present disclosure includes a generation unit that generates, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sounds arriving from a predetermined sound source, first input features that indicate a phase difference between the plurality of sound signals and second input features that indicate acoustic features, and an estimation unit that estimates the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features.

[0008] An information processing method or program according to one aspect of the present disclosure includes generating, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sounds arriving from a predetermined sound source, a first input feature that indicates a phase difference between the plurality of sound signals and a second input feature that indicates an acoustic feature, and estimating the direction of the sound source and the type of acoustic event by signal processing using the first input feature and the second input feature.

[0009] In one aspect of the present disclosure, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sound arriving from a predetermined sound source, a first input feature indicating a phase difference between the plurality of sound signals and a second input feature indicating an acoustic feature are generated, and the direction of the sound source and the type of acoustic event are estimated by signal processing using the first input feature and the second input feature.

[0010] FIG. 1 is a diagram illustrating an example of microphone placement in an environmental sound monitoring system to which the present technology is applied. FIG. 2 is a block diagram illustrating an example configuration of an environmental sound monitoring system. FIG. 3 is a diagram illustrating microphone placement conditions. FIG. 4 is a diagram illustrating an example of the distance from a microphone to a sound source. FIG. 5 is a flowchart illustrating environmental sound monitoring processing. FIG. 6 is a block diagram illustrating an application example of an environmental sound monitoring system. FIG. 7 is a diagram illustrating a configuration in which one sound signal is input. FIG. 8 is a diagram illustrating a configuration in which multiple sound signals are input. FIG. 9 is a diagram illustrating the CSP method. FIG. 10 is a diagram illustrating an input method for the CSP method. FIG. 11 is a block diagram illustrating an example configuration of a machine learning model. FIG. 12 is a block diagram illustrating an example configuration of an embodiment of a computer to which the present technology is applied.

[0011] Hereinafter, specific embodiments to which the present technology is applied will be described in detail with reference to the drawings.

[0012] <Configuration Example of Environmental Sound Monitoring System> FIG. 1 is a diagram showing an example of the arrangement of microphones 21 in an environmental sound monitoring system 11, which is an information processing system to which the present technology is applied.

[0013] As shown in FIG. 1, the environmental sound monitoring system 11 is configured by four microphones 21 - 1 to 21 - 4 provided in a vehicle 12 and connected to a control device 22 .

[0014] The microphones 21-1 to 21-4 are, for example, distributed throughout the body of the vehicle 12, capable of collecting sounds inside the vehicle 12, collect sounds inside and outside the vehicle 12, and supply sound signals corresponding to the sounds to the control device 22. Note that while Fig. 1 shows an example configuration including four microphones 21-1 to 21-4, the environmental sound monitoring system 11 may be configured to include any number (N) of microphones 21, such as at least three. Hereinafter, when there is no need to distinguish between the microphones 21-1 to 21-4, they will be referred to simply as microphones 21.

[0015] The control device 22 controls various devices (such as a speaker 31, a display 32, and a vehicle control device 33 shown in the block diagram of FIG. 2) equipped on the vehicle 12 based on the sound signals supplied from the microphones 21-1 to 21-4.

[0016] For example, in the environmental sound monitoring system 11, microphones 21-1 to 21-4 are arranged at four corners inside the vehicle 12 (at least three microphones 21 are arranged at any of the four corners) so as to be able to distinguish between sound coming from a sound source P1 inside the vehicle 12 and sound coming from a sound source P2 outside the vehicle 12. This allows the environmental sound monitoring system 11 to estimate the direction of the sound source based on the phase difference between the sound signals picked up by the microphones 21-1 to 21-4, and to distinguish between the sound source P1 inside the vehicle 12 and the sound source P2 outside the vehicle 12.

[0017] The environmental sound monitoring system 11 can then accurately estimate the direction of the sound source and the type of acoustic event based on the sounds picked up by the microphones 21-1 to 21-4 arranged inside the vehicle 12, and perform environmental sound monitoring.

[0018] FIG. 2 is a diagram showing an example of the configuration of an embodiment of the environmental sound monitoring system 11.

[0019] 2, the environmental sound monitoring system 11 is configured to include a speaker 31, a display 32, and a vehicle control device 33 in addition to the microphone 21 and the control device 22 shown in Fig. 1. The control device 22 is configured to include an input feature generation unit 41, an acoustic sound source estimation unit 42, and a control unit 43.

[0020] The input feature generation unit 41 generates input features required to be input to a deep neural network (DNN) of the acoustic sound source estimation unit 42 based on the sound signals supplied from the N microphones 21-1 to 21-N, and supplies the generated input features to the acoustic sound source estimation unit 42. For example, the input features include an input feature for acoustic event estimation indicating acoustic features for estimating the type of acoustic event, and an input feature for sound source direction estimation indicating a phase difference between a plurality of sound signals for estimating the direction of a sound source.

[0021] The acoustic sound source estimation unit 42 estimates the type of acoustic event and the direction of the sound source by signal processing using the input features supplied from the input feature generation unit 41, for example by inputting the input features into a DNN, and supplies estimation result information indicating the estimation result to the control unit 43.

[0022] When control based on the type of acoustic event and the direction of the sound source is required in accordance with the estimation result information supplied from the acoustic sound source estimation unit 42, the control unit 43 controls the speaker 31, the display 32, and the vehicle control device 33. For example, when a siren sound of an emergency vehicle is detected, the control unit 43 controls the speaker 31 and the display 32 so as to output a warning calling attention to the emergency vehicle from the speaker 31 and display it on the display 32. Furthermore, the control unit 43 controls the vehicle control device 33 according to the direction from which the siren sound of the emergency vehicle is coming so as not to interfere with the travel of the emergency vehicle.

[0023] The environmental sound monitoring system 11 is configured in this manner, and is capable of accurately estimating the type of acoustic event and the direction of the sound source, and is able to appropriately perform control based on the type of acoustic event and the direction of the sound source.

[0024] <Microphone Placement Conditions> With reference to FIGS. 3 and 4, the placement conditions of the microphone 21 required to distinguish between a sound source P1 inside the vehicle 12 and a sound source P2 outside the vehicle 12 in the environmental sound monitoring system 11 will be described.

[0025] The environmental sound monitoring system 11 needs to arrange the microphones 21-1 to 21-4 under placement conditions that enable it to estimate the direction of the sound source based on the phase difference (i.e., relative time delay / advance) between the sound signals picked up by the microphones 21-1 to 21-4 and distinguish between a sound source P1 inside the vehicle 12 and a sound source P2 outside the vehicle 12.

[0026] For example, to accommodate sounds arriving from sound sources in all directions on a two-dimensional plane, at least three or more microphones 21 are required. If the distance from the sound source to a microphone 21 is sufficiently large relative to the spacing between the microphones 21, the plane wave assumption is valid for the sound arriving from that sound source. On the other hand, if the distance from the sound source to a microphone 21 is relatively small relative to the spacing between the microphones 21, the spherical wave assumption is valid for the sound arriving from that sound source.

[0027] Therefore, it is preferable that the microphones 21-1 to 21-4 in the environmental sound monitoring system 11 are arranged so that the plane wave assumption holds for sounds coming from a sound source P2 outside the vehicle 12, and the spherical wave assumption holds for sounds coming from a sound source P1 inside the vehicle 12.

[0028] As shown in A of FIG. 3, a first arrangement example will be described in which the microphones 21-1 to 21-4 are arranged at a narrow distance of, for example, about 10 cm. In the first arrangement example, the distance from the sound source P1 inside the vehicle 12 to the microphone 21 is relatively large compared to the distance between the microphones 21-1 to 21-4, and the sound arriving at the microphones 21-1 to 21-4 from the sound source P1 inside the vehicle 12 is a plane wave. Furthermore, the sound arriving at the microphones 21-1 to 21-4 from the sound source P2 outside the vehicle 12 is also a plane wave. Therefore, in the first arrangement example, it is not possible to obtain a phase difference between multiple sound signals that allows for distinguishing between the sound source P1 inside the vehicle 12 and the sound source P2 outside the vehicle 12, making it difficult to distinguish between the sound source P1 and the sound source P2. Note that, even in the first arrangement example, it is possible to estimate the directions of the sound source P1 and the sound source P2.

[0029] As shown in B of FIG. 3, a second arrangement example will be described in which microphones 21-1 to 21-4 are arranged to surround the driver's seat and passenger seat of vehicle 12. In the second arrangement example, even if sound source P1 inside vehicle 12 is outside the frame surrounded by microphones 21-1 to 21-4, the distance from sound source P1 to microphone 21 is relatively small compared to the spacing between the multiple microphones 21, and the sound arriving from sound source P1 to microphones 21-1 to 21-4 is a spherical wave. On the other hand, the sound arriving from sound source P2 outside vehicle 12 to microphones 21-1 to 21-4 is a plane wave. Therefore, in the second arrangement example, a phase difference between multiple sound signals that allows for distinguishing between sound source P1 inside vehicle 12 and sound source P2 outside vehicle 12 can be obtained, making it possible to distinguish between sound source P1 and sound source P2.

[0030] As shown in FIG. 3C, a third arrangement example will be described in which the microphones 21-1 to 21-4 are arranged to surround the driver's seat, passenger seat, and rear seat of the vehicle 12. In the third arrangement example, a sound source P1 inside the vehicle 12 is within the frame surrounded by the microphones 21-1 to 21-4, and a sound source P2 outside the vehicle 12 is outside the frame surrounded by the microphones 21-1 to 21-4. Therefore, the microphones 21-1 to 21-4 have different characteristics in terms of the combination of the time difference (fast / slow) between the sound arriving from the sound source P1 and the sound arriving from the sound source P2. Therefore, in the third arrangement example, it is possible to reliably obtain a phase difference between multiple sound signals that allows the sound source P1 inside the vehicle 12 and the sound source P2 outside the vehicle 12 to be distinguished from each other, thereby reliably distinguishing between the sound source P1 and the sound source P2.

[0031] Referring to FIG. 4, the following describes sounds coming from sound sources P1 to P4 that are positioned at predetermined distances from microphones 21-1 to 21-4 that are positioned at the four corners inside the vehicle 12.

[0032] The sound source P1 is located inside the vehicle 12, within a frame surrounded by the microphones 21-1 to 21-4 located at the four corners inside the vehicle 12. In this case, the time difference between the sound arriving at the microphones 21-2 and 21-4 from the sound source P1 is sufficiently distinguishable, and the direction of the sound source P1 can be estimated.

[0033] The sound source P2 is located inside the vehicle 12 and outside the frame surrounded by the microphones 21-1 to 21-4. As with the sound source P1, the direction of the sound source P2 can also be estimated.

[0034] Sound source P3 is located outside vehicle 12 and at a position relatively close to vehicle 12. Sound source P4 is located outside vehicle 12 and at a position relatively far from vehicle 12. For such sound source P3 and sound source P4, the time difference between the sounds arriving at microphone 21-2 and microphone 21-4, respectively, almost disappears, and the sound signal gradually becomes equivalent to the characteristics obtained from outside and behind vehicle 12.

[0035] Therefore, the environmental sound monitoring system 11 sets the placement conditions of the microphones 21-1 to 21-4 so that a significant difference in arrival time can be formed as a sound source inside the vehicle 12 when compared with the characteristics of a specific direction outside the vehicle 12.

[0036] Here, the range in which the spherical wave assumption holds true is one in which the distance ρ from the center of the microphones 21-1 to 21-4 to the sound source P must satisfy the following formula (1) based on the interval D between the microphones 21-1 to 21-4 and the wavelength λ of the target sound source. Note that the wavelength λ can be calculated as λ = c / f using the speed of sound c and the frequency f of the target sound source.

[0037]

[0038] Therefore, if the distance D between the microphones 21-1 to 21-4 is 0.25 m and the frequency f of the target sound source is 1 kHz, then according to equation (1), the range in which the spherical wave assumption holds true is when the distance ρ from the center of the microphones 21-1 to 21-4 to the sound source P is less than 0.37 m. Similarly, if the distance D between the microphones 21-1 to 21-4 is 0.5 m and the frequency f of the target sound source is 1 kHz, then according to equation (1), the range in which the spherical wave assumption holds true is when the distance ρ from the center of the microphones 21-1 to 21-4 to the sound source P is less than 1.47 m.

[0039] Furthermore, if the distance D between the microphones 21-1 to 21-4 is 1.0 m and the frequency f of the target sound source is 1 kHz, then according to equation (1), the range in which the spherical wave assumption holds true is when the distance ρ from the center of the microphones 21-1 to 21-4 to the sound source P is less than 13.24 m. Similarly, if the distance D between the microphones 21-1 to 21-4 is 2.0 m and the frequency f of the target sound source is 1 kHz, then according to equation (1), the range in which the spherical wave assumption holds true is when the distance ρ from the center of the microphones 21-1 to 21-4 to the sound source P is less than 23.53 m.

[0040] Therefore, for example, assuming a situation in which the interior dimensions of the vehicle 12 are 2 m vertically and 1 m horizontally, and an approaching emergency vehicle is located approximately 20 m behind the vehicle 12 with one other vehicle in between, the distance D between the microphones 21-1 to 21-4 is preferably in the range of 0.5 m to 1.5 m. On the other hand, if the distance D between the microphones 21-1 to 21-4 is 0.25 m, the entire space inside the vehicle 12 cannot be considered to be a spherical wave environment, and if the distance D between the microphones 21-1 to 21-4 is 2.0 m, the sound source P2 outside the vehicle 12 cannot be considered to be a plane wave environment.

[0041] The placement conditions of the microphones 21-1 to 21-4 and the frequency bands of the target sound sources described here are only rough guidelines, as they depend on the application. For example, if the application considers the driver's mouth to be a sound source inside the vehicle 12 and only requires that it be distinguishable from sound sources outside the vehicle 12, then even if the distance D between the microphones 21-1 to 21-4 is 0.25 m, this will be sufficient to cover the target range.

[0042] FIG. 5 is a flowchart illustrating the environmental sound monitoring process executed in the environmental sound monitoring system 11.

[0043] For example, the processing starts when a control system that controls the entire vehicle 12 is activated and sound signals of sounds picked up by the microphones 21-1 to 21-N are input to the control device 22. In step S11, the input feature generation unit 41 generates input features for estimating a sound event and a sound source direction based on the sound signals supplied from the microphones 21-1 to 21-N, and supplies the input features to the acoustic sound source estimation unit 42.

[0044] In step S12, the acoustic sound source estimation unit 42 estimates the type of acoustic event and the direction of the sound source by inputting the input feature supplied from the input feature generation unit 41 in step S11 into the DNN, and supplies estimation result information indicating the estimation result to the control unit 43.

[0045] In step S13, the control unit 43 determines whether or not control based on the type of sound event and the direction of the sound source is required, in accordance with the estimation result information supplied from the sound source estimation unit 42 in step S12.

[0046] If the control unit 43 determines in step S13 that the situation requires control based on the type of acoustic event and the direction of the sound source, the process proceeds to step S14.

[0047] In step S14, the control unit 43 controls the speaker 31, the display 32, and the vehicle control device 33 as described above, based on the type of acoustic event and the direction of the sound source.

[0048] After processing in step S14, or if it is determined in step S13 that control based on the type of acoustic event and the direction of the sound source is not required, processing returns to step S11, and the same processing is repeated thereafter.

[0049] By performing the above-described environmental sound monitoring process, the environmental sound monitoring system 11 can accurately estimate the type of acoustic event and the direction of the sound source, and can appropriately perform control based on the type of acoustic event and the direction of the sound source.

[0050] Fig. 6 is a block diagram showing an application example of the environmental sound monitoring system 11. In the environmental sound monitoring system 11A shown in Fig. 6, components common to the environmental sound monitoring system 11 in Fig. 2 are denoted by the same reference numerals, and detailed description thereof will be omitted.

[0051] As shown in Fig. 6, the environmental sound monitoring system 11A has a configuration in common with the environmental sound monitoring system 11 of Fig. 2 in that it includes microphones 21-1 to 21-N, a speaker 31, a display 32, and a vehicle control device 33. However, the environmental sound monitoring system 11A is configured to include a control device 22A, an audio device 23, an agent processing unit 24, and peripheral devices 34, and the control device 22A is configured to include an input feature generation unit 41A and an acoustic sound source estimation unit 42A, which is a difference from the environmental sound monitoring system 11 of Fig. 2.

[0052] The audio device 23 reproduces and outputs sounds (such as music and audio from moving images) for in-car entertainment services, for example. To prevent malfunction of the in-car environment caused by the agent processing unit 24, the audio device 23 supplies audio signals of the sounds reproduced in the in-car entertainment services to the input feature generation unit 41A.

[0053] When generating input features and estimating the type of sound event and the direction of the sound source as described above, the input feature generation unit 41A and the sound source estimation unit 42A can feed back the audio signal supplied from the audio device 23 and perform sound processing to cancel acoustic echoes inside the vehicle 12. Then, the sound source estimation unit 42A supplies the estimation results of the type of sound event and the direction of the sound source to the agent processing unit 24.

[0054] Based on the estimation results supplied from the acoustic sound source estimation unit 42A, the agent processing unit 24 functions as an agent for the vehicle 12, for example, by assisting the driver, automatically adjusting the peripheral devices 34 that adjust various internal environments of the vehicle 12, interacting with the occupants of the vehicle 12, and controlling the vehicle 12.

[0055] The environmental sound monitoring system 11A configured in this manner can suppress the adverse effects of acoustic echo caused by the sound reproduced by the audio equipment 23 by feeding back the audio signal, and can more reliably estimate the type of acoustic event and the direction of the sound source.

[0056] <Estimation of Type of Sound Event and Direction of Sound Source> Estimation of type of sound event and direction of sound source will be described with reference to FIGS. 7 and 8. FIG.

[0057] FIG. 7 is a diagram illustrating estimation of the type of an acoustic event when one sound signal is input.

[0058] In the environmental sound monitoring system 11, an estimation method using DNN is used to associate various types of acoustic events (e.g., ambulances, fire engines, railroad crossings, etc.) that are to be estimated from sound signals with each class and register them in advance, and then the system is trained as a classification problem, making it possible to estimate (detect, identify) the types of acoustic events.

[0059] In general, the input feature generation unit 41 generates (extracts) input features by performing a fast Fourier transform (FFT) on the sound signal input from the microphone 21 and separating it into frequency components. Then, the acoustic sound source estimation unit 42 inputs the input features to a DNN, outputs the likelihood of each class between 0 and 1, and outputs the type of acoustic event (for example, an ambulance, a fire engine, or a railroad crossing) associated with a class that exceeds a predetermined threshold as an estimation result.

[0060] With such an estimation method using a DNN, for example, the types of acoustic events required for an application can be registered as classes and trained to obtain the output of the required types of acoustic events. It is also possible to configure a DNN model that makes the input sound signal multi-channel to increase robustness against wind and road noise, and that simultaneously estimates directional information.

[0061] In addition to the estimation method using DNN as described above, a pattern matching method may be used in which, for example, a reference for the sound to be estimated as the type of acoustic event is registered in advance and compared with the sound signal.

[0062] FIG. 8 is a diagram illustrating estimation of the type of acoustic event and the direction of the sound source when a plurality of sound signals are input.

[0063] In the environmental sound monitoring system 11, it is preferable to generate separate input features for each purpose, such as input features for sound event estimation used to estimate the type of sound event and input features for sound source direction estimation used to estimate the direction of a sound source. Therefore, the input feature generation unit 41 is configured to include a sound event estimation feature generation unit 51 that generates input features for sound event estimation, and a sound source direction estimation feature generation unit 52 that generates input features for sound source direction estimation.

[0064] The acoustic event estimation feature generation unit 51 generates input features for acoustic event estimation, for example, by converting sound signals input from multiple microphones 21 into frequency spectra or mel spectrograms that can easily capture acoustic features using processing such as FFT.

[0065] The sound source direction estimation feature generation unit 52 generates input features for sound source direction estimation by, for example, estimating the phase difference between sound signals input from multiple microphones 21 using the CSP (Cross-Power Spectrum Phase) method.

[0066] The acoustic sound source estimation unit 42 then inputs the input feature values ​​for acoustic event estimation and the input feature values ​​for sound source direction estimation into the DNN, and can simultaneously estimate the type of acoustic event and the direction of the sound source. Of course, the acoustic sound source estimation unit 42 can also estimate the type of acoustic event and the direction of the sound source separately. However, in order to link and output which acoustic event corresponds to which direction the sound source came from, it is expected that the acoustic sound source estimation unit 42 can further improve performance by combining (fusing) the input feature values ​​for acoustic event estimation and the input feature values ​​for sound source direction estimation inside the DNN.

[0067] For example, the acoustic sound source estimation unit 42 inputs input feature amounts for sound source direction estimation to a DNN, and estimates predictions and directions of the inside and outside of the vehicle 12 from the phase difference between sound signals input from the multiple microphones 21. In this case, it is expected that the accuracy of determining the inside and outside of the vehicle 12 can be improved by acquiring features such as the transmission characteristics of the vehicle 12 as intermediate outputs within the DNN of the input feature amounts for sound event estimation.

[0068] The acoustic sound source estimation unit 42 then outputs the likelihood for each class associated with the type of acoustic event, the direction and distance of the sound source, and the determination result of whether it is inside or outside the vehicle 12 as estimation results.

[0069] In such estimation methods using a DNN, there are various input feature amounts input to the DNN, various configurations of the DNN, and various output methods of the DNN. For example, there may be a dedicated module for determining whether a sound source is inside or outside the vehicle 12, or the type of acoustic event and the direction of the sound source may be estimated using separate machine learning models.

[0070] In addition to outputting acoustic events inside the vehicle 12 and acoustic events outside the vehicle 12 separately as shown in the figure, a configuration may also be adopted in which, for example, a dedicated output node is provided for distinguishing between the inside and outside of the vehicle 12.

[0071] <Regarding the CSP Method> The CSP method used in the environmental sound monitoring system 11 will be described with reference to FIGS.

[0072] For example, if omnidirectional microphones 21 are simply distributed inside the vehicle 12, a conventional estimation method using phase difference information in the frequency domain cannot accurately estimate the direction of the sound source because the microphone spacing, which is much larger than half the wavelength of the target sound source, affects the estimation and causes spatial aliasing. Therefore, the environmental sound monitoring system 11 uses the CSP method, which calculates the time difference in the waveform domain between sound signals input from the multiple microphones 21.

[0073] 9A shows an example of the waveforms of the sound signals input by the microphone 21 A and the microphone 21 B. As shown in the figure, the phase difference between the sound signal of the microphone 21 A and the sound signal of the microphone 21 B is measured as the time difference between the arrival of the sounds at each microphone, and the direction of the sound source can be estimated based on the phase difference.

[0074] However, in reality, the sound signals input by the microphones 21A and 21B have a noise-like waveform as shown in FIG. 9B.

[0075] Therefore, the sound source direction estimation feature generating unit 52 calculates the CSP value by the CSP method from the phase difference between the sound signals as shown in B of FIG. 9, thereby making it possible to estimate a peak in the CSP region as shown in C of FIG. 9.

[0076] Furthermore, a machine learning model can be used to determine whether a sound source is inside or outside the vehicle 12. For example, the characteristic difference between the sound signals of a sound coming from inside the distributed microphones 21 and a sound coming from outside the distributed microphones 21 can be learned, and the direction of the sound source and whether the sound source is inside or outside the vehicle 12 can be determined according to the CSP values ​​obtained between the distributed microphones 21.

[0077] Furthermore, by distributing the microphones 21 (for example, distributing them at the four corners of the vehicle 12 as shown in FIG. 1 ), it becomes easier to collect sounds inside the vehicle 12, reducing restrictions on the placement of the microphones 21. Installing a conventional microphone array for direction estimation requires, for example, space for installing a housing of approximately 10 cm × 10 cm somewhere inside the vehicle 12 or space for embedding the housing in the interior of the vehicle 12, which can easily impose design restrictions. In contrast, when distributing the microphones 21, it is sufficient to secure space for one wire and the size of a single microphone element for each microphone hole provided for each microphone 21. Furthermore, by combining signal processing using the CSP method, it is possible to estimate the type of acoustic event inside and outside the vehicle 12 and estimate the direction of a sound source even using inexpensive omnidirectional microphones 21. Furthermore, it is possible to use this system in combination with microphones 21 for voice recognition or road noise removal.

[0078] In this way, the environmental sound monitoring system 11 can use the CSP values ​​as an example of input features for sound source direction estimation to be input to the DNN. Note that the environmental sound monitoring system 11 may input sound signals input from the multiple microphones 21 directly to the DNN, expecting the DNN to acquire the necessary features during the learning process.

[0079] Furthermore, in the environmental sound monitoring system 11, as preprocessing for the machine learning model, sound signals input from the multiple microphones 21 are converted into phase differences, which makes it easier to estimate the direction of the sound source and distinguish between sound sources inside and outside the vehicle 12. For example, by utilizing the machine learning model, the acoustic characteristics of multi-channel signals that have passed through the body or windows of the vehicle 12, as well as changes in characteristics when the windows of the vehicle 12 are opened, are also taken into account during learning. This makes it possible to predict characteristics that differ from sound propagation in free space without obstructions by collecting data on environments that can actually occur.

[0080] FIG. 10 is a diagram illustrating an input method of the CSP method.

[0081] For example, the CSP method calculates the correlation between sound signals for a signal section having a certain length and calculates the time delay or time advance at the sample level in order to extract phase difference information between multiple microphones 21. Generally, the longer the window width for calculating the correlation, the more robust it is to noise, but the greater the amount of calculation.

[0082] However, because the vehicle 12 itself and the target sound source are moving, increasing the window width not only increases the amount of calculation but also causes ambiguity in the phase difference information. For example, if the window width for calculating the CSP value is set to 4 seconds and the phase difference information of the siren of an emergency vehicle approaching the vehicle 12 from behind during those 4 seconds is calculated, a CSP value that can interpret a specific direction can be calculated. On the other hand, if the window width for calculating the CSP value is set to 4 seconds and the phase difference information of the siren of an emergency vehicle approaching the vehicle 12 from ahead and passing behind during those 4 seconds is calculated, various phase differences exist during those 4 seconds, making it difficult to correctly calculate the CSP value.

[0083] Therefore, in order to overcome the advantages and disadvantages of long and short window lengths, the environmental sound monitoring system 11 employs multi-temporal CSP. That is, CSP values ​​for multiple window lengths are calculated, and samples of several milliseconds on both ends of the CSP values, which represent phase difference information between multiple microphones 21 relative to the physical sound source, are extracted to extract time lag / lead information. This time lag / lead information is then calculated for each window length and channel combination, and used as input features to be input to the DNN.

[0084] 10, CSP values ​​are calculated for eight window lengths (0.25 seconds, 0.75 seconds, 1.25 seconds, ..., 4.0 seconds), and samples of 0.0625 seconds on both ends of the CSP values ​​are extracted. Then, the combinations between the four-channel microphones 21 (x6) and the number of eight window lengths (x8) are put into a matrix, which is used as the input feature to be input to the DNN.

[0085] By using such multi-temporal CSP, CSP values ​​are calculated for multiple window lengths, and a matrix combining output intervals that are effective for estimating the direction of the sound source is extracted as input features (CSP features) and input to the DNN. This makes it possible to achieve robustness against noise and to select the optimal window length for a moving object.

[0086] <Configuration Example of Machine Learning Model> With reference to FIG. 11, a configuration example of a machine learning model that constitutes the input feature quantity generating unit 41 and the acoustic sound source estimating unit 42 will be described.

[0087] 11 , the machine learning model 61 includes a CSP processing unit 71, a DoA (Direction of Arrival) module 72, a STFT (Short-time Fourier Transform) processing unit 73, a BF (Beamforming) module 74, and an SED (Sound Event Detection) module 75. For example, the DoA module 72, the BF module 74, and the SED module 75 are functional blocks having learnable parameters.

[0088] By executing the multi-temporal CSP as described above, the CSP processing unit 71 acquires phase difference information (CSP features) indicating the phase difference between the multi-channel (M-ch) sound signals input to the machine learning model 61, and supplies the information to the DoA module 72.

[0089] The DoA module 72 uses the phase difference information supplied from the CSP processing unit 71 to output likelihoods for each class associated with the types of acoustic events inside and outside the vehicle 12, an estimated direction of a sound source inside the vehicle 12, and an estimated direction of a sound source outside the vehicle 12. Furthermore, the DoA module 72 shifts the estimated direction of a sound source inside the vehicle 12 to coordinates on the XY plane inside the vehicle 12, and shifts the estimated direction of a sound source outside the vehicle 12 to coordinates on the XY plane outside the vehicle 12, and supplies these to the BF module 74 as embedding vectors uniquely determined for each direction. Note that the coordinates on the XY plane inside the vehicle 12 and the coordinates on the XY plane outside the vehicle 12 are different from each other.

[0090] The STFT processing unit 73 performs STFT processing on the multi-channel sound signal input to the machine learning model 61 , and supplies the multi-channel sound signal that has undergone STFT processing to the BF module 74 .

[0091] The BF module 74 uses the embedding vector supplied from the DoA module 72 as auxiliary information, performs signal processing to emphasize the sound signal in the direction indicated by the auxiliary information from the multi-channel sound signal that has been subjected to STFT processing supplied from the STFT processing unit 73, and supplies the sound signal in that direction to the SED module 75. Note that, as a method of inputting the auxiliary information to the BF module 74, for example, FiLM (Feature-wise Linear Modulation), which is visual inference using a general conditioning layer, can be applied.

[0092] The SED module 75 estimates the type of acoustic event using a statistical method based on the sound signal supplied from the BF module 74 (i.e., the sound signal from the direction of the sound source that has been subjected to signal processing to enhance the sound signal). The SED module 75 then outputs the type of acoustic event inside the vehicle 12 and the type of acoustic event outside the vehicle 12.

[0093] The machine learning model 61 configured in this manner can estimate the type of acoustic event and the direction of the sound source using a multi-channel sound signal as input.

[0094] The machine learning model 61 can perform learning to simultaneously optimize the DoA module 72, the BF module 74, and the SED module 75. Alternatively, the machine learning model 61 can perform learning to optimize the remaining modules while leaving some of these modules, such as the acoustic sound source estimation unit 42, fixed as an existing model.

[0095] Furthermore, among the components that make up the machine learning model 61, components can be separated for specific functions and used as modules that provide the required functions.

[0096] For example, among the components constituting the machine learning model 61, the CSP processing unit 71 and the DoA module 72 can be separated and used as a vehicle inside / outside determination / direction estimation module 81 that determines whether a sound source is inside or outside the vehicle 12 and estimates the direction of the sound source. For example, only the vehicle inside / outside determination / direction estimation module 81 can be used in an application that displays the direction of a sound source around the vehicle 12 and alerts the driver.

[0097] Furthermore, among the components constituting the machine learning model 61, the CSP processing unit 71, the DoA module 72, the STFT processing unit 73, and the BF module 74 can be separated and used as a sound source direction emphasis module 82 that emphasizes sounds coming from the direction where the sound source is estimated to be. For example, based on the output of the sound source direction emphasis module 82, a sound source outside the vehicle 12 can be emphasized, and used for environmental monitoring, the approach of an emergency vehicle, etc., to be played back through a speaker inside the vehicle 12.

[0098] Furthermore, all of the components constituting the machine learning model 61 can be used as an interior / exterior acoustic event estimation module 83 that estimates acoustic events of sound sources inside the vehicle 12 and acoustic events of sound sources outside the vehicle 12. For example, the interior / exterior acoustic event estimation module 83 can be used to detect the siren sound of an emergency vehicle approaching from behind and to control the vehicle so that the vehicle 12 gives way.

[0099] By using it as a module in this way, it is possible to select and discard the necessary functions, which can contribute to, for example, reducing the size of the machine learning model 61 and realizing real-time processing.

[0100] Here, a method for collecting learning data in the machine learning model 61 will be described.

[0101] For example, the learning data for the machine learning model 61 can be obtained by collecting data in a real environment.

[0102] That is, in order to collect learning data in a real environment, a microphone array for determining whether a sound source is inside or outside the vehicle 12 is installed on the vehicle 12 while it is actually running, and a device required for creating ground truth data for learning is also installed. For example, a dedicated microphone array capable of detecting the direction of sound outside the vehicle 12 is installed on the outside of the vehicle 12, and an omnidirectional camera is used to capture images of the position of the target sound source.

[0103] Then, paired data of the sound source and the correct direction and distance are created. For example, among the sound signals collected by the microphone array, the sound signal of the microphone array in the section of the sound source to be learned is used as input data. Furthermore, correct data is created from information such as the direction of the target sound source captured by the omnidirectional camera and the sound signal input from the microphone array installed outside the vehicle 12, and is linked to the input data. Furthermore, by using a camera or other distance measuring sensor, the distance to the target sound source can be accurately measured, and the sound source distance can also be learned from the reverberation and sound pressure information of the sound source. Furthermore, by labeling the type of sound event (e.g., siren, railroad crossing sound, voice, music, wind, animal, etc.) for each target sound source, learning can be performed that can estimate the type of target sound source.

[0104] Alternatively, the learning data in the machine learning model 61 can be realized by collecting data in a simulation.

[0105] That is, by convolving impulse responses that can reproduce the phase difference between multiple microphones, it is possible to generate synthetic data by simulating the arrival of sounds from various sound sources. Also, training data may be generated from the sound field characteristics near the microphone array while virtually running the vehicle 12 in a virtual space and reproducing the surrounding sound field.

[0106] Although the basic principle of learning using a DNN is to create paired data of input and correct answer, it is also possible to apply, for example, semi-supervised or self-supervised methods in which the model is allowed to acquire the correct answer naturally without providing a correct answer. Also, a method can be used in which learning is performed so that the output from the input data has a distribution that can be classified into multiple clusters, and then a method is used in which it is determined in a later stage under which state each class is detected (for example, a sound source inside the vehicle 12, a sound source outside the vehicle 12).

[0107] Furthermore, when collecting learning data using the vehicle 12, for example, environmental sounds inside the vehicle 12, environmental sounds outside the vehicle 12, and noises such as rain and wind are collected as learning data. Furthermore, learning data may be collected for each type of vehicle 12, the placement of the microphone 21, and the material of the vehicle 12, or with the windows of the vehicle 12 open / closed.

[0108] When labeling the type of sound event, for example, the name of the sound event, the direction and distance of the sound source, whether it is inside or outside the vehicle 12, etc. can be labeled.

[0109] Furthermore, when generating the input feature, the CSP value may be calculated after adjusting the gain of the input sound signal, removing noise from the input sound signal, or applying a Fourier transform to the input sound signal.

[0110] In addition, when learning in the machine learning model 61, optimization learning can be performed to learn resistance to noise, learn differences between vehicles 12, learn the weather, and learn whether the windows of the vehicle 12 are open / closed.

[0111] In addition, applications can include monitoring environmental sounds, operating agents, notifying emergency vehicles, detecting objects, monitoring the interior of the vehicle 12, and controlling automatic driving.

[0112] Then, a wide variety of input signals that can occur in the real environment are collected so that the machine learning model 61 can operate under various conditions. For example, by preparing multiple patterns for the placement of the microphone 21, the size of the vehicle 12, etc., robustness can be improved.

[0113] Furthermore, a wide range of applications can be made using the output results of the machine learning model 61, for example, it is possible to notify surrounding information through dialogue with an agent, or to determine the operation of the agent itself based on the results of monitoring environmental sounds inside the vehicle 12. Other applications include emergency vehicle detection, automatic driving control in response to the detection of an emergency vehicle, and notifications for safe driving support.

[0114] The environmental sound monitoring system 11 configured as described above uses a plurality of microphones 21 distributed inside the vehicle 12 (for example, at the four corners of the vehicle 12) and can distinguish whether a sound source is inside or outside the vehicle 12 based on the phase difference between a plurality of sound signals input from these microphones 21. For example, in order to avoid spatial aliasing, the environmental sound monitoring system 11 can use input features for sound source direction estimation indicating the phase difference between a plurality of sound signals as input for signal processing using the CSP method, and construct a machine learning model 61 that takes into account the transfer characteristics passing through the body of the vehicle 12.

[0115] As a result, the environmental sound monitoring system 11 can input the input feature amounts for sound source direction estimation and the input feature amounts for sound event estimation into the DNN and learn a model that can simultaneously estimate the direction of a sound source and the type of sound event. Therefore, even when, for example, a sound source inside the vehicle 12 and a sound source outside the vehicle 12 occur simultaneously, the environmental sound monitoring system 11 can appropriately estimate the direction of the sound source and the type of sound event.

[0116] Furthermore, since the environmental sound monitoring system 11 does not require a microphone 21 to be installed outside the vehicle 12, costs can be reduced and design can be improved, and placement constraints can be eliminated by distributing individual microphones 21 throughout the vehicle 12.

[0117] <Example of Computer Configuration> Next, the above-described series of processes (information processing method) can be performed by hardware or software. When the series of processes is performed by software, a program constituting the software is installed in a general-purpose computer or the like.

[0118] FIG. 12 is a block diagram showing an example of the configuration of an embodiment of a computer in which a program for executing the above-described series of processes is installed.

[0119] In the computer, a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, and an EEPROM (Electronically Erasable and Programmable Read Only Memory) 104 are interconnected by a bus 105. An input / output interface 106 is further connected to the bus 105, and the input / output interface 106 is connected to the outside.

[0120] In a computer configured as described above, the CPU 101 performs the above-described series of processes by loading programs stored in, for example, the ROM 102 and EEPROM 104 into the RAM 103 via the bus 105 and executing the programs. In addition, the programs executed by the computer (CPU 101) can be written in advance in the ROM 102, or can be installed or updated in the EEPROM 104 from outside via the input / output interface 106.

[0121] In this specification, the processing performed by a computer according to a program does not necessarily have to be performed in chronological order according to the order described in the flowchart. In other words, the processing performed by a computer according to a program also includes processing that is executed in parallel or individually (for example, parallel processing or object-based processing).

[0122] The program may be processed by a single computer (processor), or may be distributed among multiple computers. Furthermore, the program may be transferred to and executed on a remote computer.

[0123] Furthermore, in this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0124] Also, for example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0125] Furthermore, for example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0126] Furthermore, for example, the above-described program can be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.

[0127] Also, for example, each step described in the above flowchart can be executed by one device or can be shared and executed by multiple devices. Furthermore, if one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.

[0128] In addition, the processing of the steps of a program executed by a computer may be executed in chronological order according to the order described in this specification, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the processing of each step may be executed in an order different from the order described above. Furthermore, the processing of the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0129] It should be noted that the present technologies described in this specification can be implemented independently and singly, unless a contradiction arises. Of course, any two or more of the present technologies can also be implemented in combination. For example, part or all of the present technologies described in any embodiment can be implemented in combination with part or all of the present technologies described in other embodiments. Furthermore, part or all of any of the present technologies described above can also be implemented in combination with other technologies not described above.

[0130] <Examples of Combinations of Configurations> The present technology can also be configured as follows. (1) An information processing system including: a generation unit that generates, based on sound signals input by a plurality of microphones distributed inside a vehicle, first input features indicating a phase difference between the plurality of sound signals and second input features indicating acoustic features, and an estimation unit that estimates the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features. (2) The information processing system described in (1) above, in which the estimation unit inputs the first input features and the second input features to a DNN (Deep Neural Network) to collectively estimate the direction of the sound source and the type of the acoustic event. (3) The information processing system described in (2) above, in which the generation unit generates the first input features by estimating the phase difference between the sound signals input from the plurality of microphones by the signal processing using a CSP (Cross-Power Spectrum Phase) method. (4) The information processing system according to (3) above, wherein the generation unit generates the second input feature by converting sound signals input from the plurality of microphones into a frequency spectrum or a mel spectrogram. (5) The information processing system according to (4) above, wherein at least three of the microphones are arranged at any of four corners inside the vehicle, and the estimation unit distinguishes whether the sound source is inside or outside the vehicle. (6) The information processing system according to (4) or (5) above, wherein the plurality of microphones are arranged at intervals such that a plane wave assumption holds for sounds arriving from sound sources outside the vehicle and a spherical wave assumption holds for sounds arriving from sound sources inside the vehicle. (7) The information processing system according to any of (4) to (6) above, wherein the plurality of microphones are arranged so that combinations of time differences between sounds arriving from sound sources outside the vehicle and sounds arriving from sound sources inside the vehicle have different characteristics.(8) The information processing system according to any one of (4) to (7), further comprising: a control unit that performs control based on the direction of the sound source and the type of the acoustic event when the situation requires control based on the estimation result supplied from the estimation unit. (9) The information processing system according to any one of (4) to (8), further comprising: an agent processing unit that functions as an agent for the vehicle according to the estimation result supplied from the estimation unit. (10) The information processing system according to any one of (4) to (9), wherein the generation unit performs acoustic processing to cancel acoustic echoes inside the vehicle by feeding back an audio signal supplied from an audio device that reproduces sounds output inside the vehicle. (11) The generation unit and the estimation unit are configured by a machine learning model including: a CSP processing unit that acquires phase difference information indicating a phase difference between the plurality of sound signals; a DoA (Direction of Arrival) module that uses the phase difference information to output likelihoods for each class associated with the type of the sound event inside and outside the vehicle, an estimated direction that estimates the direction of the sound source inside the vehicle, and an estimated direction that estimates the direction of the sound source outside the vehicle; an STFT processing unit that performs short-time Fourier transform (STFT) processing on the plurality of sound signals; a BF (Beamforming) module that uses an embedding vector supplied from the DoA module as auxiliary information and performs signal processing to emphasize the sound signal in the direction indicated by the auxiliary information from the plurality of sound signals that have been subjected to STFT processing supplied from the STFT processing unit; and a SED (Sound Event Detection) module that estimates the type of the sound event using a statistical method based on the sound signal that has been subjected to signal processing to emphasize in the BF module. An information processing system according to any one of (4) to (10) above.(12) The information processing system according to (11) above, wherein the CSP processing unit and the DoA module are usable as a vehicle inside / outside determination / direction estimation module that determines whether the sound source is inside or outside the vehicle and estimates the direction of the sound source. (13) The information processing system according to any of (11) or (12) above, wherein the CSP processing unit, the DoA module, the STFT processing unit, and the BF module are usable as a sound source direction enhancement module that enhances sound arriving from the direction estimated to be the sound source. (14) The information processing system according to any of (11) to (13) above, wherein the CSP processing unit, the DoA module, the STFT processing unit, the BF module, and the SED module are usable as an inside / outside vehicle acoustic event estimation module that estimates acoustic events of the sound source inside the vehicle and acoustic events of the sound source outside the vehicle. (15) An information processing method, comprising: an information processing system, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sounds arriving from a predetermined sound source, generating first input features indicating a phase difference between the plurality of sound signals and second input features indicating acoustic features, and estimating the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features. (16) A program, which causes a computer of an information processing system to execute information processing, comprising: generating first input features indicating a phase difference between the plurality of sound signals and second input features indicating acoustic features, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sounds arriving from a predetermined sound source, and estimating the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features.

[0131] It should be noted that the present embodiment is not limited to the above-described embodiment, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, the effects described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.

[0132] REFERENCE SIGNS LIST 11 Environmental sound monitoring system, 12 Vehicle, 21 Microphone, 22 Control device, 23 Audio equipment, 24 Agent processing unit, 31 Speaker, 32 Display, 33 Vehicle control device, 34 Peripheral equipment, 41 Input feature generation unit, 42 Acoustic sound source estimation unit, 43 Control unit, 51 Acoustic event estimation feature generation unit, 52 Sound source direction estimation feature generation unit, 61 Machine learning model, 71 CSP processing unit, 72 DoA module, 73 STFT processing unit, 74 BF module, 75 SED module, 81 Vehicle interior / exterior determination / direction estimation module, 82 Sound source direction enhancement module, 83 Vehicle interior / exterior acoustic event estimation module

Claims

1. An information processing system comprising: a generation unit that generates, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sounds arriving from a predetermined sound source, first input features that indicate a phase difference between the plurality of sound signals and second input features that indicate acoustic features; and an estimation unit that estimates the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features.

2. The information processing system according to claim 1, wherein the estimation unit inputs the first input feature amount and the second input feature amount into a DNN (Deep Neural Network) and simultaneously estimates the direction of the sound source and the type of the acoustic event.

3. The information processing system according to claim 2, wherein the generation unit generates the first input feature by estimating the phase difference between the sound signals input from the plurality of microphones through the signal processing using a CSP (Cross-Power Spectrum Phase) method.

4. The information processing system according to claim 3, wherein the generation unit generates the second input feature by converting sound signals input from the plurality of microphones into a frequency spectrum or a mel spectrogram.

5. The information processing system according to claim 4, wherein at least three of the microphones are arranged at any of four corners inside the vehicle, and the estimation unit distinguishes whether the sound source is inside or outside the vehicle.

6. The information processing system according to claim 4, wherein the plurality of microphones are arranged at intervals such that the plane wave assumption holds for sounds coming from sound sources outside the vehicle, and the spherical wave assumption holds for sounds coming from sound sources inside the vehicle.

7. The information processing system according to claim 4, wherein the plurality of microphones are arranged so that the combination of the time difference between the sound coming from a sound source outside the vehicle and the sound coming from a sound source inside the vehicle has different characteristics.

8. The information processing system according to claim 4, further comprising a control unit that performs control based on the direction of the sound source and the type of the acoustic event when the situation requires such control in accordance with the estimation result supplied from the estimation unit.

9. The information processing system according to claim 4, further comprising an agent processing unit that functions as an agent for said vehicle in accordance with the estimation results supplied from said estimation unit.

10. The information processing system according to claim 4, wherein the generation unit performs acoustic processing to cancel acoustic echoes inside the vehicle by feeding back an audio signal supplied from an audio device that reproduces sound output inside the vehicle.

11. The information processing system according to claim 4, wherein the generation unit and the estimation unit are configured using a machine learning model including: a CSP processing unit that acquires phase difference information indicating a phase difference between the plurality of sound signals; a DoA (Direction of Arrival) module that uses the phase difference information to output a likelihood for each class associated with the type of sound event inside and outside the vehicle, an estimated direction that estimates the direction of the sound source inside the vehicle, and an estimated direction that estimates the direction of the sound source outside the vehicle; a STFT processing unit that performs short-time Fourier transform (STFT) processing on the plurality of sound signals; a BF (Beamforming) module that uses an embedding vector supplied from the DoA module as auxiliary information and performs signal processing to emphasize the sound signal in the direction indicated by the auxiliary information from the plurality of sound signals that have been subjected to STFT processing supplied from the STFT processing unit; and a SED (Sound Event Detection) module that uses a statistical method to estimate the type of the sound event based on the sound signals that have been subjected to signal processing to emphasize in the BF module.

12. The information processing system according to claim 11, wherein the CSP processing unit and the DoA module can be used as an inside / outside vehicle determination / direction estimation module that determines whether the sound source is inside or outside the vehicle and estimates the direction of the sound source.

13. The information processing system according to claim 11, wherein the CSP processing unit, the DoA module, the STFT processing unit, and the BF module can be used as a sound source direction enhancement module that enhances sounds coming from the direction in which the sound source is estimated to be located.

14. The information processing system according to claim 11, wherein the CSP processing unit, the DoA module, the STFT processing unit, the BF module, and the SED module can be used as an interior / exterior acoustic event estimation module that estimates acoustic events of the sound source inside the vehicle and acoustic events of the sound source outside the vehicle.

15. An information processing method comprising: an information processing system generating, based on sound signals input by a plurality of microphones distributed inside a vehicle, sound coming from a predetermined sound source, first input features indicating a phase difference between the plurality of sound signals and second input features indicating acoustic features; and estimating the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features.

16. A program for causing a computer of an information processing system to execute information processing including: generating first input features indicating the phase difference between a plurality of sound signals and second input features indicating acoustic features, based on sound signals input by a plurality of microphones distributed inside a vehicle that pick up sounds arriving from a predetermined sound source; and estimating the direction of the sound source and the type of acoustic event by signal processing using the first input features and the second input features.

Citation Information

Patent Citations

  • Sound source direction estimation device and sound source direction estimation method

    JP2012042465A

  • Acoustic control method and acoustic control device

    WO2023204076A1