Pseudo-ambisonic signal generator, pseudo-ambisonic signal generation method, acoustic event presentation system, and program

The pseudo-ambisonic signal generation device and system address the challenge of arranging microphones on a fixed spherical surface by generating pseudo-ambisonic signals from human head-mounted microphones, enabling accurate sound source localization and detection with acoustic or visual presentation.

JP7835294B2Active Publication Date: 2026-03-25NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing wearable sound localization and detection devices face challenges in accurately determining the direction and type of sound sources due to the difficulty in arranging microphones on a fixed spherical surface, which is necessary for converting acoustic signals into ambisonic signals, and deriving pseudo-acoustic intensity vectors.

Method used

A pseudo-ambisonic signal generation device and system that includes a spherical coordinate acquisition unit, calculation unit, and signal extraction unit to generate pseudo-ambisonic signals using averaged spherical coordinates from microphones positioned on the human head, followed by an estimation device to determine sound source direction and type, and a presentation device to provide information to the user.

Benefits of technology

Enables accurate determination of sound source direction and type using wearable devices, allowing for effective sound localization and detection, and presentation through acoustic or visual means.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835294000013
    Figure 0007835294000013
  • Figure 0007835294000014
    Figure 0007835294000014
  • Figure 0007835294000015
    Figure 0007835294000015
Patent Text Reader

Abstract

In the present invention, a pseudo sound intensity vector is obtained by using sound signals collected by a wearable device. Accordingly, a pseudo ambisonics signal generation apparatus according to the present disclosed technology is provided with a spherical coordinates acquisition unit, a calculation unit, and a signal extraction unit. The spherical coordinates acquisition unit acquires respective spherical coordinates of microphones, by setting, as the origin, the intersection between a straight line passing the center of the left and right ears and a plane dividing the face symmetrically into left and right sides. The calculation unit calculates an average value of the radii of the spherical coordinates and replaces the respective radii of the spherical coordinates with the average value. The signal extraction unit generates a pseudo ambisonics signal by using the spherical coordinates replaced with the average value, and sound signals acquired by the microphones.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to the recording, analysis, and use of 3D acoustic information. [Background technology]

[0002] Being able to detect the type and direction of arrival of an acoustic event from an acoustic signal opens up a wide range of applications. For example, by linking the detection device with smart home devices, abnormal situations within the home can be promptly notified to the user, along with estimated event details and location information. Alternatively, by installing detection devices in self-driving cars, drivers can be notified of the occurrence of danger and the necessary actions. Alternatively, by having pedestrians carry the detection device as a wearable device, it is possible to inform pedestrians of the occurrence of danger and its precise direction.

[0003] This type of technology is called SELD (Sound Event Localization and Detection). For measuring three-dimensional sound fields, SELD primarily uses microphones called First Order Ambisonic (FOA) microphones. Figure 1 schematically shows an FOA microphone. An FOA microphone is a microphone array in which unidirectional microphones M1 to M4 are arranged at the four vertices of a regular tetrahedron.

[0004] Based on Non-Patent Document 1, we will provide an overview of the spherical harmonic expansion of acoustic signals and beamforming using ambisonic signals. The sound pressure signal p of wavenumber k observed in spherical coordinates (r,Ω) is expressed by the spherical harmonic function Y lm It can be expanded as follows using this:

number

Number

[0005] The obtained p lm is an orthogonal basis, so by weighted synthesis of these, a beamformer with an arbitrary beam pattern can be constructed. Generally, the beamformer output y can be expressed as follows.

Number

Number

Number

[0006] However, in order to obtain the signal arrival direction using Equation (7), it is necessary to calculate the signal intensity in all directions, which is not easy. Therefore, Non-Patent Document 1 proposes a method of estimating the direction of a sound source by approximately deriving a physical quantity called an acoustic intensity vector, which represents the sound propagation direction and intensity, from an ambisonic signal, taking the case of first-order ambisonics as an example. The acoustic intensity vector I is defined by the following equation using the sound pressure p and the particle velocity vector v.

Equation

Equation

Equation

Equation

Equation

[0008] [Non-Patent Document 1] DP Jarrett et al., "3D SOURCE LOCALIZATION IN THE SPHERICAL HARMONIC DOMAIN USING PSEUDOINTENSITY VECTOR," 18th European Signal Processing Conference (EUSIPCO 2010) Proceedings, pp.442-446 [Overview of the Initiative] [Problems that the invention aims to solve]

[0009] A FOA microphone, which has a total of four microphones placed at the vertices of a regular tetrahedron, is not practical for someone like a pedestrian to carry around on a daily basis, and requires some ingenuity. Making microphones wearable makes them easier for humans to carry, but it becomes difficult to arrange the microphones on the same spherical surface. In the case of a microphone array arranged on a sphere of radius R, the spherical coordinates of each microphone (R, φ) are calculated with the center of the sphere as the origin. q , θ qWhile it is possible to calculate the ambisonic signal using the formula directly, when multiple microphones are placed on the head, the spherical surface that passes through all microphone positions is generally not fixed. If the microphones are not positioned on the same spherical surface, the captured acoustic signals cannot be converted into ambisonic signals. Deriving the pseudo-acoustic intensity vectors used as input features for SELD requires signals in ambisonic format. The challenge is to be able to determine a pseudo-acoustic intensity vector using acoustic signals collected by a device attached to a person (wearable device). [Means for solving the problem]

[0010] To solve the above problems, the pseudo-ambisonic signal generation device related to the disclosed technology includes a spherical coordinate acquisition unit, a calculation unit, and a signal extraction unit. The spherical coordinate acquisition unit acquires the spherical coordinates of each microphone, using the intersection of a plane that divides the face symmetrically into left and right halves and a line passing through the centers of the left and right ears as the origin. The calculation unit calculates the average value of the radii of the spherical coordinates and replaces the radius of each spherical coordinate with the average value. The signal extraction unit generates a pseudo-ambisonic signal using spherical coordinates replaced by average values ​​and acoustic signals acquired by the microphone. Furthermore, the acoustic event presentation system relating to the disclosed technology includes at least four microphones positioned along the head of a human body, a pseudo-ambisonic signal generator, an estimation device, and a presentation device. The pseudo-ambisonic signal generator generates a pseudo-ambisonic signal from an acoustic signal acquired by a microphone. The estimation device estimates the direction and type of sound source from a pseudo-ambisonic signal. The presentation device presents the user with information about the sound source based on the estimation results. [Effects of the Invention]

[0011] According to the disclosed technology, it becomes possible to determine a pseudo-acoustic intensity vector using acoustic signals collected by a device attached to a person (wearable device), thereby realizing a wearable pseudo-ambisonic signal generator and an acoustic event presentation system. [Brief explanation of the drawing]

[0012] [Figure 1] A diagram illustrating SELD using conventional technology. [Figure 2] A functional block diagram of an acoustic event presentation system, including a pseudo-ambisonic signal generator according to the first embodiment. [Figure 3] This diagram shows an example of spherical coordinates to be set for the human head. [Figure 4] A flowchart illustrating the operation of a pseudo-ambisonic signal generator. [Figure 5] A flowchart illustrating the operation of the estimation device. [Figure 6] Functional block diagram of an acoustic presentation device. [Figure 7] A flowchart illustrating the operation of an acoustic presentation device. [Figure 8] Functional block diagram of a video display device. [Figure 9] A flowchart illustrating the operation of a video display device. [Figure 10] A diagram illustrating an example of a computer's functional configuration. [Modes for carrying out the invention]

[0013] The embodiments of the disclosed technology will be described in detail below. Components with the same function will be numbered identically, and redundant explanations will be omitted.

[0014] [First Embodiment] Figure 2 shows a functional block diagram of an example of an acoustic event presentation system, including a pseudo-ambisonic signal generator related to the disclosed technology. The acoustic event presentation system includes an acoustic information acquisition device 201, a pseudo-ambisonic signal generation device 202, an estimation device 206, and a presentation device 209.

[0015] <Acoustic information acquisition device> The acoustic information acquisition device 201 obtains a Q-channel acoustic signal x from Q microphones installed on the head or at any position on a device worn on the head. q The signal is obtained and supplied to the pseudo-ambisonic signal generator 202. Note that Q is an integer greater than or equal to 4.

[0016] <Pseudo-Ambisonics Signal Generator> The pseudo-ambisonics signal generator 202 includes a microphone coordinate acquisition unit 203, a calculation unit 204, and a signal extraction unit 205. Figure 3 shows an example of a spherical coordinate system for calculating microphone coordinates. Note that the setting of the x, y, and z axes passing through the origin in the following spherical coordinate system configuration is merely illustrative and not limiting. Let the line passing through the centers of the left and right ears be the y-axis. The origin of the spherical coordinate system is the intersection of the plane dividing the face symmetrically into left and right halves and the y-axis. Let the line passing through the origin and perpendicular to the y-axis in the vertical direction of the head be the z-axis of the spherical coordinate system. Let the line passing through the origin and perpendicular to the y-axis in the front-to-back direction of the head be the x-axis of the spherical coordinate system. Also, let the azimuth angle of the spherical coordinate system be φ and the elevation angle be θ.

[0017] Figure 4 is a flowchart illustrating the operation of the pseudo-ambisonic signal generator. The microphone coordinate acquisition unit 203 acquires the spherical coordinates p of each microphone based on the coordinate system shown in Figure 3. q =(r q , φ q , θ q Obtain (q=1,2,···,Q) (step S401). q The pseudo-ambisonic signal generator 202 may acquire values ​​measured by an external device, or it may read settings stored in the pseudo-ambisonic signal generator 202. The calculation unit 204 corrects the spherical coordinates obtained by the microphone coordinate acquisition unit. In the case of FOA microphones (or more generally, microphone arrays placed on a sphere of radius R), the spherical coordinates of each microphone (R, φ) are calculated with the center of the sphere as the origin. q , θ q ) can be used directly to calculate the ambisonic signal, but in the case of microphones placed on the head, the distance between the origin defined above and each microphone is generally not equal, and the microphone coordinates cannot be used directly to calculate the ambisonic signal. Therefore, in the first embodiment, the average value r of the distance between each microphone and the origin is calculated (step S402), p q Each r q p' replaced with r q =(r, φ q , θ q Let ) be the approximate spherical coordinates of each microphone (step S403). Next, the pseudo-ambisonics signal generator 202 receives the Q-channel acoustic signal x from the acoustic information acquisition device 201. q Obtain (step S404), and p' of group Q q and x q This method generates a pseudo-ambisonic signal. In other words, when a Q-channel microphone is placed on a rigid sphere of radius r, signal processing (such as spherical harmonic expansion) is performed to obtain the ambisonic signal and generate a pseudo-ambisonic signal.

[0018] <Estimation device> The estimation device 206 includes a pseudo-acoustic intensity vector extraction unit 207 and an estimation unit 208, and takes a pseudo-ambisonic signal as input and outputs an estimation result of the direction and type of sound source. Figure 5 is a flowchart illustrating the operation of the estimation device 206. The pseudo-acoustic intensity vector extraction unit 207 generates a pseudo-acoustic intensity vector from the pseudo-ambisonics signal using, for example, the method described in Non-Patent Document 1 (step S501). The estimation unit 208 estimates the direction of arrival of the sound source (step S502) and the type of sound source (step S503) using a pseudo-acoustic intensity vector and a pseudo-ambisonic signal. For estimation, for example, a DNN (Deep Neural Network) similar to the one described in "A. Politis et. al, “A dataset of dynamic reververant sound scenes with directional interferers for sound event localization and detection”, arXiv:2106.06999, 2021" (Reference 1) can be used, trained using the acoustic features extracted by the present invention as input. The DNN can be configured to take a pseudo-sound intensity vector and a pseudo-ambisonic signal as input, and output, for example, a 3D unit vector for the sound source direction and an integer corresponding to a label such as "bell sound" or "car driving sound" for the sound source type as estimation results.

[0019] <Presentation device> The presentation device 209 converts the estimation results into acoustic or visual information and provides it to the user.

[0020] <First presentation example> In the first presentation example, the estimation results are converted into stereophonic sound and presented to the user. Figure 6 shows a functional block diagram of the sound presentation device 601 related to the first presentation example. The sound presentation device 601 includes an HRTF search unit 602, an HRTF database 603, a voice / sound effect search unit 604, a voice / sound effect database 605, and a convolution calculation unit 606. HRTF stands for Head-related transfer function, a function that describes how sound travels from the sound source to both ears. In Japanese, it is called head-related transfer function. The HRTF database contains pre-registered HRTFs that cover all directions around the head, or HRTFs that cover all directions around the upper hemisphere, depending on the application of the acoustic event presentation system. The audio and sound effect database stores audio and sound effects corresponding to the estimated sound source types. The method for determining the correspondence between the estimated sound source type and the corresponding audio file is arbitrary; for example, an audio file containing a warning message such as "A car is approaching" can be used as the audio file corresponding to the sound source type "car".

[0021] Figure 7 is a flowchart illustrating the operation of the sound presentation device 601. The HRTF search unit 602 searches the HRTF database for the HRTF in the direction closest to the sound source direction obtained as an estimation result, and obtains the sound source direction HRTF (step S701). The audio / sound effect search unit 604 searches the audio / sound effect database for audio and sound effects corresponding to the sound source type obtained as an estimation result, and obtains an audio file corresponding to the sound source type (step S702). The convolution unit 606 convolves the sound source direction HRTF into the obtained sound source type-corresponding audio file. This generates sound that simulates the situation where the sound source type-corresponding audio file is played in the sound source direction. For example, it can present the user with three-dimensional sound where the voice saying "A car is approaching" sounds as if it were coming from the direction of the approaching car.

[0022] <Second presentation example> In the second presentation example, the estimation results are converted into a video and presented to the user. Figure 8 shows a functional block diagram of the video presentation device 801 related to the second presentation example. The video display device 801 includes a marker image acquisition unit 802, a marker image database 803, a marker image conversion unit 804, a camera image acquisition unit 805, and an estimation result synthesis unit 806. The marker image database 803 stores, for example, three-dimensional arrow images with shapes and colors corresponding to the type of sound source, as basic marker images.

[0023] Figure 9 is a flowchart illustrating the operation of the video display device 801. The marker image acquisition unit 802 acquires a basic marker image corresponding to the type of sound source from the marker image database 803 (step S901). The marker image conversion unit 804 rotates the basic marker image three-dimensionally using the estimated sound source direction to generate a modified marker image (step S902). For example, it rotates the marker image so that it appears to extend from the center of the head towards the sound source. The camera image acquisition unit 805 acquires images of the user's surroundings (step S903). The estimation result synthesis unit 806 adds the correction marker image to the image acquired by the camera image acquisition unit 805 (step S904). This allows the video display device 801 to visually present to the user the type of sound source and its direction of arrival.

[0024] Alternatively, the marker image database may be pre-registered with marker images for all sound source directions and types, and selected according to the sound source type and direction. Alternatively, a basic marker image may be generated according to the type of sound source, and the orientation of the marker image may be determined based on the direction of the sound source.

[0025] [Differentiation] In the first embodiment, the origin of the spherical coordinate system was set to the approximate center of the head (the intersection of a line passing through the centers of the left and right ears and a plane dividing the face symmetrically). However, if there are four head-mounted microphones, the sphere through which all the microphones pass can be calculated, and the center of that sphere can be set as the origin.

[0026] [Programs, recording media] The various processes described above can be carried out by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 10, and then causing the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc. to operate.

[0027] The program describing this process can be recorded on a computer-readable recording medium. Any computer-readable recording medium can be used, such as a magnetic recording device, optical disc, magneto-optical recording medium, or semiconductor memory.

[0028] Furthermore, this program may be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs or CD-ROMs on which the program is recorded. Alternatively, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.

[0029] A computer executing such a program may, for example, first store the program recorded on a portable storage medium or a program transferred from a server computer in its own storage device. Then, when processing is to be executed, the computer reads the program stored on its own storage medium and executes the processing according to the read program. Alternatively, the computer may directly read the program from the portable storage medium and execute the processing according to that program, or it may sequentially execute the processing according to the received program each time a program is transferred to it from a server computer. Furthermore, the above processing may be executed by a so-called ASP (Application Service Provider) type service, where the server computer does not transfer programs to this computer, but the processing function is realized only by execution instructions and result acquisition. In this form, the program includes information used for processing by an electronic computer that is equivalent to a program (data that is not a direct instruction to the computer but has the property of defining the processing of the computer, etc.).

[0030] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware.

Claims

1. A device for generating ambisonic signals from acoustic signals acquired by at least four microphones positioned along the head of a human body, A spherical coordinate acquisition unit that acquires the spherical coordinates of each microphone, with the origin being the intersection point of a plane that divides the face symmetrically into left and right halves and a line passing through the centers of the left and right ears, A calculation unit that calculates the average value of the radii of the aforementioned spherical coordinates and replaces the radius of each spherical coordinate with the average value, A signal extraction unit that generates a pseudo-ambisonic signal using the spherical coordinates replaced by the average value and the acoustic signal acquired by the microphone, A pseudo-ambisonic signal generator that includes [a specific component / feature].

2. A method for generating an ambisonic signal from acoustic signals acquired by at least four microphones positioned along the head of a human body, The coordinate acquisition unit performs the following steps: acquires the spherical coordinates of each microphone, with the origin being the intersection point of a plane that divides the face symmetrically into left and right halves and a line passing through the centers of the left and right ears; The calculation unit calculates the average value of the radii of the spherical coordinates and replaces the radius of each spherical coordinate with the average value. The signal extraction unit performs the steps of generating a pseudo-ambisonic signal using the spherical coordinates replaced by the average value and the acoustic signal acquired by the microphone, A pseudo-ambisonic signal generation method including [a specific component].

3. At least four microphones positioned along the head of the human body, A pseudo-ambisonic signal generator that generates a pseudo-ambisonic signal from an acoustic signal acquired by the aforementioned microphone, An estimation device for estimating the direction and type of sound source from the aforementioned pseudo-ambisonic signal, It consists of a presentation device that presents information about the sound source to the user based on the estimated direction and type of the sound source, The aforementioned pseudo-ambisonic signal generator is A spherical coordinate acquisition unit that acquires the spherical coordinates of the microphone, with the origin being the intersection point of a plane that divides the face symmetrically into left and right halves and a line passing through the centers of the left and right ears, A calculation unit that calculates the average value of the radii of the aforementioned spherical coordinates and replaces the radius of each spherical coordinate with the average value, The system includes spherical coordinates replaced by the average value, and a signal extraction unit that generates the pseudo-ambisonic signal using the acoustic signal acquired by the microphone. Audio event presentation system.

4. The acoustic event presentation system according to claim 3, The aforementioned display device presents the direction and type of sound source audibly or visually. Audio event presentation system.

5. A program for causing a computer to function as a pseudo-ambisonic signal generating device according to claim 1, or as an acoustic event presentation system according to either claim 3 or 4.