Spatial audio generation device, spatial audio generation method, and spatial audio generation program

WO2026167888A1PCT designated stage Publication Date: 2026-08-13MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2026-08-13

Smart Images

  • Figure JP2025017624_13082026_PF_FP_ABST
    Figure JP2025017624_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention comprises: an ambisonic impulse response unit (10) in which impulse responses (Impw, Impx, Impy, Impz) are stored in advance; an audio storage unit (11) in which a sound S that is an anechoic chamber recorded sound is stored in advance; an ambisonic sound generation unit (12) that, by convolving the impulse responses (Impw, Impx, Impy, Impz) acquired from the ambisonic impulse response unit (10) with the sound S acquired from the audio storage unit (11), generates ambisonic sounds (Sw, Sx, Sy, Sz), which are three-dimensional audio data imitating audio playback in a reproduction space; and an ambisonic sound playback unit (13) that plays the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Spatial audio generation device, spatial audio generation method, and spatial audio generation program

[0001] This disclosure relates to a spatial audio generation device, a spatial audio generation method, and a spatial audio generation program.

[0002] A method has been proposed that uses multiple speakers to provide location information to audio content, giving the listener a sense of sound localization. Among these, a playback method that allows users to experience sound localization (sense of direction and distance) through headphones is called spatial audio. In virtual space, it is required that the position of objects in the image and the sense of sound localization match, and that the playback method also accommodates the user's head movements. Furthermore, a playback method called "externalization of sound," where the sound from headphones is not localized inside the head but feels as if it is coming from outside the headphones, increases the sense of presence of the sound.

[0003] There are various methods for spatial audio technology that enable the externalization of sound. One of them is ambisonic sound, which allows for sound collection and reproduction in a 360° spherical sound field. Ambisonic is an example of a scene-based method that records and reproduces the physical information of the entire space surrounding the experiencer in a 360° spherical space, and has features such as high compatibility with VR (Virtual Reality).

[0004] In Ambisonic, for example, the original sound is recorded using an Ambisonic microphone with directivity directed in four directions of space. The recorded audio data is then converted into components in all directions (spherical omnidirectional), left / right (e.g., x-axis), vertical (e.g., z-axis), and front / back (e.g., y-axis), and the sound is generated considering the speaker placement and the posture of the listener.

[0005] However, Ambisonic had the problem of requiring a speaker array consisting of many speakers in the playback environment. Furthermore, improving the spatial resolution of sound in Ambisonic would require even more speakers.

[0006] Patent Document 1 discloses an invention of an audio processing device that combines ambisonic and binaural playback technology, enabling the experience of directional sound using headphones, while also reducing the processing load of convolution operations that occur during ambisonic and binaural playback technology. Binaural playback technology is a technique that presents sound synthesized with a head-related transfer function (HRTF) from a certain direction to the target sound using headphones or the like. The head-related transfer function expresses information about how sound travels from all directions surrounding the human head to the eardrums of both ears as a function of frequency and direction of arrival.

[0007] International Publication No. 2017 / 119321

[0008] However, the technology disclosed in Patent Document 1 has the problem that, in order to record ambisonic sound for the purpose of reproducing the sound of the target space, it is necessary to collect sound with an ambisonic microphone for each environment and sound source, which takes an enormous amount of time.

[0009] The purpose of this disclosure is to provide a spatial audio generation device, a spatial audio generation method, and a spatial audio generation program that enable the externalization of sound that conveys a sense of direction and distance through easy sound collection.

[0010] A spatial audio generation device according to one aspect of the present disclosure includes: an impulse response storage unit that stores the spherical harmonic function expanded impulse response of a microphone array in a reproducible space which is a space related to sound reproduction; a sound storage unit that stores an audio signal; and a playback sound data generation unit that outputs playback sound data of the audio signal generated by convolving the audio signal obtained from the sound storage unit with the impulse response obtained from the impulse response storage unit.

[0011] A spatial audio generation method according to one aspect of the present disclosure is a spatial audio generation method performed by a computer, comprising the steps of: storing the spherical harmonic function expanded impulse response of a microphone array in a reproduction space which is a space related to sound playback; storing an audio signal; and outputting playback sound data of the audio signal generated by convolving the stored audio signal with the impulse response.

[0012] A spatial audio generation program according to one aspect of the present disclosure causes a computer to perform the following steps: store the spherical harmonic function expanded impulse response of a microphone array in a reproduction space which is a space related to sound playback; store an audio signal; and output playback sound data of the audio signal generated by convolving the stored audio signal with the impulse response.

[0013] The apparatus of this disclosure provides a spatial audio generation device, a spatial audio generation method, and a spatial audio generation program that enable the externalization of sound with a sense of direction and distance through easy sound collection.

[0014] This is a block diagram showing an example of the configuration of the spatial audio generation device according to Embodiment 1. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 1. (A) is a block diagram showing an example of the spatial audio generation device according to Embodiment 1 when it is configured with hardware, and (B) is a block diagram showing an example of the spatial audio generation device according to Embodiment 1 when it is configured with a computer on which the spatial audio generation program runs. This is a block diagram showing an example of the configuration of the spatial audio generation device according to Embodiment 2. (A) is an explanatory diagram showing an example of sound collection from multiple sound sources in space using an ambisonic microphone, and (B) is an explanatory diagram showing an example of speaker arrangement when playing sound for playback with multiple speakers. This is an explanatory diagram showing an example of the calculation of a general formula for a parametric head-related transfer function. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 2. (A) is a block diagram showing an example of the spatial audio generation device according to Embodiment 2 when it is configured with hardware, and (B) is a block diagram showing an example of the spatial audio generation device according to Embodiment 2 when it is configured with a computer on which the spatial audio generation program runs. This is a block diagram showing an example of the configuration of a spatial audio generation device according to Embodiment 3. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 3. This is a block diagram showing an example of the configuration of a spatial audio generation device according to Embodiment 4. (A) is an explanatory diagram showing an example of collecting sound from multiple sound sources in space using an ambisonic microphone, and (B) is an explanatory diagram showing a case where the channel-based speaker arrangement is changed based on the angle information of the object position specification unit. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 4. This is a block diagram showing an example of the configuration of a spatial audio generation device according to Embodiment 5. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 5. This is a block diagram showing an example of the configuration of a spatial audio generation device according to Embodiment 6. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 6.This is a block diagram showing an example of the configuration of the spatial audio generation device according to Embodiment 7. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 7. This is a block diagram showing an example of the configuration of the spatial audio generation device according to Embodiment 8. This is an explanatory diagram showing an example of a configuration for estimating an adaptive filter. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 8. This is a flowchart showing an example of reverberation estimation processing. This is a block diagram showing an example of the configuration of the spatial audio generation device according to Embodiment 9. This is a flowchart showing an example of processing in the spatial audio generation device according to Embodiment 9. This is a flowchart showing an example of reverberation estimation processing.

[0015] The following describes a spatial audio generation device, spatial audio generation method, and spatial audio generation program according to an embodiment, with reference to the drawings. The following embodiments are merely examples, and it is possible to combine the embodiments as appropriate and modify each embodiment as appropriate.

[0016] <Embodiment 1> Hereinafter, Embodiment 1 of the present disclosure will be described with reference to Figure 1. The spatial audio generation device 100 according to Embodiment 1 uses the impulse response of an ambisonic microphone measured in advance in a real space corresponding to the space (reproduction space) related to sound reproduction. An impulse response is the output of a microphone array system including an ambisonic microphone when an impulse, which is a signal with a very short duration, is input to the system. For example, to obtain the impulse response, a TSP (Time Stretched Pulse) signal, which is a time-stretched pulse of sound collected three-dimensionally using an ambisonic microphone in the reproduction space, is reproduced by a speaker, and the sound reproduced from the speaker is collected by an ambisonic microphone. The reproduction and collection of the TSP signal is repeated, and the inverse signal of the TSP signal is convolved into the waveform obtained by averaging the collected waveforms to obtain the impulse response. Furthermore, because the directivity of the microphone array, including ambisonic microphones, is expanded using a so-called spherical harmonic function, a spherical harmonic function-expanded impulse response can be obtained.

[0017] As shown in Figure 1, the spatial audio generation device according to Embodiment 1 comprises an ambisonic impulse response unit (referred to as the "impulse response storage unit" in the claims) 10 which stores impulse responses recorded by an ambisonic microphone in a reproduced space or impulse responses obtained by simulation in advance; an audio storage unit (referred to as the "audio storage unit" in the claims) 11 which stores the audio signal to be heard as spatial audio; an ambisonic sound generation unit (referred to as the "playback sound data generation unit" in the claims) 12 which generates playback sound data by convolving the impulse response obtained from the ambisonic impulse response unit 10 with the sound obtained from the audio storage unit 11; and an ambisonic sound playback unit (referred to as the "audio playback unit" in the claims) 13 which plays back the generated ambisonic sound.

[0018] As shown in Figure 5(A) later, the Ambisonic Impulse Response Unit 10 stores the impulse response recorded by the Ambisonic Microphone 50 in the space where sound is reproduced, or the impulse response obtained from a simulation using known acoustic analysis. The stored impulse response is, as an example, an omnidirectional Imp w , Imp related to the left and right direction x , Imp related to the front-back direction y , Impact in the vertical direction z That is the case.

[0019] The audio storage unit 11 stores sounds to be heard as spatial audio, such as birdsong, waves, helicopter propeller sounds, and factory machinery sounds. It is desirable that the sounds stored in the audio storage unit 11 be anechoic chamber recordings, so-called dry sources, which are sounds recorded in a state without spatial acoustic characteristics.

[0020] The ambisonic sound generation unit 12 takes the sound S acquired from the sound storage unit 11 and the impulse response Imp acquired from the ambisonic impulse response unit 10. w , Imp x , Imp y , Imp zBy performing convolution, the sound S obtained from the voice storage unit 11 is generated as a sound recorded from a sound source in a reproduction space arranged in ambisonics. In other words, the ambisonic sound generation unit 12 generates three-dimensional voice data that simulates voice reproduction in the reproduction space.

[0021] The ambisonic sound reproduction unit 13 processes and reproduces the generated ambisonic sound (specifically, ambisonic B-format sound) S w , S x , S y , S z and reproduces it. The reproduction method can be anything, for example, a channel-based method using a plurality of speakers or a binaural reproduction method based on a head-related transfer function. The head-related transfer function represents the propagation of sound from the sound source to the eardrums (both ears) as a transfer function. By analyzing the head-related transfer function, the mechanism of sound image perception can be understood. By reproducing the head-related transfer function, it becomes possible to realize direction perception, but there are individual differences in the head-related transfer function, and it is difficult to define it uniformly.

[0022] The ambisonic sound generation (convolution) in the ambisonic sound generation unit 12 generates three-dimensional voice data that simulates ambisonic recorded sound by convolving the pre-measured ambisonic impulse response with the sound obtained from the voice storage unit 11 as shown in the following formula (1).

[0023]

[0024] The convolution of discrete signals f and g is defined as shown in the following formula (2).

[0025]

[0026] FIG. 2 is a flowchart showing an example of processing in the spatial audio generation device 100 according to Embodiment 1. In step S1, the ambisonic sound generation unit 12 acquires an ambisonic impulse response from the ambisonic impulse response unit 10.

[0027] In step S2, the ambisonic sound generation unit 12 acquires the sound S from the voice storage unit 11.

[0028] In step S3, the ambisonic sound generation unit 12 takes the sound S acquired from the sound storage unit 11 and the impulse response Imp acquired from the ambisonic impulse response unit 10. w , Imp x , Imp y , Imp z By folding it, ambisonic sound S w S x S y S z Generates.

[0029] In step S4, the ambisonic sound playback unit 13 generates the ambisonic sound S w S x S y S z The data is processed and played back, completing the series of processes from step S1 to step S4.

[0030] The spatial audio generation device 100 according to Embodiment 1 requires signal processing for a microphone and microphone input. The signal processing unit can be implemented using a hardware-based processing circuit 61, or a computer having a processor 65 such as a CPU (Central Processing Unit) on which the spatial audio generation program runs. The spatial audio generation device 100 may consist of not only one computer but also multiple computers.

[0031] Figure 3(A) is a block diagram showing an example of a spatial audio generation device 100 according to Embodiment 1 when it is configured in hardware. As shown in Figure 3(A), the spatial audio generation device 100 includes a processing circuit 61 that generates ambisonic sound and processes related to the playback of the generated ambisonic sound, a microphone 62 that acquires audio data, and multiple speakers or headphones 63 that play back the generated ambisonic sound, all connected to an external storage device 64.

[0032] The external storage device 64 is, for example, an HDD (hard disk drive) or SSD (solid state drive) connected directly to the spatial audio generation device 100 or via a network. The external storage device 64 contains sound S and impulse response Imp w, Imp x , Imp y , Imp z Data such as the above will be stored.

[0033] The hardware-based spatial audio generation device 100 receives the sound recorded from the microphone 62 and stores the impulse response Imp in the external storage device 64. w , Imp x , Imp y , Imp z The signal processing that convolves the sound S stored in the external storage device 64 is realized by the processing circuit 61. Alternatively, the impulse response Imp stored in the external storage device 64 is used with the sound S stored in the external storage device 64. w , Imp x , Imp y , Imp z You can fold it.

[0034] Figure 3(B) is a block diagram showing an example of a spatial audio generation device 100 according to Embodiment 1, configured with a computer on which the spatial audio generation program runs. As shown in Figure 3(A), the processor 65, and the impulse response Imp when the processor 65 is operating w , Imp x , Imp y , Imp z A memory 66 where data such as the above is stored, a microphone 62, and multiple speakers or headphones 63 are connected to an external storage device 64 via a bus 67. The processor 65 enables a spatial audio generation method that realizes the function of generating ambisonic sound and the function of playing back ambisonic sound by executing a spatial audio generation program. Furthermore, when the spatial audio generation program is running, the processor 65 operates as an ambisonic sound generation unit 12 and an ambisonic sound playback unit 13.

[0035] As described above, according to the spatial audio generation device 100 of Embodiment 1, the impulse response Imp recorded by the ambisonic microphone w , Imp x , Imp y , Imp z , or the impulse response Imp obtained by simulationw , Imp x , Imp y , Imp z By folding this into the sound S of the sound storage unit 11, it is possible to generate and play back ambisonic sound, which is sound as if the sound source were recorded in a space where an ambisonic microphone is placed.

[0036] The spatial audio generation device 100 according to Embodiment 1 does not require the collection of ambisonic sound with an ambisonic microphone for each environment and sound source, and the impulse response Imp w , Imp x , Imp y , Imp z Ambisonic sound can be generated by convolving it with sound recorded in an anechoic chamber. As a result, a spatial audio generation device, spatial audio generation method, and spatial audio generation program can be provided that enable the externalization of sound with a sense of direction and distance through easy sound collection. Furthermore, the impulse response Imp w , Imp x , Imp y , Imp z By using this method, it becomes possible to generate ambisonic sounds even from sound sources that cannot be physically captured in a real environment.

[0037] Embodiment 2 Embodiment 2 of the present disclosure will be described below with reference to Figure 4. The spatial audio generation device 200 according to Embodiment 2 differs from the spatial audio generation device 100 according to Embodiment 1 in that the ambisonic sound reproduction unit 13 is composed of a channel-based processing unit 14, a binaural sound generation unit 15, multiple speakers or headphones 16, and a head tracking adjustment unit 17. However, since the other configurations are the same as those of the spatial audio generation device 100 according to Embodiment 1, the same reference numerals are used for the same components as in Embodiment 1, and detailed descriptions are omitted.

[0038] The channel-based processing unit 14 processes the ambisonic sound S generated by the ambisonic sound generation unit 12. w S x S y S zA plurality of speaker reproduction sound signals S are generated for simulating voice reproduction in a state where a plurality of speakers are arranged in space. L , S R , S C , S SL , S SR Ambisonic sound S w , S x , S y , S z For example, as shown in FIG. 5(A), the ambisonic sound S is obtained based on the impulse responses Imp w , Imp x , Imp y , Imp z obtained from a plurality of sound sources 51 in space by an ambisonic microphone 50 and is obtained by convolution with the sound S acquired from the sound storage unit 11. The channel-based processing unit 14 generates the plurality of speaker reproduction sound signals S w , S x , S y , S z from the ambisonic sound S L , S R , S C , S SL , S SR Any method may be used, but for example, the sound S L , S R , S C , S SL , S SR output from the speakers 52 installed in space at the angles of the 5.1ch speaker arrangement defined by ITU-R (International Telecommunication Union Radiocommunication Sector) as shown in FIG. 5(B) is generated by utilizing beamforming that enhances the directivity of sound waves in a predetermined direction.

[0039] The binaural sound generation unit 15 generates binaural sound signals S L , S R , S C , S SL , S SR from the plurality of speaker reproduction sound signals S LE , S REThe binaural sound generation unit 15 generates a head-related transfer function (HTF). The HTF used in the binaural sound generation unit 15 includes a HTF measured in advance and a HTF calculated by numerical calculation. The latter is called a parametric HTF. Methods for calculating a parametric HTF include, for example, applying the boundary element method to a wave equation given the shape of the participant's head and torso as boundary conditions to calculate the HTF, and calculating the HTF using an infinite impulse response (IIR) filter.

[0040] Figure 6 is an explanatory diagram illustrating an example of the calculation of a general formula for a parametric head-related transfer function. The acoustic transfer function H is measured by a microphone placed at the entrance E of the external auditory canal from a sound source (e.g., speaker 52) located at a certain position S(r, θ, φ) of the head 53. E (S, ω) and the acoustic transfer function H measured from a sound source at position S with no head 53 present, using a microphone placed at the center of the head O. O The head-related transfer function is defined by (S, ω) as shown in equation (3) below.

[0041]

[0042] Parametric head-related transfer functions (HRPs) are difficult to optimize individually, but they often offer robust performance (there isn't a significant difference in quality between users). They also have the advantage of not requiring prior measurement of HRPs.

[0043] The head tracking adjustment unit 17 performs processing so that the sound heard from the headphones is synchronized with the direction of the user's head movement, similar to how it sounds in real space. Specifically, for example, the direction of the user's head is sensed using a rotation sensor attached to the headphones 16 worn by the user, and the sensed result (ω) is extracted. The rotation sensor may be an acceleration sensor, gyroscope, magnetic sensor, or inertial measurement unit (IMU).

[0044] The channel-based processing unit 14 displaces the sound image of the sound to be played by multiple speakers based on the direction of the head. Specifically, the channel-based processing unit 14 generates a rotation matrix based on ω, and controls the ambisonic sound Sw and S x and S y and S z is multiplied. As a result, by generating a channel-based sound obtained by rotating the speaker arrangement as a whole by ω, the binaural sound rotates in conjunction with the movement of the head of the experiencer.

[0045] FIG. 7 is a flowchart showing an example of processing in the spatial audio generation device 200 according to the second embodiment. Since the procedures from step S1 to step S3 in the processing shown in FIG. 7 are the same as those in the spatial audio generation device 100 according to the first embodiment, detailed description thereof is omitted.

[0046] In step S41, in the channel-based processing unit 14, the ambisonic sound S w and S x and S y and S z generated by the ambisonic sound generation unit 12 is used to generate a plurality of sound signals S L and S R and S C and S SL and S SR for reproduction by a plurality of speakers.

[0047] In step S42, in the binaural sound generation unit 15, based on the head-related transfer function, a plurality of sound signals S L and S R and S C and S SL and S SR are used to generate binaural sound signals S LE and S RE . The head-related transfer function used in step S42 may be a pre-measured head-related transfer function or a parametric head-related transfer function calculated by numerical calculation.

[0048] In step S43, the binaural sound signals S LE and S RE generated in step S42 are reproduced by the headphones etc.

[0049] In step S44, a rotation sensor provided in the headphones etc. detects the change in the position of the head of the experiencer.

[0050] In step S45, the ambisonic sound playback unit 13 determines whether or not the playback of binaural sound is to be stopped. If it is determined in step S45 that the playback of binaural sound is to be stopped, the series of processes from step S1 to step S44 is terminated. If it is determined in step S45 that the playback of binaural sound is not to be stopped, the procedure proceeds to step S46.

[0051] In step S46, the head tracking adjustment unit 17 senses the direction of the user's head using a rotation sensor attached to the headphones 16 or the like worn by the user, extracts the sensed result (ω), and proceeds to step S41.

[0052] In step S41, the channel-based processing unit 14 multiplies the ambisonic sound S by a rotation matrix based on ω. w S x S y S z Audio signal S for playback from multiple speakers L S R S C S SL S SR Generate the output and proceed to step S42.

[0053] Figure 8(A) is a block diagram showing an example of the spatial audio generation device 200 according to Embodiment 2 when it is configured as hardware, and Figure 8(B) is a block diagram showing an example of the spatial audio generation device 200 according to Embodiment 2 when it is configured as a computer on which the spatial audio generation program runs. The spatial audio generation device 200 according to Embodiment 2 differs from the spatial audio generation device 100 according to Embodiment 1 in that both when it is configured as hardware and when it is configured as a computer on which the spatial audio generation program runs, it has rotation sensors 68 such as an acceleration sensor, a gyro sensor, a magnetic sensor, and an inertial measuring device, but the other configurations are the same as the spatial audio generation device 100 according to Embodiment 1, so a detailed explanation is omitted.

[0054] As described above, the spatial audio generation device 200 according to Embodiment 2 uses the head-related transfer function calculated by numerical calculation to generate a binaural sound signal SLE S RE This generates a parametric head-related transfer function (HRP). The parametric HRP has robust performance that does not vary significantly in quality from one user to another, and it does not require prior measurement of the HRP.

[0055] Furthermore, the spatial audio generation device 200 according to Embodiment 2 generates ambisonic sound S based on, for example, the angle ω indicating the positional displacement of the user's head detected by the rotation sensor 68. w S x S y S z From binaural sound signal S LE S RE By generating this, it is possible to provide a spatial audio generation device, a spatial audio generation method, and a spatial audio generation program that can reproduce sound that takes into account the movement of the user's head.

[0056] <Embodiment 3> Hereinafter, Embodiment 3 of the present disclosure will be described with reference to Figure 9. The spatial audio generation device 300 according to Embodiment 3 differs from the spatial audio generation device 200 according to Embodiment 2 in that it has an object position designation unit 20. However, since the other configurations are the same as those of the spatial audio generation device 200 according to Embodiment 2, the same reference numerals as in Embodiment 2 will be used for the same configurations as in Embodiment 2, and a detailed explanation will be omitted.

[0057] The object position specification unit 20 specifies the object position, which is the position of a sound source in the reproduction space, by distance d and angle θ, and uses the impulse response corresponding to the specified distance d and angle θ. The ambisonic impulse response unit 21 according to Embodiment 2 stores the impulse responses measured in advance for each distance d and angle θ. In Embodiment 3, the method of reproducing sound in which the object position can be understood is called object audio.

[0058] The object positioning unit 20 specifies the object position, which is the position of the sound source, by, for example, visually sensing the sound source using a camera and estimating its direction and distance; estimating the direction of sound arrival based on the MUSIC method, which is a method for estimating the direction of sound arrival using an ambisonic microphone; or measuring the direction and distance of the object that is the sound source using LiDAR (Light Detection and Ranging).

[0059] The ambisonic sound generation unit 12 receives an impulse response Imp corresponding to the distance d and angle θ specified from the object position specification unit 20 from the ambisonic impulse response unit 21. w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z (d, θ) is obtained. The ambisonic sound generation unit 12, similar to the spatial audio generation device 100 according to Embodiment 1, obtains the impulse response Imp w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z By convolving (d, θ) into the sound S of the sound storage unit 11, an ambisonic sound is generated, which is a sound similar to that recorded in a space where an ambisonic microphone is placed.

[0060] Figure 10 is a flowchart showing an example of processing in the spatial audio generation device 300 according to Embodiment 3. The process shown in Figure 10 is the same as that of the spatial audio generation device 200 according to Embodiment 2 from step S2 to step S46, so a detailed explanation is omitted.

[0061] In step S51, the object position specification unit 20 specifies the distance d, which is the object position, and the angle θ.

[0062] In step S52, the ambisonic sound generation unit 12 receives an impulse response Imp from the ambisonic impulse response unit 21 that corresponds to the distance d and angle θ specified from the object positioning unit 20. w (d, θ), Imp x (d, θ), Impy (d, θ), Imp z Obtain (d, θ) and proceed to step S2. Thereafter, execute the steps from step S2 to step S46 in the same manner as in Embodiment 2.

[0063] As described above, the spatial audio generation device 300 according to Embodiment 3 generates an impulse response Imp corresponding to a specified distance d and angle θ. w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z Ambisonic sound S based on (d, θ) w S x S y S z From binaural sound signal S LE S RE By generating this, object-based audio can be realized that allows sound playback in a way that allows the listener to understand the location of the sound source.

[0064] Embodiment 4 Hereinafter, Embodiment 4 of the present disclosure will be described with reference to Figure 11. The spatial audio generation device 400 according to Embodiment 4 differs from the spatial audio generation device 300 according to Embodiment 3 in that the distance d and angle θ specified from the object position designation unit 20 are output not only to the ambisonic impulse response unit 21 but also to the channel-based processing unit 22. However, since the other configurations are the same as those of the spatial audio generation device 300 according to Embodiment 3, the same reference numerals are used for the same configurations as in Embodiment 3, and a detailed explanation is omitted.

[0065] The channel-based processing unit 22 generates sound for multi-speaker playback by changing the channel-based speaker arrangement to simulate audio playback with multiple speakers based on the angle information θ of the object position specification unit 20. As a result, the higher the order of the ambisonic sound, the more accurate the object position becomes.

[0066] Figure 12(A) is an explanatory diagram showing an example of how an ambisonic microphone 50 can collect sound from multiple sound sources 51 in space, and Figure 12(B) is an explanatory diagram showing how the channel-based speaker arrangement can be changed based on the angle information θ of the object positioning unit 20 when the speaker 52 is arranged in a 5.1ch speaker configuration as defined by ITU-R. In Figure 12(B), as an example, sound S is assumed to be produced when speaker 52L is positioned at an angle θ with respect to the axis 54 of speaker 52C which is placed in the direction in front of the person experiencing the sound. L S R S C S SL S SR This indicates that it will be generated.

[0067] Figure 13 is a flowchart showing an example of processing in the spatial audio generation device 400 according to Embodiment 4. The processing shown in Figure 13 is the same as that of the spatial audio generation device 300 according to Embodiment 3, except for the procedure in step S53, so a detailed explanation is omitted.

[0068] In step S53, the channel-based processing unit 22 processes the ambisonic sound S generated by the ambisonic sound generation unit 12. w S x S y S z Audio signal S for playback from multiple speakers L S R S C S SL S SR Generates sound S. L S R S C S SL S SR When generating the audio signal S, the channel-based processing unit 22 changes the channel-based speaker arrangement based on the angle information θ of the object positioning unit 20, for example, as shown in Figure 12(B). The channel-based processing unit 22 considers not only the angle information θ of the object positioning unit 20 but also the attenuation of sound due to distance d when generating the audio signal S for multiple speaker playback. L S R S C S SL S SRYou may generate this.

[0069] The channel-based processing unit 22 processes the audio signal S for playback by multiple speakers that was generated in step S53. L S R S C S SL S SR The output is sent to the binaural sound generation unit 15, and the procedure proceeds to step S42. Thereafter, the procedure from step S42 to step S46 is executed in the same manner as in Embodiment 3.

[0070] As described above, according to the spatial audio generation device 400 of Embodiment 4, by changing the channel-based speaker arrangement based on the angle information θ of the object position designation unit 20, the higher the order of the ambisonic sound, the more accurate the object position during playback can be.

[0071] Embodiment 5 Hereinafter, Embodiment 5 of the present disclosure will be described with reference to Figure 14. The spatial audio generation device 500 according to Embodiment 5 differs from the spatial audio generation device 300 according to Embodiment 3 in that it calculates object positions from information of data constituting free-viewpoint video stored in the database 23 and utilizes them for spatial audio. However, since the other configurations are the same as those of the spatial audio generation device 300 according to Embodiment 3, the same reference numerals are used for the same configurations as those of the spatial audio generation device 300 according to Embodiment 3, and detailed descriptions are omitted.

[0072] Database 23 stores three-dimensional spatial information, including point cloud data, NN (neural network), and sound sources such as Gaussian, which are pre-configured data for free-viewpoint video.

[0073] The object position calculation unit 24, by estimating that an image is a sound source image through image processing, determines that there is a sound source in the two-dimensional image generated from three-dimensional information such as point clouds, NNs, and Gaussians held in the database 23. In such cases, it can calculate the position of the sound source, indicated by the direction and distance to the sound source image contained in the two-dimensional image, from the three-dimensional information. For example, if the information held in the database is a point cloud, it measures the direction and distance to the points that make up the sound source image. Alternatively, if the information held in the database is a Gaussian, it may measure the direction and distance to the pixels that make up the sound source image from the two-dimensional image obtained by 3D Gaussian Platting, which generates an image from an arbitrary viewpoint by generating a 3D Gaussian from multiple original photographs and splatting it onto the two-dimensional image according to the viewpoint.

[0074] The ambisonic sound generation unit 12 receives an impulse response Imp corresponding to the distance d and angle θ output from the object position calculation unit 24 from the ambisonic impulse response unit 21. w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z Obtain (d, θ).

[0075] Figure 15 is a flowchart showing an example of processing in the spatial audio generation device 500 according to Embodiment 5. The processing shown in Figure 15 is the same as that of the spatial audio generation device 300 according to Embodiment 3, from step S52 to step S46, so a detailed explanation is omitted.

[0076] In step S61, the object position calculation unit 24 acquires three-dimensional information such as point clouds from the database 23.

[0077] In step S62, the object position calculation unit 24 calculates the object position indicating the location of the sound source from the acquired three-dimensional information, and the procedure proceeds to step S52. Thereafter, the procedure from step S2 to step S46 is executed in the same manner as in Embodiment 3.

[0078] As described above, the spatial audio generation device 500 according to Embodiment 5 calculates an object position indicating the location of a sound source from three-dimensional information stored in advance in the database 23. Then, it generates an impulse response Imp corresponding to the calculated object position. w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z Ambisonic sound S based on (d, θ) w S x S y S z From binaural sound signal S LE S RE By generating this, object-based audio can be realized that allows sound playback in a way that allows the listener to understand the location of the sound source.

[0079] Embodiment 6 Hereinafter, Embodiment 6 of the present disclosure will be described with reference to Figure 16. The spatial audio generation device 600 according to Embodiment 6 differs from the spatial audio generation device 300 according to Embodiment 3 in that it corresponds to two or more sound sources, an object position designation unit 25 specifies information on the positions of multiple objects, an ambisonic impulse response unit 26 outputs an impulse response corresponding to the multiple object positions, an ambisonic sound generation unit 28 generates an ambisonic sound corresponding to the multiple sound sources, and a sound enhancement processing unit 30 emphasizes a specific sound source (object) in the virtual space. However, since the other configurations are the same as those of the spatial audio generation device 300 according to Embodiment 3, the same reference numerals are used for the same configurations as those of the spatial audio generation device 300 according to Embodiment 3, and a detailed description is omitted.

[0080] In Figure 16, as an example, the audio storage unit 27 contains sound S 1 and sound S 2 Each of these sound sources has different sound data stored in it beforehand. The object position specification unit 25 is sound S 1 The object position (d) indicates the sound source location. 1 , θ 1 ), and sound S 2 The object position (d) indicates the sound source location. 2 , θ2 Specify ).

[0081] The ambisonic sound generation unit 28 receives an impulse response (Imp) corresponding to the specified object position from the ambisonic impulse response unit 26. w (d 1 , θ 1 ), Imp x (d 1 , θ 1 ), Imp y (d 1 , θ 1 ), Imp z (d 1 , θ 1 ), Imp w (d 2 , θ 2 ), Imp x (d 2 , θ 2 ), Imp y (d 2 , θ 2 ), Imp z (d 2 , θ 2 ) is obtained. In Figure 16, the object position (d 1 , θ 1 The impulse response related to ) is, Imp w (d 1 , θ 1 ), Imp x (d 1 , θ 1 ), Imp y (d 1 , θ 1 ), Imp z (d 1 , θ 1 A vector Imp(d) whose components are ) 1 , θ 1 ) is written. Similarly, object position (d 2 , θ 2 The impulse response related to ) is, Imp w (d 2 , θ 2 ), Imp x (d 2 , θ 2 ), Imp y (d 2 , θ 2 ), Impz (d 2 , θ 2 A vector Imp(d) whose components are ) 2 , θ 2 ) it says.

[0082] The ambisonic sound generation unit 28 generates ambisonic sound, which is three-dimensional audio data, by convolving the impulse response corresponding to the sound source's position information with the anechoic chamber recording sound of the sound source indicated by the position information corresponding to each impulse response. Specifically, the ambisonic sound generation unit 28 generates ambisonic sound, which is object position (d 1 , θ 1 The vector of the impulse response related to ) Imp(d 1 , θ 1 The component of the impulse response Imp w (d 1 , θ 1 ), Imp x (d 1 , θ 1 ), Imp y (d 1 , θ 1 ), Imp z (d 1 , θ 1 ) Sound S obtained from the sound source 1 It folds into an ambisonic sound S w (d 1 , θ 1 ), S x (d 1 , θ 1 ), S y (d 1 , θ 1 ), S z (d 1 , θ 1 A vector S whose components are ) 1 It generates the impulse response vector Imp(d 2 , θ 2 The component of the impulse response Imp w (d 2 , θ 2 ), Imp x (d 2 , θ 2 ), Imp y (d 2 , θ 2 ), Impz (d 2 , θ 2 ) Sound S obtained from the sound source 2 It folds into an ambisonic sound S w (d 2 , θ 2 ), S x (d 2 , θ 2 ), S y (d 2 , θ 2 ), S z (d 2 , θ 2 A vector S whose components are ) 2 Generates.

[0083] The sound emphasis designation unit 29 outputs an instruction to emphasize a specific object in the virtual space. Methods for emphasizing a specific object include muting other sounds or localizing a specific sound in the user's head.

[0084] The sound enhancement processing unit 30 enhances the sound of the object specified by the enhanced sound designation unit 29. Methods for enhancing the sound of a specific object include increasing the volume of the specific object, creating a sense of localization by making it appear to move from side to side to attract attention, or decreasing the volume of other sounds. Alternatively, spatial audio processing may be disabled for the sound of a specific object to localize the sound of that object within the listener's head and create a difference from other sounds.

[0085] The sound enhancement processing unit 30 generates a vector S' which enhances the sound of the specified object. 1 , S' 2 Each of these is output to the channel-based processing unit 14. Thereafter, processing by the ambisonic sound playback unit 13 is performed in the same manner as in Embodiment 3.

[0086] Figure 17 is a flowchart showing an example of processing in the spatial audio generation device 600 according to Embodiment 6. The processing shown in Figure 16 is the same as that of the spatial audio generation device 300 according to Embodiment 3 from step S41 to step S46, so a detailed explanation is omitted.

[0087] In step S71, the object positioning unit 25 controls sound S 1 The object position (d) indicates the sound source location. 1 , θ 1 ), and sound S 2 The object position (d) indicates the sound source location. 2 , θ 2 Specify ).

[0088] In step S72, the ambisonic sound generation unit 28 receives the object position (d) specified by the object position specification unit 20 from the ambisonic impulse response unit 21. 1 , θ 1 ), (d 2 , θ 2 The impulse response corresponding to ) is obtained.

[0089] In step S73, the ambisonic sound generation unit 28 generates sound S from the sound storage unit 27. 1 S 2 Obtain it.

[0090] In step S74, the ambisonic sound generation unit 28 acquires sound S from the sound storage unit 11. 1 S 2 By convolving the impulse response obtained from the ambisonic impulse response unit 26, sound S 1 S 2 It generates ambisonic sounds related to this.

[0091] In step S75, the sound enhancement processing unit 30 enhances the sound of the object specified by the sound enhancement designation unit 29 and proceeds to step S41. Thereafter, the procedures from step S41 to step S46 are executed in the same manner as in Embodiment 3.

[0092] As described above, the spatial audio generation device 600 according to Embodiment 6 can improve the sense of presence during sound reproduction by increasing the volume of a specific object.

[0093] Embodiment 7 Embodiment 7 of the present disclosure will be described below with reference to Figure 18. The spatial audio generation device 700 according to Embodiment 7 differs from the spatial audio generation device 300 according to Embodiment 3 in that it includes a reverberation estimation unit 31 that estimates the reverberation of the reproducible space, which is the space related to sound reproduction, and an ambisonic impulse response unit 32 that acquires an impulse response corresponding to the estimated reverberation. However, since the other configurations are the same as those of the spatial audio generation device 300 according to Embodiment 3, the same reference numerals are used for the same configurations as in Embodiment 3, and a detailed explanation is omitted.

[0094] The reverberation estimation unit 31 estimates the reverberation components of the reproduced space. For example, the reverberation components are estimated using the mirror image method, which calculates the reflection components from the distance to obstacles where reflections occur, using CAD (Computer-Aided Design) data of the reproduced space. The mirror image method interprets reflected sound as sound waves emitted from "mirror image sound sources" located symmetrically to the sound source across the reflective surface. First, mirror image sound sources corresponding to the reflected sound on each surface are determined, and the reverberation components are calculated by calculating the contribution of each mirror image sound source.

[0095] The Ambisonic Sound Generation Unit 12 receives an impulse response Imp from the Ambisonic Impulse Response Unit 32 that corresponds to the distance d and angle θ specified from the Object Positioning Unit 20, and the reverberation component r estimated by the Reverberation Estimation Unit 31. w (d, θ, r), Imp x (d, θ, r), Imp y (d, θ, r), Imp z (d, θ, r) is obtained. The ambisonic impulse response unit 32 has been measured for each angle θ, distance d, and reverberation component r, and the impulse response Imp w (d, θ, r), Imp x (d, θ, r), Imp y (d, θ, r), Imp z (d, θ, r) are stored beforehand.

[0096] The ambisonic sound generation unit 12, similar to the spatial audio generation device 100 according to Embodiment 3, uses an impulse response Imp w(d, θ, r), Imp x (d, θ, r), Imp y (d, θ, r), Imp z By convolving (d, θ, r) into the sound S of the sound storage unit 11, an ambisonic sound is generated, which is a sound as if the sound source were recorded in the space where the ambisonic microphone 50 is placed.

[0097] Figure 19 is a flowchart showing an example of processing in the spatial audio generation device 700 according to Embodiment 7. The processing shown in Figure 19 is the same as that of the spatial audio generation device 300 according to Embodiment 3, from step S51 to steps S2 to S46, so a detailed explanation is omitted.

[0098] In step S81, the reverberation estimation unit 31 estimates the reverberation of the reproduced space. Then, in step S82, the ambisonic sound generation unit 12 receives an impulse response Imp from the ambisonic impulse response unit 32 that corresponds to the distance d and angle θ specified from the object position specification unit 20, and the reverberation component r estimated by the reverberation estimation unit 31. w (d, θ, r), Imp x (d, θ, r), Imp y (d, θ, r), Imp z (d, θ, r) is obtained. Thereafter, the procedure from step S2 to step S46 is performed in the same manner as in Embodiment 3.

[0099] As described above, the spatial audio generation device 700 according to Embodiment 7 can generate a more realistic spatial audio generation device, spatial audio generation method, and spatial audio generation program by estimating the reverberation of the reproduced space and reproducing sound that takes into account the influence of the estimated reverberation.

[0100] Embodiment 8 Hereinafter, Embodiment 8 of the present disclosure will be described with reference to Figure 20. The spatial audio generation device 800 according to Embodiment 8 differs from the spatial audio generation device 700 according to Embodiment 7 in that the reverberation estimation unit 34, which estimates the reverberation of the reconstructed space, includes an adaptive filter 35 that estimates the reverberation component r' using recorded sound (any sound) from a microphone 33 installed in the real space corresponding to the reconstructed space, which is a virtual space, and the sound S from the sound storage unit 11, and a reverberation information extraction unit 36 ​​that updates the reverberation component r' to the reverberation component r'' using the impulse response stored in the ambisonic impulse response unit 37. However, since the other configurations are the same as those of the spatial audio generation device 700 according to Embodiment 7, the same reference numerals as in Embodiment 7 will be used for the same configurations as in Embodiment 7, and a detailed explanation will be omitted.

[0101] Figure 21 is an explanatory diagram showing an example of a configuration for estimating the adaptive filter 35. The adaptive filter 35 is a method for estimating an unknown transfer function, and generates a reverberation component r' by convolving the estimated transfer function with the sound S of the sound storage unit 11, which is the original sound. In Figure 21, the sound S of the sound storage unit 11 is the original sound x played by the speaker. n It is input as follows: Speaker playback original sound x n This is input to the transfer function 70 that represents the reverberation. Speaker playback original sound x n The transfer function 70, which represents the input reverberation, outputs an echo y(n). The microphone-recorded sound z(n) picked up by microphone 33 is added to the echo y(n) as shown in equation (4) below.

[0102]

[0103] Also, speaker playback original sound x n This signal is also input to the adaptive filter 72, which has the transfer function shown in equation (5) below, and the adaptive filter 72 outputs the pseudo-echo shown in equation (6) below.

[0104]

[0105] The difference between u(n) shown in equation (4) above and the pseudo-echo shown in equation (6), d(n), is calculated as shown in equation (7) below.

[0106]

[0107] To calculate the optimal adaptive filter 72, the adaptive filter 72 shown in equation (5) is updated so that d(n) shown in equation (7) becomes smaller. Known methods such as the LMS (Least Mean Square) method or the NLMS (Normalized Least Mean Square) method are used to update the adaptive filter 72.

[0108] The reverberation estimation unit 34 calculates the reverberation component r' by convolving the optimized adaptive filter 72 onto the speaker original sound x(n).

[0109] The reverberation information extraction unit 36 ​​retrieves an impulse response with a small time-series signal difference from among the impulse responses that have been measured in advance in multiple environments and stored in the ambisonic impulse response unit 37, and convolves it with the reverberation component r' generated by the adaptive filter 35 to generate the reverberation component r''.

[0110] The ambisonic sound generation unit 12 receives an impulse response Imp from the ambisonic impulse response unit 37 that corresponds to the distance d and angle θ specified from the object position specification unit 20, and the reverberation component r'' estimated by the reverberation estimation unit 34. w (d, θ, r''), Imp x (d, θ, r''), Imp y (d, θ, r''), Imp z (d, θ, r'') is obtained. The ambisonic impulse response unit 37 has been measured for each angle θ, distance d, and reverberation component r'', resulting in an impulse response Imp w (d, θ, r''), Imp x (d, θ, r''), Imp y (d, θ, r''), Imp z (d, θ, r'') are stored beforehand.

[0111] The ambisonic sound generation unit 12, similar to the spatial audio generation device 700 according to Embodiment 7, generates ambisonic sound, which is like sound recorded in a space where the ambisonic microphone 50 is placed, by convolving the impulse responses Impw(d, θ, r''), Impx(d, θ, r''), Impy(d, θ, r''), and Impz(d, θ, r'') into the sound S of the sound storage unit 11.

[0112] Figure 22 is a flowchart showing an example of processing in the spatial audio generation device 800 according to Embodiment 8. The processing shown in Figure 22 is the same as that of the spatial audio generation device 700 according to Embodiment 7 in steps S2, S51 and S3 to S46, so a detailed explanation is omitted.

[0113] In step S91, sound is acquired using a microphone 33 placed in the real space corresponding to the reproduced space. The sound recorded by microphone 33 can be anything; the sound source is not specified.

[0114] In step S92, the reverberation component r'' is estimated. The procedure in step S92 consists of subroutines for processing with an adaptive filter and extracting reverberation information, as will be described later.

[0115] In step S93, the impulse responses Impw(d, θ, r''), Impx(d, θ, r''), Impy(d, θ, r''), and Impz(d, θ, r'') corresponding to the distance d and angle θ specified from the object positioning unit 20, and the reverberation component r'' estimated by the reverberation estimation unit 34 are obtained. Thereafter, the procedures from step S3 to step S46 are executed in the same manner as in Embodiment 7.

[0116] Figure 23 is a flowchart showing an example of the reverberation estimation process in step S92 of Figure 22. In step S100, the optimized adaptive filter 72 is convolved onto the speaker original sound x(n) to calculate the reverberation component r'.

[0117] In step S101, the impulse response obtained from the ambisonic impulse response unit 37 is convolved into the reverberation component r' to generate the reverberation component r'' and the process is returned.

[0118] As described above, the spatial audio generation device 800 according to Embodiment 8 estimates reverberation based on sound recorded by a microphone 33 installed in a real space corresponding to the reconstructed space which is a virtual space, and reproduces sound that takes into account the effect of the estimated reverberation, thereby enabling a more realistic spatial audio generation device, spatial audio generation method, and spatial audio generation program.

[0119] Embodiment 9 Embodiment 9 of this disclosure will now be described with reference to Figure 24. The spatial audio generation device 900 according to Embodiment 9 differs from the spatial audio generation device 300 according to Embodiment 3 in that it includes a reverberation information addition unit 39 that adds reverberation to the original sound S using reverberation components r' estimated by the adaptive filter 35. However, since the other configurations are the same as those of the spatial audio generation device 300 according to Embodiment 3, the same reference numerals are used for the same configurations as in Embodiment 3, and detailed descriptions are omitted.

[0120] The reverberation estimation unit 38, like in Embodiment 8, includes an adaptive filter 35 that estimates the reverberation component r' using the sound recorded by the microphone 33 and the sound S from the sound storage unit 11. The estimated reverberation component r' is output to the reverberation information addition unit 39. The reverberation component r' may also be calculated by simulation using CAD data of the reproduced space, as in Embodiment 7.

[0121] The reverberation information addition unit 39 generates sound S' by convolving the reverberation component r' obtained by the reverberation estimation unit 38 into sound S of the sound storage unit 11, thereby simulating the sense of reverberation in space.

[0122] The ambisonic sound generation unit 40 receives the impulse response Imp from the ambisonic impulse response unit 21. w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z By convolving (d, θ) into sound S', an ambisonic sound is generated, which is a sound similar to that recorded in a space where an ambisonic microphone is placed. Thereafter, processing by the ambisonic sound playback unit 13 is performed in the same manner as in Embodiment 3.

[0123] Figure 25 is a flowchart showing an example of processing in the spatial audio generation device 900 according to Embodiment 9. The processing shown in Figure 25 is the same as that of the spatial audio generation device 300 according to Embodiment 3 in steps S2, S51, S52, and S41 to S46, so a detailed explanation is omitted.

[0124] In step S91, sound is acquired using a microphone 33 placed in the real space corresponding to the reproduced space. The sound recorded by microphone 33 can be anything; the sound source is not specified.

[0125] In step S96, the reverberation component r' is estimated. The procedure in step S96 consists of a subroutine for processing with an adaptive filter, as will be described later.

[0126] In step S97, the reverberation information addition unit 39 generates sound S' by convolving the reverberation component r' obtained by the reverberation estimation unit 38 into sound S of the sound storage unit 11.

[0127] In step S98, the impulse response Imp obtained from the ambisonic sound generation unit 40 and the ambisonic impulse response unit 21 w (d, θ), Imp x (d, θ), Imp y (d, θ), Imp z By convolving (d, θ) into sound S', an ambisonic sound is generated, which is a sound as if the sound source were recorded in a space where an ambisonic microphone was placed. Thereafter, the procedure from step S41 to step S46 is performed in the same manner as in Embodiment 3.

[0128] Figure 26 is a flowchart showing an example of the reverberation estimation process in step S96 of Figure 25. In step S94, the optimized adaptive filter 72 is convolved onto the speaker original sound x(n) to calculate the reverberation component r' and the process is returned.

[0129] As described above, the spatial audio generation device 900 according to Embodiment 9 estimates the reverberation component r' based on the sound recorded by the microphone 33 installed in the real space corresponding to the reproducible space which is a virtual space, and generates sound S' by convolving the estimated reverberation component r' into sound S of the sound storage unit 11, thereby enabling a more realistic spatial audio generation device, spatial audio generation method, and spatial audio generation program.

[0130] The spatial audio generation device, spatial audio generation method, and spatial audio generation program described herein enable realistic audio playback with externalized sound that conveys a sense of direction and distance through easy sound collection. For example, when conducting factory work training in a VR space, it becomes possible to provide feedback sounds from the device in response to the trainee's movements.

[0131] Furthermore, the spatial audio generation device, spatial audio generation method, and spatial audio generation program related to this disclosure can be applied to virtual meetings, remote collaboration where personnel dispersed in remote locations work together, and tours of facilities such as showrooms or factories located in remote locations.

[0132] Furthermore, the spatial audio generation device, spatial audio generation method, and spatial audio generation program related to this disclosure make it possible to virtually experience the sound environment of large-scale products such as elevators and to verify noise levels.

[0133] 10 Ambisonic impulse response unit, 11 Audio storage unit, 12 Ambisonic sound generation unit, 13 Ambisonic sound playback unit, 14 Channel-based processing unit, 15 Binaural sound generation unit, 16 Headphones, 17 Head tracking adjustment unit, 20 Object position specification unit, 21 Ambisonic impulse response unit, 22 Channel-based processing unit, 23 Database, 24 Object position calculation unit, 25 Object position specification unit, 26 Ambisonic impulse response unit, 27 Audio storage unit, 28 Ambisonic sound generation unit, 29 Emphasized sound specification unit, 30 Sound enhancement processing unit, 31 Echo estimation unit, 32 Ambisonic impulse response unit, 33 Microphone, 34 Echo estimation unit, 35 Adaptive filter, 36 Echo information extraction unit, 37 Ambisonic impulse response unit, 38 Echo estimation unit, 39 Echo information addition unit, 40 Ambisonic sound generation unit, 50 Ambisonic microphone, 51 Sound source, 52, 52C, 52L Speaker, 53 Head, 54 Axis, 61 Processing circuit, 62 Microphone, 63 Headphones, 64 External storage device, 65 Processor, 66 Memory, 67 Bus, 68 Rotation sensor, 70 Transfer function, 72 Adaptive filter, 100, 200, 300, 400, 500, 600, 700, 800, 900 Spatial audio generator.

Claims

1. A spatial audio generation device comprising: an impulse response storage unit that stores the spherical harmonic function expanded impulse response of a microphone array in a reproduction space which is a space related to sound reproduction; a sound storage unit that stores an audio signal; and a playback sound data generation unit that outputs playback sound data of the audio signal generated by convolving the audio signal obtained from the sound storage unit with the impulse response obtained from the impulse response storage unit.

2. A spatial audio generation device according to claim 1, comprising: an audio playback unit for playing the playback sound data, wherein the audio playback unit includes: a channel-based processing unit for generating a sound signal for multi-speaker playback that simulates audio playback with multiple speakers arranged in space from the playback sound data generated by the playback sound data generation unit; a binaural sound generation unit for generating a binaural sound signal from the multi-speaker playback sound signal based on a head-related transfer function; a speaker for playing the binaural sound signal; and a rotation sensor for detecting the direction of the head of an experiencer listening to the binaural sound signal played by the speaker, wherein the channel-based processing unit displaces the sound image of the multi-speaker playback sound signal based on the direction of the experiencer's head detected by the rotation sensor.

3. The spatial audio generation device according to claim 2, wherein the head-related transfer function is either a head-related transfer function obtained by prior measurement or a head-related transfer function calculated by numerical calculation.

4. The spatial audio generation device according to claim 2 or 3, further comprising an object positioning unit that specifies position information indicating the location of a sound source present in the reproduced space by distance and angle, wherein the playback sound data generation unit acquires an impulse response corresponding to the specified distance and angle from the impulse response storage unit.

5. The spatial audio generation device according to claim 4, wherein the channel-based processing unit generates a sound for multi-speaker playback that simulates audio playback when the multi-speaker unit is arranged based on the angle specified by the object positioning unit.

6. The spatial audio generation device according to claim 4, further comprising: a database storing three-dimensional information of a space including the sound source; and an object position calculation unit that calculates position information of the sound source included in the three-dimensional information, wherein the playback sound data generation unit acquires an impulse response from the impulse response storage unit that corresponds to the position information of the sound source calculated by the object position calculation unit.

7. The spatial audio generation device according to claim 4, further comprising: an audio storage unit that stores anechoic chamber recordings of two or more sound sources, each with different sound; an object positioning unit that specifies the position information of each of the two or more sound sources; and a playback sound data generation unit that generates three-dimensional audio data by convolving impulse responses corresponding to the position information of each of the specified two or more sound sources obtained from the impulse response storage unit with the anechoic chamber recordings of the sound sources indicated by the position information corresponding to each impulse response, the device further comprising: an emphasis sound specification unit that outputs an instruction to emphasize a specific sound source among the two or more sound sources; and a sound emphasis processing unit that emphasizes the three-dimensional audio data relating to the specific sound source indicated by the emphasis sound specification unit.

8. The spatial audio generation device according to claim 4, further comprising a reverberation estimation unit for estimating the reverberation components of the reproduced space, wherein the playback sound data generation unit acquires from the impulse response storage unit the distance and angle specified from the object position specification unit, and the impulse response corresponding to the reverberation components estimated by the reverberation estimation unit.

9. The spatial audio generation apparatus according to claim 8, comprising: an adaptive filter that estimates reverberation components using sound recorded by a microphone installed in a real space corresponding to the reproduced space and sound recorded in an anechoic chamber stored in the sound memory unit; and a reverberation information extraction unit that updates the reverberation components with an impulse response stored in the impulse response memory unit.

10. A spatial audio generation device according to any one of claims 1 to 4, further comprising: an echo estimation unit including an adaptive filter that estimates echo components using sound recorded by a microphone installed in a real space corresponding to the reproduced space and sound recorded in an anechoic chamber stored in the sound memory unit; and an echo information addition unit that generates sound by convolving the echo components estimated by the echo estimation unit into the sound recorded in the anechoic chamber stored in the sound memory unit, wherein the playback sound data generation unit convolves the impulse response obtained from the impulse response storage unit into the sound output from the echo information addition unit.

11. A spatial audio generation method performed by a computer, comprising: storing the spherical harmonic function expanded impulse response of a microphone array in a reproducible space which is a space related to sound playback; storing an audio signal; and outputting playback sound data of the audio signal generated by convolving the stored audio signal with the impulse response.

12. A spatial audio generation program that causes a computer to perform the following steps: storing the spherical harmonic function expanded impulse response of a microphone array in a reproduction space which is a space related to sound reproduction; storing an audio signal; and outputting playback sound data of the audio signal generated by convolving the stored audio signal with the impulse response.