Acoustic signal generation device, acoustic signal generation method, and acoustic signal generation program

The acoustic signal generating device uses headphones and surrounding speakers to generate and output direct and indirect sounds based on head position, addressing the challenge of sound image localization and distance perception by adjusting the direct-to-reverb ratio, thereby enhancing the sense of distance in sound localization.

WO2025253637A1PCT designated stage Publication Date: 2025-12-11NT T INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/020892
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing sound image localization techniques, such as those using head-related transfer functions and direct-to-reverb ratios, face challenges in accurately localizing sound images at desired positions due to individual head shape differences and varying reproduction methods, making it difficult to provide an absolute sense of distance.

Method used

An acoustic signal generating device that utilizes open-type headphones and speakers arranged around the listener to generate and output direct and indirect sounds based on head position information, adjusting the direct-to-reverb ratio to localize sound images at any position, including a sense of distance.

Benefits of technology

Enables easy realization of sound image localization with a sense of distance by accurately positioning sound images using headphones and surrounding speakers, overcoming individual head shape variations and reproduction method inconsistencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024020892_11122025_PF_FP_ABST
    Figure JP2024020892_11122025_PF_FP_ABST
Patent Text Reader

Abstract

An acoustic signal generation device according to an embodiment of the present invention provides a listener with sound based on a sound source by using a headphone worn by the listener and a speaker disposed around the listener. The acoustic signal generation device includes a communication circuit and a processor. The communication circuit is configured to receive position information of the head of the listener. The processor is configured to generate a first acoustic signal and a second acoustic signal for localizing a sound image at a position of the sound source on the basis of the position information, and output the generated first acoustic signal and second acoustic signal to the headphone and the speaker, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic signal generating device, acoustic signal generating method, and acoustic signal generating program

[0001] The embodiments relate to an acoustic signal generating device, an acoustic signal generating method, and an acoustic signal generating program.

[0002] In recent years, content using head-mounted displays (HMDs), such as virtual reality (VR) and augmented reality (AR), has become increasingly popular. VR technology can provide users with an experience that makes them feel as if they are actually there in a virtual space. AR technology can overlay digital content, such as CG images, onto real space, making it appear as if they exist in real space. These technologies can provide users with an experience that is comparable to the real world, thereby integrating virtual and real spaces. To improve the quality of the experience of integrating virtual and real spaces through content using HMDs, there is a need to present the sense of distance in addition to the direction from which sound is coming.

[0003] Methods for presenting a sense of distance using sound include, for example, a method that changes the volume of the sound played according to the distance, i.e., a method that presents a sense of distance using sound pressure differences, and a method that uses sound image localization technology. Methods that present a sense of distance using sound pressure differences allow the user to perceive a relative sense of distance based on the change in sound pressure level from a reference sound. On the other hand, it is difficult to allow the user to perceive absolute distance using sound pressure differences.

[0004] A commonly used sound image localization technique is a method using a head-related transfer function (HRTF) with headphones. The HRTF is a function that reproduces the amplitude difference and phase difference between the two ears based on a signal containing sound propagation information from the position of the sound source to the pinna. Sound reproduction using the HRTF can localize a sound image at any position, thereby providing an acoustic experience comparable to that of a real space compared to simple stereo reproduction. Non-Patent Document 1 also discloses a sound image localization technique that presents a sense of distance using a direct-to-reverb (DR) ratio, which is the ratio between direct sound that enters the ear directly from the sound source and indirect sound that enters the ear after reflecting off a wall or the like. For example, by presenting sound sources with different DR ratios created using measured impulse responses to a listener, it has been shown that the DR ratio contributes to the sense of distance.

[0005] Toru Kamekawa and Atsushi Marui, "The influence of spatial sound reproduction methods and musical context on the perception of distance and depth," Journal of the Acoustical Society of Japan, Vol. 72, No. 11, 2016, pp. 684-695

[0006] However, in order to localize a sound image at a desired position using a head-related transfer function, it is necessary to record the head-related transfer function of the desired position in advance, which requires a lot of effort. Also, since there are individual differences in head shape, etc., there are individual differences in the head-related transfer function. Furthermore, since the reproduction of the DR ratio is affected by the reproduction method, such as stereo reproduction or surround reproduction, no reproduction method has been established.

[0007] Therefore, the present invention has been made in light of the above circumstances, and its object is to provide an acoustic signal generating device, an acoustic signal generating method, and an acoustic signal generating program that can easily realize sound image localization including the sense of distance.

[0008] An acoustic signal generating device according to an embodiment provides a listener with sound based on a sound source using headphones worn by the listener and speakers arranged around the listener. The acoustic signal generating device includes a communication circuit and a processor. The communication circuit is configured to receive position information of the listener's head. The processor is configured to generate a first acoustic signal and a second acoustic signal that localize a sound image at the position of the sound source based on the position information, and to output the generated first acoustic signal and second acoustic signal to the headphones and speakers, respectively.

[0009] According to the embodiments, it is possible to provide an audio signal generating device, an audio signal generating method, and an audio signal generating program that can easily realize sound image localization including a sense of distance.

[0010] Fig. 1 is a block diagram showing an example of the overall configuration of an acoustic signal generation system according to an embodiment. Fig. 2 is a block diagram showing an example of the hardware configuration of an acoustic signal generation device according to an embodiment. Fig. 3 is a schematic diagram showing an example of an observation system of an acoustic signal generation system according to an embodiment. Fig. 4 is a block diagram showing an example of the functional configuration of an acoustic signal generation device according to an embodiment. Fig. 5 is a flowchart showing an example of the operation of an acoustic signal generation device according to an embodiment.

[0011] Hereinafter, embodiments will be described with reference to the drawings. The embodiments illustrate devices and methods for embodying the technical ideas of the invention. The drawings are schematic or conceptual. In the following, components having substantially the same functions and configurations are assigned the same reference numerals.

[0012] The embodiment relates to a method for generating direct sound played from open-type headphones that use the position coordinates of the listener's head as input and localize the sound source at any position, including distance, and indirect sound played from speakers placed around the listener.

[0013] <1> Configuration First, the configuration of an acoustic signal generation system 1 according to an embodiment will be described.

[0014] <1-1> Overall Configuration of Acoustic Signal Generation System 1 Fig. 1 is a block diagram showing an example of the overall configuration of an acoustic signal generation system 1 according to an embodiment. As shown in Fig. 1, the acoustic signal generation system 1 includes, for example, an acoustic signal generation device 10, headphones 20, and a speaker 30.

[0015] The acoustic signal generating device 10 is a computer that generates an output acoustic signal to be reproduced by the headphones 20 and an output acoustic signal to be reproduced by the speaker 30 based on position information of the listener's head (hereinafter referred to as head position information). Specifically, the acoustic signal generating device 10 is configured to receive the head position information as input and generate an acoustic signal capable of controlling a direct to reverb ratio (DR ratio) that localizes a sound image at the position of a sound source for the listener. The acoustic signal generating device 10 then transmits the generated output acoustic signal to the headphones 20 and the speaker 30. A detailed method for generating the acoustic signals will be described later.

[0016] The headphones 20 are open-type headphones that can be worn on the head of a listener and are configured to be able to reproduce sound based on the output sound signal received from the sound signal generating device 10. The headphones 20 also function as a head tracker that acquires head position information of the listener when worn on the listener's head. The headphones 20 can then transmit the acquired head position information to the sound signal generating device 10.

[0017] The speaker 30 is a speaker that is placed away from the listener and configured to be able to reproduce sound based on the output sound signal received from the sound signal generation device 10. Note that the sound signal generation system 1 may use multiple speakers 30. The following describes a case where the sound signal generation system 1 uses one speaker 30.

[0018] 2 is a block diagram showing an example of the hardware configuration of the acoustic signal generation device 10 according to the embodiment. As shown in FIG. 2, the acoustic signal generation device 10 includes, for example, a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a communication module 14, and a storage device 15.

[0019] The CPU 11 is a processor capable of executing various programs and controls the overall operation of the acoustic signal generating device 10. The ROM 12 is, for example, a non-volatile semiconductor memory and stores programs and control data for controlling the acoustic signal generating device 10. The RAM 13 is, for example, a volatile semiconductor memory and is used as a work area for the CPU 11. The communication module 14 is a communication circuit configured to be able to transmit acoustic signals to each of the headphones 20 and the speaker 30. The communication module 14 can also receive head position information from, for example, the headphones 20. Communication between the acoustic signal generating device 10 and each of the headphones 20 and the speaker 30 may be wireless or wired. The communication module 14 may be configured to be able to communicate with a network (not shown). The storage device 15 stores, for example, acoustic data.

[0020] <1-3> Observation System of Acoustic Signal Generation System 1 Fig. 3 is a schematic diagram showing an example of the observation system of the acoustic signal generation system 1 according to the embodiment. For simplicity, Fig. 3 uses an xyz Cartesian coordinate system and shows the observation system on the xy plane when z = 0. As shown in Fig. 3, the origin of this observation system corresponds to the initial position of the head, and 0 The position of the head of the listener 40 is represented by X (t) = [0,0,0]. H (t) = [x H , y H , z H The acoustic signal generation system 1 according to the embodiment uses the headphones 20 and the speaker 30 to generate an XS = [x S , y S , z S ], the sound source P S The headphone 20 has a driver unit 21 associated with the left ear of the listener 40 and a driver unit 22 associated with the right ear of the listener 40. The position of the driver unit 21 is X L The position of the driver unit 22 is indicated by X R is shown by

[0021] <1-4> Functional configuration of the acoustic signal generating device 10 Fig. 4 is a block diagram showing an example of the functional configuration of the acoustic signal generating device 10 according to an embodiment. As shown in Fig. 4, the acoustic signal generating device 10 includes, for example, a sound source management unit 101, an indirect sound design unit 102, a direct sound design unit 103, and an acoustic signal correction unit 104.

[0022] The sound source management unit 101 manages the sound source P S (t) and sound source P S Position X of (t) S The sound source management unit 101 manages the sound source P S (t) data and position X S and to the indirect sound design unit 102 and the direct sound design unit 103. The sound source management unit 101 also outputs the sound source P S Position X of (t) S is output to the acoustic signal correction unit 104.

[0023] The indirect sound design unit 102 determines the position X of the head of the listener 40. H (t) and sound source P S Position X of (t) S and the indirect sound P 0 Then, the indirect sound design unit 102 generates the generated indirect sound P 0 (t) is output to the sound signal correction unit 104. 0 (t) corresponds to the acoustic signals reproduced by the speakers 30 arranged around the listener 40 and indirectly reaching the left and right ears of the listener 40 .

[0024] The direct sound design unit 103 determines the position X of the head of the listener 40. H (t) is input, and the listener 40 and the sound source P S Position X of (t) S Then, the direct sound design unit 103 calculates the positional relationship between the sound source P and the sound source P, which is the sound source P where the sound image is to be localized. S Position X of (t) S Based on this, the direct sound P L,R Then, the direct sound design unit 103 generates the generated direct sound P L,R (t) is output to the acoustic signal correction unit 104. L,R (t) is the direct sound P reproduced by the driver unit 21 L (t) and the direct sound P reproduced by the driver unit 22 R (t) and the direct sound P L,R (t) corresponds to the acoustic signals reproduced by the headphones 20 and directly reaching the left and right ears of the listener 40 .

[0025] The acoustic signal correction unit 104 corrects the acoustic signal from the sound source P S Position X of (t) S Based on this, the indirect sound P 0 (t) and direct sound P L,R (t) and the time delay and amplitude value until the signals reach the listener 40. Specifically, the acoustic signal correcting unit 104 corrects the coefficient t i and correct the time delay using the coefficient c i Then, the sound signal correcting unit 104 corrects the amplitude value using the corrected indirect sound P 0 Indirect sound CP (t) 0 (t) is output to the speaker 30, and the corrected direct sound P L,R (t) is the direct sound CP L,R (t) is output to the headphones 20.

[0026] <2> Operation Next, the operation of the acoustic signal generation system 1 according to the embodiment will be described.

[0027] 5 is a flowchart showing an example of the operation of the acoustic signal generation device 10 according to the embodiment. For example, when performing sound reproduction, the acoustic signal generation device 10 executes (starts) the series of processes shown in FIG. 5 in response to a time t acquired by head tracking.

[0028] (Step ST1) First, the process of step ST1 is executed. In the process of step ST1, the indirect sound design unit 102 calculates the head position X H (t) = [x H , y H , z H ] and sound source P S Position X of (t) S = [x S , y S , z S ] and indirect sound P 0 Assuming a space with a uniform reflectance r due to the wall surfaces, the sound source P S The room transfer function G(t) calculated by the mirror method using the position information of the listener 40 and the listener 40 is expressed by the following equation (1).

[0029]

[0030] In equation (1), n ​​represents the number of reflections by the wall surface, and N (≧0) represents the maximum number of reflections that can be arbitrarily set. S,n is the sound source P when n reflections are assumed S Position X of (t) S Assuming a two-dimensional rectangular room, the position of the mirror image is X S,n is (2n+1) 2 - A matrix containing the position of one mirror image.

[0031] indirect sound P 0 (t) is P 0 (t) and the room transfer function G(t) 0 The reverberation part G with (t) removed R Specifically, it is calculated by convolving only the indirect sound P 0 (t) is expressed as the following equation (2).

[0032]

[0033] The shape of the room when the indirect sound design unit 102 calculates G(t) is not limited to a two-dimensional rectangular room, and can be set to any shape.

[0034] (Step ST2) Next, the process of step ST2 is executed. In the process of step ST2, the direct sound design unit 103 calculates the position X H (t) and sound source P S Position X of (t) S and the direct sound P reproduced from the open-type headphones 20. L,R (t) is generated. L,R (t) is the head position X H (t) The estimated positions X of the left and right ears obtained from L,R (t) and sound source P S Position X of (t) S The initial head position X is generated using the positional relationship between the head and the sound, and the amplitude and phase differences of the sound. 0 (t)=[0,0,0] T When the origin is set to , the estimated initial position of the left ear is expressed by the following equation (3), and the estimated initial position of the right ear is expressed by the following equation (4): where h is the average head width of the listener 40.

[0035]

[0036]

[0037] The positions of both ears when the head of the listener 40 moves are determined by the rotation matrix R around the x-axis. x and the rotation matrix R around the y-axis y and the rotation matrix R around the z-axis z and is expressed as the following equation (5): where α is the Euler angle around the y-axis at time t obtained by head tracking, β is the Euler angle around the x-axis at time t, and γ is the Euler angle around the z-axis at time t.

[0038]

[0039] Rotation matrix R around the x-axisx (β) is expressed as the following equation (6).

[0040]

[0041] Rotation matrix R around the y-axis y (α) is expressed as the following equation (7).

[0042]

[0043] Rotation matrix R around the z-axis z (γ) is expressed as the following equation (8).

[0044]

[0045] The acoustic signal entering both ears (direct sound P L,R (t)) is the amplitude coefficient A L,R The interaural time difference ITD(t) is calculated using the phase difference of the sound.

[0046] Amplitude coefficient A L,R (t) is calculated on the assumption that the sound pressure is inversely proportional to the distance from the sound source, and is expressed as the following equation (9).

[0047]

[0048] The interaural time difference ITD(t) is the distance D from the sound source to the left ear. L (t) = ||X S -X L (t) || and the distance D from the sound source to the right ear R (t) = ||X S -X R (t)∥ is used to express it as in the following equation (10): In equation (10), c represents the speed of sound, and fs represents the sampling frequency.

[0049]

[0050] Direct sound P L,R (t) is the direct sound part G 0 (t) and sound source P S Specifically, the direct sound P L,R(t) is the direct sound portion G(t) excluding the reverberation component of the room transfer function G(t) calculated in step ST1. 0 (t) and the amplitude coefficient A L,R The direct sound design unit 103 calculates the direct sound P L,R When generating (t), signals are synthesized using overlap-add to avoid sound interruptions due to changes in the interaural time difference ITD(t).

[0051]

[0052]

[0053]

[0054] (Step ST3) Next, the process of step ST3 is executed. In the process of step ST3, the sound signal correcting unit 104 corrects the indirect sound P obtained in step ST1. 0 (t) and the direct sound P obtained in step ST2 L,R (t) and the indirect sound CP with the time delay and amplitude corrected. 0 (t) and direct sound CP L,R (t) is generated. Details of this correction will be explained below.

[0055] When a sound source is reproduced simultaneously by the headphones 20 and the speaker 30, the direct sound P L,R (t) is the indirect sound P from the speaker 30 0 The sound reaches the listener 40 before (t). Therefore, due to the precedence effect, the sound image may be localized at the position of the headphones 20. Furthermore, when a sound source is played back using different devices, such as the headphones 20 and the speakers 30, the volume differs between the devices, making it difficult to reproduce the DR ratio.

[0056] Correct the time delay t i As a method for determining the initial position X of the listener 40, 0 The cross-correlation is maximized by using the microphones arranged at the delay time t iThe acoustic signal correction unit 104 determines t at which the cross-correlation function R shown in the following equation (14) is maximized. i In equation (14), the delay time can be corrected by head (t) and P loud (t) is a signal picked up by a microphone at the position of the listener 40 when the same sound source is reproduced from the headphones 20 and the speaker 30, respectively.

[0057]

[0058] The coefficient c that corrects the amplitude value i is expressed by the following equation (14): head (t) and P loud c so that the maximum amplitude of (t) coincides i This can be corrected by calculating

[0059]

[0060] c calculated based on the above formula (15) i As a result, the indirect sound CP output from the acoustic signal correction unit 104 is 0 (t) is expressed as the following equation (16).

[0061]

[0062] t calculated based on the above formula (14) i The direct sound CP output from the acoustic signal correction unit 104 is L,R (t) is expressed as the following equation (17).

[0063]

[0064] As described above, the acoustic signal correction unit 104 corrects the head position X obtained from the head tracker worn by the listener 40. H (t) is input, and the listener 40 receives the sound source P S Position X of (t) S When the process of step ST3 is completed, the acoustic signal generating device 10 ends the series of processes shown in FIG.

[0065] When a plurality of speakers 30 are arranged around the listener 40, the acoustic signal correction unit 104 adjusts each speaker 30 to the initial position X 0 Then, the acoustic signal correction unit 104 simultaneously reproduces the signals from the multiple speakers 30 and adjusts the phase difference to the initial position X 0 The signal collected at (t) is subjected to signal correction in step ST3.

[0066] <3> Effects of the embodiment As described above, in the acoustic signal generation system 1 according to the embodiment, the acoustic signal generation device 10 acquires the positional relationship and distance to a sound source based on the position coordinates of the head of the listener 40 acquired by head tracking, and generates an acoustic signal that becomes a direct sound. Furthermore, the acoustic signal generation device 10 generates an acoustic signal that becomes an indirect sound by the mirror method in order to present a sense of distance using a DR ratio. The acoustic signal generation device 10 then plays the generated direct sound on the headphones 20 worn on the head of the listener 40, and plays the generated indirect sound on the speakers 30 arranged around the listener 40.

[0067] As a result, the acoustic signal generating device 10 according to the embodiment can localize a sound image at any coordinates and realize sound image localization that includes a sense of distance using the headphones 20 and the speakers 30 arranged around the listener 40. In this way, the acoustic signal generating device 10 according to the embodiment can easily realize sound image localization that includes a sense of distance.

[0068] <4> Others In the acoustic signal generation system 1 according to the embodiment, the CPU 11 of the acoustic signal generation device 10 may be another circuit. For example, the acoustic signal generation device 10 may include an MPU (Micro Processing Unit) instead of a CPU. Each of the processes described in each embodiment may be realized by dedicated hardware. Each of the processes described in the above embodiments may be a mixture of processes executed by software and processes executed by hardware, or may be only one of them. The acoustic signal generation device 10 may obtain the sound source to be played via a network.

[0069] The headphones 20 described in the embodiment may be closed-type headphones if they are capable of reproducing ambient sounds. The headphones 20 may be mounted on a head-mounted display. The head-mounted display may be equipped with a head tracker function. In this case, the acoustic signal generating device 10 acquires position information of the listener 40 from the head-mounted display, generates direct sound and indirect sound based on the operation shown in FIG. 5 , and outputs the generated direct sound and indirect sound to the headphones 20 and the speaker 30, respectively. The head-mounted display may also be equipped with the functions of the acoustic signal generating device 10.

[0070] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.

[0071] REFERENCE SIGNS LIST 1 Acoustic signal generation system 10 Acoustic signal generation device 11 CPU 12 ROM 13 RAM 14 Communication module 15 Storage device 101 Sound source management unit 102 Indirect sound design unit 103 Sound design unit 104 Acoustic signal correction unit 20 Headphones 21, 22 Driver unit 30 Speaker 40 Listener

Claims

1. An acoustic signal generating device that provides a listener with sound based on a sound source using headphones worn by the listener and speakers arranged around the listener, comprising: a communication circuit configured to receive head position information of the listener; and a processor configured to generate a first acoustic signal and a second acoustic signal that localize a sound image at the position of the sound source based on the head position information and the position of the sound source, and to output the generated first acoustic signal and second acoustic signal to the headphones and the speakers, respectively.

2. The acoustic signal generating device according to claim 1, wherein the processor is further configured to calculate the second acoustic signal by convolving the sound source with a reverberation portion of a room transfer function calculated by a mirror method using the head position information and the position of the sound source, the reverberation portion being obtained by removing the direct sound portion from the room transfer function.

3. The acoustic signal generating device according to claim 2, wherein the processor is further configured to calculate the first acoustic signal by convolving the direct sound portion with the sound source.

4. The acoustic signal generating device of claim 3, wherein the first acoustic signal includes a third acoustic signal associated with the listener's left ear and a fourth acoustic signal associated with the listener's right ear, and the processor is further configured to estimate the positions of the left ear and the right ear from the head position information, and to correct the sound pressure difference and phase difference between the third acoustic signal and the fourth acoustic signal based on the positional relationship between the estimated positions of the left ear and the right ear and the position of the sound source.

5. The acoustic signal generating device according to claim 1, wherein the processor is further configured to: correct a delay time for the first acoustic signal reproduced by the headphones as a direct sound until the direct sound reaches the listener; and correct an amplitude value for the second acoustic signal reproduced by the speakers as an indirect sound until the indirect sound reaches the listener.

6. The acoustic signal generating device according to claim 5, wherein the processor is further configured to: acquire, when the same sound source is reproduced by the headphones and the speaker, a first signal obtained by picking up sound from the headphones with a microphone placed at the position of the listener, and a second signal obtained by picking up sound from the speaker with the microphone; and calculate, based on the first signal and the second signal, a first correction coefficient used to correct the delay time and a second correction coefficient used to correct the amplitude value.

7. A method for generating an acoustic signal that provides a listener with sound based on a sound source using headphones worn by the listener and speakers arranged around the listener, the method comprising: acquiring head position information of the listener; generating a first acoustic signal and a second acoustic signal that localize a sound image at the position of the sound source based on the head position information and the position of the sound source; and outputting the generated first acoustic signal and second acoustic signal to the headphones and the speakers, respectively.

8. An acoustic signal generation program that provides a listener with sound based on a sound source using headphones worn by the listener and speakers arranged around the listener, the acoustic signal generation program causing a computer to execute the following steps: acquire head position information of the listener; generate a first acoustic signal and a second acoustic signal that localize a sound image at the position of the sound source based on the head position information and the position of the sound source; and output the generated first acoustic signal and second acoustic signal to the headphones and the speakers, respectively.

Citation Information

Patent Citations

  • Head phone with functions for detecting rotating angle

    JP1996009489A

  • Audio device and audio processing method

    JP2022502886A

  • Sound output device, sound generation method, and program

    WO2017061218A1