Near-field audio rendering

By setting up a virtual speaker array and HRTF filter on a wearable head device, the high computational cost of near-field audio effects in augmented reality and mixed reality systems is solved, achieving efficient near-field audio rendering and improving the realism of the audio experience.

CN116320907BActive Publication Date: 2026-02-06MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310249063.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-01
Filing Date
2019-10-04
Publication Date
2026-02-06
Estimated Expiration
2039-10-04

AI Technical Summary

Technical Problem

In existing technologies for simulating near-field audio effects in augmented reality and mixed reality systems, the computational cost is high, making it difficult to efficiently use far-field HRTF to model near-field audio effects.

Method used

By setting up a virtual speaker array on a wearable head device, the audio signal is processed using the head correlation transfer function and source radiation filter to generate output audio signals in the user's left and right ears, and near-field audio rendering is performed by combining the virtual speaker array and HRTF filter.

Benefits of technology

It enables efficient simulation of near-field audio effects in augmented reality and mixed reality systems, improving the realism of the user's audio experience while reducing computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320907B_ABST
    Figure CN116320907B_ABST
Patent Text Reader

Abstract

A near-field audio rendering. According to an example method, a source position corresponding to an audio signal is identified. An axis of sound corresponding to the audio signal is determined. For each of respective left and right ears of a user, an angle between the axis of sound and the respective ear is determined, a virtual loudspeaker position co-linear with the source position and a position of the respective ear is determined, wherein the virtual loudspeaker position lies on a surface of a sphere concentric with a head of the user, the sphere having a first radius. A head-related transfer function (HRTF) corresponding to the virtual loudspeaker position and corresponding to the respective ear is determined; a source radiation filter is determined based on the determined angle; the audio signal is processed to generate an output audio signal for the respective ear; and the output audio signal is presented to the respective ear of the user via one or more loudspeakers associated with a wearable head device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a continuation of application with a filing date of October 4, 2019, PCT International Application No. PCT / US2019 / 054893, Chinese National Stage Application No. 201980080065.2, entitled “Near-Field Audio Rendering,” the contents of which are incorporated herein by reference in their entirety.

[0002] REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of priority to U.S. Provisional Application No. 62 / 741,677, filed October 5, 2018, and U.S. Provisional Application No. 62 / 812,734, filed March 1, 2019, the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0004] The present disclosure relates generally to systems and methods for audio signal processing, and in particular, to systems and methods for rendering audio signals in a mixed reality environment. BACKGROUND

[0005] Augmented reality and mixed reality systems place unique demands on the rendering of binaural audio signals to a user. On one hand, rendering audio signals in a realistic manner - e.g., in a manner consistent with the user’s expectations - is crucial to creating an immersive and believable augmented or mixed reality environment. On the other hand, the computational expense of processing such audio signals can be costly, particularly for mobile systems that can have limited processing capabilities and battery capacity.

[0006] One particular challenge is the modeling of near-field audio effects. Near-field effects are important to recreate the impression of sound sources that are in close proximity to the user’s head. Near-field effects can be computed using a database of head-related transfer functions (HRTFs). However, typical HRTF databases include HRTFs measured at a single distance in the far field from the user’s head (e.g., more than 1 meter from the user’s head), and can lack HRTFs at distances appropriate for near-field effects. Even if an HRTF database includes measured or simulated HRTFs for different distances from the user’s head (e.g., less than 1 meter from the user’s head), directly using a large number of HRTFs for real-time audio rendering applications can be computationally expensive. Thus, systems and methods that model near-field audio effects using far-field HRTFs in a computationally efficient manner are desirable. SUMMARY

[0007] Examples of the present disclosure describe systems and methods for presenting an audio signal to a user of a wearable head device. According to an example method, a source position corresponding to the audio signal is identified. An axis of sound corresponding to the audio signal is determined. For each of a respective left ear and a right ear of the user, an angle between the axis of sound and the respective ear is determined. For each of the respective left ear and the right ear of the user, a virtual loudspeaker position in a virtual loudspeaker array is determined that is collinear with the source position and a position of the respective ear. The virtual loudspeaker array includes a plurality of virtual loudspeaker positions, each of the plurality of virtual loudspeaker positions being located on a surface of a sphere concentric with a head of the user, the sphere having a first radius. For each of the respective left ear and the right ear of the user, a head-related transfer function (HRTF) corresponding to the virtual loudspeaker position and corresponding to the respective ear is determined; a source radiation filter is determined based on the determined angle; the audio signal is processed to generate an output audio signal for the respective ear; and the output audio signal is presented to the respective ear of the user via one or more loudspeakers associated with the wearable head device. Processing the audio signal includes applying the HRTF and the source radiation filter to the audio signal. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 An example wearable system according to some embodiments of the present disclosure is shown.

[0009] Figure 2 An example handheld controller that can be used in conjunction with the example wearable system according to some embodiments of the present disclosure is shown.

[0010] Figure 3 An example auxiliary unit that can be used in conjunction with the example wearable system according to some embodiments of the present disclosure is shown.

[0011] Figure 4 An example functional block diagram for the example wearable system according to some embodiments of the present disclosure is shown.

[0012] Figure 5 A binaural rendering system according to some embodiments of the present disclosure is shown.

[0013] Figures 6A-6C An example geometry modeling audio effects from a virtual sound source according to some embodiments of the present disclosure is shown.

[0014] Figure 7 An example of calculating a distance traveled by a sound emitted by a point sound source according to some embodiments of the present disclosure is shown.

[0015] Figures 8A-8C An example of a sound source relative to a listener's ear according to some embodiments of the present disclosure is shown.

[0016] Figures 9A-9B An example head-related transfer function (HRTF) magnitude response is shown in accordance with some embodiments of the disclosure.

[0017] Figure 10 A source radiation angle of an acoustic source relative to an acoustic axis of a user is shown in accordance with some embodiments of the disclosure.

[0018] Figure 11 An example of an acoustic source panning inside a user's head is shown in accordance with some embodiments of the disclosure.

[0019] Figure 12 An example signal flow that can be implemented to render an acoustic source in a far field is shown in accordance with some embodiments of the disclosure.

[0020] Figure 13 An example signal flow that can be implemented to render an acoustic source in a near field is shown in accordance with some embodiments of the disclosure.

[0021] Figure 14 An example signal flow that can be implemented to render an acoustic source in a near field is shown in accordance with some embodiments of the disclosure.

[0022] Figures 15A-15D An example of a head coordinate system corresponding to a user and a device coordinate system corresponding to a device is shown in accordance with some embodiments of the disclosure. DETAILED DESCRIPTION

[0023] In the following description of examples, reference is made to the accompanying drawings which form a part hereof, and in which are shown by way of illustration specific examples that can be practiced. It is to be understood that other examples can be used and structural changes can be made without departing from the scope of the disclosed examples.

[0024] Example wearable system

[0025] Figure 1An example wearable head device 100 is shown, configured to be worn on a user’s head. The wearable head device 100 can be part of a more extensive wearable system that includes one or more components, such as a head device (e.g., the wearable head device 100), a handheld controller (e.g., the handheld controller 200 described below), and / or an auxiliary unit (e.g., the auxiliary unit 300 described below). In some examples, the wearable head device 100 can be used in a virtual reality, augmented reality, or mixed reality system or application. The wearable head device 100 can include one or more displays, such as displays 110A and 110B (which can include left and right transmissive displays and associated components for coupling light from the displays to the user’s eyes, such as orthogonal pupil expansion (OPE) grating sets 112A / 112B and exit pupil expansion (EPE) grating sets 114A / 114B); left and right acoustic structures, such as speakers 120A and 120B (which can be mounted on temples 122A and 122B and positioned near the user’s left and right ears, respectively); one or more sensors, such as infrared sensors, accelerometers, GPS units, inertial measurement units (IMUs, e.g., IMU 126), acoustic sensors (e.g., microphones 150); orthogonal coil electromagnetic receivers (e.g., receiver 127 shown mounted to the left temple 122A); left and right cameras oriented away from the user (e.g., depth (time-of-flight) cameras 130A and 130B); and left and right cameras oriented toward the user (e.g., for detecting the user’s eye motion) (e.g., eye cameras 128A and 128B). However, the wearable head device 100 can incorporate any suitable display technology, as well as any suitable number, type, or combination of sensors or other components, without departing from the scope of the present disclosure. In some examples, the wearable head device 100 can incorporate one or more microphones 150 configured to detect audio signals produced by the user’s voice; such microphones can be placed adjacent to the user’s mouth. In some examples, the wearable head device 100 can incorporate networking features (e.g., Wi-Fi functionality) to communicate with other devices and systems, including other wearable systems. The wearable head device 100 can also include components such as a battery, a processor, a memory, a storage unit, or various input devices (e.g., buttons, touchpads); or can be coupled to a handheld controller (e.g., the handheld controller 200) or an auxiliary unit (e.g., the auxiliary unit 300) that includes one or more such components. In some examples, the sensors can be configured to output a set of coordinates of the headset relative to the user’s environment, and can provide input to a processor performing a simultaneous localization and mapping (SLAM) process and / or a visual odometry algorithm.In some examples, the wearable head device 100 can be coupled to a handheld controller 200 and / or an auxiliary unit 300, as further described below.

[0026] Figure 2 An example mobile handheld controller assembly 200 of an example wearable system is shown. In some examples, the handheld controller 200 can be in wired or wireless communication with the wearable head device 100 and / or the auxiliary unit 300 described below. In some examples, the handheld controller 200 includes a handle portion 220 to be held by a user and one or more buttons 240 disposed along a top surface 210. In some examples, the handheld controller 200 can be configured to act as an optical tracking target; for example, a sensor (e.g., a camera or other optical sensor) of the wearable head device 100 can be configured to detect a position and / or orientation of the handheld controller 200— by extension, this can be indicative of a position and / or orientation of a hand of a user holding the handheld controller 200. In some examples, the handheld controller 200 can include a processor, a memory, a storage unit, a display, or one or more input devices, such as described above. In some examples, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 100). In some examples, the sensors can detect a position or orientation of the handheld controller 200 relative to the wearable head device 100 or relative to another component of the wearable system. In some examples, the sensors can be positioned in the handle portion 220 of the handheld controller 200 and / or can be mechanically coupled to the handheld controller. The handheld controller 200 can be configured to provide one or more output signals, such as a signal corresponding to a depressed state of a button 240; or a position, orientation, and / or motion (e.g., via an IMU) of the handheld controller 200. Such output signals can be used as inputs to a processor of the wearable head device 100, the auxiliary unit 300, or another component of the wearable system. In some examples, the handheld controller 200 can include one or more microphones to detect sound (e.g., a user’s voice, ambient sound) and, in some cases, to provide a signal corresponding to the detected sound to a processor (e.g., a processor of the wearable head device 100).

[0027] Figure 3An example auxiliary unit 300 of an example wearable system is shown. In some examples, the auxiliary unit 300 can be in wired or wireless communication with the wearable head device 100 and / or the handheld controller 200. The auxiliary unit 300 can include a battery to provide energy to operate one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including displays, sensors, acoustic structures, processors, microphones, and / or other components of the wearable head device 100 or the handheld controller 200). In some examples, the auxiliary unit 300 can include a processor, memory, storage unit, display, one or more input devices, and / or one or more sensors, such as described above. In some examples, the auxiliary unit 300 includes a clip 310 for attaching the auxiliary unit to a user (e.g., a belt worn by the user). An advantage of using the auxiliary unit 300 to house one or more components of the wearable system is that doing so can allow large or heavy components to be carried on the user's waist, chest, or back— which are relatively well suited to support large and heavy objects— rather than being mounted to the user's head (e.g., if housed in the wearable head device 100) or carried by the user's hands (e.g., if housed in the handheld controller 200). This can be particularly advantageous for relatively heavy or bulky components, such as batteries.

[0028] Figure 4 An example functional block diagram is shown that can correspond to an example wearable system 400, such as can include the example wearable head device 100, handheld controller 200, and auxiliary unit 300 described above. In some examples, the wearable system 400 can be used for virtual reality, augmented reality, or mixed reality applications. As Figure 4As shown, wearable system 400 can include an example handheld controller 400B, referred to herein as a "totem" (and which can correspond to handheld controller 200 described above); handheld controller 400B can include a totem-to-headgear six degrees of freedom (6DOF) totem subsystem 404A. Wearable system 400 can also include an example headgear device 400A (which can correspond to wearable head device 100 described above); headgear device 400A includes a totem-to-headgear 6DOF headgear subsystem 404B. In this example, 6DOF totem subsystem 404A and 6DOF headgear subsystem 404B cooperate to determine six coordinates (e.g., three translational offsets and rotations along three axes) of handheld controller 400B relative to headgear device 400A. The six degrees of freedom can be expressed relative to a coordinate system of headgear device 400A. The three translational offsets can be expressed as X, Y, and Z offsets in such a coordinate system, can be expressed as a translation matrix, or can be expressed as some other representation. The rotational degrees of freedom can be expressed as a sequence of yaw, pitch, and roll rotations; as a vector; as a rotation matrix; as a quaternion; or as some other representation. In some examples, one or more depth cameras 444 (and / or one or more non-depth cameras) and / or one or more optical targets included in headgear device 400A (e.g., buttons 240 of handheld controller 200 described above or dedicated optical targets included in handheld controller) can be used for 6DOF tracking. In some examples, as described above, handheld controller 400B can include a camera; and headgear device 400A can include optical targets for optical tracking in conjunction with the camera. In some examples, headgear device 400A and handheld controller 400B each include a set of three orthogonally oriented solenoids for wirelessly transmitting and receiving three distinguishable signals. By measuring the relative amplitudes of the three distinguishable signals received in each coil for reception, the 6DOF of handheld controller 400B relative to headgear device 400A can be determined. In some examples, 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that can be used to provide improved accuracy and / or more timely information about rapid motion of handheld controller 400B.

[0029] In some examples involving augmented reality or mixed reality applications, it can be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the head device apparatus 400A) to an inertial coordinate space or an environmental coordinate system coordinate space. For example, such a transformation can be necessary for a display of the head device apparatus 400A to present a virtual object at an intended position and orientation relative to the real environment, rather than at a fixed position and orientation on the display (e.g., the same position in the display of the head device apparatus 400A). This can maintain the illusion that the virtual object is present in the real environment (and, for example, does not appear to be positioned unnaturally in the real environment as the head device apparatus 400A is moved and rotated). In some examples, a compensating transformation between coordinate spaces can be determined by processing images from the depth camera 444 (e.g., using a Simultaneous Localization and Mapping (SLAM) and / or visual odometry process) in order to determine a transformation of the head device apparatus 400A relative to an inertial or environmental coordinate system. In Figure 4 In the illustrated example, the depth camera 444 can be coupled to the SLAM / visual odometry block 406, and can provide images to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the images and determine a position and orientation of the user's head, which can then be used to identify a transformation between the head coordinate space and the actual coordinate space. Similarly, in some examples, an additional source of information about the user's head pose and position is obtained from the IMU 409 of the head device apparatus 400A. Information from the IMU 409 can be integrated with information from the SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information about rapid adjustments of the user's head pose and position.

[0030] In some examples, the depth camera 444 can provide 3D images to a hand gesture tracker 411, which can be implemented in a processor of the wearable head apparatus 400A. The hand gesture tracker 411 can identify a hand gesture of the user, for example, by matching the 3D images received from the depth camera 444 to stored patterns representative of hand gestures. Other suitable techniques for identifying hand gestures of a user will be apparent.

[0031] In some examples, one or more processors 416 can be configured to receive data from head device subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, microphone 450, and / or hand gesture tracker 411. Processors 416 can also send and receive control signals from 6DOF totem system 404A. Processors 416 can be wirelessly coupled to 6DOF totem system 404A, for example in examples where handheld controller 400B is not tethered. Processors 416 can further communicate with additional components such as audio-visual content memory 418, graphics processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. DSP audio spatializer 422 can be coupled to head related transfer function (HRTF) memory 425. GPU 420 can include a left channel output coupled to a left source 424 of an imagewise light modulator and a right channel output coupled to a right source 426 of the imagewise light modulator. GPU 420 can output stereoscopic image data to sources 424, 426 of the imagewise light modulator. DSP audio spatializer 422 can output audio to left speaker 412 and / or right speaker 414. DSP audio spatializer 422 can receive input from processor 419 indicating a direction vector of a vector from a user to a virtual sound source (which can be moved by the user, for example via handheld controller 400B). Based on the direction vector, DSP audio spatializer 422 can determine a corresponding HRTF (for example, by accessing an HRTF or by interpolating a plurality of HRTFs). DSP audio spatializer 422 can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. By incorporating the relative position and orientation of the user with respect to a virtual sound in a mixed reality environment— that is, by presenting a virtual sound that matches the user's expectation of what the virtual sound would sound like if it were a real sound in the real environment— the believability and realism of the virtual sound can be enhanced.

[0032] In some examples, such as Figure 4 As shown in FIG. 4C, one or more of processors 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 can be included in auxiliary unit 400C (which can correspond to auxiliary unit 300 described above). Auxiliary unit 400C can include battery 427 to power its components and / or to power head device apparatus 400A and / or handheld controller 400B. Including such components in an auxiliary unit that can be mounted to a user's waist can limit the size and weight of head device apparatus 400A, which in turn can reduce fatigue of a user's head and neck.

[0033] While Figure 4Elements corresponding to various components of the example wearable system 400 are presented, but various other suitable arrangements of these components will become apparent to those skilled in the art. For example, Figure 4 Elements associated with the auxiliary unit 400C presented in the middle can instead be associated with the head device apparatus 400A or the handheld controller 400B. Furthermore, some wearable systems can forgo the handheld controller 400B or the auxiliary unit 400C altogether. Such changes and modifications should be understood as being included in the scope of the disclosed examples.

[0034] Audio rendering

[0035] The systems and methods described below can be implemented in an augmented reality or mixed reality system, such as the one described above. For example, one or more processors (e.g., CPUs, DSPs) of the augmented reality system can be used to process audio signals or implement the steps of the computer-implemented methods described below; sensors (e.g., cameras, acoustic sensors, IMUs, LIDAR, GPS) of the augmented reality system can be used to determine the position and / or orientation of elements in the user’s environment or the user of the system; and speakers of the augmented reality system can be used to present audio signals to the user. In some embodiments, external audio playback devices (e.g., headphones, earbuds) can be used instead of the speakers of the system to deliver audio signals to the user’s ears.

[0036] In an augmented reality or mixed reality system as described above, one or more processors (e.g., DSP audio spatializer 422) can process one or more audio signals for presentation to a user of a wearable head device via one or more speakers (e.g., left speaker 412 and right speaker 414 described above). The processing of the audio signals needs to strike a balance between the realism of the perceived audio signals - e.g., the degree to which the audio signals presented to the user in a mixed reality environment match the user’s expectations of how the audio signals would sound in a real environment - and the computational overhead involved in processing the audio signals.

[0037] Modeling near-field audio effects can improve the realism of the user’s audio experience, but can be computationally expensive. In some embodiments, an integrated solution can combine a computationally efficient rendering method with one or more near-field effects for each ear. The one or more near-field effects for each ear can include, for example, a parallax angle in the simulation of sound incidence for each ear, an interaural time difference (ITD) based on object position and anthropometric data, a near-field level change due to distance, and / or an amplitude response change due to proximity to the user’s head and / or a source radiation change due to parallax angle. In some embodiments, the integrated solution can be computationally efficient so as not to increase computational cost excessively.

[0038] In the far field, as the sound source moves closer or further away from the user, the change at the user's ears can be the same for each ear and can be an attenuation of the signal for the sound source. In the near field, as the sound source moves closer or further away from the user, the change at the user's ears can be different for each ear and can be more than just an attenuation of the signal for the sound source. In some embodiments, the near field and far field boundary can be the location where the conditions change.

[0039] In some embodiments, a virtual speaker array (VSA) can be a set of discrete locations on a sphere centered at the center of the user's head. For each location on the sphere, a pair (e.g., a left-right pair) of HRTFs is provided. In some embodiments, the near field can be the region inside the VSA and the far field can be the region outside the VSA. At the VSA, either the near field method or the far field method can be used.

[0040] The distance from the center of the user's head to the VSA can be the distance at which the HRTF is obtained. For example, the HRTF filter can be measured or synthesized from a simulation. The measured / simulated distance from the VSA to the center of the user's head can be referred to as the "measured distance" (MD). The distance from the virtual sound source to the center of the user's head can be referred to as the "source distance" (SD).

[0041] Figure 5 A binaural rendering system 500 according to some embodiments is shown. In Figure 5 In the example system, a single input audio signal 501 (which can represent a virtual sound source) is split into a left signal 504 and a right signal 506 by an interaural time delay (ITD) module 502 of an encoder 503. In some examples, the left signal 504 and the right signal 506 can differ by an ITD (e.g., in milliseconds) determined by the ITD module 502. In this example, the left signal 504 is input to a left ear VSA module 510 and the right signal 506 is input to a right ear VSA module 520.

[0042] In this example, the left ear VSA module 510 can pan the left signal 504 over a set of N channels, which respectively feed a set of left ear HRTF filters 550 (L1,... L N ) in a set of left ear HRTF filters 540. The left ear HRTF filters 550 can be substantially delayless. The panning gains 512 (g L1 ,... g LN ) of the left ear VSA module can be the left incident angle (ang LThe left incident angle can indicate an incident direction of sound relative to a front-facing direction from a center of a head of a user. Although shown from a top-down perspective relative to a head of a user in the figure, the left incident angle can include an angle in three dimensions; that is, the left incident angle can include an azimuth angle and / or an elevation angle.

[0043] Similarly, in this example, the right ear VSA module 520 can pan the right signal 506 over a set of M channels that are respectively fed into a set of right ear HRTF filters 560 (R1,... R M ) of the HRTF filter set 540. The right ear HRTF filters 550 can be substantially delay-free. (Although only one HRTF filter set is shown in the figure, multiple HRTF filter sets including HRTF filters stored across a distributed system can be contemplated.) The panning gains 522 (g R1 ,... g RM ) of the right ear VSA module can be a function of the right incident angle (ang R ). The right incident angle can indicate an incident direction of sound relative to a front-facing direction from a center of a head of a user. As noted above, the right incident angle can include an angle in three dimensions; that is, the right incident angle can include an azimuth angle and / or an elevation angle.

[0044] In some embodiments, as shown in the figure, the left ear VSA module 510 can pan the left signal 504 over N channels, and the right ear VSA module can pan the right signal over M channels. In some embodiments, N and M can be equal. In some embodiments, N and M can be different. In these embodiments, the left ear VSA module can be fed into a set of left ear HRTF filters (L1,... L N ), and the right ear VSA module can be fed into a set of right ear HRTF filters (R1,... R M ), as described above. Furthermore, in these embodiments, the panning gains (g L1 ,... g LN ) of the left ear VSA module can be a function of the left incident angle (ang L ), and the panning gains (g R1 ,... g RM ) of the right ear VSA module can be a function of the right incident angle (ang R ), as described above.

[0045] The example system shows a single encoder 503 and a corresponding input signal 501. The input signal can correspond to a virtual sound source. In some embodiments, the system can include additional encoders and corresponding input signals. In these embodiments, the input signals can correspond to virtual sound sources. That is, each input signal can correspond to a virtual sound source.

[0046] In some embodiments, when rendering multiple virtual sound sources simultaneously, the system may include an encoder for each virtual sound source. In these embodiments, the mixing module (e.g., Figure 5 (530) receives output from each of the encoders, mixes the received signals, and outputs the mixed signal to the left and right HRTF filters of the HRTF filter bank.

[0047] Figure 6A The diagram illustrates geometry for modeling audio effects from a virtual sound source, according to some embodiments. The distance 630 from the virtual sound source 610 to the center 620 of the user's head (e.g., "source distance" (SD)) is equal to the distance 640 from the VSA 650 to the center of the user's head (e.g., "measurement distance" (MD)). Figure 6A As shown, the left incident angle is 65° (ang). L ) and right incident angle 654 (ang R Equal to. In some embodiments, the angle from the center 620 of the user's head to the virtual sound source 610 can be directly used to calculate the pan gain (e.g., g). L1 ,…,g LN ,g R1 ,…,g RN In the example shown, the virtual sound source position 610 is used as the position (612 / 614) for calculating the left ear tilt and right ear tilt.

[0048] Figure 6B The diagram illustrates geometry for modeling near-field audio effects from a virtual sound source, according to some embodiments. As shown, the distance 630 from the virtual sound source 610 to a reference point (e.g., "source distance" (SD)) is less than the distance 640 from the VSA 650 to the center 620 of the user's head (e.g., "measurement distance" (MD)). In some embodiments, the reference point may be the center (620) of the user's head. In some embodiments, the reference point may be the midpoint between the user's two ears. Figure 6B As shown, the left incident angle is 65° (ang). L () greater than the right incident angle 654 (ang) R ). The angle relative to each ear (e.g., left angle of incidence 65° (ang) L ) and right incident angle 654 (ang R This is different from the one at MD 640.

[0049] In some embodiments, the left incident angle 65° (ang) used to calculate the left ear signal shift is used to calculate the left ear signal shift. L This can be derived by calculating the intersection of a line passing through the virtual sound source 610 from the user's left ear with the sphere containing the VSA 650.

[0050] Similarly, in some embodiments, the right incident angle 654 (ang L ) for computing the left ear signal yaw can be derived by computing the intersection of a line passing through the location of the virtual sound source 610 from the user's right ear with a sphere containing the VSA 650. The yaw angle combination (azimuth and elevation) can be computed for the 3D environment as the spherical coordinate angles from the center 620 of the user's head to the intersection point.

[0051] In some embodiments, the intersection between the line and the sphere can be computed, e.g., by combining the equation representing the line and the equation representing the sphere.

[0052] Figure 6C Geometry for modeling the far-field audio effect from a virtual sound source is shown, according to some embodiments. The distance 630 (e.g., "source distance" (SD)) from the virtual sound source 610 to the center 620 of the user's head is greater than the distance 640 (e.g., "measurement distance" (MD)) from the VSA 650 to the center 620 of the user's head. As shown, the left incident angle 612 (ang Figure 6C L ) is less than the right incident angle 614 (ang R ). The angles relative to each ear (e.g., the left incident angle (ang L ) and the right incident angle (ang R )) are different at the MD.

[0053] In some embodiments, the left incident angle 612 (ang L ) for computing the left ear signal yaw can be derived by computing the intersection of a line passing through the location of the virtual sound source 610 from the user's left ear with a sphere containing the VSA 650.

[0054] Similarly, in some embodiments, the right incident angle 614 (ang R ) for computing the left ear signal yaw can be derived by computing the intersection of a line passing through the location of the virtual sound source 610 from the user's right ear with a sphere containing the VSA 650. The yaw angle combination (azimuth and elevation) can be computed for the 3D environment as the spherical coordinate angles from the center 620 of the user's head to the intersection point.

[0055] In some embodiments, the intersection between the line and the sphere can be computed, e.g., by combining the equation representing the line and the equation representing the sphere.

[0056] In some embodiments, the rendering scheme can not distinguish between the left incident angle 612 and the right incident angle 614, but rather assume that the left incident angle 612 and the right incident angle 614 are equal. However, in reproducing the effect as described with respect to Figure 6B ​The described near-field effect and / or as regarding the Figure 6C When the described far-field effect, it can not be applicable or acceptable to assume that the left incident angle 612 and the right incident angle 614 are equal.

[0057] Figure 7 A geometric model for calculating the distance traveled by the sound emitted by the (point) sound source 710 to the user’s ear 712 is shown, in accordance with some embodiments. In Figure 7 In the shown geometric model, the user’s head is assumed to be spherical. The same model is applied to each ear (e.g., the left ear and the right ear). The delay to each ear can be calculated by dividing the distance traveled by the sound emitted by the (point) sound source 710 to each ear (e.g., the distance A+B in Figure 7 The interaural time difference (ITD) can be the difference in delay between the user’s two ears. In some embodiments, the ITD can be applied only to the contralateral ear with respect to the location of the user’s head and the sound source 710. In some embodiments, the ITD can be applied to both ears. Figure 7 The geometric model shown in can be used for any SD (e.g., near-field or far-field), and can not take into account the position of the ears on the user’s head and / or the head size of the user’s head.

[0058] In some embodiments, Figure 7 The geometric model shown in can be used to calculate the attenuation due to the distance from the sound source 710 to each ear. In some embodiments, the ratio of the distances can be used to calculate the attenuation. The level difference with respect to a near-field source can be calculated by evaluating the ratio of the source-ear distance for the desired source location to the source-ear distance for the source corresponding to the angle calculated for the panning (e.g., as shown in Figures 6A-6C In some embodiments, the minimum distance from the ear can be used, for example, to avoid dividing by a very small number, which can be computationally expensive and / or cause numerical overflow. In these embodiments, the smaller distances can be clamped.

[0059] In some embodiments, the distances can be clamped. For example, clamping can include, for example, limiting distance values below a threshold to another value. In some embodiments, clamping can include using a finite distance value (referred to as a clamped distance value) instead of the actual distance value for the calculation. A hard clamp can include limiting distance values below a threshold to the threshold. For example, if the threshold is 5 millimeters, distance values less than the threshold will be set to the threshold, and the threshold (instead of the actual distance values less than the threshold) can be used for the calculation. A soft clamp can include limiting distance values such that they asymptotically approach the threshold when the distance values are close to or below the threshold. In some embodiments, instead of or in addition to clamping, distance values can be increased by a predetermined amount such that the distance values are never less than the predetermined amount.

[0060] In some embodiments, a first minimum distance from the listener's ear can be used to compute the gain, and a second minimum distance from the listener's ear can be used to compute other sound source position parameters, e.g., to compute the angle of the HRTF filter, the inter-aural time difference, etc. In some embodiments, the first minimum distance and the second minimum distance can be different.

[0061] In some embodiments, the minimum distance used to compute the gain can be a function of one or more properties of the sound source. In some embodiments, the minimum distance used to compute the gain can be a function of the level of the sound source (e.g., the RMS value of the signal over multiple frames), the size of the sound source, or the radiation characteristics of the sound source, etc.

[0062] Figures 8A-8C An example of a sound source relative to the listener's right ear is shown, according to some embodiments. Figure 8A A case is shown where the sound source 810 is at a distance 812 from the listener's right ear 820 that is greater than a first minimum distance 822 and a second minimum distance 824. In this embodiment, the distance 812 between the simulated sound source and the listener's right ear 820 is used to compute the gain and other sound source position parameters, and is not clamped.

[0063] Figure 8B A case is shown where the simulated sound source 810 is at a distance 812 from the listener's right ear 820 that is less than the first minimum distance 822 and greater than the second minimum distance 824. In this embodiment, the distance 812 is clamped for gain computation, but not for computing other parameters, e.g., azimuth and elevation angles or inter-aural time difference. In other words, the first minimum distance 822 is used to compute the gain, and the distance 812 between the simulated sound source 810 and the listener's right ear 820 is used to compute other sound source position parameters.

[0064] Figure 8C A case is shown where the simulated sound source 810 is closer to the ear than both the first minimum distance 822 and the second minimum distance 824. In this embodiment, the distance 812 is clamped for gain computation and for computing other sound source position parameters. In other words, the first minimum distance 822 is used to compute the gain, and the second minimum distance 824 is used to compute other sound source position parameters.

[0065] In some embodiments, instead of limiting the minimum distance used to compute the gain, the gain computed from the distance can be directly limited. In other words, the gain can be computed based on the distance as a first step, and in a second step, the gain can be clamped to not exceed a predetermined threshold.

[0066] In some embodiments, the amplitude response of a sound source can change when the sound source is closer to the listener’s head. For example, the low frequencies at the ipsilateral ear can be amplified and / or the high frequencies at the contralateral ear can be attenuated when the sound source is closer to the listener’s head. The change in amplitude response can result in a change in the interaural level difference (ILD).

[0067] Figure 9A and 9B HRTF amplitude responses at the ears 900A and 900B with respect to a (point) sound source in a horizontal plane are shown, respectively, in accordance with some embodiments. The HRTF amplitude responses can be computed as a function of azimuth using a spherical head model. Figure 9A The amplitude response 900A with respect to a (point) sound source in a far field (e.g., 1 meter from the center of the user’s head) is shown. Figure 9B The amplitude response 900B with respect to a (point) sound source in a near field (e.g., 0.25 meters from the center of the user’s head) is shown. As Figure 9A and 9B shown, the change in ILD can be most significant at low frequencies. In the far field, the amplitude response of low frequency content can be constant (e.g., independent of the angle of the source azimuth). In the near field, the amplitude response of low frequency content can be amplified for sound sources on the same side of the user’s head / ears, which can result in a higher ILD at low frequencies. In the near field, the amplitude response of high frequency content can be attenuated for sound sources on the opposite side of the user’s head.

[0068] In some embodiments, the change in amplitude response can be taken into account by, for example, considering the HRTF filters in binaural rendering. In the case of VSA, the HRTF filters can be approximated as the HRTFs corresponding to the positions used to compute the right ear panning and the positions used to compute the left ear panning (e.g., as shown in Figure 6B and Figure 6C In some embodiments, the HRTF filters can be computed using the direct MD HRTF. In some embodiments, the HRTF filters can be computed using the panned spherical head model HRTF. In some embodiments, the compensation filters can be computed independent of the parallax HRTF angles.

[0069] In some embodiments, the parallax HRTF angles can be computed and then used to compute more accurate compensation filters using their angles. For example, with reference to Figure 6B the positions used to compute the left ear panning can be compared to the virtual sound source positions used to compute the synthesis filters for the left ear, and the positions used to compute the right ear panning can be compared to the virtual sound source positions used to compute the synthesis filters for the right ear.

[0070] In some embodiments, once the attenuation due to distance is taken into account, additional signal processing can be utilized to capture the amplitude differences. In some embodiments, the additional signal processing can consist of a gain, a low shelving filter, and a high shelving filter to be applied to each ear signal.

[0071] In some embodiments, the wideband gain can be calculated for angles up to 120 degrees, e.g., according to Equation 1:

[0072] gain db = 2.5 * sin(angleMD deg * 3 / 2) (Equation 1)

[0073] where angleMD deg can be the angle of the corresponding HRTF at, e.g., the MD relative to the position of the user's ear. In some embodiments, angles other than 120 degrees can be used. In these embodiments, Equation 1 can be modified according to each angle used.

[0074] In some embodiments, the wideband gain can be calculated for angles greater than 120 degrees, e.g., according to Equation 2:

[0075] gain db = 2.5 * sin(180 + 3 * (angleMD deg - 120)) (Equation 2)

[0076] In some embodiments, angles other than 120 degrees can be used. In these embodiments, Equation 2 can be modified according to each angle used.

[0077] In some embodiments, the low shelving filter gain can be calculated, e.g., according to Equation 3:

[0078] lowshelfgain db = 2.5 * (e -angleMD_deg / 65 - e -180 / 65 ) (Equation 3)

[0079] In some embodiments, other angles can be used. In these embodiments, Equation 3 can be modified according to each angle used.

[0080] In some embodiments, the high shelving filter gain can be calculated for angles greater than 110 degrees, e.g., according to Equation 4:

[0081] highshelfgain db = 3.3 * (cos((angle deg * 180 / pi - 110) * 3) - 1) (Equation 4)

[0082] where angle deg can be the angle of the source relative to the position of the user's ear. In some embodiments, angles other than 110 degrees can be used. In these embodiments, Equation 4 can be modified according to each angle used.

[0083] The effects described above (e.g., gain, low shelving filter, and high shelving filter) can be attenuated as a function of distance. In some embodiments, a distance attenuation factor can be calculated, for example, according to Equation 5:

[0084] distanceAttenuation = (HR / (HR-MD))*(l-MD / sourceDistance_clamped) (Equation 5)

[0085] where HR is the head radius, MD is the measured distance, and sourceDistance_clamped is the source distance clamped to be at least as large as the head radius.

[0086] Figure 10 The off-axis angle (or source radiation angle) of a user relative to the sound axis 1015 of a sound source 1010 is shown, according to some embodiments. In some embodiments, the source radiation angle can be used to evaluate the magnitude response of the direct path, for example, based on the source radiation characteristics. In some embodiments, the off-axis angle can be different for each ear as the source moves closer to the user's head. In this figure, source radiation angle 1020 corresponds to the left ear; source radiation angle 1030 corresponds to the center of the head; and source radiation angle 1040 corresponds to the right ear. The different off-axis angles for each ear can result in separate direct path processing for each ear.

[0087] Figure 11A sound source 1110 that is panned inside a user's head is shown in accordance with some embodiments. To produce the in-head effect, the sound source 1110 can be processed as a crossfade between binaural rendering and stereo rendering. In some embodiments, a binaural rendering can be created for a source 1112 that is located on or outside the user's head. In some embodiments, the location of the sound source 1112 can be defined as the intersection of a line that passes through the simulated sound location 1110 from the center 1120 of the user's head with the surface 1130 of the user's head. In some embodiments, a stereo rendering can be created using amplitude-based and / or time-based panning techniques. In some embodiments, time-based panning techniques can be used to time-align the stereo and binaural signals at each ear, e.g., by applying an ITD to the contralateral ear. In some embodiments, the ITD and ILD can be scaled to zero as the sound source approaches the center 1120 of the user's head (i.e., when the source distance 1150 approaches zero). In some embodiments, a crossfade between binaural and stereo can be computed, e.g., based on the SD, and this crossfade can be normalized by the approximate radius 1140 of the user's head.

[0088] In some embodiments, a filter (e.g., an EQ filter) can be applied to a sound source that is placed at the center of the user's head. The EQ filter can be used to reduce sudden timbre changes as the sound source moves through the user's head. In some embodiments, the EQ filter can be scaled to match the amplitude response at the surface of the user's head as the simulated sound source moves from the center of the user's head to the surface of the user's head, and thus further reduce the risk of sudden amplitude response changes as the sound source goes in and out of the user's head. In some embodiments, a crossfade between the equalized signal and the unprocessed signal can be used based on the location of the sound source between the center of the user's head and the surface of the user's head.

[0089] In some embodiments, the EQ filter can be automatically computed as an average of the filters used to render the source on the surface of the user's head. The EQ filter can be exposed to the user as a set of tunable / configurable parameters. In some embodiments, the tunable / configurable parameters can include control frequencies and associated gains.

[0090] Figure 12 A signal flow 1200 that can be implemented to render a sound source in a far field is shown in accordance with some embodiments. As shown, a sound source 1210 is received at a microphone 1202. The sound source 1210 can be a source that is located in the far field, e.g., a source that is located at a distance that is greater than the distance between the user's head and the microphone 1202. In some embodiments, the sound source 1210 can be a source that is located in the near field, e.g., a source that is located at a distance that is less than the distance between the user's head and the microphone 1202. In some embodiments, the sound source 1210 can be a source that is located in the far field and the near field, e.g., a source that is located at a distance that is between the distance between the user's head and the microphone 1202. In some embodiments, the sound source 1210 can be a source that is located in the far field and the near field, e.g., a source that is located at a distance that is between the distance between the user's head and the microphone 1202. Figure 12As shown, far-field distance attenuation 1220 can be applied to input signal 1210, as described above. A common EQ filter 1230 (e.g., a source radiation filter) can be applied to the results of modeling the sound source radiation; the output of filter 1230 can be split and sent to separate left and right channels, where delay (1240A / 1240B) and VSA (1250A / 1250B) functions are applied to each channel, as described above. Figure 5 As described, to generate left and right ear signals 1290A / 1290B.

[0091] Figure 13 A signal stream 1300, which can be implemented to render a sound source in the near field according to some embodiments, is shown. Figure 13 As shown, far-field distance attenuation 1320 can be applied to input signal 1310, as described above. The output can be split into left / right channels, and separate EQ filters can be applied to each ear (e.g., left-ear near-field and source radiation filter 1330A for the left ear and right-ear near-field and source radiation filter 1330B for the right ear) to model the source radiation and near-field ILD effects, as described above. After the left-ear and right-ear signals have been separated, filters can be implemented one for each ear. It should be noted that in this case, any additional EQ applied to both ears can be folded into those filters (e.g., left-ear near-field and source radiation filter and right-ear near-field and source radiation filter) to avoid additional processing. Then, delay (1340A / 1340B) and VSA (1350A / 1350B) functions can be applied to each channel, as described above. Figure 5 As described, to generate left and right ear signals 1390A / 1390B.

[0092] In some embodiments, to optimize computational resources, the system may automatically switch between signal streams 1200 and 1300, for example, based on whether the sound source to be rendered is in the far field or near field. In some embodiments, it may be necessary to replicate filter states between filters (e.g., source radiation filter, left ear near field and source radiation filter, and right ear near field and source radiation filter) during conversion to avoid processing artifacts.

[0093] In some embodiments, the EQ filter described above can be bypassed when the setting of the EQ filter is perceptibly equivalent to a flat amplitude response with 0 dB gain. If the response is flat but has a gain other than zero, a broadband gain can be used to effectively achieve the desired result.

[0094] Figure 14 A signal stream 1400, which can be implemented to render a sound source in the near field according to some embodiments, is shown. Figure 14As shown in FIG. 14B, far-field distance attenuation 1420 can be applied to the input signal 1410, e.g., as described above. A left-ear near-field and source radiation filter 1430 can be applied to the output. The output of 1430 can be split into left / right channels, and a second filter 1440 (e.g., a right-left ear near-field and source radiation difference filter) can then be used to process the right-ear signal. The second filter models the difference between the right-ear near-field and source radiation effects and the left-ear near-field and source radiation effects. In some embodiments, the difference filter can be applied to the left-ear signal. In some embodiments, the difference filter can be applied to the contralateral ear, which can depend on the location of the sound source. Delays (1450A / 1450B) and VSA (1460A / 1460B) functions can be applied to each channel, such as described above with respect to FIG. 14A, to produce left-ear and right-ear signals 1490A / 1490B. Figure 5

[0095] A head coordinate system can be used to compute the acoustic propagation from an audio object to the listener's ears. A device coordinate system can be used by a tracking device, such as one or more sensors of a wearable head device in an augmented reality system, such as described above, to track the position and orientation of the listener's head. In some embodiments, the head coordinate system and the device coordinate system can be different. The center of the listener's head can be used as the origin of the head coordinate system and can be used to reference the position of audio objects relative to the listener, with the forward direction of the head coordinate system being defined as traveling from the center of the listener's head to the horizontal line in front of the listener. In some embodiments, an arbitrary point in space can be used as the origin of the device coordinate system. In some embodiments, the origin of the device coordinate system can be a point located between the optical lenses of a visual projection system of the tracking device. In some embodiments, the forward direction of the device coordinate system can be referenced to the tracking device itself and depend on the position of the tracking device on the listener's head. In some embodiments, the tracking device can have a non-zero pitch (i.e., tilt up or down) relative to the horizontal plane of the head coordinate system, resulting in a misalignment between the forward direction of the head coordinate system and the forward direction of the device coordinate system.

[0096] ​In some embodiments, the difference between the head coordinate system and the device coordinate system can be compensated for by applying a translation to the position of the audio object relative to the listener's head. In some embodiments, the origin difference between the head coordinate system and the device coordinate system can be compensated for by translating the position of the audio object relative to the listener's head by an amount equal to the distance between the origin of the head coordinate system and the origin of the device coordinate system reference point in three dimensions (e.g., x, y, and z). In some embodiments, the angular difference between the head coordinate system axes and the device coordinate system axes can be compensated for by applying a rotation to the position of the audio object relative to the listener's head. For example, if the tracking device is tilted N degrees downward, the position of the audio object can be rotated N degrees downward before rendering the audio output for the listener. In some embodiments, the audio object rotation compensation can be applied before the audio object translation compensation. In some embodiments, the compensations (e.g., rotations, translations, scaling, etc.) can be done together in a single transformation that includes all of the compensations (e.g., rotations, translations, scaling, etc.).

[0097] Figures 15A-15D An example of a head coordinate system 1500 corresponding to a user and a device coordinate system 1510 corresponding to a device 1512, such as a head-mounted augmented reality device as described above, is shown in accordance with an embodiment. Figure 15A A top view showing an example of a front translation offset 1520 between the head coordinate system 1500 and the device coordinate system 1500 is shown. Figure 15B A top view showing an example of a front translation offset 1520 and a rotation 1530 about a vertical axis between the head coordinate system 1500 and the device coordinate system 1510 is shown. Figure 15C A side view showing an example of both a front translation offset 1520 and a vertical translation offset 1522 between the head coordinate system 1500 and the device coordinate system 1500 is shown. Figure 15D A side view showing an example of a front translation offset 1520 and a vertical translation offset 1522 between the head coordinate system 1500 and the device coordinate system 1510 and a rotation 1530 about a left / right horizontal axis is shown.

[0098] In some embodiments, such as those depicted in Figures 15A-15D In some embodiments, such as those depicted in

[0099] Various exemplary embodiments of the present disclosure are described herein. Reference is made to these examples in a non-limiting sense. The examples are provided to illustrate more broadly applicable aspects of the present disclosure. Various changes can be made to the disclosed embodiments and equivalents can be substituted without departing from the true spirit and scope of the disclosure. In addition, many modifications can be made to adapt a particular situation, material, composition of matter, process, process act(s) or step(s) to the objective(s), spirit or scope of the present disclosure. Further, as will be apparent to those of ordinary skill in the art, each of the individual variations described and illustrated herein has separate and distinct components and features that can be readily separated from or combined with features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. All such modifications are intended to be within the scope of claims associated with the present disclosure.

[0100] The present disclosure includes methods that can be performed using the subject devices. The methods can comprise acts of providing such suitable apparatus. Such providing can be performed by an end-user. In other words, the "providing" acts merely require the end-user to obtain, access, approach, position, set, activate, turn on, or otherwise provide the necessary apparatus in the method. The methods described herein can be performed in any order that is logically possible as well as in the order recited.

[0101] The exemplary aspects of the present disclosure have been set forth above along with details regarding material selection and manufacture. Additional details regarding aspects of the underlying methods according to the present disclosure can be appreciated in connection with the patents and publications referenced above as well as generally known or understood by those skilled in the art. The same can hold true with respect to additional acts associated with the aspects of the methods of the present disclosure as where logically utilized, in general or in a logical progression, to further appreciate and practice the present disclosure.

[0102] In addition, while a number of examples have been described in relation to the various features of the present disclosure, the disclosure is not limited to the described or illustrated examples, and the disclosure can include various modifications and equivalents. The described disclosure, in its broader sense, can be directed to any of the relevant devices, apparatus, methods, techniques or procedures that are appropriately adapted to the purposes of the present disclosure. Further, any and all

[0103] Additionally, it is contemplated that any optional feature of the variations described can be set forth and claimed independently, or in combination with any one or more of the features described herein. Reference to an item in the singular, includes the possibility of more than one such item, unless otherwise stated. More specifically, as used herein and throughout the claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. In other words, to any term that is described herein in the singular, plural forms are also contemplated unless the context clearly dictates otherwise. In other words, the recitation of elements in the singular is not intended to limit the claim to a single element, but rather to encompass more than one element unless the context clearly indicates otherwise. In other words, the use of the "or" in the claims is meant to include "and / or" unless specifically indicated otherwise, or the use of "and" is specifically indicated to mean "and / or" the use of "one of' preceding the term "comprising," "including," and "containing" is meant to encompass the presence of at least one, but not necessarily more than one, of the stated item or term.

[0104] Without the use of such exclusive terminology, the term "comprising" in claims associated with the present disclosure are used to mean that the claims encompass any additional elements that do not materially alter the basic and novel characteristics of the claimed composition, method, or article of manufacture. Except as specifically defined herein, all technical and scientific terms used herein are to be given as broad as possible their ordinary meanings in the art to which the present disclosure pertains.

[0105] The breadth of the present disclosure is not limited to the provided examples and / or subject matter specification, but is only defined by the scope of the claims associated with the present disclosure.

Claims

1. A method of presenting an audio signal to a user of a wearable head device, the method comprising: identifying a source position corresponding to the audio signal; determining an acoustic axis corresponding to the audio signal; determining a reference point; for each of respective left and right ears of the user: determining an angle between the acoustic axis and the respective ear; determining a virtual loudspeaker position in a virtual loudspeaker array that is collinear with the source position and a position of the respective ear, wherein the virtual loudspeaker array comprises a plurality of virtual loudspeaker positions, each of the plurality of virtual loudspeaker positions being located on a surface of a sphere concentric with the reference point, the sphere having a first radius; determining a head-related transfer function (HRTF) corresponding to the virtual loudspeaker position and to the respective ear; determining a source radiation filter based on the determined angle; processing the audio signal to generate an output audio signal for the respective ear, wherein processing the audio signal comprises applying the HRTF and the source radiation filter to the audio signal; attenuating the audio signal based on a distance between the source position and the respective ear, wherein the distance is clamped to a minimum value; and presenting the output audio signal to the respective ear of the user via one or more loudspeakers associated with the wearable head device, wherein determining the reference point comprises: determining a position of the wearable head device based on sensors of the wearable head device, and applying a transformation to a position of an audio object relative to a head of the user to compensate for a difference between a head coordinate system corresponding to the head of the user and a device coordinate system corresponding to the wearable head device.

2. The method of claim 1, wherein, the source position is separated from the reference point by a distance that is less than the first radius.

3. The method of claim 1, wherein, the source position is separated from the reference point by a distance that is greater than the first radius.

4. The method of claim 1, wherein, the source position is separated from the reference point by a distance that is equal to the first radius.

5. The method of claim 1, further comprising: an interaural time difference is applied to the audio signal.

6. The method of claim 1, wherein, determining the HRTF corresponding to the virtual loudspeaker position comprises selecting the HRTF from a plurality of HRTFs, wherein each HRTF of the plurality of HRTFs describes a relationship between a listener and an audio source that is separated from the listener by a distance equal to the first radius.

7. The method of claim 1, wherein, the wearable head device comprises the one or more loudspeakers.

8. A system for audio signal processing, comprising: a wearable head device; one or more loudspeakers; and one or more processors configured to perform a method comprising: identifying a source position corresponding to an audio signal; determining an acoustic axis corresponding to the audio signal; determining a reference point; for each of respective left and right ears of a user of the wearable head device: determining an angle between the acoustic axis and the respective ear; ​ determining a virtual loudspeaker position in a virtual loudspeaker array that is collinear with the source position and a position of the respective ear, wherein the virtual loudspeaker array comprises a plurality of virtual loudspeaker positions, each of the plurality of virtual loudspeaker positions being located on a surface of a sphere concentric to the reference point, the sphere having a first radius; determining a head-related transfer function, HRTF, corresponding to the virtual loudspeaker position and to the respective ear; determining a source radiation filter based on the determined angle; processing the audio signal to generate an output audio signal for the respective ear, wherein processing the audio signal comprises applying the HRTF and the source radiation filter to the audio signal; attenuating the audio signal based on a distance between the source position and the respective ear, wherein the distance is clamped to a minimum value; and presenting the output audio signal to the respective ear of the user via the one or more loudspeakers, wherein determining the reference point comprises: determining a position of the wearable head device based on sensors of the wearable head device, and applying a transformation to a position of an audio object relative to a head of the user to compensate for a difference between a head coordinate system corresponding to the head of the user and a device coordinate system corresponding to the wearable head device.

9. The system of claim 8, wherein, the source position is separated from the reference point by a distance that is less than the first radius.

10. The system of claim 8, wherein, the source position is separated from the reference point by a distance that is greater than the first radius.

11. The system of claim 8, wherein, the source position is separated from the reference point by a distance that is equal to the first radius.

12. The system of claim 8, wherein, the method further comprises applying an interaural time difference to the audio signal.

13. The system of claim 8, wherein, determining the HRTF corresponding to the virtual loudspeaker position comprises selecting the HRTF from a plurality of HRTFs, wherein each of the plurality of HRTFs describes a relationship between a listener and an audio source that is separated from the listener by a distance equal to the first radius.

14. The system of claim 8, wherein, the wearable head device comprises the one or more loudspeakers.

15. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method of presenting an audio signal to a user of a wearable head device, the method comprising: identifying a source position corresponding to the audio signal; determining an acoustic axis corresponding to the audio signal; determining a reference point; for each of respective left and right ears of the user: determining an angle between the acoustic axis and the respective ear; determining a virtual loudspeaker position in a virtual loudspeaker array that is collinear with the source position and a position of the respective ear, wherein the virtual loudspeaker array comprises a plurality of virtual loudspeaker positions, each of the plurality of virtual loudspeaker positions being located on a surface of a sphere concentric to the reference point, the sphere having a first radius; determining a head-related transfer function, HRTF, corresponding to the virtual loudspeaker position and to the respective ear; determining a source radiation filter based on the determined angle; processing the audio signal to generate an output audio signal for the respective ear, wherein processing the audio signal comprises applying the HRTF and the source radiation filter to the audio signal; attenuating the audio signal based on a distance between the source position and the respective ear, wherein the distance is clamped to a minimum value; and presenting the output audio signal to the respective ear of the user via the one or more loudspeakers, wherein determining the reference point comprises: determining a position of the wearable head device based on sensors of the wearable head device, and applying a transformation to a position of an audio object relative to a head of the user to compensate for a difference between a head coordinate system corresponding to the head of the user and a device coordinate system corresponding to the wearable head device. processing the audio signal to generate an output audio signal for the respective ear, wherein processing the audio signal includes applying the HRTF and the source radiation filter to the audio signal; attenuating the audio signal based on a distance between the source position and the respective ear, wherein the distance is clamped to a minimum value; and presenting the output audio signal to the respective ear of the user via one or more speakers associated with the wearable head device, wherein determining the reference point includes: determining a position of the wearable head device based on sensors of the wearable head device, and applying a transformation to a position of an audio object relative to a head of the user to compensate for a difference between a head coordinate system corresponding to the head of the user and a device coordinate system corresponding to the wearable head device.

16. The non-transitory computer-readable medium of claim 15, wherein, the source position is separated from the reference point by a distance that is less than the first radius.

17. The non-transitory computer-readable medium of claim 15, wherein, the source position is separated from the reference point by a distance that is greater than the first radius.

18. The non-transitory computer-readable medium of claim 15, wherein, the source position is separated from the reference point by a distance that is equal to the first radius.

19. The non-transitory computer-readable medium of claim 15, wherein, the method further includes applying an interaural time difference to the audio signal.

20. The non-transitory computer-readable medium of claim 15, wherein, determining the HRTF corresponding to the virtual speaker position includes selecting the HRTF from a plurality of HRTFs, wherein each HRTF of the plurality of HRTFs describes a relationship between a listener and an audio source that is separated from the listener by a distance equal to the first radius.

Citation Information

Patent Citations

  • Sound image localization processor, Method, and program

    US20100080396A1

  • Decoupled binaural rendering

    US9992602B1