Interaural time-delay crossfader for binaural audio rendering

The method of applying delays, gain adjustments, and HRTF to audio signals in wearable headsets with crossfading techniques addresses the challenge of accurately presenting interaural time differences, reducing sound artifacts, and enhancing the immersive experience in mixed reality environments.

JP7776332B2Active Publication Date: 2025-11-26MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021518557
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-01
Filing Date
2019-10-04
Publication Date
2025-11-26
Estimated Expiration
2039-10-04

AI Technical Summary

Technical Problem

Existing audio systems in virtual, augmented, and mixed reality environments struggle to accurately present interaural time differences (ITD) between a user's ears while minimizing sound artifacts and maintaining computational efficiency, especially when rapid changes in audio signals are required to reflect object positions and orientations.

Method used

A method and system for processing audio signals using a wearable head device that applies delays, gain adjustments, and head-related transfer functions (HRTF) to generate left and right output audio signals, incorporating interaural time delays (ITD) based on source locations within the virtual environment, and employs crossfading techniques to manage transitions in delays.

Benefits of technology

This approach enhances the accuracy of audio signal presentation, reduces sound artifacts, and maintains computational efficiency, thereby improving the immersive experience by accurately simulating sound origins in mixed reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776332000001
    Figure 0007776332000001
  • Figure 0007776332000002
    Figure 0007776332000002
  • Figure 0007776332000003
    Figure 0007776332000003
Patent Text Reader

Abstract

[0006] An embodiment of the present disclosure describes a system and method for presenting audio signals to a user of a wearable head device. According to an exemplary method, a first input audio signal is received, the first input audio signal corresponding to a source location within a virtual environment presented to the user via the wearable head device. The first input audio signal is processed to generate a left output audio signal and a right output audio signal. The left output audio signal is presented to the user's left ear via a left speaker associated with the wearable head device. The right output audio signal is presented to the user's right ear via a right speaker associated with the wearable head device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 62 / 742,254, filed October 5, 2018, U.S. Provisional Application No. 62 / 812,546, filed March 1, 2019, and U.S. Provisional Application No. 62 / 742,191, filed October 5, 2018, the contents of which are incorporated herein by reference in their entireties.

[0002] The present disclosure relates generally to systems and methods for audio signal processing, and more particularly to systems and methods for presenting audio signals within a mixed reality environment. [Background technology]

[0003] Immersive and believable virtual environments require the presentation of audio signals in a manner consistent with user expectations, e.g., that an audio signal corresponding to an object in a virtual environment will be consistent with the location of that object in the virtual environment and with the visual presentation of that object. Creating rich and complex soundscapes (sound environments) in virtual reality, augmented reality, and mixed reality environments each requires the efficient presentation of numerous digital audio signals that appear to emanate from different locations / proximities and / or directions within the user's environment. A listener's brain is adapted to recognize differences in the arrival times of sound between a user's two ears (e.g., by detecting a phase shift between the two ears) and infer the spatial origin of the sound from the time difference. Thus, accurately presenting the interaural time difference (ITD) between a user's left and right ear with respect to a virtual environment can be important to the user's ability to identify audio sources within the virtual environment. However, adjusting the soundscape to credibly reflect the position and orientation of objects and the user may require rapid changes in the audio signal, which may result in undesirable sound artifacts such as "clicks" that detract from the immersive feel of the virtual environment. It is desirable for systems and methods for presenting soundscapes to a user of a virtual environment to accurately present interaural time differences to the user's ears while minimizing sound artifacts and remaining computationally efficient. Summary of the Invention [Means for solving the problem]

[0004]

[0006] An embodiment of the present disclosure describes a system and method for presenting audio signals to a user of a wearable head device. According to an exemplary method, a first input audio signal is received, the first input audio signal corresponding to a source location within a virtual environment presented to the user via the wearable head device. The first input audio signal is processed to generate a left output audio signal and a right output audio signal. The left output audio signal is presented to the user's left ear via a left speaker associated with the wearable head device. The right output audio signal is presented to the user's right ear via a right speaker associated with the wearable head device. Processing the first input audio signal includes applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal, adjusting a gain of the left audio signal, adjusting a gain of the right audio signal, applying a first head-related transfer function (HRTF) to the left audio signal to generate a left output audio signal, and applying a second HRTF to the right audio signal to generate a right output audio signal. Applying a delay process to the first input audio signal includes applying an interaural time delay (ITD) to the first input audio signal, the ITD being determined based on the source location. The present invention provides, for example, the following. (Item 1) 1. A method of presenting an audio signal to a user of a wearable head device, the method comprising: receiving a first input audio signal, the first input audio signal corresponding to a first source location within a virtual environment presented to the user via the wearable head device; processing the first input audio signal to generate a left output audio signal and a right output audio signal, wherein processing the first input audio signal comprises: applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal; adjusting a gain of the left audio signal; adjusting a gain of the right audio signal; applying a first head-related transfer function (HRTF) to the left audio signal to generate the left output audio signal; applying a second HRTF to the right audio signal to generate the right output audio signal; and presenting the left output audio signal to the user's left ear via a left speaker associated with the wearable head device; presenting the right output audio signal to the user's right ear via a right speaker associated with the wearable head device; and Including, The method, wherein applying the delay process to the first input audio signal includes applying an interaural time delay (ITD) to the first input audio signal, the ITD being determined based on the first source location. (Item 2) 2. The method of claim 1, wherein determining the ITD includes determining a first ear delay corresponding to a first ear and a second ear delay corresponding to a second ear. (Item 3) Item 3. The method of item 2, wherein the first ear delay is zero. (Item 4) the first source location corresponds to a location of a virtual object in the virtual environment at a first time; The method further includes determining a second source location corresponding to a location of the virtual object in the virtual environment at a second time; determining the first ear delay includes determining a leading first ear delay corresponding to the first time and a trailing first ear delay corresponding to the second time, and crossfading between the leading first ear delay and the trailing first ear delay to generate the first ear delay; determining the second ear delay includes determining a leading second ear delay corresponding to the first time and a trailing second ear delay corresponding to the second time, and crossfading between the leading second ear delay and the trailing second ear delay to generate the second ear delay; The method described in item 2. (Item 5) Item 5. The method of item 4, wherein applying the delay process includes applying a filter to the first input audio signal. (Item 6) Item 5. The method of item 4, further comprising applying a filter to the left audio signal. (Item 7) Item 5. The method of item 4, further comprising applying a filter to the right audio signal. (Item 8) 5. The method of claim 4, wherein applying the delay process includes transitioning from a first delay module at the first time to a second delay module at the second time, the second delay module being different from the first delay module. (Item 9) 9. The method of claim 8, wherein the first delay module is associated with applying a first one or more filters to a first one or more of the first input audio signal, the left audio signal, and the right audio signal, and the second delay module is associated with applying a second one or more filters to a second one or more of the first input audio signal, the left audio signal, and the right audio signal. (Item 10) Item 5. The method of item 4, wherein the first ear corresponds to the user's left ear and the second ear corresponds to the user's right ear. (Item 11) Item 5. The method of item 4, wherein the first ear corresponds to the user's right ear and the second ear corresponds to the user's left ear. (Item 12) Item 5. The method of item 4, wherein the first source location is closer to the first ear than to the second ear, and the second source location is closer to the second ear than to the first ear. (Item 13) Item 5. The method of item 4, wherein the second source location is closer to the first ear than to the second ear, and the first source location is closer to the second ear than to the first ear. (Item 14) Item 5. The method of item 4, wherein the first source location is closer to the first ear than the second source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear. (Item 15) Item 5. The method of item 4, wherein the second source location is closer to the first ear than the first source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear. (Item 16) 1. A system comprising: a wearable head device; a left speaker associated with the wearable head device; a right speaker associated with the wearable head device; One or more processors, the one or more processors configured to perform a method, the method comprising: receiving a first input audio signal, the first input audio signal corresponding to a first source location within a virtual environment presented to a user via the wearable head device; processing the first input audio signal to generate a left output audio signal and a right output audio signal, wherein processing the first input audio signal comprises: applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal; adjusting a gain of the left audio signal; adjusting a gain of the right audio signal; applying a first head-related transfer function (HRTF) to the left audio signal to generate the left output audio signal; applying a second HRTF to the right audio signal to generate the right output audio signal; and presenting the left output audio signal to the left ear of the user via the left speaker; presenting the right output audio signal to the user's right ear via the right speaker; Including, applying the delay process to the first input audio signal includes applying an interaural time delay (ITD) to the first input audio signal, the ITD being determined based on the first source location; and A system comprising: (Item 17) Item 17. The system of item 16, wherein determining the ITD includes determining a first ear delay corresponding to a first ear and a second ear delay corresponding to a second ear. (Item 18) Item 18. The system of item 17, wherein the first ear delay is zero. (Item 19) the first source location corresponds to a location of a virtual object in the virtual environment at a first time; The method further includes determining a second source location corresponding to a location of the virtual object in the virtual environment at a second time; determining the first ear delay includes determining a leading first ear delay corresponding to the first time and a trailing first ear delay corresponding to the second time, and crossfading between the leading first ear delay and the trailing first ear delay to generate the first ear delay; determining the second ear delay includes determining a leading second ear delay corresponding to the first time and a trailing second ear delay corresponding to the second time, and crossfading between the leading second ear delay and the trailing second ear delay to generate the second ear delay; Item 18. The system according to item 17. (Item 20) 20. The system of claim 19, wherein applying the delay process includes applying a filter to the first input audio signal. (Item 21) 20. The system of claim 19, wherein the method further comprises applying a filter to the left audio signal. (Item 22) 20. The system of claim 19, wherein the method further comprises applying a filter to the right audio signal. (Item 23) 20. The system of claim 19, wherein applying the delay process includes transitioning from a first delay module at the first time to a second delay module at the second time, the second delay module being different from the first delay module. (Item 24) 24. The system of claim 23, wherein the first delay module is associated with applying a first one or more filters to a first one or more of the first input audio signal, the left audio signal, and the right audio signal, and the second delay module is associated with applying a second one or more filters to a second one or more of the first input audio signal, the left audio signal, and the right audio signal. (Item 25) 20. The system of claim 19, wherein the first ear corresponds to the user's left ear and the second ear corresponds to the user's right ear. (Item 26) 20. The system of claim 19, wherein the first ear corresponds to the user's right ear and the second ear corresponds to the user's left ear. (Item 27) 20. The system of claim 19, wherein the first source location is closer to the first ear than to the second ear, and the second source location is closer to the second ear than to the first ear. (Item 28) 20. The system of claim 19, wherein the second source location is closer to the first ear than to the second ear, and the first source location is closer to the second ear than to the first ear. (Item 29) 20. The system of claim 19, wherein the first source location is closer to the first ear than the second source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear. (Item 30) 20. The system of claim 19, wherein the second source location is closer to the first ear than the first source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear. (Item 31) 1. A non-transitory computer-readable medium containing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for presenting an audio signal to a user of a wearable head device, the method comprising: receiving a first input audio signal, the first input audio signal corresponding to a first source location within a virtual environment presented to the user via the wearable head device; processing the first input audio signal to generate a left output audio signal and a right output audio signal, wherein processing the first input audio signal comprises: applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal; adjusting a gain of the left audio signal; adjusting a gain of the right audio signal; applying a first head-related transfer function (HRTF) to the left audio signal to generate the left output audio signal; applying a second HRTF to the right audio signal to generate the right output audio signal; and presenting the left output audio signal to the user's left ear via a left speaker associated with the wearable head device; presenting the right output audio signal to the user's right ear via a right speaker associated with the wearable head device; and Including, a first input audio signal having a first source location and a second source location, the first source location being determined based on the first source location; (Item 32) Item 32. The non-transitory computer-readable medium of item 31, wherein determining the ITD includes determining a first ear delay corresponding to a first ear and a second ear delay corresponding to a second ear. (Item 33) Item 33. The non-transitory computer-readable medium of item 32, wherein the first ear delay is zero. (Item 34) the first source location corresponds to a location of a virtual object in the virtual environment at a first time; The method further includes determining a second source location corresponding to a location of the virtual object in the virtual environment at a second time; determining the first ear delay includes determining a leading first ear delay corresponding to the first time and a trailing first ear delay corresponding to the second time, and crossfading between the leading first ear delay and the trailing first ear delay to generate the first ear delay; determining the second ear delay includes determining a leading second ear delay corresponding to the first time and a trailing second ear delay corresponding to the second time, and crossfading between the leading second ear delay and the trailing second ear delay to generate the second ear delay; Item 33. The non-transitory computer-readable medium of item 32. (Item 35) Item 35. The non-transitory computer-readable medium of item 34, wherein applying the delay process includes applying a filter to the first input audio signal. (Item 36) Item 35. The non-transitory computer-readable medium of item 34, wherein the method further comprises applying a filter to the left audio signal. (Item 37) Item 35. The non-transitory computer-readable medium of item 34, wherein the method further comprises applying a filter to the right audio signal. (Item 38) Item 35. The non-transitory computer-readable medium of item 34, wherein applying the delay process includes transitioning from a first delay module at the first time to a second delay module at the second time, the second delay module being different from the first delay module. (Item 39) Item 39. The non-transitory computer-readable medium of item 38, wherein the first delay module is associated with applying a first one or more filters to a first one or more of the first input audio signal, the left audio signal, and the right audio signal, and the second delay module is associated with applying a second one or more filters to a second one or more of the first input audio signal, the left audio signal, and the right audio signal. (Item 40) Item 35. The non-transitory computer-readable medium of item 34, wherein the first ear corresponds to the user's left ear and the second ear corresponds to the user's right ear. (Item 41) Item 35. The non-transitory computer-readable medium of item 34, wherein the first ear corresponds to the user's right ear and the second ear corresponds to the user's left ear. (Item 42) Item 35. The non-transitory computer-readable medium of item 34, wherein the first source location is closer to the first ear than to the second ear, and the second source location is closer to the second ear than to the first ear. (Item 43) Item 35. The non-transitory computer-readable medium of item 34, wherein the second source location is closer to the first ear than to the second ear, and the first source location is closer to the second ear than to the first ear. (Item 44) Item 35. The non-transitory computer-readable medium of item 34, wherein the first source location is closer to the first ear than the second source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear. (Item 45) Item 35. The non-transitory computer-readable medium of item 34, wherein the second source location is closer to the first ear than the first source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear. [Brief explanation of the drawings]

[0005] [Figure 1] FIG. 1 illustrates an exemplary audio spatialization system according to some embodiments of the present disclosure.

[0006] [Figure 2] 2A-2C illustrate an example delay module according to some embodiments of the present disclosure.

[0007] [Figure 3] 3A-3B illustrate example virtual sound sources and example corresponding delay modules for a listener, respectively, according to some embodiments of the present disclosure.

[0008] [Figure 4] 4A-4B illustrate example virtual sound sources and example corresponding delay modules, respectively, relative to a listener, according to some embodiments of the present disclosure.

[0009] [Figure 5] 5A-5B illustrate example virtual sound sources and example corresponding delay modules for a listener, respectively, according to some embodiments of the present disclosure.

[0010] [Figure 6A] FIG. 6A illustrates an exemplary crossfader according to some embodiments of the present disclosure.

[0011] [Figure 6B] 6B-6C illustrate example control signals for a crossfader, according to some embodiments of the present disclosure. [Figure 6C] 6B-6C illustrate example control signals for a crossfader, according to some embodiments of the present disclosure.

[0012] [Figure 7] 7A-7B illustrate example virtual sound sources for a listener and example corresponding delay modules including crossfaders, respectively, according to some embodiments of the present disclosure.

[0013] [Figure 8] 8A-8B illustrate example virtual sound sources for a listener and example corresponding delay modules including crossfaders, respectively, according to some embodiments of the present disclosure.

[0014] [Figure 9] 9A-9B illustrate example virtual sound sources for a listener and example corresponding delay modules including a crossfader, respectively, according to some embodiments of the present disclosure.

[0015] [Figure 10] 10A-10B illustrate example virtual sound sources for a listener and example corresponding delay modules including a crossfader, respectively, according to some embodiments of the present disclosure.

[0016] [Figure 11] 11A-11B illustrate example virtual sound sources for a listener and example corresponding delay modules including crossfaders, respectively, according to some embodiments of the present disclosure.

[0017] [Figure 12] 12A-12B illustrate example virtual sound sources for a listener and example corresponding delay modules including a crossfader, respectively, according to some embodiments of the present disclosure.

[0018] [Figure 13] 13A-13B illustrate example virtual sound sources for a listener and example corresponding delay modules including crossfaders, respectively, according to some embodiments of the present disclosure.

[0019] [Figure 14] 14A-14B illustrate example virtual sound sources for a listener and example corresponding delay modules including crossfaders, respectively, according to some embodiments of the present disclosure.

[0020] [Figure 15] 15A-15B illustrate example virtual sound sources for a listener and example corresponding delay modules including a crossfader, respectively, according to some embodiments of the present disclosure.

[0021] [Figure 16] 16A-16B illustrate example virtual sound sources for a listener and example corresponding delay modules including crossfaders, respectively, according to some embodiments of the present disclosure.

[0022] [Figure 17] FIG. 17 illustrates an example delay module according to some embodiments of the present disclosure.

[0023] [Figure 18-1] 18A-18E illustrate an example delay module according to some embodiments of the present disclosure. [Figure 18-2] 18A-18E illustrate an example delay module according to some embodiments of the present disclosure.

[0024] [Figure 19] 19-22 illustrate example processes for transitioning between delay modules according to some embodiments of the present disclosure. [Figure 20] 19-22 illustrate example processes for transitioning between delay modules according to some embodiments of the present disclosure. [Figure 21] 19-22 illustrate example processes for transitioning between delay modules according to some embodiments of the present disclosure. [Figure 22] 19-22 illustrate example processes for transitioning between delay modules according to some embodiments of the present disclosure.

[0025] [Figure 23] FIG. 23 illustrates an exemplary wearable system according to some embodiments of the present disclosure.

[0026] [Figure 24] FIG. 24 illustrates an example handheld controller that may be used in conjunction with an example wearable system, according to some embodiments of the present disclosure.

[0027] [Figure 25] FIG. 25 illustrates an example auxiliary unit that may be used in conjunction with an example wearable system, according to some embodiments of the present disclosure.

[0028] [Figure 26] FIG. 26 illustrates an example functional block diagram for an example wearable system according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0029] In the following description of the embodiments, reference is made to the accompanying drawings which form a part hereof, and in which is shown, by way of illustration, specific embodiments which may be practiced. It is to be understood that other embodiments may be used and structural changes may be made without departing from the scope of the disclosed embodiments.

[0030] Exemplary Wearable System

[0031] 23 illustrates an exemplary wearable head device 2300 configured to be worn on a user's head. The wearable head device 2300 may be part of a broader wearable system that includes one or more components, such as a head device (e.g., wearable head device 2300), a handheld controller (e.g., handheld controller 2400 described below), and / or an auxiliary unit (e.g., auxiliary unit 2500 described below). In some examples, the wearable head device 2300 can be used for virtual reality, augmented reality, or mixed reality systems or applications. The wearable head device 2300 includes one or more displays, such as displays 2310A and 2310B (which may include left and right transmissive displays and associated components for coupling light from the displays to the user's eyes, such as orthogonal pupil expansion (OPE) grating sets 2312A / 2312B and exit pupil expansion (EPE) grating sets 2314A / 2314B), left and right acoustic structures, such as speakers 2320A and 2320B (which may be mounted on temple arms 2322A and 2322B, respectively, and positioned adjacent the user's left and right ears), and infrared The wearable head device 2300 may include one or more sensors such as an accelerometer, a GPS unit, an inertial measurement unit (IMU, e.g., IMU 2326), an acoustic sensor (e.g., microphone 2350), a quadrature coil electromagnetic receiver (e.g., receiver 2327 shown mounted on left temple arm 2322A), left and right cameras oriented away from the user (e.g., depth (time-of-flight) cameras 2330A and 2330B), and left and right eye cameras oriented toward the user (e.g., for detecting the user's eye movements) (e.g., eye cameras 2328A and 2328B). However, the wearable head device 2300 may incorporate any suitable display technology and any suitable number, type, or combination of sensors or other components without departing from the scope of the invention.In some examples, the wearable head device 2300 may incorporate one or more microphones 150 configured to detect audio signals generated by the user's voice, and such microphones may be positioned adjacent to the user's mouth. In some examples, the wearable head device 2300 may incorporate networking features (e.g., Wi-Fi capabilities) for communicating with other devices and systems, including other wearable systems. The wearable head device 2300 may further include components such as a battery, a processor, memory, a storage unit, or various input devices (e.g., buttons, touchpad), or may be coupled to a handheld controller (e.g., handheld controller 2400) or auxiliary unit (e.g., auxiliary unit 2500) that includes one or more such components. In some examples, the sensors may be configured to output a set of coordinates of the head-mounted unit relative to the user's environment and may provide input to a processor to implement a simultaneous localization and mapping (SLAM) procedure and / or a visual odometry algorithm. In some embodiments, the wearable head device 2300 may be coupled to a handheld controller 2400 and / or an auxiliary unit 2500, as described further below.

[0032] FIG. 24 illustrates an exemplary mobile handheld controller component 2400 of an exemplary wearable system. In some examples, the handheld controller 2400 may communicate wired or wirelessly with the wearable head device 2300 and / or an auxiliary unit 2500, described below. In some examples, the handheld controller 2400 includes a handle portion 2420 to be held by a user and one or more buttons 2440 disposed along a top surface 2410. In some examples, the handheld controller 2400 may be configured for use as an optical tracking target; for example, a sensor (e.g., a camera or other optical sensor) of the wearable head device 2300 can be configured to detect the position and / or orientation of the handheld controller 2400, which in turn may indicate the position and / or orientation of a user's hand holding the handheld controller 2400. In some examples, the handheld controller 2400 may include a processor, memory, a storage unit, a display, or one or more input devices, such as those described above. In some examples, the handheld controller 2400 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 2300). In some examples, the sensors can detect the position or orientation of the handheld controller 2400 relative to the wearable head device 2300 or relative to another component of the wearable system. In some examples, the sensors may be positioned within a handle portion 2420 of the handheld controller 2400 and / or may be mechanically coupled to the handheld controller. The handheld controller 2400 can be configured to provide one or more output signals corresponding, for example, to a press state of a button 2440, or the position, orientation, and / or movement of the handheld controller 2400 (e.g., via an IMU).Such output signals may be used as inputs to a processor of the wearable head device 2300, to the auxiliary unit 2500, or to another component of the wearable system. In some examples, the handheld controller 2400 may include one or more microphones to detect sounds (e.g., a user's speech, environmental sounds) and, in some cases, provide signals corresponding to the detected sounds to a processor (e.g., a processor of the wearable head device 2300).

[0033] FIG. 25 illustrates an exemplary auxiliary unit 2500 of an exemplary wearable system. In some examples, the auxiliary unit 2500 may communicate wired or wirelessly with the wearable head device 2300 and / or the handheld controller 2400. The auxiliary unit 2500 may include a battery to provide energy for operating one or more components of the wearable system, such as the wearable head device 2300 and / or the handheld controller 2400 (including a display, sensors, acoustic structure, processor, microphone, and / or other components of the wearable head device 2300 or the handheld controller 2400). In some examples, the auxiliary unit 2500 may include a processor, memory, a storage unit, a display, one or more input devices, and / or one or more sensors such as those described above. In some examples, the auxiliary unit 2500 includes a clip 2510 for attaching the auxiliary unit to a user (e.g., to a belt worn by the user). An advantage of using auxiliary unit 2500 to store one or more components of the wearable system is that doing so may allow large or heavy components to be carried on the user's waist, chest, or back, which are relatively better suited to supporting large, heavy objects, rather than being mounted on the user's head (e.g., when stored in wearable head device 2300) or carried by the user's hands (e.g., when stored in handheld controller 2400). This may be particularly advantageous with respect to relatively heavy or bulky components, such as batteries.

[0034] 26 shows an example functional block diagram that may correspond to an example wearable system 2600, such as may include the example wearable head device 2300, handheld controller 2400, and auxiliary unit 2500 described above. In some examples, the wearable system 2600 may be used for virtual reality, augmented reality, or mixed reality applications. As shown in FIG. 26, the wearable system 2600 may include an example handheld controller 2600B, referred to herein as a “totem” (and which may correspond to the handheld controller 2400 described above), which may include a totem / headgear six-degree-of-freedom (6DOF) totem subsystem 2604A. The wearable system 2600 may also include an exemplary headgear device 2600A (which may correspond to the wearable head device 2300 described above), which includes a totem / headgear 6DOF headgear subsystem 2604B. In an example, the 6DOF totem subsystem 2604A and the 6DOF headgear subsystem 2604B cooperate to determine six coordinates (e.g., offsets in three translational directions and rotations along three axes) of the handheld controller 2600B relative to the headgear device 2600A. The six degrees of freedom may be expressed relative to the coordinate system of the headgear device 2600A. The three translational offsets may be expressed as X, Y, and Z offsets within such a coordinate system, a translation matrix, or some other representation. The rotational degrees of freedom may be expressed as a sequence of yaw, pitch, and roll rotations, a vector, a rotation matrix, a quaternion, or some other representation. In some embodiments, one or more depth cameras 2644 (and / or one or more non-depth cameras) included within the headgear device 2600A and / or one or more optical targets (e.g., buttons 2440 of the handheld controller 2400 as described above or dedicated optical targets included within the handheld controller) can be used for 6DOF tracking.In some embodiments, the handheld controller 2600B can include a camera as described above, and the headgear device 2600A can include an optical target for optical tracking in conjunction with the camera. In some embodiments, the headgear device 2600A and the handheld controller 2600B each include a set of three orthogonally oriented solenoids used to wirelessly transmit and receive three distinguishable signals. By measuring the relative magnitudes of the three distinguishable signals received in each of the coils used for receiving, the 6DOF of the handheld controller 2600B relative to the headgear device 2600A can be determined. In some embodiments, the 6DOF totem subsystem 2604A can include an inertial measurement unit (IMU), which is useful for providing improved accuracy and / or more timely information regarding high-speed movement of the handheld controller 2600B.

[0035] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the headgear device 2600A) to an inertial coordinate space or to an environmental coordinate space. For example, such a transformation may be necessary for the display of the headgear device 2600A to present virtual objects in an expected position and orientation relative to the real environment (e.g., a virtual person sitting in a real chair facing forward, regardless of the position and orientation of the headgear device 2600A), rather than in a fixed position and orientation on the display (e.g., at the same position on the display of the headgear device 2600A). This can maintain the illusion that the virtual objects exist in the real environment (and do not appear unnaturally positioned in the real environment, e.g., as the headgear device 2600A shifts and rotates). In some embodiments, a compensatory transformation between coordinate spaces can be determined by processing images from the depth camera 2644 (e.g., using simultaneous localization and mapping (SLAM) and / or visual odometry procedures) to determine the transformation of the headgear device 2600A relative to an inertial or environmental coordinate system. In the embodiment shown in FIG. 26, the depth camera 2644 can be coupled to the SLAM / visual odometry block 2606 and can provide images to the block 2606. The SLAM / visual odometry block 2606 implementation can include a processor configured to process the images and then determine the position and orientation of the user's head, which can be used to identify a transformation between the head coordinate space and the real coordinate space. Similarly, in some embodiments, an additional source of information regarding the user's head pose and location is obtained from the IMU 2609 of the headgear device 2600A. Information from the IMU 2609 may be integrated with information from the SLAM / Visual Odometry block 2606 to provide improved accuracy and / or more timely information regarding fast adjustments of the user's head pose and position.

[0036] In some examples, depth camera 2644 can provide 3D images to hand gesture tracker 2611, which can be implemented within a processor of headgear device 2600A. Hand gesture tracker 2611 can identify the user's hand gestures, for example, by matching 3D images received from depth camera 2644 to stored patterns representing hand gestures. Other suitable techniques for identifying the user's hand gestures will also be apparent.

[0037] In some embodiments, one or more processors 2616 may be configured to receive data from the headgear subsystem 2604B, the IMU 2609, the SLAM / visual odometry block 2606, the depth camera 2644, the microphone 2650, and / or the hand gesture tracker 2611. The processor 2616 may also send and receive control signals to the 6DOF totem system 2604A. The processor 2616 may be wirelessly coupled to the 6DOF totem system 2604A, such as in embodiments in which the handheld controller 2600B is untethered. The processor 2616 may further communicate with additional components, such as an audiovisual content memory 2618, a graphical processing unit (GPU) 2620, and / or a digital signal processor (DSP) audio spatializer 2622. The DSP audio spatializer 2622 may be coupled to a head-related transfer function (HRTF) memory 2625. The GPU 2620 may include a left channel output coupled to a left source of imagewise modulated light 2624 and a right channel output coupled to a right source of imagewise modulated light 2626. The GPU 2620 may output stereoscopic image data to the imagewise modulated light sources 2624, 2626. The DSP audio spatializer 2622 may output audio to the left speaker 2612 and / or the right speaker 2614. The DSP audio spatializer 2622 may receive an input from the processor 2616 indicating a direction vector from the user to a virtual sound source (which may be moved by the user, e.g., via the handheld controller 2600B). Based on the direction vector, the DSP audio spatializer 2622 may determine a corresponding HRTF (e.g., by accessing an HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 2622 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the believability and realism of virtual sounds by incorporating the user's relative position and orientation to the virtual sounds in the mixed reality environment, i.e., by presenting virtual sounds that match the user's expectations of what they would hear if the virtual sounds were real sounds in a real environment.

[0038] 26 , one or more of the processor 2616, GPU 2620, DSP audio spatializer 2622, HRTF memory 2625, and audio / visual content memory 2618 may be included in auxiliary unit 2600C (which may correspond to auxiliary unit 2500 described above). Auxiliary unit 2600C may include battery 2627 to power its components and / or provide power to headgear device 2600A and / or handheld controller 2600B. Including such components in an auxiliary unit, which may be mounted on the user's waist, can limit the size and weight of headgear device 2600A, which in turn can reduce fatigue in the user's head and neck.

[0039] While FIG. 26 presents elements corresponding to various components of the exemplary wearable system 2600, various other suitable arrangements of these components will be apparent to those skilled in the art. For example, elements presented in FIG. 26 as being associated with the auxiliary unit 2600C may instead be associated with the headgear device 2600A or the handheld controller 2600B. Furthermore, some wearable systems may dispense with the handheld controller 2600B or the auxiliary unit 2600C entirely. Such variations and modifications are understood to be within the scope of the disclosed embodiments.

[0040] Audio Rendering

[0041] The systems and methods described below can be implemented in an augmented reality or mixed reality system such as those described above. For example, one or more processors (e.g., CPU, DSP) of the augmented reality system can be used to process audio signals or implement steps of the computer-implemented methods described below, sensors (e.g., cameras, acoustic sensors, IMU, LIDAR, GPS) of the augmented reality system can be used to determine the position and / or orientation of a user of the system or elements in the user's environment, and speakers of the augmented reality system can be used to present audio signals to the user.

[0042] In an augmented reality or mixed reality system such as that described above, one or more processors (e.g., DSP audio spatializer 2622) can process one or more audio signals for presentation to a user of a wearable head device via one or more speakers (e.g., left and right speakers 2612 / 2614 described above). In some embodiments, the one or more speakers may reside in a unit (e.g., headphones) separate from the wearable head device. Processing of audio signals requires a trade-off between the perceived authenticity of the audio signals—e.g., the degree to which audio signals presented to a user in a mixed reality environment match the user's expectations of how the audio signals would sound in the real environment—and the computational overhead involved in processing the audio signals. Realistic spatialization of audio signals within a virtual environment can be important to creating an immersive and believable user experience.

[0043] 1 illustrates an exemplary spatialization system 100 according to some embodiments. The system 100 creates a soundscape (sound environment) by spatializing an input sound / signal. The system 100 includes an encoder 104, a mixer 106, and a decoder 110.

[0044] The system 100 receives an input signal 102. The input signal 102 may include a digital audio signal corresponding to an object to be presented within a soundscape. In some embodiments, the digital audio signal may be a pulse code modulated (PCM) waveform of audio data.

[0045] Encoder 104 receives input signal 102 and outputs one or more left gain adjustment signals and one or more right gain adjustment signals. In an embodiment, encoder 104 includes delay module 105. Delay module 105 may include a delay process that may be performed by a processor (such as the processor of the augmented reality system described above). To make objects in the soundscape appear to originate from specific locations, encoder 104 delays input signal 102 using delay module 105 accordingly and sets the values ​​of control signals (CTRL_L1 ... CRTL_LM and CTRL_R1 ... CTRL_RM) that are input to gain modules (g_L1 ... g_LM and g_R1 ... g_RM).

[0046] The delay module 105 receives the input signal 102 and outputs a left-ear delay and a right-ear delay. The left-ear delay is input to a left gain module (g_L1 ... g_LM), and the right-ear delay is input to a right gain module (g_R1 ... g_RM). The left-ear delay may be the input signal 102 delayed by a first value, and the right-ear delay may be the input signal 102 delayed by a second value. In some embodiments, the left-ear delay and / or right-ear delay may be zero, in which case the delay module 105 effectively routes the input signal 102 to the left gain module and / or right gain module, respectively. The interaural time difference (ITD) may be the difference between the left-ear delay and the right-ear delay.

[0047] One or more left control signals (CTRL_L1...CTRL_LM) are input to one or more left gain modules, and one or more right control values ​​(CTRL_R1...CTRL_RM) are input to one or more right gain modules, which output one or more left gain adjustment signals, and one or more right gain modules which output one or more right gain adjustment signals.

[0048] The one or more left gain modules each adjust a gain of a left ear delay based on a control signal value of the one or more left control signals, and the one or more right gain modules each adjust a gain of a right ear delay based on a control signal value of the one or more right control signals.

[0049] The encoder 104 adjusts the value of the control signal input to the gain modules based on the location of the object to be presented in the soundscape corresponding to the input signal 102. Each gain module may be a multiplier that multiplies the input signal 102 by a coefficient that is a function of the value of the control signal.

[0050] Mixer 106 receives the gain adjustment signals from encoder 104, mixes the gain adjustment signals, and outputs a mixed signal that is input to decoder 110, the output of which is input to left ear speaker 112A and right ear speaker 112B (hereinafter collectively referred to as "speakers 112").

[0051] The decoder 110 includes a left HRTF filter L_HRTF_1-M and a right HRTF filter R_HRTF_1-M. The decoder 110 receives the mixed signals from the mixer 106, filters and sums the mixed signals, and outputs the filtered signals to the speaker 112. A first summing block / circuit of the decoder 110 sums the left filtered signals output from the left HRTF filters, and a second summing block / circuit of the decoder 110 sums the right filtered signals output from the right HRTF filters.

[0052] In some embodiments, the decoder 110 may include a crosstalk canceller to convert left / right physical speaker positions to individual ear positions, such as that described in Jot, et al., Binaural Simulation of Complex Acoustic Scenes for Interactive Audio, Audio Engineering Society Convention Paper, presented October 5-8, 2006, the contents of which are incorporated herein by reference in their entirety.

[0053] In some embodiments, the decoder 110 may include a bank of HRTF filters. Each HRTF filter in the bank may model a specific direction relative to the user's head. These methods may be based on decomposition of HRTF data over a fixed set of spatial functions and a fixed set of basis filters. In these embodiments, each mixed signal from the mixer 106 may be mixed into the input of an HRTF filter that models the direction closest to the direction of the source. The level of the signal mixed into each of those HRTF filters is determined by the specific direction of the source.

[0054] In some embodiments, system 100 may receive multiple input signals and may include an encoder for each of the multiple input signals, where the total number of input signals may represent the total number of objects to be presented in the soundscape.

[0055] If the direction of an object being presented within the soundscape changes, not only can encoder 104A vary the values ​​of one or more left control signals and one or more right control signals input to one or more left gain modules and one or more right gain modules, but delay module 105 may vary the delay of input signal 102 to generate left ear delays and / or right ear delays in order to properly present the object within the soundscape.

[0056] 2A-2C illustrate various modes of the delay module 205, according to some embodiments. The delay module 205 may include a delay unit 216 that delays an input signal by a value, such as a time value, a sample count, and the like. One or more of the example delay modules shown in FIGS. 2A-2C may be used to implement the delay module 105 shown in the example system 100.

[0057] 2A illustrates the zero-tap delay mode of the delay module 205, according to some embodiments. In the zero-tap delay mode, the input signal 202 is divided to create a first ear delay 222 and a second ear delay 224. The delay unit 216 receives the input signal 202 but does not delay the input signal 202. In some embodiments, the delay unit 216 receives the input signal 202 and fills a buffer with samples of the input signal 202, which may then be used when the delay module 205 transitions to a one-tap delay mode or a two-tap delay mode (described below). The delay module 205 outputs the first ear delay 222 and the second ear delay 224, which are simply the input signal 202 (without any delay).

[0058] 2B illustrates a one-tap delay mode of the delay module 205, according to some embodiments. In the one-tap delay mode, the input signal 202 is divided to create a second ear delay 228. The delay unit 216 receives the input signal 202, delays the input signal 202 by a first value, and outputs a first ear delay 226. The second ear delay 228 is simply the input signal 202 (without any delay). The delay module 205 outputs a first ear delay 226 and a second ear delay 228. In some embodiments, the first ear delay 226 may be a left ear delay and the second ear delay 228 may be a right ear delay. In some embodiments, the first ear delay 226 may be a right ear delay and the second ear delay 228 may be a left ear delay.

[0059] 2C illustrates a two-tap delay mode of the delay module 205, according to some embodiments. In the two-tap delay mode, the delay unit 216 receives the input signal 202, delays the input signal 202 by a first value, and outputs a first ear delay 232, and delays the input signal 202 by a second value, and outputs a second ear delay 234. In some embodiments, the first ear delay 232 may be a left ear delay and the second ear delay 234 may be a right ear delay. In some embodiments, the first ear delay 232 may be a right ear delay and the second ear delay 234 may be a left ear delay.

[0060] In some embodiments, a soundscape (sound environment) may be presented to the user. The following discussion pertains to soundscapes with a single virtual object; however, the principles described herein may be applicable to soundscapes with many virtual objects.

[0061] 3A illustrates an environment 300 including a user 302 and a virtual object (bee) 304 on a median plane 306, according to some embodiments. The distance 308 from the left ear of the user 302 to the virtual bee 304 is equal to the distance 310 from the right ear of the user 302 to the virtual bee 304. Therefore, sound from the virtual bee 304 should take the same amount of time to reach both the left and right ears.

[0062] 3B illustrates a delay module 312 corresponding to the environment 300 of FIG. 3A , according to some embodiments. The delay module 312 may be used to implement the delay module 105 shown in the exemplary system 100. As illustrated in FIG. 3B , the delay module 312 is in zero-tap delay mode, and the input signal 314 is divided to create a left-ear delay 316 and a right-ear delay 318. The left-ear delay 316 and the right-ear delay 318 are simply the input signal 314 because the distances 308 and 310 are the same. The delay unit 320 receives the input signal 314 but does not output a signal. In some embodiments, the delay unit 316 receives the input signal 314 and fills a buffer with samples of the input signal 314, which may then be used when the delay module 312 transitions to a one-tap delay mode or a two-tap delay mode. The delay module 312 outputs the left-ear delay 316 and the right-ear delay 318.

[0063] 4A illustrates an environment 400 including a user 402 and a virtual object (bee) 404 on the left side of a median plane 406, according to some embodiments. A distance 410 from the right ear of the user 402 to the virtual bee 404 exceeds a distance 408 from the left ear of the user 402 to the virtual bee 404. Thus, sound from the virtual bee 404 should take longer to reach the right ear than the left ear.

[0064] 4B illustrates a delay module 412 corresponding to the environment 400 of FIG. 4A , according to some embodiments. The delay module 412 may be used to implement the delay module 105 shown in the exemplary system 100. As shown in FIG. 4B , the delay module 412 is in a first tap delay mode, and the input signal 414 is divided to create a left ear delay 416. A delay unit 420 receives the input signal 414, delays the input signal 414 by a time 422, and outputs a right ear delay 418. The left ear delay 416 is simply the input signal 414, and the right ear delay 418 is simply a delayed version of the input signal 414. The delay module 412 outputs the left ear delay 416 and the right ear delay 418.

[0065] 5A illustrates an environment 500 including a user 502 and a virtual object (bee) 504 to the right of a median plane 506, according to some embodiments. A distance 508 from the left ear of the user 502 to the virtual bee 504 exceeds a distance 510 from the right ear of the user 502 to the virtual bee 504. Thus, sound from the virtual bee 504 should take longer to reach the left ear than the right ear.

[0066] 5B illustrates a delay module 512 corresponding to the environment 500 of FIG. 5A , according to some embodiments. The delay module 512 may be used to implement the delay module 105 shown in the exemplary system 100. As shown in FIG. 5B , the delay module 512 is in a one-tap delay mode, where the input signal 514 is split to create a right-ear delay 518. A delay unit 520 receives the input signal 514, delays the input signal 514 by a time 522, and outputs a left-ear delay 516. The right-ear delay 518 is simply the input signal 514, and the left-ear delay 516 is simply a delayed version of the input signal 514. The delay module 512 outputs the left-ear delay 516 and the right-ear delay 518.

[0067] In some embodiments, the orientation of a virtual object in the soundscape changes relative to the user. For example, the virtual object may move from the left side of the median plane to the right side of the median plane, from the right side of the median plane to the left side of the median plane, from a first position on the right side of the median plane to a second position on the right side of the median plane, the second position being closer to the median plane than the first position, from a first position on the right side of the median plane to a second position on the right side of the median plane, the second position being farther from the median plane than the first position, from a first position on the left side of the median plane to a second position on the left side of the median plane, the second position being closer to the median plane than the first position, from a first position on the left side of the median plane to a second position on the left side of the median plane, the second position being farther from the median plane than the first position, from the right side of the median plane onto the median plane, from above the median plane to the right side of the median plane, from the left side of the median plane onto the median plane, and from the median plane to the left side of the median plane, to name a few.

[0068] In some embodiments, a change in the orientation of a virtual object in the soundscape relative to the user may require a change in the ITD (e.g., the difference between left ear delay and right ear delay).

[0069] In some embodiments, a delay module (e.g., delay module 105 shown in exemplary system 100) may vary the ITD by momentarily varying the left-ear delay and / or the right-ear delay based on changes in the orientation of the virtual object. However, momentarily varying the left-ear delay and / or the right-ear delay may result in sound artifacts. The sound artifacts may be, for example, "clicking" sounds. Minimizing such sound artifacts is desirable.

[0070] In some embodiments, a delay module (e.g., delay module 105 shown in exemplary system 100) may vary the ITD by varying the left-ear delay and / or right-ear delay using ramping or smoothing the value of the delay based on changes in the orientation of the virtual object. However, varying the left-ear delay and / or right-ear delay using ramping or smoothing the value of the delay may result in sound artifacts. The sound artifacts may be, for example, changes in pitch. Minimizing such sound artifacts is desirable. In some embodiments, varying the left-ear delay and / or right-ear delay using ramping or smoothing the value of the delay may introduce latency, for example, due to the time it takes to calculate and perform the ramping or smoothing and / or due to the time it takes for new sounds to be delivered. Minimizing such latency is desirable.

[0071] In some embodiments, a delay module (e.g., delay module 105 shown in exemplary system 100) may vary the ITD by varying the left-ear delay and / or right-ear delay using a cross-fade from a first delay to a subsequent delay. The cross-fade may reduce artifacts during transitions between delay values, for example, by avoiding stretching or compressing the signal in the time domain. Stretching or compressing the signal in the time domain may result in "clicks" or pitch shifts as described above.

[0072] 6A illustrates a crossfader 600 according to some embodiments. The crossfader 600 may be used to implement the delay module 105 shown in the exemplary system 100. The crossfader 600 receives a first ear delay 602 and a subsequent ear delay 604 as inputs and outputs a crossfaded ear delay 606. The crossfader 600 includes a first level fader (Gf) 608, a subsequent level fader (Gs) 610, and a summer 612. The first level fader 608 progressively decreases the level of the first ear delay based on changes in the control signal CTRL_Gf, and the subsequent level fader 610 progressively increases the level of the subsequent ear delay based on changes in the control signal CTRL_Gs. The summer 612 sums the outputs of the first level fader 608 and the subsequent level fader 612.

[0073] 6B illustrates a model of the control signal CTRL_Gf, according to some embodiments. In the example shown, the value of the control signal CTRL_Gf decreases from 1 to zero over a period of time (e.g., 1 at time t_0 and zero at time t_end). In some embodiments, the value of the control signal CTRL_Gf may decrease from 1 to zero linearly, exponentially, or with some other function.

[0074] 6C illustrates a model of the control signal CTRL_Gs, according to some embodiments. In the example shown, the value of the control signal CTRL_Gs increases from zero to one over a period of time (e.g., zero at time t_0 and one at time t_end). In some embodiments, the value of the control signal CTRL_Gs may increase from zero to one linearly, exponentially, or with some other function.

[0075] 7A illustrates an environment 700 including a user 702 and a virtual object (bee) 704A on the left side of a median plane 706 at a first time and a virtual object (bee) 704B on the right side of the median plane 706 at a subsequent time, according to some embodiments. At the first time, a distance 710A from the virtual bee 704A to the right ear of the user 702 exceeds a distance 708A from the virtual bee 704A to the left ear of the user 702. Thus, at the first time, sound from the virtual bee 704A will take longer to reach the right ear than the left ear. At a subsequent time, a distance 708B from the virtual bee 704B to the left ear exceeds a distance 710B from the virtual bee 704B to the right ear. Thus, at a subsequent time, sound from the virtual bee 704B will take longer to reach the left ear than the right ear.

[0076] 7B illustrates a delay module 712 corresponding to the environment 700 of FIG. 7A, according to some embodiments. The delay module 712 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 712 receives an input signal 714 and outputs a left ear delay 716 and a right ear delay 718. The delay module 712 includes a delay unit 720 and two crossfaders: a left crossfader 730A and a right crossfader 730B. The left crossfader 730A includes a first level fader (Gf) 722A, a subsequent level fader (Gs) 724A, and a summer 726A. The right crossfader 730B includes a first level fader (Gf) 722B, a subsequent level fader (Gs) 724B, and a summer 726B.

[0077] At a first time, distance 710A exceeds distance 708A. For the first time, input signal 714 is fed directly to first level fader 722A, and delay unit 720 delays input signal 714 by the first time and feeds the input signal 714 delayed by the first time to first level fader 722B.

[0078] At a subsequent time, distance 708B exceeds distance 710A. For a subsequent time, input signal 714 is fed directly to subsequent level fader 724B, and delay unit 720 delays input signal 714 by the subsequent time and feeds the delayed input signal 714 by the subsequent time to subsequent level fader 724A.

[0079] Summer 726A sums the outputs of the first level fader 722A and the subsequent level fader 724A to create left ear delay 716, and summer 726B sums the outputs of the first level fader 722B and the subsequent level fader 724B to create right ear delay 718.

[0080] Thus, the left crossfader 730A crossfades between the input signal 714 and the input signal 714 delayed by a subsequent time, and the right crossfader 730B crossfades between the input signal 714 delayed by a first time and the input signal 714.

[0081] 8A illustrates an environment 800 including a user 802 and a virtual object (bee) 804A on the right side of a median plane 806 at a first time and a virtual object (bee) 804B on the left side of the median plane 806 at a subsequent time, according to some embodiments. At the first time, a distance 808A from the virtual bee 804A to the left ear of the user 802 exceeds a distance 810A from the virtual bee 804A to the right ear of the user 802. Thus, at the first time, sound from the virtual bee 804A will take longer to reach the left ear than the right ear. At a subsequent time, a distance 810B from the virtual bee 804B to the right ear exceeds a distance 808B from the virtual bee 804B to the left ear. Thus, at a subsequent time, sound from the virtual bee 804B will take longer to reach the right ear than the left ear.

[0082] 8B illustrates a delay module 812 corresponding to the environment 800 of FIG. 8A , according to some embodiments. The delay module 812 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 812 receives an input signal 814 and outputs a left-ear delay 816 and a right-ear delay 818. The delay module 812 includes a delay unit 820 and two crossfaders: a left crossfader 830A and a right crossfader 830B. The left crossfader 830A includes a first level fader (Gf) 822A, a subsequent level fader (Gs) 824A, and a summer 826A. The right crossfader 830B includes a first level fader (Gf) 822B, a subsequent level fader (Gs) 824B, and a summer 826B.

[0083] At a first time, distance 808A exceeds distance 810A. For the first time, input signal 814 is provided directly to first level fader 822B, and delay unit 820 delays input signal 814 by the first time and provides input signal 814 delayed by the first time to first level fader 822A.

[0084] At a subsequent time, distance 810B exceeds distance 808B. For a subsequent time, input signal 814 is fed directly to subsequent level fader 824A, and delay unit 820 delays input signal 814 by the subsequent time and feeds the input signal 814, delayed by the subsequent time, to subsequent level fader 824B.

[0085] Adder 826A sums the outputs of the first level fader 822A and the subsequent level fader 824A to create left ear delay 816, and adder 826B sums the outputs of the first level fader 822B and the subsequent level fader 824B to create right ear delay 818.

[0086] Thus, the left crossfader 830A crossfades between the input signal 814 delayed by a first time and the input signal 814, and the right crossfader 830B crossfades between the input signal 814 and the input signal 814 delayed by a subsequent time.

[0087] 9A illustrates an environment 900 including a user 902 and a virtual object (bee) 904A that is further to the right of the median plane 906 at a first time and a virtual object (bee) 904B that is less to the right of the median plane 906 (e.g., closer to the median plane 906) at a subsequent time, according to some embodiments. At the first time, a distance 908A from the virtual bee 904A to the left ear of the user 902 exceeds a distance 910A from the virtual bee 904A to the right ear of the user 902. Thus, at the first time, sound from the virtual bee 904A will take longer to reach the left ear than the right ear. At a subsequent time, a distance 908B from the virtual bee 904B to the left ear exceeds a distance 910B from the virtual bee 904B to the right ear. Thus, at a subsequent time, sound from the virtual bee 904B will take longer to reach the left ear than the right ear. Comparing distances 908A and 908B, distance 908A exceeds distance 908B, so that sound from virtual bee 904A at a first time should take longer to reach the left ear than sound from virtual bee 904B at a subsequent time. Comparing distances 910A and 910B, sound from virtual bee 904A at a first time should take the same time to reach the right ear as sound from virtual bee 904B at a subsequent time.

[0088] 9B illustrates a delay module 912 corresponding to the environment 900 of FIG. 9A , according to some embodiments. The delay module 912 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 912 receives an input signal 914 and outputs a left ear delay 916 and a right ear delay 918. The delay module 912 includes a delay unit 920 and a left crossfader 930. The left crossfader 930 includes a first level fader (Gf) 922, a trailing level fader (Gs) 924, and a summer 926.

[0089] At a first time, distance 908A exceeds distance 910A. For the first time, input signal 914 is fed directly to right ear delay 918, and delay unit 920 delays input signal 914 by the first time and feeds the input signal 914 delayed by the first time to first level fader 922.

[0090] At a subsequent time, distance 908B is greater than distance 910B, and distance 908B is less than distance 908A. For a subsequent time, input signal 914 is provided directly to right ear delay 918, and delay unit 920 delays input signal 914 by the subsequent time and provides input signal 914 delayed by the subsequent time to subsequent level fader 924.

[0091] The first time delayed input signal 914 may be delayed more than the subsequent time delayed input signal 914 because distance 908A exceeds distance 908B.

[0092] A summer 926 sums the outputs of the first level fader 922 and the subsequent level fader 924 to create the left ear delay 916 .

[0093] Thus, the left crossfader 930 crossfades between the input signal 914 delayed by a first time and the input signal 914 delayed by a subsequent time.

[0094] 10A illustrates an environment 1000 including a user 1002 and a virtual object (bee) 1004A to the right of a median plane 1006 at a first time and a virtual object (bee) 1004B that is further to the right of (e.g., further from) the median plane 1006 at a subsequent time, according to some embodiments. At the first time, a distance 1008A from the virtual bee 1004A to the left ear of the user 1002 exceeds a distance 1010A from the virtual bee 1004A to the right ear of the user 1002. Thus, at the first time, sound from the virtual bee 1004A should take longer to reach the left ear than the right ear. At a subsequent time, a distance 1008B from the virtual bee 1004B to the left ear exceeds a distance 1010B from the virtual bee 1004B to the right ear. Therefore, at a subsequent time, the sound from virtual bee 1004B should take longer to reach the left ear than the right ear. Comparing distances 1008A and 1008B, distance 1008B exceeds distance 1008A, so the sound from virtual bee 1004B at a subsequent time should take longer to reach the left ear than the sound from virtual bee 1004A at the first time. Comparing distances 1010A and 1010B, the sound from virtual bee 1004A at the first time should take the same time to reach the right ear as the sound from virtual bee 1004B at a subsequent time.

[0095] 10B illustrates a delay module 1012 corresponding to the environment 1000 of FIG. 10A, according to some embodiments. The delay module 1012 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1012 receives an input signal 1014 and outputs a left ear delay 1016 and a right ear delay 1018. The delay module 1012 includes a delay unit 1020 and a left crossfader 1030. The left crossfader 1030 includes a first level fader (Gf) 1022, a trailing level fader (Gs) 1024, and a summer 1026.

[0096] At a first time, distance 1008A exceeds distance 1010A. For the first time, input signal 1014 is fed directly to right ear delay 1018, and delay unit 1020 delays input signal 1014 by a first time and feeds input signal 1014 delayed by the first time to first level fader 1022.

[0097] At a subsequent time, distance 1008B exceeds distance 1010B, and distance 1008B exceeds distance 1008A. For a subsequent time, input signal 1014 is fed directly to right ear delay 1018, and delay unit 1020 delays input signal 1014 by a subsequent time and feeds input signal 1014 delayed by a subsequent time to subsequent level fader 1024.

[0098] The first time delayed input signal 1014 may be delayed less than the subsequent time delayed input signal 1014 because distance 1008A is less than distance 1008B.

[0099] A summer 1026 sums the outputs of the first level fader 1022 and the subsequent level fader 1024 to create the left ear delay 1016 .

[0100] Thus, the left crossfader 1030 crossfades between the input signal 1014 delayed by a first time and the input signal 1014 delayed by a subsequent time.

[0101] 11A illustrates an environment 1100 including a user 1102 and a virtual object (bee) 1104A that is more to the right of the median plane 1106 at a first time and a virtual object (bee) 1104B that is less to the right of the median plane 1106 (e.g., closer to the median plane 1106) at a subsequent time, according to some embodiments. At the first time, a distance 1110A from the virtual bee 1104A to the right ear of the user 1102 exceeds a distance 1108A from the virtual bee 1104A to the left ear of the user 1102. Thus, at the first time, sound from the virtual bee 1104A should take longer to reach the right ear than the left ear. At a subsequent time, a distance 1110B from the virtual bee 1104B to the right ear exceeds a distance 1108B from the virtual bee 1104B to the left ear. Therefore, at a subsequent time, the sound from virtual bee 1104B should take longer to reach the right ear than the left ear. Comparing distances 1110A and 1110B, distance 1110A exceeds distance 1110B, so the sound from virtual bee 1104A at a first time should take longer to reach the right ear than the sound from virtual bee 1104B at a subsequent time. Comparing distances 1108A and 1108B, the sound from virtual bee 1104A at a first time should take the same time to reach the left ear as the sound from virtual bee 1104B at a subsequent time.

[0102] 11B illustrates a delay module 1112 corresponding to the environment 1100 of FIG. 11A, according to some embodiments. The delay module 1112 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1112 receives an input signal 1114 and outputs a left-ear delay 1116 and a right-ear delay 1118. The delay module 1112 includes a delay unit 1120 and a right cross-fader 1130. The right cross-fader 1130 includes a first level fader (Gf) 1122, a trailing level fader (Gs) 1124, and a summer 1126.

[0103] At a first time, distance 1110A exceeds distance 1108A. For the first time, input signal 1114 is fed directly to left ear delay 1116, and delay unit 1120 delays input signal 1114 by a first time and feeds input signal 1114 delayed by the first time to first level fader 1122.

[0104] At a subsequent time, distance 1110B is greater than distance 1108A, and distance 1110B is less than distance 1110A. For a subsequent time, input signal 1114 is provided directly to left ear delay 1116, and delay unit 1120 delays input signal 1114 by a subsequent time and provides input signal 1114 delayed by a subsequent time to subsequent level fader 1124.

[0105] The first time delayed input signal 1114 may be delayed more than the subsequent time delayed input signal 1114 because distance 1110A exceeds distance 1110B.

[0106] A summer 1126 sums the outputs of the first level fader 1122 and the subsequent level fader 1124 to create the left ear delay 1116 .

[0107] Thus, the right crossfader 1130 crossfades between the input signal 1114 delayed by a first time and the input signal 1114 delayed by a subsequent time.

[0108] 12A illustrates an environment 1200 including a user 1202 and a virtual object (bee) 1204A to the left of a median plane 1206 at a first time and a virtual object (bee) 1204B that is further to the left of (e.g., further from) the median plane 1206 at a subsequent time, according to some embodiments. At the first time, a distance 1210A from the virtual bee 1204A to the right ear of the user 1202 exceeds a distance 1208A from the virtual bee 1204A to the left ear of the user 1202. Thus, at the first time, sound from the virtual bee 1204A should take longer to reach the right ear than the left ear. At a subsequent time, a distance 1210B from the virtual bee 1204B to the right ear exceeds a distance 1208A from the virtual bee 1204B to the left ear. Therefore, at a subsequent time, the sound from virtual bee 1204B should take longer to reach the right ear than the left ear. Comparing distances 1210A and 1210B, distance 1210B exceeds distance 1210A, so the sound from virtual bee 1204B at a subsequent time should take longer to reach the right ear than the sound from virtual bee 1204A at the first time. Comparing distances 1208A and 1208B, the sound from virtual bee 1204A at the first time should take the same time to reach the left ear as the sound from virtual bee 1204B at a subsequent time.

[0109] 12B illustrates a delay module 1212 corresponding to the environment 1200 of FIG. 12A, according to some embodiments. The delay module 1212 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1212 receives an input signal 1214 and outputs a left ear delay 1216 and a right ear delay 1218. The delay module 1212 includes a delay unit 1220 and a right crossfader 1230. The right crossfader 1230 includes a first level fader (Gf) 1222, a trailing level fader (Gs) 1224, and a summer 1226.

[0110] At a first time, distance 1210A exceeds distance 1208A. For the first time, input signal 1214 is fed directly to left ear delay 1216, and delay unit 1220 delays input signal 1214 by a first time and feeds input signal 1214 delayed by the first time to first level fader 1222.

[0111] At a subsequent time, distance 1210B exceeds distance 1208B, and distance 1210B exceeds distance 1210A. For a subsequent time, input signal 1214 is provided directly to left ear delay 1216, and delay unit 1220 delays input signal 1214 by a subsequent time and provides input signal 1214 delayed by a subsequent time to subsequent level fader 1224.

[0112] The first time delayed input signal 1214 may be delayed less than the subsequent time delayed input signal 1214 because distance 1210A is less than distance 1210B.

[0113] A summer 1226 sums the outputs of the first level fader 1222 and the subsequent level fader 1224 to create the right ear delay 1216 .

[0114] Thus, the left crossfader 1230 crossfades between the input signal 1214 delayed by a first time and the input signal 1214 delayed by a subsequent time.

[0115] 13A illustrates an environment 1300 including a user 1302 and a virtual object (bee) 1304A on the right side of a median plane 1306 at a first time and a virtual object (bee) 1304B on the median plane 1306 at a subsequent time, according to some embodiments. At the first time, a distance 1308A from the virtual bee 1304A to the left ear of the user 1302 exceeds a distance 1310A from the virtual bee 1304A to the right ear of the user 1302. Thus, at the first time, sound from the virtual bee 1304A will take longer to reach the left ear than the right ear. At a subsequent time, a distance 1308B from the virtual bee 1304B to the left ear is the same as a distance 1310B from the virtual bee 1304B to the right ear. Thus, at a subsequent time, sound from the virtual bee 1304B will take the same time to reach the left ear as the right ear. Comparing distances 1308A and 1308B, distance 1308A exceeds distance 1308B, so sound from virtual bee 1304A at a first time should take longer to reach the left ear than sound from virtual bee 1304B at a subsequent time. Comparing distances 1310A and 1310B, sound from virtual bee 1304A at a first time should take the same time to reach the right ear as sound from virtual bee 1304B at a subsequent time.

[0116] 13B illustrates a delay module 1312 corresponding to the environment 1300 of FIG. 13A, according to some embodiments. The delay module 1312 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1312 receives an input signal 1314 and outputs a left ear delay 1316 and a right ear delay 1318. The delay module 1312 includes a delay unit 1320 and a left crossfader 1330. The left crossfader 1330 includes a first level fader (Gf) 1322, a trailing level fader (Gs) 1324, and a summer 1326.

[0117] At a first time, distance 1308A exceeds distance 1310A. For the first time, input signal 1314 is fed directly to right ear delay 1318, and delay unit 1320 delays input signal 1314 by a first time and feeds input signal 1314 delayed by the first time to first level fader 1322.

[0118] At a subsequent time, distance 1308B is the same as distance 1310B, and distance 1308B is less than distance 1308A. For a subsequent time, input signal 1314 is fed directly to right ear delay 1318, and input signal 1314 is fed directly to subsequent level fader 1324.

[0119] A summer 1326 sums the outputs of the first level fader 1322 and the subsequent level fader 1324 to create the left ear delay 1316 .

[0120] Thus, the left crossfader 1330 crossfades between the input signal 1314 delayed by the first time and the input signal 1314 .

[0121] 14A illustrates an environment 1400 including a user 1402 and a virtual object (bee) 1404A on a median plane 1406 at a first time and a virtual object (bee) 1404B to the right of the median plane 1406 at a subsequent time, according to some embodiments. At the first time, a distance 1408A from the virtual bee 1404A to the left ear of the user 1402 is the same as a distance 1410A from the virtual bee 1404A to the right ear of the user 1402. Therefore, at the first time, sound from the virtual bee 1404A should take the same time to reach the left ear as it does the right ear. At a subsequent time, a distance 1408B from the virtual bee 1404B to the left ear exceeds a distance 1410A from the virtual bee 1404A to the right ear. Therefore, at a subsequent time, sound from the virtual bee 1404B should take longer to reach the left ear than the right ear. Comparing distances 1408A and 1408B, distance 1408B exceeds distance 1408A, so the sound from virtual bee 1404B at the subsequent time should take longer to reach the left ear than the sound from virtual bee 1404A at the first time. Comparing distances 1410A and 1410B, the sound from virtual bee 1404A at the first time should take the same time to reach the right ear as the sound from virtual bee 1404B at the subsequent time.

[0122] 14B illustrates a delay module 1412 corresponding to the environment 1400 of FIG. 14A, according to some embodiments. The delay module 1412 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1412 receives an input signal 1414 and outputs a left ear delay 1416 and a right ear delay 1418. The delay module 1412 includes a delay unit 1420 and a left crossfader 1430. The left crossfader 1430 includes a first level fader (Gf) 1422, a trailing level fader (Gs) 1424, and a summer 1426.

[0123] At a first time, distance 1408A is the same as distance 1410A. For a first time, input signal 1414 is fed directly to right ear delay 1418, and input signal 1414 is fed directly to first level fader 1422.

[0124] At a subsequent time, distance 1408B exceeds distance 1410B. For a subsequent time, input signal 1414 is fed directly to right ear delay 1418, and delay unit 1420 delays input signal 1414 by a subsequent time and feeds input signal 1414 delayed by a subsequent time to subsequent level fader 1424.

[0125] A summer 1426 sums the outputs of the first level fader 1422 and the subsequent level fader 1424 to create the left ear delay 1416 .

[0126] Thus, the left crossfader 1430 crossfades between the input signal 1414 and the input signal 1414 delayed by a subsequent time.

[0127] 15A illustrates an environment 1500 including a user 1502 and a virtual object (bee) 1504A on the left side of a median plane 1506 at a first time and a virtual object (bee) 1504B on the median plane 1506 at a subsequent time, according to some embodiments. At the first time, a distance 1510A from the virtual bee 1504A to the right ear of the user 1502 exceeds a distance 1508A from the virtual bee 1504A to the left ear of the user 1502. Thus, at the first time, sound from the virtual bee 1504A will take longer to reach the right ear than the left ear. At a subsequent time, a distance 1508B from the virtual bee 1504B to the left ear is the same as a distance 1510B from the virtual bee 1504B to the right ear. Thus, at a subsequent time, sound from the virtual bee 1504B will take the same time to reach the left ear as the right ear. Comparing distances 1510A and 1510B, distance 1510A exceeds distance 1510B, so sound from virtual bee 1504A at a first time should take longer to reach the right ear than sound from virtual bee 1504B at a subsequent time. Comparing distances 1508A and 1508B, sound from virtual bee 1504A at a first time should take the same time to reach the left ear as sound from virtual bee 1504B at a subsequent time.

[0128] 15B illustrates a delay module 1512 corresponding to the environment 1500 of FIG. 15A, according to some embodiments. The delay module 1512 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1512 receives an input signal 1514 and outputs a left ear delay 1516 and a right ear delay 1518. The delay module 1512 includes a delay unit 1520 and a right crossfader 1530. The right crossfader 1530 includes a first level fader (Gf) 1522, a trailing level fader (Gs) 1524, and a summer 1526.

[0129] At a first time, distance 1510A exceeds distance 1508A. For the first time, input signal 1514 is fed directly to left ear delay 1516, and delay unit 1520 delays input signal 1514 by a first time and feeds input signal 1514 delayed by the first time to first level fader 1522.

[0130] At a subsequent time, distance 1508B is the same as distance 1510B, which is less than distance 1510A. At a subsequent time, input signal 1514 is fed directly to left ear delay 1516, and input signal 1514 is fed directly to subsequent level fader 1524.

[0131] A summer 1526 sums the outputs of the first level fader 1522 and the subsequent level fader 1524 to create the right ear delay 1518 .

[0132] Thus, the right crossfader 1530 crossfades between the input signal 1514 delayed by the first time and the input signal 1514 .

[0133] 16A illustrates an environment 1600 including a user 1602 and a virtual object (bee) 1604A on a median plane 1606 at a first time and a virtual object (bee) 1604B to the left of the median plane 1606 at a subsequent time, according to some embodiments. At the first time, a distance 1608A from the virtual bee 1604A to the left ear of the user 1602 is the same as a distance 1610A from the virtual bee 1604A to the right ear of the user 1602. Therefore, at the first time, sound from the virtual bee 1604A should take the same time to reach the left ear as it should take the right ear. At a subsequent time, a distance 1610B from the virtual bee 1604B to the right ear exceeds a distance 1608B from the virtual bee 1604A to the left ear. Therefore, at a subsequent time, sound from the virtual bee 1604B should take longer to reach the right ear than the left ear. Comparing distances 1610A and 1610B, distance 1610B exceeds distance 1610A, so the sound from virtual bee 1604B at the subsequent time should take longer to reach the right ear than the sound from virtual bee 1604A at the first time. Comparing distances 1608A and 1608B, the sound from virtual bee 1604A at the first time should take the same time to reach the left ear as the sound from virtual bee 1604B at the subsequent time.

[0134] 16B illustrates a delay module 1612 corresponding to the environment 1600 of FIG. 16A, according to some embodiments. The delay module 1612 may be used to implement the delay module 105 shown in the exemplary system 100. The delay module 1612 receives an input signal 1614 and outputs a left ear delay 1616 and a right ear delay 1618. The delay module 1612 includes a delay unit 1620 and a right crossfader 1630. The right crossfader 1630 includes a first level fader (Gf) 1622, a trailing level fader (Gs) 1624, and a summer 1626.

[0135] At a first time, distance 1608A is the same as distance 1610A. For a first time, input signal 1614 is fed directly to left ear delay 1616, and input signal 1614 is fed directly to first level fader 1622.

[0136] At a subsequent time, distance 1610B exceeds distance 1608B. For a subsequent time, input signal 1614 is provided directly to left ear delay 1616, and delay unit 1620 delays input signal 1614 by a subsequent time and provides input signal 1614 delayed by a subsequent time to subsequent level fader 1624.

[0137] A summer 1626 sums the outputs of the first level fader 1622 and the subsequent level fader 1624 to create the right ear delay 1618 .

[0138] Thus, the right crossfader 1630 crossfades between the input signal 1614 and the input signal 1614 delayed by a subsequent time.

[0139] 17 illustrates an example delay module 1705 that, in some embodiments, may be used to implement the delay module 105 shown in the example system 100. In some embodiments, for example, as illustrated in FIG. 17, the delay module 1705 may include one or more filters (e.g., a common filter FC 1756, a first filter F1 1752, and a second filter F2 1754). The first filter F1 1752 and the second filter F2 1754 may be used to model one or more effects of sound when a sound source is in the near field, for example. For example, the first filter F1 1752 and the second filter F2 1754 may be used to model one or more effects of sound when a sound source moves closer to or away from a speaker / ear position. The common filter FC 1756 may be used to model one or more effects, such as a sound source being obstructed by objects, air absorption, and the like, that may affect the signal to both ears. The first filter F1 1752 may apply a first effect, the second filter F2 1754 may apply a second effect, and the common filter FC 1756 may apply a third effect.

[0140] In the illustrated embodiment, an input signal 1702 is input to a delay module 1705; for example, the input signal 1702 can be applied to the input of a common filter FC 1756. The common filter FC 1756 applies one or more filters to the input signal 1702 and outputs a common filtered signal. The common filtered signal is input to both a first filter F1 1752 and a delay unit 1716. The first filter F1 1752 applies one or more filters to the common filtered signal and outputs a first filtered signal, referred to as a first ear delay 1722. The delay unit 1716 applies a delay to the common filtered signal and outputs a delayed common filtered signal. The second filter F2 1754 applies one or more filters to the delayed common filtered signal and outputs a second filtered signal, referred to as a second ear delay 1724. In some embodiments, the first ear delay 1722 may correspond to the left ear and the second ear delay 1724 may correspond to the right ear. In some embodiments, the first ear delay 1722 may correspond to the right ear and the second ear delay 1724 may correspond to the left ear.

[0141] 17 may not be required, namely, common filter FC 1756, first filter F1 1752, and second filter F2 1754. In one example, because the signal input to first filter F1 1752 and the signal input to second filter F2 1754 both have the effect of common filter FC 1756 applied to them, the common filter FC 1756 setting may be applied / added to each of first filter F1 1752 and second filter F2 1754, and common filter FC 1756 may be removed, thus reducing the total number of filters from three to two.

[0142] The delay module 1705 may be similar to the delay module 205 of FIG. 2B, with the first ear delay 226 of FIG. 2B corresponding to the second ear delay 1724 of FIG. 17, and the second ear delay 228 of FIG. 2B corresponding to the first ear delay 1722 of FIG. 17. In this embodiment, the first ear delay 1722 does not have any delay, and the second ear delay 1724 has a delay. This may be the case, for example, when a sound source is closer to the first ear than the second ear, and the first ear receives the first ear delay 1722 and the second ear receives the second ear delay 1724. Those skilled in the art will understand that while the following description primarily relates to the variant of FIG. 2B, the principles may apply to the variants of FIGS. 2A and 2C as well.

[0143] 18A-18E illustrate variations of the delay module 1805 according to some embodiments. Any of the variations of the delay module 1805 shown in FIGS. 18A-18E may be used to implement the delay module 105 shown in the exemplary system 100. FIG. 18A illustrates a delay module 1805 without any filters. The delay module 1805 may not require any filters, for example, when the sound source is in the far field. FIG. 18B illustrates a delay module 1805 with only the first filter F1 1852. The delay module 1805 may require only the first filter F1 1852, for example, when the sound source is closer to the first ear and only the first ear is obstructed by an object. FIG. 18C illustrates a delay module 1805 with only the second filter F2 1854. The delay module 1805 may require only the second filter F2 1854, for example, when the sound source is farther away from the second ear and only the second ear is obstructed by an object. Figure 18D illustrates the delay module 1805 with the first filter F1 1852 and the second filter F2 1854, where the first filter F1 1852 and the second filter F2 1854 are different. The delay module 1805 may require the first filter F1 1852 and the second filter F2 1854, for example, when the sound source is closer to the first ear and each ear is obstructed by an object of a different size. Figure 18E illustrates the delay module 1805 with only the common filter FC 1856. The delay module 1805 may only require the common filter CF 1856 when, for example, the source is in the far field and both ears are equally obstructed or air absorption is present.

[0144] In some embodiments, any one of the delay modules illustrated in Figures 18A-18E may transition to any of the other delay modules illustrated in Figures 18A-18E due to a change in the soundscape, such as the movement of an interfering object or a sound source relative to it.

[0145] Transitioning from the delay module 1805 illustrated in Figure 18A (which does not include any filters) to any of the delay modules 1805 illustrated in Figures 18B-18E (each of which includes one or more filters) may involve simply introducing one or more filters at the appropriate / desired time. Similarly, transitioning from the delay module 1805 illustrated in Figures 18B-18E to the delay module 1805 illustrated in Figure 18A may involve simply removing one or more filters at the appropriate / desired time.

[0146] 18B (including the first filter F1 1852) to the delay module 1805 (including the second filter F2 1854) illustrated in FIG. 18C may involve removing the first filter F1 1852 and adding the second filter F2 1854 at the appropriate / desired time. Similarly, the transition from the delay module 1805 (including the second filter F2 1854) illustrated in FIG. 18C to the delay module 1805 (including the first filter F1 1852) illustrated in FIG. 18B may involve removing the second filter F2 1854 and adding the first filter F1 1852 at the appropriate / desired time.

[0147] 18B (including the first filter F1 1852) to the delay module 1805 (including the first filter F1 1852 and the second filter F2 1854) illustrated in FIG. 18D may include adding, at the appropriate / desired time, the second filter F2 1854. Similarly, the transition from the delay module 1805 (including the first filter F1 1852) and the second filter F2 1854 illustrated in FIG. 18D to the delay module 1805 (including the first filter F1 1852) illustrated in FIG. 18B may include removing, at the appropriate / desired time, the second filter F2 1854.

[0148] 18B (including the first filter F1 1852) to the delay module 1805 (including the common filter FC 1856) illustrated in FIG. 18E may include adding the common filter 1856 at the appropriate / desired time, copying the state of the first filter F1 1852 to the common filter FC 1856, and removing the first filter F1 1852. Similarly, the transition from the delay module 1805 (including the common filter FC 1856) illustrated in FIG. 18E to the delay module 1805 (including the first filter F1 1852) illustrated in FIG. 18B may include adding the first filter F1 1852 at the appropriate / desired time, copying the state of the common filter FC 1856 to the first filter F1 1852, and removing the common filter FC 1856.

[0149] 18C (including second filter F2 1854) to the delay module 1805 (including first filter F1 1852 and second filter F2 1854) illustrated in FIG. 18D may include adding, at the appropriate / desired time, the first filter F1 1852. Similarly, the transition from the delay module 1805 (including first filter F1 1852) and second filter F2 1854 illustrated in FIG. 18D to the delay module 1805 (including second filter F2 1854) illustrated in FIG. 18C may include removing the first filter F1 1852 at the appropriate / desired time.

[0150] The transition from the delay module 1805 illustrated in FIG. 18C (including the second filter F2 1854) to the delay module 1805 illustrated in FIG. 18E (including the common filter FC 1856) may include performing a process such as that illustrated by the example of FIG. 19. In step 1902 of the example process, the common filter FC 1856 is added and the state of the second filter F2 1854 is copied to the common filter FC 1856. This may occur at time T1. In 1904, the system waits a delay time. The delay time is the amount of time the delay unit 1816 delays the signal. In 1906, the second filter F2 1854 is removed. This may occur at time T2.

[0151] The delay unit 1816 includes a first-in, first-out buffer. Before time T1, the buffer of the delay unit 1816 is filled with the input signal 1802. The second filter F2 1854 filters the output of the delay unit 1816, which contains only the input signal 1802 from before time T1. Between time T1 and time T2, the common filter FC 1856 filters the input signal 1802, and the buffer of the delay unit 1816 is filled with both the input signal 1802 from before T1 and the filtered input signal from between time T1 and time T2. The second filter F2 1854 filters the output of the delay unit 1816, which contains only the input signal 1802 from before time T1. At time T2, the second filter 1854 is removed, and the delay unit 1816 is filled with only the filtered input signal starting at time T1.

[0152] In some embodiments, the transition from the delay module 1805 illustrated in FIG. 18C (including the second filter F2 1854) to the delay module 1805 illustrated in FIG. 18E (including the common filter FC 1856) may include processing all samples in the delay unit 1816 with the second filter F2 1854 (or with another filter having the same settings as the second filter F2 1854), writing the processed samples to the delay unit 1816, adding a filter for the common filter FC 1856, copying the state of the second filter F2 1854 to the common filter FC 1856, and removing the second filter F2 1854. In some embodiments, all of the aforementioned steps may occur at time T1. That is, all of the aforementioned steps may occur simultaneously (or nearly simultaneously). In some embodiments, the delay unit 1816 includes a first-in, first-out buffer. In these embodiments, when processing all samples in delay unit 1816, processing may proceed from the end of the buffer to the beginning (ie, from oldest sample to newest).

[0153] The transition from the delay module 1805 illustrated in FIG. 18E (including the common filter FC 1856) to the delay module 1805 illustrated in FIG. 18C (including the second filter F2 1854) may include performing a process such as that illustrated by the embodiment of FIG. 20. At 2002, the state of the common filter FC 1856 is saved. This may occur at time T1. At 2004, the system waits a delay time. The delay time is the amount of time the delay unit 1816 delays the signal. At 2006, the second filter F2 1854 is added, the saved state of the common filter FC 1856 is copied to the second filter F2 1854, and the common filter FC 1856 is removed. This may occur at time T2.

[0154] The delay unit 1816 includes a first-in, first-out buffer. Before time T1, the common filter FC 1856 filters the input signal 1802, and the buffer of the delay unit 1816 fills with the filtered input signal. Between time T1 and time T2, the common filter FC 1856 continues to filter the input signal 1802, and the buffer of the delay unit 1816 continues to fill with the filtered input signal. At time T2, a second filter F2 1854 is added, the saved state of the common filter FC 1856 is copied to the second filter F2 1854, and the common filter FC 1856 is removed.

[0155] The transition from the delay module 1805 illustrated in FIG. 18D (including the first filter F1 1852 and the second filter F2 1854) to the delay module 1805 illustrated in FIG. 18E (including the common filter FC 1856) may include performing the process illustrated by the example of FIG. 21. At 2102, the common filter FC 1856 is added, the state of the first filter F1 1852 is copied to the common filter FC 1856, and the first filter F1 1852 is removed. This may occur at time T1. At 2104, the system waits a delay time. The delay time is the amount of time the delay unit 1816 delays the signal. At 2106, the second filter F2 1854 is removed. This may occur at time T2.

[0156] The delay unit 1816 includes a first-in, first-out buffer. Before time T1, the buffer of the delay unit 1816 is filled with the input signal 1802. The second filter F2 1854 filters the output of the delay unit 1816, which contains only the input signal 1802 from before time T1. Between time T1 and time T2, the common filter FC 1856 filters the input signal 1802, and the buffer of the delay unit 1816 is filled with both the input signal 1802 from before T1 and the filtered input signal from between time T1 and time T2. The second filter F2 1854 filters the output of the delay unit 1816, which contains only the input signal 1802 from before time T1. At time T2, the second filter 1854 is removed, and the delay unit 1816 is filled with only the filtered input signal starting at time T1.

[0157] The transition from the delay module 1805 (including the common filter FC 1856) illustrated in FIG. 18E to the delay module 1805 (including the first filter F1 1852 and the second filter F2 1854) illustrated in FIG. 18E may include performing the process illustrated by the example in FIG. 22. At 2202, the state of the common filter FC 1856 is saved. This may occur at time T1. At 2204, the system waits a delay time. The delay time is the amount of time the delay unit 1816 delays the signal. At 2206, a first filter F1 1852 is added, the saved state of common filter FC 1856 is copied to first filter F1 1852, a second filter F2 1854 is added, the saved state of common filter FC 1856 is copied to second filter F2 1854, and common filter FC 1856 is removed. This may occur at time T2.

[0158] The delay unit 1816 includes a first-in, first-out buffer. Before time T1, the common filter FC 1856 filters the input signal 1802, and the buffer of the delay unit 1816 fills with the filtered input signal. Between time T1 and time T2, the common filter FC 1856 continues to filter the input signal 1802, and the buffer of the delay unit 1816 continues to fill with the filtered input signal. At time T2, a first filter F1 1852 is added, the saved state of the common filter FC 1856 is copied to the first filter 1852, a second filter F2 1854 is added, the saved state of the common filter FC 1856 is copied to the second filter F2 1854, and the common filter FC 1856 is removed.

[0159] Various exemplary embodiments of the present disclosure are described herein. These examples are referred to in a non-limiting sense. They are provided to illustrate the more broadly applicable aspects of the present disclosure. Various changes may be made to the disclosed embodiments, and equivalents may be substituted without departing from the true spirit and scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation, material, composition, process, process act, or step to the objective, spirit, or scope of the present disclosure. Furthermore, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein has discrete components and features that can be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. All such modifications are intended to be within the scope of the claims associated with this disclosure.

[0160] The present disclosure includes methods that may be implemented using the subject devices. The methods may include the act of providing such a suitable device. Such provisioning may be performed by an end user. In other words, the act of "providing" merely requires the end user to obtain, access, approach, locate, configure, activate, power on, or otherwise act to provide the device required in the subject methods. The methods recited herein may occur in any order of the recited events and recited sequence of events that is logically possible.

[0161] Exemplary aspects of the present disclosure are described above, along with details regarding material selection and manufacturing. As for other details of the present disclosure, these may be understood in connection with the above-referenced patents and publications and are generally known or may be understood by those skilled in the art. The same may be true with respect to the method-based aspects of the present disclosure in terms of additional acts as commonly or logically adopted.

[0162] Additionally, while the present disclosure has been described with reference to several embodiments incorporating various features, the present disclosure is not limited to that described or illustrated, as each variation of the disclosure is discussed. Various modifications may be made to the present disclosure as described, and equivalents (whether recited herein or not included for purposes of brevity to some extent) may be substituted without departing from the true spirit and scope of the present disclosure. Additionally, when a range of values ​​is provided, it is understood that all intervening values ​​between the upper and lower limits of that range, and any other stated value or intervening values ​​within the stated range, are encompassed within the present disclosure.

[0163] It is also contemplated that any optional features of the described inventive variations may be set forth and claimed independently or in combination with any one or more of the features described herein. Reference to a singular item includes the possibility that multiple identical items are present. More specifically, as used in this specification and the claims associated therewith, the singular forms "a," "an," "said," and "the" include plural references unless specifically stated otherwise. In other words, the use of articles allows for "at least one" of the items of the present subject matter in the above description and in the claims associated with this disclosure. Furthermore, it should be noted that such claims may be drafted to exclude any optional element. Accordingly, this language is intended to serve as a predicate for the use of exclusive terminology such as "solely," "only," and the like in connection with the recitation of claim elements, or the use of a "negative" limitation.

[0164] Without the use of such exclusive terminology, the term "comprising" in a claim associated with this disclosure shall be deemed to permit the inclusion of any additional elements, regardless of whether a given number of elements are recited in such a claim, or the addition of features may be deemed to change the nature of the elements recited in such a claim. Except as specifically defined herein, all technical and scientific terms used herein should be given the broadest commonly understood meaning possible while maintaining claim legitimacy.

[0165] The scope of the present disclosure is not intended to be limited to the examples provided and / or this specification, but rather is intended to be limited only by the scope of the terms of the claims associated with this disclosure.

Claims

1. 1. A method of presenting an audio signal to a user of a wearable head device, the method comprising: receiving a first input audio signal, the first input audio signal corresponding to a first source location within a virtual environment presented to the user via the wearable head device, the first source location corresponding to a first location of a virtual object within the virtual environment at a first time; processing the first input audio signal to generate a left output audio signal and a right output audio signal, wherein processing the first input audio signal comprises: applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal, wherein applying the delay process includes applying a first filter to the first input audio signal to generate a first filtered signal, applying a second filter to the first filtered signal to generate one of the left audio signal and the right audio signal, applying a delay process to the first filtered signal to generate a delayed signal, and applying a third filter to the delayed signal to generate the other of the left audio signal and the right audio signal; adjusting a gain of the left audio signal; adjusting a gain of the right audio signal; applying a first head-related transfer function (HRTF) to the left audio signal to generate the left output audio signal; applying a second HRTF to the right audio signal to generate the right output audio signal; and presenting the left output audio signal to the user's left ear via a left speaker associated with the wearable head device; presenting the right output audio signal to the user's right ear via a right speaker associated with the wearable head device; determining a second source location corresponding to a second location of the virtual object in the virtual environment at a second time, wherein the virtual object is at the first location in the virtual environment at the first time and the virtual object is at the second location in the virtual environment at the second time, and the first location in the virtual environment and the second location in the virtual environment are different locations in the virtual environment relative to the user's position; determining a first ear delay, wherein determining the first ear delay includes determining a preceding first ear delay corresponding to the first time and a subsequent first ear delay corresponding to the second time, and crossfading between the preceding first ear delay and the subsequent first ear delay to generate the first ear delay; determining a second ear delay, wherein determining the second ear delay includes determining a preceding second ear delay corresponding to the first time and a subsequent second ear delay corresponding to the second time, and crossfading between the preceding second ear delay and the subsequent second ear delay to generate the second ear delay; Including, 11. The method of claim 10, wherein applying the delay process to the first input audio signal includes applying an interaural time delay (ITD) to the first input audio signal, the ITD being determined based on the first source location, the first ear delay, and the second ear delay.

2. The method of claim 1 , wherein the first ear delay is zero.

3. The method of claim 1 , further comprising applying a filter to the left audio signal.

4. The method of claim 1 , further comprising applying a filter to the right audio signal.

5. 2. The method of claim 1, wherein applying the delay process includes transitioning from a first delay module at the first time to a second delay module at the second time, the second delay module being different from the first delay module.

6. 6. The method of claim 5, wherein the first delay module is associated with applying first one or more filters to a first one or more of the first input audio signal, the left audio signal, and the right audio signal, and the second delay module is associated with applying second one or more filters to a second one or more of the first input audio signal, the left audio signal, and the right audio signal.

7. The method of claim 1 , wherein the first ear corresponds to the user's left ear and the second ear corresponds to the user's right ear.

8. The method of claim 1 , wherein the first ear corresponds to the user's right ear and the second ear corresponds to the user's left ear.

9. The method of claim 1 , wherein the first source location is closer to the first ear than to the second ear, and the second source location is closer to the second ear than to the first ear.

10. The method of claim 1 , wherein the second source location is closer to the first ear than to the second ear, and the first source location is closer to the second ear than to the first ear.

11. 2. The method of claim 1, wherein the first source location is closer to the first ear than the second source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear.

12. 2. The method of claim 1, wherein the second source location is closer to the first ear than the first source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear.

13. 1. A system comprising: a wearable head device; a left speaker associated with the wearable head device; a right speaker associated with the wearable head device; One or more processors, the one or more processors configured to perform a method, the method comprising: receiving a first input audio signal, the first input audio signal corresponding to a first source location within a virtual environment presented to a user via the wearable head device, the first source location corresponding to a first location of a virtual object within the virtual environment at a first time; processing the first input audio signal to generate a left output audio signal and a right output audio signal, wherein processing the first input audio signal comprises: applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal, wherein applying the delay process includes applying a first filter to the first input audio signal to generate a first filtered signal, applying a second filter to the first filtered signal to generate one of the left audio signal and the right audio signal, applying a delay process to the first filtered signal to generate a delayed signal, and applying a third filter to the delayed signal to generate the other of the left audio signal and the right audio signal; adjusting a gain of the left audio signal; adjusting a gain of the right audio signal; applying a first head-related transfer function (HRTF) to the left audio signal to generate the left output audio signal; applying a second HRTF to the right audio signal to generate the right output audio signal; and presenting the left output audio signal to the left ear of the user via the left speaker; presenting the right output audio signal to the user's right ear via the right speaker; determining a second source location corresponding to a second location of the virtual object in the virtual environment at a second time, wherein the virtual object is at the first location in the virtual environment at the first time and the virtual object is at the second location in the virtual environment at the second time, and the first location in the virtual environment and the second location in the virtual environment are different locations in the virtual environment relative to the user's position; determining a first ear delay, wherein determining the first ear delay includes determining a preceding first ear delay corresponding to the first time and a subsequent first ear delay corresponding to the second time, and crossfading between the preceding first ear delay and the subsequent first ear delay to generate the first ear delay; determining a second ear delay, wherein determining the second ear delay includes determining a preceding second ear delay corresponding to the first time and a subsequent second ear delay corresponding to the second time, and crossfading between the preceding second ear delay and the subsequent second ear delay to generate the second ear delay; one or more processors, Equipped with applying the delay process to the first input audio signal includes applying an interaural time delay (ITD) to the first input audio signal, the ITD being determined based on the first source location, the first ear delay, and the second ear delay.

14. The system of claim 13 , wherein the first ear delay is zero.

15. The system of claim 13 , wherein the method further comprises applying a filter to the left audio signal.

16. The system of claim 13 , wherein the method further comprises applying a filter to the right audio signal.

17. 14. The system of claim 13, wherein applying the delay process includes transitioning from a first delay module at the first time to a second delay module at the second time, the second delay module being different from the first delay module.

18. 18. The system of claim 17, wherein the first delay module is associated with applying first one or more filters to a first one or more of the first input audio signal, the left audio signal, and the right audio signal, and the second delay module is associated with applying second one or more filters to a second one or more of the first input audio signal, the left audio signal, and the right audio signal.

19. The system of claim 13 , wherein the first ear corresponds to the user's left ear and the second ear corresponds to the user's right ear.

20. The system of claim 13 , wherein the first ear corresponds to the user's right ear and the second ear corresponds to the user's left ear.

21. The system of claim 13 , wherein the first source location is closer to the first ear than to the second ear, and the second source location is closer to the second ear than to the first ear.

22. The system of claim 13 , wherein the second source location is closer to the first ear than to the second ear, and the first source location is closer to the second ear than to the first ear.

23. 14. The system of claim 13, wherein the first source location is closer to the first ear than the second source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear.

24. 14. The system of claim 13, wherein the second source location is closer to the first ear than the first source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear.

25. 1. A non-transitory computer-readable medium containing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for presenting an audio signal to a user of a wearable head device, the method comprising: receiving a first input audio signal, the first input audio signal corresponding to a first source location within a virtual environment presented to the user via the wearable head device, the first source location corresponding to a first location of a virtual object within the virtual environment at a first time; processing the first input audio signal to generate a left output audio signal and a right output audio signal, wherein processing the first input audio signal comprises: applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal, wherein applying the delay process includes applying a first filter to the first input audio signal to generate a first filtered signal, applying a second filter to the first filtered signal to generate one of the left audio signal and the right audio signal, applying a delay process to the first filtered signal to generate a delayed signal, and applying a third filter to the delayed signal to generate the other of the left audio signal and the right audio signal; adjusting a gain of the left audio signal; adjusting a gain of the right audio signal; applying a first head-related transfer function (HRTF) to the left audio signal to generate the left output audio signal; applying a second HRTF to the right audio signal to generate the right output audio signal; and presenting the left output audio signal to the user's left ear via a left speaker associated with the wearable head device; presenting the right output audio signal to the user's right ear via a right speaker associated with the wearable head device; determining a second source location corresponding to a second location of the virtual object in the virtual environment at a second time, wherein the virtual object is at the first location in the virtual environment at the first time and the virtual object is at the second location in the virtual environment at the second time, and the first location in the virtual environment and the second location in the virtual environment are different locations in the virtual environment relative to the user's position; determining a first ear delay, wherein determining the first ear delay includes determining a preceding first ear delay corresponding to the first time and a subsequent first ear delay corresponding to the second time, and crossfading between the preceding first ear delay and the subsequent first ear delay to generate the first ear delay; determining a second ear delay, wherein determining the second ear delay includes determining a preceding second ear delay corresponding to the first time and a subsequent second ear delay corresponding to the second time, and crossfading between the preceding second ear delay and the subsequent second ear delay to generate the second ear delay; Including, a first input audio signal processing unit configured to process the first input audio signal based on the first source location, the first ear delay, and the second ear delay;

26. 26. The non-transitory computer-readable medium of claim 25, wherein the first ear delay is zero.

27. 26. The non-transitory computer-readable medium of claim 25, wherein the method further comprises applying a filter to the left audio signal.

28. 26. The non-transitory computer-readable medium of claim 25, wherein the method further comprises applying a filter to the right audio signal.

29. 26. The non-transitory computer-readable medium of claim 25, wherein applying the delay process includes transitioning from a first delay module at the first time to a second delay module at the second time, the second delay module being different from the first delay module.

30. 30. The non-transitory computer-readable medium of claim 29, wherein the first delay module is associated with applying first one or more filters to a first one or more of the first input audio signal, the left audio signal, and the right audio signal, and the second delay module is associated with applying second one or more filters to a second one or more of the first input audio signal, the left audio signal, and the right audio signal.

31. 26. The non-transitory computer-readable medium of claim 25, wherein the first ear corresponds to the user's left ear and the second ear corresponds to the user's right ear.

32. 26. The non-transitory computer-readable medium of claim 25, wherein the first ear corresponds to the user's right ear and the second ear corresponds to the user's left ear.

33. 26. The non-transitory computer-readable medium of claim 25, wherein the first source location is closer to the first ear than to the second ear and the second source location is closer to the second ear than to the first ear.

34. 26. The non-transitory computer-readable medium of claim 25, wherein the second source location is closer to the first ear than to the second ear and the first source location is closer to the second ear than to the first ear.

35. 26. The non-transitory computer-readable medium of claim 25, wherein the first source location is closer to the first ear than the second source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear.

36. 26. The non-transitory computer-readable medium of claim 25, wherein the second source location is closer to the first ear than the first source location, the first source location is closer to the first ear than the second ear, and the second source location is closer to the first ear than the second ear.

37. determining a first sound effect; determining the first one or more filters to apply according to the first sound effect; determining a second sound effect; determining the second one or more filters to apply in accordance with the second sound effect; and The method of claim 6 further comprising:

Citation Information

Patent Citations

  • JPP4499358B

  • Augmeted reality system with spatialized audio tied to user manipulated virtual object

    WO2018183390A1