Dual listener positions for mixed reality

The method and system for stereo audio presentation in mixed reality environments address VR's immersion challenges by using sensors to simulate realistic sound propagation, enhancing user experience through accurate sound localization and reduced computational demands.

JP2025105787AActive Publication Date: 2025-07-10MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025071289
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-02-15
Filing Date
2025-04-23
Publication Date
2025-07-10
Estimated Expiration
2039-02-15

AI Technical Summary

Technical Problem

Conventional VR systems face challenges such as motion sickness, disorientation, high computational burden, and inability to utilize real-world sensory data, while AR and MR systems can reduce these issues by maintaining the real environment but struggle with authentic audio cues and immersive sound localization.

Method used

A method and system for presenting stereo audio signals in a mixed reality environment by identifying the positions of a user's ears and virtual sound sources, applying filters and attenuation based on real and virtual objects to simulate realistic sound propagation, using sensors like depth cameras to determine object characteristics.

Benefits of technology

Enhances user immersion by accurately localizing sound sources and simulating realistic audio cues, reducing computational burden and improving the authenticity of the mixed reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105787000001_ABST
    Figure 2025105787000001_ABST
Patent Text Reader

Abstract

To provide a method of presenting audio signals in a mixed reality environment.SOLUTION: A method comprises: identifying a position of a first ear of a listener in a mixed reality environment; identifying a position of a second ear of the listener in the mixed reality environment; identifying a first virtual sound source in the mixed reality environment; identifying a first object in the mixed reality environment; determining a first audio signal in the mixed reality environment; determining a second audio signal in the mixed reality environment; determining a third audio signal on the basis of the second audio signal and the first object; presenting the first audio signal to a first ear of a user via a first speaker; and presenting the third audio signal to a second ear of the user via a second speaker.SELECTED DRAWING: Figure 5B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Field) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 631,422, filed on February 15, 2018, which is incorporated herein by reference in its entirety.

[0002] The present disclosure generally relates to systems and methods for presenting audio signals, and more particularly, to systems and methods for presenting stereo audio signals to a user of an augmented reality system.

Background Art

[0003] (Background) Virtual environments are ubiquitous in computing environments and find use in video games (where the virtual environment may represent a game world), maps (where the virtual environment may represent terrain to be navigated), simulations (where the virtual environment may simulate a real-world environment), digital storytelling (where virtual characters may interact with each other within the virtual environment), and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the user experience with virtual environments can be limited by the technology used to present the virtual environment. For example, conventional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may not be able to realize a virtual environment so as to attract people and create a realistic and immersive experience.

[0004] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present sensory information to a user of an XR system that corresponds to a virtual environment represented by data within a computer system. This disclosure takes into account the specificities among VR, AR, and MR systems (although some systems may be categorized as VR in one aspect (e.g., the visual aspect) and at the same time be categorized as AR or MR in another aspect (e.g., the audio aspect)). As used herein, a VR system presents a virtual environment that replaces a user's real environment in at least one aspect. For example, a VR system may present a view of a virtual environment to a user while simultaneously obscuring that view of the real environment using, e.g., a light-blocking head-mounted display. Similarly, a VR system may present audio corresponding to the virtual environment to a user while simultaneously blocking (attenuating) audio from the real environment.

[0005] A VR system can suffer from various drawbacks arising from replacing the user's real environment with a virtual environment. One drawback can occur when the user's field of view within the virtual environment no longer corresponds to the state of their inner ear that detects its balance and orientation (in the real environment, not the virtual environment), resulting in motion sickness. Similarly, the user can experience a sense of disorientation within the VR environment if their body and limbs (the view of which the user relies on to feel "grounded" in the real environment) are not directly visible. Another drawback is the computational burden (e.g., memory, processing power) imposed on VR systems, especially in real-time applications that attempt to immerse the user in the virtual environment and must present a fully 3D virtual environment. Similarly, such an environment tends to be sensitive to even minor imperfections within the virtual environment, any of which can disrupt the user's sense of immersion and thus may need to achieve a very high level of realism to be considered immersive. Additionally, another drawback of VR systems is that such uses of the system cannot utilize a wide range of sensory data within the real environment, such as various sights and sounds experienced in the real world. A related drawback is that VR systems can struggle to create a shared environment in which multiple users can interact because users sharing physical space within the real environment may not be able to directly see or interact with each other within the virtual environment.

[0006] As used herein, an AR system presents a virtual environment that overlaps or overlays a real environment in at least one aspect. For example, while presenting a displayed image, an AR system can present a view of a virtual environment overlaid on the user's view of the real environment to the user using a transmissive head-mounted display or the like that allows light to pass through the display into the user's eyes. Similarly, an AR system can present audio corresponding to the virtual environment to the user while simultaneously mixing the audio from the real environment. Similarly, as used herein, an MR system presents a virtual environment that overlaps or overlays a real environment in at least one aspect, as in the case of an AR system, and in addition, enables the virtual environment within the MR system to interact with the real environment in at least one aspect. For example, a virtual character within the virtual environment may switch a lighting switch within the real environment and turn on or off the corresponding light bulb within the real environment. As another example, the virtual character may react to an audio signal within the real environment (using, for example, facial expressions). By maintaining the presentation of the real environment, AR and MR systems can avoid some of the aforementioned drawbacks of VR systems. For example, motion sickness in the user may be reduced because visual cues from the real environment (including the user's own body) can remain visible and such systems do not need to present a fully realized 3D environment to the user because they are immersive. Further, AR and MR systems can create new applications that utilize real-world sensory inputs (e.g., scenery, objects, and the views and sounds of other users) to augment that input.

[0007] An XR system can provide various ways for a user to interact with a virtual environment. For example, the XR system may include various sensors (such as cameras, microphones, etc.) to detect the user's position and orientation, facial expressions, speech, and other characteristics, and present this information as input to the virtual environment. Some XR systems may incorporate sensor-equipped input devices such as virtual "mallets" and may be configured to detect the position, orientation, or other characteristics of the input device.

[0008] An XR system can provide a unique high level of immersion and realism by combining virtual visual and audio cues with real-world scenery and sounds. For example, it may be desirable to present audio cues to a user of an XR system to mimic aspects of the user's sensory experience, particularly subtle aspects. The present invention is directed to presenting a stereo audio signal originating from a single sound source to a user within a mixed reality environment such that the user can identify the position and orientation of the sound source within the mixed reality environment based on the differences in the signals received by the user's left and right ears. By using audio cues to identify the position and orientation of a sound source within a mixed reality environment, the user can experience a high level of perception of virtual sounds arising from that position and orientation. Additionally, the immersion of the user within the mixed reality environment can be enhanced by presenting not only stereo audio corresponding directly to the audio signal but also a fully immersive soundscape generated using a 3D propagation model. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0009] Embodiments of the present disclosure describe systems and methods for presenting audio signals within a mixed reality environment. In one embodiment, the method includes identifying a position of a first ear of a listener within the mixed reality environment; identifying a position of a second ear of the listener within the mixed reality environment; identifying a first virtual sound source within the mixed reality environment; identifying a first object within the mixed reality environment; determining a first audio signal within the mixed reality environment, the first audio signal originating at the first virtual sound source and intersecting the position of the listener's first ear; determining a second audio signal within the mixed reality environment, the second audio signal originating at the first virtual sound source, intersecting the first object, and intersecting the position of the listener's second ear; determining a third audio signal based on the second audio signal and the first object; presenting the first audio signal to the user's first ear via a first speaker; and presenting the third audio signal to the user's second ear via a second speaker. This specification also provides, for example, the following items. (Item 1) A method for presenting an audio signal within a mixed reality environment, the method comprising: identifying a position of a first ear of a listener within the mixed reality environment; identifying a position of a second ear of the listener within the mixed reality environment; identifying a first virtual sound source within the mixed reality environment; identifying a first object within the mixed reality environment; determining a first audio signal within the mixed reality environment, the first audio signal originating at the first virtual sound source and intersecting the position of the listener's first ear; determining a second audio signal within the mixed reality environment, the second audio signal originating at the first virtual sound source, intersecting the first object, and intersecting the position of the listener's second ear; Determining a third audio signal based on the second audio signal and the first object; Presenting the first audio signal to a first ear of the user via a first speaker; Presenting the third audio signal to a second ear of the user via a second speaker A method comprising. (Item 2) Determining the third audio signal from the second audio signal includes applying a low - pass filter to the second audio signal, the low - pass filter having parameters based on the first object, the method according to item 1. (Item 3) Determining the third audio signal from the second audio signal includes applying attenuation to the second audio signal, the intensity of the attenuation being based on the first object, the method according to item 1. (Item 4) Identifying the first object includes identifying a real object, the method according to item 1. (Item 5) Identifying the real object includes using a sensor to determine the position of the real object with respect to the user within the mixed reality environment, the method according to item 4. (Item 6) The sensor comprises a depth camera, the method according to item 5. (Item 7) Further comprising generating helper data corresponding to the real object, the method according to item 4. (Item 8) Further comprising generating a virtual object corresponding to the real object, the method according to item 4. (Item 9) Further comprising identifying a second virtual object, the first audio signal intersecting the second virtual object, and a fourth audio signal being determined based on the second virtual object, the method according to item 1. (Item 10) A system, A wearable head device, A display for presenting a composite reality environment to a user, the display comprising a transmissive eyepiece through which a real environment is visible, a display, and A first speaker configured to present an audio signal to the first ear of the user, A second speaker configured to present an audio signal to the second ear of the user Comprising a wearable head device, and One or more processors, Identifying the position of the first ear of the listener within the composite reality environment, Identifying the position of the second ear of the listener within the composite reality environment, Identifying a first virtual sound source within the composite reality environment, Identifying a first object within the composite reality environment, Determining a first audio signal within the composite reality environment, the first audio signal originating at the first virtual sound source and intersecting the position of the first ear of the listener, Determining a second audio signal within the composite reality environment, the second audio signal originating at the first virtual sound source, intersecting the first object, and intersecting the position of the second ear of the listener, Determining a third audio signal based on the second audio signal and the first object, Presenting the first audio signal to the first ear via the first speaker, Presenting the third audio signal to the second ear via the second speaker One or more processors configured to perform, and Comprising a system. (Item 11) Determining the third audio signal from the second audio signal includes applying a low-pass filter to the second audio signal, the low-pass filter having parameters based on the first object, the system of item 10. (Item 12) Determining the third audio signal from the second audio signal includes applying attenuation to the second audio signal, the intensity of the attenuation being based on the first object, the system of item 10. (Item 13) Identifying the first object includes identifying a physical object, the system of item 10. (Item 14) The wearable head device further includes a sensor, and identifying the physical object includes using the sensor to determine the position of the physical object relative to the user within the mixed reality environment, the system of item 13. (Item 15) The sensor includes a depth camera, the system of item 14. (Item 16) The one or more processors are further configured to generate helper data corresponding to the physical object, the system of item 13. (Item 17) The one or more processors are further configured to generate a virtual object corresponding to the physical object, the system of item 13. (Item 18) The one or more processors are further configured to identify a second virtual object, the first audio signal intersecting the second virtual object, and a fourth audio signal being determined based on the second virtual object, the system of item 10.

Brief Description of the Drawings

[0010]

Figure 1A

Figure 1B

Figure 1C

[0011]

Figure 2A

Figure 2B

Figure 2C

Figure 2D

[0012]

Figure 3A

[0013]

Figure 3B

[0014]

Figure 4

[0015]

Figure 5A

Figure 5B

[0016]

Figure 6

[0017]

Figure 7

[0018] In the following description of embodiments, reference is made to the accompanying drawings that form a part hereof and in which are shown, by way of illustration, specific embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the disclosed embodiments.

[0019] Mixed Reality Environment

[0020] As with all people, a user of a mixed reality system is present within a physical environment, i.e., a three-dimensional portion of the "real world" and all of its contents are perceivable by the user. For example, the user perceives the physical environment using normal human senses, i.e., sight, sound, touch, taste, smell, and interacts with the physical environment by moving his or her body within the physical environment. Locations within the physical environment can be described as coordinates within a coordinate space. For example, the coordinates can include latitude, longitude, and altitude above sea level, distances in three orthogonal dimensions from a reference point, or other suitable values. Similarly, a vector can describe a quantity having a direction and magnitude within a coordinate space.

[0021] A computing device can maintain a representation of a virtual environment, for example, within a memory associated with the device. As used herein, a virtual environment is a calculated representation of a three-dimensional space. The virtual environment can include a representation of any object, action, signal, parameter, coordinate, vector, or other characteristic associated with that space. In some embodiments, a circuit (e.g., a processor) of the computing device can maintain and update the state of the virtual environment. That is, the processor can determine the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or input provided by a user at a first time t0. For example, if an object within the virtual environment is located at a first coordinate at time t0, has certain programmed physical parameters (e.g., mass, coefficient of friction), and the input received from the user indicates that a force should be applied to the object in a certain direction vector, the processor can apply the laws of kinematics and use basic mechanics to determine the location of the object at time t1. The processor can use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at time t1. When maintaining and updating the state of the virtual environment, the processor can execute any suitable software, including software related to creating and deleting virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling inputs and outputs, software for implementing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.

[0022] Output devices such as displays or speakers can present to the user any or all aspects of the virtual environment. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, light, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, line of sight, and frustum), and render on the display a visible scene of the virtual environment corresponding to that view. Any suitable rendering technique may be used for this purpose. In some embodiments, the visible scene may include only some of the virtual objects within the virtual environment and may exclude other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects within the virtual environment may generate sounds arising from the location coordinates of the objects (e.g., a virtual character can speak or produce sound effects), or the virtual environment may be associated with music cues or ambient sounds, which may or may not be associated with a particular location. The processor can determine an audio signal corresponding to "listener" coordinates, e.g., an audio signal that is mixed and processed to simulate an audio signal that would be heard by the listener at the listener coordinates and corresponding to the synthesis of sound within the virtual environment, and present the audio signal to the user via one or more speakers.

[0023] Since the virtual environment only exists as a computational structure, a user cannot directly perceive the virtual environment using normal senses. Instead, the user can only indirectly perceive the virtual environment as presented to the user, for example, by a display, speakers, a tactile output device, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data to a processor via an input device or sensor to update the virtual environment using device or sensor data. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use that data to appropriately respond to the object within the virtual environment.

[0024] A mixed reality system can present a user with a mixed reality environment (“MRE”) that combines aspects of the real environment and the virtual environment, for example, using a see-through display and / or one or more speakers (e.g., that can be incorporated into a wearable head device). In some embodiments, the one or more speakers may be external to the wearable head device. As used herein, an MRE is a simultaneous representation of the real environment and the corresponding virtual environment. In some examples, the corresponding real and virtual environments share a single coordinate space. In some examples, the real coordinate space and the corresponding virtual coordinate space are related to each other by a transformation matrix (or other suitable representation). Thus, a single coordinate (along with a transformation matrix in some examples) can define a first location within the real environment and also a second corresponding location within the virtual environment, and vice versa.

[0025] In MRE, a virtual object (e.g., within a virtual environment associated with MRE) may correspond to a real object (e.g., within a real environment associated with MRE). For example, if the real environment of MRE includes a real street lamp post (real object) at certain location coordinates, the virtual environment of MRE may include a virtual street lamp post (virtual object) at the corresponding location coordinates. As used herein, a real object, in combination with its corresponding virtual object, together constitute a "composite reality object". It is not necessary for the virtual object to perfectly match or align with the corresponding real object. In some embodiments, the virtual object can be a simplified version of the corresponding real object. For example, if the real environment includes a real street lamp post, the corresponding virtual object may include a cylinder generally of the same height and radius as the real street lamp post (reflecting that the street lamp post may be of a generally cylindrical shape). Simplifying the virtual object in this way can enable computational efficiency and simplify the calculations to be performed on such virtual objects. Further, in some embodiments of MRE, not all real objects within the real environment may be associated with corresponding virtual objects. Similarly, in some embodiments of MRE, not all virtual objects within the virtual environment may be associated with corresponding real objects. That is, some virtual objects may exist only within the virtual environment of MRE without any real-world counterpart.

[0026] In some embodiments, the virtual object may have characteristics (sometimes significantly different and different from those of the corresponding real object). For example, the real environment within the MRE may include a cactus with two green branches extending, i.e., an inanimate object covered with thorns, while the corresponding virtual object within the MRE may have the characteristics of a virtual character with two green arms accompanied by human facial features and a surly attitude. In this embodiment, the virtual object is similar to its corresponding real object in some characteristics (color, number of arms), but different from the real object in other characteristics (facial features, personality). Thus, the virtual object has the potential to represent the real object or endow the real object, which would otherwise be inanimate, with behavior (e.g., human personality) in a creative, abstract, exaggerated, or fictional style. In some embodiments, the virtual object may be a purely fictional creation without a real-world counterpart (e.g., perhaps a virtual monster within the virtual environment at a location corresponding to a void within the real environment).

[0027] Compared to a VR system that presents a virtual environment to a user while obscuring the real environment, a mixed reality (MR) system that presents an MR environment has the advantage that the real environment remains perceivable while the virtual environment is presented. Thus, a user of an MR system can use visual and audio cues associated with the real environment to experience and interact with the corresponding virtual environment. As an example, while a user of a VR system may struggle to perceive or interact with virtual objects displayed within the virtual environment since, as described above, the user cannot directly perceive or interact with the virtual environment, a user of an MR system may find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching corresponding real objects within their own real environment. This level of interaction can enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by presenting the real and virtual environments simultaneously, an MR system can reduce negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. The MR system also presents many possibilities for applications that can extend or modify our experience of the real world.

[0028] Figure 1A illustrates an exemplary real - world environment 100 in which a user 110 uses a mixed - reality system 112. The mixed - reality system 112 may include a display (e.g., a transmissive display) and one or more speakers, and one or more sensors (e.g., a camera) as described below. The illustrated real - world environment 100 includes a rectangular room 104A in which the user 110 stands, and real objects 122A (a lamp), 124A (a table), 126A (a sofa), and 128A (a painting). The room 104A further includes location coordinates 106, which may be regarded as the origin of the real - world environment 100. As shown in Figure 1A, an environment / world coordinate system 108 (comprising an x - axis 108X, a y - axis 108Y, and a z - axis 108Z) associated with the origin at point 106 (world coordinates) may define a coordinate space for the real - world environment 100. In some embodiments, the origin 106 of the environment / world coordinate system 108 may correspond to the location where the power of the mixed - reality system 112 is turned on. In some embodiments, the origin 106 of the environment / world coordinate system 108 may be reset during operation. In some examples, the user 110 may be regarded as a real object within the real - world environment 100. Similarly, the body parts (e.g., hands, feet) of the user 110 may be regarded as real objects within the real - world environment 100. In some examples, a user / listener / head coordinate system 114 (comprising an x - axis 114X, a y - axis 114Y, and a z - axis 114Z) associated with the origin at point 115 (e.g., user / listener / head coordinates) may define a coordinate space for the user / listener / head on which the mixed - reality system 112 is located. The origin 115 of the user / listener / head coordinate system 114 may be defined with respect to one or more components of the mixed - reality system 112. For example, the origin 115 of the user / listener / head coordinate system 114 may be defined with respect to the display of the mixed - reality system 112 during initial calibration of the mixed - reality system 112 and the like. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrices) or other suitable representation can characterize the transformation between the user / listener / head coordinate system 114 space and the environment / world coordinate system 108 space.In some embodiments, the left ear coordinates 116 and the right ear coordinates 117 may be defined relative to the origin 115 of the user / listener / head coordinate system 114. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the left ear coordinates 116 and the right ear coordinates 117 and the user / listener / head coordinate system 114 space. The user / listener / head coordinate system 114 can simplify the representation of the location of the user's head or the wearable head device relative to, for example, the environment / world coordinate system 108. Using simultaneous localization and mapping (SLAM), visual odometry, or other techniques, the transformation between the user coordinate system 114 and the environment coordinate system 108 can be determined and updated in real time.

[0029] FIG. 1B illustrates an exemplary virtual environment 130 corresponding to the real environment 100. The illustrated virtual environment 130 includes a virtual rectangular room 104B corresponding to the real rectangular room 104A, a virtual object 122B corresponding to the real object 122A, a virtual object 124B corresponding to the real object 124A, and a virtual object 126B corresponding to the real object 126A. The metadata associated with the virtual objects 122B, 124B, 126B can include information derived from the corresponding real objects 122A, 124A, 126A. The virtual environment 130 additionally includes a virtual monster 132, which does not correspond to any real object within the real environment 100. The real object 128A within the real environment 100 does not correspond to any virtual object within the virtual environment 130. A persistent coordinate system 133 (comprising an x-axis 133X, a y-axis 133Y, and a z-axis 133Z) with its origin at point 134 (persistent coordinates) can define a coordinate space for the virtual content. The origin 134 of the persistent coordinate system 133 may be defined relative to / with respect to one or more real objects such as the real object 126A. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrices) or other suitable representation can characterize the transformation between the persistent coordinate system 133 space and the environment / world coordinate system 108 space. In some embodiments, the virtual objects 122B, 124B, 126B, and 132 may each have their own persistent coordinate points relative to the origin 134 of the persistent coordinate system 133. In some embodiments, there may be multiple persistent coordinate systems, and the virtual objects 122B, 124B, 126B, and 132 may each have their own persistent coordinate points relative to one or more of the persistent coordinate systems.

[0030] With respect to FIGS. 1A and 1B, the environment / world coordinate system 108 defines a shared coordinate space for both the real environment 100 and the virtual environment 130. In the illustrated embodiment, the coordinate space has its origin at point 106. Further, the coordinate space is defined by the same three orthogonal axes (108X, 108Y, 108Z). Thus, a first location within the real environment 100 and a second corresponding location within the virtual environment 130 can be described with respect to the same coordinate space. This simplifies the steps of identifying and displaying corresponding locations within the real and virtual environments since the same coordinates can be used to identify both locations. However, in some embodiments, the corresponding real and virtual environments need not use a shared coordinate space. For example, in some embodiments (not shown), a matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.

[0031] FIG. 1C illustrates an exemplary MRE 150 that simultaneously presents side views of the real environment 100 and the virtual environment 130 to the user 110 via the mixed reality system 112. In the illustrated embodiment, the MRE 150 simultaneously presents to the user 110 real objects 122A, 124A, 126A, and 128A from the real environment 100 (e.g., via the transmissive portion of the display of the mixed reality system 112) and virtual objects 122B, 124B, 126B, and 132 from the virtual environment 130 (e.g., via the active display portion of the display of the mixed reality system 112). As described above, the origin 106 acts as the origin for the coordinate space corresponding to the MRE 150, and the coordinate system 108 defines the x-axis, y-axis, and z-axis for the coordinate space.

[0032] In the illustrated embodiments, the composite reality object comprises a corresponding pair of real and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) that occupy corresponding locations within the coordinate space 108. In some embodiments, both the real object and the virtual object may be visible to the user 110 simultaneously. This may be desirable, for example, in instances where the virtual object presents information designed to augment the view of the corresponding real object (such as in museum applications where the virtual object presents missing parts of an ancient damaged statue). In some embodiments, the virtual objects (122B, 124B, and / or 126B) may be displayed so as to occlude the corresponding real objects (122A, 124A, and / or 126A) (e.g., via active pixelated occlusion using a pixelated occlusion shutter). This may be desirable, for example, in instances where the virtual object acts as a visual replacement for the corresponding real object (such as in two-way storytelling applications where an inanimate real object becomes a "living" character).

[0033] In some embodiments, the real objects (e.g., 122A, 124A, 126A) may be associated with virtual content or helper data that does not necessarily constitute the virtual object. The virtual content or helper data can facilitate the processing or handling of the virtual object within the composite reality environment. For example, such virtual content may include a two-dimensional representation of the corresponding real object, a custom asset type associated with the corresponding real object, or statistical data associated with the corresponding real object. This information can enable or facilitate calculations involving the real object without incurring unnecessary computational overhead.

[0034] In some embodiments, the presentation described above may also incorporate an audio aspect. For example, in the MRE 150, the virtual monster 132 may be associated with one or more audio signals, such as a footstep effect, that are generated as the monster walks around the MRE 150. As further described below, the processor of the mixed reality system 112 calculates an audio signal corresponding to the mixing and processed synthesis of all such sounds within the MRE 150 and can present the audio signal to the user 110 via one or more speakers included within the mixed reality system 112 and / or one or more external speakers.

[0035] Exemplary Mixed Reality System

[0036] The exemplary mixed reality system 112 can include a display (which can include left and right transmissive displays, which can be near-eye displays, and associated components for coupling light from the display to the user's eyes), left and right speakers (e.g., positioned adjacent to the user's left and right ears, respectively), an inertial measurement unit (IMU) (e.g., mounted on the arm of the head device's stem), an orthogonal coil electromagnetic receiver (e.g., mounted on the left stem component), left and right cameras (e.g., depth (time-of-flight) cameras) oriented away from the user, and left and right eye cameras (e.g., for detecting the user's eye movements) oriented towards the user, in a wearable head device (e.g., a wearable augmented reality or mixed reality head device). However, the mixed reality system 112 can incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic). Additionally, the mixed reality system 112 can incorporate networking features (e.g., Wi-Fi capabilities) and communicate with other devices and systems, including other mixed reality systems. The mixed reality system 112 can further include a battery (which may be mounted in an auxiliary unit such as a belt pack designed to be worn around the user's waist), a processor, and a memory. The wearable head device of the mixed reality system 112 can include a tracking component, such as an IMU or other suitable sensor, configured to output a coordinate set of the wearable head device with respect to the user's environment. In some embodiments, the tracking component may provide an input to the processor and perform simultaneous localization and mapping (SLAM) and / or visual odometry algorithms. In some embodiments, the mixed reality system 112 may also include a handheld controller 300 and / or an auxiliary unit 320, which can be a wearable belt pack, as further described below.

[0037] Figures 2A-2D illustrate components of an exemplary mixed reality system 200 (which may correspond to mixed reality system 112) that may be used to present an MRE (which may correspond to MRE150) or other virtual environment to a user. FIG. 2A illustrates a perspective view of a wearable head device 2102 included within the exemplary mixed reality system 200. FIG. 2B illustrates a top view of the wearable head device 2102 worn on a user's head 2202. FIG. 2C illustrates a front view of the wearable head device 2102. FIG. 2D illustrates an edge view of an exemplary eyepiece 2110 of the wearable head device 2102. As shown in FIGS. 2A-2C, the exemplary wearable head device 2102 includes an exemplary left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an exemplary right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 can include a transmissive element through which the real environment is visible and a display element for presenting a display (e.g., via light modulated for each image) that overlays the real environment. In some embodiments, such a display element can include a surface diffractive optical element for controlling the flow of light modulated for each image. For example, the left eyepiece 2108 can include a left internal coupling grating set 2112, a left orthogonal pupil expansion (OPE) grating set 2120, and a left exit (output) pupil expansion (EPE) grating set 2122. Similarly, the right eyepiece 2110 can include a right internal coupling grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. Light modulated for each image can be transferred to the user's eyes via the internal coupling gratings 2112 and 2118, the OPEs 2114 and 2120, and the EPEs 2116 and 2122. Each internal coupling grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 can be designed to gradually deflect light downward toward its associated EPE 2122, 2116, thereby horizontally extending the formed exit pupil.Each EPE 2122, 2116 can be configured to gradually and outwardly redirect at least a portion of the light received from its corresponding OPE grating set 2120, 2114 to a user eye box position (not shown) defined behind the eyepieces 2108, 2110, so as to vertically extend the exit pupil formed at the eye box. Alternatively, instead of the internal coupling grating sets 2112 and 2118, the OPE grating sets 2114 and 2120, and the EPE grating sets 2116 and 2122, the eyepieces 2108 and 2110 can include gratings and / or other arrays of refractive and reflective features for controlling the coupling of light modulated for each image to the user's eyes.

[0038] In some embodiments, the wearable head device 2102 can include a left arm 2130 of the strap and a right arm 2132 of the strap. The left arm 2130 of the strap includes a left speaker 2134, and the right arm 2132 of the strap includes a right speaker 2136. The orthogonal coil electromagnetic receiver 2138 can be located in the left temple component or another suitable location within the wearable head device 2102. The inertial measurement unit (IMU) 2140 can be located in the right arm 2132 of the strap or another suitable location within the wearable head device 2102. The wearable head device 2102 can also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142, 2144 can preferably be oriented in different directions so as to together cover a wider field of view.

[0039] In the embodiments shown in FIGS. 2A-2D, the left source of the light 2124 modulated for each image can be optically coupled into the left eyepiece lens 2108 through the left internal coupling grating set 2112, and the right source of the light 2126 modulated for each image can be optically coupled into the right eyepiece lens 2110 through the right internal coupling grating set 2118. The sources of the light 2124, 2126 modulated for each image include, for example, a projector including an electro-optical modulator such as an optical fiber scanner, a digital light processing (DLP) chip, or a liquid crystal on silicon (LCoS) modulator, or a light-emitting display such as a micro light-emitting diode (μLED) or a micro organic light-emitting diode (μOLED) panel that is coupled into the internal coupling grating sets 2112, 2118 using one or more lenses per side. The input coupling grating sets 2112, 2118 can deflect the light from the sources of the light 2124, 2126 modulated for each image at an angle exceeding the critical angle for total internal reflection (TIR) for the eyepiece lenses 2108, 2110. The OPE grating sets 2114, 2120 gradually deflect the propagating light downward toward the EPE grating sets 2116, 2122 by TIR. The EPE grating sets 2116, 2122 gradually couple the light toward the user's face, including the pupil of the user's eye.

[0040] In some embodiments, as shown in FIG. 2D, the left eyepiece lens 2108 and the right eyepiece lens 2110 each include a plurality of waveguides 2402. For example, each eyepiece lens 2108, 2110 can include a plurality of individual waveguides, each dedicated to an individual color channel (e.g., red, blue, and green). In some embodiments, each eyepiece lens 2108, 2110 can include a plurality of sets of such waveguides, each set configured to impart a different wavefront curvature to the emitted light. The wavefront curvature can be convex with respect to the user's eye, for example, to present a virtual object positioned at a certain distance in front of the user (e.g., a distance corresponding to the reciprocal of the wavefront curvature). In some embodiments, the EPE grating sets 2116, 2122 can include curved grating grooves to provide a convex wavefront curvature by modifying the Poynting vector of the light exiting across each EPE.

[0041] In some embodiments, to create the perception that the displayed content is three-dimensional, the stereoscopically adjusted left and right eye images can be presented to the user through the light modulators 2124, 2126 and the eyepiece lenses 2108, 2110 for each image. The perceived realism of the presentation of the three-dimensional virtual object can be improved by selecting the waveguides (and thus the corresponding wavefront curvature) such that the virtual object is displayed at a distance approximating the distance indicated by the stereoscopic left and right images. This technique can also reduce motion sickness experienced by some users, which can be caused by the difference between the depth perception cues provided by the stereoscopic left and right eye images and the automatic focusing of the human eye (e.g., object distance-dependent focusing).

[0042] FIG. 2D illustrates an edge view from above the right eyepiece lens 2110 of the exemplary wearable head device 2102. As shown in FIG. 2D, the plurality of waveguides 2402 can include a first subset of three waveguides 2404 and a second subset of three waveguides 2406. The two subsets of waveguides 2404, 2406 can be distinguished by different EPE gratings characterized by different grating line curvatures to impart different wavefront curvatures to the emitted light. Within each of the subsets of waveguides 2404, 2406, each waveguide can be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although not shown in FIG. 2D, the structure of the left eyepiece lens 2108 is similar to the structure of the right eyepiece lens 2110.)

[0043] Figure 3A illustrates an exemplary hand-held controller component 300 of the mixed reality system 200. In some embodiments, the hand-held controller 300 includes a gripping portion 346 and one or more buttons 350 disposed along an upper surface 348. In some embodiments, the buttons 350 may be configured to be used as optical tracking targets for tracking the six degrees of freedom (6DOF) movement of the hand-held controller 300, in conjunction with, for example, a camera or other optical sensor (which may be mounted within a head unit of the mixed reality system 200, such as the wearable head device 2102). In some embodiments, the hand-held controller 300 includes a tracking component (such as an IMU or other suitable sensor) for detecting a position or orientation, such as a position or orientation relative to the wearable head device 2102. In some embodiments, such a tracking component may be positioned within the handle of the hand-held controller 300 and / or may be mechanically coupled to the hand-held controller. The hand-held controller 300 can be configured to provide one or more output signals corresponding to one or more of a button pressed state, or the position, orientation, and / or movement of the hand-held controller 300 (e.g., via an IMU). Such output signals may be used as inputs to a processor of the mixed reality system 200. Such inputs may correspond to the position, orientation, and / or movement of the hand-held controller, and further, the position, orientation, and / or movement of the hand of the user holding the controller. Such inputs may also correspond to the user pressing the button 350.

[0044] FIG. 3B illustrates an exemplary auxiliary unit 320 of the mixed reality system 200. The auxiliary unit 320 can include a battery that provides energy to operate the system 200 and can include a processor that executes a program to operate the system 200. As shown, the exemplary auxiliary unit 320 includes a clip 2128 for attaching the auxiliary unit 320 to the user's belt or the like. It will also be apparent that other form factors are suitable for the auxiliary unit 320 and can include form factors that do not involve mounting the unit on the user's belt. In some embodiments, the auxiliary unit 320 is coupled to the wearable head device 2102 through a multi-tubular cable that can include, for example, electrical wires and optical fibers. A wireless connection between the auxiliary unit 320 and the wearable head device 2102 can also be used.

[0045] In some embodiments, the mixed reality system 200 includes one or more microphones that can detect sound and provide corresponding signals to the mixed reality system. In some embodiments, the microphone may be attached to or integrated with the wearable head device 2102 and may be configured to detect the user's voice. In some embodiments, the microphone may be attached to or integrated with the handheld controller 300 and / or the auxiliary unit 320. Such a microphone may be configured to detect ambient sound, background noise, the voice of the user or a third party, or other sounds.

[0046] FIG. 4 shows an exemplary functional block diagram corresponding to an exemplary mixed reality system such as the mixed reality system 200 described above (which may correspond to the mixed reality system 112 of FIG. 1). As shown in FIG. 4, an exemplary handheld controller 400B (which may correspond to the handheld controller 300 (“Totem”)) includes a totem / wearable head device 6 degrees of freedom (6DOF) totem subsystem 404A, and an exemplary wearable head device 400A (which may correspond to the wearable head device 2102) includes a totem / wearable head device 6DOF subsystem 404B. In an embodiment, the 6DOF totem subsystem 404A and the 6DOF subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom may be represented relative to the coordinate system of the wearable head device 400A. The three translational offsets may be represented as X, Y, and Z offsets, a translation matrix, or some other representation within such a coordinate system. The rotational degrees of freedom may be represented as a sequence of yaw, pitch, and roll rotations, a rotation matrix, a quaternion, or some other representation. In some embodiments, the wearable head device 400A, one or more depth cameras 444 (and / or one or more non-depth cameras) included within the wearable head device 400A, and / or one or more optical targets (e.g., a button 350 of the handheld controller 400B as described above or a dedicated optical target included within the handheld controller 400B) may be used for 6DOF tracking. In some embodiments, the handheld controller 400B may include a camera as described above, and the wearable head device 400A may include an optical target for optical tracking in conjunction with the camera. In some embodiments, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids, which are used to wirelessly transmit and receive three distinguishable signals.The 6DOF of the wearable head device 400A relative to the handheld controller 400B can be determined by measuring the relative magnitudes of three distinguishable signals received within each of the coils for reception. Additionally, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that is useful for providing improved accuracy and / or more timely information regarding fast movements of the handheld controller 400B.

[0047] In some embodiments, for example, to compensate for the movement of the wearable head device 400A relative to the coordinate system 108, it may be necessary to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment). For example, such a transformation is such that the display of the wearable head device 400A presents virtual objects at their expected positions and orientations relative to the real environment rather than at fixed positions and orientations on the display (e.g., the same position in the lower right corner of the display), and that the virtual objects appear to exist within the real environment (and, for example, do not appear unnaturally positioned within the real environment as the wearable head device 400A shifts and rotates). In some embodiments, the compensating transformation between coordinate spaces can be determined by processing images from the depth camera 444 using SLAM and / or visual odometry procedures to determine the transformation of the wearable head device 400A relative to the coordinate system 108. In the embodiment shown in FIG. 4, the depth camera 444 is coupled to the SLAM / visual odometry block 406 and can provide the images to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the images and then determine the position and orientation of the user's head, which can be used to identify the transformation between the head coordinate space and another coordinate space (e.g., the inertial coordinate space). Similarly, in some embodiments, an additional source of information regarding the user's head pose and location is obtained from the IMU 409. The information from the IMU 409 can be integrated with the information from the SLAM / visual odometry block 406 to provide more timely information regarding improved accuracy and / or fast adjustment of the user's head pose and position.

[0048] In some embodiments, the depth camera 444 can supply 3D images to a hand gesture tracker 411 that can be implemented within the processor of the wearable head device 400A. The hand gesture tracker 411 can identify a user's hand gesture, for example, by matching the 3D images received from the depth camera 444 to stored patterns representing hand gestures. Other suitable techniques for identifying a user's hand gesture will be apparent.

[0049] In some embodiments, one or more processors 416 may be configured to receive data from the 6DOF wearable head device subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, and / or hand gesture tracker 411 of the wearable head device. The processor 416 may also be able to send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A in embodiments where the handheld controller 400B is not tethered. The processor 416 may further communicate with additional components such as the audiovisual content memory 418, graphical processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to the head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to the left source of the light 424 modulated per image and a right channel output coupled to the right source of the light 426 modulated per image. The GPU 420 may output stereoscopic image data to the sources of the light 424, 426 modulated per image, as described above with respect to FIGS. 2A-2D, for example. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or right speaker 414. The DSP audio spatializer 422 may receive, from the processor 419, an input indicating a direction vector from the user to a virtual sound source (e.g., movable by the user via the handheld controller 320). Based on the direction vector, the DSP audio spatializer 422 may be able to determine the corresponding HRTF (e.g., by accessing the HRTF or interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the credibility and realism of virtual sounds by incorporating the user's relative position and orientation with respect to the virtual sounds in the composite reality environment, i.e., by presenting virtual sounds that match the user's expectations of what the virtual sounds would sound like if they were real sounds in the real environment.

[0050] In some embodiments, such as those shown in FIG. 4, one or more of the processor 416, GPU 420, DSP audio spatialization device 422, HRTF memory 425, and audio / visual content memory 418 may be included within an auxiliary unit 400C (which may correspond to the auxiliary unit 320 described above). The auxiliary unit 400C includes a battery 427, powers its components, and / or supplies power to the wearable head device 400A or the handheld controller 400B. Including such components within an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0051] FIG. 4 presents elements corresponding to various components of an exemplary composite reality system, but various other suitable arrangements of these components will be apparent to those skilled in the art. For example, the elements presented in FIG. 4 as being associated with the auxiliary unit 400C could instead be associated with the wearable head device 400A or the handheld controller 400B. Additionally, some composite reality systems may eliminate the handheld controller 400B or the auxiliary unit 400C entirely. Such changes and modifications should be understood to be within the scope of the disclosed embodiments.

[0052] Virtual sound source

[0053] As described above, MRE (such as experienced through a composite reality system, e.g., the composite reality system 200 described above) can present an audio signal to the user that can correspond to "listener" coordinates such that the audio signal represents what can be heard by the user at those listener coordinates. Some audio signals may correspond to the position and / or orientation of sound sources within the MRE. That is, the signals may appear to originate from the positions of sound sources within the MRE and be presented to propagate in the direction of the orientation of the sound sources within the MRE. In some cases, such audio signals may be considered virtual in that they correspond to virtual content within a virtual environment and not necessarily to real sounds within the real environment. Sounds associated with virtual content may be synthesized or generated by processing stored sound samples. The virtual audio signals can be presented to the user as real audio signals detectable by the human ear, for example, generated via speakers 2134 and 2136 of the wearable head device 2102 in FIGS. 2A-2D.

[0054] The sound source may correspond to a real object and / or a virtual object. For example, a virtual object (e.g., the virtual monster 132 in FIG. 1C) can emit an audio signal within the MRE, which is represented as a virtual audio signal within the MRE and presented to the user as a real audio signal. For example, the virtual monster 132 in FIG. 1C can emit a virtual sound or sound effect corresponding to the speech (e.g., dialogue) of the monster. Similarly, a real object (e.g., the real object 122A in FIG. 1C) can also emit a virtual sound within the MRE, which is represented as a virtual audio signal within the MRE and presented to the user as a real audio signal. For example, the real lamp 122A can emit a virtual sound corresponding to the sound effect of the lamp being switched on or off, even if the lamp cannot be switched on or off in the real environment. (The brightness of the lamp can be virtually generated using the eyepieces 2108, 2110 and the light sources 2124, 2126 modulated for each image.) The virtual sound can correspond to the position and orientation of the sound source (whether real or virtual). For example, when the virtual sound is presented to the user as a real audio signal (e.g., via speakers 2134 and 2136), the user can perceive the virtual sound as originating from the position of the sound source and traveling in the direction of the orientation of the sound source. (The sound source, although the sound source itself can correspond to a real object as described above, can be referred to as a "virtual sound source" in this specification.)

[0055] In some virtual or mixed reality environments, when a user is presented with an audio signal as described above, it is an intuitive and natural ability to identify the audio source in the real environment, but it can be difficult to quickly and accurately identify the source of the audio signal in the virtual environment. It is desirable to improve the user's ability to perceive the position or orientation of the sound source within the MRE so that the user's experience within the virtual or mixed reality environment closely resembles the user's experience in the real world.

[0056] Similarly, some virtual or augmented reality environments suffer from the perception that the environment does not feel real or authentic. One reason for this perception is that audio and visual cues do not always match each other within the virtual environment. For example, if a user is positioned behind a large brick wall within an MRE, the user might expect the sound originating from behind the brick wall to be quieter and more muffled than the sound originating directly next to the user. This expectation is based on the user's own auditory experience in the real world that sound can become quieter and muffled when blocked by a large, dense object. When the user is presented with an audio signal that is supposed to originate from behind the brick wall but is not muffled and is presented at full volume, the illusion that the user is behind the brick wall or that the sound is originating from behind it is broken. Since the entire virtual experience does not fully conform to the user's expectations based on real-world interactions, it can feel fake and not authentic. Additionally, in some cases, the "uncanny valley" problem can occur where even a slight difference between the virtual and physical experiences can cause a sense of discomfort. In an MRE, it is desirable to improve the user's experience by presenting audio signals that appear to interact realistically with the objects within the user's environment, even in small ways. The more consistent such audio signals are with the user's expectations based on real-world experiences, the more immersive and engaging the user's MRE experience will be.

[0057] One way the human brain detects the location and orientation of a sound source is by interpreting the differences between the sounds received by the left and right ears. For example, if an audio signal in a real environment reaches the user's left ear before it reaches the right ear (which the human auditory system can determine, for example, by identifying the time delay or phase shift between the left ear signal and the right ear signal), the brain may recognize that the source of the audio signal is to the left of the user. Similarly, since the effective strength of an audio signal generally decreases with distance and can be blocked by the user's own head, if the audio signal appears louder in the left ear than in the right ear, the brain may recognize that the source is to the left of the user. Similarly, our brains recognize that differences in the frequency characteristics between the left ear signal and the right ear signal can indicate the location of the source or the direction in which the audio signal is traveling.

[0058] The above techniques that the human brain unconsciously performs act by processing stereo audio signals, specifically, when applicable, by analyzing the differences (e.g., amplitude, phase, frequency characteristics) between the individual audio signals generated by a single sound source and received at the left and right ears. As humans, we necessarily rely on these stereo auditory techniques to quickly and accurately identify where the sounds in our real environment are originating and the directions in which they are traveling. We also rely on such stereo techniques to better understand the surrounding world, for example, whether a sound source is on the other side of a nearby wall and, if applicable, the thickness of that wall and the material from which it is made.

[0059] It may be desirable to place virtual sound sources within the MRE in a manner that is persuasive in a way that the user can quickly localize, utilizing the same natural stereo techniques that our brains use in the real world. Similarly, using these same techniques, it may be desirable to improve the sense of coexistence of such virtual sound sources with real and virtual content within the MRE by presenting stereo audio signals corresponding to those sound sources that behave, for example, like stereo audio signals in the real world. By presenting an audio experience that evokes the audio experiences of our daily lives to the user of the MRE, the MRE can improve the user's sense of immersion and connection when engaging with the MRE.

[0060] Figures 5A and 5B depict a perspective view and a top view, respectively, of an exemplary mixed reality environment 500 (which may correspond to the mixed reality environment 150 of FIG. 1C). In the MRE 500, the user 501 has a left ear 502 and a right ear 504. In the illustrated embodiment, the user 501 is wearing a wearable head device 510 (which may correspond to the wearable head device 2102) that includes a left speaker 512 and a right speaker 514 (which may correspond to speakers 2134 and 2136, respectively). The left speaker 512 is configured to present an audio signal to the left ear 502, and the right speaker 514 is configured to present an audio signal to the right ear 504.

[0061] The exemplary MRE 500 includes a virtual sound source 520, which may have a position and orientation within the coordinate system of the MRE 500. In some embodiments, the virtual sound source 520 may be a virtual object (e.g., the virtual object 122A in FIG. 1C) or may be associated with a real object (e.g., the real object 122B in FIG. 1C). Thus, the virtual sound source 520 may have any or all of the characteristics described above with respect to virtual objects.

[0062] In some embodiments, the virtual sound source 520 may be associated with one or more physical parameters such as size, shape, mass, or material. In some embodiments, the orientation of the virtual sound source 520 may correspond to one or more such physical parameters. For example, in an embodiment where the virtual sound source 520 corresponds to a speaker with a speaker cone, the orientation of the virtual sound source 520 may correspond to the axis of the speaker cone. In embodiments where the virtual sound source 520 is associated with a real object, the physical parameters associated with the virtual sound source 520 may be derived from one or more physical parameters of the real object. For example, if the real object is a speaker with a 12-inch speaker cone, the virtual sound source 520 may have physical parameters corresponding to the 12-inch speaker cone (e.g., the virtual object 122B may derive physical parameters or dimensions from the corresponding real object 122A of the MRE150).

[0063] In some embodiments, the virtual sound source 520 may be associated with one or more virtual parameters, which may affect the audio signal or other signals or properties associated with the virtual sound source. Virtual parameters can include spatial properties (e.g., position, orientation, shape, dimensions) within the coordinate space of the MRE, visual properties (e.g., color, transparency, reflectivity), physical properties (e.g., density, elasticity, tensile strength, temperature, smoothness, wetness, resonance, conductivity), or other suitable properties of the object. The mixed reality system can determine such parameters and thus generate virtual objects having those parameters. These virtual objects can be rendered to the user according to these parameters (e.g., by the wearable head device 510).

[0064] In one embodiment of the MRE500, the virtual audio signal 530 is emitted by the virtual sound source 520 at the position of the virtual sound source and propagates outward from the virtual sound source. In one instance, an anisotropic directivity pattern (e.g., exhibiting frequency-dependent anisotropy) can be associated with the virtual sound source, and the virtual audio signal emitted in a certain direction (e.g., the direction towards the user 501) can be determined based on the directivity pattern. The virtual audio signal is not directly perceptible by the user of the MRE, but can be converted into a real audio signal by one or more speakers (e.g., speakers 512 or 514), which generates a real audio signal that can be heard by the user. For example, the virtual audio signal can be converted into an analog signal via a digital-audio converter, for example, by a processor and / or memory associated with the MRE, and then amplified and used to drive the speakers to generate sound perceptible by the listener. It may be a calculated representation of digital audio data. Such a calculated representation can include, for example, coordinates within the MRE where the virtual audio signal is generated, a vector within the MRE along which the virtual audio signal propagates, directivity, the time at which the virtual audio signal is generated, the speed at which the virtual audio signal propagates, or other suitable characteristics.

[0065] The MRE may also include, respectively, a representation of one or more listener coordinates corresponding to a location (a "listener") in a coordinate system where a virtual audio signal can be perceived. In some embodiments, the MRE may also include a representation of one or more listener vectors representing the orientation of the listener (e.g., for use in determining an audio signal that may be affected by the direction the listener is facing). Within the MRE, the listener coordinates may correspond to the actual location of the user's ears, which can be determined using SLAM, visual odometry, and / or an IMU (e.g., the IMU 409 described above with respect to FIG. 4). In some embodiments, the MRE can include left and right listener coordinates corresponding, respectively, to the locations of the user's left and right ears within the MRE's coordinate system. By determining the vector of the virtual audio signal from the virtual sound source to the listener coordinates, an actual audio signal can be determined that corresponds to how a human listener with ears at those coordinates would perceive the virtual audio signal.

[0066] In some embodiments, the virtual audio signal comprises base audio data (e.g., a computer file representing an audio waveform) and one or more parameters that can be applied to the base audio data. Such parameters may correspond to attenuation of the base sound (e.g., volume drop), filtering of the base sound (e.g., low-pass filter), time delay of the base sound (e.g., phase shift), reverberation sound parameters for applying artificial reverberation and echo effects, voltage-controlled oscillator (VCO) parameters for applying time-based modulation effects, pitch modulation of the base sound (e.g., to simulate the Doppler effect), or other suitable parameters. In some embodiments, these parameters can be a function of the relationship of the listener coordinates of the virtual audio source. For example, the parameter can be defined as a decreasing function of the distance from the listener coordinates to the position of the virtual audio source for the attenuation of the actual audio signal. That is, as the distance from the listener to the virtual audio source increases, the gain of the audio signal decreases. As another example, the parameter can be defined as a function of the distance from the listener coordinates (and / or the angle of the listener vector) to the propagation vector of the virtual audio signal for the low-pass filter applied to the virtual audio signal. For example, a listener far from the virtual audio signal may perceive less frequency power of the signal than a listener closer to the signal. As a further example, the parameter can be defined such that the time delay (e.g., phase shift) is applied based on the distance between the listener coordinates and the origin of the virtual audio signal. In some embodiments, the processing of the virtual audio signal can be calculated using the DSP audio spatialization device 422 of FIG. 4, which can present the audio signal based on the position and orientation of the user's head using HRTF.

[0067] Virtual audio signal parameters can be affected by virtual or real objects, i.e., sound occluders through which the virtual audio signal passes on its way to the listener coordinates. (As used herein, virtual or real objects include any suitable representation of virtual or real objects within the MRE.) For example, if a virtual audio signal intersects (e.g., is blocked by) a virtual wall within the MRE, the MRE can apply attenuation to the virtual audio signal (resulting in a signal that appears quieter to the listener). The MRE can also apply a low-pass filter to the virtual audio signal as the high-frequency components roll off, resulting in a signal that appears more muffled. These effects are consistent with our expectations when hearing sound from behind a wall, where the nature of the wall in the real environment causes the sound from the other side of the wall to be quieter and have fewer high-frequency components because the wall blocks the sound waves originating on the opposite side of the wall from the listener. The application of such parameters to the audio signal can be based on the nature of the virtual wall. For example, a virtual wall corresponding to a thicker or denser material can result in a greater degree of attenuation or low-pass filtering than a virtual wall corresponding to a thinner or less dense material. In some cases, the virtual object may apply a phase shift or additional effects to the virtual audio signal. The effect that a virtual object has on a virtual audio signal can be determined by the physical modeling of the virtual object. For example, if the virtual object corresponds to a particular material (e.g., brick, aluminum, water), the effect can be applied based on the known transmission characteristics of the audio signal in the presence of that material in the real world.

[0068] In some embodiments, the virtual objects at which the virtual audio signals intersect may correspond to real objects (e.g., real objects 122A, 124A, and 126A correspond to virtual objects 122B, 124B, and 126B in FIG. 1C). In some embodiments, such virtual objects may not correspond to real objects (e.g., virtual monster 132 in FIG. 1C, etc.). When the virtual objects correspond to real objects, the virtual objects may adopt parameters (e.g., dimensions, materials) corresponding to the properties of those real objects.

[0069] In some embodiments, the virtual audio signals may intersect real objects that do not have corresponding virtual objects. For example, the characteristics of the real objects (e.g., position, orientation, dimensions, materials) can be determined by sensors (such as being attached to the wearable head device 510), and those characteristics can be used to process virtual audio signals as described above with respect to the virtual object occluder.

[0070] Stereo effect

[0071] As described above, by determining the vector of the virtual audio signal from the virtual sound source at the listener coordinates, a real audio signal can be determined that corresponds to how a human listener with ears at those listener coordinates would perceive the virtual audio signal. In some embodiments, the left and right stereo listener coordinates (corresponding to the left and right ears) are simply used instead of a single listener coordinate and the effect of the real object on the audio signal, which would be determined separately for each ear, can enable attenuation or filtering based on the interaction of the audio signal and the real object. This can improve the realism of the virtual environment by mimicking a real-world stereo audio experience, and receiving different audio signals at each ear can help in understanding the sounds in our surroundings. Such an effect, where the audio signal is experienced as being affected differently by the left and right ears, can be particularly prominent when the real object can approach the user closely. For example, if user 501 is peeking at a meowing virtual cat from around the corner of a real object, the sound of the cat's meow can be determined and presented differently for each ear. That is, the sound for the ear positioned behind the real object can reflect that the real object between the cat and the ear can attenuate and filter the sound of the cat as it is heard by that ear, while the sound for the other ear positioned beyond the real object can reflect that the real object does not perform such attenuation or filtering. Such sounds can be presented via the ears 512, 514 of the user of the wearable head device 510.

[0072] The desirable stereo-auditory effects as described above can be simulated by determining two such vectors, one for each ear, and identifying a unique virtual audio signal for each ear. These two unique virtual audio signals can then each be converted into real audio signals and presented to the individual ears via the speakers associated with that ear. The user's brain will process those real audio signals in the same way it would process normal stereo audio signals in the real world, as described above.

[0073] This is illustrated by the exemplary MRE500 in FIGS. 5A and 5B. The MRE500 includes a wall 540 that is between the virtual sound source 520 and the user 501. In some embodiments, the wall 540 may be a real object that is no different from the real object 126A of FIG. 1C. In some embodiments, the wall 540 may be a virtual object such as the virtual object 122B of FIG. 1C. Further, in some such embodiments, the virtual object may correspond to a real object such as the real object 122A of FIG. 1C.

[0074] In an embodiment where the wall 540 is a real object, the wall 540 may be detected, for example, using a depth camera or other sensors of the wearable head device 510. This can identify one or more characteristics of the real object, such as its position, orientation, visual properties, or material properties. These characteristics can be associated with the wall 540 and included when updating and maintaining the MRE 500, as described above. These characteristics can then be used to process the virtual audio signal according to how their virtual audio signals will be affected by the wall 540, as described below. In some embodiments, virtual content, such as helper data, may be associated with the real object to facilitate the processing of virtual audio signals affected by the real object. For example, the helper data may include geometric primitives similar to the real object, two-dimensional image data associated with the real object, or custom asset types that identify one or more properties associated with the real object.

[0075] In some embodiments where the wall 540 is a virtual object, the virtual object may be calculated to correspond to a physical object that can be detected as described above. For example, with respect to FIG. 1C, as described above, the physical object 122A may be detected by the wearable head device 510, and the virtual object 122B may be generated to correspond to one or more characteristics of the physical object 122A. Additionally, one or more characteristics may be associated with the virtual object that are not derived from its corresponding physical object. The advantage of identifying a virtual object associated with a corresponding physical object is that the virtual object can be used to simplify calculations associated with the wall 540. For example, the virtual object may be geometrically simpler than its corresponding physical object. However, in some embodiments where the wall 540 is a virtual object, there may be no corresponding physical object, and the wall 540 may be determined by software (e.g., a software script that defines the presence of the wall 540 at a particular location and orientation). The characteristics associated with the wall 540 can be included when maintaining and updating the MRE 500 as described above. These characteristics can then be used to process the virtual audio signal according to how their virtual audio signals will be affected by the wall 540, as described below.

[0076] The wall 540 can be considered an acoustic occluder, whether physical or virtual, as described above. As seen in the top view shown in FIG. 5B, two vectors 532 and 534 can represent the individual paths of the virtual audio signals 530 from the virtual sound source 520 within the MRE 500 to the user's left ear 502 and right ear 504. The vectors 532 and 534 can each correspond to unique left and right audio signals to be presented to the left and right ears, respectively. As shown in the embodiment, the vector 534 (corresponding to the right ear 504) intersects the wall 540, while the vector 532 (corresponding to the left ear 502) cannot. Thus, the wall 540 can impart different characteristics to the right audio signal than to the left audio signal. For example, the right audio signal can be made to undergo attenuation and low-pass filtering to correspond to the wall 540, while the left audio signal is not. In some embodiments, the left audio signal can be phase-shifted or time-shifted relative to the right audio signal to correspond to the greater distance from the right ear 504 to the virtual sound source 520 than from the left ear 502 to the virtual sound source 520 (which would result in the audio signal from that source arriving at the left ear 502 slightly later than at the right ear 504). The user's auditory system can, as in the real world, interpret the phase or time shift and use it to identify that the virtual sound source 520 is on one side (e.g., the right side) of the user within the MRE 500.

[0077] The relative importance of these stereo differences can depend on the differences in the frequency spectrum of the signals. For example, phase shift can be more useful for localizing high-frequency signals than for localizing low-frequency audio signals (i.e., signals with wavelengths on the order of the width of the listener's head). For such low-frequency signals, the difference in arrival times between the left and right ears can be useful for localizing the sources of these signals.

[0078] In some embodiments not shown in FIGS. 5A-5B, an object such as wall 540 (whether physical or virtual) need not be between user 501 and virtual sound source 520. In such embodiments, such as when wall 540 is behind the user, the wall can impart different characteristics to the left and right audio signals via reflection of virtual audio signal 530 towards left and right ears 502 and 504 with respect to wall 540.

[0079] An advantage of MRE500 over some environments such as video games presented by conventional display monitors and room speakers is that the actual location of the user's ears within MRE500 can be determined. As described above with respect to FIG. 4, wearable head device 510 can be configured to identify the location of user 501, for example, through the use of sensors and measurement hardware such as SLAM, visual odometry techniques, and / or IMU. In some embodiments, wearable head device 510 may be configured to directly detect the individual locations of the user's ears (e.g., via sensors associated with ears 502 and 504, speakers 512 and 514, or vine arms such as vine arms 2130 and 2132 shown in FIGS. 2A-2D). In some embodiments, wearable head device 510 may be configured to detect the position of the user's head and approximate the individual locations of the user's ears based on that position (e.g., by estimating or detecting the width of the user's head and identifying ear locations that are located along the circumference of the head and separated by the width of the head). By identifying the location of the user's ears, audio signals can be presented to the ears corresponding to those specific locations. Determining the location of the ears and presenting audio signals based on that location, as compared to techniques that determine audio signals based on audio receiver coordinates (e.g., origin coordinates of a virtual camera within a virtual 3D environment), which may or may not correspond to the user's actual ears, can improve the user's sense of immersion and connection within the MRE.

[0080] By presenting unique and distinct left and right audio signals via speakers 512 and 514, corresponding to left and right listener positions (e.g., the locations of the user's ears 502 and 504 within MRE 500), respectively, user 501 is able to identify the location and / or orientation of virtual sound source 520. This is because the user's auditory system naturally attributes differences (e.g., in gain, frequency, and phase) between the left and right audio signals to the location and orientation of virtual sound source 520, along with the presence of sound occluders such as wall 540. Thus, these stereo audio cues improve user 501's perception of virtual sound source 520 and wall 540 within MRE 500. This, in turn, can enhance the user's sense of engagement with MRE 500. For example, if virtual sound source 520 corresponds to an important object within MRE 500, such as a virtual character speaking to user 501, user 501 can quickly identify the location of that object using the stereo audio signals. This can, in turn, reduce the cognitive burden on user 501 for identifying the location of the object, and also reduce the computational burden on MRE 501. For example, the processor and / or memory (e.g., processor 416 and / or memory 418 of FIG. 4) may no longer need to present high-fidelity visual cues (e.g., via high-resolution assets such as 3D models and textures and lighting effects) to user 501 for identifying the location of the object, as the audio cues are taking on much of that work.

[0081] The asymmetric occlusion effect as described above can be particularly prominent in situations where a real or virtual object such as wall 540 is physically close to the user's face, or where a real or virtual object occludes one ear but not the other (such as when the center of the user's face is aligned with the edge of wall 540 as seen in FIG. 5B). These situations can be utilized to be effective. For example, in MRE500, user 501 can hide behind the edge of wall 540 and peek out from the corner to locate a virtual object (such as corresponding to virtual sound source 520) based on the stereo audio effect imparted to the sound emission of that object (such as virtual audio signal 530) by the wall. This can enable, for example, tactical gameplay within a game environment based on MRE500, where user 501 checks the appropriate acoustics in different regions of a virtual room for architectural design purposes, or for educational or creative benefits where user 501 explores the interaction of various audio sources (such as the chirping of virtual birds) with their environment.

[0082] In some embodiments, the left and right audio signals may not be determined independently of each other, but may be based on other or common audio sources. For example, if a single audio source generates both the left and right audio signals, the left and right audio signals may not be entirely independent, but may be considered acoustically related to each other through the single audio source.

[0083] FIG. 6 shows an exemplary process 600 for presenting left and right audio signals to a user of an MRE such as user 501 of MRE500. The exemplary process 600 may be implemented by a processor of the wearable head device 510 (such as corresponding to processor 416 of FIG. 4) and / or a DSP module (such as corresponding to DSP audio spatializer 422 of FIG. 4).

[0084] At stage 605 of process 600, the individual locations (e.g., listener coordinates and / or vectors) of the first ear (e.g., the user's left ear 502) and the second ear (e.g., the user's right ear 504) are determined. These locations can be determined using the sensors of the wearable head device 510 as described above. Such coordinates can be with respect to the local user coordinate system of the wearable head device (e.g., the user coordinate system 114 described above with respect to FIG. 1A). In such a user coordinate system, the origin of such a coordinate system approximately corresponds to the center of the user's head and can simplify the representation of the locations of the left and right virtual listeners. Using SLAM, visual odometry, and / or IMU, the displacement and rotation (e.g., 6 degrees of freedom) of the user coordinate system 114 with respect to the environmental coordinate system 108 can be updated in real time.

[0085] At stage 610, a first virtual sound source can be defined that may correspond to the virtual sound source 520. In some embodiments, the virtual sound source may correspond to a virtual or real object, which may be identified and located via the depth camera or sensors of the wearable head device 510. In some embodiments, the virtual object may correspond to a real object as described above. For example, the virtual object may have one or more characteristics (e.g., position, orientation, material, visual properties, acoustic properties) of the corresponding real object. The location of the virtual sound source can be established within the coordinate system 108 (FIGS. 1A-1C).

[0086] In stage 620A, a first virtual audio signal that propagates along vector 532 and intersects a first virtual listener (e.g., a first approximate ear position) and that may correspond to virtual audio signal 530 can be identified. For example, in response to a determination that an audio signal was generated by a first virtual sound source at a first time t, a vector from the first sound source to the first virtual listener can be calculated. The first virtual audio signal can be associated with base audio data (e.g., a waveform file) and optionally one or more parameters for modifying the base audio data as described above. Similarly, in stage 620B, a second virtual audio signal that propagates along vector 534 and intersects a second virtual listener (e.g., a second approximate ear position) and that may correspond to virtual audio signal 530 can be identified.

[0087] In stage 630A, a real or virtual object intersected by the first virtual audio signal, one of which may correspond to, for example, wall 540, is identified. For example, a trace can be calculated along the vector from the first sound source in MRE 500 to the first virtual listener, and a real or virtual object intersecting the trace can be identified (in some embodiments, together with parameters of the intersection point such as the position and vector at which the real or virtual object is intersected). In some cases, such a real or virtual object may not exist. Similarly, in stage 630B, a real or virtual object intersected by the second virtual audio signal is identified. Again, in some cases, such a real or virtual object may not exist.

[0088] In some embodiments, the physical objects identified at stage 630A or stage 630B can be identified using a depth camera or other sensors associated with the wearable head device 510. In some embodiments, the virtual objects identified at stage 630A or stage 630B may correspond to the physical objects and physical objects 122A, 124A, and 126A, and corresponding virtual objects 122B, 124B, and 126B, as described with respect to FIG. 1C. In such embodiments, such physical objects can be identified using a depth camera or other sensors associated with the wearable head device 510, and the virtual objects can be generated to correspond to those physical objects as described above.

[0089] In step 640A, each real or virtual object identified in step 630A is processed to identify any signal modification parameters associated with that real or virtual object in step 650A. For example, as described above, such signal modification parameters may include attenuation, filtering, phase shift, time-based effects (e.g., delay, reverberation, modulation), and / or functions for determining other effects to be applied to the first virtual audio signal. As described above, these parameters may depend on other parameters associated with the real or virtual object, such as the size, shape, or material of the real or virtual object. In step 660A, those signal modification parameters are applied to the first virtual audio signal. For example, if the signal modification parameter is specified to be attenuated by a factor that linearly increases with the distance between the listener coordinates and the audio source of the first virtual audio signal, that factor is calculated (i.e., by calculating the distance between the first ear and the first virtual sound source within the MRE500) in step 660A and can be applied to the first virtual audio signal (i.e., by multiplying the amplitude of the signal by the resulting gain factor). In some embodiments, the signal modification parameter can be determined or applied using the DSP audio spatialization device 422 of FIG. 4, which can modify the audio signal based on the position and orientation of the user's head as described above using HRTF. When all real or virtual objects identified in step 630A are applied in step 660A, the processed first virtual audio signal (e.g., representing all signal modification parameters of the identified real or virtual objects) is output by step 640A. Similarly, in step 640B, each real or virtual object identified in step 630B is processed, signal modification parameters are identified (step 650B), and those signal modification parameters are applied to the second virtual audio signal (step 660B).When all the real or virtual objects identified in stage 630B are applied in stage 660B, the processed first virtual audio signal (e.g., representing all the signal modification parameters of the identified real or virtual objects) is output by stage 640B.

[0090] In stage 670A, the processed first virtual audio signal output from stage 640A can be used to determine a first audio signal (e.g., a left-channel audio signal) that can be presented to the first ear. For example, in stage 670A, the first virtual audio signal can be mixed with other left-channel audio signals (e.g., other virtual audio signals, music, or dialogue). In some embodiments, such as in a simple virtual reality environment without other sounds, stage 670A may perform little or no processing to determine the first audio signal from the processed first virtual audio signal. Stage 670A can incorporate any suitable stereo mixing technique. Similarly, in stage 680A, the processed second virtual audio signal output from stage 640B can be used to determine a second audio signal (e.g., a right-channel audio signal) that can be presented to the second ear.

[0091] In stages 680A and 680B, the audio signals output by stages 670A and 670B are presented to the first ear and the second ear, respectively. For example, the left and right stereo signals can be converted into left and right analog signals (e.g., by the DSP audio spatializer 422 of FIG. 4) that are amplified and presented to the left and right speakers 512 and 514, respectively. When the left and right speakers 512 and 514 are configured to acoustically couple to the left and right ears 502 and 504, respectively, the left and right ears 502 and 504 may be presented with their respective left and right stereo signals in sufficient isolation from other stereo signals that produce a stereo effect.

[0092] FIG. 7 shows a functional block diagram of an exemplary augmented reality processing system 700 that can be used to implement one or more of the embodiments described above. The exemplary system 700 can be implemented within a mixed reality system such as the mixed reality system 112 described above. FIG. 7 shows an aspect of the audio architecture of system 700. In the embodiment shown, a game engine 702 generates virtual 3D content 704 and simulates events with the virtual 3D content 704 (the events can include interactions between the virtual 3D content 704 and real objects). The virtual 3D content 704 can include, for example, static virtual objects, virtual objects with functionality, such as virtual musical instruments, virtual animals, and virtual people. In the embodiment shown, the virtual 3D content 704 includes localized virtual sound sources 706. The localized virtual sound sources 706 can include sound sources corresponding to, for example, the songs of virtual birds, sounds emitted by virtual musical instruments played by a user or virtual person, or the voices of virtual people.

[0093] The exemplary augmented reality processing system 700 can integrate the virtual 3D content 704 into the real world with a high degree of realism. For example, the audio associated with the localized virtual sound sources can be located at a distance from the user and in a location that would be partially blocked by real objects if the audio were a real audio signal. However, in the exemplary system 700, the audio can be output by left and right speakers 412, 414, 2134, 2136 (which can belong to the wearable head device 400A of the mixed reality system 112, for example). The audio that travels only a short distance from the speakers 2134, 2136 into the user's ears is not physically affected by obstacles. However, the system 700 can modify the audio and account for the effects of obstacles, as described below.

[0094] In the exemplary system 700, the user coordinate determination subsystem 708 can preferably be physically stored within the wearable head devices 200, 400A. The user coordinate determination subsystem 708 can maintain information about the position (e.g., X, Y, and Z coordinates) and orientation (e.g., roll, pitch, yaw; quaternions) of the wearable head device relative to the real-world environment. Virtual content is defined within the environmental coordinate system 108 (FIGS. 1A - 1C), which is generally fixed relative to the real world. However, in embodiments, the same virtual content is typically fixed to the wearable head devices 200, 400A and is output via the eyepieces 408, 410 and speakers 412, 414, 2134, 2136 that move relative to the real world as the user's head moves. As the wearable head devices 200, 400A are displaced or rotated, the spatialization of the virtual audio may be adjusted, and the visual display of the virtual content should be re-rendered to account for the displacement and / or rotation. The user coordinate determination subsystem 708 can include an inertial measurement unit (IMU) 710, which can include a set of three orthogonal accelerometers (not shown in FIG. 7) that provide measurements of acceleration from which displacement can be determined by integration, and three orthogonal gyroscopes (not shown in FIG. 7) that provide measurements of rotation from which orientation can be determined by integration. To adjust for drift errors in displacement and orientation obtained from the IMU 710, a simultaneous localization and mapping (SLAM) and / or visual odometry block 406 can be included within the user coordinate determination system 708. As shown in FIG. 4, a depth camera 444 can be coupled to the SLAM and / or visual odometry block 406 to provide image input therefor.

[0095] The sensor subsystem 712 for spatially prominent actual occluding objects (the "occlusion subsystem") is included within the exemplary augmented reality processing system 700. The occlusion subsystem 712 can include, for example, a depth camera 444, a non-depth camera (not shown in FIG. 7), an acoustic navigation and ranging (sonar) sensor (not shown in FIG. 7), and / or a light detection and ranging (LIDAR) sensor (not shown in FIG. 7). The occlusion subsystem 712 can have sufficient spatial resolution to identify obstacles that affect the virtual propagation paths corresponding to the left and right listener positions. For example, if a user of the wearable head devices 200, 400A is secretly viewing a virtual sound-emitting virtual object (e.g., an enemy in a virtual game where a wall forming an angle blocks the direct line of sight to the user's left ear but not the right ear) from behind a real corner, the occlusion subsystem 712 can sense the obstacle with sufficient resolution to determine that only the direct path to the left ear will be occluded. In some embodiments, the occlusion subsystem 712 may have better spatial resolution and may be able to determine the size (or the solid angle thereto) of the occluding real object and the distance thereto.

[0096] In the embodiment shown in FIG. 7, occlusion subsystem 712 is coupled to an intersection and obstacle range calculator (hereinafter, “obstacle calculator”) 714 for each channel (i.e., left and right audio channels). In an embodiment, user coordinate determination system 708 and game engine 702 are also coupled to obstacle calculator 714. Obstacle calculator 714 can receive the coordinates of the virtual audio sources from game engine 702, the user coordinates from user coordinate determination system 708, and information indicating the coordinates of obstacles (e.g., optionally, angular coordinates including distance) from occlusion subsystem 712. By applying geometries, obstacle calculator 714 can determine whether there are blocked or unblocked lines of sight from each virtual audio source to each of the left and right listener positions. In FIG. 7, although shown as a separate block, obstacle calculator 714 can be integrated with game engine 702. In some embodiments, occlusion may first be sensed by occlusion subsystem 712 based on information from user coordinate determination system 708 within a user-centered coordinate system, and the occlusion coordinates are transformed to environment coordinate system 108 for the purpose of analyzing obstacle geometries. In some embodiments, the coordinates of the virtual sound sources may be transformed to a user-centered coordinate system for the purpose of calculating obstacle geometries. In some embodiments that provide spatially decomposed information about the object being occluded by occlusion subsystem 712, obstacle calculator 714 can determine the range of the solid angle centered on the line of sight blocked by the occluding object. Obstacles having a larger solid angle range can be considered by applying a larger attenuation and / or attenuation in a larger range of high frequency components.

[0097] In some embodiments, the located virtual sound source 706 can include a mono audio signal or left and right spatialized audio signals. Such left and right spatialized audio signals can be determined by applying left and right head-related transfer functions (HRTFs) that can be selected based on the coordinates of the located virtual sound source relative to the user. In embodiment 700, the game engine 702 is coupled to the user coordinate determination system 708 and receives the coordinates of the user (e.g., position and orientation). The game engine 702 itself can determine the coordinates of the virtual sound source (e.g., in response to user input) and, in response to receiving the user coordinates, can determine the coordinates of the sound source relative to the user based on the geometry.

[0098] In the embodiment shown in FIG. 7, the obstacle computer 714 is coupled to the filter activation and control device 716. In some embodiments, the filter activation and control device 716 is coupled to the left control input 718A of the left filter bypass switch 718 and the right control input 720A of the right filter bypass switch 720. In some embodiments, like other components of the exemplary system 700, the bypass switches 718, 720 can be implemented in software. In the embodiment shown, the left filter bypass switch 718 receives the left channel of the spatialized audio from the game engine 702, and the right filter bypass switch 720 receives the right spatialized audio from the game engine 704. In some embodiments where the game engine 702 outputs a mono audio signal, both bypass switches 718, 720 can receive the same mono audio signal.

[0099] In the embodiment shown in FIG. 7, the first output 718B of the left bypass switch 718 is coupled through the left obstacle filter 722 to the left digital / analog converter (“left D / A”) 724, and the second output 718C of the left bypass switch 718 is coupled to the left D / A 724 (bypassing the left obstacle filter 722). Similarly, in the embodiment, the first output 720B of the right bypass switch 720 is coupled through the right obstacle filter 726 to the right digital / analog converter (“right D / A”) 728, and the second output 720C is coupled to the right D / A 728 (bypassing the right obstacle filter 726).

[0100] In the embodiment shown in FIG. 7, a set of filter configurations 730 can be used to configure the left obstacle filter 722 and / or the right obstacle filter based on the output of the per-channel intersection and obstacle range calculator 722 (e.g., by the filter activation and control device 716). In some embodiments, instead of providing the bypass switches 718, 720, a non-filtering pass-through configuration of the obstacle filters 722, 726 can be used. The obstacle filters 722, 726 can be time-domain or frequency-domain filters. In embodiments where the filter is a time-domain filter, each filter configuration can include a set of tap coefficients. In embodiments where the filter is a frequency-domain filter, each filter configuration can include a set of frequency band weightings. In some embodiments, instead of a set of a predetermined number of filter configurations, the filter activation and control device 716 can be configured to define a filter having a certain level of attenuation depending on the size of the obstacle (e.g., programmatically). The filter activation and control device 716 can select or define a filter configuration (e.g., a configuration that attenuates more for larger obstacles), and / or can select or define a filter that attenuates higher frequency bands (e.g., to a greater extent for larger obstacles to simulate the effect of real obstacles).

[0101] In the embodiment shown in FIG. 7, the filter activation and control device 716 is coupled to the control input 722A of the left obstacle filter 722 and the control input 726A of the right obstacle filter 726. The filter activation and control device 716 can separately configure the left obstacle filter 722 and the right obstacle filter 726 using a configuration selected from the filter configurations 730 based on the output from the per-channel intersection and obstacle range calculator 714.

[0102] In the embodiment shown in FIG. 7, the left D / A 724 is coupled to the input 732A of the left audio amplifier 732, and the right D / A 728 is coupled to the input 734A of the right audio amplifier 734. In the embodiment, the output 732B of the left audio amplifier 732 is coupled to the left speakers 2134, 412, and the output 734B of the right audio amplifier 734 is coupled to the right speakers 2136, 414.

[0103] Note that the elements of the exemplary functional block diagram shown in FIG. 7 can be arranged in any suitable order, not necessarily in the order shown. Further, some of the elements shown in the embodiment in FIG. 7 (e.g., bypass switches 718, 720) can be omitted as needed. The present disclosure is not limited to any particular order or arrangement of the functional components shown in the embodiments.

[0104] Some embodiments of the present disclosure are methods for presenting audio signals in an augmented reality environment, including identifying the position of a listener's first ear in the augmented reality environment; identifying the position of the listener's second ear in the augmented reality environment; identifying a first virtual sound source in the augmented reality environment; identifying a first object in the augmented reality environment; determining a first audio signal in the augmented reality environment, where the first audio signal originates at the first virtual sound source and intersects the position of the listener's first ear; determining a second audio signal in the augmented reality environment, where the second audio signal originates at the first virtual sound source, intersects the first object, and intersects the position of the listener's second ear; determining a third audio signal based on the second audio signal and the first object; presenting the first audio signal to the user's first ear via a first speaker; and presenting the third audio signal to the user's second ear via a second speaker. In addition to or as an alternative to one or more of the embodiments disclosed above, in some embodiments, determining the third audio signal from the second audio signal includes applying a low-pass filter to the second audio signal, where the low-pass filter has parameters based on the first virtual object. In addition to or as an alternative to one or more of the embodiments disclosed above, in some embodiments, determining the third audio signal from the second audio signal includes applying attenuation to the second audio signal, where the intensity of the attenuation is based on the first object. In addition to or as an alternative to one or more of the embodiments disclosed above, in some embodiments, identifying the first object includes identifying a real object. In addition to or as an alternative to one or more of the embodiments disclosed above, in some embodiments, identifying a real object includes using sensors to determine the position of the real object relative to the user in the augmented reality environment.In addition to, or as an alternative to, one or more of the embodiments disclosed above, in some embodiments, the sensor comprises a depth camera. In addition to, or as an alternative to, one or more of the embodiments disclosed above, in some embodiments, the method further comprises generating helper data corresponding to the real object. In addition to, or as an alternative to, one or more of the embodiments disclosed above, in some embodiments, the method further comprises generating a virtual object corresponding to the real object. In addition to, or as an alternative to, one or more of the embodiments disclosed above, in some embodiments, the method further comprises identifying a second virtual object, wherein the first audio signal intersects the second virtual object and a fourth audio signal is determined based on the second virtual object.

[0105] Some embodiments of the present disclosure are systems that include a wearable head device. The wearable head device includes a display for presenting a mixed reality environment to a user, the display including a transmissive eyepiece through which a real environment is visible; a first speaker configured to present an audio signal to the user's first ear; and a second speaker configured to present an audio signal to the user's second ear. The system also includes one or more processors configured to perform the following steps: identify the position of the listener's first ear within the mixed reality environment; identify the position of the listener's second ear within the mixed reality environment; identify a first virtual sound source within the mixed reality environment; identify a first object within the mixed reality environment; determine a first audio signal within the mixed reality environment, the first audio signal originating at the first virtual sound source and intersecting the position of the listener's first ear; determine a second audio signal within the mixed reality environment, the second audio signal originating at the first virtual sound source, intersecting the first object, and intersecting the position of the listener's second ear; determine a third audio signal based on the second audio signal and the first object; present the first audio signal to the first ear via the first speaker; and present the third audio signal to the second ear via the second speaker. In addition to or instead of one or more of the embodiments disclosed above, in some embodiments, the step of determining the third audio signal from the second audio signal includes applying a low-pass filter to the second audio signal, the low-pass filter having parameters based on the first object. In addition to or instead of one or more of the embodiments disclosed above, in some embodiments, the step of determining the third audio signal from the second audio signal includes applying attenuation to the second audio signal, the intensity of the attenuation being based on the first object.In addition to, or alternatively to, one or more of the embodiments disclosed above, in some embodiments, the step of identifying the first object includes the step of identifying a physical object. In addition to, or alternatively to, one or more of the embodiments disclosed above, in some embodiments, the wearable head device further comprises a sensor, and the step of identifying a physical object includes the step of using the sensor to determine the position of the physical object relative to the user within the mixed reality environment. In addition to, or alternatively to, one or more of the embodiments disclosed above, in some embodiments, the sensor comprises a depth camera. In addition to, or alternatively to, one or more of the embodiments disclosed above, in some embodiments, the one or more processors are further configured to perform the step of generating helper data corresponding to the physical object. In addition to, or alternatively to, one or more of the embodiments disclosed above, in some embodiments, the one or more processors are further configured to perform the step of generating a virtual object corresponding to the physical object. In addition to, or alternatively to, one or more of the embodiments disclosed above, in some embodiments, the one or more processors are further configured to perform the step of identifying a second virtual object, the first audio signal intersects the second virtual object, and the fourth audio signal is determined based on the second virtual object.

[0106] It should be noted that the disclosed embodiments have been fully described with reference to the accompanying drawings, and various changes and modifications will be apparent to those skilled in the art. For example, the elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications should be understood to be included within the scope of the disclosed embodiments as defined by the appended claims.

Claims

1. A method, the method comprising: determining, via a sensor, a listener's position within an augmented reality environment; identifying a virtual sound source within the augmented reality environment; identifying an object within the augmented reality environment, wherein the object is associated with one or more properties; identifying the object including identifying a location of the object and the one or more properties; the one or more properties including at least one of a visual property and a material property; determining a first audio signal within the augmented reality environment, the first audio signal originating at the virtual sound source and intersecting the listener's position; in accordance with a determination that the first audio signal intersects the object, determining a second audio signal based on the first audio signal and the object, and presenting the second audio signal to a user's ear via a speaker; and in accordance with a determination that the first audio signal does not intersect the object, omitting determining the second audio signal, and presenting the first audio signal to the user's ear via the speaker. A method comprising the above.

2. The method according to claim 1, wherein a wearable head device comprises the sensor.

3. The method according to claim 1, wherein the sensor includes an inertial measurement unit.

4. The method according to claim 1, wherein the sensor includes a camera.

5. The method according to claim 1, wherein the listener's position within the augmented reality environment is further determined via a second sensor.

6. The object is associated with a property, and the second audio signal is further determined based on the property.

7. The method according to claim 6, wherein the property is associated with at least one of attenuation, filtering, phase shift, delay, reverberation, and modulation.

8. The method according to claim 1, wherein determining the second audio signal based on the first audio signal and the object includes applying at least one of a filter and attenuation to the first audio signal. ​

9. The method according to claim 1, wherein identifying the object includes identifying a virtual object.

10. The method according to claim 1, wherein identifying the object includes identifying a physical object.

11. A system, comprising: a speaker; a sensor; and one or more processors configured to execute a method, the method comprising: determining a listener's position in a mixed reality environment via the sensor; identifying a virtual sound source in the mixed reality environment; identifying an object in the mixed reality environment, wherein the object is associated with one or more properties, identifying the object including identifying a location of the object and the one or more properties, wherein the one or more properties include at least one of a visual property and a material property; determining a first audio signal in the mixed reality environment, the first audio signal originating at the virtual sound source and intersecting the listener's position; determining a second audio signal based on the first audio signal and the object according to a determination that the first audio signal intersects the object, and presenting the second audio signal to the user's ear via the speaker; eliminating determination of the second audio signal and presenting the first audio signal to the user's ear via the speaker according to a determination that the first audio signal does not intersect the object. A system comprising the above.

12. The system according to claim 11, further comprising a wearable head device, the wearable head device comprising the sensor.

13. The system according to claim 11, wherein the sensor includes an inertial measurement unit.

14. The system according to claim 11, wherein the sensor includes a camera.

15. The system according to claim 11, further comprising a second sensor, wherein the listener's position in the mixed reality environment is further determined via the second sensor.

16. The object is associated with a property. ​ ​ ​ ​ ​ The system according to claim 11, wherein the second audio signal is determined further based on the property.

17. Determining the second audio signal based on the first audio signal and the object includes applying at least one of filtering and attenuation to the first audio signal, the system according to claim 11.

18. Identifying the object includes identifying a virtual object, the system according to claim 11.

19. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to execute a method, wherein the method is determining a listener's position in a mixed reality environment via a sensor, identifying a virtual sound source in the mixed reality environment, identifying an object in the mixed reality environment, wherein the object is associated with one or more properties, identifying the object includes identifying the location of the object and the one or more properties, the one or more properties include at least one of a visual property and a material property, determining a first audio signal in the mixed reality environment, the first audio signal originating at the virtual sound source and intersecting the listener's position, in accordance with a determination that the first audio signal intersects the object, determining a second audio signal based on the first audio signal and the object, and presenting the second audio signal to the user's ear via a speaker and in accordance with a determination that the first audio signal does not intersect the object, omitting determining the second audio signal, and presenting the first audio signal to the user's ear via the speaker and including a non-transitory computer-readable medium.

Citation Information

Patent Citations

  • Video game processing device and video game processing program

    JP2012054698A

  • Stereophonic sound calculation method, apparatus, program, recording medium, stereophonic sound presentation system, and virtual reality space presentation system

    JP2013201577A

  • Video game processing device and video game processing program

    JP2014226201A

  • Video analysis-assisted generation of multichannel audio data

    JP2016513410A

  • Information processing device, information processing method, and program

    WO2017183346A1