Multi-application audio rendering
The system addresses auditory inconsistencies in XR systems by using environment-aware audio models to simulate realistic sounds, enhancing user immersion and comfort in mixed reality environments.
Patent Information
- Application Number
- JP2025081001
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-02-14
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-01
AI Technical Summary
Existing XR systems fail to provide realistic and immersive audio experiences by neglecting the user's physical environment, leading to auditory inconsistencies that can cause discomfort and degrade the user experience.
A system and method for rendering audio that considers the user's physical environment by using shared and individual audio models to simulate sounds that match the acoustic properties of the real world, incorporating direct-path, reflective, and reverberant behaviors.
Enhances the user's sense of connection to the mixed reality environment by providing spatially aware and realistic audio that adapts to the user's movement and surroundings, reducing motion sickness and improving overall immersion.
Smart Images

Figure 2025113285000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Application No. 62 / 976,569, filed on February 14, 2020, the content of which is incorporated herein by reference in its entirety.
[0002] The present disclosure generally relates to systems and methods for managing and storing audio data, and more particularly, to systems and methods for managing and storing audio data within a mixed - reality environment.
Background Art
[0003] Virtual environments are prevalent in computing environments and are found in use in video games (where the virtual environment can represent a game world), maps (where the virtual environment can represent terrain to be navigated), simulations (where the virtual environment can simulate a real environment), digital storytelling (where virtual characters can interact with each other within the virtual environment), and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the user experience with virtual environments can be limited by the technology used to present the virtual environment. For example, conventional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may not be able to realize a virtual environment so as to attract people and create a realistic and immersive experience.
[0004] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present sensory information to a user of an XR system corresponding to a virtual environment represented by data within a computer system. Such systems can uniquely increase immersion and realism by combining virtual visual and audio cues with real scenery and sounds. Thus, it may be desirable to present digital sound to a user of an XR system such that the sound is perceived as if it were occurring naturally and consistent with the user's expectations of the sound within the user's real environment. Generally, a user expects that virtual sounds will have the acoustic properties of the real environment in which they are heard. For example, a user of an XR system within a large concert hall expects that the virtual sound of the XR system will have a sound wave quality like that of a large cave, and conversely, a user within a small apartment will expect the sound to be more attenuated, closer, and immediate. In addition to matching the acoustic properties of the virtual sound with the real and / or virtual environment, realism is further enhanced by spatializing the virtual sound. For example, a virtual object may visually fly behind and over the user, and the user may expect the corresponding virtual sound to similarly reflect the spatial movement of the virtual object relative to the user.
[0005] Existing technologies often lack these expectations by presenting virtual audio, for example, that can lead to a sense of inauthenticity, which does not consider the user's surroundings and does not support the spatial movement of virtual objects, potentially degrading the user experience. Observation of XR system users indicates that while users may be relatively tolerant of visual inconsistencies between virtual content and the real environment (e.g., lighting mismatches), they can be more sensitive to auditory inconsistencies. Our own auditory experiences, which are continuously refined throughout our lives, can make us acutely aware of how our physical environment affects the sounds we hear, and we can be very sensitive to sounds that do not match those expectations. In XR systems, such inconsistencies can be unpleasant and can turn an immersive and engaging experience into a gimmicky imitation. In extreme cases, auditory inconsistencies can cause motion sickness and other adverse effects because the inner ear is unable to reconcile the auditory stimuli with their corresponding visual cues.
[0006] A system architecture is required to orchestrate and manage a system for generating virtual audio. The step of generating virtual audio may involve the step of managing and storing information about the user's environment so that the information can be used to produce realistic virtual sounds. The audio system architecture may thus need to interface with other systems and receive and utilize information related to the audio engine. Furthermore, it may be desirable to have an audio system architecture that can present realistic sounds without interruption during use. An audio system architecture that can update the audio engine without interruption can produce an immersive user experience where the auditory signals continuously reflect the user's environment.
[0007] By considering the characteristics of the user's physical environment, the systems and methods described herein can simulate what a user would hear as if the virtual sound were real sound naturally generated within that environment. By presenting virtual sound in a manner faithful to how sound behaves in the real world, the user can experience an enhanced sense of connection to the mixed reality environment. Similarly, by presenting location-aware virtual content that responds to the user's movement and environment, the content becomes more subjective, two-way, and realistic. For example, the user's experience at point A can be entirely different from their experience at point B. This enhanced reality and interaction can provide a basis for new uses of mixed reality, such as enabling new forms of gameplay, social features, or two-way behavior using spatially aware audio. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0008] Embodiments of the present disclosure describe systems and methods for efficiently rendering audio. According to embodiments of the present disclosure, the method includes receiving a request to present a first audio track, the first audio track being based on a first audio model comprising a shared model component and a first model component; receiving a request to present a second audio track, the second audio track being based on a second audio model comprising a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and presenting an audio signal comprising the rendered sound via one or more speakers. The present invention provides, for example, the following. (Item 1) A system comprising One or more speakers, and One or more processors, wherein Receiving a request to present a first audio track, wherein the first audio track is based on a first audio model comprising a shared model component and a first model component; and Receiving a request to present a second audio track, wherein the second audio track is based on a second audio model comprising a shared model component and a second model component; and Rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and Presenting an audio signal comprising the rendered sound via the one or more speakers One or more processors configured to execute a method comprising the above; and A system comprising the above. (Item 2) The system according to item 1, wherein rendering the sound comprises rendering the first audio track and the second audio track based on the shared model component. (Item 3) The system according to item 1, wherein rendering the sound comprises rendering the first audio track based on a first audio component. (Item 4) The system according to item 1, wherein rendering the sound comprises rendering the second audio track based on a second audio component. (Item 5) The system according to item 1, wherein the one or more speakers are speakers of a head-mounted device. (Item 6) The first audio model is the system according to item 5, based on the physical environment of the head-mounted device. (Item 7) The shared model component is the system according to item 1, corresponding to direct-path acoustic behavior. (Item 8) The shared model component is the system according to item 1, corresponding to reflective acoustic behavior. (Item 9) The shared model component is the system according to item 1, corresponding to reverberant acoustic behavior. (Item 10) A method comprising: Receiving a request to present a first audio track, wherein the first audio track is based on a first audio model comprising a shared model component and a first model component; Receiving a request to present a second audio track, wherein the second audio track is based on a second audio model comprising a shared model component and a second model component; Rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; Presenting an audio signal comprising the rendered sound via one or more speakers. A method. (Item 11) The rendering of the sound includes rendering the first audio track and the second audio track based on the shared model component, according to the method of item 10. (Item 12) The rendering of the sound includes rendering the first audio track based on a first audio component, according to the method of item 10. (Item 13) Rendering the sound includes rendering the second audio track based on a second audio component, according to the method of item 10. (Item 14) Presenting the audio signal includes presenting the audio signal via one or more speakers of a head-mounted device, according to the method of item 10. (Item 15) The first audio model is based on the physical environment of the head-mounted device, according to the method of item 14. (Item 16) The shared model component corresponds to direct-path acoustic behavior, according to the method of item 10. (Item 17) The shared model component corresponds to reflective acoustic behavior, according to the method of item 10. (Item 18) The shared model component corresponds to reverberant acoustic behavior, according to the method of item 10. (Item 19) A non-transitory computer-readable medium that stores instructions which, when executed by one or more processors, cause the one or more processors to Receive a request to present a first audio track, the first audio track being based on a first audio model comprising a shared model component and a first model component, and Receive a request to present a second audio track, the second audio track being based on a second audio model comprising a shared model component and a second model component, and Render sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component, and Presenting an audio signal comprising the rendered sound via one or more speakers A non-transitory computer-readable medium for causing the execution of a method including the above. (Item 20) The non-transitory computer-readable medium according to item 19, wherein rendering the sound includes rendering the first audio track and the second audio track based on the shared model component. (Item 21) The non-transitory computer-readable medium according to item 19, wherein rendering the sound includes rendering the first audio track based on a first audio component. (Item 22) The non-transitory computer-readable medium according to item 19, wherein rendering the sound includes rendering the second audio track based on a second audio component. (Item 23) The non-transitory computer-readable medium according to item 19, wherein presenting the audio signal includes presenting the audio signal via one or more speakers of a head-mounted device. (Item 24) The non-transitory computer-readable medium according to item 23, wherein the first audio model is based on the physical environment of the head-mounted device. (Item 25) The non-transitory computer-readable medium according to item 19, wherein the shared model component corresponds to direct-path acoustic behavior. (Item 26) The non-transitory computer-readable medium according to item 19, wherein the shared model component corresponds to reflective acoustic behavior. (Item 27) The non-transitory computer-readable medium according to item 19, wherein the shared model component corresponds to reverberant acoustic behavior.
Brief Description of the Drawings
[0009]
Figure 1A
Figure 1B
Figure 1C
[0010]
Figure 2A
Figure 2B
Figure 2C
Figure 2D
[0011]
Figure 3A
[0012]
Figure 3B
[0013]
Figure 4
[0014]
Figure 5
[0015]
Figure 6
[0016]
Figure 7
[0017]
Figure 8
[0018]
Figure 9A
Figure 9B
[0019] Detailed Description In the following description of the embodiments, reference is made to the accompanying drawings which form a part hereof, and in which are shown by way of illustration specific embodiments in which the invention may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the disclosed embodiments.
[0020] Mixed Reality Environment
[0021] Like all people, users of a mixed reality system are present within a physical environment, i.e., the three-dimensional portion of the "real world" and all of its contents are perceivable by the user. For example, the user uses normal human senses, i.e., vision, hearing, touch, taste, and smell, to perceive the physical environment and interacts with the physical environment by moving their body within the physical environment. Locations within the physical environment can be described as coordinates within a coordinate space. For example, the coordinates can include latitude, longitude, and altitude relative to sea level, distances in three orthogonal dimensions from a reference point, or other suitable values. Similarly, a vector can describe a quantity having a direction and magnitude within a coordinate space.
[0022] A computing device can maintain a representation of a virtual environment, for example, within a memory associated with the device. As used herein, a virtual environment is a calculated representation of a three-dimensional space. The virtual environment can include a representation of any object, action, signal, parameter, coordinate, vector, or other characteristic associated with that space. In some embodiments, a circuit (e.g., a processor) of the computing device can maintain and update the state of the virtual environment. That is, the processor can determine the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or input provided by a user at a first time t0. For example, if an object within the virtual environment is located at a first coordinate at time t0, has certain programmed physical parameters (e.g., mass, coefficient of friction), and the input received from the user indicates that a force should be applied to the object in a certain direction vector, the processor can apply the laws of kinematics and use basic mechanics to determine the location of the object at time t1. The processor can use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at time t1. When maintaining and updating the state of the virtual environment, the processor can execute any suitable software, including software related to creating and deleting virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling inputs and outputs, software for implementing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.
[0023] Output devices such as displays or speakers can present to the user any or all aspects of the virtual environment. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, light, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, viewing axis, and frustum), and render on the display a visible scene of the virtual environment corresponding to that view. Any suitable rendering technique may be used for this purpose. In some embodiments, the visible scene may include only some of the virtual objects within the virtual environment and may exclude other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects within the virtual environment may generate sounds arising from the location coordinates of the objects (e.g., a virtual character can speak or produce sound effects), or the virtual environment may be associated with music cues or ambient sounds that may or may not be associated with a particular location. The processor can determine an audio signal corresponding to "listener" coordinates, e.g., an audio signal that is mixed and processed to simulate an audio signal that would be heard by the listener at the listener coordinates corresponding to the synthesis of sounds within the virtual environment, and present the audio signal to the user via one or more speakers.
[0024] Since the virtual environment only exists as a computational structure, a user cannot directly perceive the virtual environment using normal senses. Instead, a user can only indirectly perceive the virtual environment as presented to the user, for example, by a display, speakers, a haptic output device, etc. Similarly, a user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data to a processor via an input device or sensor to update the virtual environment using device or sensor data. For example, a camera sensor can provide optical data indicating that a user is attempting to move an object in the virtual environment, and the processor can use that data to appropriately respond with the object within the virtual environment.
[0025] A mixed reality system can present a user with a mixed reality environment ("MRE") that combines aspects of the real and virtual environments using, for example, a see-through display and / or one or more speakers (e.g., that can be incorporated within a wearable head device). In some embodiments, one or more speakers may be external to the head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of the real environment and the corresponding virtual environment. In some examples, the corresponding real and virtual environments share a single coordinate space. In some examples, the real coordinate space and the corresponding virtual coordinate space are related to each other by a transformation matrix (or other suitable representation). Thus, a single coordinate (along with, in some examples, a transformation matrix) can define a first location within the real environment and also a second corresponding location within the virtual environment, and vice versa.
[0026] In MRE, a virtual object (e.g., within a virtual environment associated with MRE) can correspond to a real object (e.g., within a real environment associated with MRE). For example, if the real environment of MRE includes a real street lamp post (real object) at a certain location coordinate, the virtual environment of MRE may include a virtual street lamp post (virtual object) at the corresponding location coordinate. As used herein, a real object, in combination with its corresponding virtual object, constitutes a "composite reality object". It is not necessary for the virtual object to perfectly match or align with the corresponding real object. In some embodiments, the virtual object can be a simplified version of the corresponding real object. For example, if the real environment includes a real street lamp post, the corresponding virtual object may include a cylinder of generally the same height and radius as the real street lamp post (reflecting that the street lamp post can be of a generally cylindrical shape). Simplifying the virtual object in this way can enable calculation efficiency and simplify the calculations to be performed on such virtual objects. Further, in some embodiments of MRE, not all real objects in the real environment need to be associated with corresponding virtual objects. Similarly, in some embodiments of MRE, not all virtual objects in the virtual environment need to be associated with corresponding real objects. That is, some virtual objects can exist only within the virtual environment of MRE without any real-world counterpart.
[0027] In some embodiments, a virtual object may sometimes have significantly different characteristics from those of the corresponding real object. For example, the real environment within the MRE may include a cactus with two green branches extending, i.e., an inanimate object covered with thorns, while the corresponding virtual object within the MRE may have the characteristics of a virtual character with two green arms accompanied by human facial features and a surly attitude. In this embodiment, the virtual object is similar to its corresponding real object in some characteristics (color, number of arms), but different from the real object in other characteristics (facial features, personality). Thus, virtual objects have the potential to represent real objects in a creative, abstract, exaggerated, or fictional manner, or to endow real objects that would otherwise be inanimate with behavior (e.g., human personality). In some embodiments, the virtual object may be a purely fictional creation without a real-world counterpart (e.g., perhaps a virtual monster within the virtual environment at a location corresponding to a void within the real environment).
[0028] Compared to a VR system that presents a virtual environment to a user while obscuring the real environment, a mixed reality (MR) system that presents an MR environment provides the advantage that the real environment remains perceivable while the virtual environment is presented. Thus, a user of an MR system can experience and interact with a corresponding virtual environment using visual and audio cues associated with the real environment. As an example, a user of a VR system may struggle to perceive or interact with virtual objects displayed within the virtual environment as, as described above, the user cannot directly perceive or interact with the virtual environment, whereas a user of an MR system may find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching corresponding real objects within their own real environment. This level of interaction can enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by presenting the real and virtual environments simultaneously, an MR system can reduce negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. The MR system also offers many possibilities for applications that can extend or modify our experience of the real world.
[0029] Figure 1A illustrates an exemplary real - world environment 100 in which a user 110 uses a mixed - reality system 112. The mixed - reality system 112 may include a display (e.g., a transmissive display) and one or more speakers, and one or more sensors (e.g., cameras) as described below. The illustrated real - world environment 100 includes a rectangular room 104A in which the user 110 stands, and real objects 122A (lamp), 124A (table), 126A (sofa), and 128A (painting). The room 104A further includes location coordinates 106, which may be considered the origin of the real - world environment 100. As shown in Figure 1A, an environment / world coordinate system 108 (comprising an x - axis 108X, a y - axis 108Y, and a z - axis 108Z) associated with the origin at point 106 (world coordinates) may define a coordinate space for the real - world environment 100. In some embodiments, the origin 106 of the environment / world coordinate system 108 may correspond to the location where the power of the mixed - reality system 112 is turned on. In some embodiments, the origin 106 of the environment / world coordinate system 108 may be reset during operation. In some examples, the user 110 may be considered a real object within the real - world environment 100. Similarly, the body parts of the user 110 (e.g., hands, feet) may be considered real objects within the real - world environment 100. In some examples, a user / listener / head coordinate system 114 (comprising an x - axis 114X, a y - axis 114Y, and a z - axis 114Z) associated with the origin at point 115 (e.g., user / listener / head coordinates) may define a coordinate space for the user / listener / head on which the mixed - reality system 112 is located. The origin 115 of the user / listener / head coordinate system 114 may be defined with respect to one or more components of the mixed - reality system 112. For example, the origin 115 of the user / listener / head coordinate system 114 may be defined with respect to the display of the mixed - reality system 112 during initial calibration etc. of the mixed - reality system 112. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrices) or other suitable representation can characterize the transformation between the user / listener / head coordinate system 114 space and the environment / world coordinate system 108 space.In some embodiments, the left ear coordinates 116 and the right ear coordinates 117 may be defined relative to the origin 115 of the user / listener / head coordinate system 114. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the left ear coordinates 116 and the right ear coordinates 117 and the user / listener / head coordinate system 114 space. The user / listener / head coordinate system 114 can simplify the representation of the location of the user's head or head-mounted device relative to, for example, the environment / world coordinate system 108. Using simultaneous localization and mapping (SLAM), visual odometry, or other techniques, the transformation between the user coordinate system 114 and the environment coordinate system 108 can be determined and updated in real time.
[0030] FIG. 1B illustrates an exemplary virtual environment 130 corresponding to the real environment 100. The virtual environment 130 shown includes a virtual rectangular room 104B corresponding to the real rectangular room 104A, a virtual object 122B corresponding to the real object 122A, a virtual object 124B corresponding to the real object 124A, and a virtual object 126B corresponding to the real object 126A. The metadata associated with the virtual objects 122B, 124B, 126B can include information derived from the corresponding real objects 122A, 124A, 126A. The virtual environment 130 additionally includes a virtual monster 132, which does not correspond to any real object within the real environment 100. The real object 128A within the real environment 100 does not correspond to any virtual object within the virtual environment 130. A persistent coordinate system 133 (comprising an x-axis 133X, a y-axis 133Y, and a z-axis 133Z) with its origin at point 134 (persistent coordinates) can define a coordinate space for the virtual content. The origin 134 of the persistent coordinate system 133 may be defined relative to / with respect to one or more real objects such as the real object 126A. A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the persistent coordinate system 133 space and the environment / world coordinate system 108 space. In some embodiments, the virtual objects 122B, 124B, 126B, and 132 may each have their own persistent coordinate points relative to the origin 134 of the persistent coordinate system 133. In some embodiments, there may be multiple persistent coordinate systems, and the virtual objects 122B, 124B, 126B, and 132 may each have their own persistent coordinate points relative to one or more persistent coordinate systems.
[0031] With respect to FIGS. 1A and 1B, the environment / world coordinate system 108 defines a shared coordinate space for both the real environment 100 and the virtual environment 130. In the illustrated embodiment, the coordinate space has its origin at point 106. Further, the coordinate space is defined by the same three orthogonal axes (108X, 108Y, 108Z). Thus, a first location within the real environment 100 and a second corresponding location within the virtual environment 130 can be described with respect to the same coordinate space. This simplifies identifying and presenting corresponding locations within the real and virtual environments since the same coordinates can be used to identify both locations. However, in some embodiments, the corresponding real and virtual environments need not use a shared coordinate space. For example, in some embodiments (not shown), a matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.
[0032] FIG. 1C illustrates an exemplary MRE 150 that simultaneously presents side views of the real environment 100 and the virtual environment 130 to the user 110 via the mixed reality system 112. In the illustrated embodiment, the MRE 150 simultaneously presents to the user 110 real objects 122A, 124A, 126A, and 128A from the real environment 100 (e.g., via the transmissive portion of the display of the mixed reality system 112) and virtual objects 122B, 124B, 126B, and 132 from the virtual environment 130 (e.g., via the active display portion of the display of the mixed reality system 112). As described above, the origin 106 acts as the origin for the coordinate space corresponding to the MRE 150, and the coordinate system 108 defines the x-axis, y-axis, and z-axis for the coordinate space.
[0033] In the illustrated embodiments, the composite reality object includes a corresponding pair of real and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) that occupy corresponding locations within the coordinate space 108. In some embodiments, both the real object and the virtual object may be visible to the user 110 simultaneously. This may be desirable, for example, in instances where the virtual object presents information designed to augment the view of the corresponding real object (such as in museum applications where the virtual object presents the missing portions of an ancient damaged statue). In some embodiments, the virtual objects (122B, 124B, and / or 126B) may be displayed so as to occlude the corresponding real objects (122A, 124A, and / or 126A) (e.g., via active pixelated occlusion using a pixelated occlusion shutter). This may be desirable, for example, in instances where the virtual object acts as a visual replacement for the corresponding real object (such as in two-way storytelling applications where an inanimate real object becomes a "living" character).
[0034] In some embodiments, a real object (e.g., 122A, 124A, 126A) may be associated with virtual content or helper data that does not necessarily constitute a virtual object. The virtual content or helper data can facilitate the processing or handling of virtual objects within the composite reality environment. For example, such virtual content may include a two-dimensional representation of the corresponding real object, a custom asset type associated with the corresponding real object, or statistical data associated with the corresponding real object. This information can enable or facilitate calculations involving the real object without incurring unnecessary computational overhead.
[0035] In some embodiments, the presentation described above may also incorporate an audio aspect. For example, in MRE150, the virtual monster 132 may be associated with one or more audio signals, such as a footstep effect, that are generated as the monster walks around MRE150. As further described below, a processor of the mixed reality system 112 may calculate an audio signal corresponding to the mixing and processed synthesis of all such sounds within MRE150, and present the audio signal to the user 110 via one or more speakers included within the mixed reality system 112 and / or one or more external speakers.
[0036] Exemplary Mixed Reality System
[0037] The exemplary mixed reality system 112 can include a display (which can include left and right transmissive displays, which can be near-eye displays, and associated components for coupling light from the display to the user's eyes), left and right speakers (e.g., positioned adjacent to the user's left and right ears, respectively), an inertial measurement unit (IMU) (e.g., mounted on the arm of the head device's stem), an orthogonal coil electromagnetic receiver (e.g., mounted on the left stem component), left and right cameras (e.g., depth (time-of-flight) cameras) oriented away from the user, and left and right eye cameras (e.g., for detecting the user's eye movements) oriented towards the user, in a wearable head device (e.g., a wearable augmented or mixed reality head device). However, the mixed reality system 112 can incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic). Additionally, the mixed reality system 112 can incorporate networking features (e.g., Wi-Fi capabilities) and communicate with other devices and systems, including other mixed reality systems. The mixed reality system 112 can further include a battery (which may be mounted within an auxiliary unit such as a belt pack designed to be worn around the user's waist), a processor, and a memory. The wearable head device of the mixed reality system 112 can include a tracking component, such as an IMU or other suitable sensor, configured to output a coordinate set of the wearable head device relative to the user's environment. In some embodiments, the tracking component can provide an input to the processor and perform simultaneous localization and mapping (SLAM) and / or visual odometry algorithms. In some embodiments, the mixed reality system 112 can also include a handheld controller 300 and / or an auxiliary unit 320, which can be a wearable belt pack, as further described below.
[0038] Figures 2A-2D illustrate components of an exemplary mixed reality system 200 (which may correspond to mixed reality system 112) that can be used to present an MRE (which may correspond to MRE150) or other virtual environment to a user. FIG. 2A illustrates a perspective view of a wearable head device 2102 included within the exemplary mixed reality system 200. FIG. 2B illustrates a top view of the wearable head device 2102 worn on a user's head 2202. FIG. 2C illustrates a front view of the wearable head device 2102. FIG. 2D illustrates an edge view of an exemplary eyepiece 2110 of the wearable head device 2102. As shown in FIGS. 2A-2C, the exemplary wearable head device 2102 includes an exemplary left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an exemplary right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 can include a transmissive element through which a real environment is visible and a display element for presenting a display (e.g., via light modulated for each image) that overlays the real environment. In some embodiments, such a display element can include a surface diffractive optical element for controlling the flow of light modulated for each image. For example, the left eyepiece 2108 can include a left internal coupling grating set 2112, a left orthogonal pupil expansion (OPE) grating set 2120, and a left exit (output) pupil expansion (EPE) grating set 2122. Similarly, the right eyepiece 2110 can include a right internal coupling grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. Light modulated for each image can be transferred to the user's eyes through the internal coupling gratings 2112 and 2118, the OPEs 2114 and 2120, and the EPEs 2116 and 2122. Each internal coupling grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 can be designed to gradually deflect light downward toward its associated EPE 2122, 2116, thereby horizontally extending the formed exit pupil.Each EPE 2122, 2116 can be configured to gradually and outwardly redirect at least a portion of the light received from its corresponding OPE grating set 2120, 2114 to a user eye box position (not shown) defined behind the eyepieces 2108, 2110, such that the exit pupil formed at the eye box extends vertically. Alternatively, instead of the internal coupling grating sets 2112 and 2118, the OPE grating sets 2114 and 2120, and the EPE grating sets 2116 and 2122, the eyepieces 2108 and 2110 can include gratings and / or other arrangements of refractive and reflective features for controlling the coupling of light modulated for each image to the user's eyes.
[0039] In some embodiments, the wearable head device 2102 can include a left arm 2130 and a right arm 2132, the left arm 2130 including a left speaker 2134 and the right arm 2132 including a right speaker 2136. The orthogonal coil electromagnetic receiver 2138 can be located in the left temple component or another suitable location within the wearable head unit 2102. The inertial measurement unit (IMU) 2140 can be located in the right arm 2132 or another suitable location within the wearable head device 2102. The wearable head device 2102 can also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142, 2144 can preferably be oriented in different directions such that together they cover a wider field of view.
[0040] In the embodiment shown in FIGS. 2A-2D, the left source 2124 of light modulated for each image can be optically coupled into the left eyepiece lens 2108 through the left internal coupling grating set 2112, and the right source 2126 of light modulated for each image can be optically coupled into the right eyepiece lens 2110 through the right internal coupling grating set 2118. The sources 2124, 2126 of light modulated for each image can include, for example, a projector including an electro-optical modulator such as an optical fiber scanner, a digital light processing (DLP) chip, or a liquid crystal on silicon (LCoS) modulator, or a light-emitting display such as a micro light-emitting diode (μLED) or a micro organic light-emitting diode (μOLED) panel that is coupled into the internal coupling grating sets 2112, 2118 using one or more lenses per side. The input coupling grating sets 2112, 2118 can deflect the light from the sources 2124, 2126 of light modulated for each image at an angle greater than the critical angle for total internal reflection (TIR) for the eyepiece lenses 2108, 2110. The OPE grating sets 2114, 2120 gradually deflect the propagating light downward toward the EPE grating sets 2116, 2122 by TIR. The EPE grating sets 2116, 2122 gradually couple the light toward the user's face, including the pupil of the user's eye.
[0041] In some embodiments, as shown in FIG. 2D, the left eyepiece lens 2108 and the right eyepiece lens 2110 each include a plurality of waveguides 2402. For example, each eyepiece lens 2108, 2110 can include a plurality of individual waveguides, each dedicated to an individual color channel (e.g., red, blue, and green). In some embodiments, each eyepiece lens 2108, 2110 can include a plurality of sets of such waveguides, each set configured to impart a different wavefront curvature to the emitted light. The wavefront curvature can be convex with respect to the user's eye, for example, to present a virtual object positioned at a certain distance in front of the user (e.g., a distance corresponding to the reciprocal of the wavefront curvature). In some embodiments, the EPE grating sets 2116, 2122 can include curved grating grooves to provide a convex wavefront curvature by modifying the Poynting vector of the light exiting across each EPE.
[0042] In some embodiments, to create the perception that the displayed content is three-dimensional, the stereoscopically adjusted left and right eye images can be presented to the user through the light modulators 2124, 2126 and the eyepiece lenses 2108, 2110 for each image. The perceived realism of the presentation of the three-dimensional virtual object can be improved by selecting the waveguides (and thus the corresponding wavefront curvature) such that the virtual object is displayed at a distance approximating the distance indicated by the stereoscopic left and right images. The technique can also reduce motion sickness experienced by some users, which can be caused by differences between the depth perception cues provided by the stereoscopic left and right eye images and the automatic focusing of the human eye (e.g., object distance-dependent focusing).
[0043] FIG. 2D illustrates an edge view from above the right eyepiece lens 2110 of the exemplary wearable head device 2102. As shown in FIG. 2D, the plurality of waveguides 2402 can include a first subset 2404 of three waveguides and a second subset 2406 of three waveguides. The two subsets 2404, 2406 of waveguides can be distinguished by different EPE gratings, each characterized by a different grating line curvature, to impart different wavefront curvatures to the emitted light. Within each of the subsets 2404, 2406 of waveguides, each waveguide can be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although not shown in FIG. 2D, the structure of the left eyepiece lens 2108 is similar to the structure of the right eyepiece lens 2110.)
[0044] Figure 3A illustrates an exemplary handheld controller component 300 of the mixed reality system 200. In some embodiments, the handheld controller 300 includes a grip portion 346 and one or more buttons 350 disposed along an upper surface 348. In some embodiments, the buttons 350 may be configured to be used as optical tracking targets for tracking the six degrees of freedom (6DOF) movement of the handheld controller 300 in conjunction with, for example, a camera or other optical sensor (which may be mounted within a head unit of the mixed reality system 200, such as the wearable head device 2102). In some embodiments, the handheld controller 300 includes a tracking component (such as an IMU or other suitable sensor) for detecting a position or orientation, such as a position or orientation relative to the wearable head device 2102. In some embodiments, such a tracking component may be positioned within the handle of the handheld controller 300 and / or may be mechanically coupled to the handheld controller. The handheld controller 300 can be configured to provide one or more output signals corresponding to one or more of a button press state, or the position, orientation, and / or movement of the handheld controller 300 (e.g., via an IMU). Such output signals may be used as inputs to a processor of the mixed reality system 200. Such inputs may correspond to the position, orientation, and / or movement of the handheld controller (and further, the position, orientation, and / or movement of the user's hand holding the controller). Such inputs may also correspond to the user pressing the button 350.
[0045] Figure 3B illustrates an exemplary auxiliary unit 320 of the mixed reality system 200. The auxiliary unit 320 can include a battery that provides energy and operates the system 200, and can include a processor that executes a program and operates the system 200. As shown, the exemplary auxiliary unit 320 includes a clip 2128 for attaching the auxiliary unit 320 to the user's belt or the like. It will also be apparent that other form factors are suitable for the auxiliary unit 320 and can include form factors that do not involve mounting the unit on the user's belt. In some embodiments, the auxiliary unit 320 is coupled to the wearable head device 2102 through a multi-tubular cable that can include, for example, electrical wires and optical fibers. A wireless connection between the auxiliary unit 320 and the wearable head device 2102 can also be used.
[0046] In some embodiments, the mixed reality system 200 includes one or more microphones that can detect sound and provide a corresponding signal to the mixed reality system. In some embodiments, the microphone may be attached to or integrated with the wearable head device 2102 and may be configured to detect the user's voice. In some embodiments, the microphone may be attached to or integrated with the handheld controller 300 and / or the auxiliary unit 320. Such a microphone may be configured to detect ambient sound, ambient noise, the voice of the user or a third party, or other sounds.
[0047] FIG. 4 shows an exemplary functional block diagram corresponding to an exemplary mixed reality system such as the mixed reality system 200 described above (which may correspond to the mixed reality system 112 of FIG. 1). As shown in FIG. 4, the exemplary handheld controller 400B (which may correspond to the handheld controller 300 (“Totem”)) includes a totem / wearable head device 6 degrees of freedom (6DOF) totem subsystem 404A, and the exemplary wearable head device 400A (which may correspond to the wearable head device 2102) includes a totem / wearable head device 6DOF subsystem 404B. In an embodiment, the 6DOF totem subsystem 404A and the 6DOF subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom may be represented relative to the coordinate system of the wearable head device 400A. The three translational offsets may be represented as X, Y, and Z offsets, a translation matrix, or some other representation within such a coordinate system. The rotational degrees of freedom may be represented as a sequence of yaw, pitch, and roll rotations, as a rotation matrix, as a quaternion, or as some other representation. In some embodiments, the wearable head device 400A, one or more depth cameras 444 (and / or one or more non-depth cameras) included within the wearable head device 400A, and / or one or more optical targets (e.g., the button 350 of the handheld controller 400B as described above or a dedicated optical target included within the handheld controller 400B) may be used for 6DOF tracking. In some embodiments, the handheld controller 400B may include a camera as described above, and the wearable head device 400A may include an optical target for optical tracking in conjunction with the camera. In some embodiments, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids, which are used to wirelessly transmit and receive three distinguishable signals.The 6DOF of the wearable head device 400A relative to the handheld controller 400B can be determined by measuring the relative magnitudes of three distinguishable signals received within each of the coils for reception. Additionally, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that is useful for providing improved accuracy and / or more timely information regarding fast movement of the handheld controller 400B.
[0048] In some embodiments, for example, in order to compensate for the movement of the wearable head device 400A relative to the coordinate system 108, it may be necessary to convert coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment). For example, such a conversion may cause the display of the wearable head device 400A to present virtual objects at their expected positions and orientations relative to the real environment rather than at fixed positions and orientations on the display (e.g., the same position at the lower right corner of the display), and to preserve the illusion that the virtual objects exist within the real environment (and, for example, do not appear unnaturally positioned within the real environment as the wearable head device 400A drifts and rotates). In some embodiments, the compensation transformation between coordinate spaces can be determined by processing images from the depth camera 444 using SLAM and / or visual odometry procedures to determine the transformation of the wearable head device 400A relative to the coordinate system 108. In the embodiment shown in FIG. 4, the depth camera 444 can be coupled to the SLAM / visual odometry block 406 and provide images to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the present image and then determine the position and orientation of the user's head, which can be used to identify the transformation between the head coordinate space and another coordinate space (e.g., the inertial coordinate space). Similarly, in some embodiments, an additional source of information regarding the user's head pose and location is obtained from the IMU 409. The information from the IMU 409 can be integrated with the information from the SLAM / visual odometry block 406 to provide more timely information regarding improved accuracy and / or fast adjustment of the user's head pose and position.
[0049] In some embodiments, the depth camera 444 can supply a 3D image to a hand gesture tracker 411 that can be implemented within the processor of the wearable head device 400A. The hand gesture tracker 411 can identify a user's hand gesture, for example, by matching the 3D image received from the depth camera 444 to a stored pattern representing the hand gesture. Other suitable techniques for identifying a user's hand gesture will also be apparent.
[0050] In some embodiments, one or more processors 416 may be configured to receive data from the 6DOF headgear subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, and / or hand gesture tracker 411 of the wearable head device. The processor 416 may also be able to send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A in embodiments where the handheld controller 400B is not tethered. The processor 416 may further communicate with additional components such as the audio / visual content memory 418, graphical processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to the head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to the left source 424 of light modulated per image and a right channel output coupled to the right source 426 of light modulated per image. The GPU 420 may output stereoscopic image data to the sources 424, 426 of light modulated per image, as described above with respect to FIGS. 2A-2D, for example. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or right speaker 414. The DSP audio spatializer 422 may receive, from the processor 419, an input indicating a direction vector from the user to a virtual sound source (e.g., that can be moved by the user via the handheld controller 320). Based on the direction vector, the DSP audio spatializer 422 may be able to determine the corresponding HRTF (e.g., by accessing the HRTF or interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the credibility and realism of virtual sounds by incorporating the user's relative position and orientation with respect to the virtual sounds within the composite reality environment, i.e., by presenting virtual sounds that match the user's expectations of what the virtual sounds would sound like if they were real sounds within the real environment.
[0051] In some embodiments, such as those shown in FIG. 4, one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 may be included within an auxiliary unit 400C (which may correspond to the auxiliary unit 320 described above). The auxiliary unit 400C may include a battery 427 to power its components and / or supply power to the wearable head device 400A or the handheld controller 400B. Including such components within an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.
[0052] FIG. 4 presents elements corresponding to various components of an exemplary composite reality system, but various other suitable arrangements of these components will be apparent to those skilled in the art. For example, the elements presented in FIG. 4 as being associated with the auxiliary unit 400C may instead be associated with the wearable head device 400A or the handheld controller 400B. Further, some composite reality systems may eliminate the handheld controller 400B or the auxiliary unit 400C entirely. Such changes and modifications should be understood to be within the scope of the disclosed embodiments.
[0053] Ambient Acoustic Persistence
[0054] As described above, MRE (such as, experienced through a mixed reality system 112, which may include components such as the wearable head unit 200, the handheld controller 300, or the auxiliary unit 320, etc., described above) can present an audio signal to the user of the MRE that appears to occur at a sound source with origin coordinates within the MRE. That is, the user can perceive these audio signals as if they were actual audio signals originating from the origin coordinates of the sound source.
[0055] In some cases, the audio signals can be considered virtual in that they correspond to calculated signals within a virtual environment. The virtual audio signals can be presented to the user as actual audio signals that are detectable by a human ear, for example, generated through the speakers 2134 and 2136 of the wearable head unit 200 in FIG. 2.
[0056] The sound source can correspond to a real object and / or a virtual object. For example, a virtual object (e.g., the virtual monster 132 in FIG. 1C) can emit an audio signal within the MRE, which is represented as a virtual audio signal within the MRE and presented to the user as a real audio signal. For example, the virtual monster 132 in FIG. 1C can emit a virtual sound corresponding to the speech (e.g., dialogue) or sound effect of the monster. Similarly, a real object (e.g., the real object 122A in FIG. 1C) can also be made to appear to emit a virtual audio signal within the MRE, which is represented as a virtual audio signal within the MRE and presented to the user as a real audio signal. For example, the real lamp 122A can emit a virtual sound corresponding to the sound effect of the lamp being switched on or off even when the lamp is not actually switched on or off in the real environment. The virtual sound can correspond to the position and orientation of the sound source (whether real or virtual). For example, when the virtual sound is presented to the user as a real audio signal (e.g., via speakers 2134 and 2136), the user can perceive the virtual sound as originating from the position of the sound source. Even if the underlying object itself that is seemingly made to emit sound corresponds to a real or virtual object as described above, in this specification, it is referred to as a "virtual sound source".
[0057] Some virtual or augmented reality environments suffer from the perception that the environment does not feel real or authentic. One reason for this perception is that audio and visual cues do not always match each other within such an environment. For example, if a user is positioned behind a brick wall within an MRE, the user might expect the sound originating from behind the brick wall to be quieter and more muffled than the sound originating from right next to the user. This expectation is based on the user's auditory experience in the real world where sound becomes quieter and more muffled when passing behind large, dense objects. When the user is presented with an audio signal that is supposed to originate from behind the brick wall but is not muffled and is presented at full volume, the illusion that the sound is originating from behind the brick wall is impaired. Since the entire virtual experience does not fully match the user's expectations based on real-world interactions, it can feel fake and inauthentic. Additionally, in some cases, the "uncanny valley" problem can occur, where even a slight difference between the virtual and real experiences can cause an increased sense of discomfort. It is desirable to improve the user's experience by presenting audio signals within the MRE that appear to interact, even slightly, with the objects within the user's environment. The more such audio signals are consistent with the user's expectations based on real-world experiences, the more immersive and engaging the user's experience within the MRE can become.
[0058] One way for a user to perceive and understand their surrounding environment is through audio cues. In the real world, the actual audio signals that are audible to the user are affected by the location where those audio signals originate and the objects with which the audio signals interact. For example, all other factors being equal, a sound that occurs at a distance from the user (e.g., a dog barking in the distance) will appear quieter than the same sound that occurs at a short distance from the user (e.g., a dog barking in the same room as the user). The user can thus identify the location of the dog in the real environment, based in part on the perceived volume of its bark. Similarly, all other factors being equal, a sound that is moving away from the user (e.g., the voice of a person facing outward from the user) will not be as clear and will appear more muffled (i.e., low-pass filtered) than the same sound that is moving toward the user (e.g., the voice of a person facing the user). The user can thus identify the orientation of the person in the real environment, based on the perceived characteristics of their voice.
[0059] The user's perception of the actual audio signal can also be affected by the presence of objects in the environment with which the audio signal interacts. That is, the user can perceive not only the audio signal generated by the sound source, but also the reflection of that audio signal off neighboring objects and the reverberation signature imparted by the surrounding acoustic space. For example, when a person is speaking in a small room with nearby walls, those walls can produce short natural reverberation signals as the person's voice reflects off the walls. The user can infer from those reverberations that they are in a small room with nearby walls. Similarly, a large concert hall or cathedral can produce longer reverberations from which the user can infer that they are in a large, spacious room. Similarly, the reverberations of the audio signal can have various acoustic characteristics based on the position or orientation of the surfaces off which those signals reflect, or the materials of those surfaces. For example, the reverberations off a tiled wall will sound different from those off a brick, carpet, drywall, or other material. These reverberation characteristics can be used by the user to acoustically understand the size, shape, and material composition of the space in which they are located.
[0060] The above examples illustrate how audio cues can provide information about the user's surrounding environment. These cues can act in combination with visual cues. For example, if a user can see a dog in the distance, the user can expect the dog's bark to correspond to that distance (if not, they may be confused or disoriented, as in some virtual environments). In some embodiments, in low-light environments or with respect to visually impaired users, for example, visual cues may be limited or unavailable, and in such cases, audio cues can take on particular importance and serve as the primary means for the user to understand the environment.
[0061] A system architecture may be useful for compiling, storing, retrieving, and / or managing information necessary to present realistic virtual audio. For example, an MR system (e.g., MR systems 112, 200) may manage environmental information such as the physical environment in which a user may be present, the acoustic properties the physical environment may have, and / or the location within that physical environment where the user may be located. The MR system may further manage information about objects in the physical and / or virtual environment (e.g., objects that may affect the general acoustic properties of the physical environment and / or objects that may affect the acoustic properties of virtual sound sources that interact with the objects). The MR system may also manage information about virtual sound sources. For example, the location where a virtual sound source is located may be relevant when rendering realistic virtual audio.
[0062] In addition to managing the virtual audio system, it may also be necessary to manage other systems simultaneously in order to present a complete MR experience. For example, a complete MR experience may require a virtual visual system, which may manage information used to render virtual objects. A complete MR experience may require a simultaneous localization and mapping system (“SLAM”), which may construct, update, and / or maintain a three-dimensional model of the user's environment. An MR system (e.g., MR systems 112, 200) may manage these systems and may further present a complete MR experience in addition to the virtual audio system. The virtual audio system architecture may be useful for managing the interaction between these systems and facilitating data transfer, management, storage, and / or security.
[0063] In some embodiments, the system (e.g., a virtual audio system) may interact with other, higher-level systems. In some embodiments, a lower-level system (e.g., a virtual audio system) may interact more closely with hardware-level inputs and / or outputs, while a higher-level system (e.g., an application) may interface with the lower-level system. The higher-level system may utilize the lower-level system to perform its functions (e.g., a game application may rely on a lower-level virtual audio system to render realistic virtual audio). The virtual audio system may benefit from a system architecture designed to manage interactions with higher-level systems while maintaining the integrity of the virtual audio system. For example, multiple higher-level systems (e.g., multiple third-party applications) may interface with the virtual audio system simultaneously or substantially simultaneously. In some embodiments, maintaining a single virtual audio system capable of rendering virtual audio may be computationally more efficient than having each higher-level system maintain a separate audio system. For example, in some embodiments, a single digital reverberator may be used to process sound objects from multiple higher-level systems when those objects are intended to be present within the same virtual or physical acoustic space (e.g., a room in which the user is present). A well-designed system architecture may also protect the integrity of information (e.g., from data corruption and / or unauthorized disclosure) that may be used in other applications.
[0064] In some embodiments, it may be advantageous to design the system architecture such that changes can be made in real time without interfering with services to other systems (e.g., higher-level systems). For example, a virtual audio system may store, maintain, or otherwise manage an audio model that takes into account the acoustic properties of the real environment (e.g., a room). If the user changes the real environment (e.g., moves to a different room), the audio model may be updated to account for the change in the real environment. If an MR system is currently in use (e.g., the MR system is presenting virtual vision and / or virtual audio to the user), it may be necessary to update the audio model when a different system (e.g., a higher-level system) is still using the audio model to render virtual audio.
[0065] In some embodiments, it may be advantageous to design the system architecture to propagate changes to other systems. For example, some systems may maintain separate copies of the audio model, or some systems may store specific repeating sound effects that are rendered using the audio model maintained by the virtual audio system. Thus, it may be advantageous to propagate changes made to the virtual audio system to other systems. For example, if the user changes the environment (e.g., moves rooms) and the new audio model may be more accurate, the virtual audio system may modify the audio model and notify any clients (e.g., systems that use and / or may rely on the virtual audio system) of the change. In some embodiments, the clients may then query the virtual audio system and update their internal data as appropriate.
[0066] Figure 5 illustrates an exemplary virtual audio system according to some embodiments. The virtual audio system 500 can include a persistence module 502. A module (e.g., the persistence module 502) can include one or more computer systems configured to execute instructions and / or store one or more data structures. In some embodiments, a module (e.g., the persistence module 502) can be configured to execute processes, sub-processes, threads, and / or services managed by an audio service 522 (e.g., instructions executed by the persistence module 502 can be launched within the audio service 522), which may be launched on one or more computer systems. In some embodiments, the audio service 522 can be a process that can be launched within a runtime environment, and instructions executed by a module (e.g., the persistence module 502) can be components of the audio service 522 (e.g., instructions executed by the persistence module 502 can be sub-processes of the audio service 522). In some embodiments, the audio service 522 can be a sub-process of a parent process. Instructions executed by a module (e.g., the persistence module 502) can include one or more components (e.g., processes, sub-processes, threads, and / or services executed by a location identification status sub-module 506, an acoustic data sub-module 508, and / or an audio model sub-module 510). In some embodiments, instructions executed by a module (e.g., the persistence module 502) can be launched as a sub-process of the audio service 522 and / or as a separate process at a location different from other components of the audio service 522. For example, instructions executed by a module (e.g., the persistence module 502) can be launched within a general-purpose processor, and one or more other components of the audio service 522 can be launched within an audio-specific processor (e.g., a DSP).In some embodiments, the instructions executed by a module (e.g., persistent module 502) may be launched within a different process address space and / or memory space than other components of the audio service 522. In some embodiments, the instructions executed by a module (e.g., persistent module 502) may be launched as one or more threads within the audio service 522. In some embodiments, the instructions executed by a module (e.g., persistent module 502) may be instantiated within the audio service 522. In some embodiments, the instructions executed by a module (e.g., persistent module 502) may share a process address and / or memory space with other components of the audio service 522.
[0067] In some embodiments, the persistence module 502 can include a localization status sub-module 506. The localization status sub-module 506 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the instructions executed by the localization status sub-module 506 can be sub-processes of the persistence module 502. In some embodiments, the localization status sub-module 506 can indicate whether localization has been achieved (e.g., the localization status sub-module 506 may indicate whether the MR system has identified the real environment and / or localized itself within the real environment). In some embodiments, the localization status sub-module 506 can interface with a localization system (e.g., via an API). The localization system may determine the location of the MR system (and / or the user using the MR system). In some embodiments, the localization system can create a 3D model of the real environment and estimate the location of the system (and / or the user) within the environment using techniques such as SLAM. In some embodiments, the localization system can rely on a passable world system (described in more detail below) and one or more sensors (e.g., of MR systems 112, 200) to estimate the location of the MR system (and / or the user) within the environment. In some embodiments, the localization status sub-module 506 can query the localization system to determine whether localization is currently achieved with respect to the MR system. Similarly, the localization system may notify the localization status sub-module 506 of the success of the localization. The localization status (e.g., of MR systems 112, 200) may be used to determine whether the audio model should be updated (e.g., because the user's real environment has changed).
[0068] In some embodiments, the persistence module 502 can include an acoustic data sub-module 508. The acoustic data sub-module 508 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the instructions executed by the acoustic data sub-module 508 can be sub-processes of the persistence module 502. In some embodiments, the acoustic data sub-module 508 can store one or more data structures representing acoustic data that can be used to create an audio model. In some embodiments, the acoustic data sub-module 508 can interface with a passable world system (e.g., via an API). The passable world system can include information about known real-world environments (e.g., rooms, buildings, and / or outdoor spaces) and associated real and / or virtual objects. In some embodiments, the passable world system can include a persistent coordinate frame and / or an anchor point. The persistent coordinate frame and / or the anchor point can be known to the MR system (e.g., by a unique identifier) and can be points fixed in space. The virtual object can be positioned relative to one or more persistent coordinate frames and / or anchor points and can enable object persistence (e.g., the virtual object can appear to stay in the same location within the real-world environment regardless of the person viewing the virtual object and regardless of any movement of the user). The persistent coordinate frame and / or the anchor point can be particularly advantageous when two or more users with separate MR systems utilize different world coordinate frames (e.g., the location of each user is designated as the origin with respect to their individual world coordinate frame). Object persistence across users can be achieved by converting between individual world coordinate frames and a universal persistent coordinate frame and installing / referencing the virtual object relative to the persistent coordinate frame.In some embodiments, the passable world system can manage and maintain persistent coordinate frames by, for example, remapping known areas, creating new persistent coordinate frames, reconciling new persistent coordinate frames with previously determined persistent coordinate frames, and / or associating persistent coordinate frames with identifiable information (e.g., locations and / or neighboring objects). In some embodiments, the acoustic data sub-module 508 may query a separate system (e.g., the passable world system) regarding one or more persistent coordinate frames. In some embodiments, the acoustic data sub-module 508 can facilitate reading, accessing, and managing associated persistent coordinate frames (e.g., persistent coordinate frames within a threshold radius of the user's position) including creation, modification, and / or deletion of associated acoustic data.
[0069] In some embodiments, the acoustic data stored within the acoustic data submodule 508 can be organized into physically correlated modular units (e.g., a room may be represented by a modular unit, and a chair within the room may be represented by another modular unit). For example, the modular units may include physical and / or perceptual relevance properties of the physical environment (e.g., a room). The physical and / or perceptual relevance properties can include properties that can affect the acoustic characteristics of the room (e.g., the dimensions and / or shape of the room). In some embodiments, the physical and / or perceptual relevance properties can include functional and / or behavioral properties, which can be interpreted by the rendering engine (e.g., whether a source outside the room should be occluded). In some embodiments, the physical and / or perceptual relevance properties can include properties of known and / or recognized objects. For example, the geometric shapes of fixed objects (e.g., floors, walls, furniture, etc.) and / or movable objects (e.g., a mug) can be stored as physical and / or perceptual relevance properties and associated with a particular environment. In some embodiments, the physical and / or perceptual relevance properties can include transmission loss, scattering coefficient, and / or absorption coefficient. In some embodiments, the modular units can include physical and / or perceptual relevance connections (e.g., between other modular units or within a modular unit). For example, the physical and / or perceptual relevance connections may describe how two or more rooms are connected and how the rooms can interact with each other (e.g., the cross-coupling gain level between digital reverberators simulating the rooms and / or the line-of-sight path between two spaces).
[0070] In some embodiments, the physical and / or perceptual related properties can include acoustic properties such as reverberation time, reverberation delay, and / or reverberation gain. The reverberation time may include the length of time required for sound to decay by a certain amount (e.g., 60 decibels). The sound decay may be the result of sound reflecting from surfaces (e.g., walls, floors, furniture, etc.) in the real environment while losing energy due to, for example, sound absorption by the boundaries of the room (e.g., walls, floors, ceilings, etc.), objects inside the room (e.g., chairs, furniture, people, etc.), and the air within the room. The reverberation time can be affected by environmental factors. For example, absorptive surfaces (e.g., cushions) can absorb sound in addition to geometric diffusion, and as a result, the reverberation time can be reduced. In some embodiments, it may not be necessary to have information about the original source in order to estimate the reverberation time of the environment. The reverberation gain can include the ratio of the direct / source / original energy of the sound to the reverberant energy of the sound (e.g., the energy of the reverberation resulting from the direct / source / original sound) where the listener and the source are substantially co-located (e.g., a user can produce a source sound that can be considered to be substantially co-located with one or more microphones mounted on a head-mounted MR system when the user claps their hands). For example, an impulse (e.g., a tap sound) can have energy associated with the impulse, and the reverberant sound from the impulse can have energy associated with the reverberation of the impulse. The ratio of the original / source energy to the reverberant energy can be the reverberation gain. The reverberation gain of a real environment can be affected by, for example, absorptive surfaces that can absorb sound and thereby reduce the reverberant energy.
[0071] In some embodiments, the acoustic data can include metadata (e.g., metadata of physical and / or perceptual related properties). For example, information about the time and / or location where the acoustic data was collected may be included in the acoustic data. In some embodiments, reliability data (e.g., the estimated measurement accuracy and / or the number of repeated measurements) associated with the acoustic data may be included as metadata. In some embodiments, the type of modular unit (e.g., modular unit related to a room or the connection between modular units) and / or data versioning may be included as metadata. In some embodiments, a unique identifier associated with the acoustic data, persistent coordinate frame, and / or anchor point may be included as metadata. In some embodiments, the relative transformation from the persistent coordinate frame and / or anchor point and the associated virtual object may be included as metadata. In some embodiments, the metadata can be stored as a single bundle together with the acoustic data.
[0072] In some embodiments, the acoustic data may be organized by a persistent coordinate frame and / or an anchor point, and the persistent coordinate frame and / or the anchor point may be organized in a map. In some embodiments, the audio model can consider acoustic data organized by a persistent coordinate frame and / or an anchor point, which may correspond to locations within the environment. In some embodiments, the acoustic data may be loaded into the acoustic data sub-module 508 in response to the success of a localization event (which may be indicated by the localization status sub-module 506). In some embodiments, all available acoustic data may be loaded into the acoustic data sub-module 508. In some embodiments, only relevant acoustic data may be loaded into the acoustic data sub-module 508 (e.g., acoustic data related to the persistent coordinate frame and / or anchor point within a certain distance of the location of the MR system).
[0073] In some embodiments, the acoustic data can include different states that can vary according to changes in the real environment. For example, the acoustic data regarding a given modular unit that can represent a room can include the acoustic data regarding the room in an empty state and the acoustic data regarding the room in an occupied state. In some embodiments, changes in the furniture arrangement of the room can be reflected by changes in the states within the acoustic data associated with the room. In some embodiments, the modular unit (e.g., representing a room) can include different acoustic data regarding the state when the door is opened and closed. The state can be represented as a binary value (e.g., 0 or 1) or a continuous value (e.g., the degree to which the door is open, the degree to which the room is occupied, etc.).
[0074] In some embodiments, the persistence module 502 can include an audio model sub-module 510. The audio model sub-module 510 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the audio model sub-module 510 can include one or more data structures representing an audio model regarding a real and / or virtual environment. In some embodiments, the audio model can be generated, at least in part, by the acoustic data stored within the acoustic data sub-module 508. The audio model can represent how sound behaves within a particular environment. For example, virtual sound generated by an MR system (e.g., MR systems 112, 200) can be modified by the audio model within the audio model sub-module 510 to reflect the acoustic characteristics of the environment. A virtual concert presented to a user sitting in a spacious concert hall can have similar acoustic properties to a real concert presented in the same concert hall. The MR system can localize itself relative to the concert hall, load the relevant acoustic data, generate an audio model, and model the acoustic properties of the concert hall.
[0075] In some embodiments, the audio model may be used to model sound propagation in an environment. For example, the propagation effects can include occlusion, obstacles, early reflections, diffraction, time-of-flight delays, Doppler effects, and other effects. In some embodiments, the audio model can take into account frequency-dependent absorption rates and / or transmission losses (e.g., based on the acoustic data loaded in the acoustic data sub-module 508). In some embodiments, the audio model stored within the audio model sub-module 510 may inform other aspects of the audio engine. For example, the audio model may use the acoustic data to synthesize audio (e.g., collisions between virtual and / or real objects) during the process.
[0076] In some embodiments, the audio rendering service 522 may include a rendering track module 514. The rendering track module 514 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the rendering track module 514 may include audio information that can be later presented to the user. In some embodiments, the MR system may present virtual sounds that include several sound sources to be mixed together (e.g., the sound source of two swords colliding and the sound source of a shouting person). The rendering track module 514 may store one or more tracks that can be mixed with other tracks for presentation to the user. In some embodiments, the rendering track module 514 can include information about spatial sources. For example, the rendering track module 514 may include information about where the sound source is located, which may be considered in the audio model and / or the rendering algorithm. In some embodiments, the rendering track module 514 may include information about modular units and / or the relationships between sound sources. For example, one or more rendering tracks and / or audio models may be associated together as a single group.
[0077] In some embodiments, the audio rendering service 522 may include a location manager module 516. The location manager module 516 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the location manager module 516 may manage location information related to the audio engine (e.g., the current location of the MR system within the physical environment). In some embodiments, the location manager module 516 can include a perception wrapper sub-module. The perception wrapper sub-module may be a wrapper around perception data (e.g., what the MR system has detected or is detecting). In some embodiments, the perception wrapper may interface and / or transform between the perception data and the location manager module 516. In some embodiments, the location manager module 516 may include a head pose sub-module, which may include head pose data. The head pose data may include the location and / or orientation of the MR system (or corresponding user) within the physical environment. In some embodiments, the head pose may be determined based on the perception data.
[0078] In some embodiments, the audio rendering service 522 may include an audio model module 518. The audio model module 518 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the audio model module 518 may include an audio model, which may be the same audio model included within the audio model sub-module 510. In some embodiments, modules 510 and 518 may maintain duplicate copies of the same audio model. For example, it may be advantageous to maintain more than one copy of the audio model when the audio model has been updated but sound is to be presented to the user. When the model is currently in use, it may be advantageous to update a copy of the model and then, when the old model becomes unavailable (e.g., is no longer in use), update the old model. In some embodiments, the audio model can be transferred between modules 510 and 518 through serialization. The audio model within module 510 can be serialized and deserialized to facilitate data transfer to module 518. Serialization can facilitate data transfer between processors (e.g., general processors and audio-specific processors) such that typed memory does not need to be shared.
[0079] In some embodiments, the audio rendering service 522 can include a rendering algorithm module 520. The rendering algorithm module 520 can include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the rendering algorithm module 520 can include an algorithm for rendering virtual sounds such that they can be presented to a user (e.g., via one or more speakers of the MR system). The rendering algorithm module 520 can take into account an audio model of a particular environment (e.g., the audio model within modules 510 and / or 518).
[0080] In some embodiments, the audio service 522 can be a process, sub-process, thread, and / or service that runs on one or more computer systems (e.g., within the MR systems 112, 200). In some embodiments, a separate system (e.g., a third-party application) may require that an audio signal be presented (e.g., via one or more speakers of the MR systems 112, 200). Such a requirement may take any suitable form. In some embodiments, the requirement for an audio signal to be presented can include software instructions for presenting the audio signal, and in some embodiments, such a requirement may be hardware-driven. The requirement may be issued with or without user involvement. Further, such a requirement may be received via local hardware (e.g., from the MR system itself), via external hardware (e.g., a separate computer system communicating with the MR system), via the Internet (e.g., via a cloud server), or via any other suitable source or combination of sources. In some embodiments, the audio service 522 may receive the requirement, render the requested audio signal (e.g., through a rendering algorithm 520 that may consider the audio model from blocks 510 and / or 518), and present the requested audio signal to the user. In some embodiments, the audio service 522 may be a process that continuously runs (e.g., in the background) while the operating system of the MR system is running. In some embodiments, the audio service 522 can be an instantiation of a parent background service that can serve as a host process for one or more background processes and / or sub-processes. In some embodiments, the audio service 522 may be part of the operating system of the MR system. In some embodiments, the audio service 522 may be accessible to an application that can be launched on the MR system.In some embodiments, a user of the MR system may not directly provide input to the audio service 522. For example, the user may provide the input (e.g., movement command) to an application (e.g., a role-playing game) launched on the MR system. The application may provide the input to the audio service 522 (e.g., for rendering footsteps), and the audio service 522 may provide the output (e.g., rendered footsteps) to the user (e.g., via a speaker) and / or other processes and / or services.
[0081] FIG. 6 illustrates an exemplary process for updating an audio model according to some embodiments. At step 606, localization may be determined (e.g., the MR system may be able to properly identify its location within the environment). At step 607, which may occur within the persistence module 602 (which may correspond to the persistence module 502), a notification of successful localization may be issued. In some embodiments, the notification of successful localization may trigger a process to update the audio model (e.g., because the previous audio model may no longer be applicable to the current location).
[0082] In step 608, it can be determined whether to start the call of the acoustic data. It may be desirable to set one or more conditions for starting the call so that the audio model is not updated too frequently. For example, if a user using the MR system moves only slightly within the room, it may not be desirable to update the audio model (e.g., because the updated model may not be perceptually distinguishable from the existing model and / or continuously updating the model may be computationally expensive). In some embodiments, the threshold condition in step 608 may be time-based. For example, the call may be started only if the call has not already been started for 5 seconds. In some embodiments, the threshold condition in step 608 may be location-based. For example, the call may be started only if the user changes the threshold distance amount position. Note that other threshold conditions may be used as well. In some embodiments, steps 607 and / or 608 may occur within a location status sub-module (e.g., location status sub-module 506).
[0083] If it is determined that the call should be started, a persistent coordinate frame may be read in step 610. In some embodiments, only a subset of the available persistent coordinate frames may be read in step 610. For example, only the persistent coordinate frames near the location may be read.
[0084] In step 612, the acoustic data may be read. In some embodiments, the acoustic data read in step 612 may correspond to the acoustic data stored in the acoustic data sub-module 508. In some embodiments, only a subset of the available acoustic data may be read in step 612. For example, the acoustic data associated with one or more persistent coordinate frames and / or anchor points may be read. In some embodiments, step 612 and / or 614 may occur within the acoustic data sub-module (e.g., acoustic data sub-module 508).
[0085] In step 614, the audio model may be constructed and / or modified. In some embodiments, the audio model may consider the acoustic data read in step 612, and the audio model may model the acoustic characteristics of a particular environment. In some embodiments, step 614 may occur within the audio model sub-module (e.g., audio model sub-module 510).
[0086] In step 616, it can be determined whether a copy of the audio model should be updated. To avoid interfering with the service (e.g., presenting audio to the user), it may be desirable to set one or more conditions for updating the copy of the audio model. For example, one condition may be to evaluate whether a copy of the audio model exists (e.g., within the audio rendering service 604 but outside the persistence module 602). If the copy of the audio model does not exist, the audio model may be read by the audio rendering service 604 (which may correspond to the audio rendering service 522). In some embodiments, the condition may be to evaluate whether the copy of the audio model is currently in use (e.g., whether the audio model is being used to render and present audio to the user). If the copy of the audio model is not in use, the updated audio model (e.g., the audio model generated in step 614) may be propagated to the copy.
[0087] In step 618, the audio rendering service 604 may read a copy of the audio model (e.g., the audio model generated in step 614). The audio model may be transferred using data serialization and / or deserialization.
[0088] In step 620, the obsolete audio model may optionally be deleted and / or invalidated. For example, the audio rendering service 604 may have a first existing audio model that was previously used. The audio rendering service 604 may read a second updated audio model (e.g., from the persistence module 602) and delete and / or invalidate the first existing audio model.
[0089] In step 622, a notification may be issued regarding the new audio model. In some embodiments, the notification may include a callback function to a client (e.g., a third-party application) that can be subscribed to receive the notification when the audio model changes.
[0090] FIG. 7 illustrates an exemplary process for updating an audio model according to some embodiments. In step 706, audio data may be received (e.g., via one or more sensors of MR systems 112, 200). In some embodiments, the audio data may be manually entered (e.g., a user and / or developer may manually enter reverberation time, reverberation delay, reverberation gain, etc.). In step 708, which may occur within persistence module 702 (which may correspond to persistence module 502), the associated environment may be identified. The associated environment may be identified by metadata that may be associated with the audio data (e.g., the metadata may carry information about one or more persistent coordinate frames and / or anchor points that may be known to the MR system).
[0091] In step 710, it can be determined whether the associated environment is new. For example, if the associated environment cannot be identified and / or is associated with an unknown identifier, it may be determined that the audio data is associated with a new environment. If it is determined that the associated environment is not new, a copy of the official audio model may be updated (e.g., with room properties that may be derived from the audio data). If it is determined that the associated environment is new, the new environment may be added to a copy of the official audio model, and the copy audio model may be updated as appropriate. In some embodiments, the new environment may be represented by a new modular unit.
[0092] In step 716, metadata associated with the new environment (e.g., metadata associated with the new modular unit) may be initialized. For example, metadata associated with measurement numbers, reliability, or other information may be created and bundled with the new modular unit.
[0093] In step 718, the official audio model within the persistent module 702 may be updated. For example, the official audio model may be replicated from the updated copy audio model. In some embodiments, the official audio model may be locked in step 718 to prevent further changes from being made to the official audio model. In some embodiments, the changes may still be made to a copy of the official audio model (which may still exist within the persistent module 702) while the official audio model is locked.
[0094] In step 720, acoustic data associated with the new audio data may be saved. For example, the new modular unit associated with the new room may be saved and / or passed through the passable world system (and may be made accessible to the MR system in the future as needed). In some embodiments, step 720 may occur within the acoustic data submodule 508. In some embodiments, step 720 may occur sequentially following step 718. In some embodiments, step 720 may occur at an independent time as determined by other components, e.g., based on the availability of the passable world system in which the acoustic data can be stored.
[0095] In step 722, an audio rendering service 704 (which may correspond to the audio rendering service 522) may read a copy of the audio model. In some embodiments, the copy of the audio model may be the same audio model updated in step 718. In some embodiments, data transfer may occur through serialization and deserialization. In some embodiments, the audio model may be locked while the serialization process is being executed, which may prevent the audio model from being changed while a snapshot of the audio model is being created.
[0096] In step 724, the copy of the outdated audio model may be deleted and / or invalidated.
[0097] In step 726, a notification may be issued regarding the new audio model. The notification can be a callback function to clients subscribed to receive notifications when the model is updated.
[0098] In step 728, the official audio model may be released, which may indicate that the serialized bundle corresponding to the audio model can be deleted. In some embodiments, it may be desirable to lock the official audio model in the persistence module 702 while a copy of the official audio model is being transferred to the audio rendering service 704. Once the audio rendering service 704 finishes reading the copy of the official audio model, it may be desirable to unlock the official audio model so that the official audio model can continue to be updated.
[0099] In some embodiments, an audio rendering service (e.g., audio rendering service 522) may not manage the interaction between the persistence module (e.g., persistence module 502) and the rendering algorithm (e.g., rendering algorithm module 520). For example, the rendering algorithm 520 may directly communicate with the persistence module 502 and read the updated audio model. In some embodiments, the rendering algorithm 520 may include its own copy of the audio model. In some embodiments, the rendering algorithm 520 may access the audio model within the persistence module 502.
[0100] Multi-application audio rendering
[0101] An MR system (e.g., MR systems 112, 200) can utilize various on-board sensors to develop a customized audio model for a user's MRE. The MR system can consider the actual features (e.g., floors, walls, people) within the user's MRE and virtual objects (e.g., virtual benches) within the user's MRE and generate a customized audio model that can appropriately reflect the actual and / or virtual physics of the user's MRE. For example, a virtual sound source may be located within the user's MRE. The audio model can identify the direct path from the virtual sound source to the user and determine whether the virtual sound should be occluded by any actual or virtual objects within the direct path. In some embodiments, the audio model can determine how sound from the virtual sound source can reflect off actual and / or virtual objects within the user's MRE (e.g., a hard surface can reflect more sound, a certain surface can reflect more high-frequency sound, a rough surface can scatter sound in multiple directions, etc.). In some embodiments, the audio model can determine how sound from the virtual sound source can reverberate within the user's MRE. In some embodiments, it may be desirable to enable customization of the model such that it can be "accurate" with respect to the actual and / or virtual physics of the user's MRE. For example, an MR application can start with the user within its current physical space and slowly transform the user's environment into an exotic environment. As the MR application transforms the user's environment (e.g., by adding virtual trees, virtual tree leaves, etc.), the MR application can adapt the "official" audio model (e.g., by removing physical walls and ceilings from the official audio model) to suit the desired user experience.
[0102] However, problems can occur when more than one MR application is launched concurrently on the MR system. In some embodiments, each application may desire to present sound to the user using a different audio model. For example, the MR system may launch a web browsing application and a podcast application concurrently. In some embodiments, the web browser may be playing an informational video, and the web browser may use the official audio model generated by the MR system. Since the web browser can use the official audio model, the informational video may sound as if the speaker is placed within the user's MRE (e.g., the speaker's voice can be appropriately occluded, reflected, and reverberated based on real and / or virtual objects within the user's MRE). In some embodiments, the podcast application may be playing a podcast placed within a cave. Thus, it may be desirable for the podcast application to present sound as highly reflective and / or reverberant.
[0103] However, if a web browsing application and a podcast application attempt to present audio using different audio models simultaneously, conflicts can occur. In some embodiments, the sound tracks from each application may be rendered using separate audio models separately and then later mixed together to present the sound (e.g., an audio signal comprising one or more rendered sounds) to the user. However, a completely separate rendering pipeline can be computationally expensive and, in some cases, can double the computational resources required to render a sound track from a single application. The computational resources can become even more distorted as more applications are launched in parallel, and computationally, for example, rendering sound tracks from 10 different simultaneously launched applications that utilize separate customized audio models may not be feasible. Thus, it may be desirable to design systems and methods to adapt to simultaneously launched MR applications that can utilize different audio models.
[0104] FIG. 8 illustrates a model management architecture according to some embodiments. In some embodiments, an MR system 803 (which may correspond to MR systems 112, 200) can include one or more computer systems configured to execute instructions. The MR system 803 can include an audio service module 802 (which may correspond to audio service 522). The audio service module 802 may manage audio rendering and can include one or more computer systems configured to execute one or more processes, sub-processes, threads, and / or services. In some embodiments, the audio service module 802 can be configured to execute a service (e.g., a background service) as part of an operating system of one or more computer systems.
[0105] In some embodiments, a separate system (e.g., a third-party application) may require that an audio signal be presented (e.g., via one or more speakers of MR systems 112, 200). Such a requirement may take any suitable form. In some embodiments, the requirement that an audio signal be presented can include one or more software instructions for presenting the audio signal. In some embodiments, such a requirement may be hardware-driven. The requirement may be issued with or without user involvement. Further, such a requirement may be received via local hardware (e.g., the MR system itself), via external hardware (e.g., a separate computer system communicating with the MR system), via the Internet (e.g., via a cloud server), or via any other suitable source or combination of sources. In some embodiments, the audio service 802 may receive the requirement, render the requested audio signal (e.g., through one or more rendering algorithms that may utilize an audio model managed by the model layer 804), and present the requested audio signal to the user. In some embodiments, the audio service 802 may be configured to continuously start (e.g., in the background) while the operating system of the MR system is running and execute a process. In some embodiments, the audio service 802 may be configured to execute an instantiation of a parent background service, which may serve as a host process for one or more background processes and / or sub-processes. In some embodiments, the audio service 802 may be configured to execute instructions as part of the operating system of the MR system. In some embodiments, the audio service 802 may be accessible to an application that can be launched on the MR system. In some embodiments, the user of the MR system may not be required to directly provide input to the audio service 802.For example, a user may provide an input (e.g., a movement command) to an application (e.g., a role-playing game) that is launched on the MR system. The application may provide the input to an audio service 802 (e.g., for rendering the sound of footsteps), and the audio service 802 may provide an output (e.g., an audio signal comprising the rendered footsteps) to the user (e.g., via a speaker) and / or other processes and / or services.
[0106] In some embodiments, the audio service 802 can include a model layer 804, which may be configured to manage one or more audio models. In some embodiments, the model layer 804 can include one or more computer systems configured to execute one or more processes, sub-processes, threads, and / or services. In some embodiments, the model layer 804 can include one or more computer systems configured to store information. For example, the model layer 804 may store one or more audio models 806a, 806b, and / or 806c. In some embodiments, an audio model can include one or more abstract representations of data that are used to render audio. In some embodiments, an audio model can include one or more data structures configured to store information. In some embodiments, an audio model (e.g., audio model 806a) can include one or more audio model components (e.g., components 808a and 808b). In some embodiments, an audio model component can include one or more data structures configured to store information.
[0107] In some embodiments, the audio model component can store one or more abstract representations of data used to render audio. For example, the audio model component can include data regarding one or more real and / or virtual objects. In some embodiments, the audio model component can include position data regarding one or more real and / or virtual objects. In some embodiments, the audio model component can include material data (e.g., audio reflectivity, transmittance, absorption, scattering, diffraction, etc.) on one or more real and / or virtual objects. In some embodiments, one audio model component can represent one real and / or virtual object and its associated properties.
[0108] In some embodiments, the audio model can include data regarding one or more virtual sound sources. For example, the audio model can include position data regarding one or more virtual sound sources. In some embodiments, one audio model component can represent one virtual sound source and its associated properties.
[0109] In some embodiments, the audio model can include audio parameter data (e.g., sound emission properties, volume, etc.) regarding one or more virtual sound sources. In some embodiments, the audio model can include one or more audio parameters of the environment. For example, the audio model may include the dimensions of the environment (e.g., the room in which the user is located). In some embodiments, the audio model may include the reverberation time and / or reverberation gain of the environment. In some embodiments, one audio model component can represent one audio parameter of the environment.
[0110] In some embodiments, the model layer 804 can be configured to manage different audio models. For example, model 806a can be considered an official model (e.g., model 806a can be configured to represent realistic physical interactions between virtual sounds and real and / or virtual objects in the user's environment). In some embodiments, model 806a can include a model component 808a that represents a virtual wall (which can reside within the user's environment). In some embodiments, model 806a can include a model component 808b that represents the reverberation time (e.g., of the user's environment). In some embodiments, model 806a may be used by application 814a, which can be configured to be launched on MR system 803. In some embodiments, application 814a can be an MR application that causes the requested virtual audio to be presented to the user. In some embodiments, the association between application 814a and model 806a can be stored within model layer 804 (e.g., via metadata associated with model 806a). In some embodiments, application 814a can store data (e.g., a model identifier), and application 814a may define that a particular audio model is to be used with the data.
[0111] In some embodiments, the model layer 804 can be configured to manage a model 806b. In some embodiments, the model 806b may be different from the model 806a. For example, the model 806b may include a model component 808c, which may correspond to and / or be identical to the model component 808a (e.g., both model components 808a and 808c may represent the same virtual wall). However, the model 806b may include a model component 808d, which may not correspond to and / or not be identical to the model component 808b. For example, the model component 806d may represent a second response time, which may be longer than the response time corresponding to the model component 806b.
[0112] In some embodiments, the model 806b can be a full audio model. For example, the model 806b may be used to render audio without dependencies on other models (e.g., the model 806a). In some embodiments, the model 806b can include one or more dependencies on other models. For example, the model 806b may include one or more pointers to the model component 808a within the model 806a (e.g., the model component 808c may correspond to the model component 808a). In some embodiments, storing an audio model with dependencies may require less memory and / or storage on a memory device than storing a completely independent audio model.
[0113] In some embodiments, model 806c may share components with models 806a and / or 806b. For example, model component 808e may represent a virtual sofa, model component 808f may represent a virtual whiteboard, and model component 808g may represent a third reverberation time, which may be shorter than the second and first reverberation times respectively associated with model components 808b and 808d. In some embodiments, model 806c may be associated with application 814b, which may be configured to be launched on MR system 803.
[0114] FIG. 8 depicts model layer 804 to manage one or more audio models as part of audio service 802, although other embodiments are contemplated as well. In some embodiments, each application may manage its own one or more audio models. For example, application 814a may locally store model 806a within application 814a. In some embodiments, model 806a may be an official model, and application 814a may store a reference to model 806a within model layer 804. In some embodiments, application 814b may locally store model 806c, and model 806c may not be stored within model layer 804. In some embodiments, model layer 804 may generate an audio model. For example, a user of the MR system may change the physical environment, and the new audio model may more accurately represent the acoustic characteristics of the new environment. In some embodiments, an application (e.g., application 814a) may require that a new audio model be generated because the application may utilize one or more customizations and / or because the application may utilize a fully custom audio model.
[0115] In some embodiments, the model layer 804 may store an audio model associated with an application currently running on the MR system 803. For example, the MR system 803 may currently be running only applications 814a and 814b, and the model 806b may not be associated with application 814a or 814b. In some embodiments, the model 806b may be removed from the model layer 804 (e.g., to save memory). In some embodiments, the official model may be continuously stored within the model layer 804. In some embodiments, the model layer 804 may remove the audio model from memory if it has not been used for a threshold amount of time. In some embodiments, an application may request that the model layer 804 retain the audio model in memory.
[0116] In some embodiments, the audio service 802 may include a rendering layer 812, which may be configured to manage one or more audio models. In some embodiments, the rendering layer 812 (which may correspond to the rendering algorithm module 520) may include one or more computer systems configured to execute one or more processes, sub-processes, threads, and / or services. The rendering layer 812 may render audio from one or more inputs (e.g., model components, sound source information, etc.). In some embodiments, the rendering layer 812 may render a stream of audio information, and the audio stream may be rendered in real time. In some embodiments, the rendering layer 812 may render audio from a fixed amount of input data. The rendered audio may be presented to the user as one or more audio signals comprising the rendered audio. The audio signal may be presented via one or more speakers of the device (e.g., speakers 412 and 414 shown in FIG. 4).
[0117] FIG. 8 depicts three models with variable model components, although it is contemplated that any number of models may be managed by the audio service 802. It is also contemplated that any model may include any number of model components.
[0118] FIG. 9A illustrates a rendering architecture according to some embodiments. In some embodiments, the rendering layer 902 (which may correspond to the rendering layer 812) can include one or more computer systems configured to execute one or more processes, sub-processes, threads, and / or services. In some embodiments, the rendering layer 902 can be configured to render audio from one or more inputs (e.g., model components, sound source information, etc.). In some embodiments, the rendering layer 902 may render a stream of audio information, and the audio stream may be rendered in real time. In some embodiments, the rendering layer 902 may render audio from a fixed amount of input data.
[0119] The rendering layer 902 may include a direct layer 904. In some embodiments, the direct layer 904 can include one or more computer systems configured to execute one or more processes, sub-processes, threads, and / or services. In some embodiments, the direct layer 904 can be configured to render a direct path between a sound source and a listening node (e.g., a user). The direct layer 904 may determine whether sound should be occluded (e.g., because physical and / or virtual objects block the direct path from the sound source to the listening node). In some embodiments, the direct layer 904 may determine whether sound should be attenuated (e.g., because of the distance between the sound source and the listening node).
[0120] The rendering layer 902 may include a reflection layer 906. In some embodiments, the reflection layer 906 can include one or more computer systems configured to execute processes, sub - processes, threads, and / or services. In some embodiments, the reflection layer 906 can be configured to render one or more reflected audio paths between a sound source and a listening node (e.g., a user). For example, sound from a sound source can reflect off a wall before reaching the listening node, and the reflection layer 906 may render one or more effects of the reflection (e.g., the reflected sound can be delayed compared to the direct sound, the reflected sound can be attenuated compared to the direct sound). In some embodiments, the reflection layer 906 can be configured to apply one or more filters to the audio to approximate the behavior of reflected sound in a general environment. In some embodiments, the reflection layer 906 can be configured for audio ray tracing, which may calculate reflection paths for one or more audio rays.
[0121] The rendering layer 902 may include a reverberation layer 908. In some embodiments, the reverberation layer 908 may include one or more computer systems configured to execute processes, sub - processes, threads, and / or services. In some embodiments, the reverberation layer 908 can be configured to render the reverberation behavior of audio (e.g., delayed reverberation behavior). In some embodiments, the delayed reverberation behavior may be affected and / or determined by environmental properties such as the reverberation time and / or reverberation gain of the environment.
[0122] The rendering layer 902 may include a visualizer 910. In some embodiments, the visualizer 910 may include one or more computer systems configured to execute processes, sub-processes, threads, and / or services. In some embodiments, the visualizer 910 may be configured to render audio to one or more virtual speakers. For example, an MR system may present audio to a user such that it originates from one of six virtual speakers arranged within a speaker array around the user's head. In some embodiments, the visualizer 910 may be configured to determine which sound or combination of sounds should be reproduced by which virtual speaker or combination of speakers within a virtual speaker array.
[0123] In the exemplary embodiment illustrated in FIG. 9A, audio stream 912 and audio stream 914 may be rendered by rendering layer 902. In some embodiments, audio stream 912 may correspond to an audio stream originating from a first application, and audio stream 914 may correspond to an audio stream originating from a second application. In some embodiments, the first application may use the same audio model as the second application, and the first and second applications may be launched in parallel on the MR system. In some embodiments, rendering both audio stream 912 and audio stream 914 may be more efficient than rendering the audio streams separately (e.g., because both audio streams may rely on the same audio model). For example, the direct paths corresponding to audio stream 912 and audio stream 914 may be rendered and passed through visualizer 910. In some embodiments, audio stream 912 and audio stream 914 may be rendered using the same direct path calculation (e.g., because the same real and / or virtual objects may sometimes block the direct path between one or more sound sources and one or more listening nodes).
[0124] In some embodiments, audio stream 912 and audio stream 914 can be rendered together through reflection layer 906. In some embodiments, reflection layer 906 can receive one or more input audio streams directly from layer 904. In some embodiments, reflection layer 906 can receive one or more input audio streams directly from the application (e.g., without first passing through layer 904 directly). In some embodiments, rendering audio stream 912 and audio stream 914 together can be more efficient than rendering the audio streams separately. For example, a single set of filters can be applied to both audio streams (and / or a single audio stream corresponding to a mix of the two audio streams) to approximate reflective behavior. In some embodiments, the rendered reflection can be passed to visualizer 910.
[0125] In some embodiments, audio stream 912 and audio stream 914 can be rendered together through reverberation layer 908. In some embodiments, reverberation layer 908 can receive one or more input audio streams from reflection layer 906 and / or directly from layer 904. In some embodiments, reverberation layer 908 can receive one or more input audio streams directly from the application (e.g., without passing through reflection layer 906 and / or directly through layer 904). In some embodiments, rendering audio stream 912 and audio stream 914 together can be more efficient than rendering the audio streams separately. For example, a delayed reverberation behavior can be determined according to a shared audio model, and the delayed reverberation behavior can be applied to both audio streams. In some embodiments, the rendered reverberation behavior can be passed to visualizer 910.
[0126] Figure 9B illustrates a rendering architecture according to some embodiments. In some embodiments, multiple audio streams, relying on different audio models, may be rendered together. For example, audio stream 912 may originate from a first application using a first audio model, and audio stream 914 may originate from a second application using a second audio model. In some embodiments, the first audio model may not be present within the second audio model (e.g., the first application introduces virtual objects that the second application cannot utilize), and may include model components corresponding to the virtual objects. In some embodiments, the virtual objects introduced by the first application may affect the direct rendering path for audio stream 912 (e.g., direct sound may be occluded as a result of obstacles), but may not affect the direct rendering path for audio stream 914. In some embodiments, audio stream 912 and audio stream 914 may be rendered in separate instances in the direct layer 904. In some embodiments, audio stream 912 and audio stream 914 may each be rendered independently in the direct layer 904.
[0127] In some embodiments, other model components may also be shared between audio stream 912 and audio stream 914. For example, virtual objects may be configured not to interfere with audio reflections, and audio stream 912 and audio stream 914 may be rendered using the same reflection calculations at the reflection layer 906. In some embodiments, virtual objects may not affect the late reverberation of the MRE, and audio streams 912 and 914 may be rendered using the same reverberation calculations at the reverberation layer 908. In some embodiments, both audio streams 912 and 914 may be mixed together in the visualizer 910.
[0128] In some embodiments, efficiency can be leveraged when audio streams 912 and 914 rely on shared model components (even if there may be one or more differences in the model components). For example, reflection, reverberation, and visualizer calculations may be shared across the audio streams, even when the direct path may require independent calculations across the two audio streams.
[0129] FIG. 9B illustrates a rendering architecture with different direct path instances, although other embodiments are also contemplated. For example, audio stream 912 may correspond to a first application that changes the virtual object material composition to more reflective audio (compared to audio stream 914, which may correspond to a second application that does not modify the virtual object material composition). In some embodiments, audio stream 912 and audio stream 914 may both be rendered in direct layer 904, but may be rendered separately (e.g., in separate instances) in reflection layer 906. In some embodiments, audio stream 912 and audio stream 914 can again both be rendered in reverberation layer 908 (e.g., because a model component corresponding to delayed reverberation is shared between the audio models relied upon by audio streams 912 and 914).
[0130] FIGS. 9A-9B illustrate a rendering architecture that includes a direct layer 904, a reflection layer 906, a reverberation layer 908, and a visualizer 910, although other architectures may equally well be used. In some embodiments, one or more additional layers may be included in the rendering architecture (e.g., due to computational limitations). In some embodiments, one or more additional layers may be added to more realistically model acoustic behavior.
[0131] Exemplary systems, methods, and computer-readable media are disclosed. According to some embodiments, a system includes one or more speakers and a processor configured to execute a method including receiving a request to present a first audio track, the first audio track being based on a first audio model including a shared model component and a first model component; receiving a request to present a second audio track, the second audio track being based on a second audio model including a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and presenting an audio signal comprising the rendered sound via the one or more speakers. In some embodiments, the step of rendering sound includes rendering the first audio track and the second audio track based on the shared model component. In some embodiments, the step of rendering sound includes rendering the first audio track based on the first audio component. In some embodiments, the step of rendering sound includes rendering the second audio track based on the second audio component. In some embodiments, the one or more speakers are speakers of a head-mounted device. In some embodiments, the first audio model is based on the physical environment of the head-mounted device. In some embodiments, the shared model component corresponds to direct-path acoustic behavior. In some embodiments, the shared model component corresponds to reflective acoustic behavior. In some embodiments, the shared model component corresponds to reverberant acoustic behavior.
[0132] According to some embodiments, the method comprises receiving a request to present a first audio track, the first audio track being based on a first audio model comprising a shared model component and a first model component; receiving a request to present a second audio track, the second audio track being based on a second audio model comprising a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and presenting an audio signal comprising the rendered sound via one or more speakers. In some embodiments, the step of rendering sound comprises rendering the first audio track and the second audio track based on the shared model component. In some embodiments, the step of rendering sound comprises rendering the first audio track based on the first audio component. In some embodiments, the step of rendering sound comprises rendering the second audio track based on the second audio component. In some embodiments, the step of presenting the audio signal comprises presenting the audio signal via one or more speakers of a head-mounted device. In some embodiments, the first audio model is based on the physical environment of the head-mounted device. In some embodiments, the shared model component corresponds to direct-path acoustic behavior. In some embodiments, the shared model component corresponds to reflective acoustic behavior. In some embodiments, the shared model component corresponds to reverberant acoustic behavior.
[0133] According to some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform steps including receiving a request to present a first audio track, where the first audio track is based on a first audio model comprising a shared model component and a first model component; receiving a request to present a second audio track, where the second audio track is based on a second audio model comprising a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and presenting an audio signal comprising the rendered sound via one or more speakers. In some embodiments, the step of rendering sound includes rendering the first audio track and the second audio track based on the shared model component. In some embodiments, the step of rendering sound includes rendering the first audio track based on the first audio component. In some embodiments, the step of rendering sound includes rendering the second audio track based on the second audio component. In some embodiments, the step of presenting the audio signal includes presenting the audio signal via one or more speakers of a head-mounted device. In some embodiments, the first audio model is based on the physical environment of the head-mounted device. In some embodiments, the shared model component corresponds to direct-path acoustic behavior. In some embodiments, the shared model component corresponds to reflective acoustic behavior. In some embodiments, the shared model component corresponds to reverberant acoustic behavior.
[0134] The disclosed embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. For example, one or more elements of an implementation may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications should be understood to be included within the scope of the disclosed embodiments as defined by the appended claims.
Claims
【Claim 1】 The invention described in this specification.