Multi-application audio rendering

CN115398936BActive Publication Date: 2026-09-22MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180027939.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-14
Filing Date
2021-02-11
Publication Date
2026-09-22
Estimated Expiration
2041-02-11

AI Technical Summary

Technical Problem

[0006]现有技术往往达不到这些期望,诸如通过呈现不考虑用户周围环境或不与虚拟对象的空间运动相对应的虚拟音频,导致可能损害用户体验的不真实的感觉

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115398936B_ABST
    Figure CN115398936B_ABST
Patent Text Reader

Abstract

Systems and methods for efficiently rendering audio are disclosed herein. One method can include receiving a request to render a first audio track, wherein the first audio track is based on a first audio model that includes a shared model component and a first model component; receiving a request to render a second audio track, wherein the second audio track is based on a second audio model that includes the shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and presenting, via one or more speakers, an audio signal that includes the rendered sound.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 976,569, filed February 14, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to systems and methods for managing and storing audio data, and in particular, to systems and methods for managing and storing audio data in mixed reality environments. Background Technology

[0004] Virtual environments are ubiquitous in computing environments and are used in video games (where a virtual environment can represent a game world); maps (where a virtual environment can represent terrain to be navigated); simulations (where a virtual environment can simulate a real environment); digital storytelling (where virtual characters can interact with each other in a virtual environment); and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the user experience with virtual environments can be limited by the technologies used to render them. For example, conventional displays (e.g., 2D displays) and audio systems (e.g., fixed speakers) may not be able to render virtual environments in a way that produces a convincing, realistic, and immersive experience.

[0005] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share sensory information presented to users of XR systems that corresponds to a virtual environment represented by data in a computer system. Such systems can provide a uniquely enhanced sense of immersion and realism by combining virtual visual and audio cues with real-world sight and sound. Therefore, it may be desirable to present digital sound to users of XR systems in such a way that the sound appears natural and consistent with the user's expectations of sound in their real-world environment. Generally, users expect virtual sound to reproduce the acoustic characteristics of the environment in which they hear it. For example, users of an XR system in a large concert hall would expect the virtual sound to have a large, spongy quality; conversely, users in a small apartment would expect the sound to be more attenuated, closer, and more immediate. In addition to matching virtual sound to the acoustic characteristics of real and / or virtual environments, the sense of realism is further enhanced by spatializing virtual sound. For example, a virtual object might visually fly past the user from behind, and the user might expect the corresponding virtual sound to similarly reflect the virtual object's spatial movement relative to the user.

[0006] Existing technologies often fall short of these expectations, such as by presenting virtual audio that doesn't consider the user's surroundings or correspond to the spatial movement of virtual objects, leading to an unrealistic feeling that can harm the user experience. Observations on XR system users suggest that while users may be relatively tolerant of visual mismatches between virtual content and the real environment (e.g., inconsistent lighting), they may be more sensitive to auditory mismatches. Our own auditory experience, refined throughout our lives, makes us acutely aware of how our physical environment affects the sounds we hear; we can be highly perceptive of sounds that are inconsistent with expectations. With XR systems, such inconsistencies can be jarring and can turn an immersive and engaging experience into a fancy imitation. In extreme cases, auditory inconsistencies can lead to motion sickness and other adverse effects because the inner ear cannot coordinate auditory stimuli with their corresponding visual cues.

[0007] A system architecture is needed to organize and manage the systems used to generate virtual audio. Generating virtual audio may involve managing and storing information about the user's environment so that this information can be used to generate realistic virtual sounds. Therefore, the audio system architecture may need to interface with other systems to receive and utilize information related to the audio engine. An audio system architecture that can deliver realistic sound without interruption during use may also be desired. An audio system architecture that can update the audio engine without interrupting the user can produce an immersive user experience, where auditory signals continuously reflect the user's environment.

[0008] By considering the characteristics of the user's physical environment, the systems and methods described in this paper can simulate what a user would hear if the virtual sound were a real sound naturally generated in that environment. By presenting virtual sound in a way that faithfully reflects how sound behaves in the real world, users can experience a high degree of engagement with the mixed reality environment. Similarly, by presenting location-aware virtual content that responds to user movement and the environment, the content becomes more subjective, interactive, and realistic—for example, a user's experience at point A may be completely different from their experience at point B. This enhanced realism and interactivity can provide the foundation for new applications of mixed reality, such as using spatially aware audio to enable novel forms of gameplay, social features, or interactive behaviors. Summary of the Invention

[0009] Examples of this disclosure describe systems and methods for efficiently rendering audio. According to examples of this disclosure, a method may include: receiving a request to render a first audio track, wherein the first audio track is based on a first audio model including a shared model component and a second model component; receiving a request to render a second audio track, wherein the second audio track is based on a second audio model including the shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and rendering an audio signal including the rendered sound via one or more speakers. Attached Figure Description

[0010] Figures 1A to 1C An example mixed reality environment is shown according to some embodiments.

[0011] Figures 2A to 2D Components of an example mixed reality environment, which can be used to generate and interact with a mixed reality environment according to some embodiments, are shown.

[0012] Figure 3A An example mixed reality handheld controller, according to some embodiments, is shown that can be used to provide input to a mixed reality environment.

[0013] Figure 3B An example auxiliary unit, which can be used with an example mixed reality system according to some embodiments, is shown.

[0014] Figure 4 An example functional block diagram for an example mixed reality system is shown according to some embodiments.

[0015] Figure 5 An example of a virtual audio system according to some embodiments is shown.

[0016] Figure 6 An example process for updating an audio model is shown according to some embodiments.

[0017] Figure 7 An example process for updating an audio model is shown according to some embodiments.

[0018] Figure 8 An exemplary model management architecture according to some embodiments is shown.

[0019] Figures 9A to 9B An exemplary rendering architecture according to some embodiments is shown. Detailed Implementation

[0020] In the following description of the examples, reference is made to the accompanying drawings, which form part of them and illustrate, by way of illustration, examples that can be practiced. It should be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.

[0021] Mixed Reality Environment

[0022] Like everyone else, users of mixed reality systems exist within a real-world environment—that is, the three-dimensional portion of the "real world" and all of its content are perceptible to the user. For example, users perceive the real world using their ordinary human senses—sight, sound, touch, taste, smell—and interact with the real environment by moving their bodies within it. A location within the real environment can be described as coordinates in a coordinate space; for example, coordinates may include latitude, longitude, and altitude relative to sea level; distance from a reference point in three orthogonal dimensions; or other suitable values. Similarly, vectors can describe quantities having direction and magnitude in coordinate space.

[0023] A computing device may maintain a representation of a virtual environment, for example, in memory associated with the device. As used herein, a virtual environment is a computational representation of a three-dimensional space. A virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other properties associated with that space. In some examples, the circuitry of the computing device (e.g., a processor) may maintain and update the state of the virtual environment; that is, the processor may determine the state of the virtual environment at a second time t1 at a first time t0 based on data associated with the virtual environment and / or input provided by the user. For example, if an object in the virtual environment is located at a first coordinate at time t0 and has some programmed physical parameter (e.g., mass, coefficient of friction); and input received from the user indicates that a force should be applied to the object in a direction vector; the processor may apply laws of kinematics to determine the position of the object at time t1 using fundamental mechanics. The processor may use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at time t1. When maintaining and updating the state of the virtual environment, the processor may execute any suitable software, including software related to the creation and deletion of virtual objects in the virtual environment; software for defining the behavior of virtual objects or characters in the virtual environment (e.g., scripts); software for defining the behavior of signals (e.g., audio signals) in the virtual environment; software for creating and updating parameters associated with the virtual environment; software for generating audio signals in the virtual environment; software for processing inputs and outputs; software for implementing network operations; software for applying asset data (e.g., animation data of virtual objects moving over time); or many other possibilities.

[0024] Output devices (such as displays or speakers) can present any or all aspects of a virtual environment to a user. For example, a virtual environment may include representations of virtual objects (which may include inanimate objects; people; animals; light; etc.). A processor can determine a view of the virtual environment (e.g., corresponding to a “camera” with a coordinate origin, viewing axis, and viewing cone); and render a visual scene of the virtual environment corresponding to that view to the display. Any suitable rendering technique can be used for this purpose. In some examples, the visual scene may include only some virtual objects in the virtual environment and exclude certain other virtual objects. Similarly, a virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects in the virtual environment may generate sounds originating from the object's location coordinates (e.g., a virtual character can speak or cause sound effects); or the virtual environment may be associated with musical cues or ambient sounds that may or may not be associated with a specific location. The processor can determine the audio signal corresponding to the “listener” coordinates—for example, the audio signal corresponding to a composite of sounds in a virtual environment, and mix and process it to simulate the audio signal that will be heard by the listener at the listener coordinates—and present the audio signal to the user via one or more speakers.

[0025] Because the virtual environment exists only as a computational structure, users cannot directly perceive it using their ordinary senses. Instead, users can perceive the virtual environment indirectly, such as through displays, speakers, haptic output devices, etc. Similarly, users cannot directly touch, manipulate, or otherwise interact with the virtual environment; however, they can provide input data to a processor that can update the virtual environment using input devices or sensors. For example, a camera sensor can provide optical data indicating that a user is attempting to move an object in the virtual environment, and the processor can use this data to make the object react accordingly in the virtual environment.

[0026] Mixed reality systems can present a mixed reality environment (“MRE”) to a user, combining aspects of a real and a virtual environment, for example, using a transmissive display and / or one or more speakers (which may be incorporated, for example, into a wearable head-mounted device). In some embodiments, one or more speakers may be external to the head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of a real environment and a corresponding virtual environment. In some examples, the corresponding real and virtual environments share a single coordinate space; in some examples, the real coordinate space and the corresponding virtual coordinate space are correlated with each other through a transformation matrix (or other suitable representation). Thus, a single coordinate (in some examples, together with the transformation matrix) can define a first position in the real environment and a corresponding second position in the virtual environment; and vice versa.

[0027] In MRE, virtual objects (e.g., in a virtual environment associated with the MRE) can correspond to real objects (e.g., in a real environment associated with the MRE). For example, if the real environment of the MRE includes a real lamppost (real object) at location coordinates, then the virtual environment of the MRE can include a virtual lamppost (virtual object) at the corresponding location coordinates. As used herein, real objects and their corresponding virtual objects are combined to form a "mixed reality object." Virtual objects do not need to perfectly match or align with their corresponding real objects. In some examples, virtual objects can be simplified versions of their corresponding real objects. For example, if the real environment includes a real lamppost, the corresponding virtual object can include a cylinder with roughly the same height and radius as the real lamppost (reflecting in its roughly cylindrical shape). Simplifying virtual objects in this way allows for computational efficiency and simplifies the computations performed on such virtual objects. Furthermore, in some examples of the MRE, not all real objects in the real environment can be associated with corresponding virtual objects. Similarly, in some examples of the MRE, not all virtual objects in the virtual environment can be associated with corresponding real objects. That is, some virtual objects can exist only in the virtual environment of MRE without any real-world counterparts.

[0028] In some examples, virtual objects can possess characteristics that differ (sometimes drastically) from their corresponding real-world counterparts. For instance, while a real-world environment in an MRE might include a green, two-armed cactus—a thorny, inanimate object—the corresponding virtual object in the MRE might include characteristics of a green, two-armed virtual character with human facial features and aggressive behavior. In this example, the virtual object resembles its real-world counterpart in some characteristics (color, number of arms); but differs from the real object in other characteristics (facial features, personality). In this way, virtual objects have the potential to represent real-world objects in a creative, abstract, exaggerated, or imaginative way; or to impart behavior (e.g., human personality) to other inanimate real-world objects. In some examples, virtual objects can be purely imaginary creations without any real-world counterpart (e.g., a virtual monster in a virtual environment, perhaps in a location corresponding to an empty space in a real environment).

[0029] Compared to VR systems, which present a virtual environment while blurring the real environment, mixed reality (MRE) systems offer the advantage of maintaining the perception of the real environment while the virtual environment is presented. Therefore, users of MRE systems can experience and interact with the corresponding virtual environment using visual and audio cues associated with the real environment. For example, while a VR user might strive to perceive or interact with virtual objects displayed in the virtual environment—because, as mentioned above, the user cannot directly perceive or interact with the virtual environment—an MR user might find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching corresponding real objects in their own real environment. This level of interactivity can enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by simultaneously presenting the real and virtual environments, MRE systems can reduce the negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. Furthermore, MRE systems offer numerous possibilities for applications that can enhance or alter our real-world experiences.

[0030] Figure 1A An example real-world environment 100 is shown where user 110 uses a mixed reality system 112. The mixed reality system 112 may include a display (e.g., a transmissive display) and one or more speakers, as well as one or more sensors (e.g., cameras), as described below. The real-world environment 100 shown includes a rectangular room 104A where user 110 is standing; and real-world objects 122A (lamp), 124A (table), 126A (sofa), and 128A (painting). Room 104A also includes positional coordinates 106, which may be referred to as the origin of the real-world environment 100. Figure 1AAs shown, an environment / world coordinate system 108 (including x-axis 108X, y-axis 108Y, and z-axis 108Z) with an origin at point 106 (world coordinates) can define the coordinate space for the real environment 100. In some embodiments, the origin 106 of the environment / world coordinate system 108 can correspond to the location where the mixed reality environment 112 is powered. In some embodiments, the origin 106 of the environment / world coordinate system 108 can be reset during operation. In some examples, user 110 can be considered a real object in the real environment 100; similarly, body parts of user 110 (e.g., hands, feet) can be considered real objects in the real environment 100. In some examples, a user / audience / head coordinate system 114 (including x-axis 114X, y-axis 114Y, and z-axis 114Z) with an origin at point 115 (e.g., user / audience / head coordinates) can define the coordinate space for the user / audience / head where the mixed reality system 112 is located. The origin 115 of the user / audience / head coordinate system 114 can be defined relative to one or more components of the mixed reality system 112. For example, the origin 115 of the user / audience / head coordinate system 114 can be defined relative to the display of the mixed reality system 112, such as during the initial calibration of the mixed reality system 112. Matrices (which may include translation matrices and quaternion matrices or other rotation matrices) or other suitable representations can characterize the transformation between the user / audience / head coordinate system 114 space and the environment / world coordinate system 108 space. In some embodiments, left ear coordinates 116 and right ear coordinates 117 can be defined relative to the origin 115 of the user / audience / head coordinate system 114. Matrices (which may include translation matrices and quaternion matrices or other rotation matrices) or other suitable representations can characterize the transformation between the left ear coordinates 116 and right ear coordinates 117 and the user / audience / head coordinate system 114 space. The user / audience / head coordinate system 114 can simplify the representation of position relative to the user's head or head-mounted device (e.g., relative to the environment / world coordinate system 108). Using simultaneous localization and mapping (SLAM), visual odometry, or other techniques, the transformation between the user coordinate system 114 and the environment coordinate system 108 can be determined and updated in real time.

[0031] Figure 1BAn example virtual environment 130 corresponding to the real environment 100 is shown. The virtual environment 130 shown includes a virtual rectangular room 104B corresponding to the real rectangular room 104A; a virtual object 122B corresponding to the real object 122A; a virtual object 124B corresponding to the real object 124A; and a virtual object 126B corresponding to the real object 126A. Metadata associated with virtual objects 122B, 124B, and 126B may include information derived from the corresponding real objects 122A, 124A, and 126A. The virtual environment 130 additionally includes a virtual monster 132, which does not correspond to any real object in the real environment 100. Real object 128A in the real environment 100 does not correspond to any virtual object in the virtual environment 130. A persistent coordinate system 133 (including x-axis 133X, y-axis 133Y, and z-axis 133Z) with an origin at point 134 (persistent coordinates) can define the coordinate space used for the virtual content. The origin 134 of the persistent coordinate system 133 can be defined relative to / about one or more real objects, such as real object 126A. Matrices (which may include translation matrices and quaternion matrices or other rotation matrices) or other suitable representations can characterize the transformation between the space of the persistent coordinate system 133 and the environment / world coordinate system 108. In some embodiments, each of the virtual objects 122B, 124B, 126B, and 132 can have its own persistent coordinate point relative to the origin 134 of the persistent coordinate system 133. In some embodiments, multiple persistent coordinate systems may exist, and each of the virtual objects 122B, 124B, 126B, and 132 can have its own persistent coordinate point relative to one or more persistent coordinate systems.

[0032] Compared to Figure 1A and Figure 1B The environment / world coordinate system 108 defines a shared coordinate space for the real environment 100 and the virtual environment 130. In the example shown, the coordinate space has an origin at point 106. Furthermore, the coordinate space is defined by the same three orthogonal axes (108X, 108Y, 108Z). Therefore, a first position in the real environment 100 and a corresponding second position in the virtual environment 130 can be described relative to the same coordinate system. This simplifies the identification and display of corresponding positions in the real and virtual environments, as the same coordinates can be used to identify both positions. However, in some examples, the corresponding real and virtual environments do not need to use a shared coordinate space. For example, in some examples (not shown), matrices (which may include translation matrices and quaternion matrices or other rotation matrices) or other suitable representations can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.

[0033] Figure 1CAn example MRE 150 is shown that simultaneously presents aspects of a real environment 100 and a virtual environment 130 to a user via a mixed reality system 112. In the example shown, the MRE 150 simultaneously presents to the user 110 real objects 122A, 124A, 126A, and 128A from the real environment 100 (e.g., via the transmissive portion of the display of the mixed reality system 112); and virtual objects 122B, 124B, 126B, and 132 from the virtual environment 130 (e.g., via the active display portion of the display of the mixed reality system 112). As described above, the origin 106 serves as the origin for the coordinate space corresponding to the MRE 150, and the coordinate system 108 defines the x-axis, y-axis, and z-axis for the coordinate space.

[0034] In the illustrated example, the mixed reality object comprises corresponding pairs of real and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) occupying corresponding positions in coordinate space 108. In some examples, both the real and virtual objects can be simultaneously visible to the user 110. This is desirable in instances where the virtual object is presented to augment the view of the corresponding real object (such as in a museum application where the virtual object presents a missing, damaged ancient sculpture). In some examples, the virtual object (122B, 124B, and / or 126B) can be displayed (e.g., via active pixelation occlusion using a pixelation occlusion shutter) to occlude the corresponding real object (122A, 124A, and / or 126A). This is desirable in instances where the virtual object acts as a visual replacement for the corresponding real object (such as in an interactive narrative application where an inanimate real object becomes a "living" character).

[0035] In some examples, real-world objects (e.g., 122A, 124A, 126A) may be associated with virtual content or helper data that may not necessarily constitute a virtual object. Virtual content or helper data can facilitate the processing or disposal of virtual objects in a mixed reality environment. For example, such virtual content may include a two-dimensional representation of: the corresponding real-world object; a custom asset type associated with the corresponding real-world object; or statistical data associated with the corresponding real-world object. This information can enable or facilitate computations involving real-world objects without incurring unnecessary computational overhead.

[0036] In some examples, the presentation described above may also include audio aspects. For example, in MRE 150, the virtual monster 132 may be associated with one or more audio signals, such as footsteps generated as the monster moves around in MRE 150. As further described below, the processor of mixed reality system 112 may calculate and process a composite audio signal corresponding to the mixing of all such sounds in MRE 150, and present the audio signal to user 110 via one or more speakers included in mixed reality system 112 and / or one or more external speakers.

[0037] Example Mixed Reality System

[0038] Example mixed reality system 112 may include a wearable head-mounted device (e.g., a wearable augmented reality or mixed reality head-mounted device) comprising: a display (which may include a left transmissive display and a right transmissive display, which may be near-eye displays, and associated components for coupling light from the displays to the user's eyes); a left speaker and a right speaker (e.g., positioned adjacent to the user's left and right ears, respectively); an inertial measurement unit (IMU) (e.g., mounted to a support arm of the head-mounted device); an orthogonal coil electromagnetic receiver (e.g., mounted to the left support); a left camera and a right camera (e.g., a depth (time-of-flight) camera) oriented away from the user; and a left-eye camera and a right-eye camera oriented towards the user (e.g., for detecting the user's eye movements). However, mixed reality system 112 may incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LiDAR, EOG, GPS, magnetic). Additionally, mixed reality system 112 may include network features (e.g., Wi-Fi capability) for communicating with other devices and systems (including other mixed reality systems). The mixed reality system 112 may also include a battery (which may be housed in an auxiliary unit, such as a belt pouch designed to be worn around the user's waist), a processor, and memory. The wearable head-mounted device of the mixed reality system 112 may include tracking components, such as an IMU or other suitable sensors, configured to output a set of coordinates of the wearable head-mounted device relative to the user's environment. In some examples, the tracking components may provide input to a processor performing Simultaneous Localization and Mapping (SLAM) and / or visual odometry. In some examples, the mixed reality system 112 may also include a handheld controller 300 and / or an auxiliary unit 320, which may be a wearable belt pouch, as further described below.

[0039] Figure 2A-2D Components of an example mixed reality system 200 (which may correspond to mixed reality system 112) that can be used to present an MRE (which may correspond to MRE 150) or other virtual environment to a user are shown. Figure 2A A perspective view of a wearable head-mounted device 2102 included in an example mixed reality system 200 is shown. Figure 2B A top view of a wearable head device 2102 worn on a user's head 2202 is shown. Figure 2C A front view of the wearable head device 2102 is shown. Figure 2D A side view of an example eyepiece 2110 of a wearable head device 2102 is shown. Figure 2A-2C As shown, the example wearable head device 2102 includes an example left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an example right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 may include: a transmission element through which the real environment can be visible; and a display element for presenting a display superimposed on the real environment (e.g., via imaging modulated light). In some examples, such a display element may include surface diffraction optics for controlling the imaging modulated light flow. For example, the left eyepiece 2108 may include a left-coupled-in grating set 2112, a left orthogonal pupil expansion (OPE) grating set 2120, and a left-exit (output) pupil expansion (EPE) grating set 2122. Similarly, the right eyepiece 2110 may include a right-coupled-in grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. Imaging modulated light can be delivered to the user's eye via coupling gratings 2112 and 2118, OPEs 2114 and 2120, and EPEs 2116 and 2122. Each coupling grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 can be designed to deflect light downwards in an incremental manner toward its associated EPEs 2122, 2116, thereby horizontally extending the resulting exit pupil. Each EPE 2122, 2116 can be configured to incrementally redirect at least a portion of the light received from its corresponding OPE grating set 2120, 2114 outwards to a user's eyebox location (not shown) defined behind eyepieces 2108, 2110, vertically extending the resulting exit pupil at the eyebox location. Alternatively, instead of coupling grating sets 2112 and 2118, OPE grating sets 2114 and 2120, and EPE grating sets 2116 and 2122, eyepieces 2108 and 2110 may include gratings and / or other arrangements for controlling the refractive and reflective characteristics of coupling imaging modulated light to the user's eye.

[0040] In some examples, the wearable head device 2102 may include a left support arm 2130 and a right support arm 2132, wherein the left support arm 2130 includes a left speaker 2134 and the right support arm 2132 includes a right speaker 2136. An orthogonal coil electromagnetic receiver 2138 may be positioned in the left support or in another suitable location within the wearable head unit 2102. An inertial measurement unit (IMU) 2140 may be positioned in the right support arm 2132 or in another suitable location within the wearable head device 2102. The wearable head device 2102 may also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142 and 2144 may be suitably oriented in different directions to cover a wider field of view together.

[0041] exist Figure 2A-2D In the example shown, the left imaging modulation light source 2124 can be optically coupled to the left eyepiece 2108 via the left coupling grating set 2112, and the right imaging modulation light source 2126 can be optically coupled to the right eyepiece 2110 via the right coupling grating set 2118. The imaging modulation light sources 2124, 2126 may include, for example, fiber optic scanners; projectors, including electro-optic modulators such as digital light processing (DLP) chips or liquid crystal on silicon (LCoS) modulators; or emitting displays, such as micro-LED or micro-organic light-emitting diode (μOLED) panels, each coupled to the coupling grating sets 2112, 2118 using one or more lenses on each side. The input coupling grating sets 2112, 2118 can deflect light from the imaging modulation light sources 2124, 2126 to an angle greater than the critical angle for total internal reflection (TIR) ​​for the eyepieces 2108, 2110. OPE grating sets 2114 and 2120 incrementally deflect light propagating through TIR toward EPE grating sets 2116 and 2122. EPE grating sets 2116 and 2122 incrementally couple light toward the user's face, including the pupils of the user's eyes.

[0042] In some examples, such as Figure 2DAs shown, each of the left eyepiece 2108 and the right eyepiece 2110 includes multiple waveguides 2402. For example, each eyepiece 2108, 2110 may include multiple individual waveguides, each dedicated to a corresponding color channel (e.g., red, blue, and green). In some examples, each eyepiece 2108, 2110 may include multiple sets of such waveguides, wherein each set is configured to impart a different wavefront curvature to the emitted light. The wavefront curvature may be convex relative to the user's eye, for example, to present a virtual object positioned at a distance in front of the user (e.g., by a distance corresponding to the reciprocal of the wavefront curvature). In some examples, the EPE grating sets 2116, 2122 may include curved grating recesses that achieve the convex wavefront curvature by varying the Poynting vector across the emitted light from each EPE.

[0043] In some examples, to create a perception that the displayed content is three-dimensional, stereoscopically accommodating left-eye and right-eye images can be presented to the user via imaging light modulators 2124, 2126 and eyepieces 2108, 2110. The perceptual realism of the presentation of three-dimensional virtual objects can be enhanced by selecting waveguides (and thus corresponding wavefront curvatures) so that the virtual objects are displayed at a distance approximately indicated by the stereoscopic left and right images. This technique can also reduce some of the motion sickness experienced by users, which can be caused by the difference between the depth perception cues provided by the stereoscopic left and right eye images and the autoacoustic function of the human eye (e.g., object distance-related focus).

[0044] Figure 2D An edge-facing view is shown from the top of the right eyepiece 2110 of the example wearable head device 2102. Figure 2D As shown, multiple waveguides 2402 may include a first subset of three waveguides 2404 and a second subset of three waveguides 2406. The two subsets of waveguides 2404 and 2406 can be distinguished by different EPE gratings characterized by different grating line curvatures that impart different wavefront curvatures to the outgoing light. Within each subset of waveguides 2404 and 2406, each waveguide can be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although not explicitly stated in the original text...) Figure 2D (As shown in the diagram, the structure of the left eyepiece 2108 is similar to that of the right eyepiece 2110.)

[0045] Figure 3AAn example handheld controller component 300 of a mixed reality system 200 is shown. In some examples, the handheld controller 300 includes a handle 346 and one or more buttons 350 disposed along a top surface 348. In some examples, the buttons 350 may be configured to serve as optical tracking targets, for example, for tracking the six degrees of freedom (6DOF) motion of the handheld controller 300 in conjunction with a camera or other optical sensors (which may be mounted in the head unit of the mixed reality system 200, such as wearable head device 2102). In some examples, the handheld controller 300 includes tracking components (e.g., an IMU or other suitable sensors) for detecting position or orientation (such as position or orientation relative to wearable head device 2102). In some examples, such tracking components may be positioned in the handle of the handheld controller 300 and / or may be mechanically coupled to the handheld controller. The handheld controller 300 may be configured to provide one or more output signals corresponding to one or more of the button press states; or the position, orientation, and / or motion of the handheld controller 300 (e.g., via an IMU). Such an output signal can be used as an input to the processor of the mixed reality system 200. This input can correspond to the position, orientation, and / or movement of the handheld controller (e.g., the position, orientation, and / or movement of the user's hand holding the controller). Such input can also correspond to user buttons 350.

[0046] Figure 3B An example auxiliary unit 320 of a mixed reality system 200 is shown. The auxiliary unit 320 may include a battery that provides power to the operating system 200 and may include a processor for executing programs on the operating system 200. As shown, the example auxiliary unit 320 includes a chip 2128, such as for attaching the auxiliary unit 320 to a user's belt. Other form factors suited to the auxiliary unit 320 and will be apparent, including form factors not involving mounting the unit to the user's belt. In some examples, the auxiliary unit 320 is coupled to a wearable head device 2102 via a multi-conduit cable, which may include, for example, wires and optical fibers. A wireless connection between the auxiliary unit 320 and the wearable head device 2102 may also be used.

[0047] In some examples, the mixed reality system 200 may include one or more microphones that detect sound and provide corresponding signals to the mixed reality system. In some examples, the microphone may be attached to or integrated with a wearable head device 2102 and configured to detect the user's voice. In some examples, the microphone may be attached to or integrated with a handheld controller 300 and / or auxiliary unit 320. Such a microphone may be configured to detect ambient sound, ambient noise, the user's or a third party's voice, or other sounds.

[0048] Figure 4 An example functional block diagram is shown that can correspond to an example mixed reality system, such as the mixed reality system 200 described above (which can correspond to the mixed reality system 112 relative to Figure 1). Figure 4 As shown, an example handheld controller 400B (which may correspond to handheld controller 300 (“Totem”)) includes a totem-to-wear head-mounted device six-DOF (6DOF) totem subsystem 404A, and an example wearable head-mounted device 400A (which may correspond to wearable head-mounted device 2102) includes a totem-to-wearable head-mounted device 6DOF subsystem 404B. In the example, the 6DOF totem subsystem 404A and 6DOF subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head-mounted device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom can be represented relative to the coordinate system of the wearable head-mounted device 400A. The three translational offsets can be represented as X, Y, and Z offsets in such a coordinate system, translation matrices, or some other representation. The rotational degrees of freedom can be represented as a sequence of yaw, pitch, and roll rotations, rotation matrices, quaternions, or some other representation. In some examples, a wearable head device 400A; one or more depth cameras 444 (and / or one or more non-depth cameras) included in the wearable head device 400A; and / or one or more optical targets (e.g., a button 450 of a handheld controller 400B as described above, or a dedicated optical target included in the handheld controller 400B) can be used for 6DOF tracking. In some examples, the handheld controller 400B may include a camera as described above; and the wearable head device 400A may include an optical target for combined camera optical tracking. In some examples, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids for wirelessly transmitting and receiving three distinguishable signals. The 6DOF of the wearable head device 400A relative to the handheld controller 400B can be determined by measuring the relative magnitudes of the three distinguishable signals received in each of the coils used for receiving. In addition, the 6DOF totem subsystem 404A may include an inertial measurement unit (IMU) that can be used to provide improved accuracy and / or more timely information about the rapid movements of the handheld controller 400B.

[0049] In some examples, it may become necessary to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment), for example, to compensate for motion of the wearable head device 400A relative to coordinate system 108. For example, such a transformation may be necessary for the display of the wearable head device 400A to present virtual objects at a desired position and orientation relative to the real environment (e.g., a virtual person sitting in a real chair, facing forward, regardless of the position and orientation of the wearable head device), rather than at a fixed position and orientation on the display (e.g., the same position in the lower right corner of the display), to maintain the illusion that the virtual objects exist in the real environment (and, for example, not appear unnaturally positioned in the real environment when the wearable head device 400A moves and rotates). In some examples, the compensation transformation between coordinate spaces can be determined by processing images from depth camera 444 using SLAM and / or visual odometry procedures to determine the transformation of the wearable head device 400A relative to coordinate system 108. Figure 4 In the example shown, a depth camera 444 is coupled to a SLAM / visual odometry block 406 and can provide imagery to the block 406. SLAM / visual odometry block 406 implementations may include a processor configured to process the imagery and determine the position and orientation of the user's head, which can then be used to identify transformations between the head coordinate space and another coordinate space (e.g., inertial coordinate space). Similarly, in some examples, an additional source of information about the user's head pose and position is obtained from an IMU 409. Information from the IMU 409 can be integrated with information from the SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information regarding the user's head pose and position through rapid adjustment.

[0050] In some examples, depth camera 444 can feed 3D images to gesture tracker 411, which can be implemented in the processor of wearable head device 400A. Gesture tracker 411 can identify a user's gestures, for example, by matching the 3D images received from depth camera 444 with stored patterns representing gestures. Other suitable techniques for identifying user gestures will be apparent.

[0051] In some examples, one or more processors 416 may be configured to receive data from a 6DOF helmet subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, and / or gesture tracker 411 of a wearable head-mounted device. Processor 416 may also send and receive control signals from a 6DOF totem system 404A. Processor 416 may be wirelessly coupled to the 6DOF totem system 404A, such as in examples not limited to a handheld controller 400B. Processor 416 may also communicate with additional components, such as an audio-visual content memory 418, a graphics processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer. DSP audio spatializer 422 may be coupled to a head-related transfer function (HRTF) memory 425. GPU 420 may include a left channel output coupled to a left imaging modulation light source 424 and a right channel output coupled to a right imaging modulation light source 426. GPU 420 can output stereo image data to imaging modulation light sources 424 and 426, for example, as mentioned above relative to... Figure 2A-2D As described, the DSP audio spatializer 422 can output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 can receive input from the processor 419 indicating a direction vector from the user to a virtual sound source (which can be moved by the user, e.g., via a handheld controller 320). Based on this direction vector, the DSP audio spatializer 422 can determine the corresponding HRTF (e.g., by accessing the HRTF or by interpolating multiple HRTFs). The DSP audio spatializer can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. This can improve the credibility and realism of the virtual sound by interpolating the user's relative position and orientation to the virtual sound in the mixed reality environment—that is, by presenting a virtual sound that matches the user's expectation that the virtual sound will sound like a real sound in a real environment.

[0052] In some examples, such as Figure 4 As shown, one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 may be included in the auxiliary unit 400C (which may correspond to the auxiliary unit 320 described above). The auxiliary unit 400C may include a battery 427 that powers its components and / or powers the wearable head device 400A or the handheld controller 400B. Including such components in an auxiliary unit that can be mounted to the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0053] Although Figure 4 Elements corresponding to various components of the example mixed reality system are presented, but various other suitable arrangements of these components will become apparent to those skilled in the art. For example, in Figure 4 The components presented as associated with the auxiliary unit 400C may conversely be associated with the wearable head device 400A or the handheld controller 400B. Furthermore, some mixed reality systems may completely omit the handheld controller 400B or the auxiliary unit 400C. Such changes and modifications will be understood to be included within the scope of the disclosed examples.

[0054] Ambient acoustic durability

[0055] As described above, an MRE (such as one experienced via a mixed reality system, like mixed reality system 112, which may include components such as the wearable head unit 200, handheld controller 300, or auxiliary unit 320 described above) can present audio signals to the MRE user as if they originated from a sound source with the MRE's origin coordinates. That is, the user can perceive these audio signals as if they were real audio signals originating from the sound source's origin coordinates.

[0056] In some cases, audio signals can be considered virtual because they correspond to computational signals in a virtual environment. Virtual audio signals can be presented to the user as real audio signals that can be detected by the human ear, for example, as generated by speakers 2134 and 2136 of the wearable head unit 200 in Figure 2.

[0057] The sound source can correspond to a real object and / or a virtual object. For example, a virtual object (e.g., Figure 1C The virtual monster (132) can emit audio signals in the MRE, which are represented as virtual audio signals in the MRE and presented to the user as real audio signals. For example, Figure 1C The virtual monster 132 can emit virtual sounds corresponding to the monster's voice (e.g., dialogue) or sound effects. Similarly, real objects (e.g., Figure 1CThe real object 122A appears to emit a virtual audio signal in the MRE, which is represented as a virtual audio signal in the MRE and presented to the user as a real audio signal. For example, a real light 122A can emit a virtual sound corresponding to the sound effect of the light being turned on or off—even if the light is not turned on or off in the real environment. The virtual sound can correspond to the location and orientation of a sound source (whether real or virtual). For example, if a virtual sound is presented to the user as a real audio signal (e.g., via speakers 2134 and 2136), the user can perceive the virtual sound as originating from the location of the sound source. The sound source is referred to herein as a "virtual sound source," even though the underlying object that obviously emits the sound can itself correspond to a real or virtual object, as described above.

[0058] Some virtual or mixed reality environments suffer from the problem of the environment not feeling real or believable. One reason for this perception is that audio and visual cues don't always match in such environments. For example, if a user is positioned behind a large brick wall in an MRE, the user might expect a sound coming from behind the wall to be quieter and less audible than a sound originating next to the user. This expectation is based on the user's auditory experience in the real world, where sound becomes quiet and inaudible as it passes through large, dense objects. When the user is presented with an audio signal that is supposedly originating from behind the brick wall but is strong and presented at full volume, the illusion of the sound originating from behind the brick wall is compromised. The entire virtual experience may feel fake and unreal, partly because it is not based on the user's expectation that real-world interactions fit into their expectations. Furthermore, in some cases, the "uncanny valley" problem arises, where even subtle differences between virtual and real experiences can cause heightened feelings of discomfort. Improving the user experience by presenting audio signals in the MRE that appear to interact with objects in the user's environment realistically—even subtly. The more consistent the audio signal is with the user's expectations, the more immersive and engaging the user's MRE experience can be, based on real-world experience.

[0059] One way users perceive and understand their surroundings is through audio cues. In the real world, the actual audio signals a user hears are influenced by where those signals originate and what objects they interact with. For example, all other things being equal, a sound originating from a long distance away (e.g., a dog barking in the distance) will sound quieter than the same sound originating from a short distance (e.g., a dog barking in the same room as the user). Therefore, a user can identify the location of a dog in a real-world environment, partly based on the perceived volume of the bark. Similarly, all other things being equal, a sound moving away from the user (e.g., the voice of a person moving away from the user) will sound less distinct and less audible (i.e., low-pass filtered) than the same sound moving towards the user (e.g., the voice of a person facing the user). Therefore, a user can identify the orientation of a person in a real-world environment based on the perceived frequency characteristics of that person's voice.

[0060] The user's perception of a real audio signal can also be influenced by the presence of objects in the environment to which the audio signal interacts. That is, the user can perceive not only the audio signal generated by the sound source, but also its reflection from nearby objects and the reverberation characteristics imparted by the surrounding acoustic space. For example, if a user speaks in a small room with enclosed walls, those walls may produce a short, natural reverberation signal because the human voice reflects off the walls. The user can infer from that reverberation that they are in a small, enclosed room. Similarly, a large concert hall or church may produce a longer reverberation, from which the user can infer that they are in a large, spacious room. Likewise, the reverberation of an audio signal can exhibit various acoustic characteristics based on the location or orientation of the surfaces from which the signal reflects, or the materials of those surfaces. For example, the reverberation of a patchwork wall will sound different from the reverberation of brick, carpet, drywall, or other materials. These reverberation characteristics can be used by the user to understand—acoustically—the size, shape, and material composition of the space they inhabit.

[0061] The examples above illustrate how audio cues can inform a user's perception of their surroundings. These cues can work in conjunction with visual cues: for example, if we see a dog in the distance, we expect the sound of its bark to correspond to that distance (and if it doesn't, we might feel confused or disoriented, as in some virtual environments). In some examples, such as in low-light environments, or relative to visually impaired users, visual cues may be limited or unavailable; in such cases, audio cues can take on specific importance and can serve as the primary means for users to understand their environment.

[0062] The system architecture may facilitate the organization, storage, retrieval, and / or management of information required to render realistic virtual audio. For example, an MR system (e.g., MR systems 112, 200) can manage environmental information, such as the real-world environment in which a user might be, the acoustic characteristics of that environment, and / or the user's location within that environment. MR systems can further manage information about objects in the real and / or virtual environments (e.g., objects that may affect the general acoustic characteristics of the real environment and / or objects that may affect the acoustic characteristics of virtual sound sources interacting with those objects). MR systems can also manage information about virtual sound sources. For example, the location of virtual sound sources may be relevant to rendering realistic virtual audio.

[0063] In addition to managing the virtual audio system, other systems may also need to be managed to deliver a complete MR experience. For example, a complete MR experience might require a virtual vision system that manages information used to render virtual objects. A complete MR experience might also require a Simultaneous Localization and Mapping (“SLAM”) system that builds, updates, and / or maintains a 3D model of the user's environment. Besides the virtual audio system, MR systems (e.g., MR systems 112, 200) can manage these and more systems to deliver a complete MR experience. The virtual audio system architecture may facilitate the management of interactions between these systems to enable data transfer, management, storage, and / or preservation.

[0064] In some embodiments, a system (e.g., a virtual audio system) can interact with other higher-level systems. In some embodiments, lower-level systems (e.g., virtual audio systems) can interact more closely with hardware-level inputs and / or outputs, while higher-level systems (e.g., applications) can interact with lower-level systems. Higher-level systems can utilize lower-level systems to perform their functions (e.g., a gaming application might rely on a lower-level virtual audio system to render realistic virtual audio). Virtual audio systems can benefit from system architectures designed to manage interactions with higher-level systems while maintaining the integrity of the virtual audio system. For example, multiple higher-level systems (e.g., multiple third-party applications) can interact with the virtual audio system simultaneously or substantially simultaneously. In some embodiments, maintaining a single virtual audio system capable of rendering virtual audio may be computationally more efficient than having each higher-level system maintain a separate audio system. For example, in some embodiments, a single digital reverb can be used to process sound objects from multiple higher-level systems when those objects are intended to be located in the same virtual or real acoustic space (e.g., the user's room). A well-designed system architecture can also protect the integrity of information that might be used in other applications (e.g., preventing data corruption and / or tampering).

[0065] In some embodiments, it may be advantageous to design a system architecture that allows for real-time changes without interrupting services to other systems (e.g., higher-level systems). For example, a virtual audio system may store, maintain, or otherwise manage an audio model that takes into account the acoustic properties of a real-world environment (e.g., a room). If a user changes the real-world environment (e.g., moves to a different room), the audio model may be updated to reflect these changes. If an MR system is currently in use (e.g., the MR system is presenting virtual visuals and / or virtual audio to a user), the audio model may need to be updated while other systems (e.g., higher-level systems) are still using it to render virtual audio.

[0066] In some embodiments, it may be advantageous to design a system architecture that allows changes to be propagated to other systems. For example, some systems may maintain a separate copy of the audio model, or some systems may store specific repetitive sound effects rendered using an audio model maintained by a virtual audio system. Therefore, it may be advantageous to propagate changes made in the virtual audio system to other systems. For example, if a user changes the environment (e.g., moves the room) and a new audio model may be more accurate, the virtual audio system can modify its audio model and notify any clients (e.g., systems that use and / or depend on the virtual audio system) of the change. In some embodiments, the client can then query the virtual audio system and update its internal data accordingly.

[0067] Figure 5An exemplary virtual audio system according to some embodiments is illustrated. The virtual audio system 500 may include a persistent module 502. A module (e.g., persistent module 502) may include one or more computer systems configured to execute instructions and / or store one or more data structures. In some embodiments, the module (e.g., persistent module 502) may be configured to execute processes, subprocesses, threads, and / or services managed by an audio service 522 (e.g., instructions executed by persistent module 502 may run within the audio service 522), which may run on one or more computer systems. In some embodiments, the audio service 522 may be a process that can run in a runtime environment, and the instructions executed by the module (e.g., persistent module 502) may be components of the audio service 522 (e.g., instructions executed by persistent module 502 may be child processes of the audio service 522). In some embodiments, the audio service 522 may be a child process of a parent process. Instructions executed by a module (e.g., persistent module 502) may include one or more components (e.g., processes, subprocesses, threads, and / or services executed by the positioning state submodule 506, acoustic data submodule 508, and / or audio model submodule 510). In some embodiments, instructions executed by a module (e.g., persistent module 502) may run as a subprocess of audio service 522 and / or as a separate process in a location different from other components of audio service 522. For example, instructions executed by a module (e.g., persistent module 502) may run in a general-purpose processor, and one or more other components of audio service 522 may run in an audio-specific processor (e.g., DSP). In some embodiments, instructions executed by a module (e.g., persistent module 502) may run in a process address space and / or memory space different from other components of audio service 522. In some embodiments, instructions executed by a module (e.g., persistent module 502) may run as one or more threads within audio service 522. In some embodiments, instructions executed by a module (e.g., persistent module 502) may be instantiated within the audio service 522. In some embodiments, instructions executed by a module (e.g., persistent module 502) may share process addresses and / or memory space with other components of the audio service 522.

[0068] In some embodiments, the persistence module 502 may include a location state submodule 506. The location state submodule 506 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the instructions executed by the location state submodule 506 may be a subprocess of the persistence module 502. In some embodiments, the location state submodule 506 may indicate whether location has been achieved (e.g., the location state submodule 506 may indicate whether the MR system has identified the real environment and / or located itself in the real environment). In some embodiments, the location state submodule 506 may interact with the location system (e.g., via an API). The location system may determine the location for the MR system (and / or the user using the MR system). In some embodiments, the location system may utilize SLAM-like techniques to create a 3D model of the real environment and estimate the system's (and / or the user's) location within the environment. In some embodiments, the location system may rely on a connectable world system (described further in detail below) and one or more sensors (e.g., one or more sensors of MR systems 112, 200) to estimate the MR system's (and / or the user's) location within the environment. In some embodiments, the positioning status submodule 506 may query the positioning system to determine whether positioning for the MR system has been achieved. Similarly, the positioning system may notify the positioning status submodule 506 of successful positioning. (For example, the positioning status of MR systems 112, 200) can be used to determine whether the audio model should be updated (e.g., because the user's real-world environment has changed).

[0069] In some embodiments, the persistence module 502 may include an acoustic data submodule 508. The acoustic data submodule 508 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, instructions executed by the acoustic data submodule 508 may be subprocesses of the persistence module 502. In some embodiments, the acoustic data submodule 508 may store one or more data structures representing acoustic data that can be used to create audio models. In some embodiments, the acoustic data submodule 508 may interact with a connectable world system (e.g., via an API). The connectable world system may include information about a known real-world environment (e.g., rooms, buildings, and / or external spaces) and associated real and / or virtual objects. In some embodiments, the connectable world system may include a persistent coordinate system and / or anchor points. The persistent coordinate system and / or anchor points may be fixed points in space that the MR system may know (e.g., by unique identifiers). Virtual objects may be positioned relative to one or more persistent coordinate systems and / or anchor points to enable object persistence (e.g., virtual objects may appear to remain in the same location in the real environment regardless of who is viewing the virtual object or any movement by the user). Persistent coordinate systems and / or anchor points can be particularly advantageous when two or more users with separate MR systems use different world coordinate systems (e.g., each user's location is designated as the origin for their respective world coordinate system). Object persistence across users can be achieved by converting between individual world coordinate systems and a universal persistent coordinate system, and by placing / referencing virtual objects relative to the persistent coordinate system. In some embodiments, a connectable world system can manage and maintain a persistent coordinate system by, for example, mapping new areas and creating new persistent coordinate systems; remapping known areas and harmonizing new persistent coordinate systems with previously determined persistent coordinate systems; and / or associating persistent coordinate systems with identifiable information (e.g., location and / or nearby objects). In some embodiments, the acoustic data submodule 508 can query one or more persistent coordinate systems from a separate system (e.g., a connectable world system). In some embodiments, the acoustic data submodule 508 can retrieve associated persistent coordinate systems (e.g., persistent coordinate systems within a threshold radius of a user's location) to facilitate access and management, including the creation, modification, and / or deletion of associated acoustic data.

[0070] In some embodiments, acoustic data stored in the acoustic data submodule 508 may be organized into physically relevant modular units (e.g., a room may be represented by a modular unit, and a chair within the room may be represented by another modular unit). For example, a modular unit may include physical and / or perceived properties of the physical environment (e.g., the room). Physical and / or perceived properties may include attributes that may affect the acoustic characteristics of the room (e.g., the room's size and / or shape). In some embodiments, physical and / or perceived properties may include functional and / or behavioral attributes that can be interpreted by a rendering engine (e.g., whether sources outside the room should be occluded). In some embodiments, physical and / or perceived properties may include attributes of known and / or identified objects. For example, the geometry of fixed (e.g., floors, walls, furniture, etc.) and / or movable (e.g., cups) objects may be stored as physical and / or perceived properties and may be associated with a specific environment. In some embodiments, physical and / or perceived properties may include transmission loss, scattering coefficient, and / or absorption coefficient. In some embodiments, a modular unit may include physically and / or perceived links (e.g., between other modular units or within a modular unit). For example, physical and / or sensory-related links can link two or more rooms together and describe how the rooms can interact with each other (e.g., cross-coupling gain levels between digital reverberators simulating the rooms and / or line-of-sight paths between the two spaces).

[0071] In some embodiments, physical and / or perceived properties may include acoustic properties such as reverberation time, reverberation delay, and / or reverberation gain. Reverberation time may include the length of time required for sound to decay by a certain amount (e.g., 60 dB). Sound decay may be a result of sound reflection from surfaces in the real environment (e.g., walls, floors, furniture, etc.) while energy is lost due to sound absorption caused by factors such as room boundaries (e.g., walls, floors, ceilings, etc.), objects within the room (e.g., chairs, furniture, people, etc.), and the air within the room. Reverberation time may be affected by environmental factors. For example, absorbent surfaces (e.g., cushions) may absorb sound in addition to geometric expansion and thus reduce reverberation time. In some embodiments, information about the origin may not be required to estimate the reverberation time of the environment. Reverberation gain may include the ratio of the direct / source / original energy of the sound to the reverberant energy of the sound (e.g., the energy of the reverberation produced by the direct / source / original sound), where the listener and source are substantially co-located (e.g., a user clapping their hands produces a source sound that can be considered substantially co-located with one or more microphones mounted on a head-mounted MR system). For example, a pulse (e.g., clapping) can have energy associated with that pulse, and the reverberant sound from that pulse can have energy associated with the reverberation of that pulse. The ratio of the original / source energy to the reverberant energy can be the reverberation gain. The reverberation gain in a real environment may be affected by, for example, absorbing surfaces that can absorb sound and thus reduce reverberation energy.

[0072] In some embodiments, acoustic data may include metadata (e.g., metadata about physical and / or perceived properties). For example, information about when and / or where acoustic data was collected may be included in the acoustic data. In some embodiments, confidence data associated with the acoustic data (e.g., estimated measurement accuracy and / or counts of repeated measurements) may be included as metadata. In some embodiments, a type of modular unit (e.g., a modular unit for a room or a link between modular units) and / or a data version may be included as metadata. In some embodiments, unique identifiers associated with the acoustic data, persistent coordinate system, and / or anchor points may be included as metadata. In some embodiments, relative transformations from the persistent coordinate system and / or anchor points and associated virtual objects may be included as metadata. In some embodiments, metadata may be stored as a single bundle along with the acoustic data.

[0073] In some embodiments, acoustic data may be organized using persistent coordinate systems and / or anchor points, and these persistent coordinate systems and / or anchor points may be organized into a map. In some embodiments, the audio model may take into account acoustic data organized using persistent coordinate systems and / or anchor points, which may correspond to locations within the environment. In some embodiments, acoustic data may be loaded into the acoustic data submodule 508 upon a successful localization event (which may be indicated by the localization status submodule 506). In some embodiments, all available acoustic data may be loaded into the acoustic data submodule 508. In some embodiments, only relevant acoustic data may be loaded into the acoustic data submodule 508 (e.g., acoustic data for persistent coordinate systems and / or anchor points within a specific distance of the location in the MR system).

[0074] In some embodiments, acoustic data may include different states that can change according to variations in the real environment. For example, acoustic data representing a given modular unit of a room may include acoustic data of a room in an empty state and acoustic data of a room in an occupied state. In some embodiments, changes in the arrangement of furniture in a room may be reflected by changes in the state of acoustic data associated with the room. In some embodiments, a modular unit (e.g., representing a room) may include different acoustic data for the state when a door is open or closed. The state may be represented as a binary value (e.g., 0 or 1) or a continuous value (e.g., the degree of door opening, the occupancy of the room, etc.).

[0075] In some embodiments, persistence module 502 may include audio model submodule 510. Audio model submodule 510 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, audio model submodule 510 may include one or more data structures representing an audio model for a real and / or virtual environment. In some embodiments, the audio model may be generated at least in part from acoustic data stored in acoustic data submodule 508. The audio model may represent how sound behaves in a particular environment. For example, virtual sounds generated by an MR system (e.g., MR systems 112, 200) may be modified by the audio model in audio model submodule 510 to reflect the acoustic properties of the environment. A virtual concert presented to a user sitting in a large concert hall may have similar acoustic properties to a real concert presented in the same concert hall. The MR system may locate itself to the identified concert hall, load relevant acoustic data, and generate an audio model to model the acoustic properties of the concert hall.

[0076] In some embodiments, the audio model can be used to model sound propagation in an environment. For example, propagation effects may include occlusion, blockage, early reflection, diffraction, time-of-flight delay, Doppler effect, and other effects. In some embodiments, the audio model may take into account frequency-related absorption and / or transmission losses (e.g., based on acoustic data loaded into the acoustic data submodule 508). In some embodiments, the audio model stored in the audio model submodule 510 may inform other aspects of the audio engine. For example, the audio model may use acoustic data to procedurally synthesize audio (e.g., collisions between virtual and / or real objects).

[0077] In some embodiments, the audio rendering service 522 may include a rendering track module 514. The rendering track module 514 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the rendering track module 514 may include audio information that can later be presented to a user. In some embodiments, the MR system may present virtual sounds comprising several sound sources mixed together (e.g., the sound of two swords clashing and the sound of a person shouting). The rendering track module 514 may store one or more tracks that can be mixed with other tracks to be presented to a user. In some embodiments, the rendering track module 514 may include information about spatial sources. For example, the rendering track module 514 may include information about where the sound sources are located, which may be taken into account in the audio model and / or rendering algorithm. In some embodiments, the rendering track module 514 may include information about the relationships between modular units and / or sound sources. For example, one or more rendering tracks and / or audio models may be associated together as a single group.

[0078] In some embodiments, the audio rendering service 522 may include a position manager module 516. The position manager module 516 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the position manager module 516 may manage location information associated with the audio engine (e.g., the current location of the MR system in a real-world environment). In some embodiments, the position manager module 516 may include a perception wrapper submodule. The perception wrapper submodule may be a wrapper around perception data (e.g., what the MR system has detected or is detecting). In some embodiments, the perception wrapper may communicate and / or transform between the perception data and the position manager module 516. In some embodiments, the position manager module 516 may include a head pose submodule, which may include head pose data. The head pose data may include the location and / or orientation of the MR system (or the corresponding user) in a real-world environment. In some embodiments, head pose may be determined based on the perception data.

[0079] In some embodiments, audio rendering service 522 may include audio model module 518. Audio model module 518 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, audio model module 518 may include an audio model, which may be the same audio model included in audio model submodule 510. In some embodiments, modules 510 and 518 may maintain duplicate copies of the same audio model. For example, maintaining more than one copy of the audio model may be advantageous when updating the audio model but needing to present sound to the user. It may be advantageous to update the copy of the model when the model is currently in use and then update the outdated model when it becomes available (e.g., no longer in use). In some embodiments, the audio model may be transferred between modules 510 and 518 via serialization. The audio model in module 510 may be serialized and deserialized to facilitate data transfer to module 518. Serialization may facilitate data transfer between processors (e.g., general-purpose processors and audio-specific processors) so that shared typed memory is not required.

[0080] In some embodiments, the audio rendering service 522 may include a rendering algorithm module 520. The rendering algorithm module 520 may include one or more computer systems configured to execute instructions and / or store one or more data structures. For example, the rendering algorithm module 520 may include algorithms for rendering virtual sounds so that they can be presented to a user (e.g., via one or more speakers of an MR system). The rendering algorithm module 520 may take into account an audio model specific to an environment (e.g., the audio model in modules 510 and / or 518).

[0081] In some embodiments, audio service 522 may be a process, subprocess, thread, and / or service running on one or more computer systems (e.g., in MR systems 112, 200). In some embodiments, a separate system (e.g., a third-party application) may request the presentation of an audio signal (e.g., via one or more speakers of MR system 112, 200). Such a request may take any suitable form. In some embodiments, the request to present an audio signal may include software instructions to present an audio signal; in some embodiments, such a request may be hardware-driven. The request may be made with or without user involvement. Further, such requests may be received via local hardware (e.g., from the MR system itself), via external hardware (e.g., a separate computer system communicating with the MR system), via the Internet (e.g., via a cloud server), or via any other suitable source or combination of sources. In some embodiments, audio service 522 may receive the request, render the requested audio signal (e.g., via rendering algorithm 520, which may take into account the audio models from blocks 510 and / or 518), and present the requested audio signal to the user. In some embodiments, audio service 522 may be a process that runs continuously (e.g., in the background) while the operating system of the MR system is running. In some embodiments, audio service 522 may be an instance of a parent background service, which may serve as the host process for one or more background processes and / or child processes. In some embodiments, audio service 522 may be part of the operating system of the MR system. In some embodiments, applications that can run on the MR system may access audio service 522. In some embodiments, the user of the MR system may not provide input directly to audio service 522. For example, the user may provide input (e.g., movement commands) to an application running on the MR system (e.g., a role-playing game). The application may provide input to audio service 522 (e.g., to render footsteps), and audio service 522 may provide output (e.g., the rendered footsteps) to the user (e.g., via a speaker) and / or other processes and / or services.

[0082] Figure 6 An exemplary process for updating an audio model according to some embodiments is shown. At step 606, a location may be determined (e.g., the MR system can successfully identify its location within the environment). At step 607, which may occur within persistence module 602 (which may correspond to persistence module 502), a notification of successful location may be issued. In some embodiments, the notification of successful location may trigger the process of updating the audio model (e.g., because the previous audio model may no longer be applicable to the current location).

[0083] At step 608, it can be determined whether to initiate a call to the acoustic data. It may be desirable to set one or more conditions for initiating the call so that the audio model is not updated too frequently. For example, if a user using the MR system only moves slightly within the room, updating the audio model may not be desirable (e.g., because an updated model may not be perceptually distinguishable from the existing model, and / or continuous model updates may be computationally expensive). In some embodiments, the threshold condition at step 608 may be time-based. For example, a call may be initiated only if it has not been initiated within the previous 5 seconds. In some embodiments, the threshold condition at step 608 may be location-based. For example, a call may be initiated only if the user has changed their location by a threshold distance. It should be noted that other threshold conditions may also be used. In some embodiments, steps 607 and / or 608 may occur within a location state submodule (e.g., location state submodule 506).

[0084] If it is determined that a call should be initiated, a persistent coordinate system can be retrieved at step 610. In some embodiments, only a subset of available persistent coordinate systems can be retrieved at step 610. For example, only persistent coordinate systems near the location can be retrieved.

[0085] At step 612, acoustic data may be retrieved. In some embodiments, the acoustic data retrieved at step 612 may correspond to acoustic data stored in acoustic data submodule 508. In some embodiments, only a subset of the available acoustic data may be retrieved at step 612. For example, acoustic data associated with one or more persistent coordinate systems and / or anchor points may be retrieved. In some embodiments, steps 612 and / or 614 may occur within the acoustic data submodule (e.g., acoustic data submodule 508).

[0086] At step 614, an audio model may be constructed and / or modified. In some embodiments, the audio model may take into account the acoustic data retrieved at step 612, and the audio model may model the acoustic characteristics of a particular environment. In some embodiments, step 614 may occur within an audio model submodule (e.g., audio model submodule 510).

[0087] At step 616, it can be determined whether a copy of the audio model should be updated. It may be desirable to set one or more conditions for updating the copy of the audio model to avoid service interruption (e.g., presenting audio to the user). For example, a condition may assess whether a copy of the audio model exists (e.g., within audio rendering service 604 but outside of persistence module 602). If a copy of the audio model does not exist, the audio model can be retrieved by audio rendering service 604 (which may correspond to audio rendering service 522). In some embodiments, a condition may assess whether a copy of the audio model is currently in use (e.g., whether the audio model is being used to render audio for presentation to the user). If a copy of the audio model is not in use, an updated audio model (e.g., the audio model generated at step 614) can be propagated to that copy.

[0088] At step 618, the audio rendering service 604 can retrieve a copy of the audio model (e.g., the audio model generated in step 614). The audio model can be transmitted using data serialization and / or deserialization.

[0089] At step 620, outdated audio models may optionally be deleted and / or disabled. For example, audio rendering service 604 may have a first existing audio model that it has previously used. Audio rendering service 604 may retrieve a second updated audio model (e.g., from persistence module 602) and delete and / or disable the first existing audio model.

[0090] At step 622, a notification of the new audio model may be issued. In some embodiments, the notification may include a callback function for a client (e.g., a third-party application) that can subscribe to hear the audio model when it changes.

[0091] Figure 7 An exemplary process for updating an audio model according to some embodiments is illustrated. At step 706, audio data may be received (e.g., via one or more sensors of MR systems 112, 200). In some embodiments, audio data may be manually entered (e.g., a user and / or developer may manually enter reverberation time, reverberation delay, reverberation gain, etc.). At step 708, which may occur in persistence module 702 (which may correspond to persistence module 502), the associated environment may be identified. The associated environment may be identified by metadata that may accompany the audio data (e.g., the metadata may carry information about one or more persistent coordinate systems and / or anchor points that may be known to the MR system).

[0092] At step 710, it can be determined whether the associated environment is new. For example, if the associated environment may not be identified and / or associated with an unknown identifier, it can be determined that the audio data is associated with a new environment. If it is determined that the associated environment is not new, the copy of the official audio model can be updated (e.g., utilizing room properties that can be derived from the audio data). If it is determined that the associated environment is new, the new environment can be added to the copy of the official audio model, and the copy audio model can be updated accordingly. In some embodiments, the new environment may be represented by a new modular unit.

[0093] At step 716, metadata associated with the new environment (e.g., metadata associated with the new modular unit) can be initialized. For example, metadata associated with measurement counts, confidence levels, or other information can be created and bundled with the new modular unit.

[0094] At step 718, the official audio model within the persistence module 702 can be updated. For example, the official audio model can be copied from the updated copy audio model. In some embodiments, the official audio model can be locked at step 718 to prevent further changes to the official audio model. In some embodiments, changes can still be made to the copy of the official audio model (which may still exist within the persistence module 702) even when the official audio model is locked.

[0095] At step 720, acoustic data associated with the new audio data can be saved. For example, a new modular unit associated with a new room can be saved and / or transferred to a connectable world system (which can then be accessed by the MR system as needed in the future). In some embodiments, step 720 can occur within the acoustic data submodule 508. In some embodiments, step 720 can occur sequentially after step 718. In some embodiments, step 720 can occur at an independent time, as determined by other components, for example, based on the availability of a connectable world system where acoustic data can be saved.

[0096] At step 722, audio rendering service 704 (which may correspond to audio rendering service 522) may retrieve a copy of the audio model. In some embodiments, the copy of the audio model may be the same audio model updated at step 718. In some embodiments, data transfer may occur via serialization and deserialization. In some embodiments, the audio model may be locked during the serialization process, which may prevent the audio model from being changed when a snapshot of the audio model is created.

[0097] At step 724, outdated copies of the audio model can be deleted and / or disabled.

[0098] At step 726, a notification about the new audio model can be issued. This notification can be a callback function for subscribed clients to hear when the model is updated.

[0099] At step 728, a formal audio model can be published, which may indicate that the serialized package corresponding to the audio model can be deleted. In some embodiments, it may be desirable to lock the formal audio model within the persistence module 702 while a copy of the formal audio model is being transmitted to the audio rendering service 704. Once the audio rendering service 704 has finished retrieving a copy of the formal audio model, it may be desirable to publish the lock on the formal audio model so that the formal audio model can continue to be updated.

[0100] In some embodiments, the audio rendering service (e.g., audio rendering service 522) may not manage the interaction between the persistence module (e.g., persistence module 502) and the rendering algorithm (e.g., rendering algorithm module 520). For example, the rendering algorithm 520 may communicate directly with the persistence module 502 to retrieve updated audio models. In some embodiments, the rendering algorithm 520 may include a copy of its own audio model. In some embodiments, the rendering algorithm 520 may access the audio model within the persistence module 502.

[0101] Multi-application audio rendering

[0102] MR systems (e.g., MR systems 112, 200) can utilize various onboard sensors to develop customized audio models for a user's MRE. The MR system can interpret real-world features within the user's MRE (e.g., floors, walls, people) and virtual objects within the user's MRE (e.g., a virtual sofa) to generate a customized audio model that appropriately reflects the real and / or virtual physical phenomena of the user's MRE. For example, a virtual sound source may be located within the user's MRE. The audio model can identify the direct path from the virtual sound source to the user to determine whether virtual sounds should be occluded by real or virtual objects in the direct path. In some embodiments, the audio model can determine how sound from a virtual sound source is reflected from real and / or virtual objects within the user's MRE (e.g., hard surfaces reflect more sound, certain surfaces reflect higher frequencies, rough surfaces scatter sound in multiple directions, etc.). In some embodiments, the audio model can determine how sound from a virtual sound source reverberates within the user's MRE. In some embodiments, it may be desirable to allow for customization of the model, which may be "accurate" to the real and / or virtual physical phenomena of the user's MRE. For example, an MR application can start in the user's current physical space and gradually transform the user's environment into a different, exotic one. As the MR application transforms the user's environment (e.g., by adding virtual trees, virtual leaves, etc.), it can adapt to the "formal" audio model (e.g., by removing physical walls and ceilings from the formal audio model) to suit the desired user experience.

[0103] However, problems can arise when more than one MR application is running simultaneously on an MR system. In some embodiments, each application may want to use a different audio model to present sound to the user. For example, an MR system may run a web browsing application and a podcast application simultaneously. In some embodiments, the web browser may play an informational video, and the web browser may use a formal audio model generated by the MR system. Because the web browser can use a formal audio model, the informational video may sound as if the speaker is located in the user's MRE (e.g., the speaker's sound may be appropriately masked, reflected, and reverberated based on real and / or virtual objects in the user's MRE). In some embodiments, a podcast application may play a podcast located in a cave. Therefore, it may be desirable for the podcast application to present sound as highly reflected and / or reverberant.

[0104] However, conflicts can arise if web browsing and podcast applications simultaneously attempt to render audio using different audio models. In some embodiments, audio tracks from each application can be rendered separately using a separate audio model and later mixed together to present the sound to the user (e.g., including one or more audio signals that render the sound). However, completely separate rendering pipelines can be computationally expensive and, in some cases, may double the computational resources required to render audio tracks from a single application. As more applications run concurrently, computational resources can become even more strained, and rendering audio tracks from, for example, ten different concurrently running applications using separate custom audio models may be computationally infeasible. Therefore, it may be desirable to design systems and methods to accommodate MR applications that can utilize different audio models running concurrently.

[0105] Figure 8 A model management architecture according to some embodiments is illustrated. In some embodiments, MR system 803 (which may correspond to MR system 112, 200) may include one or more computer systems configured to execute instructions. MR system 803 may include audio service module 802 (which may correspond to audio service 522). Audio service module 802 may manage audio rendering and may include one or more computer systems configured to execute one or more processes, subprocesses, threads, and / or services. In some embodiments, audio service module 802 may be configured to execute services (e.g., background services) as part of the operating system of one or more computer systems.

[0106] In some embodiments, a standalone system (e.g., a third-party application) may request the presentation of an audio signal (e.g., via one or more speakers of the MR system 112, 200). Such a request may take any suitable form. In some embodiments, the request to present an audio signal may include one or more software instructions to present an audio signal; in some embodiments, such a request may be hardware-driven. The request may be made with or without user involvement. Further, such requests may be received via local hardware (e.g., the MR system itself), via external hardware (e.g., a separate computer system communicating with the MR system), via the Internet (e.g., via a cloud server), or via any other suitable source or combination of sources. In some embodiments, audio service 802 may receive the request, render the requested audio signal (e.g., via one or more rendering algorithms that may utilize an audio model managed by model layer 804), and present the requested audio signal to the user. In some embodiments, audio service 802 may be configured to execute a continuously running (e.g., in the background) process while the MR system's operating system is running. In some embodiments, audio service 802 may be configured to execute an instance of a parent background service, which may serve as the host process for one or more background processes and / or child processes. In some embodiments, audio service 802 may be configured to execute instructions that are part of the operating system of the MR system. In some embodiments, applications that can run on the MR system can access audio service 802. In some embodiments, the user of the MR system may not provide input directly to audio service 802. For example, the user may provide input (e.g., movement commands) to an application running on the MR system (e.g., a role-playing game). The application may provide input to audio service 802 (e.g., to render footsteps), and audio service 802 may provide output (e.g., an audio signal including the rendered footsteps) to the user (e.g., via a speaker) and / or other processes and / or services.

[0107] In some embodiments, audio service 802 may include model layer 804, which may be configured to manage one or more audio models. In some embodiments, model layer 804 may include one or more computer systems configured to execute one or more processes, subprocesses, threads, and / or services. In some embodiments, model layer 804 may include one or more computer systems configured to store information. For example, model layer 804 may store one or more audio models 806a, 806b, and / or 806c. In some embodiments, an audio model may include one or more abstract representations of data for rendering audio. In some embodiments, an audio model may include one or more data structures configured to store information. In some embodiments, an audio model (e.g., audio model 806a) may include one or more audio model components (e.g., components 808a and 808b). In some embodiments, audio model components may include one or more data structures configured to store information.

[0108] In some embodiments, an audio model component may store one or more abstract representations of data used to render audio. For example, an audio model component may include data about one or more real and / or virtual objects. In some embodiments, an audio model component may include location data about one or more real and / or virtual objects. In some embodiments, an audio model component may include material data (e.g., audio reflectivity, transmittance, absorption, scattering, diffraction, etc.) about one or more real and / or virtual objects. In some embodiments, an audio model component may represent a real and / or virtual object and its associated properties.

[0109] In some embodiments, the audio model may include data about one or more virtual sound sources. For example, the audio model may include location data about one or more virtual sound sources. In some embodiments, an audio model component may represent a virtual sound source and its associated characteristics.

[0110] In some embodiments, the audio model may include audio parameter data (e.g., sound radiation characteristics, volume, etc.) about one or more virtual sound sources. In some embodiments, the audio model may include one or more audio parameters of the environment. For example, the audio model may include the size of the environment (e.g., the room the user is occupying). In some embodiments, the audio model may include the reverberation time and / or reverberation gain of the environment. In some embodiments, an audio model component may represent an audio parameter of the environment.

[0111] In some embodiments, model layer 804 may be configured to manage different audio models. For example, model 806a may be considered a formal model (e.g., model 806a may be configured to represent realistic physical interactions between virtual sounds and real and / or virtual objects in the user's environment). In some embodiments, model 806a may include model component 808a representing a virtual wall (which may reside in the user's environment). In some embodiments, model 806a may include model component 808b representing reverberation time (e.g., in the user's environment). In some embodiments, model 806a may be used by application 814a, which may be configured to run on MR system 803. In some embodiments, application 814a may be an MR application that has requested to present virtual audio to a user. In some embodiments, the association between application 814a and model 806a may be stored in model layer 804 (e.g., via metadata associated with model 806a). In some embodiments, application 814a may store data (e.g., model identifiers), and application 814a may specify the use of this data to use a particular audio model.

[0112] In some embodiments, model layer 804 may be configured to manage model 806b. In some embodiments, model 806b may be different from model 806a. For example, model 806b may include model component 808c, which may correspond to and / or be the same as model component 808a (e.g., model components 808a and 808c may both represent the same virtual wall). However, model 806b may include model component 808d, which may not correspond to and / or be different from model component 808b. For example, model component 806d may represent a second reverberation time, which may be a longer reverberation time than the reverberation time corresponding to model component 806b.

[0113] In some embodiments, model 806b may be a fully audio model. For example, model 806b may be used to render audio without depending on other models (e.g., model 806a). In some embodiments, model 806b may include one or more dependencies on other models. For example, model 806b may include one or more pointers to model component 808a in model 806a (e.g., because model component 808c corresponds to model component 808a). In some embodiments, storing a dependent audio model may place lower requirements on memory and / or storage devices compared to storing a completely independent audio model.

[0114] In some embodiments, model 806c may not share components with models 806a and / or 806b. For example, model component 808e may represent a virtual sofa, model component 808f may represent a virtual whiteboard, and model component 808g may represent a third reverberation time, which may be shorter than the second and first reverberation times associated with model components 808b and 808d, respectively. In some embodiments, model 806c may be associated with application 814b, which may be configured to run on MR system 803.

[0115] although Figure 8 Model layer 804 is depicted as managing one or more audio models as part of audio service 802, but other embodiments are also contemplated. In some embodiments, each application may manage its own one or more audio models. For example, application 814a may store model 806a locally within application 814a. In some embodiments, model 806a may be a formal model, and application 814a may store references to model 806a in model layer 804. In some embodiments, application 814b may store model 806c locally, and model 806b may not be stored in model layer 804. In some embodiments, model layer 804 may generate audio models. For example, a user of an MR system may change the physical environment, and a new audio model may more accurately represent the acoustic characteristics of the new environment. In some embodiments, an application (e.g., application 814a) may request the generation of a new audio model because the application may utilize one or more customized and / or fully customized audio models.

[0116] In some embodiments, model layer 804 may store audio models associated with applications currently running on MR system 803. For example, MR system 803 may currently only be running applications 814a and 814b, and model 806b may not be associated with application 814a or 814b. In some embodiments, model 806b may be removed from model layer 804 (e.g., to save memory). In some embodiments, formal models may be stored contiguously in model layer 804. In some embodiments, model layer 804 may remove an audio model from memory if a threshold amount of time has not been used. In some embodiments, an application may request model layer 804 to retain an audio model in memory.

[0117] In some embodiments, audio service 802 may include a rendering layer 812, which may be configured to manage one or more audio models. In some embodiments, rendering layer 812 (which may correspond to rendering algorithm module 520) may include one or more computer systems configured to execute one or more processes, subprocesses, threads, and / or services. Rendering layer 812 may render audio from one or more inputs (e.g., model components, source information, etc.). In some embodiments, rendering layer 812 may render an audio stream and may render the audio stream in real time. In some embodiments, rendering layer 812 may render audio from a fixed amount of input data. The rendered audio may be presented to the user as one or more audio signals including the rendered audio. The audio signals may be delivered via one or more speakers of the device (e.g., Figure 4 The speakers 412 and 414 shown are presented.

[0118] although Figure 8 Three models with different numbers of model components are depicted, but it should be envisioned that any number of models can be managed by the audio service 802. It is also envisioned that any model can include any number of model components.

[0119] Figure 9A A rendering architecture according to some embodiments is illustrated. In some embodiments, rendering layer 902 (which may correspond to rendering layer 812) may include one or more computer systems configured to execute one or more processes, subprocesses, threads, and / or services. In some embodiments, rendering layer 902 may be configured to render audio from one or more inputs (e.g., model components, source information, etc.). In some embodiments, rendering layer 902 may render an audio stream, and may render the audio stream in real time. In some embodiments, rendering layer 902 may render audio from a fixed amount of input data.

[0120] Rendering layer 902 may include direct layer 904. In some embodiments, direct layer 904 may include one or more computer systems configured to execute processes, subprocesses, threads, and / or services. In some embodiments, direct layer 904 may be configured to render the direct path between a sound source and a listening node (e.g., a user). Direct layer 904 may determine whether sound should be occluded (e.g., because real and / or virtual objects obstruct the direct path from the sound source to the listening node). In some embodiments, direct layer 904 may determine whether sound should be attenuated (e.g., due to the distance between the sound source and the listening node).

[0121] Rendering layer 902 may include reflection layer 906. In some embodiments, reflection layer 906 may include one or more computer systems configured to execute processes, subprocesses, threads, and / or services. In some embodiments, reflection layer 906 may be configured to render one or more reflected audio paths between a sound source and a listening node (e.g., a user). For example, sound from a sound source may reflect off a wall before reaching the listening node, and reflection layer 906 may render one or more effects of the reflection (e.g., the reflected sound may be delayed compared to the direct sound, and the reflected sound may be attenuated compared to the direct sound). In some embodiments, reflection layer 906 may be configured to apply one or more filters to audio to approximate the behavior of reflected sound in a general environment. In some embodiments, reflection layer 906 may be configured for audio ray tracing, which may compute reflection paths for one or more audio rays.

[0122] Rendering layer 902 may include reverb layer 908. In some embodiments, reverb layer 908 may include one or more computer systems configured to execute processes, subprocesses, threads, and / or services. In some embodiments, reverb layer 908 may be configured to render the reverb behavior of audio (e.g., post-reverb behavior). In some embodiments, post-reverb behavior may be influenced and / or determined by environmental characteristics such as the reverb time and / or reverb gain of the environment.

[0123] Rendering layer 902 may include virtualizer 910. In some embodiments, virtualizer 910 may include one or more computer systems configured to execute processes, subprocesses, threads, and / or services. In some embodiments, virtualizer 910 may be configured to present audio to one or more virtual speakers. For example, an MR system may present audio to a user originating from one of six virtual speakers arranged in a speaker array around the user's head. In some embodiments, virtualizer 910 may be configured to determine which sound or combination of sounds should be played at which virtual speaker or speaker combination in the virtual speaker array.

[0124] exist Figure 9AIn the example embodiment shown, audio streams 912 and 914 can be rendered by rendering layer 902. In some embodiments, audio stream 912 may correspond to an audio stream originating from a first application, and audio stream 914 may correspond to an audio stream originating from a second application. In some embodiments, the first application may use the same audio model as the second application, and the first and second applications may run simultaneously on an MR system. In some embodiments, rendering audio streams 912 and 914 together may be more efficient than rendering them separately (e.g., because the two audio streams may depend on the same audio model). For example, the direct paths corresponding to audio streams 912 and 914 can be rendered and passed to virtualizer 910. In some embodiments, the same direct path calculation can be used to render audio streams 912 and 914 (e.g., because the same real and / or virtual objects may obstruct the direct path between one or more sound sources and one or more listening nodes).

[0125] In some embodiments, audio streams 912 and 914 can be rendered together via reflection layer 906. In some embodiments, reflection layer 906 can receive one or more input audio streams from direct layer 904. In some embodiments, reflection layer 906 can receive one or more input audio streams directly from the application (e.g., without first passing through direct layer 904). In some embodiments, rendering audio streams 912 and 914 together may be more efficient than rendering the audio streams separately. For example, the same set of filters can be applied to both audio streams (and / or a single audio stream corresponding to a mixture of the two audio streams) to approximate reflection behavior. In some embodiments, the rendered reflections can be passed to virtualizer 910.

[0126] In some embodiments, audio streams 912 and 914 can be rendered together via a reverb layer 908. In some embodiments, the reverb layer 908 can receive one or more input audio streams from a reflection layer 906 and / or a direct layer 904. In some embodiments, the reverb layer 908 can receive one or more input audio streams directly from the application (e.g., without going through the reflection layer 906 and / or the direct layer 904). In some embodiments, rendering audio streams 912 and 914 together may be more efficient than rendering them separately. For example, post-reverb behavior can be determined based on a shared audio model, and post-reverb behavior can be applied to both audio streams. In some embodiments, the rendered reverb behavior can be passed to the virtualizer 910.

[0127] Figure 9BA rendering architecture according to some embodiments is illustrated. In some embodiments, multiple audio streams depending on different audio models can be rendered together. For example, audio stream 912 may originate from a first application using a first audio model, and audio stream 914 may originate from a second application using a second audio model. In some embodiments, the first audio model may include model components corresponding to virtual objects that may not exist in the second audio model (e.g., because the first application has introduced virtual objects that the second application may not utilize). In some embodiments, virtual objects introduced by the first application may affect the direct rendering path for audio stream 912 (e.g., direct sound may be occluded due to obstacles), but may not affect the direct rendering path for audio stream 914. In some embodiments, audio streams 912 and 914 may be rendered in separate instances of direct layer 904. In some embodiments, audio streams 912 and 914 may each be rendered independently at direct layer 904.

[0128] In some embodiments, other model components may be shared between audio stream 912 and audio stream 914. For example, virtual objects may be configured not to interfere with audio reflections, and audio streams 912 and 914 may be rendered using the same reflection calculations at reflection layer 906. In some embodiments, virtual objects may not affect the post-reverberation of the MRE, and audio streams 912 and 914 may be rendered using the same reverberation calculations at reverberation layer 908. In some embodiments, audio streams 912 and 914 may be blended together at virtualizer 910.

[0129] In some embodiments, efficiency can be utilized when audio streams 912 and 914 depend on shared model components (even if one or more differences may exist in the model components). For example, reflection, reverberation, and virtualizer calculations can be shared across audio streams, even if the direct path might require independent calculations across the two audio streams.

[0130] although Figure 9B Rendering architectures with different direct path instances are illustrated, but other embodiments are also envisioned. For example, audio stream 912 could correspond to a first application that modifies the virtual object material composition to better reflect the audio (in contrast to audio stream 914, which could correspond to a second application that does not modify the virtual object material composition). In some embodiments, audio streams 912 and 914 could be rendered together at direct layer 904, but could be rendered separately (e.g., in separate instances) at reflection layer 906. In some embodiments, audio streams 912 and 914 could again be rendered together at reverberation layer 908 (e.g., because the model component corresponding to post-reverberation is shared between the audio models on which audio streams 912 and 914 depend).

[0131] although Figures 9A to 9B A rendering architecture including a direct layer 904, a reflection layer 906, a reverberation layer 908, and a virtualizer 910 is shown, but other architectures may also be used. In some embodiments, one or more layers may not be included in the rendering architecture (e.g., due to computational limitations). In some embodiments, one or more layers may be added to simulate acoustic behavior more realistically.

[0132] Example systems, methods, and computer-readable media are disclosed. According to some examples, a system includes: one or more speakers; and one or more processors configured to perform a method comprising: receiving a request to render a first audio track, wherein the first audio track is based on a first audio model including a shared model component and a second model component; receiving a request to render a second audio track, wherein the second audio track is based on a second audio model including a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and rendering an audio signal including the rendered sound via the one or more speakers. In some examples, rendering sound includes: rendering the first audio track and the second audio track based on the shared model component. In some examples, rendering sound includes: rendering the first audio track based on the first audio component. In some examples, rendering sound includes: rendering the second audio track based on the second audio component. In some examples, the one or more speakers are speakers of a head-mounted device. In some examples, the first audio model is based on the physical environment of the head-mounted device. In some examples, the shared model component corresponds to direct path acoustic behavior. In some examples, the shared model component corresponds to reflective acoustic behavior. In some examples, the shared model component corresponds to reverberant acoustic behavior.

[0133] According to some examples, a method includes: receiving a request to render a first audio track, wherein the first audio track is based on a first audio model including a shared model component and a second model component; receiving a request to render a second audio track, wherein the second audio track is based on a second audio model including a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and rendering an audio signal including the rendered sound via one or more speakers. In some examples, rendering the sound includes: rendering the first and second audio tracks based on the shared model component. In some examples, rendering the sound includes: rendering the first audio track based on the first audio component. In some examples, rendering the sound includes: rendering the second audio track based on the second audio component. In some examples, rendering the audio signal includes: rendering the audio signal via one or more speakers of a head-mounted device. In some examples, the first audio model is based on the physical environment of the head-mounted device. In some examples, the shared model component corresponds to direct path acoustic behavior. In some examples, the shared model component corresponds to reflection acoustic behavior. In some examples, the shared model component corresponds to reverberant acoustic behavior.

[0134] According to some examples, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause one or more processors to perform a method comprising: receiving a request to render a first audio track, wherein the first audio track is based on a first audio model including a shared model component and a second model component; receiving a request to render a second audio track, wherein the second audio track is based on a second audio model including a shared model component and a second model component; rendering sound based on the first audio track, the second audio track, the shared model component, the first model component, and the second model component; and rendering an audio signal including the rendered sound via one or more speakers. In some examples, rendering sound includes: rendering the first and second audio tracks based on the shared model component. In some examples, rendering sound includes: rendering the first audio track based on the first audio component. In some examples, rendering sound includes: rendering the second audio track based on the second audio component. In some examples, rendering the audio signal includes: rendering the audio signal via one or more speakers of a head-mounted device. In some examples, the first audio model is based on the physical environment of the head-mounted device. In some examples, the shared model component corresponds to direct path acoustic behavior. In some examples, the shared model component corresponds to reflective acoustic behavior. In some examples, the shared model component corresponds to reverberant acoustic behavior.

[0135] While the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. For example, elements of one or more embodiments may be combined, deleted, modified, or supplemented to form further embodiments. Such changes and modifications will be understood to be included within the scope of the disclosed examples as defined by the appended claims.

Claims

1. A system for managing and storing audio data, comprising: Head-mounted devices; One or more speakers; as well as One or more processors are configured to perform a method for managing and storing audio data, the method comprising: The audio service receives a request from a first application associated with the headset to present a first audio track; Based on the location of the head-mounted device in the mixed reality environment, an audio model including shared model components is updated, wherein the audio model is associated with the acoustic properties of the mixed reality environment. The updates mentioned include: Determine whether to update the copy of the audio model; Update the copy of the audio model based on the determination of the copy to be updated; Determine the first audio model in the audio model corresponding to the first application; The first sound is rendered based on the first audio track and also based on the first audio model that includes the shared model component and further includes the first model component; The audio service receives a request to present a second audio track from a second application associated with the headset. Determine the second audio model in the audio model corresponding to the second application; The second sound is rendered based on the second audio track and also based on the second audio model that includes the shared model component and further includes the second model component; Based on the determination of updating the copy of the audio model, if the copy of the audio model is not in use, then the copy of the audio model is updated; and An audio signal including the first sound and the second sound is presented via the one or more speakers.

2. The system according to claim 1, wherein, The one or more speakers are speakers of the head-mounted device.

3. The system according to claim 2, wherein, The first audio model is based on the physical environment of the head-mounted device.

4. The system according to claim 1, wherein, The shared model components correspond to direct path acoustic behavior.

5. The system according to claim 1, wherein, The shared model components correspond to reflective acoustic behavior.

6. The system according to claim 1, wherein, The shared model components correspond to reverberant acoustic behavior.

7. A method for managing and storing audio data, comprising: The audio service receives a request from the first application associated with the head-mounted device to present the first audio track; Based on the location of the head-mounted device in the mixed reality environment, an audio model including shared model components is updated, wherein the audio model is associated with the acoustic properties of the mixed reality environment. The updates mentioned include: Determine whether to update the copy of the audio model; Update the copy of the audio model based on the determination of the copy to be updated; Determine the first audio model in the audio model corresponding to the first application; The first sound is rendered based on the first audio track and also based on the first audio model that includes the shared model component and further includes the first model component; The audio service receives a request to present a second audio track from a second application associated with the headset. Determine the second audio model in the audio model corresponding to the second application; The second sound is rendered based on the second audio track and also based on the second audio model that includes the shared model component and further includes the second model component; Based on the determination of updating the copy of the audio model, if the copy of the audio model is not in use, then the copy of the audio model is updated; and An audio signal comprising the first sound and the second sound is presented via one or more speakers of the head-mounted device.

8. The method according to claim 7, wherein, The first audio model is based on the physical environment of the head-mounted device.

9. The method according to claim 7, wherein, The shared model components correspond to direct path acoustic behavior.

10. The method according to claim 7, wherein, The shared model components correspond to reflective acoustic behavior.

11. The method according to claim 7, wherein, The shared model components correspond to reverberant acoustic behavior.

12. A non-transitory computer-readable medium storing instructions, said instructions, when executed by one or more processors, causing said one or more processors to perform a method for managing and storing audio data, said method comprising: The audio service receives a request from the first application associated with the head-mounted device to present the first audio track; Based on the location of the head-mounted device in the mixed reality environment, an audio model including shared model components is updated, wherein the audio model is associated with the acoustic properties of the mixed reality environment. The updates mentioned include: Determine whether to update the copy of the audio model; Update the copy of the audio model based on the determination of the copy to be updated; Determine the first audio model in the audio model corresponding to the first application; The first sound is rendered based on the first audio track and also based on the first audio model that includes the shared model component and further includes the first model component; The audio service receives a request to present a second audio track from a second application associated with the headset. Determine the second audio model in the audio model corresponding to the second application; The second sound is rendered based on the second audio track and also based on the second audio model that includes the shared model component and further includes the second model component; Based on the determination of updating the copy of the audio model, if the copy of the audio model is not in use, then the copy of the audio model is updated; and An audio signal comprising the first sound and the second sound is presented via one or more speakers of the head-mounted device.

13. The non-transitory computer-readable medium according to claim 12, wherein, The first audio model is based on the physical environment of the head-mounted device.

14. The non-transitory computer-readable medium according to claim 12, wherein, The shared model components correspond to direct path acoustic behavior.

15. The non-transitory computer-readable medium according to claim 12, wherein, The shared model components correspond to reflective acoustic behavior.

16. The non-transitory computer-readable medium according to claim 12, wherein, The shared model components correspond to reverberant acoustic behavior.

Citation Information

Patent Citations

  • Rendering operations using sparse volumetric data

    US20190180499A1

  • Methods, apparatus, systems, computer programs for enabling consumption of virtual content for mediated reality

    WO2019002668A1