Mixed Reality Spatial Audio
By determining acoustic parameters based on the user's real-world environment and applying a transfer function to audio signals, the system addresses auditory inconsistencies in XR systems, creating a more immersive and realistic mixed reality experience.
Patent Information
- Application Number
- JP2024059444
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-15
- Filing Date
- 2024-04-02
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2038-10-17
AI Technical Summary
Existing XR systems fail to present virtual audio that accurately reflects the user's real-world environment, leading to auditory inconsistencies that can cause discomfort and motion sickness, as they do not account for the acoustic properties of the user's surroundings.
A system and method that determines acoustic parameters based on the user's real-world environment, applying a transfer function to audio signals to simulate how sound behaves in that environment, enhancing the realism and interactivity of mixed reality experiences.
This approach creates a more immersive and realistic mixed reality experience by simulating audio that aligns with the user's real-world acoustic environment, reducing discomfort and enhancing user engagement.
Smart Images

Figure 0007770456000001 
Figure 0007770456000002 
Figure 0007770456000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 62 / 573,448, filed October 17, 2017, and U.S. Provisional Patent Application No. 62 / 631,418, filed February 15, 2018, the contents of both of which are incorporated herein by reference in their entirety for all purposes.
[0002] The present disclosure relates generally to systems and methods for presenting audio signals, and more particularly to systems and methods for presenting audio signals to a user of a mixed reality environment. [Background technology]
[0003] Virtual environments are ubiquitous in computing environments, finding use in video games (where a virtual environment may represent a game world), maps (where a virtual environment may represent a terrain to be navigated), simulations (where a virtual environment may simulate a real environment), digital storytelling (where virtual characters may interact with one another within a virtual environment), and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, a user's experience with a virtual environment may be limited by the technology for presenting the virtual environment. For example, traditional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may be unable to realize a virtual environment in a way that creates a convincingly realistic and immersive experience.
[0004] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively “XR”) share the ability to present sensory information corresponding to a virtual environment represented by data within a computer system to a user of the XR system. Such systems can provide a uniquely enhanced sense of immersion and realism by combining virtual visual and audio cues with actual visual and audio. It may therefore be desirable to present digital audio to a user of an XR system so that the audio appears to be occurring naturally and consistently within the user's real environment and in accordance with the user's audio expectations. Generally speaking, users expect virtual audio to take on the acoustic qualities of the real environment in which they are heard. For example, a user of an XR system in a large concert hall would expect the XR system's virtual audio to have a spacious, cavernous quality; conversely, a user in a small apartment would expect the audio to be more attenuated, closer, and more immediate.
[0005] Existing technologies often fall short of these expectations by presenting virtual audio that does not take into account the user's surroundings, leading to a sense of inauthenticity that can compromise the user experience. Observations of users of XR systems indicate that while users can be relatively tolerant of visual inconsistencies between virtual content and their real-world environment (e.g., lighting inconsistencies), they can be more sensitive to auditory inconsistencies. Our unique auditory experiences, which are continually refined throughout our lives, can actually make us aware of the extent to which our physical environment influences the sounds we hear, and we can be highly aware of sounds that contradict these expectations. With XR systems, such inconsistencies can be uncomfortable, turning an immersive and compelling experience into a cosmetic imitation. In extreme examples, auditory inconsistencies can cause motion sickness and other adverse effects due to the inner ear's inability to reconcile auditory stimuli with their corresponding visual cues.
[0006] The present invention is directed to addressing these shortcomings by presenting virtual audio to a user with an audio presentation that incorporates one or more playback parameters based on aspects of the user's real-world environment. For example, the presentation can incorporate simulated reverberation effects, in which one or more parameters of the reverberation depend on attributes of the user's real-world environment, such as the volume of the room or the material of the room's walls. By considering the characteristics of the user's physical environment, the systems and methods described herein can simulate what the user would hear if the virtual audio were an actual audio naturally generated in that environment. By presenting virtual audio in a manner that is faithful to the way audio behaves in the real world, the user may experience an enhanced sense of connection to the mixed reality environment. Similarly, by presenting location-aware virtual content that responds to the user's movement and environment, the content becomes more subjective, interactive, and realistic; for example, a user's experience at point A may be completely different from their experience at point B. This enhanced realism and interactivity can provide the basis for new applications of mixed reality, such as those that use spatially aware audio and enable novel forms of gameplay, social features, or interactive behaviors. Summary of the Invention [Means for solving the problem]
[0007] A system and method for presenting an audio signal to a user of a mixed reality environment is disclosed. According to an exemplary method, an audio event associated with the mixed reality environment is detected. The audio event is associated with a first audio signal. A location of the user relative to the mixed reality environment is determined. An acoustic region associated with the user's location is identified. First acoustic parameters associated with the first acoustic region are determined. A transfer function is determined using the first acoustic parameters. The transfer function is applied to the first audio signal to generate a second audio signal, which is then presented to the user. The present specification also provides, for example, the following items: (Item 1) 1. A method for presenting an audio signal to a user of a mixed reality environment, the method comprising: Detecting an audio event associated with the mixed reality environment, the audio event being associated with a first audio signal; determining a location of the user relative to the mixed reality environment; identifying a first acoustic region associated with the user's location; determining a first acoustic parameter associated with the first acoustic region; determining a transfer function using the first acoustic parameters; applying the transfer function to the first audio signal to generate a second audio signal; presenting the second audio signal to the user; and A method comprising: (Item 2) Item 10. The method of item 1, wherein the first audio signal comprises a waveform audio file. (Item 3) Item 10. The method of claim 1, wherein the first audio signal comprises a live audio stream. (Item 4) the user is associated with a wearable system comprising one or more sensors and a display configured to present a view of the mixed reality environment; 2. The method of claim 1, wherein identifying the first acoustic region includes detecting a first sensor input from the one or more sensors and identifying the first acoustic region based on the first sensor input. (Item 5) 5. The method of claim 4, wherein determining the first acoustic parameter includes detecting a second sensor input from the one or more sensors and determining that the first acoustic parameter is based on the second sensor input. (Item 6) 6. The method of claim 5, wherein determining the first acoustic parameter further includes identifying a geometric characteristic of the acoustic field based on the second sensor input and determining that the first acoustic parameter is based on the geometric characteristic. (Item 7) 6. The method of claim 5, wherein determining the first acoustic parameter further includes identifying an associated material of the acoustic region based on the second sensor input and determining that the first acoustic parameter is based on the material. (Item 8) Item 5. The method of item 4, wherein the wearable system further comprises a microphone, the microphone being located within the first acoustic region, and the first acoustic parameter being determined based on a signal detected by the microphone. (Item 9) Item 10. The method of item 1, wherein the first acoustic parameter corresponds to a reverberation parameter. (Item 10) Item 10. The method of item 1, wherein the first acoustic parameter corresponds to a filtering parameter. (Item 11) identifying a second acoustic region, the second acoustic region acoustically coupled to the first acoustic region; determining a second acoustic parameter associated with the second acoustic region; and further comprising Item 10. The method of item 1, wherein the transfer function is determined using the second acoustic parameters. (Item 12) 1. A system comprising: 1. A wearable headgear unit, comprising: a display configured to present a view of the mixed reality environment; A speaker and one or more sensors; A circuit, the circuit comprising: Detecting an audio event associated with the mixed reality environment, the audio event being associated with a first audio signal; determining a location of the wearable headgear unit relative to the mixed reality environment based on the one or more sensors; and identifying a first acoustic region associated with the user's location; determining a first acoustic parameter associated with the first acoustic region; determining a transfer function using the first acoustic parameters; applying the transfer function to the first audio signal to generate a second audio signal; presenting the second audio signal to the user via the speaker; a circuit configured to perform a method including: a wearable headgear unit, A system comprising: (Item 13) Item 13. The system of item 12, wherein the first audio signal comprises a waveform audio file. (Item 14) Item 13. The system of item 12, wherein the first audio signal comprises a live audio stream. (Item 15) Item 13. The system of item 12, wherein identifying the first acoustic region includes detecting a first sensor input from the one or more sensors and identifying the first acoustic region based on the first sensor input. (Item 16) Item 16. The system of item 15, wherein determining the first acoustic parameter includes detecting a second sensor input from the one or more sensors and determining that the first acoustic parameter is based on the second sensor input. (Item 17) Item 17. The system of item 16, wherein determining the first acoustic parameter further includes identifying a geometric characteristic of the acoustic field based on the second sensor input and determining that the first acoustic parameter is based on the geometric characteristic. (Item 18) 17. The system of claim 16, wherein determining the first acoustic parameter further includes identifying an associated material of the acoustic region based on the second sensor input and determining that the first acoustic parameter is based on the material. (Item 19) Item 16. The system of item 15, wherein the first acoustic parameter is determined based on a signal detected by the microphone. (Item 20) Item 13. The system of item 12, wherein the first acoustic parameter corresponds to a reverberation parameter. (Item 21) Item 13. The system of item 12, wherein the first acoustic parameter corresponds to a filtering parameter. (Item 22) identifying a second acoustic region, the second acoustic region acoustically coupled to the first acoustic region; determining a second acoustic parameter associated with the second acoustic region; and further comprising Item 13. The system of item 12, wherein the transfer function is determined using the second acoustic parameters. (Item 23) 1. An augmented reality system, comprising: a localization subsystem configured to determine an identity of a first space in which the augmented reality system is located; and a communications subsystem configured to communicate an identification of the first space in which the augmented reality system is located and further configured to receive audio parameters associated with the first space; an audio output subsystem configured to process an audio segment based on the audio parameters and further configured to output the audio segment; and An augmented reality system comprising: (Item 24) 1. An augmented reality system, comprising: a sensor subsystem configured to determine information associated with an acoustic property of a first space corresponding to a location of the augmented reality system; and an audio processing subsystem configured to process an audio segment based on the information, the audio processing subsystem communicatively coupled to the sensor subsystem; and a speaker for presenting the audio segments, the speaker being coupled to the audio processing subsystem; and An augmented reality system comprising: (Item 25) Item 25. The augmented reality system of item 24, wherein the sensor subsystem is further configured to determine geometric information for the first space. (Item 26) Item 25. The augmented reality system of item 24, wherein the sensor subsystem comprises a camera. (Item 27) Item 25. The augmented reality system of item 24, wherein the sensor subsystem comprises an object recognition device configured to recognize objects having distinct sound absorption properties. (Item 28) Item 25. The augmented reality system of item 24, wherein the sensor subsystem comprises a microphone. [Brief explanation of the drawings]
[0008] [Figure 1A] 1A-1C illustrate an example mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 1B] 1A-1C illustrate an example mixed reality environment in accordance with one or more embodiments of the present disclosure. [Figure 1C] 1A-1C illustrate an example mixed reality environment in accordance with one or more embodiments of the present disclosure.
[0009] [Figure 2] FIG. 2 illustrates an example wearable head unit of an example mixed reality system in accordance with one or more embodiments of the present disclosure.
[0010] [Figure 3A] FIG. 3A illustrates an example mixed reality handheld controller that can be used to provide input to a mixed reality environment, in accordance with one or more embodiments of the present disclosure.
[0011] [Figure 3B] FIG. 3B illustrates an example auxiliary unit that may be included in an example mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0012] [Figure 4] FIG. 4 illustrates an example functional block diagram of an example mixed reality system in accordance with one or more embodiments of the present disclosure.
[0013] [Figure 5] FIG. 5 illustrates an example configuration of components of an example mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0014] [Figure 6] FIG. 6 illustrates a flowchart of an example process for presenting an audio signal in a mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0015] [Figure 7] 7-8 illustrate flowcharts of example processes for determining acoustic parameters of a room in a mixed reality system, in accordance with one or more embodiments of the present disclosure. [Figure 8] 7-8 illustrate flowcharts of example processes for determining acoustic parameters of a room in a mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0016] [Figure 9]FIG. 9 illustrates an example of an acoustically coupled room in a mixed reality environment, in accordance with one or more embodiments of the present disclosure.
[0017] [Figure 10] FIG. 10 illustrates an example of an acoustic graph structure in accordance with one or more embodiments of the present disclosure.
[0018] [Figure 11] FIG. 11 illustrates a flowchart of an example process for determining composite acoustic parameters of an acoustic environment of a mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0019] [Figure 12] 12-14 illustrate components of an example wearable mixed reality system in accordance with one or more embodiments of the present disclosure. [Figure 13] 12-14 illustrate components of an example wearable mixed reality system in accordance with one or more embodiments of the present disclosure. [Figure 14] 12-14 illustrate components of an example wearable mixed reality system in accordance with one or more embodiments of the present disclosure.
[0020] [Figure 15] FIG. 15 illustrates an example configuration of components of an example mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0021] [Figure 16] 16-20 illustrate flowcharts of example processes for presenting an audio signal to a user of a mixed reality system, in accordance with one or more embodiments of the present disclosure. [Figure 17] 16-20 illustrate flowcharts of example processes for presenting an audio signal to a user of a mixed reality system, in accordance with one or more embodiments of the present disclosure. [Figure 18]16-20 illustrate flowcharts of example processes for presenting an audio signal to a user of a mixed reality system, in accordance with one or more embodiments of the present disclosure. [Figure 19] 16-20 illustrate flowcharts of example processes for presenting an audio signal to a user of a mixed reality system, in accordance with one or more embodiments of the present disclosure. [Figure 20] 16-20 illustrate flowcharts of example processes for presenting an audio signal to a user of a mixed reality system, in accordance with one or more embodiments of the present disclosure.
[0022] [Figure 21] FIG. 21 illustrates a flowchart of an example process for determining a location of a user of a mixed reality system, in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0023] In the following description of the embodiments, reference is made to the accompanying drawings which form a part hereof, and in which is shown, by way of illustration, specific embodiments which may be practiced. It is to be understood that other embodiments may be used and structural changes may be made without departing from the scope of the disclosed embodiments.
[0024] Mixed Reality Environment
[0025] Like all people, users of mixed reality systems exist in a real environment, i.e., the three-dimensional portion of the "real world" and all of its content that is perceptible by the user. For example, users perceive the real environment using their normal human senses, i.e., sight, sound, touch, taste, and smell, and interact with the real environment by moving their own body within the real environment. Locations within the real environment can be described as coordinates within a coordinate space; for example, coordinates can comprise latitude, longitude, and altitude relative to sea level, distance in three orthogonal dimensions from a reference point, or other suitable values. Similarly, vectors can describe quantities that have direction and magnitude within the coordinate space.
[0026] A computing device may maintain a representation of a virtual environment, for example, in a memory associated with the device. As used herein, a virtual environment is a computerized representation of a three-dimensional space. A virtual environment may include representations of objects, actions, signals, parameters, coordinates, vectors, or other properties associated with that space. In some examples, the circuitry (e.g., a processor) of the computing device may maintain and update the state of the virtual environment; for example, the processor may determine the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or input provided by a user at a first time t0. For example, if an object in the virtual environment is located at a first coordinate at time t0 and has certain programmed physical parameters (e.g., mass, coefficient of friction), and input received from the user indicates that a force should be applied to the object in a directional vector, the processor may apply the laws of kinematics and use basic mechanics to determine the location of the object at time t1. The processor may use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at time t1. In maintaining and updating the state of the virtual environment, the processor may execute any suitable software, including software related to the creation and deletion of virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling input and output, software for performing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.
[0027] Output devices, such as a display or speakers, can present aspects of the virtual environment to the user. For example, the virtual environment may include virtual objects (which may include representations of objects, people, animals, lights, etc.) that can be visually presented to the user. A processor can determine a field of view of the virtual environment (e.g., corresponding to a camera with an origin coordinate, a viewing axis, and a frustum) and render a viewable scene of the virtual environment corresponding to that field of view on the display. Any suitable rendering technique may be used for this purpose. In some examples, the viewable scene may include only a subset of the virtual objects in the virtual environment and exclude certain other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, a virtual object in the virtual environment may generate spatial audio resulting from the object's location coordinates (e.g., a virtual character may speak or trigger a sound effect), or the virtual environment may be associated with musical cues or ambient sounds that may or may not be associated with a particular location. The processor can determine an audio signal corresponding to a "user" coordinate, e.g., an audio signal corresponding to a composition of sounds in the virtual environment and rendered to simulate the audio signal that would be heard by the user at the user coordinate, and present the audio signal to the user via one or more speakers. In some examples, a user can be associated with two or more listener coordinates, e.g., first and second listener coordinates corresponding to the user's left and right ears, respectively, and audio signals can be rendered separately for each listener coordinate.
[0028] Because the virtual environment exists only as a computational construct, the user cannot directly perceive the virtual environment using their normal senses. Instead, the user can indirectly perceive the virtual environment, for example, as presented to the user by a display, speakers, haptic feedback device, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data via input devices or sensors to a processor, which can use the device or sensor data to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is about to touch an object in the virtual environment, and the processor can use that data to cause the object to respond accordingly in the virtual environment.
[0029] A mixed reality system can present a user with a mixed reality environment (“MRE”) that combines aspects of a real environment and a virtual environment, for example, using a see-through display and / or one or more speakers integrated into a head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of a real environment and a corresponding virtual environment. In some examples, the corresponding real and virtual environments share a single coordinate space, and in some examples, the real coordinate space and the corresponding virtual coordinate space are related to each other by a transformation matrix (or other suitable representation). Thus, a single coordinate (in some examples, together with the transformation matrix) can define a first location in the real environment and a second corresponding location in the virtual environment, and vice versa.
[0030] In an MRE, a virtual object (e.g., in a virtual environment associated with the MRE) may correspond to a real object (e.g., in a real environment associated with the MRE). For example, if the real environment of the MRE comprises a real lamppost (real object) at a location coordinate, the virtual environment of the MRE may comprise a virtual lamppost (virtual object) at a corresponding location coordinate. As used herein, a real object in combination with its corresponding virtual object together constitutes a "mixed reality object." A virtual object need not perfectly match or match the corresponding real object. In some embodiments, a virtual object may be a simplified version of the corresponding real object. For example, if the real environment includes a real lamppost, the corresponding virtual object may comprise a cylinder of approximately the same height and radius as the real lamppost (reflecting that a lamppost may be approximately cylindrical). Simplifying virtual objects in this manner may enable computational efficiency and simplify the calculations to be performed on such virtual objects. Furthermore, in some embodiments of an MRE, not all real objects in the real environment may be associated with a corresponding virtual object. Similarly, in some embodiments of the MRE, not all virtual objects in the virtual environment may be associated with corresponding real-world objects, i.e., some virtual objects may exist solely within the virtual environment of the MRE without any real-world counterpart.
[0031] In some embodiments, virtual objects may have characteristics that differ, sometimes dramatically, from those of the corresponding real objects. For example, while the real environment in an MRE may comprise a green, two-armed cactus, i.e., a thorny, inanimate object, the corresponding virtual object in the MRE may have the characteristics of a green, two-armed virtual character with human facial features and a surly attitude. In this embodiment, the virtual object resembles its corresponding real object in some characteristics (color, number of arms) but differs from the real object in other characteristics (facial features, personality). In this manner, virtual objects have the potential to represent real objects in a creative, abstract, exaggerated, or fantastical manner, or to impart behavior (e.g., human personality) to otherwise inanimate real objects. In some embodiments, a virtual object may be a purely fantasy creation with no real-world counterpart (e.g., a virtual monster in a virtual environment in a location corresponding to empty space in the real environment).
[0032] Compared to VR systems that present a virtual environment to a user while obscuring the real environment, mixed reality systems that present an MRE allow the real environment to remain perceptible while the virtual environment is presented. Thus, a user of a mixed reality system can experience and interact with the corresponding virtual environment using visual and audio cues associated with the real environment. As an example, as noted above, because a user cannot directly perceive or interact with the virtual environment, a user of a VR system may struggle to perceive or interact with virtual objects displayed in the virtual environment, whereas a user of an MR system may find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching corresponding real objects in their own real environment. This level of interactivity can enhance a user's immersion, connection, and engagement with the virtual environment. Similarly, by simultaneously presenting a real environment and a virtual environment, a mixed reality system can reduce negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. Mixed reality systems also offer many potential applications that can augment or alter our experience of the real world.
[0033] 1A illustrates an exemplary real-world environment 100 in which a user 110 uses a mixed reality system 112. The mixed reality system 112 may include, for example, a display (e.g., a see-through display), one or more speakers, and one or more sensors (e.g., a camera), as described below. The illustrated real-world environment 100 includes a rectangular room 104A in which the user 110 is standing and real-world objects 122A (a lamp), 124A (a table), 126A (a sofa), and 128A (a painting). The room 104A further includes a corner 106A, which may be considered the origin of the real-world environment 100. As shown in FIG. 1A, an environmental coordinate system 108 (comprising an x-axis 108X, a y-axis 108Y, and a z-axis 108Z) with its origin at the corner 106A can define a coordinate space for the real-world environment 100. In some examples, user 110 may be considered an actual object in real environment 100, and similarly, body parts (e.g., hands, feet) of user 110 may be considered actual objects in real environment 100. In some examples, a user coordinate system 114 relative to mixed reality system 112 may be defined. This may simplify the representation of locations relative to the user's head or head-worn device. Using SLAM, visual odometry, or other techniques, a transformation between user coordinate system 114 and environment coordinate system 108 may be determined and updated in real time.
[0034] 1B illustrates an exemplary virtual environment 130 corresponding to real environment 100. The illustrated virtual environment 130 includes a virtual rectangular room 104B corresponding to real rectangular room 104A, a virtual object 122B corresponding to real object 122A, a virtual object 124B corresponding to real object 124A, and a virtual object 126B corresponding to real object 126A. Metadata associated with virtual objects 122B, 124B, and 126B may include information derived from the corresponding real objects 122A, 124A, and 126A. The virtual environment 130 additionally includes a virtual monster 132 that does not correspond to any real object in real environment 100. Similarly, real object 128A in real environment 100 does not correspond to any virtual object in virtual environment 130. Virtual room 104B includes a corner 106B that corresponds to corner 106A of real room 104A and may be considered the origin of virtual environment 130. As shown in FIG. 1B, a coordinate system 108 (comprising an x-axis 108X, a y-axis 108Y, and a z-axis 108Z) with its origin at corner 106B may define a coordinate space for virtual environment 130.
[0035] 1A and 1B, coordinate system 108 defines a shared coordinate space for both real environment 100 and virtual environment 130. In the illustrated embodiment, the coordinate space has its origin at corner 106A in real environment 100 and corner 106B in virtual environment 130. Furthermore, the coordinate space is defined by three orthogonal axes (108X, 108Y, 108Z) that are identical in both real environment 100 and virtual environment 130. Thus, a first location in real environment 100 and a second corresponding location in virtual environment 130 can be described with respect to the same coordinate space. This simplifies identifying and displaying corresponding locations in the real and virtual environments because the same coordinates can be used to identify both locations. However, in some embodiments, corresponding real and virtual environments need not use a shared coordinate space. For example, in some embodiments (not shown), a matrix (or other suitable representation) can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.
[0036] 1C illustrates an exemplary MRE 150 that simultaneously presents aspects of real environment 100 and virtual environment 130 to user 110 via mixed reality system 112. In the example shown, MRE 150 simultaneously presents to user 110 real objects 122A, 124A, 126A, and 128A from real environment 100 (e.g., via a see-through portion of the display of mixed reality system 112) and virtual objects 122B, 124B, 126B, and 132 from virtual environment 130 (e.g., via an active display portion of the display of mixed reality system 112). As described above, room corner 106A / 106B acts as the origin of a coordinate space corresponding to MRE 150, and coordinate system 108 defines the x-, y-, and z-axes for the coordinate space.
[0037] In the example shown, the mixed reality objects comprise corresponding pairs of real and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) that occupy corresponding locations in coordinate space 108. In some examples, both real and virtual objects may be visible to user 110 simultaneously. This may be desirable, for example, in instances where a virtual object presents information designed to enhance the view of the corresponding real object (such as in a museum application where a virtual object presents a missing piece of an ancient, damaged statue). In some examples, the virtual objects (122B, 124B, and / or 126B) may be displayed so as to occlude the corresponding real object (122A, 124A, and / or 126A) (e.g., via active pixelated occlusion using a pixelated occlusion shutter). This may be desirable, for example, in instances where a virtual object acts as a visual substitute for the corresponding real object (such as in an interactive storytelling application where an inanimate real object becomes an “animate” character).
[0038] In some examples, real-world objects (e.g., 122A, 124A, 126A) may be associated with virtual content or helper data that may not necessarily constitute virtual objects. The virtual content or helper data may facilitate processing or handling of the virtual objects within a mixed reality environment. For example, such virtual content may include a two-dimensional representation of the corresponding real-world object, a custom asset type associated with the corresponding real-world object, or statistical data associated with the corresponding real-world object. This information may enable or facilitate computations involving the real object without incurring the computational overhead associated with creating and associating a corresponding virtual object with the real-world object.
[0039] In some embodiments, the presentations described above may also incorporate audio aspects. For example, in MRE 150, virtual monster 132 may be associated with one or more audio signals, such as footstep effects, that are generated as the monster walks around MRE 150. As described further below, a processor in mixed reality system 112 may calculate an audio signal corresponding to a mixed and processed composite of all such sounds within MRE 150 and present the audio signal to user 110 via speakers included within mixed reality system 112.
[0040] Exemplary Mixed Reality System
[0041] An exemplary mixed reality system 112 can include a wearable head-mounted unit (e.g., a wearable augmented reality or mixed reality headgear unit) that includes a display (which may include left and right see-through displays, which may be near-eye displays, and associated components for coupling light from the displays to the user's eyes), left and right speakers (e.g., positioned adjacent the user's left and right ears, respectively), an inertial measurement unit (IMU) (e.g., mounted on the device's temple arms), a quadrature coil electromagnetic receiver (e.g., mounted on the left temple component), left and right cameras (e.g., depth (time-of-flight) cameras) oriented away from the user, and left and right cameras oriented toward the user (e.g., for detecting the user's eye movements). However, the mixed reality system 112 can incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic). Additionally, mixed reality system 112 may incorporate networking features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other mixed reality systems. Mixed reality system 112 may further include a battery (which may be mounted in an auxiliary unit, such as a beltpack designed to be worn around the user's waist), a processor, and memory. The head-mounted unit of mixed reality system 112 may include a tracking component, such as an IMU or other suitable sensor, configured to output a set of coordinates of the head-mounted unit relative to the user's environment. In some examples, the tracking component may provide input to a processor that implements simultaneous localization and mapping (SLAM) and / or visual odometry algorithms. In some examples, mixed reality system 112 may also include auxiliary unit 320, which may be handheld controller 300 and / or a wearable beltpack, as described further below.
[0042] 2, 3A, and 3B together illustrate an example mixed reality system (which may correspond to mixed reality system 112) that may be used to present an MRE (which may correspond to MRE 150) to a user. Figure 2 illustrates an example wearable head unit 200 of the example mixed reality system, which may be a head-mountable system configured to be worn on a user's head. In the example shown, wearable head unit 200 (which may be, for example, a wearable augmented reality or mixed reality headgear unit) includes a display (which may include left and right see-through displays and associated components for coupling light from the displays to the user's eyes), left and right acoustic structures (e.g., speakers positioned adjacent the user's left and right ears, respectively), one or more sensors, such as a radar sensor (including a transmitting and / or receiving antenna), an infrared sensor, an accelerometer, a gyroscope, a magnetometer, a GPS unit, an inertial measurement unit (IMU), an acoustic sensor, etc., a quadrature coil electromagnetic receiver (e.g., mounted in the left temple component), left and right cameras oriented away from the user (e.g., depth (time-of-flight) cameras), and left and right cameras oriented toward the user (e.g., for detecting the user's eye movements). However, wearable head unit 200 may incorporate any suitable display technology and any suitable number, type, or combination of components without departing from the scope of the present invention. In some examples, the wearable head unit 200 may incorporate one or more microphones configured to detect audio signals generated by the user's voice, and such microphones may be positioned within the wearable head unit adjacent to the user's mouth. In some examples, the wearable head unit 200 may incorporate networking or wireless features (e.g., Wi-Fi capabilities, Bluetooth®) to communicate with other devices and systems, including other wearable systems. The wearable head unit 200 may further include a battery (which may be mounted in an auxiliary unit such as a belt pack designed to be worn around the user's waist), a processor, and memory.In some embodiments, the tracking component of the wearable head unit 200 may provide input to a processor that implements simultaneous localization and mapping (SLAM) and / or visual odometry algorithms. The wearable head unit 200 may be a first component of a mixed reality system that includes additional system components. In some embodiments, such a wearable system may also include an auxiliary unit 320, which may be a handheld controller 300 and / or a wearable belt pack, as described further below.
[0043] 3A illustrates example handheld controller components 300 of an example mixed reality system. In some examples, handheld controller 300 includes a grip portion 346 and one or more buttons 350 disposed along a top surface 348. In some examples, button 350 may be configured for use as an optical tracking target in conjunction with a camera or other optical sensor (which, in some examples, may be mounted within wearable head unit 200), e.g., to track six degrees of freedom (6DOF) movement of handheld controller 300. In some examples, handheld controller 300 includes a tracking component (e.g., an IMU, a radar sensor (including transmit and / or receive antennas), or other suitable sensor or circuitry) for detecting a position or orientation, such as a position or orientation relative to the wearable head unit or belt pack. In some examples, such tracking components may be positioned within the handle of the handheld controller 300, facing outward from a surface of the handheld controller 300 (e.g., grip portion 346, top surface 348, and / or bottom surface 352), and / or may be mechanically coupled to the handheld controller. The handheld controller 300 can be configured to provide one or more output signals corresponding to one or more of a button press state, or the position, orientation, and / or movement of the handheld controller 300 (e.g., via an IMU). Such output signals may be used as inputs to a processor of the wearable head unit 200, the handheld controller 300, or another component of the mixed reality system (e.g., a wearable mixed reality system). Such inputs may correspond to the position, orientation, and / or movement of the handheld controller (or, for that matter, the position, orientation, and / or movement of a user's hand holding the controller). Such inputs may also correspond to a user pressing a button 350. In some embodiments, handheld controller 300 may include a processor, memory, or other suitable computer system components.The processor of the handheld controller 300 may be used, for example, to execute any suitable processes disclosed herein.
[0044] 3B illustrates an example auxiliary unit 320 of a mixed reality system, such as a wearable mixed reality system. The auxiliary unit 320 may include, for example, one or more batteries to provide energy to operate the wearable head unit 200 and / or handheld controller 300, including the display and / or acoustic structures within those components, a processor (which may execute any suitable process disclosed herein), memory, or any other suitable component of a wearable system. Compared to a head-mounted unit (e.g., wearable head unit 200) or a handheld unit (e.g., handheld controller 300), the auxiliary unit 320 may be more suitable for housing large or heavy components (e.g., batteries) because it is relatively sturdy and may be more easily positioned on a part of the user's body, such as the waist or back, that is less easily fatigued by heavy items.
[0045] In some examples, sensing and / or tracking components may be located within the auxiliary unit 320. Such components may include, for example, one or more IMUs and / or radar sensors (including transmit and / or receive antennas). In some examples, the auxiliary unit 320 can use such components to determine the position and / or orientation (e.g., 6DOF location) of the handheld controller 300, the wearable head unit 200, or the auxiliary unit itself. As shown in the example, the auxiliary unit 320 may include a clip 2128 for attaching the auxiliary unit 320 to a user's belt. Other form factors are suitable for the auxiliary unit 320 and will be apparent, including form factors that do not involve mounting the unit on a user's belt. In some examples, the auxiliary unit 320 can be coupled to the wearable head unit 200 through a multi-conduit cable, which may include, for example, electrical wires and optical fibers. A wireless connection to and from the auxiliary unit 320 can also be used (e.g., Bluetooth, Wi-Fi, or any other suitable wireless technology).
[0046] FIG. 4 shows an example functional block diagram that may correspond to an example mixed reality system (e.g., a mixed reality system including one or more of the components described above with respect to FIGS. 2, 3A, and 3B). As shown in FIG. 4, an example handheld controller 400B (which may correspond to handheld controller 300 (“totem”)) may include a totem headgear six degrees of freedom (6DOF) totem subsystem 404A and a sensor 407, and an example augmented reality headgear 400A (which may correspond to wearable head unit 200) may include a totem headgear 6DOF headgear subsystem 404B. In an example, the 6DOF totem subsystem 404A and the 6DOF headgear subsystem 404B may individually or collectively determine three position coordinates and three rotation coordinates of the handheld controller 400B relative to the augmented reality headgear 400A (e.g., with respect to the coordinate system of the augmented reality headgear 400A). The three positions may be represented as X, Y, and Z values within such a coordinate system, as a transformation matrix, or as some other representation. The position coordinates may be determined through any suitable positioning technique, such as involving radar, sonar, GPS, or other sensors. The rotational coordinates may be represented as a series of yaw, pitch, and roll rotations, as a rotation matrix, as a quaternion, or as some other representation.
[0047] In some examples, the wearable head unit 400A, one or more depth cameras 444 (and / or one or more non-depth cameras) included within the wearable head unit 400A, and / or one or more optical targets (e.g., buttons 350 of the handheld controller 400B as described above, or dedicated optical targets included within the handheld controller 400B) can be used for 6DOF tracking. In some examples, the handheld controller 400B can include a camera, as described above, and the wearable head unit 400A can include an optical target for optical tracking in conjunction with the camera.
[0048] In some examples, it may be necessary to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head unit 400A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment). For example, such a transformation may be necessary for the display of the wearable head unit 400A to present a virtual object (e.g., a virtual person seated in a real chair facing forward in the real environment, regardless of the position and orientation of the headgear) in an expected position and orientation relative to the real environment, rather than in a fixed position and orientation on the display (e.g., at the same location in the lower right corner of the display). This can preserve the illusion that the virtual object exists in the real environment (e.g., does not shift or rotate unnaturally in the real environment as the wearable head unit 400A shifts and rotates). In some examples, a compensatory transformation between coordinate spaces can be determined by processing images from the depth camera 444 (e.g., using SLAM and / or visual odometry techniques) to determine the transformation of the headgear relative to the coordinate system. 4, depth camera 444 can be coupled to SLAM / visual odometry block 406 and can provide images to block 406. SLAM / visual odometry block 406 implementations can include a processor configured to process the images and then determine the user's head position and orientation, which can be used to identify a transformation between head coordinate space and real-world coordinate space. Similarly, in some embodiments, an additional source of information about the user's head pose and location is obtained from IMU 409 (or another suitable sensor, such as an accelerometer or gyroscope). Information from IMU 409 can be integrated with information from SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information in response to rapid adjustments of the user's head pose and position.
[0049] In some examples, depth camera 444 can provide 3D images to hand gesture tracker 411, which can be implemented within a processor of wearable head unit 400A. Hand gesture tracker 411 can identify the user's hand gestures, for example, by matching the 3D images received from depth camera 444 to stored patterns representing hand gestures. Other suitable techniques for identifying the user's hand gestures will be apparent.
[0050] In some embodiments, one or more processors 416 may be configured to receive data from the wearable head unit's headgear subsystem 404B, radar sensor 408, IMU 409, SLAM / visual odometry block 406, depth camera 444, microphone 450, and / or hand gesture tracker 411. The processor 416 may also send and receive control signals from the totem system 404A. The processor 416 may be wirelessly coupled to the totem system 404A, such as in embodiments in which the handheld controller 400B is not tethered to other system components. The processor 416 may further communicate with additional components, such as an audiovisual content memory 418, a graphical processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to a head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to a left image-wise dimming light source 424 and a right channel output coupled to a right image-wise dimming light source 426. The GPU 420 may output stereoscopic image data to the image-wise dimming light sources 424, 426. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 may receive input from the processor 419 indicating a direction vector from the user to a virtual sound source (which may be moved by the user, e.g., via the handheld controller 320). Based on the direction vector, the DSP audio spatializer 422 may determine a corresponding HRTF (e.g., by accessing an HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 422 may apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the verisimilitude and realism of virtual sounds by incorporating the user's relative position and orientation to the virtual sound within the mixed reality environment, i.e., by presenting a virtual sound that matches the user's expectations of what the virtual sound would sound like if it were an actual sound in the real environment.
[0051] 4 , one or more of processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audiovisual content memory 418 may be included within auxiliary unit 400C (which may correspond to auxiliary unit 320 described above). Auxiliary unit 400C may include battery 427 for powering its components and / or for supplying power to other system components, such as wearable head unit 400A and / or handheld controller 400B. Including such components within an auxiliary unit that may be mounted on the user's waist can limit the size and weight of wearable head unit 400A, which may in turn reduce fatigue in the user's head and neck.
[0052] While FIG. 4 presents elements corresponding to various components of an exemplary mixed reality system, various other suitable arrangements of these components will be apparent to those skilled in the art. For example, elements presented in FIG. 4 as being associated with auxiliary unit 400C may instead be associated with wearable head unit 400A and / or handheld controller 400B. And, one or more of wearable head unit 400A, handheld controller 400B, and auxiliary unit 400C may include a processor that can perform one or more of the methods disclosed herein. Furthermore, some mixed reality systems may dispense with handheld controller 400B or auxiliary unit 400C altogether. Such variations and modifications should be understood as being within the scope of the disclosed embodiments.
[0053] 5 shows an example configuration in which client device 510 (which may be a component of a mixed reality system, including a wearable mixed reality system) communicates with server 520 via communications network 530. Client device 510 may comprise, for example, one or more of wearable head unit 200, handheld controller 300, and auxiliary unit 320 as described above. Server 520 may comprise one or more dedicated server machines (which may include, for example, one or more cloud servers), but in some examples may comprise one or more of wearable head unit 200, handheld controller 300, and / or auxiliary unit 320 that may act as a server. Server 520 may communicate with one or more client devices, including client component 510, via communications network 530 (e.g., via the Internet and / or via a wireless network). Server 520 may maintain a persistent world state with which one or many users may interact (e.g., via a client device corresponding to each user). Additionally, server 520 may perform computationally intensive operations that would be prohibitively expensive to execute on "thin" client hardware. Other client-server topologies will be apparent in addition to the example shown in FIG. 5. For example, in some embodiments, a wearable system may act as a server to other wearable system clients. Additionally, in some embodiments, a wearable system may communicate and share information via a peer-to-peer network. The present disclosure is not limited to any particular topology of networked components. Furthermore, embodiments of the disclosure herein may be implemented on any suitable combination of client and / or server components, including processors residing in client and server devices.
[0054] Virtual Audio
[0055] As described above, an MRE (e.g., as experienced via mixed reality system 112, which may include components such as wearable head unit 200, handheld controller 300, or auxiliary unit 320 described above) can present a user of the MRE with audio signals that appear to originate from a sound source with origin coordinates within the MRE and travel in the direction of an orientation vector within the MRE. That is, the user can perceive these audio signals as if they were actual audio signals originating from the origin coordinates of the sound source and traveling along the orientation vector.
[0056] In some cases, audio signals may be considered virtual in that they correspond to computational signals in a virtual environment. The virtual audio signals, when generated, for example, via speakers 2134 and 2136 of wearable head unit 200 of FIG. 2, can be presented to a user as actual audio signals detectable by the human ear.
[0057] The sound source may correspond to a real object and / or a virtual object. For example, a virtual object (e.g., virtual monster 132 in FIG. 1C ) can emit an audio signal within the MRE that is represented in the MRE as a virtual audio signal and presented to the user as a real audio signal. For example, virtual monster 132 in FIG. 1C can emit a virtual sound corresponding to the monster's utterances (e.g., dialogue) or sound effects. Similarly, a real object (e.g., real object 122A in FIG. 1C ) can be created to appear to emit a virtual audio signal within the MRE that is represented in the MRE as a virtual audio signal and presented to the user as a real audio signal. For example, real lamp 122A can emit a virtual sound corresponding to the sound effect of a lamp being turned on or off, even when the lamp is not turned on or off in the real environment. The virtual sound can correspond to the position and orientation of the sound source (whether real or virtual). For example, if a virtual sound is presented to a user as a real audio signal (e.g., via speakers 2134 and 2136), the user may perceive the virtual sound as emanating from the location of the sound source and traveling in the direction of the sound source's orientation. The sound source is referred to herein as a "virtual sound source," even though the underlying object apparently made to emit the sound may itself correspond to a real or virtual object as described above.
[0058] Some virtual or mixed reality environments suffer from the perception that the environment does not feel real or authentic. One reason for this perception is that audio and visual cues do not always match each other within such environments. For example, if a user is positioned behind a large brick wall in an MRE, the user may expect sounds originating from behind the brick wall to be quieter and more attenuated than sounds originating immediately next to the user. This expectation is based on the user's auditory experience in the real world, where sounds become quieter and attenuated when passing through large, dense objects. When the user is presented with an audio signal that intentionally originates from the brick wall but is presented at full volume without attenuation, the illusion that the sounds originate from behind the brick wall is violated. The entire virtual experience may feel fake and inauthentic, in part because it does not conform to the user's expectations based on real-world interactions. Furthermore, in some cases, an "uncanny valley" problem arises, in which even subtle differences between virtual and real experiences can cause heightened discomfort. It is desirable to improve the user's experience by presenting audio signals within the MRE that are believed to interact realistically, even subtly, with objects in the user's environment. The more such audio signals match the user's expectations, based on real-world experiences, the more immersive and engaging the user's experience within the MRE can be.
[0059] One way a user perceives and understands their surrounding environment is through audio cues. In the real world, the actual audio signals a user hears are affected by the location where those audio signals originate, the direction they propagate, and the objects with which they interact. For example, all other factors being equal, a sound originating at a greater distance from the user (e.g., a dog barking in the distance) will be perceived as quieter than the same sound originating at a closer distance from the user (e.g., a dog barking in the same room as the user). The user can therefore identify the dog's location within the real environment based in part on the perceived volume of the bark. Similarly, all other factors being equal, a sound traveling away from the user (e.g., the voice of an individual facing away from the user) will be perceived as less clear and more attenuated (i.e., low-pass filtered) than the same sound traveling toward the user (e.g., the voice of a person facing toward the user). The user can therefore identify the orientation of an individual within the real environment based on the perceived characteristics of that individual's voice.
[0060] A user's perception of an actual audio signal may also be affected by the presence of objects in the environment with which the audio signal interacts. That is, a user may perceive not only the audio signal generated by the sound source, but also the reverberation of the audio signal relative to nearby objects. For example, if an individual speaks in a small room with close walls, these walls may result in a short, natural reverberation signal as the individual's voice reflects off the walls. From these reverberations, a user may infer that the user is in a small room with closed walls. Similarly, a large concert hall or cathedral may cause longer reverberations, from which a user may infer that the user is in a large, spacious room. Similarly, the reverberation of audio signals may take on different sonic characteristics based on the position or orientation of the surfaces from which they reflect, or the material of these surfaces. For example, reverberation relative to a tiled wall will sound different from reverberation relative to brick, carpet, drywall, or other materials. These reverberation characteristics can be used by a user to acoustically understand the size, shape, and material composition of the space in which they reside.
[0061] The above examples illustrate how audio cues can inform a user's perception of their surrounding environment. These cues can work in combination with visual cues; for example, if a user sees a dog in the distance, the user may expect the dog's sound to match that distance (and may feel confused or disoriented if this is not the case, as in some virtual environments). In some examples, such as in low-light environments or for visually impaired users, visual cues may be limited or unavailable; in such cases, audio cues may take on particular importance and serve as the user's primary means of understanding their environment.
[0062] It may be desirable to present virtual audio signals to a user within an MRE in a manner that incorporates realistic reverberation effects based on objects within the MRE so that the user can understand the virtual audio signals as realistically present in that physical space. Some mixed reality systems may create a dissonance between the user's auditory experience within the MRE and the user's auditory experience in the real world, such that the audio signals within the MRE appear less accurate (e.g., the "uncanny valley" problem). Compared to other mixed reality audio systems, the present disclosure may enable a more nuanced and realistic presentation of audio signals by considering the user's position, orientation, the nature of objects in the user's environment, the nature of the user's environment, and other characteristics of the audio signal and the environment. By presenting an audio experience to a user of an MRE that evokes the audio experience of their daily life, the MRE can enhance the user's sense of immersion and connectedness when engaging with the MRE.
[0063] FIG. 6 illustrates an example process 600 for presenting a virtual audio signal to a user of a mixed reality environment (e.g., mixed reality environment 150 of FIG. 1C ), according to some embodiments. The user may be using a wearable mixed reality system such as described above with respect to FIGS. 1-4 . According to process 600, an audio event 610 may be identified. The audio event 610 may be associated with one or more audio assets (e.g., a waveform audio file or a live audio stream from a microphone or from a network) and may have a position and orientation within the coordinate system of the MRE. Audio events that are within the user's acoustic space (e.g., close enough to the user to be heard) may be presented to the user via speakers, such as speakers 2134 and 2136, of wearable head unit 200.
[0064] According to example process 600, such an audio event can be presented to a user of a wearable mixed reality system as follows: In stage 620, one or more raw audio assets associated with audio event 610 can be loaded into the wearable system's memory or otherwise prepared for presentation via the wearable system (e.g., by loading a portion of an audio stream into a streaming audio buffer). The raw audio assets can include one or more static audio files or portions of such audio files (e.g., one or more samples of a file), and / or may include a real-time audio feed, such as the output of a microphone, or an audio stream received over the Internet. In some examples, it may be preferable for such raw audio assets to be "dry," with minimal effects or processing applied to the raw audio assets.
[0065] At stage 630, one or more acoustic parameters can be determined that, when applied to the raw audio asset at stage 640 to create a processed audio signal, can enhance the audio asset by adding sonic characteristics consistent with the user's current acoustic environment (e.g., the current "room"). These acoustic parameters can correspond to the acoustic effects that the room would impart to the underlying audio generated in that room. Such acoustic parameters can include, for example, parameters corresponding to attenuation (e.g., volume reduction) of the underlying audio, filtering (e.g., low-pass filtering), phase shifting of the underlying audio, pitch modulation of the underlying audio, or other acoustic effects. The acoustic parameters can also include input parameters (e.g., wet / dry levels, attack / decay times) of a reverberation engine for applying reverberation and reflection effects to the underlying audio. Thus, the processed audio signal output by stage 640 can incorporate a simulation of reverberation, attenuation, filtering, or other effects that would be imparted to the raw audio asset by the walls, surfaces, and / or objects in the room. The application of the acoustic parameters in stage 640 can be described as convolution of one or more transfer functions based on the acoustic parameters (e.g., transfer function H(t)) with the raw audio asset to generate a processed audio signal. This process can be performed by an audio engine, which may include a reverberation engine, to which the raw audio asset and appropriate input parameters are supplied. The determination of the acoustic parameters in stage 630 is described in more detail below.
[0066] The audio signal generated in stage 640 may be a virtual audio signal that is not directly perceptible by the user but can be converted into an actual audio signal by one or more speakers (e.g., speakers 2134 and / or 2136) so that it can be heard by the user. For example, the audio signal may be a computational representation that includes coordinates in the mixed reality environment where the processed audio signal originates, a vector in the MRE along which the processed audio signal propagates, a time at which the processed audio signal originates, a velocity at which the processed audio signal propagates, or other suitable characteristics. In stage 650, the one or more virtual audio signals can be mixed down into one or more channels corresponding to the speaker configuration of wearable head unit 200. For example, in stage 650, the virtual audio signal may be mixed down into the left and right channels of a stereo speaker configuration. In step 660, these mixed-down signals are output via speakers, e.g., digital audio data that may be converted to analog signals via a digital-to-analog converter (e.g., as part of DSP audio spatializer 422 of FIG. 4), then amplified and used to drive speakers to produce sounds perceptible by the user.
[0067] FIG. 7 illustrates an example process 700 for determining acoustic parameters for an audio event as described above with respect to stage 630 of example process 600. Example process 700 can be executed on one or more processors of, for example, wearable head unit 200 and / or a server, such as server 520 described above. As described above, such acoustic parameters can represent the acoustic characteristics of the room in which the audio event occurs. These acoustic characteristics, and thus the acoustic parameters, are determined in large part based on / with respect to the physical dimensions of the room, the objects present in the room and the size and shape of these objects, the materials of the room's surfaces and any objects within the room, and the like. Because these characteristics of a room may remain constant over time, it may be beneficial to associate each individual room in the MRE with a set of acoustic parameters (an “acoustic fingerprint”) that describe the acoustic characteristics of that room. This configuration has several potential advantages. By creating, storing, and retrieving acoustic fingerprints on a room-by-room basis, acoustic parameters can be easily and efficiently managed, exchanged, and updated without having to recreate such parameters each time a user enters a room. Additionally, as described below, this configuration can simplify the process of generating acoustic parameters that describe the composition of two or more rooms. Furthermore, allowing the acoustic parameters of a room to persist over time can improve immersion because an individual's auditory experience of the physical space within the MRE remains constant over time (as in real-world auditory spaces). And because the same set of acoustic parameters can be provided to multiple users, multiple users in a single shared space can receive a common auditory experience, improving the sense of connectedness among those users.
[0068] Exemplary process 700 describes a system in which acoustic parameters are stored on a room-by-room basis (although other suitable configurations are possible and within the scope of this disclosure). At stage 710 of process 700, a room is identified for an audio event, and this room can determine a set of audio parameters to be applied to the audio event. The room can be identified using one or more sensors of the mixed reality system (e.g., sensors of wearable head unit 200). For example, a GPS module of wearable head unit 200 can identify the user's location, which can be used to determine the room corresponding to that location. In some examples, the user's location can be determined by triangulation based on the location of nearby Wi-Fi receivers or cellular antennas. In some examples, sensors such as LIDAR, depth cameras, RGB cameras, and / or the like can be used to identify the user's current surroundings, and the sensor output can be compared against a room database to identify the room corresponding to the sensor output. Determining the room from the user's location can, in some embodiments, be performed based on mapping data and / or architectural records, such as floor plan records, which may be stored on a server, such as server 520 described above. Other techniques for identifying the room corresponding to the user's current location will be apparent to those skilled in the art.
[0069] In exemplary process 700, a query can be made as to whether a set of acoustic parameters exists and can be retrieved. At stage 720, a client device (e.g., client device 510 described above, which may include a wearable head unit) can be queried for acoustic parameters corresponding to the current room. If such a set is determined to be stored on the client device (stage 730), it can be retrieved and output for use (stage 770). If the set of acoustic parameters is not stored on the client device, a server (e.g., server 520 described above) can be queried for the acoustic parameters at stage 740. As above, if such a set is determined to be stored on the server (stage 750), it can be retrieved and output for use (stage 770). If the current set of acoustic parameters for the room is not available on either the client device or the server, a new set of acoustic parameters is created for the room in step 760, as described in more detail below, with the resulting acoustic parameter output (step 770) for use, and can potentially be stored on the client device or server device for subsequent retrieval, as described below.
[0070] FIG. 8 illustrates an example process 800 for determining the placement of acoustic parameters of a room, as may be implemented in stage 760 of example process 700. Example process 800 may employ any combination of suitable techniques for determining such acoustic parameters. One such technique includes determining the acoustic parameters based on data from sensors of a wearable device, such as wearable head unit 200. In stage 810, such sensor data may be provided as input to the example process. The sensor data may include data from a depth camera (e.g., depth camera 444), an RGB camera, a LIDAR module, a sonar module, a radar module, a GPS receiver, an orientation sensor (e.g., an IMU, a gyroscope, or an accelerometer), and / or a microphone (e.g., microphone 450). In stage 820, from the sensor input, the current room geometry may be determined. Such geometry may include the size, shape, position, and orientation of one or more surfaces (e.g., walls, floor, ceiling) and / or objects within the room. This data can affect the acoustics of sounds within a room. For example, large, cavernous spaces can cause longer and more noticeable reverberation than smaller spaces. Similarly, rooms filled with acoustically damping objects (e.g., curtains, sofas) can dampen sounds within those rooms.
[0071] The room geometry information can be determined based on sensor inputs (e.g., camera images showing light reflected by the geometry, LIDAR data providing spatial coordinates corresponding to the geometry) using techniques familiar to those skilled in the art. In some embodiments, the room geometry may be retrieved from a database that associates the room geometry with geographic coordinates, such as may be provided by a GPS receiver in step 810. Similarly, in some embodiments, the GPS coordinates can be used to retrieve architectural data (e.g., a floor plan) corresponding to the GPS coordinates, and the room geometry can be determined using the architectural data.
[0072] In addition to the room geometry determined in step 820, materials corresponding to that geometry can be determined in step 830. Such materials can exhibit acoustic properties that affect sound within the room. For example, a wall made of tile will be acoustically reflective and exhibit a bright reverberation, while a carpet-covered floor will exhibit a damping effect. Such materials can be determined using sensor input provided in step 810. For example, an RGB camera can be used to identify surface materials based on their visual appearance. Other suitable techniques will be apparent to those skilled in the art. As noted above, in some embodiments, surface materials may be retrieved from a database that associates surface materials with geographic coordinates, such as may be provided by a GPS receiver in step 810, or from building data corresponding to those coordinates.
[0073] In step 840, the room geometry determined in step 820 and / or the surface materials determined in step 830 can be used to determine corresponding acoustic parameters of the room, which describe the acoustic effect the room geometry and / or surface materials may have on sound within the room. Various techniques can be used to determine such acoustic parameters. As one example, reverberation engine input parameters (e.g., decay time, mix level, attack time, or selection index into the reverberation algorithm) can be determined based on a known relationship to the room's volume. As another example, a physical representation of the room can be constructed based on sensor inputs, with an acoustic response model of the room mathematically determined from the representation. As another example, a lookup table can be maintained that associates reverberation or filter parameters with surface material types. In cases where a room contains multiple materials with different acoustic parameters, a composite set of acoustic parameters can be determined by blending parameters, for example, based on the relative surface area of the room covered by each individual material. Other suitable exemplary techniques are described, for example, in L. Savioja et al., Creating Interactive Virtual Acoustic Environments, 47 J. Audio Eng. Soc. 675, 705 n. 9 (1999) and will be familiar to those skilled in the art.
[0074] Another technique for determining the acoustic characteristics of a room involves presenting a known test audio signal through speakers in the room, recording a "wet" test signal through microphones in the room, and presenting the test signal (850) and the wet signal (860) for comparison in step 840. The comparison of the test signal and the wet signal is described, for example, in A. Deb et al., Time Invariant System Identification: Via 'Deconvolution', in Analysis and Identification of Time-Invariant Systems, Time-Varying Systems, and Multi-Delay Systems using Orthogonal Hybrid Functions 319-330 (Springer, 1999). st ed. 2016), a transfer function can be generated that characterizes the room's acoustic effect on the test signal. In some embodiments, a "blind" estimation technique may be employed to retrieve the room's acoustic parameters by recording only the wet signal, as described, for example, in J. Jot et al., Blind Estimation of the Reverberation Fingerprint of Unknown Acoustic Environments, Audio Engineering Society Convention Paper 9905 (October 18-21, 2017).
[0075] In some embodiments, such as exemplary process 800, multiple techniques for determining acoustic parameters can be combined. For example, the acoustic parameters determined from the test signal and the wet signal as described above with respect to steps 850 and 860 can be refined using the room geometry and surface materials determined in steps 820 and 830, respectively, and / or vice versa.
[0076] In response to determining the set of acoustic parameters for the room in step 840, the set of acoustic parameters may be stored for subsequent retrieval so as to avoid the need to recalculate such parameters (which may incur significant computational overhead). The set of acoustic parameters may be stored on a client device (e.g., client device 510, as retrieved as described above with respect to step 720 of process 700), on a server device (e.g., server device 520, as retrieved as described above with respect to step 740 of process 700), in another suitable storage location, or in some combination of the above.
[0077] In some examples, it may be desirable to obtain a more realistic acoustic model by subjecting an audio signal to acoustic parameters associated with more than one room (e.g., in stage 640 of example process 600). For example, in an acoustic environment that includes more than one acoustic region or room, the audio signal may take on the acoustic properties of multiple rooms. Also, in MRE, one or more of such rooms may be virtual rooms that correspond to acoustic regions that do not necessarily exist in the real environment.
[0078] FIG. 9 illustrates an example room 900 that includes multiple acoustically connected areas. In FIG. 9, area 910 corresponds to a living room with various objects within the room. A doorway 914 connects living room 910 with a second room, dining room 960. A sound source 964 is located within dining room 960. In this real-world environment, a sound generated by sound source 964 in dining room 960 and heard by a user in living room 910 would take on the acoustic characteristics of both dining room 960 and living room 910. In an MRE corresponding to indoor scene 900, if the virtual sound similarly adopted the acoustic characteristics of these multiple rooms, a more realistic acoustic experience would result.
[0079] A multi-room acoustic environment, such as the example room 900 of FIG. 9 , can be represented by an acoustic graph structure that describes the acoustic relationships between rooms in the environment. FIG. 10 shows an example acoustic graph structure 1000 that may describe rooms in a house corresponding to the example indoor scene 900. Each room in the acoustic graph structure may have its own unique acoustic characteristics. In some embodiments, the acoustic graph structure 1000 may be stored on a server, such as server 520, which may be accessed by one or more client devices, such as client device 510. In the example acoustic graph structure 1000, the living room 910 shown in FIG. 9 is represented by a corresponding room data structure 1010. The room data structure 1010 may be associated with one or more data elements that describe aspects of the living room 910 (e.g., the size and shape of the room, objects in the room, and the like). In an embodiment, an acoustic parameter data structure 1012 is associated with the room data structure 1010 and may describe a set of acoustic parameters associated with the corresponding living room 910. This set of acoustic parameters may correspond, for example, to the set of acoustic parameters as described above with respect to FIGS.
[0080] In acoustic graph structure 1000, rooms in a house may be acoustically coupled (e.g., via windows, doorways, or objects through which sound waves may travel). These acoustic connections are indicated via lines connecting room data structures in acoustic graph structure 1000. For example, acoustic graph structure 1000 includes a room data structure 1060 corresponding to dining room 960 in FIG. 9. In the illustration, dining room data structure 1060 is connected to living room data structure 1010 by a line, reflecting that dining room 960 and living room 910 are acoustically coupled via doorway 964, as shown in FIG. 9. Like living room data structure 1010, dining room data structure 1060 is associated with an acoustic parameter data structure 1062, which may describe a set of acoustic parameters associated with the corresponding dining room 960. Similarly, acoustic graph structure 1000 includes representations of other rooms in the house (e.g., basement 1030, kitchen 1040, den 1050, bedroom 1020, bathroom 1070, garage 1080, office 1090) and their associated acoustic parameters (e.g., 1032, 1042, 1052, 1022, 1072, 1082, and 1090, corresponding to basement 1030, kitchen 1040, den 1050, bedroom 1020, bathroom 1070, garage 1080, and office 1090, respectively). As shown in the figure, these rooms and their associated data may be represented using a hash table. The lines connecting the room data structures represent the acoustic connections between the rooms. The parameters describing the acoustic connections between the rooms can be represented, for example, by data structures associated with the lines, within the acoustic parameter data structures (e.g., 1012, 1062) described above, or via some other data structure. Such parameters may include, for example, the size of the opening between the rooms (e.g., doorway 914), the thickness and material of the walls between the rooms, etc. This information can be used to determine the extent to which the acoustical properties of one room affect the sounds produced or heard in an acoustically connected room.
[0081] An acoustic graph structure, such as exemplary acoustic graph structure 1000, can be created or modified using any suitable technique. In some examples, rooms can be added to the acoustic graph structure based on sensor input from a wearable system (e.g., input from sensors such as a depth camera, an RGB camera, LIDAR, sonar, radar, and / or GPS). The sensor input can be used to identify rooms, room geometries, room materials, objects, object materials, and the like, as described above, and to determine whether (and how) rooms are acoustically connected. In some examples, the acoustic graph structure can be manually modified, such as when a mixed reality designer desires to add a virtual room (which may not have a real-world counterpart) to one or more existing rooms.
[0082] 11 shows an example process 1100 for determining a set of composite acoustic parameters associated with two or more acoustically connected rooms for audio that may be presented (by a sound source) in a first room and heard (by a user) in a second room different from the first room. Example process 1100 can be used to retrieve acoustic parameters for application to an audio signal and can be implemented, for example, in step 630 of example process 600 described above. In step 1110, a room corresponding to a user's location can be identified as described above with respect to step 710 of example process 700. This user's room may correspond, for example, to living room 910 described above. In step 1120, the acoustic parameters of the user's room are determined, for example, as described above with respect to FIGS. 7-8 and steps 720 through 760. In the example described above, these parameters may be described by acoustic parameters 1012.
[0083] At step 1130, a room corresponding to the location of the sound can be identified as described above with respect to step 710 of exemplary process 700. For example, this sound source may correspond to sound source 964 described above, and the sound source room may correspond to dining room 960 described above (which is acoustically connected to living room 910). At step 1140, acoustic parameters of the sound source room are determined, for example, as described above with respect to FIGS. 7-8 and steps 720 through 760. In the embodiment described above, these parameters may be described by acoustic parameters 1062.
[0084] At stage 1150 of example process 1100, an acoustic graph describing the acoustic relationship between the user's room and the sound source's room can be determined. The acoustic graph can correspond to acoustic graph structure 1000 described above. In some embodiments, this acoustic graph can be retrieved in a manner similar to the process described with respect to FIG. 7 for retrieving acoustic parameters; for example, the acoustic graph can be selected from a set of acoustic graphs that can be stored on the client device and / or server based on sensor input.
[0085] In response to determining the acoustic graph, rooms that may be acoustically connected to the sound source's room and / or the user's room, and the acoustic effects that these rooms may have on the presented sound, can be determined from the acoustic graph. For example, using acoustic graph structure 1000 as an example, the acoustic graph indicates that living room 1010 and dining room 1060 are directly connected by a first path, and the acoustic graph further indicates that living room 1010 and dining room 1060 are also indirectly connected via a second path that includes kitchen 1040. In stage 1160, acoustic parameters of such intermediate rooms can be determined (e.g., as described above with respect to Figures 7-8 and stages 720 to 760). In addition, stage 1160 can determine parameters that describe the acoustic relationships between these rooms (such as the size and shape of objects or passageways between the rooms), as described above.
[0086] The outputs of stages 1120, 1140, and 1160, i.e., the acoustic parameters corresponding to the user's room, the source's room, or any intermediate rooms, respectively, along with parameters describing their acoustic connections, can be presented to stage 1170, at which point they can be combined into a single composite set of acoustic parameters that can be applied to the audio, for example, as described in J. Jot et al., Binaural Simulation of Complex Acoustic Scenes for Interactive Audio, Audio Engineering Society Convention Paper 6950 (October 1, 2006). In some embodiments, the composite set of parameters can be determined based on the acoustic relationships between the rooms, as may be represented by an acoustic graph. For example, in some embodiments, if the user's room and the source's room are separated by a thick wall, the acoustic parameters of the user's room may dominate the composite set of acoustic parameters relative to the acoustic parameters of the source's room. However, in some embodiments, if the rooms are separated by a large doorway, the acoustic parameters of the source's room may be more prominent. The composite parameters may also be determined based on the user's location relative to the room, for example, if the user is located near an adjacent room, the acoustic parameters of that room may be more prominent than if the user were located farther away from the room. In response to determining a composite set of acoustic parameters, the composite set can be applied to the audio to impart the acoustic characteristics of not just a single room, but the entire connected acoustic environment as described by the acoustic graph.
[0087] 12, 13, and 14 illustrate components of an exemplary wearable system that may correspond to one or more embodiments described above. For example, the exemplary wearable head unit 12-100 shown in FIG. 12, the exemplary wearable head unit 13-100 shown in FIG. 13, and / or the exemplary wearable head unit 14-100 shown in FIG. 14 may correspond to the wearable head unit 200, the exemplary handheld controller 12-200 shown in FIG. 12 may correspond to the handheld controller 300, and the exemplary auxiliary unit 12-300 shown in FIG. 12 may correspond to the auxiliary unit 320. As illustrated in FIG. 12, the wearable head unit 12-100 (also referred to as augmented reality glasses) may include an eyepiece, a camera (e.g., a depth camera, an RGB camera, and the like), a stereo image source, an inertial measurement unit (IMU), and a speaker. Briefly referring to FIG. 13, the wearable head unit 13-100 (also referred to as augmented reality glasses) may include left and right eyepieces, an image source (e.g., a projector), left and right cameras (e.g., depth cameras, RGB cameras, and the like), and left and right speakers. The wearable head unit 13-100 may be worn on a user's head. Briefly referring to FIG. 14, the wearable head unit 14-100 (also referred to as augmented reality glasses) may include left and right eyepieces, each eyepiece including one or more internal coupling gratings, orthogonal pupil expansion gratings, and exit pupil expansion gratings. Referring back to FIG. 12, the wearable head unit 12-100 may be communicatively coupled to an auxiliary unit 12-300 (also referred to as a battery / computer), for example, by a wired or wireless connection. The handheld controller 12-200 may be communicatively coupled to the wearable head unit 12-100 and / or the auxiliary unit 12-300, for example, by a wired or wireless connection.
[0088] FIG. 15 illustrates an exemplary configuration of an exemplary wearable system that may correspond to one or more embodiments described above. For example, the exemplary augmented reality user gear 15-100 may include a wearable head unit and may correspond to the client device 510 described above with reference to FIG. 5, the cloud server 15-200 may correspond to the server device 520 described above with reference to FIG. 5, and the communication network 15-300 may correspond to the communication network 530 described above with reference to FIG. 5. The cloud server 15-200 may include, among other components / elements / modules, an audio reverberation analysis engine. The communication network 15-300 may be, for example, the Internet. The augmented reality user gear 15-100 may include, for example, a wearable head unit 200. The augmented reality user gear 15-100 may include a vision system, an audio system, and a positioning system. The vision system may include left and right stereo image sources that provide images to left and right augmented reality eyepieces, respectively. The vision system may further include one or more cameras (e.g., depth cameras, RGB cameras, and / or the like). The audio system may include one or more speakers and one or more microphones. The location system may include sensors such as one or more cameras, Wi-Fi, GPS, and / or other wireless receivers.
[0089] FIG. 16 illustrates a flowchart of an example process 16-100 for presenting an audio signal to a user of a mixed reality system, which may correspond to one or more embodiments described above. For example, one or more aspects of example process 16-100 may correspond to one or more of the example processes described above with respect to FIGS. 6, 7, and / or 8. The mixed reality system to which example process 16-100 refers may include a mixed reality device. After starting the example process, a room identification is determined. It is determined whether the room's reverberation characteristics / parameters are stored locally, for example, on the mixed reality device (sometimes referred to as a client device). If the room's reverberation characteristics / parameters are stored locally, the locally stored reverberation characteristics / patterns are accessed, and audio associated with the virtual content with the room's reverberation characteristics / parameters is processed. If the room's reverberation characteristics / parameters are not stored locally, the room identification is sent to a cloud server along with a request for the room's reverberation characteristics / parameters. It is determined whether the reverberation characteristics / parameters of the room are immediately available from the cloud server. If the reverberation characteristics / parameters of the room are immediately available from the cloud server, the room reverberation characteristics / parameters of the room are received from the cloud server, and audio associated with the virtual content with the room reverberation characteristics / parameters is processed. If the reverberation characteristics / parameters of the room are not immediately available from the cloud server, the geometry of the room is mapped, materials affecting audio in the room are detected, and audio signals in the room are recorded. The identification of the room, the mapped geometry of the room, materials affecting audio in the room, and the recorded audio signals in the room are sent to the cloud server. The reverberation characteristics / parameters of the room are received from the cloud server, and audio associated with the virtual content with the room reverberation characteristics / parameters is processed. After the audio associated with the virtual content with the room reverberation characteristics / parameters is processed, the audio associated with the virtual content is output through the mixed reality system (e.g., via mixed / augmented reality user gear).
[0090] 17-19 illustrate flowcharts of example processes 17-100, 18-100, and 19-100, respectively, for presenting an audio signal to a user of a mixed reality system, which may correspond to one or more embodiments described above. For example, one or more aspects of example processes 17-100, 18-100, and / or 19-100 may correspond to one or more of the example processes described above with respect to FIGS. 6, 7, and / or 8.
[0091] In some embodiments, one or more steps of process 17-100 of Figure 17 may be performed by a cloud server. After initiating example process 17-100, an identification of a room is received along with a request for the room's reverberation characteristics / parameters. The room's reverberation characteristics / parameters are transmitted to a first mixed / augmented reality user gear.
[0092] In some embodiments, one or more steps of process 18-100 of FIG. 18 may be performed by a cloud server. After initiating example process 18-100, an identification of a particular room is received. The persistent world model graph is checked to identify adjacent connected rooms. The reverberation characteristics / parameters of any adjacent connected rooms are accessed. The reverberation characteristics / parameters of any adjacent connected rooms are transmitted to the mixed / augmented reality user gear.
[0093] In some embodiments, one or more steps of process 19-100 of FIG. 19 may be implemented by a cloud server. After initiating example process 19-100, a room identification is received from a first mixed / augmented reality user gear along with room data. The room data may include, for example, a mapped geometry of the room, materials affecting audio in the room, and a recorded audio signal in the room. In some embodiments, reverberation characteristics / parameters are calculated based on the room geometry and materials affecting audio in the room. In some embodiments, the recorded audio signal in the room is processed to extract the room reverberation characteristics / parameters. The room reverberation characteristics / parameters associated with the room identification are stored in the cloud server. The room reverberation characteristics / parameters are transmitted to the first mixed / augmented reality user gear.
[0094] FIG. 20 illustrates a flowchart of an exemplary process 20-100 for presenting an audio signal to a user of a mixed reality system based on parameters of acoustically connected spaces, which may correspond to one or more embodiments described above. For example, one or more aspects of exemplary process 20-100 may correspond to the exemplary process described above with respect to FIG. 11. After starting exemplary process 20-100, acoustic parameters of a space in which mixed / augmented reality user gear is operated are received. Using the acoustic parameters of the space in which mixed / augmented reality user gear is operated, visual and / or audio emitting virtual content is generated within the space in which mixed / augmented reality user gear is operated. World graph information is accessed to identify adjacent connected spaces. Acoustic parameters of the adjacent connected spaces are received. Virtual content is moved into the adjacent connected spaces. Audio segments for the virtual content are processed using the acoustic parameters of the adjacent connected spaces. The processed audio segments can then be presented as output.
[0095] FIG. 21 illustrates a flowchart of an example process 21-100 for determining a user's location in a mixed reality system, which may correspond to one or more embodiments described above. For example, the example process 21-100 may be implemented at step 710 of the example process 700 described above with respect to FIG. 7. In an example embodiment, after starting, it is determined whether a sufficient number of GPS satellites are within range. If a sufficient number of GPS satellites are within range, the GPS receiver is operated to determine a position. If a sufficient number of GPS satellites are not within range, the identification of nearby Wi-Fi receivers is received. Stored information about the location of nearby Wi-Fi receivers is accessed. One or more images of the space in which the device is located are captured. The images are assembled into a composite. A pixel intensity histogram of the composite is determined. The determined histogram is matched to a set of pre-stored histograms, each linked to a location node in the world graph. Based on the identification of accessible Wi-Fi networks, GPS networks, and / or histogram matches, the current location of the mixed / augmented reality user gear can be estimated.
[0096] In some embodiments, the augmented reality user gear may include a localization subsystem for determining an identification of a space in which the augmented reality user gear is located, a communication subsystem for communicating an identification of the space in which the augmented reality gear is located to receive at least one audio parameter associated with the identification of the space, and an audio output subsystem for processing an audio segment based on the at least one parameter and outputting the audio segment. Alternatively or additionally, in some embodiments, the augmented reality user gear may include a sensor subsystem for acquiring information regarding acoustical properties of a first space in which the augmented reality gear is located, an audio processing subsystem for processing the audio segment based on the information regarding the acoustical properties of the first space, the audio processing subsystem communicatively coupled to the sensor subsystem, and an audio speaker for outputting the audio segment, the audio speaker coupled to the audio processing subsystem for receiving the audio segment. Alternatively or additionally, in some embodiments, the sensor subsystem is configured to acquire geometric information for the first space. Alternatively or additionally, in some embodiments, the sensor subsystem includes a camera. Alternatively or additionally, in some embodiments, the camera includes a depth camera. Alternatively or additionally, in some embodiments, the sensor subsystem includes a stereo camera. Alternatively or additionally, in some embodiments, the sensor subsystem includes an object recognizer configured to recognize distinct objects having distinct sound absorption properties. Alternatively or additionally, in some embodiments, the object recognizer is configured to recognize at least one object selected from the group consisting of a carpet, a curtain, and a sofa. Alternatively or additionally, in some embodiments, the sensor subsystem includes a microphone.Alternatively or additionally, in some embodiments, the augmented reality gear further includes a localization subsystem for determining an identification of a first space in which the augmented reality user gear is located, and a communication subsystem for communicating the identification of the first space in which the augmented reality gear is located and for transmitting information regarding the acoustic properties of the first space in which the augmented reality gear is located. Alternatively or additionally, in some embodiments, the communication subsystem is further configured to receive information derived from the acoustic properties of the second space. Alternatively or additionally, in some embodiments, the augmented reality gear further includes a localization subsystem for determining that the virtual sound source is located in the second space, and an audio processing subsystem for processing an audio segment associated with the virtual sound source based on the information regarding the acoustic properties of the second space.
[0097] Although the disclosed embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are to be understood as being included within the scope of the disclosed embodiments as defined by the appended claims.
Claims
1. 1. A method for presenting audio to a user of a mixed reality environment, the method being performed by a mixed reality system comprising one or more processors, the method comprising: the one or more processors determining a first location of the user relative to the mixed reality environment; determining, by the one or more processors, a first acoustic parameter associated with the first location by considering one or more characteristics of a physical environment at the first location; the one or more processors determining a second location of a virtual object relative to the mixed reality environment; The virtual object is identified via an inertial measurement unit and a camera of a wearable head device; and the one or more processors determining an acoustic relationship between the first location and the second location, the first location being in a first room and the second location being in a second room, the acoustic relationship describing an acoustic relationship between the first room and the second room; the one or more processors determining a transfer function based on the first acoustic parameter and the acoustic relationship, the transfer function for simulating propagation of sound from the second location to the first location based on the one or more characteristics of the physical environment; and the one or more processors detecting an audio event associated with the mixed reality environment; the first audio signal includes audio assets for the audio event; the virtual object is associated with the audio event; and the one or more processors applying the transfer function to the first audio signal to generate a second audio signal; the one or more processors presenting to the user via a speaker an audible sound generated based on the second audio signal; A method comprising:
2. The method further includes the one or more processors communicating the first acoustic parameter to a memory at a first time; 2. The method of claim 1, wherein determining the transfer function includes the one or more processors receiving the first acoustic parameters from the memory at a second time that is later than the first time.
3. Determining the acoustic relationship between the first location and the second location includes: the one or more processors identifying whether a path exists for sound propagation between the first location and the second location; if the path for the sound propagation exists, the one or more processors determining that the first location is acoustically coupled to the second location; The method of claim 1 , comprising:
4. Determining the acoustic relationship between the first location and the second location includes:
2. The method of claim 1, further comprising: determining that the first location and the second location are within a common acoustic area and that the first room is the same room as the second room.
5. 2. The method of claim 1, wherein determining the first acoustic parameter includes the one or more processors detecting spatial properties of the first location via sensors of a wearable device associated with the mixed reality environment.
6. The method of claim 1 , wherein determining the second location includes the one or more processors determining the second location based on an output of a sensor of a wearable device associated with the mixed reality environment.
7. The method of claim 1, further comprising the one or more processors determining a second acoustic parameter associated with the second location, and the transfer function being determined further based on the second acoustic parameter.
8. The method of claim 1, further comprising the one or more processors storing the first acoustic parameters and the first location in an acoustic graph, the acoustic graph describing a relationship between two or more locations in the mixed reality environment.
9. one or more sensors; one or more speakers; one or more processors configured to execute a method for presenting audio to a user of a mixed reality environment; 1. A wearable system comprising: determining a first location of the user relative to the mixed reality environment; determining a first acoustic parameter associated with the first location by considering one or more characteristics of a physical environment at the first location; determining a second location of a virtual object relative to the mixed reality environment via the one or more sensors; The virtual object is identified via an inertial measurement unit and a camera of a wearable head device; and determining an acoustic relationship between the first location and the second location, the first location being in a first room and the second location being in a second room, the acoustic relationship describing an acoustic relationship between the first room and the second room; determining a transfer function based on the first acoustic parameter and the acoustic relationship, the transfer function for simulating propagation of sound from the second location to the first location of the user based on the one or more characteristics of the physical environment; and Detecting an audio event associated with the mixed reality environment, the first audio signal includes audio assets for the audio event; the virtual object is associated with the audio event; and applying the transfer function to the first audio signal to generate a second audio signal; presenting to the user via the one or more speakers an audible sound generated based on the second audio signal; and A wearable system comprising:
10. The method further includes communicating the first acoustic parameter to a memory at a first time; The wearable system of claim 9 , wherein determining the transfer function includes receiving the first acoustic parameter from the memory at a second time that is later than the first time.
11. Determining the acoustic relationship between the first location and the second location includes: identifying whether a path exists for sound propagation between the first location and the second location; determining that the first location is acoustically coupled to the second location if the path for the sound propagation exists; and The wearable system of claim 9 , comprising:
12. 10. The wearable system of claim 9, wherein determining the acoustic relationship between the first location and the second location includes determining that the first location and the second location are within a common acoustic area and that the first room is the same room as the second room.
13. The wearable system of claim 9 , wherein determining the first acoustic parameter comprises detecting a spatial property of the first location via the one or more sensors.
14. 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for presenting audio to a user of a mixed reality environment, the method comprising: determining a first location of the user relative to the mixed reality environment; determining a first acoustic parameter associated with the first location by considering one or more characteristics of a physical environment at the first location; determining a second location of a virtual object relative to the mixed reality environment; The virtual object is identified via an inertial measurement unit and a camera of a wearable head device; and determining an acoustic relationship between the first location and the second location, the first location being in a first room and the second location being in a second room, the acoustic relationship describing an acoustic relationship between the first room and the second room; determining a transfer function based on the first acoustic parameter and the acoustic relationship, the transfer function for simulating propagation of sound from the second location to the first location of the user based on the one or more characteristics of the physical environment; and Detecting an audio event associated with the mixed reality environment, the first audio signal includes audio assets for the audio event; the virtual object is associated with the audio event; and applying the transfer function to the first audio signal to generate a second audio signal; presenting to the user via a speaker an audible sound generated based on the second audio signal; and 1. A non-transitory computer-readable medium comprising:
15. 15. The non-transitory computer-readable medium of claim 14, wherein determining the acoustic relationship between the first location and the second location comprises determining that the first location and the second location are in a common acoustic area and that the first room is the same room as the second room.
16. 15. The non-transitory computer-readable medium of claim 14, wherein determining the first acoustic parameter comprises detecting a spatial property of the first location via a sensor of a wearable device associated with the mixed reality environment.
17. 15. The non-transitory computer-readable medium of claim 14, wherein determining the second location comprises determining the second location based on an output of a sensor of a wearable device associated with the mixed reality environment.
Citation Information
Patent Citations
Audio system and its operating method
JP2014505420A
System and method for high-precision 3-dimensional audio for augmented reality
US20120093320A1
Head mounted display and method for providing audio content by using same
US20160088417A1
Spatial audio with remote speakers
US20160212538A1
Method and apparatus for the simulation of complex audio environments
US7099482B1