A method, system and medium for determining and processing audio information

By receiving audio signals in wearable head devices and estimating the reverberation time difference, the problem of insufficient integration between audio signal simulation and the real environment in virtual reality systems is solved, thereby improving immersion and interactivity.

CN114586382BActive Publication Date: 2025-09-23MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202080074331.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-25
Filing Date
2020-10-23
Publication Date
2025-09-23
Estimated Expiration
2040-10-23

AI Technical Summary

Technical Problem

Existing virtual reality systems face problems such as high computational burden, inability to utilize sensory data from the real environment, and difficulty in multi-user interaction when creating immersive experiences, especially the insufficient integration of audio signal simulation with the real environment.

Method used

The audio signal is received through the microphone of the wearable head device, the reverberation time difference is estimated to determine the environmental changes, and based on this, a virtual audio signal is presented to simulate the acoustic characteristics of the real environment.

Benefits of technology

It improves the immersion and interactivity of the virtual environment, reduces motion sickness, enhances the perception of the real environment, and achieves a natural fusion of virtual audio content and real audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114586382B_ABST
    Figure CN114586382B_ABST
Patent Text Reader

Abstract

Examples of the present disclosure describe systems and methods for estimating acoustic properties of an environment. In the example method, a first audio signal is received via a microphone of a wearable head device. An envelope of the first audio signal is determined, and a first reverberation time is estimated based on the envelope of the first audio signal. A difference between the first reverberation time and a second reverberation time is determined. A change in the environment is determined based on the difference between the first reverberation time and the second reverberation time. A second audio signal is presented via a speaker of the wearable head device, wherein the second audio signal is based on the second reverberation time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 926,330, filed October 25, 2019, which is hereby incorporated by reference in its entirety for all purposes. Technical Field

[0003] The present disclosure relates generally to systems and methods for determining and processing audio information, and more particularly, to systems and methods for determining and processing audio information in a mixed reality environment. Background Art

[0004] Virtual environments are ubiquitous in computing environments and are used in video games (where a virtual environment can represent a game world); maps (where a virtual environment can represent a terrain to be navigated); simulations (where a virtual environment can simulate a real-world environment); digital storytelling (where virtual characters can interact in a virtual environment); and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, a user's experience of a virtual environment can be limited by the technology used to render the virtual environment. For example, traditional displays (e.g., 2D screens) and audio systems (e.g., fixed speakers) may not be able to realize a virtual environment in a way that creates an engaging, realistic, and immersive experience.

[0005] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present sensory information to a user of an XR system, the sensory information corresponding to a virtual environment represented by data in a computer system. This disclosure contemplates distinctions between VR, AR, and MR systems (although some systems may be classified as VR in one aspect (e.g., visual) but simultaneously classified as AR or MR in another aspect (e.g., audio). As used herein, a VR system presents a virtual environment that replaces the user's real environment in at least one aspect; for example, a VR system may present a view of the virtual environment to a user while obscuring his or her view of the real environment, such as with a light-blocking head-mounted display. Similarly, a VR system may present audio corresponding to the virtual environment to a user while blocking (attenuating) audio from the real environment.

[0006] VR systems can suffer from various drawbacks caused by replacing a user's real-world environment with a virtual one. One drawback is the feeling of motion sickness that can occur when a user's field of view in the virtual environment no longer corresponds to the state of their inner ears, which monitor a person's balance and orientation in the real (not virtual) environment. Similarly, users can experience disorientation in VR environments where they cannot directly see their body and limbs (the view they use to feel "grounded" in the real world). Another drawback is the computational burden (e.g., storage, processing power) placed on VR systems that must render a full 3D virtual environment, particularly in real-time applications that seek to immerse users in the virtual environment. Similarly, such environments must meet very high standards of realism to be considered immersive, as users are often sensitive to even the slightest imperfections in the virtual environment—any flaw can disrupt the user's sense of immersion. Furthermore, another drawback of VR systems is that these applications are unable to utilize the extensive sensory data of the real environment, such as the various sights and sounds that people experience in the real world. A related disadvantage is that VR systems have difficulty creating shared environments in which multiple users can interact, because users who share physical space in the real world may not be able to directly see or interact with each other in the virtual environment.

[0007] As used herein, an AR system presents a virtual environment that overlaps or overlays the real environment in at least one aspect. For example, an AR system may present a view of the virtual environment to a user that is overlaid on the user's view of the real environment, such as using a transmissive head-mounted display that presents a display image while allowing light to pass through the display into the user's eyes. Similarly, an AR system may present audio corresponding to the virtual environment to the user while being mixed in with audio from the real environment. Similarly, as used herein, like an AR system, an MR system presents a virtual environment that overlaps or overlays the real environment in at least one aspect, and may further allow the virtual environment in the MR system to interact with the real environment in at least one aspect. For example, a virtual character in a virtual environment may flip a light switch in the real environment, causing a corresponding light bulb in the real environment to turn on or off. As another example, the virtual character may react to an audio signal in the real environment (such as with a facial expression). By maintaining a representation of the real world, AR and MR systems can avoid some of the aforementioned drawbacks of VR systems; for example, motion sickness is alleviated because visual cues from the real world (including the user's own body) remain visible, and such systems can immerse the user without presenting a fully realized 3D environment. Furthermore, AR and MR systems can leverage sensory input from the real world (e.g., views and sounds of scenery, objects, and other users) to create new applications that augment that input.

[0008] Ideally, MR systems would interact with as many human senses as possible to create an immersive mixed reality environment for the user. The visual display of virtual content may be important for the mixed reality experience, but audio signals are also valuable for creating immersion in a mixed reality environment. Similar to virtual content that is displayed visually, virtual audio content can also be used to simulate sounds from the real environment. For example, virtual audio content presented in a real environment with echoes can also be presented as echoes, even though the virtual audio content may not actually have echoes in the real environment. This adaptation can help blend virtual content with real content so that the difference between the two is not obvious or even imperceptible to the end user. In order to effectively mix virtual audio content with real audio content, it may be necessary to understand the acoustic characteristics of the real environment so that the virtual audio content can simulate the characteristics of the real audio content. Summary of the Invention

[0009] Examples of the present disclosure describe systems and methods for estimating acoustic properties of an environment. In the example method, a first audio signal is received via a microphone of a wearable head device. An envelope of the first audio signal is determined, and a first reverberation time is estimated based on the envelope of the first audio signal. A difference between the first reverberation time and a second reverberation time is determined. A change in the environment is determined based on the difference between the first reverberation time and the second reverberation time. A second audio signal is presented via a speaker of the wearable head device, wherein the second audio signal is based on the second reverberation time. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1A-Figure 1C An example mixed reality environment is shown in accordance with one or more embodiments of the present disclosure.

[0011] Figure 2A-2D Components of an example mixed reality system that may be used to generate and interact with a mixed reality environment in accordance with one or more embodiments of the present disclosure are shown.

[0012] Figure 3A An example mixed reality handheld controller that can be used to provide input to a mixed reality environment is shown in accordance with one or more embodiments of the present disclosure.

[0013] Figure 3B An example auxiliary unit that may be used with an example mixed reality system in accordance with one or more embodiments of the present disclosure is shown.

[0014] Figure 4 An example functional block diagram of an example mixed reality system is shown in accordance with one or more embodiments of the present disclosure.

[0015] Figure 5An example of estimating a reverberation fingerprint according to one or more embodiments of the present disclosure is shown.

[0016] Figure 6 An example of estimating reverberation time according to one or more embodiments of the present disclosure is shown.

[0017] Figure 7 An example of estimating reverberation time according to one or more embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0018] In the following description of the examples, reference is made to the accompanying drawings which form a part hereof, and in which is shown by way of illustration specific examples that may be practiced. It is to be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.

[0019] Mixed reality environment

[0020] Like all people, users of mixed reality systems exist in a real-world environment—the three-dimensional portion of the "real world" and all its perceptible contents. For example, users perceive the real-world environment using their normal human senses—sight, hearing, touch, taste, and smell—and interact with it by moving their bodies within it. Positions in the real-world environment can be described as coordinates in a coordinate space; for example, coordinates can include latitude, longitude, and altitude relative to sea level; distance in three orthogonal dimensions from a reference point; or other suitable values. Similarly, vectors can describe quantities that have both direction and magnitude in a coordinate space.

[0021] A computing device may, for example, maintain a representation of a virtual environment in a memory associated with the device. As used herein, a virtual environment is a computational representation of a three-dimensional space. The virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other features associated with the space. In some examples, circuitry (e.g., a processor) of a computing device may maintain and update the state of the virtual environment; that is, the processor may determine, at a first time t0, the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or input provided by a user. For example, if an object in the virtual environment is located at a first coordinate at time t0 and has certain programmed physical parameters (e.g., mass, coefficient of friction); and input received from the user indicates that a force should be applied to the object along a direction vector; the processor may apply the laws of kinematics to determine the position of the object at time t1 using basic mechanics. The processor may use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at time t1. In maintaining and updating the state of the virtual environment, the processor may execute any suitable software, including: software relating to creating and deleting virtual objects in the virtual environment; software (e.g., scripts) for defining the behavior of virtual objects or characters in the virtual environment; software for defining the behavior of signals (e.g., audio signals) in the virtual environment; software for creating and updating parameters associated with the virtual environment; software for generating audio signals in the virtual environment; software for processing input and output; software for implementing network operations; software for applying asset data (e.g., animation data that moves a virtual object over time); or many other possibilities.

[0022] An output device, such as a display or speakers, can present any or all aspects of the virtual environment to the user. For example, the virtual environment can include virtual objects (which can include inanimate objects; people; animals; representations of lights, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" having origin coordinates, viewing axes, and a frustum); and present the visual scene of the virtual environment corresponding to that view to the display. Any suitable rendering technique can be used to achieve this. In some examples, the visual scene may include only some virtual objects in the virtual environment and not certain other virtual objects. Similarly, the virtual environment can include audio aspects that can be presented to the user as one or more audio signals. For example, a virtual object in the virtual environment can generate sounds that originate from the object's location coordinates (e.g., a virtual character can speak or cause a sound effect); or the virtual environment can be associated with music cues or background sounds that may or may not be associated with a specific location. The processor can determine an audio signal corresponding to the "listener" coordinates—e.g., a synthesized audio signal corresponding to sounds in a virtual environment, mixed and processed to simulate an audio signal that would be heard by a listener located at the listener coordinates—and present the audio signal to the user via one or more speakers.

[0023] Because the virtual environment exists only as a computational construct, the user cannot directly perceive the virtual environment with their ordinary senses. Instead, the user can only perceive the virtual environment indirectly, for example, by being presented to the user through a display, speakers, tactile output devices, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment; however, input data can be provided to the processor via input devices or sensors, and the processor can use the device or sensor data to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use this data to cause the object in the virtual environment to respond accordingly.

[0024] A mixed reality system may, for example, use a transmissive display and / or one or more speakers (e.g., which may be incorporated into a wearable head device) to present a mixed reality environment ("MRE") to a user that combines aspects of a real environment and a virtual environment. In some embodiments, the one or more speakers may be located external to the head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of a real environment and a corresponding virtual environment. In some examples, the corresponding real environment and the virtual environment share a coordinate space; in some examples, the real coordinate space and the corresponding virtual coordinate space are related to each other by a transformation matrix (or other suitable representation). Thus, a single coordinate (in some examples together with a transformation matrix) may define a first position in the real environment and a corresponding second position in the virtual environment, or vice versa.

[0025] In an MRE, a virtual object (e.g., in a virtual environment associated with the MRE) can correspond to a real object (e.g., in the real environment associated with the MRE). For example, if the real environment of the MRE includes a real lamppost (real object) located at a certain location coordinate, the virtual environment of the MRE can include a virtual lamppost (virtual object) located at the corresponding location coordinate. As used herein, a real object and its corresponding virtual object, combined together, constitute a "mixed reality object." A virtual object does not necessarily have to perfectly match or align with the corresponding real object. In some examples, a virtual object can be a simplified version of the corresponding real object. For example, if the real environment includes a real lamppost, the corresponding virtual object can include a cylinder with approximately the same height and radius as the real lamppost (reflecting that the shape of the lamppost may be roughly cylindrical). Simplifying virtual objects in this way can improve computational efficiency and simplify calculations performed on such virtual objects. Furthermore, in some examples of MREs, not all real objects in the real environment are associated with corresponding virtual objects. Similarly, in some examples of MREs, not all virtual objects in the virtual environment are associated with corresponding real objects. That is, some virtual objects may exist only in the virtual environment of the MRE and have no real-world counterparts.

[0026] In some examples, the characteristics of virtual objects may be different, sometimes drastically different, from those of corresponding real-world objects. For example, when the real-world environment in an MRE includes a green, two-armed cactus (an inanimate object with thorns), the corresponding virtual object in the MRE includes the characteristics of a green, two-armed virtual character with human-like facial features and rude behavior. In this example, the virtual object is similar to its corresponding real-world object in some characteristics (color, number of arms); but different from the real object in other characteristics (human-like facial features, personality). This makes it possible for virtual objects to represent real-world objects in creative, abstract, exaggerated, or fantasy ways; or to assign behaviors (e.g., human personalities) to otherwise inanimate real-world objects. In some examples, virtual objects may be purely fantasy creations with no real-world counterparts (e.g., virtual monsters in a virtual environment, perhaps in locations that correspond to empty spaces in the real environment).

[0027] Compared to VR systems, which present a virtual environment to the user while blurring the real environment, mixed reality systems that present MREs offer the advantage of presenting a virtual environment while maintaining the perceptual presence of the real environment. Consequently, users of mixed reality systems are able to experience and interact with the corresponding virtual environment using visual and auditory cues associated with the real environment. For example, while users of VR systems may have difficulty perceiving or interacting with virtual objects displayed in a virtual environment—because, as mentioned above, users cannot directly perceive or interact with the virtual environment—users of MR systems may find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching the corresponding real objects in their own real environment. This level of interactivity can enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by simultaneously presenting the real and virtual environments, mixed reality systems can reduce the negative psychological experiences (e.g., cognitive dissonance) and negative physical experiences (e.g., motion sickness) associated with VR systems. Mixed reality systems further offer numerous possibilities for applications that may enhance or alter our experience of the real world.

[0028] Figure 1A An example real world environment 100 is shown in which a user 110 uses a mixed reality system 112. The mixed reality system 112 can include a display (e.g., a transmissive display) and one or more speakers; and one or more sensors (e.g., a camera), such as described below. The real world environment 100 shown includes a rectangular room 104A in which the user 110 is standing; and real objects 122A (a lamp), 124A (a table), 126A (a sofa), and 128A (a painting). The room 104A also includes position coordinates 106, which can be considered the origin of the real world environment 100. Figure 1AAs shown, an environment / world coordinate system 108 (including an x-axis 108X, a y-axis 108Y, and a z-axis 108Z) with point 106 (world coordinates) as its origin can define the coordinate space of real environment 100. In some embodiments, origin 106 of environment / world coordinate system 108 can correspond to the location where mixed reality system 112 is turned on. In some embodiments, origin 106 of environment / world coordinate system 108 can be reset during operation. In some examples, user 110 can be considered a real object in real environment 100; similarly, body parts of user 110 (e.g., hands, feet) can be considered real objects in real environment 100. In some examples, a user / listener / head coordinate system 114 (including an x-axis 114X, a y-axis 114Y, and a z-axis 114Z) with point 115 (e.g., user / listener / head coordinates) as its origin can define the coordinate space of the user / listener / head in which mixed reality system 112 resides. The origin 115 of the user / listener / head coordinate system 114 can be defined relative to one or more components of the mixed reality system 112. For example, the origin 115 of the user / listener / head coordinate system 114 can be defined relative to a display of the mixed reality system 112, such as during initial calibration of the mixed reality system 112. A matrix (which can include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can represent the transformation between the space of the user / listener / head coordinate system 114 and the space of the environment / world coordinate system 108. In some embodiments, left ear coordinates 116 and right ear coordinates 117 can be defined relative to the origin 115 of the user / listener / head coordinate system 114. A matrix (which can include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can represent the transformation between the left ear coordinates 116, right ear coordinates 117, and the space of the user / listener / head coordinate system 114. The user / listener / head coordinate system 114 can simplify the representation of a position relative to a user's head or a wearable headset (e.g., relative to the environment / world coordinate system 108). The transformation between the user coordinate system 114 and the environment coordinate system 108 may be determined and updated in real time using simultaneous localization and mapping (SLAM), visual odometry, or other techniques.

[0029] Figure 1BAn example virtual environment 130 corresponding to real environment 100 is shown. Virtual environment 130 includes: a virtual rectangular room 104B corresponding to real rectangular room 104A; a virtual object 122B corresponding to real object 122A; a virtual object 124B corresponding to real object 124A; and a virtual object 126B corresponding to real object 126A. Metadata associated with virtual objects 122B, 124B, and 126B may include information derived from the corresponding real objects 122A, 124A, and 126A. Virtual environment 130 also includes a virtual monster 132, which does not correspond to any real object in real environment 100. Real object 128A in real environment 100 does not correspond to any virtual object in virtual environment 130. A persistent coordinate system 133 (including an x-axis 133X, a y-axis 133Y, and a z-axis 133Z) with point 134 (a persistent coordinate) as its origin may define a coordinate space for virtual content. The origin 134 of the permanent coordinate system 133 can be defined relative to / with respect to one or more real objects (such as real object 126A). A matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can represent the transformation between the space of the permanent coordinate system 133 and the space of the environment / world coordinate system 108. In some embodiments, each of the virtual objects 122B, 124B, 126B, and 132 can have its own permanent coordinate point relative to the origin 134 of the permanent coordinate system 133. In some embodiments, there can be multiple permanent coordinate systems and each of the virtual objects 122B, 124B, 126B, and 132 can have its own permanent coordinate point relative to one or more permanent coordinate systems.

[0030] about Figure 1A and Figure 1B , the environment / world coordinate system 108 defines a shared coordinate space for the real environment 100 and the virtual environment 130. In the example shown, the origin of the coordinate space is located at point 106. In addition, the coordinate space is defined by the same three orthogonal axes (108X, 108Y, 108Z). Therefore, a first position in the real environment 100 and a corresponding second position in the virtual environment 130 can be described with respect to the same coordinate space. This simplifies identifying and displaying corresponding positions in the real and virtual environments because the same coordinates can be used to identify both positions. However, in some examples, the corresponding real and virtual environments do not need to use a shared coordinate space. For example, in some examples (not shown), a matrix (which may include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.

[0031] Figure 1CAn example MRE 150 is shown that simultaneously presents aspects of a real environment 100 and a virtual environment 130 to a user 110 via a mixed reality system 112. In the example shown, the MRE 150 simultaneously presents real objects 122A, 124A, 126A, and 128A from the real environment 100 (e.g., via a transmissive portion of the display of the mixed reality system 112); and virtual objects 122B, 124B, 126B, and 132 from the virtual environment 130 (e.g., via an actively displayed portion of the display of the mixed reality system 112) to the user 110. As described above, the origin 106 serves as the origin of a coordinate space corresponding to the MRE 150, and the coordinate system 108 defines the x-axis, y-axis, and z-axis of the coordinate space.

[0032] In the example shown, the mixed reality objects include corresponding pairs of real objects and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) that occupy corresponding positions in coordinate space 108. In some examples, both the real objects and the virtual objects can be visible to user 110 simultaneously. This is desirable, for example, where the virtual objects present information designed to enhance the view of the corresponding real objects (such as, in a museum application, where the virtual objects present missing portions of an ancient damaged sculpture). In some examples, the virtual objects (122B, 124B, and / or 126B) can be displayed (e.g., via active pixelated occlusion using a pixelated occlusion shutter) to occlude the corresponding real objects (122A, 124A, and / or 126A). This is desirable, for example, where the virtual objects serve as visual stand-ins for the corresponding real objects (such as in an interactive storytelling application where inanimate real objects become "living" characters).

[0033] In some examples, real objects (e.g., 122A, 124A, 126A) can be associated with virtual content or auxiliary data that may not necessarily constitute a virtual object. The virtual content or auxiliary data can facilitate the processing or handling of virtual objects in a mixed reality environment. For example, such virtual content can include a two-dimensional representation of the corresponding real object; a custom asset type associated with the corresponding real object; or statistical data associated with the corresponding real object. This information can enable or facilitate computations involving real objects without incurring unnecessary computational overhead.

[0034] In some examples, the above-described presentation can also incorporate audio aspects. For example, in MRE 150, virtual monster 132 can be associated with one or more audio signals, such as the sound effects of footsteps produced when the monster walks around within MRE 150. As described further below, the processor of mixed reality system 112 can calculate an audio signal corresponding to the mixed and processed synthesis of all such sounds in MRE 150 and present the audio signal to user 110 via one or more speakers included in mixed reality system 112 and / or one or more external speakers.

[0035] Mixed Reality System Example

[0036] An example mixed reality system 112 may include a wearable head device (e.g., a wearable augmented reality or mixed reality head device) that includes a display (which may include a left transmissive display and a right transmissive display that may be near-eye displays, and associated components for coupling light from the displays to the user's eyes); left and right speakers (e.g., located near the user's left and right ears, respectively); an inertial measurement unit (IMU) (e.g., mounted on the temples of the head device); an orthogonal coil electromagnetic receiver (e.g., mounted on the left temple); left and right cameras oriented away from the user (e.g., depth (time of flight) cameras); and left and right eye cameras oriented toward the user (e.g., for detecting eye movements of the user). However, the mixed reality system 112 may incorporate any suitable display technology and any suitable sensor (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic sensors). In addition, the mixed reality system 112 may incorporate network features (e.g., Wi-Fi capabilities) to communicate with other devices and systems (including other mixed reality systems). The mixed reality system 112 may also include a battery (which may be mounted in an auxiliary unit, such as a belt pack designed to be worn around the user's waist), a processor, and memory. The wearable head device of the mixed reality system 112 may include a tracking component, such as an IMU or other suitable sensor, which is configured to output a set of coordinates of the wearable head device relative to the user's environment. In some examples, the tracking component may provide input to a processor that performs simultaneous localization and mapping (SLAM) and / or visual odometry. In some examples, the mixed reality system 112 may also include a handheld controller 300 and / or an auxiliary unit 320, which may be a wearable belt pack, as further described below.

[0037] Figure 2A-2D Components of an example mixed reality system 200 (which may correspond to mixed reality system 112 ) that may be used to present an MRE (which may correspond to MRE 150 ) or other virtual environment to a user are shown. Figure 2AA perspective view of a wearable head device 2102 included in an example mixed reality system 200 is shown. Figure 2B A top view of a wearable head device 2102 worn on a user's head 2202 is shown. Figure 2C A front view of the wearable head device 2102 is shown. Figure 2D A side view of an example eyepiece 2110 of a wearable head device 2102 is shown. Figure 2A-2C As shown, the example wearable head device 2102 includes an example left eyepiece (e.g., a left transparent waveguide group eyepiece) 2108 and an example right eyepiece (e.g., a right transparent waveguide group eyepiece) 2110. Each eyepiece 2108 and 2110 may include a transmissive element through which a real environment is viewed, and a display element for presenting a display that overlaps with the real environment (e.g., via imaging modulated light). In some examples, such a display element may include a surface diffraction optical element for controlling the optical flow of the imaging modulated light. For example, the left eyepiece 2108 may include a left coupling grating group 2112, a left orthogonal pupil expansion (OPE) grating group 2120, and a left exit (output) pupil expansion (EPE) grating group 2122. Similarly, the right eyepiece 2110 may include a right coupling grating group 2118, a right OPE grating group 2114, and a right EPE grating group 2116. The imaging modulated light can be transmitted to the user's eyes via the coupling-in gratings 2112 and 2118, the OPEs 2114 and 2120, and the EPEs 2116 and 2122. Each coupling-in grating group 2112, 2118 can be configured to deflect light toward its corresponding OPE grating group 2120, 2114. Each OPE grating group 2120, 2114 can be designed to gradually direct light downward toward its associated EPE 2122, 2116 sheet, thereby horizontally expanding the exit pupil being formed. Each EPE 2122, 2116 can be configured to gradually redirect at least a portion of the light received from its corresponding OPE grating group 2120, 2114 outward toward the user's eye-tracking position (not shown) defined behind the eyepieces 2108, 2110, vertically expanding the exit pupil formed within the eye-tracking range. Alternatively, instead of coupling-in grating groups 2112 and 2118, OPE grating groups 2114 and 2120, and EPE grating groups 2116 and 2122, the eyepieces 2108 and 2110 may include other arrangements of gratings and / or refractive and reflective features for controlling the coupling of imaging modulated light to the user's eye.

[0038] In some examples, the wearable head unit 2102 may include a left temple 2130 and a right temple 2132, wherein the left temple 2130 includes a left speaker 2134 and the right temple 2132 includes a right speaker 2136. An orthogonal coil electromagnetic receiver 2138 may be located in the left temple piece, or in another suitable location on the wearable head unit 2102. An inertial measurement unit (IMU) 2140 may be located in the right temple 2132, or in another suitable location on the wearable head unit 2102. The wearable head unit 2102 may also include a left depth (e.g., time of flight) camera 2142 and a right depth camera 2144. The depth cameras 2142 and 2144 may be appropriately oriented in different directions so as to collectively cover a wider field of view.

[0039] exist Figure 2A-2D In the example shown, a left imaging modulated light source 2124 can be optically coupled into the left eyepiece 2108 via a left incoupling grating set 2112, and a right imaging modulated light source 2126 can be optically coupled into the right eyepiece 2110 via a right incoupling grating set 2118. The imaging modulated light sources 2124 and 2126 can include, for example, a fiber scanner; a projector including an electronic light modulator, such as a digital light processing (DLP) chip or a liquid crystal on silicon (LCoS) modulator; or an emissive display, such as a micro-light emitting diode (μLED) or micro-organic light emitting diode (μOLED) panel, coupled to the incoupling grating sets 2112 and 2118 using one or more lenses on each side. The incoupling grating sets 2112 and 2118 can deflect light from the imaging modulated light sources 2124 and 2126 to an angle greater than the critical angle for total internal reflection (TIR) ​​of the eyepieces 2108 and 2110. The OPE grating groups 2114, 2120 gradually deflect light propagating via TIR downwardly towards the EPE grating groups 2116, 2122. The EPE grating groups 2116, 2122 gradually couple light into the user's face, including the pupils of the user's eyes.

[0040] In some examples, such as Figure 2DAs shown, each of the left eyepiece 2108 and the right eyepiece 2110 includes a plurality of waveguides 2402. For example, each eyepiece 2108, 2110 can include a plurality of separate waveguides, each dedicated to a corresponding color channel (e.g., red, blue, and green). In some examples, each eyepiece 2108, 2110 can include multiple groups of such waveguides, each group of waveguides being configured to impart a different wavefront curvature to the emitted light. The wavefront curvature can be convex relative to the user's eye to, for example, present a virtual object located a distance in front of the user (e.g., a distance corresponding to the inverse of the wavefront curvature). In some examples, the EPE grating groups 2116, 2122 can include curved grating grooves to achieve a convex wavefront curvature by changing the Poynting vector of the exiting light passing through each EPE.

[0041] In some examples, to create the perception that the displayed content is three-dimensional, stereoscopically adjusted left-eye and right-eye images can be presented to the user via imaging light modulators 2124, 2126 and eyepieces 2108, 2110. The realism of the three-dimensional virtual object presentation can be enhanced by selecting the waveguides (and therefore the corresponding wavefront curvature) so that the virtual objects appear at a distance approximately equal to the distance indicated by the stereoscopic left and right images. This technique can also reduce motion sickness experienced by some users, which is caused by the difference between the depth perception cues provided by the stereoscopic left and right eye images and the human eye's autonomous accommodation (e.g., focusing dependent on object distance).

[0042] Figure 2D 2 shows a side view from the top of the right eyepiece 2110 of the example wearable head device 2102. Figure 2D As shown, the plurality of waveguides 2402 may include a first subset 2404 of three waveguides and a second subset 2406 of three waveguides. The two subsets 2404 and 2406 of waveguides may be distinguished by different EPE gratings, each of which has a different grating line curvature to impart a different wavefront curvature to the outgoing light. Within each waveguide subset 2404 and 2406, each waveguide may be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although Figure 2D (Not shown, but the structure of the left eyepiece 2108 is similar to that of the right eyepiece 2110.)

[0043] Figure 3AAn example handheld controller component 300 of the mixed reality system 200 is shown. In some examples, the handheld controller 300 includes a grip portion 346 and one or more buttons 350 disposed along a top surface 348. In some examples, the buttons 350 can be configured to function as optical tracking targets, for example, in conjunction with a camera or other optical sensor (which can be mounted in a head unit of the mixed reality system 200 (e.g., wearable head device 2102)) to track six degrees of freedom (6DOF) motion of the handheld controller 300. In some examples, the handheld controller 300 includes a tracking component (e.g., an IMU or other suitable sensor) for detecting position or orientation, such as relative to the wearable head device 2102. In some examples, such a tracking component can be located in the handle of the handheld controller 300 and / or can be mechanically coupled to the handheld controller. The handheld controller 300 can be configured to provide one or more output signals (e.g., via an IMU) corresponding to one or more of a button press; or a position, orientation, and / or movement of the handheld controller 300. Such output signals can be used as input to a processor of mixed reality system 200. Such input can correspond to the position, orientation, and / or movement of a handheld controller (and, by extension, the position, orientation, and / or movement of a user's hand holding the controller). Such input can also correspond to a user pressing button 350.

[0044] Figure 3B An example auxiliary unit 320 of the mixed reality system 200 is shown. The auxiliary unit 320 can include a battery to provide energy to operate the system 200 and can include a processor for executing programs to operate the system 200. As shown, the example auxiliary unit 320 includes a clip 2128, such as for attaching the auxiliary unit 320 to a user's belt. Other form factors are suitable for the auxiliary unit 320 and will be apparent, including form factors that do not involve mounting the unit to a user's belt. In some examples, the auxiliary unit 320 is coupled to the wearable head device 2102 via, for example, a multi-conductor fiber optic cable, which can include both electrical wires and optical fibers. A wireless connection between the auxiliary unit 320 and the wearable head device 2102 can also be used.

[0045] In some examples, the mixed reality system 200 may include one or more microphones to detect sound and provide corresponding signals to the mixed reality system. In some examples, the microphone may be attached to or integrated with the wearable head device 2102 and may be configured to detect the user's voice. In some examples, the microphone may be attached to or integrated with the handheld controller 300 and / or the auxiliary unit 320. Such microphones may be configured to detect ambient sound, background noise, the user's or a third party's voice, or other sounds.

[0046] Figure 4 1 shows an example functional block diagram that may correspond to an example mixed reality system, such as mixed reality system 200 described above (which may correspond to mixed reality system 112 with respect to FIG. 1 ). Figure 4As shown, an example handheld controller 400B (which may correspond to handheld controller 300 ("Totem")) includes a totem to wearable head device six degrees of freedom (6DOF) totem subsystem 404A, and an example wearable head device 400A (which may correspond to wearable head device 2012) includes a totem to wearable head device 6DOF subsystem 404B. In this example, the 6DOF totem subsystem 404A and the 6DOF subsystem 404B cooperate to determine six coordinates (e.g., offsets in three translational directions and rotations along three axes) of the handheld controller 400B relative to the wearable head device 400A. The six degrees of freedom can be expressed relative to the coordinate system of the wearable head device 400A. The three translational offsets can be expressed as X, Y, and Z offsets in such a coordinate system, as a translation matrix, or as some other representation. The rotational degrees of freedom can be expressed as a sequence of yaw, pitch, and roll rotations, as a rotation matrix, as a quaternion, or as some other representation. In some examples, the wearable head device 400A; one or more depth cameras 444 (and / or one or more non-depth cameras) included in the wearable head device 400A; and / or one or more optical targets (e.g., the button 350 of the handheld controller 400B as described above, or a dedicated optical target included in the handheld controller 400B) can be used for 6DOF tracking. In some examples, the handheld controller 400B can include a camera, as described above; and the wearable head device 400A can include an optical target for optical tracking in conjunction with the camera. In some examples, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids for wirelessly sending and receiving three distinguishable signals. By measuring the relative amplitudes of the three distinguishable signals received in each coil used for reception, the 6DOF of the wearable head device 400A relative to the handheld controller 400B can be determined. Additionally, the 6DOF totem subsystem 404A may include an inertial measurement unit (IMU) that may be used to provide improved accuracy and / or more timely information regarding rapid movements of the handheld controller 400B.

[0047] In some examples, it may be necessary to transform coordinates from a local coordinate space (e.g., a coordinate space that is fixed relative to the wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space that is fixed relative to the real environment), for example, to compensate for movement of the wearable head device 400A relative to the coordinate system 108. For example, such a transformation may be necessary for the display of the wearable head device 400A to present a virtual object (e.g., a virtual person sitting in a real chair, facing forward, regardless of the position and orientation of the wearable head device) at an expected position and orientation relative to the real environment, rather than at a fixed position and orientation on the display (e.g., the same position in the lower right corner of the display), thereby maintaining the illusion that the virtual object exists in the real environment (and does not appear to be unnaturally positioned in the real environment when the wearable head device 400A moves and rotates). In some examples, the compensating transformation between coordinate spaces can be determined by processing imagery from the depth camera 444 using SLAM and / or visual odometry programs to determine the transformation of the wearable head device 400A relative to the coordinate system 108. Figure 4 In the example shown, a depth camera 444 is coupled to the SLAM / visual odometry block 406 and can provide imagery to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the imagery and determine the position and orientation of the user's head, which can then be used to identify a transformation between the head coordinate space and another coordinate space (e.g., an inertial coordinate space). Similarly, in some examples, an additional source of information about the user's head pose and position is obtained from the IMU 409. Information from the IMU 409 can be combined with information from the SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information for rapid adjustments to the user's head pose and position.

[0048] In some examples, depth camera 444 can provide 3D imagery to gesture tracker 411, which can be implemented in a processor of wearable head device 400A. Gesture tracker 411 can recognize a user's gestures, for example, by matching the 3D imagery received from depth camera 444 with stored patterns representing gestures. Other suitable techniques for recognizing user gestures will be apparent.

[0049] In some examples, one or more processors 416 can be configured to receive data from the wearable head device's 6DOF head-mounted subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, and / or gesture tracker 411. Processor 416 can also send and receive control signals from 6DOF totem system 404A. Processor 416 can be wirelessly coupled to 6DOF totem system 404A, such as in an example where handheld controller 400B is disconnected. Processor 416 can further communicate with other components, such as audio-visual content storage 418, graphics processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. DSP audio spatializer 422 can be coupled to head-related transfer function (HRTF) storage 425. GPU 420 can include a left channel output coupled to a left imaging modulated light source 424 and a right channel output coupled to a right imaging modulated light source 426. GPU 420 can output stereoscopic image data to imaging modulated light sources 424, 426, for example as described above with respect to Figure 2A-2D As described. The DSP audio spatializer 422 can output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 can receive an input from the processor 419 indicating a direction vector from the user to the virtual sound source (which can be moved by the user, for example, via the handheld controller 320). Based on the direction vector, the DSP audio spatializer 422 can determine a corresponding HRTF (e.g., by accessing the HRTF, or by interpolating multiple HRTFs). The DSP audio spatializer 422 can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. This can enhance the credibility and realism of the virtual sound by incorporating the user's relative position and orientation with respect to the virtual sound in the mixed reality environment, that is, by matching the presented virtual sound with the user's expectation that the virtual sound will sound like a real sound in the real environment.

[0050] In some examples, such as Figure 4 As shown, one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio-visual content memory 418 may be included in an auxiliary unit 400C (which may correspond to the auxiliary unit 320 described above). The auxiliary unit 400C may include a battery 427 to power its components and / or power the wearable head device 400A or the handheld controller 400B. By including these components in an auxiliary unit that can be mounted on the user's waist, the size and weight of the wearable head device 400A can be limited, thereby reducing fatigue on the user's head and neck.

[0051] Although Figure 4 Elements corresponding to various components of an example mixed reality system are presented, but various other suitable arrangements of these components will become apparent to those skilled in the art. For example, Figure 4 Elements presented as being associated with auxiliary unit 400C may instead be associated with wearable head device 400A or handheld controller 400B. In addition, some mixed reality systems may completely forgo handheld controller 400B or auxiliary unit 400C. Such changes and modifications are to be understood as being included within the scope of the disclosed examples.

[0052] Reverberation fingerprint estimation

[0053] Presenting virtual audio content to the user facilitates creating an immersive augmented / mixed reality experience. By presenting convincing audio and convincing video, the immersive augmented / mixed reality experience can further blend real content with virtual content. Displaying convincing virtual video content (e.g., aligned with and / or inseparable from real content) can include constructing a map of the real (sometimes unknown) environment while estimating the position and orientation of the MR system within the real environment to accurately display the virtual video content within the real environment. Displaying convincing virtual video content can also include rendering two identical sets of virtual video content from two different perspectives, so that stereo images can be presented to the user to simulate three-dimensional virtual video content. Similar to displaying convincing virtual video content, presenting virtual audio content in a convincing manner can also involve complex analysis of the real environment. For example, it may be necessary to understand the acoustic characteristics of the real environment in which the MR system is used so that the virtual audio content can be presented in a manner that simulates the real audio content. The MR system (e.g., MR system 112, 200) can use the acoustic characteristics of the real environment to modify the rendering algorithm so that the virtual audio content sounds as if it originates from or otherwise belongs to the real environment. For example, an MR system used in a room with a hard floor and exposed walls may produce virtual audio content that simulates the echoes that real audio content may have. As the user changes the real environment (which may have different acoustic properties), playing the virtual audio content in a static manner may reduce the immersiveness of the experience. If the real audio content and the virtual audio content can interact with each other (for example, the user can talk to a virtual companion, and the virtual companion can talk back to the user), to do this, the MR system can determine the acoustic properties of the real environment and apply these acoustic properties to the virtual audio content (for example, by changing the rendering algorithm of the virtual audio content). Additional details can be found in U.S. patent application No. 16 / 163,529, the entire contents of which are incorporated herein by reference.

[0054] One parameter that can characterize the acoustic properties of a real-world environment is reverberation time (e.g., T60 time). Reverberation time can include the length of time required for a sound to decay by a certain amount (e.g., by 60 decibels). Sound decay is the result of sound reflecting off surfaces in a real-world environment (e.g., walls, floors, furniture, etc.) while losing energy due to, for example, geometric propagation. Reverberation time can be affected by environmental factors. For example, absorptive surfaces (e.g., cushions) can absorb sound in addition to geometric propagation, thereby reducing reverberation time. In some embodiments, it may not be necessary to have information about the original source to estimate the reverberation time of the environment.

[0055] Another parameter that can characterize the acoustic properties of a real-world environment is reverberation gain. The reverberation gain can include the ratio of the direct / source / original energy of a sound to the reverberation energy of the sound (e.g., the energy of the reverberation produced by the direct / source / original sound), where the listener and the source are substantially co-located (e.g., a user can clap their hands, generating a source sound that can be considered to be substantially co-located with one or more microphones mounted on a head-mounted MR system). For example, an impulse (e.g., a clap) can have energy associated with the impulse, and the reverberation sound from the impulse can have energy associated with the reverberation of the impulse. The ratio of the original / source energy to the reverberation energy can be the reverberation gain. The reverberation gain of a real-world environment can be affected by, for example, absorptive surfaces that can absorb sound and thereby reduce the reverberation energy.

[0056] Reverberation time and reverberation gain can be collectively referred to as a reverberation fingerprint. In some embodiments, the reverberation fingerprint can be passed as one or more input parameters to an audio rendering algorithm, which can allow the audio rendering algorithm to render virtual audio content with the same or similar characteristics as real audio content in a real environment.

[0057] Reverberation fingerprints can be useful because they can characterize the acoustic properties of a real-world environment, regardless of the location and / or orientation of the sound sources in the real-world environment. For example, a standard indoor room with four walls, a floor, and a ceiling can exhibit the same (or substantially the same) reverberation time and / or reverberation gain, regardless of whether the source is located in a corner of the room, in the center of the room, or along any wall / edge of the room. As another example, according to the reverberation fingerprint of the real-world environment, sound sources directly facing the corner of the room, the center of the room, or a wall of the room all behave the same (or substantially the same). Reverberation fingerprints are also useful because they can characterize the acoustic properties of a real-world environment, regardless of the characteristics of the sound source. For example, low-frequency, mid-frequency, or high-frequency sound sources (e.g., a person speaking) can all behave the same (or substantially the same) according to the reverberation time and / or reverberation gain of the real-world environment. Similarly, impulse sound sources (e.g., clapping) and non-impulse sound sources can behave the same (or substantially the same) according to the reverberation fingerprint (e.g., reverberation time and / or reverberation gain) of the real-world environment. As another example, loud sound sources and quiet sound sources (e.g., in terms of amplitude) may appear identical (or substantially identical) based on the reverberation fingerprint of the real environment (e.g., reverberation time and / or reverberation gain). The fact that the reverberation fingerprint is independent of the characteristics and / or location of the sound source may make the reverberation fingerprint a useful tool for rendering virtual audio content in a computationally efficient manner (e.g., as long as the user does not change the environment by moving to a different room, the rendering algorithm can be the same). In some embodiments, the reverberation fingerprint is applicable to "regular" rooms (e.g., a standard indoor room with four walls, a floor, and a ceiling) and is not applicable to "irregular" rooms (e.g., long corridors) that may have special acoustic properties.

[0058] In some embodiments, it may be desirable to perform a "blind" estimation of the reverberation fingerprint of a real-world environment. A blind estimation is an estimation of the reverberation fingerprint in which no information about the sound source is required. For example, the reverberation fingerprint can be estimated simply based on human conversation, where information about the original speech may not be provided to the estimation algorithm. Pauses during human conversation can provide sufficient time to estimate the reverberation fingerprint using blind estimation. Performing a blind estimation is beneficial because it can be accomplished without requiring lengthy setup procedures and / or user interaction. In some embodiments, the reverberation time can be estimated blindly and may not require information about the original sound source. In some embodiments, blind estimation of the reverberation gain may not be performed, and the reverberation gain may include information about the original sound source.

[0059] Figure 5An example process 500 for estimating a reverberation fingerprint according to some embodiments is shown. The example process shown can be implemented using one or more components of a mixed reality system, such as the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above, or by a system in communication with the mixed reality system 200 (e.g., a system including a cloud server). In step 502 of process 500, input 501 can be separated into one or more filtered components, which can then be processed separately. For example, in step 502, a bandpass filter can be applied to input 501, which can be an audio signal from one or more microphones (e.g., one or more microphones mounted on an MR system). The bandpass filter can preferentially allow certain frequency ranges to pass through the filter and / or suppress frequencies outside of the frequency range. The bandpass filter can decompose the signal into smaller, more easily processed components to improve computational efficiency. The bandpass filter can also improve the signal-to-noise ratio of the signal by removing unwanted noise at frequencies outside of the frequency range. In some embodiments, the bandpass filter can be used to separate the audio signal into six frequency ranges. A reverberation fingerprint (e.g., reverberation time and reverberation gain) can be estimated for each frequency range. This can be used to create a continuous frequency response curve so that each frequency can have an associated reverberation time and / or reverberation gain (e.g., the reverberation time and / or reverberation gain can be interpolated from calculated values ​​centered within the frequency range separated by the bandpass filter). Although six frequency ranges are discussed, the audio signal can be divided into any number of frequency ranges (e.g., using any number of bandpass filters). In some embodiments, an octave filter can be applied to the input signal. In some embodiments, a 1 / 3 octave filter can be applied to the input signal. In some embodiments, signals with frequencies that are too low (e.g., less than 100 Hz) may not be analyzed for reverberation fingerprinting (e.g., because the low frequencies are not reverberant enough for reverberation fingerprint analysis).

[0060] At step 504, frequency band boosting may optionally be applied. Frequency band boosting may be applied to low frequencies that have a low signal-to-noise ratio (e.g., less than 500 Hz), but the signal-to-noise ratio may still be high enough to determine a reverberation fingerprint (e.g., the signal-to-noise ratio may be higher than the signal-to-noise ratio for frequencies less than 100 Hz). Frequency band boosting may be applied to other frequency bands, or not applied at all.

[0061] At step 506, an energy estimation may be performed on the signal. The energy estimation may be performed in the frequency domain, the time domain, the spectral domain, and / or any other suitable domain. The signal energy may be estimated in the time domain by determining the area under the squared magnitude of the signal or by using other suitable methods.

[0062] At step 508, envelope detection can be performed on the signal and the envelope detection can be based on (an estimate of) the running energy of the signal. The signal envelope can be a representation of the peaks and / or valleys of the signal and can define the upper and / or lower boundaries of a signal (e.g., an oscillating signal). Envelope detection can be performed using a Hilbert transform, a leaky integrator-based RMS detector, and / or other suitable methods.

[0063] Peak picking may be run on the signal envelope at step 510. Peak picking may identify local peaks in the signal envelope based on the magnitude of previously detected peaks and / or based on local maxima.

[0064] At step 512, a free decay region estimation can be run on the signal envelope. A free decay region can be a region of the signal envelope where the envelope decreases (e.g., after a local peak). This can be a result of reverberation, where new sounds may not be detected, only previous sounds continue to reverberate in the real environment, causing the signal envelope to decrease. At step 512, a linear fit can be determined for each of the one or more free decay regions in the signal. A linear fit may be appropriate where the signal envelope is measured on a decibel scale due to the exponential decay of acoustic energy and the decibel scale is measured on a logarithmic scale.

[0065] At step 514, reverberation time can be estimated. The reverberation time can be estimated based on the free decay region or the portion of the free decay region with the fastest decay slope, which can be determined based on a linear fit determined for each free decay region (or portion of the free decay region). In some embodiments, a threshold amount of time (e.g., 50 ms) after a local peak can be ignored when determining the linear fit. This helps avoid short-term reverberation (which may behave differently) and / or helps ensure that the regression is only fitted to the reverberant sound and not the source sound. The linear fit slope can represent the amount of decrease (in decibels) of the signal envelope per unit time (e.g., per second).

[0066] In some embodiments, multiple linear fits can be applied to a single free decay region. For example, a linear regression can be applied only within a time range where the regression is sufficiently accurate (e.g., a correlation of 97% or greater). If the linear regression no longer fits the remaining portion of the free decay region duration, one or more additional / alternative linear regressions can be applied. The accuracy of the reverberation time estimate can be increased by using only the fastest decay slope within the free decay region, as the relevant portion of the free decay region can most accurately represent only the reverberant sound. For example, a portion of the free decay region with a slower decay slope may capture a small amount of non-reverberant (e.g., original / source) sound, which may artificially slow the measured decay rate. Based on the fastest decaying linear fit slope, the reverberation time (which may be the time required for the signal to decay by 60 dB) can be extrapolated.

[0067] Figure 6 An example process 600 for estimating reverberation time is shown. Example process 600 may correspond to step 514 of example process 500 described above. Example process 600 may be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of example mixed reality system 200 described above, or by a system in communication with mixed reality system 200 (e.g., a system including a cloud server). In step 602 of example process 600, a local peak (e.g., a local peak from a signal envelope) may be determined. In step 604, a linear regression may be fitted to part or all of a free decay region. A free decay region may be a region of the signal envelope where the envelope decreases (e.g., after a local peak). In some embodiments, the linear regression may not consider a portion of time after a local peak (e.g., 50 ms after the local peak). In step 608, a determination may be made as to whether the linear fit is sufficiently accurate (e.g., having a sufficiently low root mean square error). If the linear fit is determined to be insufficiently accurate, the next free decay region or portion of the free decay region may be examined in step 609. If the linear fit is determined to be sufficiently accurate, then a determination may be made at step 610 as to whether the decay region occurs for a sufficiently long period of time (e.g., >400 ms). If it is determined that the decay region does not occur for a sufficiently long period of time, then the next free decay region or portion of a free decay region may be examined at step 609. If it is determined that the decay region does occur for a sufficiently long period of time, then a determination may be made at step 612 as to whether the decay slope from the linear regression is the fastest decay slope of the entire free decay region. If it is determined that the decay slope is not the fastest decay slope of the entire free decay region, then the next free decay region or portion of a free decay region may be examined at step 609. If it is determined that the decay slope is the fastest decay slope of the entire free decay region, then the reverberation time may be extrapolated based on the fastest decay slope at step 614.

[0068] In some embodiments, a convergence (or near-convergence) measure can be used to estimate reverberation time. For example, the reverberation time can be declared after the decay slopes of a threshold number of consecutive free decay regions are within a threshold of each other. An average decay slope can then be determined and declared as the reverberation time. In some embodiments, the decay slopes associated with the free decay regions can be weighted based on the quality estimate of each measured decay slope. In some embodiments, a decay slope can be determined to be more accurate when the relevant portion of the free decay region lasts for a threshold amount of time (e.g., 400 ms), which can increase the accuracy of the decay slope estimate. In some embodiments, a decay slope can be determined to be more accurate if the decay slope has a relatively accurate linear fit (e.g., a low root mean square error). The more accurate decay slope can be assigned a higher weight in the weighted average to determine the reverberation time. In some embodiments, the single decay slope determined to be most accurate (e.g., based on decay length and / or linear fit accuracy) can be used to determine the reverberation time, which can be the reverberation time for a given frequency range (e.g., the frequency range selected by the bandpass filter in step 502).

[0069] Return Reference Figure 5 Following process 500, at step 514, a confidence value can be determined and associated with the reverberation time. The confidence value can be determined based on various factors. For example, the confidence value can be based on the number of converged decay slopes, the linear fit accuracy of the decay slopes utilized, the decay length of the decay slopes utilized, the difference between the new reverberation time estimate and the previous reverberation time estimate, or any combination of these and / or other factors. In some embodiments, if the confidence value is below a threshold (e.g., because insufficient free decay region for convergence is detected), the reverberation time estimate with the associated confidence value may not be announced. If a reverberation time estimate is not announced, other reverberation time estimates for other frequency ranges (e.g., the frequency range separated using the bandpass filter in step 502) may still be announced (e.g., if those reverberation time estimates have sufficiently high confidence values). The reverberation time estimates for the missing frequency ranges can be interpolated based on the reverberation times announced for the other frequency ranges.

[0070] At step 516, a direct sound energy estimation may be performed. The direct sound energy estimation may utilize information about the direct / source sound. For example, if the direct / source sound is known, the direct sound energy estimation may estimate the energy of the direct / source sound (e.g., by integrating the area under the peak of the signal envelope that includes the direct / source sound). This may be accomplished using impulse sounds, which may make it easier to separate the direct / source sound from the reverberant sound. In some embodiments, the user may be prompted (e.g., via the MR system) to clap their hands to produce an impulse sound. In some embodiments, a speaker, such as a speaker mounted on the MR system, may play the impulse sound. In some embodiments, the impulse sound may be used to estimate the direct sound energy and reverberation time estimate. In some embodiments, the direct sound estimation may be estimated blindly (e.g., if the blind estimation can separate the direct / source sound from the reverberant sound without prior knowledge of the direct / source sound).

[0071] The reverberant sound energy may be estimated at step 518. The reverberant sound energy may be estimated by integrating the signal envelope from the direct / source sound end until the reverberant sound is no longer detected and / or the reverberant sound is below a certain gain threshold (eg, -90 dB).

[0072] At step 520, a reverberation gain can be estimated based on the direct sound energy estimate and the reverberation energy estimate. In some embodiments, the reverberation gain is calculated by taking the ratio of the reverberation energy to the direct sound energy. In some embodiments, the reverberation gain is calculated by taking the ratio of the direct sound energy to the reverberation energy. The reverberation gain estimate can be declared (e.g., passed to an audio rendering algorithm). In some embodiments, a confidence level can be associated with the reverberation gain estimate. For example, if a peak is detected in the reverberation energy estimate, it may indicate that a new direct / source sound has been introduced and the reverberation gain estimate may no longer be accurate. In some embodiments, the reverberation gain estimate can only be declared if the confidence level is at or above a certain threshold.

[0073] In addition to using reverberation fingerprints to more realistically render virtual audio content, reverberation fingerprints can also be used to identify real-world environments and / or identify changes in real-world environments. For example, a user may calibrate an MR system in a first room (e.g., a first acoustic environment) and then move to a second room. The second room may have different acoustic properties than the first room (e.g., a different reverberation time and / or a different reverberation gain). The MR system may blindly estimate the reverberation time in the second room, determine that the reverberation time is sufficiently different from the previously announced reverberation time, and conclude that the user has changed rooms. The MR system may then announce a new reverberation time and / or a new reverberation gain (e.g., by asking the user to clap their hands again, playing an impulse through an external speaker, and / or blindly estimating the reverberation gain). As another example, a user may calibrate the MR system in a room, and the MR system may determine the reverberation fingerprint of the room. The MR system may then identify the room based on the reverberation fingerprint and / or other factors (e.g., location determined via GPS and / or WiFi networks, or via one or more sensors such as those described above with respect to the example mixed reality system 200). The MR system can access a remote database of previously constructed rooms and use reverberation fingerprints and / or other factors to identify the room as a previously constructed room. The MR system can download assets associated with the room (e.g., a previously generated 3D map of the room).

[0074] Figure 7An example process for identifying changes in the acoustic properties of a real environment is shown. The example process shown can be implemented using one or more components of a mixed reality system, such as one or more of the wearable head device 2102, handheld controller 300, and auxiliary unit 320 of the example mixed reality system 200 described above, or by a system in communication with the mixed reality system 200 (e.g., a system including a cloud server). In step 702 of the example process, a new reverberation time can be determined (e.g., using process 500 and / or process 600). In step 704, the new reverberation time can be compared to the previously announced reverberation time. In step 706, a determination can be made as to whether the new reverberation time is sufficiently different from the previously announced reverberation time. The difference can be assessed in a variety of ways. For example, if the new reverberation time for a frequency range differs from the announced reverberation time for the frequency range by more than a specified threshold (e.g., 10%, which may be sufficiently different for a human listener to perceive the difference), then the difference may be sufficient. As another example, a sufficient difference may be determined if the reverberation time for a given frequency range differs by a threshold amount from the announced reverberation time for that frequency range. As another example, the absolute value of the difference between the new frequency response curve (which may include interpolated points between the declared reverberation times for the tested frequency range) and the declared frequency response curve may be integrated. If the integrated area is above a certain threshold, it may be determined that the new reverberation time is sufficiently different from the declared reverberation time.

[0075] If it is determined that the new reverberation time is not sufficiently different from the announced reverberation time, the MR system may continue to determine a new reverberation time at step 702. If it is determined that the new reverberation time is sufficiently different from the announced reverberation time, it may be determined at step 708 whether a sufficient number of sufficiently different reverberation times have been detected. For example, for a given frequency range, three consecutive reverberation time estimates that are all different from the announced reverberation item may be a sufficient number of sufficiently different reverberation times. Other thresholds (e.g., three out of five most recent reverberation time estimates) may also be used. If it is determined that a sufficient number of sufficiently different reverberation times have not been detected, the MR system may continue to determine a new reverberation time at step 702. If it is determined that a sufficient number of sufficiently different reverberation times have been detected, the new reverberation time may be announced at step 710. In some embodiments, step 710 may also include initiating a new reverberation gain estimate, which may prompt the user to clap their hands or play an impulse sound from an external speaker. In some embodiments, step 710 may also include accessing a remote database to identify a new real environment based on the new reverberation fingerprint and / or other information available to the MR system (e.g., a location determined via a GPS and / or WiFi connection, or via one or more sensors such as described above with respect to the example mixed reality system 200).

[0076] Although the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. For example, one or more elements of the implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are understood to be included within the scope of the disclosed examples as defined by the appended claims.

Claims

1. A method for determining and processing audio information, comprising: receiving, at a first time, a first audio signal via a microphone of a wearable head device configured to present a view of a virtual environment; determining an envelope of the first audio signal; estimating a first reverberation time based on the envelope of the first audio signal; determining, based on the estimated first reverberation time, that a position of the wearable head device at the first time corresponds to a first area of ​​the virtual environment; receiving, via the microphone of the wearable head device, a second audio signal at a second time; determining an envelope of a second time signal; estimating a second reverberation time based on the envelope of the second audio signal; as well as determining, based on the estimated second reverberation time, that the position of the wearable head device at the second time corresponds to a second area of ​​the virtual environment, the second area being different from the first area; as well as Based on a difference between the estimated first reverberation time and the estimated second reverberation, it is determined that the first area is in a first room and the second area is in a second room different from the first room.

2. The method according to claim 1, wherein The estimating the first reverberation time includes determining whether the envelope of the first audio signal decays for a time greater than a threshold amount of time.

3. The method according to claim 1, wherein The estimating the first reverberation time comprises: determining a linear fit of a decaying region in the envelope of the first audio signal; and A determination is made as to whether the correlation of the linear fit is greater than a threshold correlation.

4. The method according to claim 1, further comprising: A first reverberation gain is estimated based on the envelope of the first audio signal, wherein determining that the position of the wearable head device at the first time corresponds to the first area of ​​the virtual environment is further based on the estimated first reverberation gain.

5. The method according to claim 4, wherein The estimating the first reverberation gain includes prompting the user to clap.

6. The method according to claim 4, wherein: The estimating the first reverberation gain includes rendering an impulse sound via a speaker of the wearable head device.

7. The method according to claim 1, wherein Determining the envelope of the first audio signal comprises: applying a bandpass filter to the first audio signal; and Determining the envelope of the second audio signal includes applying the bandpass filter to the second audio signal.

8. A system for determining and processing audio information, comprising: a wearable head device configured to present a view of a virtual environment; a microphone of the wearable head device; One or more processors configured to perform a method comprising: receiving, via the microphone of the wearable head device, a first audio signal at a first time; determining an envelope of the first audio signal; estimating a first reverberation time based on the envelope of the first audio signal; Based on the estimated first reverberation time, determining that the position of the wearable head device at the first time corresponds to a first area of ​​the virtual environment; and receiving, via the microphone of the wearable head device, a second audio signal at a second time; determining an envelope of the second audio signal; estimating a second reverberation time based on the envelope of the second audio signal; and determining, based on the estimated second reverberation time, that the position of the wearable head device at the second time corresponds to a second area of ​​the virtual environment, the second area being different from the first area; and Based on a difference between the estimated first reverberation time and the estimated second reverberation, it is determined that the first area is in a first room and the second area is in a second room different from the first room.

9. The system according to claim 8, wherein: The estimating the first reverberation time includes determining whether the envelope of the first audio signal decays for a time greater than a threshold amount of time.

10. The system according to claim 8, wherein The estimating the first reverberation time comprises: determining a linear fit of a decay region of the envelope of the first audio signal; and A determination is made as to whether the correlation of the linear fit is greater than a threshold correlation.

11. The system of claim 8, wherein the method comprises: A first reverberation gain is estimated based on the envelope of the first audio signal, wherein determining that the position of the wearable head device at the first time corresponds to a first area of ​​the virtual environment is further based on the estimated first reverberation gain.

12. The system according to claim 11, wherein The estimating the first reverberation gain includes prompting the user to clap.

13. The system according to claim 11, wherein: The estimating the first reverberation gain includes rendering an impulse sound via a speaker of the wearable head device.

14. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for determining and processing audio information, the method comprising: receiving, at a first time, a first audio signal via a microphone of a wearable head device configured to present a view of a virtual environment; determining an envelope of the first audio signal; estimating a first reverberation time based on the envelope of the first audio signal; determining, based on the estimated first reverberation time, that a position of the wearable head device at the first time corresponds to a first area of ​​the virtual environment; receiving, via the microphone of the wearable head device, a second audio signal at a second time; determining an envelope of the second audio signal; estimating a second reverberation time based on the envelope of the second audio signal; as well as determining, based on the estimated second reverberation time, that the position of the wearable head device at the second time corresponds to a second area of ​​the virtual environment, the second area being different from the first area; as well as Based on a difference between the estimated first reverberation time and the estimated second reverberation, it is determined that the first area is in a first room and the second area is in a second room different from the first room.

15. The non-transitory computer-readable medium of claim 14, wherein: The estimating the first reverberation time includes determining whether the envelope of the first audio signal decays for a time greater than a threshold amount of time.

16. The non-transitory computer readable medium of claim 14, wherein: The estimating the first reverberation time comprises: determining a linear fit of a decay region of the envelope of the first audio signal; and A determination is made as to whether the correlation of the linear fit is greater than a threshold correlation.

17. The non-transitory computer-readable medium of claim 14, wherein: The method further includes estimating a first reverberation gain based on the envelope of the first audio signal, wherein determining that the position of the wearable head device at the first time corresponds to the first area of ​​the virtual environment is further based on the estimated first reverberation gain.

18. The non-transitory computer-readable medium of claim 17, wherein: The estimating the first reverberation gain includes prompting the user to clap.

19. The non-transitory computer-readable medium of claim 17, wherein: The estimating the first reverberation gain includes rendering an impulse sound via a speaker of the wearable head device.

20. The non-transitory computer readable medium of claim 14, wherein: Determining the envelope of the first audio signal comprises: applying a bandpass filter to the first audio signal; and Determining the envelope of the second audio signal includes applying the bandpass filter to the second audio signal.

Citation Information

Patent Citations

  • Mixed reality spatial audio

    US10616705B2

  • Method and device for measuring random incidence sound-absorbing coefficient / sound-absorbing quantity of material

    CN103675104A

  • Instrument for measuring sound field

    JP2005012784A

  • Estimating room acoustic properties using microphone arrays

    US10440498B1

  • Virtual, augmented, and mixed reality

    US20180020312A1