Spatial Audio for a Two-Way Audio Environment

The system addresses the challenge of simulating realistic acoustic properties in virtual environments by determining intermediate audio signals based on location and environment properties, enhancing the immersive experience through accurate sound reproduction.

JP7715771B2Active Publication Date: 2025-07-30MAGIC LEAP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023143435
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-18
Filing Date
2023-09-05
Publication Date
2025-07-30
Estimated Expiration
2039-06-18

AI Technical Summary

Technical Problem

Conventional audio systems in virtual environments fail to realistically simulate the acoustic properties of a room, limiting the immersive experience by not accurately reproducing virtual sounds according to the user's physical environment.

Method used

A system and method for presenting an output audio signal that determines individual intermediate audio signals based on the location and acoustic properties of sound sources within a virtual environment, using sensors and buses to simulate reflections and reverberations, and applying filters to enhance the audio experience.

Benefits of technology

Enhances the realism of virtual audio by accurately simulating the acoustic properties of the virtual environment, providing a more convincing and immersive experience by matching the user's expectations of sound quality based on their physical surroundings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715771000006
    Figure 0007715771000006
  • Figure 0007715771000007
    Figure 0007715771000007
  • Figure 0007715771000008
    Figure 0007715771000008
Patent Text Reader

Abstract

To provide a favorable spatial audio for interactive audio environments.SOLUTION: Systems and methods of presenting an output audio signal to a listener located at a first location in a virtual environment are disclosed. According to embodiments of a method, an input audio signal is received. For each sound source of multiple sound sources in the virtual environment, a respective first intermediate audio signal corresponding to the input audio signal is determined based on a location of the respective sound source in the virtual environment, and the respective first intermediate audio signal is associated with a first bus. For each sound sources of the multiple sound sources in the virtual environment, a respective second intermediate audio signal is determined. The respective second intermediate audio signal corresponds to a reflection of the input audio signal in a surface of the virtual environment. The respective second intermediate audio signal is determined based on a location of the respective sound source and further on an acoustic property of the virtual environment. The respective second intermediate audio signal is associated with a second bus. The output audio signal is presented to the listener via the first bus and the second bus.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority of U.S. Provisional Application No. 62 / 686,655, filed on Jun. 18, 2018, the content of which is incorporated herein by reference in its entirety. This application also claims the benefit of priority of U.S. Provisional Application No. 62 / 686,665, filed on Jun. 18, 2018, the content of which is incorporated herein by reference in its entirety.

[0002] The present disclosure generally relates to spatial audio rendering, and more specifically, to spatial audio rendering for virtual sound sources within a virtual acoustic environment.

Background Art

[0003] Virtual environments are prevalent in computing environments and find use in video games (where the virtual environment can represent a game world), maps (where the virtual environment can represent terrain to be navigated), simulations (where the virtual environment can simulate a real environment), digital storytelling (where virtual characters can interact with each other within the virtual environment), and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the user experience with virtual environments can be limited by the technology used to present the virtual environment. For example, conventional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may not be able to realize a virtual environment so as to attract people and create a realistic and immersive experience.

[0004] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present sensory information to a user of an XR system that corresponds to a virtual environment represented by data within a computer system. Such systems can provide a uniquely enhanced sense of immersion and presence by combining virtual visual and audio cues with real-world sights and sounds. Thus, it can be desirable to present digital sound to a user of an XR system such that the sound appears to occur naturally within the user's physical environment and consistently with what the user expects. Generally speaking, a user expects that virtual sounds will have the acoustic properties of the physical environment in which they are heard. For example, a user of an XR system within a large concert hall expects the virtual sounds of the XR system to have a sound quality similar to that of a large cave, and conversely, a user within a small apartment expects the sound to be more attenuated, proximate, and immediate.

[0005] Digital or artificial reverberators can be used in audio and music signal processing to simulate the perceived effect of room reverberation in an enclosed space. In an XR environment, it is desirable to use a digital reverberator to realistically simulate the acoustic properties of a room within the XR environment. Such a convincing simulation of acoustic properties can give credibility and immersion to the XR environment. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0006] A system and method for presenting an output audio signal to a listener located at a first location within a virtual environment are disclosed. According to an embodiment of a method, an input audio signal is received. For each sound source of a plurality of sound sources within the virtual environment, an individual first intermediate audio signal corresponding to the input audio signal is determined based on the location of the individual sound source within the virtual environment, and the individual first intermediate audio signal is associated with a first bus. For each sound source of the plurality of sound sources within the virtual environment, an individual second intermediate audio signal is determined. The individual second intermediate audio signal corresponds to a reflection of the input audio signal within the surface of the virtual environment. The individual second intermediate audio signal is determined based on the location of the individual sound source and further based on the acoustic properties of the virtual environment. The individual second intermediate audio signal is associated with a second bus. The output audio signal is presented to the listener via the first bus and the second bus. This specification also provides, for example, the following items. (Item 1) A method for presenting an output audio signal to a listener located at a first location within a virtual environment, the method comprising: Receiving an input audio signal; For each sound source among a plurality of sound sources within the virtual environment, Determining an individual first intermediate audio signal corresponding to the input audio signal based on the location of the individual sound source within the virtual environment; Associating the individual first intermediate audio signal with a first bus; Determining an individual second intermediate audio signal based on the location of the individual sound source and further based on the acoustic properties of the virtual environment, wherein the individual second intermediate audio signal corresponds to reflections of the input audio signal within the surface of the virtual environment; Associating the individual second intermediate audio signal with a second bus; Presenting the output audio signal to the listener via the first bus and the second bus A method comprising. (Item 2) The method according to item 1, wherein the acoustic properties of the virtual environment are determined via one or more sensors associated with the listener. (Item 3) The method according to item 2, wherein the one or more sensors comprise one or more microphones. (Item 4) The one or more sensors are associated with a wearable head unit configured to be worn by the listener, The output signal is presented to the listener via one or more speakers associated with the wearable head unit. The method according to item 2. (Item 5) The method according to item 4, wherein the wearable head unit comprises a display configured to display a view of the virtual environment to the listener in parallel with the presentation of the output signal. (Item 6) The method according to item 4, further comprising reading the acoustic properties from a database, the acoustic properties including the acoustic properties determined via one or more sensors of the wearable head unit. (Item 7) Reading the acoustic properties comprises: Determining the location of the listener based on the output of the one or more sensors; Identifying the acoustic properties based on the location of the listener The method according to item 6, comprising. (Item 8) A wearable device comprising: ​ ​ ​ ​ Receiving an input audio signal, for each sound source of a plurality of sound sources in the virtual environment, determining an individual first intermediate audio signal corresponding to the input audio signal based on the location of the individual sound source in the virtual environment; associating the individual first intermediate audio signal with a first bus; determining an individual second intermediate audio signal based on the location of the individual sound source and further based on the acoustic properties of the virtual environment, wherein the individual second intermediate audio signal corresponds to the reflection of the input audio signal within a surface in the virtual environment; associating the individual second intermediate audio signal with a second bus; presenting the output audio signal to a listener via the speaker and via the first bus and the second bus; one or more processors configured to implement a method comprising: A wearable device comprising: (Item 9) The wearable device according to item 8, wherein the acoustic properties of the virtual environment are determined via the one or more sensors. (Item 10) The wearable device according to item 8, wherein the one or more sensors comprise one or more microphones. (Item 11) The method according to item 8, further comprising displaying a view of the virtual environment via the display in parallel with the presentation of the output signal. (Item 12) The method according to item 8, further comprising reading the acoustic properties from a database, the acoustic properties including acoustic properties determined via one or more sensors. (Item 13) Reading the acoustic properties comprises: determining the location of the listener based on the output of the one or more sensors; identifying the acoustic properties based on the location of the listener; The method according to item 12, comprising: (Item 14) for each sound source of the plurality of sound sources in the virtual environment, determining an individual third intermediate audio signal based on the location of the individual sound source and further based on a second acoustic property of the virtual environment, wherein the individual third intermediate audio signal corresponds to the reverberation of the input audio signal in the virtual environment; associating the individual third intermediate audio signal with the second bus; further comprising: The second bus comprises a reflection bus and a reverberation bus. Associating the individual second intermediate audio signal with the second bus includes associating the individual second intermediate audio signal with the reflection bus, Associating the individual third intermediate audio signal with the second bus includes associating the individual third intermediate audio signal with the reverberation bus, The method according to item 1. (Item 15) The method further includes, For each sound source of the plurality of sound sources in the virtual environment, Determining an individual third intermediate audio signal based on the location of the individual sound source and further based on a second acoustic property of the virtual environment, wherein the individual third intermediate audio signal corresponds to the reverberation of the input audio signal in the virtual environment, Associating the individual third intermediate audio signal with the second bus Including, The second bus includes a reflection bus and a reverberation bus, Associating the individual second intermediate audio signal with the second bus includes associating the individual second intermediate audio signal with the reflection bus, Associating the individual third intermediate audio signal with the second bus includes associating the individual third intermediate audio signal with the reverberation bus, The wearable device according to item 8. (Item 16) Determining the individual first intermediate audio signal includes applying a first individual filter to the input audio signal, the first individual filter comprising one or more of a sound source directivity model, a distance model, and an orientation model, the method according to item 1. (Item 17) Determining the individual first intermediate audio signal further includes applying one or more of an individual gain and an individual panning process to the input audio signal, the method according to item 16. (Item 18) The individual panning process includes panning the input audio signal based on the geometry of the loudspeaker array, the method according to item 17. (Item 19) Determining the individual second intermediate audio signal includes applying a second individual filter to the input audio signal, the second individual filter comprising a sound source directivity model, the method according to item 1. (Item 20) Determining the individual second intermediate audio signal further includes applying one or more of an individual delay, an individual gain, and an individual panning process to the input audio signal, the method according to item 19. (Item 21) The individual panning process includes encoding the input audio signal into an ambisonic signal including three channels, the method according to item 20. (Item 22) The individual panning process includes panning the reflection of the input audio signal based on one or more of an azimuth angle and a spatial focus parameter, the method according to item 20. (Item 23) Determining the individual first intermediate audio signal includes applying a first individual filter to the input audio signal, the first individual filter comprising one or more of a sound source directivity model, a distance model, and an orientation model, the wearable device according to item 8. (Item 24) Determining the individual first intermediate audio signal further includes applying one or more of an individual gain and an individual panning process to the input audio signal, the wearable device according to item 23. (Item 25) The individual panning process includes panning the input audio signal based on the geometry of an array of loudspeakers, the wearable device according to item 24. (Item 26) Determining the individual second intermediate audio signal includes applying a second individual filter to the input audio signal, the second individual filter comprising a sound source directivity model, the wearable device according to item 8. (Item 27) Determining the individual second intermediate audio signal further includes applying one or more of an individual delay, an individual gain, and an individual panning process to the input audio signal, the wearable device according to item 26. (Item 28) The individual panning process includes encoding the input audio signal into an ambisonic signal including three channels, the wearable device according to item 27. (Item 29) The wearable device of item 27, wherein the individual panning process includes panning reflections of the input audio signal based on one or more of an azimuth angle and a spatial focus parameter.

Brief Description of the Drawings

[0007]

Figure 1

[0008]

Figure 2

[0009]

Figure 3

[0010]

Figure 4

[0011]

Figure 5

[0012]

Figure 6

[0013]

Figure 7

[0014]

Figure 8

[0015]

Figure 9

[0016]

Figure 10

[0017]

Figure 11

[0018]

Figure 12

[0019]

Figure 13

[0020]

Figure 14

[0021]

Figure 15

[0022]

Figure 16

[0023]

Figure 17

[0024]

Figure 18

[0025] In the following description of the embodiments, reference is made to the accompanying drawings which form a part hereof and in which are shown by way of illustration specific embodiments in which the invention may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the disclosed embodiments.

[0026] Exemplary Wearable System

[0027] FIG. 1 illustrates an exemplary wearable head device 100 configured to be worn on a user's head. The wearable head device 100 may be part of a broader wearable system that includes one or more components such as a head device (e.g., the wearable head device 100), a handheld controller (e.g., the handheld controller 200 described below), and / or an auxiliary unit (e.g., the auxiliary unit 300 described below). In some embodiments, the wearable head device 100 can be used for virtual reality, augmented reality, or mixed reality systems or applications. The wearable head device 100 may include one or more displays such as displays 110A and 110B (left and right transmissive displays, and associated components for coupling light from the displays to the user's eyes such as orthogonal pupil expansion (OPE) grating sets 112A / 112B and exit pupil expansion (EPE) grating sets 114A / 114B), left and right acoustic structures such as speakers 120A and 120B (mounted on respective arm stems 122A and 122B and positioned adjacent to the user's left and right ears), one or more sensors such as an infrared sensor, an accelerometer, a GPS unit, an inertial measurement unit (IMU) (e.g., IMU 126), an acoustic sensor (e.g., microphone 150), an orthogonal coil electromagnetic receiver (e.g., receiver 127 shown mounted on the left arm stem 122A), left and right cameras oriented away from the user (e.g., depth (time-of-flight) cameras 130A and 130B), and left and right eye cameras oriented towards the user (e.g., for detecting the user's eye movements) (e.g., eye cameras 128 and 128B). However, the wearable head device 100 can incorporate any suitable display technology and any suitable number, type, or combination of sensors or other components without departing from the scope of the present invention.In some embodiments, the wearable head device 100 may incorporate one or more microphones 150 configured to detect an audio signal generated by the user's voice, such microphones may be positioned within the wearable head device adjacent to the user's mouth. In some embodiments, the wearable head device 100 may incorporate networking features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other wearable systems. The wearable head device 100 may further include components such as a battery, a processor, a memory, a storage unit, or various input devices (e.g., buttons, touch pads), or may be coupled to a handheld controller (e.g., handheld controller 200) or an auxiliary unit (e.g., auxiliary unit 300) comprising one or more such components. In some embodiments, the sensor may be configured to output a set of coordinates of the head-mounted unit relative to the user's environment, provide the input to a processor, and perform a simultaneous localization and mapping (SLAM) procedure and / or a visual odometry algorithm. In some embodiments, the wearable head device 100 may be further coupled to the handheld controller 200 and / or the auxiliary unit 300, as further described below.

[0028] Figure 2 illustrates an exemplary mobile handheld controller component 200 of an exemplary wearable system. In some embodiments, the handheld controller 200 may communicate, either wired or wirelessly, with the wearable head device 100 and / or an auxiliary unit 300 described below. In some embodiments, the handheld controller 200 includes a handle portion 220 to be held by a user and one or more buttons 240 disposed along an upper surface 210. In some embodiments, the handheld controller 200 may be configured for use as an optical tracking target; for example, sensors (e.g., cameras or other optical sensors) of the wearable head device 100 may be configured to detect the position and / or orientation of the handheld controller 200, which in turn may indicate the position and / or orientation of the hand of the user holding the handheld controller 200. In some embodiments, the handheld controller 200 may include one or more input devices such as a processor, memory, storage unit, display, or those described above. In some embodiments, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 100). In some embodiments, the sensors may be capable of detecting the position or orientation of the handheld controller 200 relative to the wearable head device 100 or another component of the wearable system. In some embodiments, the sensors may be positioned within the handle portion 220 of the handheld controller 200 and / or may be mechanically coupled to the handheld controller. The handheld controller 200 may be configured to provide one or more output signals corresponding to, for example, the depressed state of the button 240 or the position, orientation, and / or movement (e.g., via an IMU) of the handheld controller 200. Such output signals may be used as input to a processor of the wearable head device 100, to the auxiliary unit 300, or to another component of the wearable system.In some embodiments, the handheld controller 200 can include one or more microphones to detect sound (e.g., the user's speech, ambient sound) and, in some cases, provide a signal corresponding to the detected sound to a processor (e.g., the processor of the wearable head device 100).

[0029] FIG. 3 illustrates an exemplary auxiliary unit 300 of an exemplary wearable system. In some embodiments, the auxiliary unit 300 may communicate with the wearable head device 100 and / or the handheld controller 200, either wired or wirelessly. The auxiliary unit 300 can include a battery to provide energy for operating one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including a display, sensors, acoustic structures, processors, microphones, and / or other components of the wearable head device 100 or the handheld controller 200). In some embodiments, the auxiliary unit 300 may include a processor, a memory, a storage unit, a display, one or more input devices, and / or one or more sensors such as those described above. In some embodiments, the auxiliary unit 300 includes a clip 310 (e.g., a belt worn by the user) for attaching the auxiliary unit to the user. The advantage of using the auxiliary unit 300 to store one or more components of the wearable system is that doing so can allow large or heavy components to be carried on the user's waist, chest, or back, which is relatively well-suited for supporting large and heavy objects, rather than being mounted on the user's head (e.g., if stored within the wearable head device 100) or carried by the user's hand (e.g., if stored within the handheld controller 200). This can be particularly advantageous for relatively heavy or bulky components such as batteries.

[0030] Figure 4 shows an exemplary functional block diagram corresponding to an exemplary wearable system 400 that may include, among other things, the exemplary wearable head device 100, the handheld controller 200, and the auxiliary unit 300 described above. In some embodiments, the wearable system 400 may be used for virtual reality, augmented reality, or mixed reality applications. As shown in Figure 4, the wearable system 400 can include an exemplary handheld controller 400B, herein referred to as a "totem" (and corresponding to the handheld controller 200 described above), which can include a totem / headgear 6 degrees of freedom (6DOF) totem subsystem 404A. The wearable system 400 can also include an exemplary wearable head device 400A (corresponding to the wearable headgear device 100 described above), which includes a totem / headgear 6DOF headgear subsystem 404B. In an embodiment, the 6DOF totem subsystem 404A and the 6DOF headgear subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom may be represented relative to the coordinate system of the wearable head device 400A. The three translational offsets may be represented as X, Y, and Z offsets, a translation matrix, or some other representation within such a coordinate system. The rotational degrees of freedom may be represented as a sequence of yaw, pitch, and roll rotations, a vector, a rotation matrix, a quaternion, or some other representation. In some embodiments, one or more depth cameras 444 (and / or one or more non-depth cameras) included within the wearable head device 400A and / or one or more optical targets (e.g., buttons 240 of the handheld controller 200 as described above or dedicated optical targets included within the handheld controller) can be used for 6DOF tracking.In some embodiments, the handheld controller 400B can include a camera as described above, and the headgear 400A can include an optical target for optical tracking in conjunction with the camera. In some embodiments, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids, which are used to wirelessly transmit and receive three distinguishable signals. The six degrees of freedom (6DOF) of the handheld controller 400B relative to the wearable head device 400A may be determined by measuring the relative magnitudes of the three distinguishable signals received in each of the coils used for reception. In some embodiments, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that is useful for providing improved accuracy and / or more timely information regarding fast movement of the handheld controller 400B.

[0031] In some embodiments involving augmented or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space fixed with respect to the wearable head device 400A) to an inertial coordinate space or to an environmental coordinate space. For example, such a transformation may be necessary so that the display of the wearable head device 400A presents virtual objects at expected positions and orientations with respect to the real environment rather than at fixed positions and orientations on the display (e.g., at the same position on the display of the wearable head device 400A). For example, a virtual person sitting on a real chair facing forward (regardless of the position and orientation of the wearable head device 400A). This can maintain the illusion that the virtual object exists within the real environment (and, for example, does not appear unnaturally positioned within the real environment as the wearable head device 400A shifts and rotates). In some embodiments, a compensatory transformation between coordinate spaces can be determined by processing images from the depth camera 444 (e.g., using simultaneous localization and mapping (SLAM) and / or visual odometry procedures) to determine the transformation of the wearable head device 400A with respect to the inertial or environmental coordinate system. In the embodiment shown in FIG. 4, the depth camera 444 can be coupled to the SLAM / visual odometry block 406 and provide the image to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the present image and then determine the position and orientation of the user's head, which can be used to identify the transformation between the head coordinate space and the real coordinate space. Similarly, in some embodiments, an additional source of information regarding the user's head pose and location is obtained from the IMU 409 of the wearable head device 400A. The information from the IMU 409 can be integrated with the information from the SLAM / visual odometry block 406 to provide more timely information regarding improved accuracy and / or faster adjustment of the user's head pose and position.

[0032] In some embodiments, the depth camera 444 can supply 3D images to a hand gesture tracker 411, which can be implemented within the processor of the wearable head device 400A. The hand gesture tracker 411 can identify a user's hand gesture, for example, by matching the 3D images received from the depth camera 444 to stored patterns representing hand gestures. Other suitable techniques for identifying a user's hand gesture will be apparent.

[0033] In some embodiments, one or more processors 416 may be configured to receive data from the headgear subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, microphone (not shown), and / or hand gesture tracker 411. The processor 416 may also be able to send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A in embodiments where the handheld controller 400B is not tethered. The processor 416 may further communicate with additional components such as the audiovisual content memory 418, graphics processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to the head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to the left source 424 of light modulated per image and a right channel output coupled to the right source 426 of light modulated per image. The GPU 420 may output stereoscopic image data to the sources 424, 426 of light modulated per image. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or right speaker 414. The DSP audio spatializer 422 may receive from the processor 416 an input indicating a direction vector from the user to a virtual sound source (e.g., movable by the user via the handheld controller 400B). Based on the direction vector, the DSP audio spatializer 422 may be able to determine the corresponding HRTF (e.g., by accessing the HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal such as an audio signal corresponding to the virtual sound generated by the virtual object.This can improve the credibility and realism of virtual sounds by incorporating the user's relative position and orientation with respect to the virtual sounds within the mixed reality environment, i.e., by presenting virtual sounds that match the user's expectations of what the virtual sounds would sound like if they were real sounds within the real environment.

[0034] In some embodiments, such as those shown in FIG. 4, one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 may be included within an auxiliary unit 400C (which may correspond to the auxiliary unit 300 described above). The auxiliary unit 400C includes a battery 427, powers its components, and / or supplies power to the wearable head device 400A and / or the handheld controller 400B. Including such components within an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0035] FIG. 4 presents elements corresponding to various components of the exemplary wearable system 400, but various other suitable arrangements of these components will be apparent to those skilled in the art. For example, the elements presented in FIG. 4 as being associated with the auxiliary unit 400C could instead be associated with the wearable head device 400A or the handheld controller 400B. Additionally, some wearable systems may eliminate the handheld controller 400B or the auxiliary unit 400C entirely. Such changes and modifications are understood to be within the scope of the disclosed embodiments.

[0036] Mixed reality environment

[0037] Like all people, users of a mixed reality system exist within the physical environment, i.e., the three-dimensional portion of the "real world" and all of its contents that are perceivable by the user. For example, the user perceives the physical environment using their normal human senses, i.e., vision, hearing, touch, taste, and smell, and interacts with the physical environment by moving their own body within the physical environment. Locations within the physical environment can be described as coordinates within a coordinate space, e.g., the coordinates can include latitude, longitude, and altitude relative to sea level, distances in three orthogonal dimensions from a reference point, or other suitable values. Similarly, a vector can describe a quantity having a direction and magnitude within a coordinate space.

[0038] A computing device can maintain a representation of a virtual environment, for example, within a memory associated with the device. As used herein, a virtual environment is a computer representation of a three-dimensional space. The virtual environment can include representations of any object, action, signal, parameter, coordinate, vector, or other characteristic associated with that space. In some embodiments, the circuitry (e.g., a processor) of the computing device can maintain and update the state of the virtual environment, i.e., the processor can determine the state of the virtual environment at a second time based on data associated with the virtual environment and / or input provided by a user at a first time. For example, if an object within the virtual environment is located at a first coordinate at a certain time, has certain programmed physical parameters (e.g., mass, coefficient of friction), and the input received from the user indicates that a force should be applied to the object in a certain direction vector, the processor can apply the laws of kinematics and use basic mechanics to determine the location of the object at that time. The processor can use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at a given time. When maintaining and updating the state of the virtual environment, the processor can execute any suitable software, including software related to the creation and deletion of virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling inputs and outputs, software for implementing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.

[0039] Output devices such as displays or speakers can present any or all aspects of the virtual environment to the user. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, light, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, viewing axis, and frustum), and render a visible scene of the virtual environment corresponding to that view on the display. Any suitable rendering technique may be used for this purpose. In some embodiments, the visible scene may include only some of the virtual objects within the virtual environment and may exclude other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects within the virtual environment may generate sounds arising from the location coordinates of the objects (e.g., a virtual character may speak or cause sound effects), or the virtual environment may be associated with musical cues or ambient sounds that may or may not be associated with a particular location. The processor can determine an audio signal corresponding to "listener" coordinates, e.g., an audio signal that is mixed and processed to simulate an audio signal that would be heard by a listener at the listener coordinates and that corresponds to a composite of the sounds within the virtual environment, and present the audio signal to the user via one or more speakers.

[0040] Since a virtual environment exists only as a computer construct, a user cannot directly perceive the virtual environment using their normal senses. Instead, a user can only indirectly perceive the virtual environment as presented to the user, for example, by way of a display, speakers, a haptic output device, and the like. Similarly, a user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data to a processor that can use the device or sensor data to update the virtual environment via an input device or sensor. For example, a camera sensor can provide optical data indicating that a user is attempting to move an object within the virtual environment, and the processor can use that data to cause the object to respond accordingly within the virtual environment.

[0041] Reflection and reverberation

[0042] Aspects of a listener's audio experience within the space (e.g., a room) of a virtual environment include the listener's perception of the direct sound, the listener's perception of the reflection of that direct sound off the surfaces of the room, and the listener's perception of the reverberation ( "reverb") of the direct sound within the room. FIG. 5 illustrates a geometric room representation 500 according to some embodiments. The geometric room representation 500 shows exemplary propagation paths for direct sound (502), reflection (504), and reverberation (506). These paths represent the paths that an audio signal can take from a source to a listener within the room. The room shown in FIG. 5 can be any suitable type of environment associated with one or more acoustic properties. For example, room 500 can be a concert hall and can include a stage with a piano player and an audience section with an audience. As shown, direct sound is the sound that originates at the source (e.g., the piano player) and travels directly towards the listener (e.g., the audience). Reflection is the sound that originates at the source, reflects off a surface (e.g., the wall of the room), and travels towards the listener. Reverberation is the sound that includes decaying signals that include many reflections that arrive in close proximity to each other at a certain time.

[0043] FIG. 6 illustrates an exemplary model 600 of a room response measured from an indoor source to a listener, according to some embodiments. The model of the room response shows the amplitudes of the direct sound (610), the reflections of the direct sound (620), and the reverberation of the direct sound (630) from the perspective of a listener at a distance from the direct sound source. As illustrated in FIG. 6, the direct sound generally arrives at the listener before the reflections (with a reflection delay (622) in the figure indicating the difference in time between the direct sound and the reflection), which in turn arrives before the reverberation (with a reverberation delay (632) in the figure indicating the difference in time between the direct sound and the reverberation). The reflections and the reverberation can be perceptually different to the listener. The reflections can be modeled separately from the reverberation, for example, to better control the time, attenuation, spectral shape, and arrival direction of the individual reflections. The reflections may be modeled using a reflection model, and the reverberation may be modeled using a reverberation model, which may be different from the reflection model.

[0044] The reverberation properties (e.g., reverberation decay) for the same sound source can vary between two different acoustic environments (e.g., rooms) for the same sound source, and it is desirable to realistically reproduce the sound source according to the properties of the current room within the listener's virtual environment. That is, when a virtual sound source is presented in a composite reality system, the reflection and reverberation properties of the listener's real environment should be accurately reproduced. L. Savioja, J. Huopaniemi, T. Lokki, and R. Vaananen, "Creating Interactive Virtual Acoustic Environments," J. Audio Eng. Soc. 47(9): 675 - 705 (1999) describes methods for reproducing the direct path, individual reflections, and acoustic reverberation in a real - time virtual 3D audio reproduction system for video games, simulations, or AR / VR. In the methods disclosed by Savioja et al., the arrival direction, delay, amplitude, and spectral equalization of each individual reflection are derived from a geometric and physical model of the room (e.g., a real room, a virtual room, or some combination thereof), which can require a complex rendering system. These methods are computationally complex and may be prohibitively complex for mobile applications where computing resources may be limited.

[0045] In some room - acoustic simulation algorithms, reverberation can be implemented by down - mixing all sound sources to a mono signal and sending the mono signal to a reverberation simulation module. The gains used for down - mixing and sending can depend on dynamic parameters such as source distance and manual parameters such as reverberation gain.

[0046] The sound source directivity or radiation pattern can refer to a measure of the amount of energy that the sound source emits in different directions. The sound source directivity affects all parts of the room impulse response (e.g., direct, reflected, and reverberant). Different sound sources can exhibit different directivities. For example, human speech can have a different directivity pattern than a trumpet performance. A room simulation model can take into account the sound source directivity when generating an accurate simulation of an acoustic signal. For example, a model that incorporates the sound source directivity can include a function of the direction of the line from the sound source to the listener relative to the front direction (or the main acoustic axis) of the sound source. The directivity pattern is axisymmetric about the main acoustic axis of the sound source. In some embodiments, a parametric gain model may be defined using a frequency-dependent filter. In some embodiments, the average of the diffusion power of the sound source may be calculated (e.g., by integrating over a sphere centered on the acoustic center of the sound source) to determine the amount of audio from a given sound source that should be sent into the reverberation bus.

[0047] Bidirectional audio engines and sound design tools can make assumptions about the acoustic system to be modeled. For example, some bidirectional audio engines can model the sound source directivity as a function independent of frequency, which can have two potential drawbacks. First, this can ignore the frequency-dependent attenuation for the direct sound propagation from the sound source to the listener. Second, this can ignore the frequency-dependent attenuation for the reflected and reverberant transmissions. These effects can be important from a psychoacoustic perspective, and failure to reproduce them can lead to a room simulation that is perceived as unnatural and different from what the listener is accustomed to experiencing in a real acoustic environment.

[0048] In some cases, a room simulation system or a two-way audio engine may not fully separate the sound source, the listener, and acoustic environment parameters such as reflections and reverberations. Instead, the room simulation system may be adjusted as a whole for a specific virtual environment and may not be suitable for different playback scenarios. For example, the reverberation within the simulated environment may not match the environment physically present when the user / listener is listening to the rendered content.

[0049] In augmented or mixed reality applications, computer-generated audio objects can be rendered via an acoustically transparent playback system so as to blend with the physical environment naturally heard by the user / listener. This may require binaural artificial reverberation processing to match the local environmental acoustics, and thus the synthetic audio objects may not be distinguishable from naturally occurring or reproduced sounds over loudspeakers. For example, approaches involving measurement or calculation of the room impulse response based on estimating the geometry of the environment may be limited in the consumer environment due to practical obstacles and complexities. Additionally, physical models may not necessarily provide the most engaging listening experience as they do not consider the psychoacoustic acoustic principles or may not provide suitable audio scene parameterization for the sound designer to fine-tune the listening experience.

[0050] Matching some specific physical properties of the target acoustic environment may not provide a simulation that perceptually closely matches the listener's environment or the intentions of the application designer. A perceptually relevant model of the target acoustic environment that can be characterized using a practical audio environment description interface may be desired.

[0051] For example, a rendering model that separates the contributions of the source, listener, and room properties may be desired. The rendering model that separates the contributions may enable components to be adapted or swapped at runtime according to the local environment and the nature of the end user. For example, the listener may be present in a physical room that has acoustic characteristics different from the virtual environment in which the content was originally created. Modifying the early reflections and / or reverberation of the simulation to match the listening environment can lead to a more convincing listening experience. Matching the listening environment can be particularly important in mixed reality applications where the desired effect may be that the listener cannot distinguish between the simulated sounds around them and the sounds present in the actual surrounding environment.

[0052] It may be desired to create a convincing effect without requiring detailed knowledge of the geometry of the actual surrounding environment and / or the acoustic properties of the surrounding surfaces. Detailed knowledge of the actual surrounding environment properties may not be available, or they may be particularly complex to estimate, especially on portable devices. Instead, models based on perception and psychoacoustic principles can be much more practical tools for characterizing the acoustic environment.

[0053] FIG. 7 illustrates Table 700, which includes several objective acoustic and geometric parameters characterizing each segment in a binaural room impulse model that differentiates the nature of the source, listener, and room according to several embodiments. Several source properties, including free-field and diffuse-field transfer functions, may be independent of the way and location in which the content will be rendered, while other properties, including position and orientation, may need to be updated dynamically during playback. Similarly, several listener properties, including free-field and diffuse-field head-related transfer functions or diffuse-field interaural coherence (IACC), may be independent of the location in which the content will be rendered, while other properties, including position and orientation, may be updated dynamically during playback. Several room properties, particularly those contributing to late reverberation, may be completely environment-dependent. Representations of the reverberation decay rate and room cubical volume may be for adapting a spatial audio rendering system to the listener's playback environment.

[0054] The ears of the source and listener may be modeled as emission and reception transducers, respectively, characterized by a set of direction-dependent free-field transfer functions that include the listener's head-related transfer function (HRTF).

[0055] FIG. 8 illustrates an exemplary audio mixing system 800 for rendering a plurality of virtual sound sources within a virtual room such as within an XR environment, according to some embodiments. For example, the audio mixing architecture may include a rendering engine for room acoustic simulation of a plurality of virtual sound sources 810 (i.e., objects 1-N). System 800 includes a room transmission bus 830 that feeds a module 850 (e.g., a shared reverberation and reflection module) that renders reflections and reverberations. Aspects of this general process are described, for example, in IA-SIG 3D Audio Rendering Guidelines (Level 2), www.iasig.net (1999). The room transmission bus combines contributions from all sources, e.g., sound sources 810, each processed by a corresponding module 820, to derive an input signal for the room module. The room transmission bus may comprise a mono room transmission bus. The format of the main mixing bus 840 may be a two-channel or multi-channel format that matches the final output rendering method, which may include, for example, a binaural renderer for headphone playback, an ambisonic decoder, and / or a multi-channel loudspeaker system. The main mixing bus combines contributions from all sources with the room module output to derive an output rendering signal 860.

[0056] Referring to the exemplary system 800, each of the N objects may represent a virtual sound source signal and may be assigned an apparent location in the environment, such as by a panning algorithm. For example, each object can be assigned an angular position on a sphere centered on the position of the virtual listener. The panning algorithm may calculate the contribution of each object to each channel of the main mix. This general process is described, for example, in J.-M. Jot, V. Larcher, and J.-M. Pernaux, "A comparative study of 3-D audio encoding and rendering techniques", Proc. AES 16th International Conference on Spatial Sound Reproduction (1999). Each object may be input to a pan, gain module 820, which can implement a panning algorithm and perform additional signal processing such as adjusting the gain level for each object.

[0057] In some embodiments, the system 800 may assign (e.g., via module 820) to each virtual sound source an apparent distance from the position of the virtual listener from which the rendering engine can derive a source-specific direct gain and a source-specific room gain for each object. The direct and room gains can affect the audio signal power contributed to the main mix bus 840 and the room send bus 830, respectively, by the virtual sound source. A minimum distance parameter may be assigned to each virtual sound source, and the direct and room gains may roll off at different rates as the distance increases beyond this minimum distance.

[0058] In some embodiments, system 800 of FIG. 8 may be used for audio recording and generation of two-way audio applications targeting a conventional two-channel front stereo loudspeaker playback system. However, when applied in a binaural or immersive 3D audio system that allows for a spatially diffuse distribution of simulated reverberations and reflections, system 800 may not provide sufficiently convincing auditory image localization cues when rendering virtual sound sources, particularly those far from the listener. This can be addressed by the inclusion of a clustered reflection rendering module shared among the virtual sound sources 810 while supporting source-by-source control of the spatial distribution of reflections. It is desirable for such a module to incorporate an early reflection processing algorithm for each source and dynamic control of the early reflection parameters based on the virtual sound source and listener positions.

[0059] In some embodiments, it may be desirable to have a spatial audio processing model / system and method that can accurately reproduce position-dependent room acoustic cues without the computationally complex rendering of individual early reflections for each virtual sound source or detailed descriptions of the acoustic reflector geometries and physical properties.

[0060] The reflection processing model may dynamically consider the positions of the listener and virtual sound sources in an actual or virtual room / environment without the associated physical and geometric descriptions. A perceptual model for source-by-source clustered reflection panning and control of the early reflection processing parameters may be efficiently implemented.

[0061] FIG. 9 illustrates an audio mixing system 900 for rendering a plurality of virtual sound sources in a virtual room according to some embodiments. For example, system 900 may include a rendering engine for room acoustic simulation of a plurality of virtual sound sources 910 (e.g., objects 1-N). Compared to system 800 described above, system 900 can include separate control of reverberation and reflection transmission channels for each virtual sound source. Each object may be input to an individual source processing module 920, and the room transmission bus 930 may feed the room processing module 950.

[0062] FIG. 10 illustrates a source processing module 1020 according to some embodiments. Module 1020 can correspond to one or more of modules 920 shown in FIG. 9 and exemplary system 900. The source processing module 1020 can perform processing specific to an individual source (e.g., 1010 which can correspond to one of sources 910) of the overall system (e.g., system 900). The source processing module may include a direct processing path (e.g., 1030A) and / or a room processing path (e.g., 1030B).

[0063] In some embodiments, individual direct and room filters may be applied separately for each sound source. Applying filters separately can allow for more refined and accurate control over how each source radiates sound towards the listener and into the surrounding environment. In contrast to broadband gain, the use of filters can allow for matching a desired sound radiation pattern as a function of frequency. This is beneficial because the radiation properties can vary across sound source types and can be frequency-dependent. The angle between the main acoustic axis of the sound source and the listener's position can affect the sound pressure level perceived by the listener. Additionally, the source radiation characteristics can affect the average of the source's diffusion power.

[0064] In some embodiments, the frequency-dependent filter may be implemented using the double-shelving approach disclosed in U.S. Patent Application No. 62 / 678259, entitled "INDEX SCHEMING FOR FILTER PARAMETERS", the content of which is incorporated by reference in its entirety. In some embodiments, the frequency-dependent filter may be applied in the frequency domain and / or using a finite impulse response filter.

[0065] As shown in the example, the direct processing path may include a direct transmission filter 1040, followed by a direct panning module 1044. The direct transmission filter 1040 may model one or more acoustic effects such as one or more of sound source directivity, distance, and / or orientation. The direct panning module 1044 can spatialize the audio signal to correspond to an apparent position in the environment (e.g., a 3D location in a virtual environment such as an XR environment). The direct panning module 1044 may be amplitude and / or intensity based and may depend on the geometry of the loudspeaker array. In some embodiments, the direct processing path may include a direct transmission gain 1042, along with the direct transmission filter and the direct panning module. The direct panning module 1044 can output to a main mixing bus 1090, which may correspond to the main mixing bus 940 described above with respect to the exemplary system 900.

[0066] In some embodiments, the room processing path includes a room delay 1050 and a room transmit filter 1052, followed by a reflection path (e.g., 1060A) and a reverberation path (e.g., 1060B). The room transmit filter may be used to model the effect of source directivity on signals traveling through the reflection and reverberation paths. The reflection path may include a reflection transmit gain 1070 and may transmit signals to a reflection transmit bus 1074 via a reflection pan module 1072. The reflection pan module 1072 may be similar to the direct pan module 1044 in that it can spatialize the audio signal and operate on the reflections instead of the direct signal. The reverberation path 1060B may include a reverberation gain 1080 and may transmit signals to a reverberation transmit bus 1084. The reflection transmit bus 1074 and the reverberation transmit bus 1084 may be grouped into a room transmit bus 1092, which may correspond to the room transmit bus 930 described above with respect to the exemplary system 900.

[0067] FIG. 11 illustrates an example of a per-source reflection pan module 1100 that may correspond to the reflection pan module 1072 described above, according to some embodiments. As shown in the figure, the input signal may be encoded into a 3-channel ambisonic B-format signal, as described, for example, in J.-M. Jot, V. Larcher, and J.-M. Pernaux, “A comparative study of 3-D audio encoding and rendering techniques,” Proc. AES 16th International Conference on Spatial Sound Reproduction (1999). The encoding coefficients 1110 can be calculated according to Equations 1-3.

Chemical formula

[0068] In Equation 1-3, k is

Chemical formula

[0069] Az can be an azimuth angle defined by the projection of the main arrival direction of reflection onto the head-relative horizontal plane (e.g., perpendicular to the "up" vector of the listener's head and the plane containing the listener's ears). The spatial focus parameter F can indicate the spatial concentration of the reflected signal energy arriving at the listener. When F is zero, the spatial distribution of the reflected energy arrival can be uniform around the listener. As F increases, the spatial distribution can become increasingly concentrated around the main direction determined by the azimuth angle Az. The maximum theoretical value of F is 1.0, which can indicate that all the energy is arriving from the main direction determined by the azimuth angle Az.

[0070] In an embodiment of the present invention, the spatial focus parameter F may be defined, for example, as the magnitude of the Garson energy vector as described in J.-M. Jot, V. Larcher, and J.-M. Pernaux, "A comparative study of 3-D audio encoding and rendering techniques", Proc. AES 16th International Conference on Spatial Sound Reproduction (1999).

[0071] The output of the reflection panning module 1100 can be provided to a reflection transmission bus 1174, which may correspond to the reflection transmission bus 1074 described above with respect to FIG. 10 and the exemplary processing module 1020.

[0072] FIG. 12 illustrates an exemplary room processing module 1200 according to some embodiments. The room processing module 1200 can correspond to the room processing module 950 described above with respect to FIG. 9 and the exemplary system 900. As shown in FIG. 9, the room processing module 1200 may include a reflection processing path 1210A and / or a reverberation processing path 1210B.

[0073] The reflection processing path 1210A may receive a signal from a reflection transmission bus 1202 (which may correspond to the reflection transmission bus 1074 described above) and output the signal into a main mixing bus 1290 (which may correspond to the main mixing bus 940 described above). The reflection processing path 1210A may include a reflection global gain 1220, a reflection global delay 1222, and / or a reflection module 1224 that can simulate / render reflections.

[0074] The reverberation processing path 1210B may receive a signal from a reverberation transmission bus 1204 (which may correspond to the reverberation transmission bus 1084 described above) and output the signal into the main mixing bus 1290. The reverberation processing path 1210B may include a reverberation global gain 1230, a reverberation global delay 1232, and / or a reverberation module 1234.

[0075] FIG. 13 illustrates an exemplary reflection module 1300 according to some embodiments. The input 1310 of the reflection module can be output by a reflection pan module 1100 such as those described above and presented to the reflection module 1300 via the reflection transmission bus 1174. The reflection transmission bus may carry a three-channel ambisonic B-format signal that combines contributions from all virtual sound sources (e.g., the sound sources 910 (objects 1-N) described above with respect to FIG. 9). In the illustrated embodiment, the three channels represented as (W, X, Y) are fed to an ambisonic decoder 1320. According to an embodiment, the ambisonic decoder generates six output signals, which are each fed to six mono input / output basic reflection modules 1330 (R1-R6) to generate a set of six reflected output signals 1340 (s1-s6). (The embodiment shows six signals and reflection modules, but any suitable number may be used.) The reflected output signals 1340 are presented to a main mixing bus 1350, which may correspond to the main mixing bus 940 described above.

[0076] FIG. 14 illustrates a spatial distribution 1400 of the apparent arrival directions of reflections as detected by a listener 1402 according to some embodiments. For example, the illustrated reflections can be generated by the reflection module 1300 described above with respect to sound sources assigned specific values of the reflection pan parameters Az and F, such as those described above with respect to FIG. 11.

[0077] As shown in FIG. 14, the effect of the reflection module 1300 combined with the reflection pan module 1100 is to generate a series of reflections, each of which can arrive from each of the virtual loudspeaker directions 1410 (e.g., 1411 - 1416 corresponding to the reflection output signals s1 - s6 described above) at different times (e.g., as shown in the model 600). The effect of the reflection pan module 1100 combined with the ambisonic decoder 1320 is to adjust the relative magnitude of the reflection output signal 1340 for the listener so that the reflections generate the sensation of emanating from the main direction angle Az with a spatial distribution (e.g., more or less concentrated around its main direction) determined by the setting of the spatial focus parameter F.

[0078] In some embodiments, the reflection main direction angle Az coincides, for each source, with the apparent arrival direction of the direct path, which can be controlled for each source by the direct pan module 1020. The simulated reflections can enhance the perception of the directional position of the virtual sound source perceived by the listener.

[0079] In some embodiments, the main mixing bus 940 and the direct pan module 1020 can enable three - dimensional reproduction of the sound direction. In these embodiments, the reflection main direction angle Az can coincide with the projection of the apparent direction onto the plane on which the reflection main angle Az is measured.

[0080] FIG. 15 illustrates a model 1500 of exemplary direct gain, reflection gain, and reverberation gain as a function of distance (e.g., to the listener) according to some embodiments. The model 1500 illustrates an example of the variation of the direct, reflection, and reverberation transmission gains, such as shown in FIG. 10, with respect to the source distance. As shown in the figure, the direct sound, its reflections, and its reverberations can have significantly different fall - off curves with respect to distance. In some cases, per - source processing such as that described above can enable a faster distance - based roll - off for reflections than for reverberations. Psychophysically, this can enable robust directivity perception and distance perception, particularly for distant sources.

[0081] FIG. 16 illustrates an exemplary model 1600 of the spatial focus-to-source distance for direct and reflected components, according to some embodiments. In this example, the direct pan module 1020 is configured to produce a maximum spatial concentration of the direct path component in the direction of the sound source, regardless of the distance. On the other hand, the reflected spatial focus parameter F may be set to an exemplary value of 2 / 3 in a realistic manner to enhance the directivity perception for all distances longer than a limit distance (e.g., the reflection minimum distance 1610). As illustrated by the exemplary model 1600, the reflected spatial focus parameter value decreases towards zero as the source approaches the listener.

[0082] FIG. 17 shows an exemplary model 1700 of the amplitude of an audio signal as a function of time. As described above, the reflected processing path (e.g., 1210A) may receive a signal from the reflected transmission bus and output the signal onto the main mixing bus. The reflected processing path may include a reflected global gain (e.g., 1220), a reflected global delay (e.g., 1222) for controlling a parameter Der as shown in the model 1700, and / or a reflected module (e.g., 1224), such as those described above.

[0083] As described above, the reverberation processing path (e.g., 1210B) may receive a signal from the reverberation transmission bus and output the signal into the main mixing bus. The reverberation processing path 1210B may include a reverberation global gain (e.g., 1230) for controlling parameter Lgo as shown in model 1700, a reverberation global delay (e.g., 1232) for controlling parameter Drev as shown in model 1700, and / or a reverberation module (e.g., 1234). The processing blocks within the reverberation processing path may be implemented in any suitable order. Examples of reverberation modules are described in U.S. Patent Application No. 62 / 685235 entitled "REVERBERATION GAIN NORMALIZATION" and U.S. Patent Application No. 62 / 684086 entitled "LOW-FREQUENCY INTERCHANNEL COHERENCE CONTROL", the contents of each of which are hereby incorporated by reference in their entirety.

[0084] The model 1700 of FIG. 17 illustrates a way in which source-specific parameters including distance and reverberation delay can be considered to dynamically adjust reverberation delay and level, according to some embodiments. In the figure, Dtof represents the delay due to the time of flight for a given object, i.e., Dtof = ObjDist / c, where ObjDist is the object distance from the center of the listener's head and c is the speed of sound in air. Drm represents the room delay per object. Dobj represents the total delay per object, i.e., Dobj = Dtof + Drm. Der represents the global early reflection delay. Drev represents the global reverberation delay. Dtotal represents the total delay for a given object, i.e., Dtotal = Dobj + Dglobal.

[0085] Lref represents the level of the response with respect to Dtotal = 0. Lgo represents the global level offset due to global delay, which can be calculated according to Equation 10, where T60 is the reverberation time of the reverberation algorithm. Loo represents the level offset per object due to global delay, which can be calculated according to Equation 11. Lto represents the total level offset for a given object and can be calculated according to Equation 12 (assuming dB values).

Chem.

Chem.

[0086] In some embodiments, the reverberation level is calibrated independently of object position, reverberation time, and other user-controllable parameters. Thus, Lrev can be the extrapolated level of the decaying reverberation at the initial time of sound emission. Lrev can be the same quantity as the Reverberation Initial Power (RIP) defined in U.S. Patent Application No. 62 / 685235, titled "REVERBERATION GAIN NORMALIZATION", the content of which is hereby incorporated by reference in its entirety. Lrev can be calculated according to Equation 13.

Chem.

[0087] In some embodiments, T60 may be a function of frequency. Accordingly, Lgo, Loo, and as a result, Lto are frequency-dependent.

[0088] FIG. 18 illustrates an exemplary system 1800 for determining spatial audio properties based on an acoustic environment. The exemplary system 1800 can be used to determine spatial audio properties related to reflections and / or reverberations such as those described above. As an example, such properties may include the volume of the room, the reverberation time as a function of frequency, the position of the listener relative to the room, the presence of indoor objects (e.g., sound attenuation objects), surface materials, or other suitable properties. In some embodiments, these spatial audio properties may be locally read out by capturing a single impulse response using microphones and loudspeakers freely positioned within the local environment, or may be adaptively derived by continuously monitoring and analyzing the sound captured by a mobile device microphone. In some embodiments, such as when the acoustic environment can be sensed via sensors of an XR system (e.g., an extended reality system including one or more of the wearable head unit 100, the handheld controller 200, and the auxiliary unit 300 described above), the location of the user can be used to present audio reflections and reverberations corresponding to the environment presented to the user (e.g., via a display).

[0089] In the exemplary system 1800, an acoustic environment perception module 1810 identifies the spatial audio properties of the acoustic environment, such as those described above. In some embodiments, the acoustic environment perception module 1810 can capture data corresponding to the acoustic environment (step 1812). For example, the data captured at step 1812 can include audio data from one or more microphones, camera data from a camera such as an RGB camera or a depth camera, LIDAR data, sonar data, radar data, GPS data, or other suitable data that can convey information about the acoustic environment. In some instances, the data captured at step 1812 can include data related to the user, such as the position or orientation of the user with respect to the acoustic environment. The data captured at step 1812 can be captured via one or more sensors of a wearable device such as the wearable head unit 100 described above.

[0090] In some embodiments, the local environment in which a head-mounted display device exists may include one or more microphones. In some embodiments, one or more microphones may be employed, mounted on a mobile device, positioned in the environment, or both. The benefits of such an arrangement can include collecting directional information about the room's reverberation or reducing the poor signal quality of any one of the one or more microphones. The signal quality can be poor on a given microphone, for example, due to occlusion, overload, wind noise, transducer damage, and the like.

[0091] At step 1814 of module 1810, features can be extracted from the data captured at step 1812. For example, the dimensions of the room can be determined from sensor data such as camera data, LIDAR data, sonar data, etc. The features extracted at step 1814 can be used to determine one or more acoustic properties of the room, such as the frequency-dependent reverberation time, and these properties can be stored at step 1816 and associated with the current acoustic environment.

[0092] In some embodiments, module 1810 can communicate with database 1840 to store and retrieve acoustic properties regarding an acoustic environment. In some embodiments, the database may be stored locally on the memory of the device. In some embodiments, the database may be stored online as a cloud-based service. The database may assign geographic locations to room properties for easy access at a later time based on the location of the listener. In some embodiments, the database may contain additional information to identify the location of the listener and / or to determine reverberation properties in the database that are a close approximation of the environmental properties of the listener. For example, room properties may be classified by room type, so that a set of parameters can be used as soon as it is identified that the listener is present within a known type of room (e.g., a bedroom or a living room), even if the absolute geographic location cannot be ascertained.

[0093] The storage of reverberation properties into the database may be related to U.S. Patent Application No. 62 / 573,448, titled "PERSISTENT WORLD MODEL SUPPORTING AUGMENTED REALITY AND INCLUDING AUDIO COMPONENT", the content of which is incorporated herein by reference in its entirety.

[0094] In some embodiments, system 1800 can include a reflection adaptation module 1820 for reading out acoustic properties of a room and applying those properties to audio reflections (e.g., audio reflections presented to the user of the wearable head unit 100 via headphones or via speakers). At stage 1822, the user's current acoustic environment can be determined. For example, GPS data can indicate the user's location within the GPS coordinates, which in turn can indicate the user's current acoustic environment (e.g., the room located at those GPS coordinates). As another example, camera data combined with optical recognition software can be used to identify the user's current environment. The reflection adaptation module 1820 can then communicate with a database 1840 to read out the acoustic properties associated with the determined environment, and those acoustic properties are used at stage 1824 to update the audio rendering accordingly. That is, the acoustic properties related to the reflection (e.g., the directivity pattern or the fall-off curve such as those described above) can be applied to the reflected audio signal presented to the user such that the presented reflected audio signal incorporates those acoustic properties.

[0095] Similarly, in some embodiments, system 1800 can include a reflection adaptation module 1830 for reading out the acoustic properties of a room and applying those properties to audio reverberation (e.g., audio reflections presented to the user of the wearable head unit 100 via headphones or via speakers). The prominent acoustic properties regarding reverberation can be different from those regarding reflections as described above (e.g., in Table 700 regarding FIG. 7). At step 1832, as described above, the user's current acoustic environment can be determined. For example, GPS data can indicate the location of the user within the GPS coordinates, which in turn can indicate the user's current acoustic environment (e.g., the room located at those GPS coordinates). As another example, camera data combined with optical recognition software can be used to identify the user's current environment. The reverberation adaptation module 1830 can then communicate with a database 1840 to read out the acoustic properties associated with the determined environment, and those acoustic properties can be used at step 1824 to update the audio rendering accordingly. That is, the acoustic properties related to reverberation (e.g., the reverberation decay time as described above) can be applied to the reverberation audio signal presented to the user such that the presented reverberation audio signal incorporates those acoustic properties.

[0096] Regarding the systems and methods described above, the elements of the present systems and methods can be implemented, as appropriate, by one or more computer processors (e.g., a CPU or DSP). The present disclosure is not limited to any particular configuration of computer hardware that is used to implement these elements. In some cases, multiple computer systems can be employed to implement the systems and methods described above. For example, a first computer processor (e.g., the processor of a wearable device coupled to a microphone) can be utilized to receive input microphone signals and perform initial processing of those signals (e.g., signal conditioning and / or segmentation such as that described above). A second (and perhaps more computationally powerful) processor can then be utilized to perform more computationally intensive processing such as the determination of probability values associated with the utterance segments of those signals. Another computer device such as a cloud server can host the speech recognition engine and the input signals are ultimately provided thereto. Other suitable configurations will also become apparent and are within the scope of the present disclosure.

[0097] It should be noted that the disclosed embodiments have been fully described with reference to the accompanying drawings, but various changes and modifications will be apparent to those skilled in the art. For example, the elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are understood to be included within the scope of the disclosed embodiments as defined by the appended claims.

Claims

1. A method for presenting an output audio signal to a listener located at a first location within a virtual environment, the method comprising: Receiving an input audio signal; For each sound source of a plurality of sound sources within the virtual environment, Determining an individual first intermediate audio signal corresponding to the input audio signal based on the location of the individual sound source within the virtual environment; Associating the individual first intermediate audio signal with a first bus; Determining an individual second intermediate audio signal based on the location of the individual sound source and further based on the acoustic properties of the virtual environment, the individual second intermediate audio signal corresponding to reflections of the input audio signal within the surface of the virtual environment, the acoustic properties of the virtual environment being determined via one or more sensors located on a wearable head unit worn by the listener, the one or more sensors being configured to determine the head pose of the listener, and the azimuth angle between the axis of the listener and the direction of the individual sound source being determined based on the head pose; and Associating the individual second intermediate audio signal with a second bus Performing; Presenting the output audio signal to the listener via the first bus and the second bus and further via a speaker of the wearable head unit; Including, the acoustic properties being read from a database, and reading the acoustic properties includes: Determining the first location of the listener based on the output of the one or more sensors; and Identifying the acoustic properties based on the first location of the listener; The output audio signal is determined via decoding the individual second intermediate audio signal and further via adjusting the magnitude of the individual second intermediate signal according to the azimuth angle. A method.

2. For each sound source of the plurality of sound sources within the virtual environment, Determining an individual third intermediate audio signal based on the location of the individual sound source and further based on a second acoustic property of the virtual environment, the individual third intermediate audio signal corresponding to reverberation of the input audio signal within the virtual environment; and Associating the individual third intermediate audio signal with the second bus Further including performing, The second bus comprises a reflection bus and a reverberation bus. Associating the individual second intermediate audio signals with the second bus includes associating the individual second intermediate audio signals with the reflection bus. Associating the individual third intermediate audio signals with the second bus includes associating the individual third intermediate audio signals with the reverberation bus. The method according to claim 1.

3. The method according to any one of claims 1 or 2, wherein the one or more sensors comprise one or more microphones and a LIDAR sensor.

4. The output signal is presented to the listener via one or more speakers associated with the wearable head unit. The method according to any one of claims 1 or 2.

5. The method according to claim 4, wherein the wearable head unit comprises a display configured to display a view of the virtual environment to the listener in parallel with the presentation of the output signal.

6. A wearable device worn by a listener located at a first location within a virtual environment, a display configured to display a view of the virtual environment; one or more sensors located on the wearable device; one or more speakers; one or more processors, receiving an input audio signal; for each sound source of a plurality of sound sources within the virtual environment, determining an individual first intermediate audio signal corresponding to the input audio signal based on the location of the individual sound source within the virtual environment; associating the individual first intermediate audio signal with a first bus; determining an individual second intermediate audio signal based on the location of the individual sound source and further based on the acoustic properties of the virtual environment, wherein the individual second intermediate audio signal corresponds to a reflection of the input audio signal within a surface within the virtual environment, the acoustic properties of the virtual environment are determined via the one or more sensors, the one or more sensors are configured to determine the head pose of the listener, and the azimuth angle between the axis of the listener and the direction of the individual sound source is determined based on the head pose; assigning a minimum distance parameter to the individual sound source; determining the distance between the location of the individual sound source and the listener; determining a gain for the individual sound source based on the minimum distance parameter and the distance between the location of the individual sound source and the listener. When the distance is not greater than the minimum distance parameter, the gain has a constant value, when the distance is greater than the minimum distance parameter, the gain has a value rolled off from the constant value, and the amount of roll-off is determined based on the distance and is further determined based on the acoustic properties of the virtual environment, applying the gain to one or more of the first intermediate audio signal and the second intermediate audio signal, and associating the individual second intermediate audio signal with a second bus performing, presenting an output audio signal to the listener via the speaker and via the first bus and the second bus configured to implement a method including, wherein the acoustic properties are read from a database, and reading the acoustic properties includes determining the first location of the listener based on the output of the one or more sensors, and identifying the acoustic properties based on the first location of the listener including one or more processors, and the output audio signal is determined via decoding the individual second intermediate audio signal and further via adjusting the magnitude of the individual second intermediate signal according to the azimuth angle. A wearable device. **Claim 7** The method further includes, for each sound source of the plurality of sound sources in the virtual environment, determining an individual third intermediate audio signal based on the location of the individual sound source and further based on second acoustic properties of the virtual environment, wherein the individual third intermediate audio signal corresponds to the reverberation of the input audio signal in the virtual environment, and associating the individual third intermediate audio signal with the second bus including performing, wherein the second bus includes a reflection bus and a reverberation bus, associating the individual second intermediate audio signal with the second bus includes associating the individual second intermediate audio signal with the reflection bus, and associating the individual third intermediate audio signal with the second bus includes associating the individual third intermediate audio signal with the reverberation bus. The wearable device according to claim 6. **Claim 8** The wearable device according to any one of claims 6 or 7, wherein the one or more sensors comprise one or more microphones and a LIDAR sensor.

9. The wearable device according to any one of claims 6 or 7, wherein the method further comprises displaying a view of the virtual environment via the display in parallel with the presentation of the output signal.

10. Determining the individual first intermediate audio signal includes applying a first individual filter to the input audio signal, the first individual filter comprising one or more of a sound source directivity model, a distance model, and an orientation model. The method according to any one of claims 1 or 2.

11. Determining the individual first intermediate audio signal further includes applying one or more of an individual gain and an individual panning process to the input audio signal. The method according to claim 10.

12. The individual panning process includes panning the input audio signal based on the geometry of the loudspeaker array. The method according to claim 11.

13. Determining the individual second intermediate audio signal includes applying a second individual filter to the input audio signal, the second individual filter comprising a sound source directivity model. The method according to any one of claims 1 or 2.

14. Determining the individual second intermediate audio signal further includes applying one or more of an individual delay, an individual gain, and an individual panning process to the input audio signal. The method according to claim 13.

15. The individual panning process includes encoding the input audio signal into an ambisonic signal including three channels. The method according to claim 14.

16. The individual panning process includes panning the reflection of the input audio signal based on one or more of an azimuth angle and a spatial focus parameter. The method according to claim 14.

17. Determining the individual first intermediate audio signal includes applying a first individual filter to the input audio signal, the first individual filter comprising one or more of a sound source directivity model, a distance model, and an orientation model, the wearable device according to any one of claims 6 or 7.

18. Determining the individual first intermediate audio signal further includes applying one or more of an individual gain and an individual panning process to the input audio signal, the wearable device according to claim 17.

19. The individual panning process includes panning the input audio signal based on the geometry of the loudspeaker array, the wearable device according to claim 18.

20. Determining the individual second intermediate audio signal includes applying a second individual filter to the input audio signal, the second individual filter comprising a sound source directivity model, the wearable device according to any one of claims 6 or 7.

21. Determining the individual second intermediate audio signal further includes applying one or more of an individual delay, an individual gain, and an individual panning process to the input audio signal, the wearable device according to claim 20.

22. The individual panning process includes encoding the input audio signal into an ambisonic signal including three channels, the wearable device according to claim 21.

23. The individual panning process includes panning the reflection of the input audio signal based on one or more of an azimuth angle and a spatial focus parameter, the wearable device according to claim 21.

24. Determining the individual second intermediate audio signal includes converting the individual second intermediate audio signal from a first format to a second format, and encoding the processed input audio signal based on the location of the individual sound source and the individual second intermediate audio signal is associated with a first plurality of channels, the second bus is associated with the first plurality of channels and the first format, The output audio signal is determined via decoding the individual second intermediate audio signals, The decoded individual second intermediate audio signals are associated with a second plurality of channels, The method according to any one of claims 1 or 2, wherein the first bus is associated with the second plurality of channels and the second format. **Claim 25** The one or more sensors comprise an inertial measurement unit (IMU) and a camera, The method according to any one of claims 1 or 2, wherein the first location of the listener is determined based on the output of the IMU and the output of the camera. **Claim 26** A first sound source is associated with a moving virtual object in the virtual environment, A first acoustic property of the acoustic properties of the sound source is associated with the moving virtual object, Identifying the acoustic property includes further identifying the first acoustic property based on the location of the moving virtual object, and the location relative to the moving virtual object is determined based on a second output of the one or more sensors. The method according to any one of claims 1 or 2.

Citation Information

Patent Citations

  • Determining and Using Room-Optimized Transfer Functions

    JP2017522771A

  • Digital camera with audio, visual and motion analysis

    US20170345216A1

  • Method and apparatus for the simulation of complex audio environments

    US7099482B1

  • Mixed reality system with spatialized audio

    WO2018026828A1