Spatial audio for interactive audio environments

CN116156411BActive Publication Date: 2026-09-01MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211572588.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-18
Filing Date
2019-06-18
Publication Date
2026-09-01
Estimated Expiration
2039-06-18

AI Technical Summary

Technical Problem

然而,用户在虚拟环境中的体验可能会受到用于呈现虚拟环境的技术的限制

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116156411B_ABST
    Figure CN116156411B_ABST
Patent Text Reader

Abstract

A system and method for presenting an output audio signal to a listener located at a first position in a virtual environment are disclosed. According to an embodiment of the method, an input audio signal is received. For each of a plurality of sound sources in the virtual environment, a corresponding first intermediate audio signal corresponding to the input audio signal is determined based on the position of the corresponding sound source in the virtual environment, and the corresponding first intermediate audio signal is associated with a first bus. For each of the plurality of sound sources in the virtual environment, a corresponding second intermediate audio signal is determined. The corresponding second intermediate audio signal corresponds to a reflection of the input audio signal on a surface of the virtual environment. The corresponding second intermediate audio signal is determined based on the position of the corresponding sound source and further based on the acoustic characteristics of the virtual environment. The corresponding second intermediate audio signal is associated with a second bus. An output audio signal is presented to the listener via the first bus and the second bus.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application filed on June 18, 2019, with application number 201980053576.5 and invention title "Spatial Audio for Interactive Audio Environment".

[0002] Cross-reference of related applications

[0003] This application claims priority to U.S. Provisional Application No. 62 / 686,655, filed June 18, 2018, the entire contents of which are incorporated herein by reference. This application also claims priority to U.S. Provisional Application No. 62 / 686,665, filed June 18, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0004] This disclosure generally relates to spatial audio rendering, and particularly to spatial audio rendering for virtual sound sources in a virtual acoustic environment. Background Technology

[0005] Virtual environments are ubiquitous in computing environments and are found in video games (where a virtual environment can represent a game world), maps (where a virtual environment can represent terrain to be navigated), simulations (where a virtual environment can simulate a real environment), digital storytelling (where virtual characters can interact with each other in a virtual environment), and many other applications. Modern computer users can generally perceive and interact with virtual environments easily. However, the user experience in a virtual environment can be limited by the technologies used to render it. For example, traditional displays (e.g., 2D displays) and audio systems (e.g., fixed speakers) may not be able to create a compelling, realistic, and immersive virtual environment.

[0006] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present users of XR systems with sensory information corresponding to a virtual environment represented by data in a computer system. By combining virtual visual and audio cues with real visuals and sounds, the system can provide a uniquely enhanced sense of immersion and realism. Therefore, it is expected that digital sound will be presented to users of XR systems in a way that makes the sound appear to occur naturally in the user's real environment and conforms to the user's expectations of sound. Generally, users expect virtual sound to reproduce the acoustic characteristics of the real environment they hear. For example, users of an XR system in a large concert hall will expect the virtual sound of the XR system to have a vast, cavernous quality; conversely, users in a small apartment will expect the sound to be softer, more intimate, and more direct.

[0007] Digital or artificial reverberators are used in audio and music signal processing to simulate the perceived effects of diffuse acoustic reverberation in a room. In XR environments, it is desirable to use digital reverberators to realistically simulate the acoustic characteristics of a room within the XR environment. A convincing simulation of such acoustic characteristics can bring realism and immersion to the XR environment. Summary of the Invention

[0008] A system and method for presenting an output audio signal to a listener located at a first position in a virtual environment are disclosed. According to an embodiment of the method, an input audio signal is received. For each of a plurality of sound sources in the virtual environment, a corresponding first intermediate audio signal corresponding to the input audio signal is determined based on the position of the corresponding sound source in the virtual environment, and the corresponding first intermediate audio signal is associated with a first bus. For each of the plurality of sound sources in the virtual environment, a corresponding second intermediate audio signal is determined. The corresponding second intermediate audio signal corresponds to a reflection of the input audio signal on a surface of the virtual environment. The corresponding second intermediate audio signal is determined based on the position of the corresponding sound source and further based on the acoustic characteristics of the virtual environment. The corresponding second intermediate audio signal is associated with a second bus. An output audio signal is presented to the listener via the first bus and the second bus. Attached Figure Description

[0009] Figure 1 An example wearable system according to some embodiments is shown.

[0010] Figure 2 An example handheld controller, according to some embodiments, is shown that can be used in conjunction with an example wearable system.

[0011] Figure 3 An example auxiliary unit, according to some embodiments, can be used in conjunction with an example wearable system.

[0012] Figure 4 An example functional block diagram for an example wearable system is shown according to some embodiments.

[0013] Figure 5 An example geometric room representation is shown according to some embodiments.

[0014] Figure 6 An example model of room response measured from the source to the listener in the room, according to some embodiments, is shown.

[0015] Figure 7 Examples of factors that affect a user’s perception of direct sound, reflections, and reverberation, according to some embodiments, are shown.

[0016] Figure 8An example audio mixing architecture for rendering multiple virtual sound sources in a virtual room is shown according to some embodiments.

[0017] Figure 9 An example audio mixing architecture for rendering multiple virtual sound sources in a virtual room is shown according to some embodiments.

[0018] Figure 10 An example per-source processing module is shown according to some embodiments.

[0019] Figure 11 An example per-source reflection pan module is shown according to some embodiments.

[0020] Figure 12 An example room processing algorithm according to some embodiments is shown.

[0021] Figure 13 An example reflection module according to some embodiments is shown.

[0022] Figure 14 Examples of spatial distributions showing apparent directions of reflection according to some embodiments are shown.

[0023] Figure 15 Examples of direct gain, reflection gain, and reverberation gain as a function of distance are shown according to some embodiments.

[0024] Figure 16 Examples of the relationship between distance and spatial focus are shown according to some embodiments.

[0025] Figure 17 Examples of the relationship between time and signal amplitude are shown according to some embodiments.

[0026] Figure 18 An example system for processing spatial audio is shown according to some embodiments. Detailed Implementation

[0027] In the following illustrative description, reference is made to the accompanying drawings, which form part of the description, and specific examples that can be practiced are illustrated in the drawings. It will be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.

[0028] Example wearable system

[0029] Figure 1An example wearable head device 100 configured to be worn on a user's head is shown. The wearable head device 100 may be part of a broader wearable system that includes one or more components, such as a head device (e.g., wearable head device 100), a handheld controller (e.g., handheld controller 200 described below), and / or an auxiliary unit (e.g., auxiliary unit 300 described below). In some examples, the wearable head device 100 may be used in virtual reality, augmented reality, or mixed reality systems or applications. The wearable head-mounted device 100 may include one or more displays, such as displays 110A and 110B (which may include left and right transmissive displays, and associated components for coupling light from the displays to the user's eyes, such as orthogonal pupil dilation (OPE) grating groups 112A / 112B and outgoing pupil dilation (EPE) grating groups 114A / 114B); left and right acoustic structures, such as speakers 120A and 120B (which may be mounted on temples 122A and 122B and positioned adjacent to the user's left and right ears); and one or more sensors, such as infrared sensors, accelerometers, GPS units, and inertial measurement units (IMUs) (e.g., IMUs). 126) Acoustic sensors (e.g., microphone 150); orthogonal coil electromagnetic receivers (e.g., receiver 127 shown mounted to the left temple 122A); left and right cameras facing away from the user (e.g., depth (time-of-flight) cameras 130A and 130B); and left and right eye cameras facing the user (e.g., for detecting the user's eye movements) (e.g., eye cameras 128 and 128B). However, the wearable head device 100 may incorporate any suitable display technology and any suitable number, type, or combination of sensors or other components without departing from the scope of the invention. In some examples, the wearable head device 100 may incorporate one or more microphones 150 configured to detect audio signals generated by the user's speech; the microphone may be located near the user's mouth in the wearable head device. In some examples, the wearable head device 100 may incorporate networking features (e.g., Wi-Fi functionality) to communicate with other devices and systems, including other wearable systems. The wearable head-mounted device 100 may further include components such as a battery, processor, memory, storage unit, or various input devices (e.g., buttons, touchpads); or may be coupled to a handheld controller (e.g., handheld controller 200) or auxiliary unit (e.g., auxiliary unit 300) that includes one or more such components. In some examples, sensors may be configured to output a set of coordinates of the head-mounted unit relative to the user's environment and may provide input to a processor that performs simultaneous localization and mapping (SLAM) processes and / or visual ranging algorithms.In some examples, as further described below, the wearable head device 100 may be coupled to a handheld controller 200 and / or an auxiliary unit 300.

[0030] Figure 2 An example mobile handheld controller assembly 200 of an example wearable system is shown. In some examples, the handheld controller 200 can communicate wirelessly with the wearable head device 100 and / or auxiliary unit 300 described below. In some examples, the handheld controller 200 includes a handle portion 220 to be held by a user, and one or more buttons 240 disposed along a top surface 210. In some examples, the handheld controller 200 can be configured to serve as an optical tracking target; for example, sensors of the wearable head device 100 (e.g., cameras or other optical sensors) can be configured to detect the position and / or orientation of the handheld controller 200, which, by extension, can indicate the position and / or orientation of the user's hand holding the handheld controller 200. In some examples, the handheld controller 200 may include a processor, memory, storage unit, display, or one or more input devices, such as those described above. In some examples, the handheld controller 200 includes one or more sensors (e.g., any sensors or tracking components described above with respect to the wearable head device 100). In some examples, the sensor can detect the position or orientation of the handheld controller 200 relative to the wearable head device 100 or relative to another component of the wearable system. In some examples, the sensor may be placed in the handle portion 220 of the handheld controller 200 and / or may be mechanically coupled to the handheld controller. The handheld controller 200 may be configured to provide one or more output signals, such as corresponding to a pressed state of button 240, or the position, orientation, and / or movement of the handheld controller 200 (e.g., via an IMU). This output signal may be used as an input to the processor of the wearable head device 100, the auxiliary unit 300, or another component of the wearable system. In some examples, the handheld controller 200 may include one or more microphones for detecting sound (e.g., the user's voice, ambient sound), and in some cases, provide a signal corresponding to the detected sound to a processor (e.g., the processor of the wearable head device 100).

[0031] Figure 3An example assistive unit 300 of an example wearable system is shown. In some examples, the assistive unit 300 may be wired or wirelessly connected to the wearable head device 100 and / or the handheld controller 200. The assistive unit 300 may include a battery to provide power for operating one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including displays, sensors, acoustic structures, processors, microphones, and / or other components of the wearable head device 100 or the handheld controller 200)). In some examples, the assistive unit 300 may include a processor, memory, storage unit, display, one or more input devices, and / or one or more sensors, as described above. In some examples, the assistive unit 300 includes a clip 310 for attaching the assistive unit to a user (e.g., a strap worn by the user). The advantage of using the auxiliary unit 300 to house one or more components of the wearable system is that it allows large or heavy components to be carried by the user's waist, chest, or back (which are relatively well-suited for supporting large and heavy objects), rather than mounted on the user's head (e.g., placed in the wearable headpiece 100) or carried by the user's hand (e.g., placed in the handheld controller 200). This can be particularly advantageous for relatively heavy or bulky components, such as batteries.

[0032] Figure 4 An example functional block diagram is shown that corresponds to an example wearable system 400 (such as the example wearable head device 100, handheld controller 200, and auxiliary unit 300 described above). In some examples, the wearable system 400 can be used for virtual reality, augmented reality, or mixed reality applications. Figure 4As shown, the wearable system 400 may include an example handheld controller 400B, referred to herein as a "totem" (and may correspond to the handheld controller 200 described above); the handheld controller 400B may include a totem-to-headband six-degree-of-freedom (6DOF) totem subsystem 404A. The wearable system 400 may also include an example wearable head device 400A (which may correspond to the wearable headband device 100 described above); the wearable head device 400A includes a totem-to-headband 6DOF headband subsystem 404B. In the example, the 6DOF totem subsystem 404A and the 6DOF headband subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom can be represented relative to the coordinate system of the wearable head device 400A. In this coordinate system, the three translational offsets can be represented as X, Y, and Z offsets, as translation matrices, or as some other representation. Rotational degrees of freedom can be represented as a sequence of yaw, pitch, and roll rotations; as vectors; as rotation matrices; as quaternions; or as some other representation. In some examples, one or more depth cameras 444 (and / or one or more non-depth cameras) included in the wearable head-mounted device 400A; and / or one or more optical targets (e.g., buttons 240 of the handheld controller 200 as described above, or dedicated optical targets included in the handheld controller) can be used for 6DOF tracking. In some examples, as described above, the handheld controller 400B may include a camera; and the headband 400A may include optical targets used in conjunction with the camera for optical tracking. In some examples, both the wearable head-mounted device 400A and the handheld controller 400B include a set of three orthogonally oriented solenoids for wirelessly transmitting and receiving three distinguishable signals. The 6DOF of the handheld controller 400B relative to the wearable head device 400A can be determined by measuring the relative magnitudes of the three distinguishable signals received in each coil used for reception. In some examples, the 6DOF totem subsystem 404A may include an inertial measurement unit (IMU) that can be used to provide improved accuracy and / or more timely information about the rapid movements of the handheld controller 400B.

[0033] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a fixed coordinate space relative to the wearable head-mounted device 400A) to an inertial coordinate space or to an environmental coordinate space. For example, such a transformation may be necessary for the display of the wearable head-mounted device 400A to present virtual objects (e.g., a virtual person sitting in a real chair facing forward, regardless of the position and orientation of the wearable head-mounted device 400A) at a desired position and orientation relative to the real environment, rather than at a fixed position and orientation on the display (e.g., the same position within the display of the wearable head-mounted device 400A). This preserves the illusion of the virtual objects existing in the real environment (and, for example, does not appear unnaturally in the real environment as the wearable head-mounted device 400A moves and rotates). In some examples, the compensation transformation between coordinate spaces can be determined by processing images from the depth camera 444 (e.g., using simultaneous localization and mapping (SLAM) and / or visual ranging processes) to determine the transformation of the wearable head-mounted device 400A relative to the inertial or environmental coordinate system. Figure 4 In the example shown, depth camera 444 can be coupled to SLAM / visual ranging module 406 and can provide images to frame 406. Implementation of SLAM / visual ranging module 406 may include a processor configured to process the image and determine the position and orientation of the user's head, which can then be used to identify transformations between head coordinate space and ground coordinate space. Similarly, in some examples, additional information about the user's head pose and position is obtained from IMU 409 of wearable head device 400A. Information from IMU 409 can be integrated with information from SLAM / visual ranging module 406 to provide improved accuracy and / or more timely information for rapid adjustments to the user's head pose and position.

[0034] In some examples, depth camera 444 can provide 3D images to gesture tracker 411, which can be implemented in the processor of wearable head device 400A. Gesture tracker 411 can recognize a user's gestures, for example, by matching the 3D images received from depth camera 444 with stored patterns representing gestures. Other suitable techniques for recognizing user gestures will be readily apparent.

[0035] In some examples, one or more processors 416 may be configured to receive data from a headband subsystem 404B, an IMU 409, a SLAM / visual ranging module 406, a depth camera 444, a microphone (not shown), and / or a gesture tracker 411. Processor 416 may also send and receive control signals from a 6DOF totem system 404A. In examples where the handheld controller 400B is unrestricted, processor 416 may be wirelessly coupled to the 6DOF totem system 404A. Processor 416 may further communicate with additional components such as an audiovisual content memory 418, a graphics processing unit (GPU) 420, and / or a digital signal processor (DSP) audio field locator 422. The DSP audio field locator 422 may be coupled to a head-related transfer function (HRTF) memory 425. GPU 420 may include a left-channel output coupled to a left source 424 of per-image modulated light and a right-channel output coupled to a right source 426 of per-image modulated light. GPU 420 can output stereoscopic image data to sources 424 and 426 of per-image modulated light. DSP audio field locator 422 can output audio to left speaker 412 and / or right speaker 414. DSP audio field locator 422 can receive input from processor 416 indicating a direction vector from the user to a virtual sound source (which can be moved by the user, for example, via handheld controller 400B). Based on the direction vector, DSP audio field locator 422 can determine the corresponding HRTF (e.g., by accessing the HRTF or by interpolating multiple HRTFs). DSP audio field locator 422 can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. This can enhance the credibility and realism of the virtual sound by incorporating the user's relative position and orientation to the virtual sound in the mixed reality environment (i.e., by presenting a virtual sound that matches the user's expectation of sounding like a real sound in a real environment).

[0036] In some examples, such as Figure 4 As shown, one or more of the processor 416, GPU 420, DSP audio sound field locator 422, HRTF memory 425, and audio / visual content memory 418 may be included in the auxiliary unit 400C (which may correspond to the auxiliary unit 300 described above). The auxiliary unit 400C may include a battery 427 to power its components and / or power the wearable head device 400A and / or the handheld controller 400B. Including such components in an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0037] exist Figure 4 While presenting the components corresponding to the various components of the example wearable system 400, various other suitable arrangements of these components will become apparent to those skilled in the art. For example, in Figure 4 The elements shown as associated with the auxiliary unit 400C may alternatively be associated with the wearable head device 400A or the handheld controller 400B. Furthermore, some wearable systems may completely omit the handheld controller 400B or the auxiliary unit 400C. It is understood that such changes and modifications are included within the scope of the disclosed examples.

[0038] Mixed Reality Environment

[0039] Like everyone else, users of mixed reality systems exist within a real-world environment—that is, the three-dimensional portion of the "real world" and all its contents that the user can perceive. For example, users perceive the real-world environment using ordinary human senses (sight, sound, touch, taste, smell) and interact with it by moving their bodies within it. A position in the real-world environment can be described as coordinates in a coordinate space; for example, coordinates can include latitude, longitude, and altitude relative to sea level; distance from a reference point in three orthogonal dimensions; or other suitable values. Similarly, vectors can describe quantities having direction and magnitude in coordinate space.

[0040] A computing device may maintain a representation of a virtual environment in, for example, memory associated with the device. As used herein, a virtual environment is a computational representation of a three-dimensional space. A virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other features associated with that space. In some examples, the circuitry of the computing device (e.g., a processor) may maintain and update the state of the virtual environment; that is, the processor may determine the state of the virtual environment at a second time based on data associated with the virtual environment and / or user-provided input at a first time. For example, if an object in the virtual environment is currently located at a first coordinate and has specifically programmed physical parameters (e.g., mass, coefficient of friction); and input received from the user indicates that a force with a direction vector should be applied to the object; then the processor may apply laws of kinematics to determine the object's position at that time using fundamental mechanics. The processor may use any appropriate known information about the virtual environment and / or any appropriate input to determine the state of the virtual environment at a given time. While maintaining and updating the state of the virtual environment, the processor may execute any suitable software, including software related to creating and deleting virtual objects in the virtual environment; software (e.g., scripts) for defining the behavior of virtual objects or characters in the virtual environment; software for defining the behavior of signals (e.g., audio signals) in the virtual environment; software for creating and updating parameters associated with the virtual environment; software for generating audio signals in the virtual environment; software for processing inputs and outputs; software for implementing network operations; software for applying asset data (e.g., animation data for moving virtual objects over time); or many other possibilities.

[0041] Output devices (such as displays or speakers) can present any or all aspects of a virtual environment to a user. For example, a virtual environment can include virtual objects that can be presented to the user (which may include representations of inanimate objects, people, animals, lights, etc.). The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" with original coordinates, view axes, and a frustum); and render the visible scene of the virtual environment corresponding to that view to the display. Any suitable rendering technique can be used for this purpose. In some examples, the visible scene may include only some virtual objects in the virtual environment, excluding certain other virtual objects. Similarly, a virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects in the virtual environment may generate sounds derived from the object's location coordinates (e.g., a virtual character can speak or cause sound effects); or the virtual environment may be associated with musical cues or ambient sounds (which may or may not be associated with a specific location). The processor can determine the audio signal corresponding to the "listener" coordinates (e.g., the audio signal corresponding to the synthesis of multiple sounds in a virtual environment, and mix and process it to simulate the audio signal heard by the listener at the listener coordinates), and present the audio signal to the user via one or more speakers.

[0042] Because the virtual environment exists solely as a computational structure, users cannot directly perceive it using their ordinary senses. Instead, they can only perceive it indirectly, for example, through displays, speakers, haptic output devices, etc. Similarly, users cannot directly touch, manipulate, or otherwise interact with the virtual environment; however, they can provide input data to a processor that can update the virtual environment using this data via input devices or sensors. For example, a camera sensor can provide optical data indicating that a user is attempting to move an object in the virtual environment, and the processor can use this data to make the object respond accordingly in the virtual environment.

[0043] Reflection and reverberation

[0044] Aspects of the listener’s audio experience in a virtual environment’s space (e.g., a room) include the listener’s perception of direct sound, the listener’s perception of the reflection of that direct sound on the room’s surfaces, and the listener’s perception of the reverberation of the direct sound in the room. Figure 5 A geometric room representation 500 is shown according to some embodiments. The geometric room representation 500 shows example propagation paths of direct sound (502), reflections (504), and reverberation (506). These paths represent the paths by which an audio signal can propagate from a source to a listener in the room. Figure 5The room shown can be any suitable type of environment associated with one or more acoustic properties. For example, room 500 could be a concert hall and could include a stage for the pianist and seating areas for the audience. As shown, direct sound is sound originating from a source (e.g., the pianist) and propagating directly toward the listener (e.g., the audience members). Reflection is sound emitted from a source, reflected by a surface (e.g., the walls of the room), and propagating to the listener. Reverberation is sound that includes a decaying signal comprising numerous reflections that are close to each other in time.

[0045] Figure 6 An example model 600 of the room response, measured from the source to a listener in the room, is shown according to some embodiments. The model of the room response shows the amplitude of the direct sound (610), the reflection of the direct sound (620), and the reverberation (630) of the direct sound from the listener at an angle at a distance from the direct sound source. Figure 6 As shown, the direct sound typically arrives at the listener before the reflection (reflection_delay (622) in the figure indicates the time difference between the direct sound and the reflection), which in turn arrives before the reverberation (reverberation_delay (632) in the figure indicates the time difference between the direct sound and the reverberation). Reflections and reverberation may be perceived differently by the listener. Reflections can be modeled separately from reverberation, for example, to better control the timing, decay, spectral shape, and direction of arrival of individual reflections. Reflections can be modeled using a reflection model, and reverberation can be modeled using a reverberation model that may differ from the reflection model.

[0046] The reverberation characteristics (e.g., reverberation attenuation) of the same sound source may differ between two different acoustic environments (e.g., rooms) for the same sound source, and it is desirable to realistically reproduce the sound source based on the characteristics of the current room in the listener's virtual environment. That is, when presenting a virtual sound source in a mixed reality system, the reflection and reverberation characteristics of the listener's real environment should be accurately reproduced. In the Journal of the Audio Engineering Society (J.Audio Eng.Soc.) 47(9):675–705 (1999), L. Savioja, J. Huopanemi, T. Lokki and The paper "Creating Interactive Virtual Acoustic Environments" describes methods for reproducing direct paths, individual reflections, and sound reverberation in real-time virtual 3D audio reproduction systems for video games, simulations, or AR / VR. In the methods disclosed by Savioja et al., the arrival direction, delay, amplitude, and spectral equalization of individual reflections are derived from the geometry and physical model of a room (e.g., a real room, a virtual room, or some combination thereof), which may require a complex rendering system. These methods can be computationally complex and may be prohibitively expensive for mobile applications where computational resources can be very valuable.

[0047] In some room acoustics simulation algorithms, reverberation can be achieved by downmixing all sound sources into a single-channel signal and sending the single-channel signal to a reverberation simulation module. The gain used for downmixing and sending can depend on dynamic parameters (such as source distance) and manual parameters (such as reverberation gain).

[0048] Sound source directionality, or radiation pattern, can refer to a measure of how much energy a sound source emits in different directions. Sound source directionality has an effect on all parts of the room's impulse response (e.g., direct, reflected, and reverberant). Different sound sources can exhibit different directions of directionality; for example, human speech can have a different directionality pattern than trumpet playing. Room simulation models may take sound source directionality into account when generating accurate simulations of acoustic signals. For example, a model incorporating sound source directionality may include a function of the direction of the path from the sound source to the listener relative to the forward direction (or main acoustic axis) of the sound source. The directionality pattern is axisymmetric about the main acoustic axis of the sound source. In some embodiments, the parametric gain model can be defined using a frequency-dependent filter. In some embodiments, to determine how much audio from a given sound source should be sent to the reverberation bus, the average diffusion power of the sound source can be calculated (e.g., by integrating over a sphere centered at the acoustic center of the sound source).

[0049] Interactive audio engines and sound design tools can make assumptions about the acoustic system being modeled. For example, some interactive audio engines can model sound source directionality as a frequency-independent function, which can have two potential drawbacks. First, it may ignore frequency-dependent attenuation in the direct sound propagation from the sound source to the listener. Second, it may ignore frequency-dependent attenuation in reflections and reverberation. From a psychoacoustic perspective, these effects can be significant, and failing to reproduce them can lead to room simulations being perceived as unnatural and different from simulations that listeners are accustomed to experiencing in real acoustic environments.

[0050] In some cases, room simulation systems or interactive audio engines may not be able to completely separate sound sources, listeners, and acoustic environment parameters such as reflections and reverberation. Instead, room simulation systems may make holistic adjustments for a specific virtual environment and may not adapt to different playback scenarios. For example, reverberation in a simulated environment may not match the environment in which the user / listener is physically present when listening to rendered content.

[0051] In augmented reality or mixed reality applications, computer-generated audio objects can be rendered via a transparent playback system to blend with the physical environment that the user / listener naturally hears. This may require binaural artificial reverberation processing to match the local ambient acoustics, making the synthesized audio objects indistinguishable from naturally generated or loudspeaker-reproduced sounds. Methods involving measuring or calculating room impulse responses (e.g., based on estimates of the environment's geometry) may be limited by practical obstacles and complexities in consumer environments. Furthermore, physical models may not provide the most compelling auditory experience because they may not account for psychoacoustic principles or offer audio scene parameterization suitable for sound designers to adjust the auditory experience.

[0052] Matching certain physical characteristics of the target acoustic environment may not provide a simulation that perceptually closely matches the listener's environment or the application designer's intent. A perceptually relevant model of the target acoustic environment, which can be characterized by an interface that describes the actual audio environment, may be needed.

[0053] For example, a rendering model that separates the contributions of the source, the listener, and room characteristics might be needed. A rendering model that separates contributions allows components to be adapted or swapped at runtime based on the characteristics of the end user and the local environment. For instance, the listener might be in a physical room with acoustic characteristics different from the virtual environment in which the content was originally created. Modifying the early reflections and / or reverberation portions of the simulation to match the listening environment can result in a more convincing listening experience. Matching the listening environment can be particularly important in mixed reality applications, where the desired effect might be that the listener cannot distinguish which sounds around them are simulated and which are present in their real surroundings.

[0054] It may be desirable to produce convincing results without requiring detailed knowledge of the geometry and / or acoustic properties of the surrounding environment. Detailed knowledge of the characteristics of the real-world environment may be unavailable, or estimating it may be complex, especially on portable devices. Instead, models based on perceptual and psychoacoustic principles may be a more practical tool for characterizing the acoustic environment.

[0055] Figure 7Table 700, according to some embodiments, is shown. This table includes objective acoustic and geometric parameters that characterize each part of a binaural room impulse model, thus distinguishing the characteristics of the source, listener, and room. Some source characteristics may be independent of how and where the content is rendered (including free-field and diffuse-field transfer functions), while others may need to be dynamically updated during playback (including position and orientation). Similarly, some listener characteristics may be independent of the position where the content is rendered (including free-field and diffuse-field head-related transfer functions or diffuse-field interaural coherence (IACC)), while others may be dynamically updated during playback (including position and orientation). Some room characteristics (particularly those contributing to post-reverberation) may be entirely environmentally dependent. Representations of reverberation decay rate and room cube product allow the spatial audio rendering system to adapt to the listener's playback environment.

[0056] The source and the listener's ear can be modeled as transmitting and receiving transducers, each characterized as a set of direction-dependent free-field transfer functions, including the listener's head-related transfer function (HRTF).

[0057] Figure 8 An example audio mixing system 800 for rendering multiple virtual sound sources in a virtual room (such as an XR environment) according to some embodiments is shown. For example, the audio mixing architecture may include a rendering engine for room acoustic simulation of multiple virtual sound sources 810 (i.e., objects 1 to N). System 800 includes a room send bus 830 that feeds rendering reflections and reverberation modules 850 (e.g., shared reverberation and reflection modules). Aspects of this general process are described, for example, in the IA-SIG 3D Audio Rendering Guide (Level 2), www.iasig.net (1999). The room send bus combines contributions from all sources (e.g., sound source 810, each processed by a corresponding module 820) to obtain the input signal for the room modules. The room send bus may include a mono room send bus. The format of the main mixing bus 840 may be a two-channel or multi-channel format matched to the method of rendering the final output, which may, for example, include a binaural renderer, a surround sound decoder, and / or a multi-channel speaker system for headphone playback. The main hybrid bus combines contributions from all sources with the room module output to obtain the output rendering signal 860.

[0058] Referring to example system 800, each of the N objects can represent a virtual sound source signal and can be assigned a definite location in the environment, such as through a panning algorithm. For example, each object can be assigned an angular position on a sphere centered on the virtual listener's location. The panning algorithm can calculate the contribution of each object to each channel of the main mix. This general process is described, for example, in "A comparative study of 3-D audio encoding and rendering techniques" by J.-M. Jot, V. Larcher, and J.-M. Pernaux at the 16th International Conference on Spatial Sound Reproduction (Proc. AES 16th International Conference on Spatial Sound Reproduction) (1999). Each object can be input to a panning gain module 820, which implements the panning algorithm and performs additional signal processing, such as adjusting the gain level for each object.

[0059] In some embodiments, system 800 (e.g., via module 820) can assign a significant distance relative to the location of a virtual listener to each virtual sound source, from which the rendering engine can derive the per-source direct gain and per-source room gain for each object. The direct gain and room gain may affect the audio signal power contributed by the virtual sound sources to the main mixing bus 840 and the room transmit bus 830, respectively. A minimum distance parameter can be assigned to each virtual sound source, and as the distance increases beyond this minimum distance, the direct gain and room gain can roll off at different rates.

[0060] In some examples, Figure 8 System 800 can be used for audio recording and interactive audio applications for traditional two-channel front stereo speaker playback systems. However, when applying System 800 in a binaural or immersive 3D audio system to enable the spatial diffusion of simulated reverberation and reflection distribution, System 800 may not provide sufficiently convincing auditory localization cues when rendering virtual sound sources (especially those far from the listener). This can be addressed by including a clustered reflection rendering module shared among virtual sound sources 810, while supporting per-source control of the spatial distribution of reflections. This module is expected to combine an early reflection processing algorithm for each source with dynamic control of early reflection parameters based on the virtual sound source and listener positions.

[0061] In some embodiments, it is desirable to have a spatial audio processing model / system and method that can accurately reproduce location-related room acoustic cues without computationally complex rendering of individual early reflections of each virtual sound source or detailed descriptions of the geometry and physical properties of the acoustic reflectors.

[0062] Reflection processing models can dynamically describe the positions of listeners and virtual sound sources in real or virtual rooms / environments without requiring associated physical and geometric descriptions. Reflection translation for each source cluster and a perceptual model for controlling early reflection processing parameters can be efficiently implemented.

[0063] Figure 9 An audio mixing system 900 for rendering multiple virtual sound sources in a virtual room is illustrated according to some embodiments. For example, system 900 may include a rendering engine for room acoustic simulation of multiple virtual sound sources 910 (e.g., objects 1 to N). Compared to system 800 described above, system 900 may include separate control for reverberation and reflection transmission channels for each virtual sound source. Each object may be input to a corresponding per-source processing module 920, and a room transmission bus 930 may feed to a room processing module 950.

[0064] Figure 10 A per-source processing module 1020 is shown according to some embodiments. Module 1020 may correspond to Figure 9 One or more of the example systems 900 and modules 920 shown. Each source processing module 1020 can perform processing specific to a single source (e.g., 1010, which may correspond to one of the sources 910) within the entire system (e.g., system 900). Each source processing module may include a direct processing path (e.g., 1030A) and / or a room processing path (e.g., 1030B).

[0065] In some embodiments, separate direct filters and room filters can be applied to each sound source. Applying filters individually allows for finer and more precise control over how each source radiates sound to the listener and to the surrounding environment. Unlike broadband gain, using filters allows for matching the desired sound radiation pattern according to frequency. This is advantageous because radiation characteristics can vary across sound source types and can be frequency-dependent. The angle between the main acoustic axis of the sound source and the listener's position can affect the sound pressure level perceived by the listener. Furthermore, source radiation characteristics can affect the average diffuse power of the source.

[0066] In some embodiments, the frequency-correlated filter may be implemented using the dual-shelf method disclosed in U.S. Patent Application No. 62 / 678259 entitled "Index Scheduling for Filter Parameters," the entire contents of which are incorporated herein by reference. In some embodiments, the frequency-correlated filter may be applied in the frequency domain and / or using a finite impulse response filter.

[0067] As shown in the example, the direct processing path may include a direct send filter 1040 followed by a direct translation module 1044. The direct send filter 1040 may model one or more acoustic effects, such as source directionality, distance, and / or orientation. The direct translation module 1044 may spatialize the audio signal to correspond to a tangible location in the environment (e.g., a 3D location in a virtual environment such as an XR environment). The direct translation module 1044 may be amplitude- and / or intensity-based and may depend on the geometry of the speaker array. In some embodiments, the direct processing path may include a direct send gain 1042, along with the direct send filter and the direct translation module. The direct translation module 1044 may output to a main mixing bus 1090, which may correspond to the main mixing bus 940 described above with respect to the example system 900.

[0068] In some embodiments, the room processing path includes a room delay 1050 and a room transmit filter 1052, followed by a reflection path (e.g., 1060A) and a reverberation path (e.g., 1060B). The room transmit filter can be used to model the effect of sound source directivity on signals destined for the reflection and reverberation paths. The reflection path may include a reflection transmit gain 1070 and may transmit the signal to a reflection transmit bus 1074 via a reflection translation module 1072. The reflection translation module 1072 may be similar to the direct translation module 1044 because it can spatialize the audio signal, but it can operate on the reflection rather than the direct signal. The reverberation path 1060B may include a reverberation gain 1080 and may transmit the signal to a reverberation transmit bus 1084. The reflection transmit bus 1074 and the reverberation transmit bus 1084 may be grouped into a room transmit bus 1092, which may correspond to the room transmit bus 930 described above with respect to example system 900.

[0069] Figure 11Examples of per-source reflection translation modules 1100, corresponding to the aforementioned reflection translation module 1072, are shown according to some embodiments. As shown, for example, as described in "A Comparative Study of Three-Dimensional Audio Coding and Rendering Techniques" by J.-M. Jot, V. Larcher, and J.-M. Pernaux at the 16th International Conference on Spatial Sound Reproduction (1999), the input signal can be encoded as a three-channel surround sound B-format signal. The coding coefficients 1110 can be calculated according to Equations 1-3.

[0070]

[0071] gX=k*cos(Az) Equation 2

[0072] gY=k*sin(Az) Equation 3

[0073] In equation 1-3, k can be calculated as Where F is the spatial focus parameter with a value between [0, 2 / 3], and Az is the angle between [0, 360]. The encoder can encode the input signal into a three-channel surround sound B-format signal.

[0074] Az can be an azimuth angle defined by projecting the principal direction of arrival of the reflected signal onto a horizontal plane relative to the head (e.g., a plane perpendicular to the "up" vector of the listener's head and including the listener's ears). The spatial focus parameter F indicates the spatial concentration of the reflected signal energy arriving at the listener. When F is zero, the spatial distribution of the reflected energy arrival may be uniform around the listener. As F increases, the spatial distribution may become increasingly concentrated around the principal direction determined by the azimuth angle Az. The maximum theoretical value of F can be 1.0, indicating that all energy arrives from the principal direction determined by the azimuth angle Az.

[0075] In embodiments of the present invention, the spatial focus parameter F can be defined as the magnitude of the Gerzon energy vector, as described, for example, in “Comparative Study of Three-Dimensional Audio Coding and Rendering Techniques” by J.-M. Jot, V. Larcher, and J.-M. Pernaux at the 16th International Conference on Spatial Sound Reproduction (1999).

[0076] The output of the reflection translation module 1100 can be provided to the reflection transmission bus 1174, which can correspond to the above regarding Figure 10 The reflection transmission bus 1074 and the example processing module 1020 are described.

[0077] Figure 12 An example room processing module 1200 according to some embodiments is shown. The room processing module 1200 may correspond to the above-described... Figure 9The room processing module 950 and example system 900 are described. Figure 9 As shown, the room processing module 1200 may include a reflection processing path 1210A and / or a reverberation processing path 1210B.

[0078] The reflection processing path 1210A can receive signals from the reflection transmission bus 1202 (which may correspond to the reflection transmission bus 1074 described above) and output signals to the main hybrid bus 1290 (which may correspond to the main hybrid bus 940 described above). The reflection processing path 1210A may include a reflection global gain 1220, a reflection global delay 1222, and / or a reflection module 1224, which can simulate / render reflections.

[0079] The reverb processing path 1210B can receive signals from the reverb transmit bus 1204 (which may correspond to the reverb transmit bus 1084 described above) and output signals to the main mixing bus 1290. The reverb processing path 1210B may include a reverb global gain 1230, a reverb global delay 1232 and / or a reverb module 1234.

[0080] Figure 13 An example reflection module 1300 according to some embodiments is shown. As described above, the input 1310 of the reflection module can be output by the reflection translation module 1100 and presented to the reflection module 1300 via the reflection transmission bus 1174. The reflection transmission bus can carry a 3-channel surround sound B-format signal, which combines signals from all virtual sound sources (e.g., those mentioned above). Figure 9 The contribution of the described sound source 910 (objects 1 to N) is shown. In the example shown, three channels (represented by (W, X, Y)) are fed to the surround sound decoder 1320. According to this example, the surround sound decoder produces six output signals, which are fed to six single-input / output basic reflection modules 1330 (R1 to R6), respectively, producing a set of six reflected output signals 1340 (s1 to s6) (although this example shows six signals and reflection modules, any suitable number can be used). The reflected output signals 1340 are presented to the main mixing bus 1350, which may correspond to the main mixing bus 940 described above.

[0081] Figure 14 The illustration shows a spatial distribution 1400 of apparent arrival directions of reflections detected by listener 1402 according to some embodiments. For example, the reflections shown could be reflections generated by the aforementioned reflection module 1300, for example, for those assigned with the above-mentioned... Figure 11 The sound source is described by specific values ​​of the reflection translation parameters Az and F.

[0082] like Figure 14As shown, the effect of combining the reflection module 1300 with the reflection translation module 1100 is to generate a series of reflections, each of which can arrive at different times (e.g., as shown in model 600) and originate from each virtual speaker direction 1410 (e.g., 1411 to 1416, which may correspond to the aforementioned reflection output signals s1 to s6). The effect of combining the reflection translation module 1100 with the surround sound decoder 1320 is to adjust the relative amplitude of the reflection output signal 1340 to create for the listener the sensation of reflections emanating from the principal direction angle Az, the spatial distribution of which is determined by the setting of the spatial focus parameter F (e.g., more or less concentrated around that principal direction).

[0083] In some embodiments, for each source, the principal direction angle of reflection Az coincides with the apparent direction of arrival of the direct path, which can be controlled by the direct translation module 1020 for each source. The simulated reflection can emphasize the perceived directional location of the virtual sound source by the listener.

[0084] In some embodiments, the main hybrid bus 940 and the direct translation module 1020 can achieve three-dimensional reproduction of the sound direction. In these embodiments, the principal reflection direction angle Az can coincide with the projection of the apparent direction onto the plane on which the principal reflection angle Az is measured.

[0085] Figure 15 Model 1500 illustrates example direct gain, reflection gain, and reverberation gain as functions of distance (e.g., to the listener) according to some embodiments. Model 1500 illustrates, for example... Figure 10 The diagram illustrates examples of variations in direct, reflected, and reverberant send gain relative to source distance. As shown, direct sound, its reflections, and its reverberation can have significantly different attenuation curves with respect to distance. In some cases, per-source processing as described above can allow for a faster distance-based roll-off for reflections than for reverberation. Psychologically, this can enable robust direction and distance perception, especially for distant sources.

[0086] Figure 16 An example model 1600 is shown, illustrating the spatial focus and source distance for both direct and reflected components according to some embodiments. In this example, the direct translation module 1020 is configured to generate maximum spatial concentration of the direct path component in the direction of the sound source, regardless of its distance. On the other hand, for all distances greater than a limiting distance (e.g., the minimum reflection distance 1610), the reflected spatial focus parameter F can be set to an example value of 2 / 3, thereby realistically enhancing direction perception. As shown in example model 1600, the reflected spatial focus parameter value decreases toward zero as the source approaches the listener.

[0087] Figure 17An example model 1700 is shown, illustrating the amplitude of an audio signal as a function of time. As described above, the reflection processing path (e.g., 1210A) can receive a signal from the reflection transmit bus and output the signal to the main mixing bus. As described above, the reflection processing path may include a reflection global gain (e.g., 1220), a reflection global delay (e.g., 1222) for controlling the parameter Der shown in model 1700, and / or a reflection module (e.g., 1224).

[0088] As described above, a reverb processing path (e.g., 1210B) can receive signals from the reverb transmit bus and output signals to the main mixing bus. The reverb processing path 1210B may include a reverb global gain (e.g., 1230) for controlling parameter Lgo, as shown in Model 1700, a reverb global delay (e.g., 1232) for controlling parameter Drev, as shown in Model 1700, and / or a reverb module (e.g., 1234). The processing modules within the reverb processing path can be implemented in any suitable order. Examples of reverb modules are described in U.S. Patent Application No. 62 / 685235 entitled "Reverb Gain Normalization" and U.S. Patent Application No. 62 / 684086 entitled "Low-Frequency Interchannel Coherence Control," the entire contents of which are incorporated herein by reference.

[0089] Figure 17 Model 1700 illustrates, according to some embodiments, how parameters for each source (including distance and reverberation delay) can be taken into account to dynamically adjust reverberation delay and level. In the figure, Dtof represents the delay due to the flight time of a given object: Dtof = ObjDist / c, where ObjDist is the object distance from the center of the listener's head, and c is the speed of sound in air. Drm represents the per-object room delay. Dobj represents the total per-object delay: Dobj = Dtof + Drm. Der represents the global early reflection delay. Drev represents the global reverberation delay. Dtotal represents the total delay for a given object: Dtotal = Dobj + Dglobal.

[0090] Lref represents the reverberation level when Dtotal = 0. Lgo represents the global level offset due to global latency, which can be calculated according to Equation 10, where T60 is the reverberation time of the reverberation algorithm. Loo represents the per-object level offset due to global latency, which can be calculated according to Equation 11. Lto represents the total level offset for a given object, and can be calculated according to Equation 12 (assuming dB values).

[0091] Lgo = Dglobal / T60 * 60 (dB) Equation 10

[0092] Loo = Dobj / T60*60 (dB) Equation 11

[0093] Lto = Lgo + Loo Equation 12

[0094] In some embodiments, the reverberation level is calibrated independently of object location, reverberation time, and other user-controllable parameters. Therefore, Lrev can be an extrapolated level of decaying reverberation at the initial time of sound emission. Lrev can have the same quantity as the initial reverberation power (RIP) defined in U.S. Patent Application No. 62 / 685235 entitled “Reverberation Gain Normalization,” the entire contents of which are incorporated herein by reference. Lrev can be calculated according to Equation 13.

[0095] Lrev = Lref + Lto Equation 13

[0096] In some embodiments, T60 can be a function of frequency. Therefore, Lgo, Loo, and thus Lto are frequency-dependent.

[0097] Figure 18 An example system 1800 for determining spatial audio characteristics based on the acoustic environment is shown. The example system 1800 can be used to determine spatial audio characteristics such as reflections and / or reverberation as described above. As an example, these characteristics may include the volume of the room, reverberation time as a function of frequency, the listener's position relative to the room, the presence of objects in the room (e.g., objects that attenuate sound), surface materials, or other suitable characteristics. In some examples, these spatial audio characteristics can be obtained locally by acquiring single impulse responses using microphones and speakers freely placed in the local environment, or adaptively derived by continuously monitoring and analyzing sound acquired by a mobile device microphone. In some examples, where the acoustic environment can be sensed via sensors of an XR system (e.g., an augmented reality system including one or more of the wearable head unit 100, handheld controller 200, and auxiliary unit 300 described above), the user's position can be used to present audio reflections and reverberation corresponding to the environment presented to the user (e.g., via a display).

[0098] In example system 1800, as described above, the acoustic environment sensing module 1810 identifies the spatial audio characteristics of the acoustic environment. In some examples, the acoustic environment sensing module 1810 may acquire data corresponding to the acoustic environment (stage 1812). For example, the data acquired at stage 1812 may include audio data from one or more microphones, camera data from a camera (such as an RGB camera or a depth camera), LiDAR data, sonar data, radar data, GPS data, or other suitable data that may convey information about the acoustic environment. In some cases, the data acquired at stage 1812 may include user-related data, such as the user's location or orientation relative to the acoustic environment. The data acquired at stage 1812 may be acquired via one or more sensors of a wearable device (such as the wearable head unit 100 described above).

[0099] In some embodiments, the local environment in which the head-mounted display device is located may include one or more microphones. In some embodiments, one or more microphones may be employed, and may be microphones mounted on a mobile device, microphones located in the environment, or both. The benefits of this arrangement may include collecting directional information about the room's reverberation, or mitigating poor signal quality from any of the one or more microphones. For example, the signal quality at a given microphone may be poor due to occlusion, overload, wind noise, transducer damage, etc.

[0100] At stage 1814 of module 1810, features can be extracted from the data acquired at stage 1812. For example, the dimensions of a room can be determined from sensor data (such as camera data, LiDAR data, sonar data, etc.). The features extracted at stage 1814 can be used to determine one or more acoustic characteristics of the room (e.g., frequency-dependent reverberation time), and these characteristics can be stored at stage 1816 and associated with the current acoustic environment.

[0101] In some examples, module 1810 can communicate with database 1840 to store and retrieve acoustic characteristics of the acoustic environment. In some embodiments, the database may be stored locally in the device's memory. In some embodiments, the database may be stored online as a cloud-based service. The database may assign geographic locations to room characteristics based on the listener's location for easy access later. In some embodiments, the database may contain additional information to identify the listener's location and / or identify reverberation characteristics in the database that closely resemble the listener's environmental characteristics. For example, room characteristics may be categorized by room type, so once it is determined that the listener is in a room of a known type (e.g., a bedroom or living room), a set of parameters can be used even if the absolute geographic location may not be known.

[0102] Storing reverberation characteristics in a database may relate to U.S. Patent Application No. 62 / 573448 entitled “Persistent World Model Supporting Augmented Reality and Incorporating Audio Components,” the entire contents of which are incorporated herein by reference.

[0103] In some examples, system 1800 may include a reflection adaptation module 1820 for acquiring the acoustic characteristics of a room and applying these characteristics to audio reflections (e.g., audio reflections presented to a user of wearable head unit 100 via headphones or speakers). At stage 1822, the user's current acoustic environment can be determined. For example, GPS data can indicate the user's location in GPS coordinates, which in turn can indicate the user's current acoustic environment (e.g., the room located at those GPS coordinates). As another example, camera data combined with optical recognition software can be used to identify the user's current environment. The reflection adaptation module 1820 can then communicate with database 1840 to acquire acoustic characteristics associated with the determined environment, and these acoustic characteristics can be used to update the audio rendering at stage 1824. That is, acoustic characteristics related to reflection (e.g., directional patterns or roll-off curves as described above) can be applied to the reflected audio signal presented to the user such that the presented reflected audio signal incorporates these acoustic characteristics.

[0104] Similarly, in some examples, system 1800 may include a reflection adaptation module 1830 for acquiring the acoustic characteristics of the room and applying these characteristics to audio reverberation (e.g., audio reflections presented to the user of the wearable head unit 100 via headphones or speakers). As described above (e.g., in relation to...) Figure 7(Table 700) The acoustic characteristics of interest for reverberation may differ from those of interest for reflection. As described above, at stage 1832, the user's current acoustic environment can be determined. For example, GPS data can indicate the user's location in GPS coordinates, which in turn can indicate the user's current acoustic environment (e.g., the room located at those GPS coordinates). As another example, camera data combined with optical recognition software can be used to identify the user's current environment. The reverberation adaptation module 1830 can then communicate with the database 1840 to obtain acoustic characteristics associated with the determined environment, and these acoustic characteristics can be used accordingly to update the audio rendering at stage 1824. That is, reverberation-related acoustic characteristics (e.g., reverberation decay time as described above) can be applied to the reverberated audio signal presented to the user, such that the presented reverberated audio signal incorporates these acoustic characteristics.

[0105] Regarding the systems and methods described above, the elements of these systems and methods may be suitably implemented by one or more computer processors (e.g., CPUs or DSPs). This disclosure is not limited to any particular configuration of the computer hardware (including computer processors) used to implement these elements. In some cases, multiple computer systems may be employed to implement the systems and methods described above. For example, a first computer processor (e.g., a processor of a wearable device coupled to a microphone) may be employed to receive input microphone signals and perform initial processing on these signals (e.g., signal conditioning and / or segmentation, as described above). A second (and perhaps more powerful) processor may then be employed to perform more computationally intensive processing, such as determining probability values ​​associated with speech segments of these signals. Another computer device (such as a cloud server) may host a speech recognition engine, ultimately providing the input signals to that speech recognition engine. Other suitable configurations will be apparent and are within the scope of this disclosure.

[0106] Although the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications should be understood to be included within the scope of the disclosed examples as defined by the appended claims.

Claims

1. A method for spatial audio rendering, comprising: Based on the location of the sound source in the virtual environment, a first intermediate audio signal corresponding to the input audio signal is determined, wherein the first intermediate audio signal is the input audio signal after being processed by a direct processing path including a direct transmission filter, a direct transmission gain and a direct translation module; Associate the first intermediate audio signal with the first bus; Based on the location of the sound source and further based on the acoustic characteristics of the virtual environment, a second intermediate audio signal is determined, the second intermediate audio signal corresponding to the reflection of the input audio signal from the surface of the virtual environment; The determination of the second intermediate audio signal includes encoding the input audio signal to be processed into a three-channel surround sound B-format signal, wherein the input audio signal to be processed is the input audio signal after being processed by a room processing path including room delay, room send filtering, and reflection send gain, and The second intermediate audio signal is associated with three channels; Associate the second intermediate audio signal with a second bus, wherein the second bus is associated with the three channels; and The output audio signal is presented to the listener via the first bus and the second bus. The output audio signal is determined by decoding the second intermediate audio signal. The decoded second intermediate audio signal is associated with six channels, and The first bus is associated with the six channels.

2. The method according to claim 1, wherein, The acoustic characteristics of the virtual environment are determined via one or more sensors associated with the listener.

3. The method according to claim 2, wherein, The one or more sensors include one or more microphones.

4. The method according to claim 2, wherein: The one or more sensors are associated with a wearable head-mounted device configured to be worn by the listener, and The output audio signal is presented to the listener via one or more speakers associated with the wearable head device.

5. The method of claim 1, further comprising displaying a view of the virtual environment to the listener simultaneously with the presentation of the output audio signal.

6. The method of claim 1, further comprising obtaining the acoustic characteristics from a database, wherein, The acoustic properties include acoustic properties determined by one or more sensors.

7. The method according to claim 6, wherein, Obtaining the aforementioned acoustic characteristics includes: Based on the output of the one or more sensors, the location of the listener is determined; and The acoustic characteristics are identified based on the listener's location.

8. A wearable device, comprising: A monitor configured to display a view of a virtual environment; One or more sensors; One or more speakers; One or more processors are configured to perform a method comprising: Based on the location of the sound source in the virtual environment, a first intermediate audio signal corresponding to the input audio signal is determined, wherein the first intermediate audio signal is the input audio signal after being processed by a direct processing path including a direct transmission filter, a direct transmission gain and a direct translation module; Associate the first intermediate audio signal with the first bus; Based on the location of the sound source and further based on the acoustic characteristics of the virtual environment, a second intermediate audio signal is determined, the second intermediate audio signal corresponding to the reflection of the input audio signal from the surface of the virtual environment; The determination of the second intermediate audio signal includes encoding the input audio signal to be processed into a three-channel surround sound B-format signal, wherein the input audio signal to be processed is the input audio signal after being processed by a room processing path including room delay, room send filtering, and reflection send gain, and The second intermediate audio signal is associated with three channels; Associate the second intermediate audio signal with a second bus, wherein the second bus is associated with the three channels; and The output audio signal is presented to the listener via the one or more speakers, the first bus, and the second bus, wherein... The output audio signal is determined by decoding the second intermediate audio signal. The decoded second intermediate audio signal is associated with six channels, and The first bus is associated with the six channels.

9. The wearable device according to claim 8, wherein, The acoustic characteristics of the virtual environment are determined by one or more sensors associated with the listener.

10. A non-transitory computer-readable storage medium storing one or more programs, said programs including instructions that, when executed by an electronic device having one or more processors and memory, cause the device to perform a method, said method comprising: Based on the location of the sound source in the virtual environment, a first intermediate audio signal corresponding to the input audio signal is determined, wherein the first intermediate audio signal is the input audio signal after being processed by a direct processing path including a direct transmission filter, a direct transmission gain and a direct translation module; Associate the first intermediate audio signal with the first bus; Based on the location of the sound source and further based on the acoustic characteristics of the virtual environment, a second intermediate audio signal is determined, the second intermediate audio signal corresponding to the reflection of the input audio signal from the surface of the virtual environment; The determination of the second intermediate audio signal includes encoding the input audio signal to be processed into a three-channel surround sound B-format signal, wherein the input audio signal to be processed is the input audio signal after being processed by a room processing path including room delay, room send filtering, and reflection send gain, and The second intermediate audio signal is associated with three channels; Associate the second intermediate audio signal with a second bus, wherein the second bus is associated with the three channels; and The output audio signal is presented to the listener via the first bus and the second bus, wherein, The output audio signal is determined by decoding the second intermediate audio signal. The decoded second intermediate audio signal is associated with six channels, and The first bus is associated with the six channels.

Citation Information

Patent Citations

  • System and method for high-precision 3-dimensional audio for augmented reality

    US20120093320A1

  • Method and apparatus for the simulation of complex audio environments

    US7099482B1