Efficient rendering of virtual sound field

The modified virtual speaker panning method dynamically selects active speakers based on proximity and energy levels, optimizing FIR filter usage to enhance the efficiency of spatial audio rendering in virtual environments.

JP2025131850APending Publication Date: 2025-09-09MAGIC LEAP INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025101212
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-06-12
Filing Date
2025-06-17
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing virtual speaker-based spatial audio systems are inefficient in rendering virtual environments, particularly when dealing with a large number of sound sources, leading to high computational requirements and resource consumption.

Method used

A method and system that uses modified virtual speaker panning, dynamically selecting a subset of fixed virtual speakers based on proximity and energy levels to reduce processing blocks, thereby optimizing FIR filter usage.

Benefits of technology

Reduces computational complexity and resource consumption by selectively bypassing processing blocks, enhancing the efficiency of spatial audio rendering in virtual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131850000001_ABST
    Figure 2025131850000001_ABST
Patent Text Reader

Abstract

To provide a system and method for increasing the efficiency of virtual speaker-based spatial audio systems.SOLUTION: A method of spatially rendering audio signals includes: determining a model of a virtual environment; determining a spatial configuration of the virtual environment; determining a plurality of signals associated with the spatial configuration and the model; determining whether energy levels of the signals exceed a predetermined threshold; decoding one or more signals according to a determination that the energy level of one or more signals exceeds the predetermined threshold; and rendering an audio signal based on the one or more decoded signals.SELECTED DRAWING: Figure 5A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 684,093, filed June 12, 2018, which is incorporated herein by reference in its entirety. (Technical field)

[0002] This disclosure relates generally to spatial audio rendering and associated systems, and more particularly to systems and methods for increasing the efficiency of virtual speaker-based spatial audio systems. [Background technology]

[0003] Virtual environments are ubiquitous in computing environments, finding use in video games (where a virtual environment may represent a game world), maps (where a virtual environment may represent a terrain to be navigated), simulations (where a virtual environment may simulate a real environment), digital storytelling (where virtual characters may interact with one another within a virtual environment), and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, a user's experience of a virtual environment may be limited by the technology for presenting the virtual environment. For example, traditional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may not be capable of realizing a virtual environment in a way that creates a compelling, realistic, and immersive experience.

[0004] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present users of XR systems with sensory information corresponding to a virtual environment represented by data in a computer system. Such systems can provide a uniquely enhanced sense of immersion and presence by combining virtual visual and audio cues with real sights and sounds. Therefore, it may be desirable to present digital sounds to users of XR systems so that the sounds appear to occur naturally in the user's real environment and consistent with the sounds the user expects. Generally speaking, users expect virtual sounds to take on the acoustic characteristics of the real environment in which they are heard. For example, a user of an XR system in a large concert hall would expect the XR system's virtual sounds to have a sound quality similar to a large cavern, while a user in a small apartment would conversely expect the sounds to be more attenuated, close, and direct. In addition, users expect the virtual sounds to be presented without delay.

[0005] Ambisonics and non-ambisonics, among other techniques, may be used to generate spatial audio. With a large number of source objects, Ambisonics and non-ambisonics, due to their design and architecture, may be an efficient way to render spatial audio. This may be especially true when reflections are modeled. Ambisonics and non-ambisonics multi-channel-based spatial audio systems may render audio signals through several steps. Exemplary steps may include a per-source encoding step, a fixed overhead sound field decoding step, and / or a fixed speaker virtualization step. One or more hardware components may perform the steps.

[0006] In a first method for rendering audio signals, each sound source can have its own pair of finite impulse response (FIR) filters. In such a system, the perceived location of a sound is changed by changing the filter coefficients of the FIR filters. In some embodiments, each sound can use multiple (e.g., two pairs) FIR filters. Each pair can use two filters (i.e., four FIR filters). As the sound moves around the virtual environment, the FIR filters can be cross-faded. In some embodiments, four FIR filters can be used for each sound.

[0007] In a second method for rendering an audio signal, virtual speaker panning may be implemented using a fixed number of virtual speakers. Each sound source may be panned across a fixed number of virtual speakers. In some embodiments, multiple (e.g., two) FIR filters may be used for each virtual speaker. Virtual speaker panning may be efficient for certain applications and may use very little computing resources.

[0008] In some embodiments, one method may have higher efficiency than another depending on the number of sounds played simultaneously. For example, 30 sounds may be played simultaneously. If four FIR filters are used for each sound source, 120 FIR filters (30 sound sources × 4 FIR filters per sound source = 120 FIR filters) may be required for the first method. If two FIR filters are used for each virtual speaker, only 32 FIR filters may be required for the second method (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters).

[0009] As another example, only one sound may be played: A first method may require only four FIR filters (one sound source x four FIR filters per sound source = four FIR filters), while a second method may require 32 FIR filters (16 virtual speakers x two FIR filters per virtual speaker = 32 FIR filters).

[0010] As illustrated through the above examples, the first method may be beneficial for a small number of sounds, and the second method may be beneficial for a large number of sounds. Thus, audio systems and methods that increase efficiency based on the number of sound sources at a given time may be desirable. Summary of the Invention [Means for solving the problem]

[0011] An audio system and method for rendering an audio signal is disclosed, where the system uses modified virtual speaker panning. The audio system may include a fixed number F of virtual speakers, and the modified virtual speaker panning may dynamically select and use a subset P of the fixed virtual speakers. Each sound source may be panned across the subset P of virtual speakers. In some embodiments, multiple (e.g., two) FIR filters may be used for each virtual speaker in the subset P. The subset P of virtual speakers may be selected based on one or more factors, such as proximity to the sound source. The subset P of virtual speakers may be referred to as active speakers.

[0012] The modified virtual speaker panning method can be compared with the first and second methods disclosed above as an example. If three sounds are played simultaneously and the audio system has 16 fixed virtual speakers, the first method may require 12 FIR filters (3 sound sources × 4 FIR filters per sound source = 12 FIR filters), and the second method may require 32 FIR filters (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters). On the other hand, the modified virtual speaker panning method may dynamically select three virtual speakers to be active virtual speakers as part of subset P. The modified virtual speaker panning method may require six FIR filters, i.e., two FIR filters for each active virtual speaker (3 virtual speakers × 2 FIR filters = 6 FIR filters). The present invention provides, for example, the following. (Item 1) 1. A method for spatially rendering an audio signal, said method comprising: Using a spatial modeler to model a virtual environment; Distributing signals from the spatial modeler across multiple virtual speakers using a spatial encoder; representing a spatial configuration of the virtual environment using an internal spatial representation; decoding the signal from said internal spatial representation using a decoder / virtualizer; introducing virtual sounds into the decoded signal using a decoder / virtualizer; Selectively bypassing one or more processing blocks associated with inactive virtual speakers within the decoder / virtualizer; combining the signals from said decoder / virtualizer; outputting the combined signal as the audio signal; A method comprising: (Item 2) determining an energy level associated with said signal from a sound field decoder; determining whether each of the detected energy levels is less than an energy threshold; further comprising the selective bypassing of the one or more processing blocks includes, in accordance with a determination that the detected energy level of at least one of the virtual speakers is less than the energy threshold, bypassing head-related transfer function (HRTF) processing of the corresponding signal from the sound field decoder; Item 1. The method of item 1, wherein the sound field decoder is included within the decoder / virtualizer. (Item 3) 3. The method of claim 2, further comprising: performing HRTF processing of the corresponding signal from the sound field decoder according to a determination that the detected energy level of at least one of the virtual speakers is not less than the energy threshold. (Item 4) determining whether the number of sound sources is greater than or equal to a predetermined sound source threshold; the selective bypassing of the one or more processing blocks includes bypassing a plurality of detectors and passing signals from a sound field decoder directly to a plurality of HRTF blocks when the number of sound sources is equal to or greater than the predetermined sound source threshold; Item 10. The method of item 1, wherein the plurality of detectors and the plurality of HRTF blocks are included within the decoder / virtualizer. (Item 5) Item 5. The method of item 4, further comprising, in response to a determination that the number of sound sources is not greater than or equal to the predetermined sound source threshold, passing signals from the sound field decoder directly to the plurality of detectors. (Item 6) determining the location of each sound source; determining which of the plurality of virtual speakers is located proximate to the respective sound source; Item 1, the method of claim 1 further comprising: (Item 7) 7. The method of claim 6, wherein the determination of which of the plurality of virtual speakers is located proximate to the respective sound source is performed for every video frame. (Item 8) 7. The method of claim 6, wherein the selective bypassing of the one or more processing blocks within the decoder / virtualizer includes bypassing all of the one or more processing blocks associated with at least one speaker within the decoder / virtualizer that is not located proximate to the respective sound source. (Item 9) - introducing a representation of a translation associated with said audio signal using a rotation / translation representation; determining whether the amplitude of a signal from said rotation / translation representation is greater than or equal to a predetermined amplitude threshold; further comprising the selective bypassing of the one or more processing blocks within the decoder / virtualizer includes bypassing a sound field decoder and a plurality of HRTF blocks when the amplitude of the signal from the rotated / translated representation is not greater than or equal to the predetermined amplitude threshold; Item 1. The method according to item 1, wherein the sound field decoder and the plurality of HRTF blocks are included within the decoder / virtualizer. (Item 10) upon determining that the amplitude of the signal from the rotated / translated representation is greater than or equal to the predetermined amplitude threshold; decoding a signal from said rotation / translation representation; determining a head-related transfer function (HRTF) and applying it to the decoded signal; Item 10. The method of item 9, further comprising: (Item 11) Item 10. The method of claim 1, wherein the plurality of virtual speakers includes the inactive virtual speaker and an active virtual speaker at a first time, and at least one of the active virtual speakers at the first time is designated as inactive at a second time while a signal is being processed. (Item 12) 1. A system comprising: a wearable head device configured to provide an audio signal to a user; a circuit configured to spatially render the audio signal; and Equipped with The circuit comprises: a spatial modeler configured to model the virtual environment; a spatial encoder configured to distribute signals from the spatial modeler across a plurality of virtual speakers; an internal spatial representation configured to represent the spatial configuration of the virtual environment; a decoder / virtualizer configured to decode a signal from the internal spatial representation and to introduce virtual sounds into the decoded signal; Including, The decoder / virtualizer a rotation / translation representation configured to introduce a representation of a movement associated with said audio signal; a sound field decoder configurable to decode signals from said rotation / translation representation; a plurality of head-related transfer function (HRTF) blocks, each HRTF block configured to determine an HRTF corresponding to an input signal thereof and to apply the corresponding HRTF to the input signal; a plurality of combiners configured to combine signals from the plurality of HRTF blocks and output the audio signal; Including, the system. (Item 13) a plurality of detectors configured to receive signals from the sound field decoder and determine energy levels associated with the signals from the sound field decoder; a plurality of first switches configured to pass the signal from the sound field decoder to the plurality of HRTF blocks when the determined energy level is not less than an energy threshold; Item 13. The system of item 12, further comprising: (Item 14) The device further includes a second switch, the second switch comprising: receiving the signal from the sound field decoder; selectively passing the signal from the sound field decoder directly through the plurality of detectors or through the plurality of HRTF blocks; Item 14. The system of item 13, configured to perform the following: (Item 15) and a sound field decoding decision, the sound field decoding decision comprising: determining whether the amplitude of a signal from the rotation / translation representation is greater than a predetermined amplitude threshold; passing the signal from the rotated / translated representation to the sound field decoder in accordance with a determination that the amplitude of the signal from the rotated / translated representation is greater than the predetermined amplitude threshold; Item 13. The system according to item 12, configured to perform the following: [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 illustrates an exemplary wearable system, according to some embodiments.

[0014] [Figure 2] FIG. 2 illustrates an example handheld controller that may be used with an example wearable system, according to some embodiments.

[0015] [Figure 3] FIG. 3 illustrates an example auxiliary unit that may be used in conjunction with an example wearable system, according to some embodiments.

[0016] [Figure 4] FIG. 4 illustrates an example functional block diagram for an example wearable system, according to some embodiments.

[0017] [Figure 5A]FIG. 5A illustrates a block diagram of an exemplary spatial audio system, according to some embodiments.

[0018] [Figure 5B] FIG. 5B illustrates a flow of an exemplary method for operating the system of FIG. 5A, according to some embodiments.

[0019] [Figure 5C] FIG. 5C illustrates a flow of an example method for operating an example decoder / virtualizer according to some embodiments.

[0020] [Figure 6] FIG. 6 illustrates an example configuration of sound sources and speakers, according to some embodiments.

[0021] [Figure 7A] FIG. 7A illustrates a block diagram of an exemplary decoder / virtualizer including multiple detectors, according to some embodiments.

[0022] [Figure 7B] FIG. 7B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 7A according to some embodiments.

[0023] [Figure 8A] FIG. 8A illustrates a block diagram of an exemplary decoder / virtualizer, according to some embodiments.

[0024] [Figure 8B] FIG. 8B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 8A according to some embodiments.

[0025] [Figure 9] FIG. 9 illustrates an example configuration of sound sources and speakers, according to some embodiments.

[0026] [Figure 10A] FIG. 10A illustrates a block diagram of an exemplary decoder / virtualizer for use in a system including active speakers, according to some embodiments.

[0027] [Figure 10B] FIG. 10B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 10A according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0028] In the following description of examples, reference is made to the accompanying drawings, which form a part hereof, and in which is shown, by way of illustration, specific examples which may be practiced. It is to be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.

[0029] (Exemplary Wearable System)

[0030] 1 illustrates an exemplary wearable head device 100 configured to be worn on a user's head. Wearable head device 100 may be part of a broader wearable system that includes one or more components, such as a head device (e.g., wearable head device 100), a handheld controller (e.g., handheld controller 200 described below), and / or an auxiliary unit (e.g., auxiliary unit 300 described below). In some examples, wearable head device 100 can be used for virtual reality, augmented reality, or mixed reality systems or applications. The wearable head device 100 includes one or more displays, such as displays 110A and 110B (which may comprise left and right transmissive displays and associated components for coupling light from the displays to the user's eyes, such as orthogonal pupil expansion (OPE) grating sets 112A / 112B and exit pupil expansion (EPE) grating sets 114A / 114B), left and right acoustic structures, such as speakers 120A and 120B (which may be mounted on temple arms 122A and 122B, respectively, and positioned adjacent the user's left and right ears), and an infrared sensor. The wearable head device 100 may include one or more sensors, such as a microphone, an accelerometer, a GPS unit, an inertial measurement unit (IMU) (e.g., IMU 126), an acoustic sensor (e.g., microphone 150), a quadrature coil electromagnetic receiver (e.g., receiver 127 shown mounted on left temple arm 122A), left and right cameras pointed away from the user (e.g., depth (time-of-flight) cameras 130A and 130B), and left and right eye cameras pointed towards the user (e.g., to detect the user's eye movements) (e.g., eye cameras 128 and 128B). However, the wearable head device 100 may incorporate any suitable display technology and any suitable number, type, or combination of sensors or other components without departing from the scope of the invention.In some examples, wearable head device 100 may include one or more microphones 150 configured to detect audio signals generated by the user's voice, and such microphones may be positioned within the wearable head device adjacent to the user's mouth. In some examples, wearable head device 100 may incorporate networking features (e.g., Wi-Fi capabilities) for communicating with other devices and systems, including other wearable systems. Wearable head device 100 may further include components such as a battery, a processor, memory, a storage unit, or various input devices (e.g., buttons, touchpad), or may be coupled to a handheld controller (e.g., handheld controller 200) or auxiliary unit (e.g., auxiliary unit 300) equipped with one or more such components. In some examples, sensors may be configured to output a set of coordinates of the head-mounted unit relative to the user's environment and provide input to a processor to implement a simultaneous localization and mapping (SLAM) procedure and / or a visual odometry algorithm. In some examples, the wearable head device 100 may be coupled to a handheld controller 200 and / or an auxiliary unit 300, as further described below.

[0031] 2 illustrates an exemplary mobile handheld controller component 200 of an exemplary wearable system. In some examples, handheld controller 200 may communicate wired or wirelessly with wearable head device 100 and / or auxiliary unit 300, described below. In some examples, handheld controller 200 includes a handle portion 220 to be held by a user and one or more buttons 240 disposed along a top surface 210. In some examples, handheld controller 200 may be configured for use as an optical tracking target; for example, a sensor (e.g., a camera or other optical sensor) of wearable head device 100 can be configured to detect the position and / or orientation of handheld controller 200, which in turn may indicate the position and / or orientation of a user's hand holding handheld controller 200. In some examples, handheld controller 200 may include a processor, memory, a storage unit, a display, or one or more input devices, such as those described above. In some examples, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 100). In some examples, the sensors can detect the position or orientation of the handheld controller 200 relative to the wearable head device 100 or relative to another component of the wearable system. In some examples, the sensors may be positioned within the handle portion 220 of the handheld controller 200 and / or may be mechanically coupled to the handheld controller. The handheld controller 200 can be configured to provide one or more output signals corresponding, for example, to a pressed state of the button 240, or the position, orientation, and / or movement of the handheld controller 200 (e.g., via an IMU). Such output signals may be used as input to a processor of the wearable head device 100, an input to the auxiliary unit 300, or an input to another component of the wearable system.In some examples, the handheld controller 200 may include one or more microphones to detect sounds (e.g., a user's speech, environmental sounds) and, in some cases, provide signals corresponding to the detected sounds to a processor (e.g., a processor of the wearable head device 100).

[0032] 3 illustrates an exemplary auxiliary unit 300 of an exemplary wearable system. In some examples, the auxiliary unit 300 may communicate wired or wirelessly with the wearable head device 100 and / or the handheld controller 200. The auxiliary unit 300 may include a battery to provide energy for operating one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including a display, sensors, an acoustic structure, a processor, a microphone, and / or other components of the wearable head device 100 or the handheld controller 200). In some examples, the auxiliary unit 300 may include a processor, memory, a storage unit, a display, one or more input devices, and / or one or more sensors, such as those described above. In some examples, the auxiliary unit 300 includes a clip 310 for attaching the auxiliary unit to a user (e.g., a belt worn by the user). An advantage of using auxiliary unit 300 to store one or more components of a wearable system is that doing so may allow large or heavy components to be carried on the user's waist, chest, or back, which are relatively well suited to supporting large, heavy objects, rather than being mounted on the user's head (e.g., when stored in wearable head device 100) or carried by the user's hand (e.g., when stored in handheld controller 200). This may be particularly advantageous with respect to relatively heavy or bulky components, such as batteries.

[0033] 4 shows an example functional block diagram that may correspond to an example wearable system 400, such as may include the example wearable head device 100, handheld controller 200, and auxiliary unit 300 described above. In some examples, the wearable system 400 may be used for virtual reality, augmented reality, or mixed reality applications. As shown in FIG. 4 , the wearable system 400 may include an example handheld controller 400B, referred to herein as a “totem” (and which may correspond to the handheld controller 200 described above), which may include a totem / headgear six-degree-of-freedom (6DOF) totem subsystem 404A. The wearable system 400 may also include an example wearable head device 400A (which may correspond to the wearable headgear device 100 described above), which includes a totem / headgear 6DOF headgear subsystem 404B. In an example, the 6DOF totem subsystem 404A and the 6DOF headgear subsystem 404B cooperate to determine six coordinates (e.g., offsets in three translational directions and rotations along three axes) of the handheld controller 400B relative to the wearable head device 400A. The six degrees of freedom may be expressed relative to the coordinate system of the wearable head device 400A. The three translational offsets may be expressed as X, Y, and Z offsets within such a coordinate system, a translation matrix, or some other representation. The rotational degrees of freedom may be expressed as a sequence of yaw, pitch, and roll rotations, a vector, a rotation matrix, a quaternion, or some other representation. In some examples, one or more depth cameras 444 (and / or one or more non-depth cameras) and / or one or more optical targets (e.g., buttons 240 of the handheld controller 200 as described above or dedicated optical targets included in the handheld controller) included within the wearable head device 400A can be used for 6DOF tracking.In some examples, the handheld controller 400B can include a camera as described above, and the headgear 400A can include optical targets for optical tracking in conjunction with the camera. In some examples, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids used to wirelessly transmit and receive three distinguishable signals. By measuring the relative magnitudes of the three distinguishable signals received at each of the coils used for receiving, the 6DOF of the handheld controller 400B relative to the wearable head device 400A can be determined. In some examples, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU), which is useful for providing improved accuracy and / or more timely information regarding high-speed movement of the handheld controller 400B.

[0034] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space that is fixed relative to the wearable head device 400A) to an inertial coordinate space or to an environmental coordinate space. For example, such a transformation may be necessary for the display of the wearable head device 400A to present virtual objects in an expected position and orientation relative to the real environment (e.g., a virtual person sitting in a real chair facing forward, regardless of the position and orientation of the wearable head device 400A), rather than in a fixed position and orientation on the display (e.g., at the same position on the display of the wearable head device 400A). This can maintain the illusion that the virtual objects exist in the real environment (and do not appear unnaturally positioned in the real environment, for example, as the wearable head device 400A shifts and rotates). In some examples, a compensatory transformation between coordinate spaces can be determined by processing images from the depth camera 444 (e.g., using simultaneous localization and mapping (SLAM) and / or visual odometry procedures) to determine a transformation of the wearable head device 400A relative to an inertial or environmental coordinate system. In the example shown in FIG. 4, the depth camera 444 can be coupled to the SLAM / visual odometry block 406 and can provide images to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the images and then determine the position and orientation of the user's head, which can be used to identify a transformation between the head coordinate space and the real coordinate space. Similarly, in some examples, an additional source of information regarding the user's head pose and location is obtained from the IMU 409 of the wearable head device 400A. Information from the IMU 409 can be integrated with information from the SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information regarding rapid adjustments of the user's head pose and position.

[0035] In some examples, the depth camera 444 can provide 3D images to a hand gesture tracker 411, which can be implemented within a processor of the wearable head device 400A. The hand gesture tracker 411 can identify the user's hand gestures, for example, by matching the 3D images received from the depth camera 444 to stored patterns representing hand gestures. Other suitable techniques for identifying the user's hand gestures will be apparent.

[0036] In some examples, one or more processors 416 may be configured to receive data from the headgear subsystem 404B, the IMU 409, the SLAM / visual odometry block 406, the depth camera 444, a microphone (not shown), and / or the hand gesture tracker 411. The processor 416 may also send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A, such as in examples where the handheld controller 400B is not tethered. The processor 416 may further communicate with additional components, such as an audiovisual content memory 418, a graphical processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to a head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to a left source of imagewise modulated light 424 and a right channel output coupled to a right source of imagewise modulated light 426. The GPU 420 may output stereoscopic image data to the sources of imagewise modulated light 424, 426. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 may receive an input from the processor 416 indicating a direction vector from the user to a virtual sound source (which may be moved by the user, e.g., via the handheld controller 400B). Based on the direction vector, the DSP audio spatializer 422 may determine a corresponding HRTF (e.g., by accessing an HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the believability and realism of virtual sounds by incorporating the user's relative position and orientation to the virtual sounds in the mixed reality environment, i.e., by presenting virtual sounds that match the user's expectations of what they would hear if the virtual sounds were real sounds in a real environment.

[0037] 4 , one or more of the processor 416, the GPU 420, the DSP audio spatializer 422, the HRTF memory 425, and the audio / visual content memory 418 may be included within the auxiliary unit 400C (which may correspond to the auxiliary unit 320 described above). The auxiliary unit 400C may include a battery 427 to power its components and / or it may provide power to the wearable head device 400A and / or the handheld controller 400B. Including such components within an auxiliary unit that may be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0038] While FIG. 4 presents elements corresponding to various components of exemplary wearable system 400, various other suitable arrangements of these components will be apparent to those skilled in the art. For example, elements presented in FIG. 4 as associated with auxiliary unit 400C may instead be associated with wearable head device 400A or handheld controller 400B. Furthermore, some wearable systems may dispense with handheld controller 400B or auxiliary unit 400C entirely. Such variations and modifications should be understood as falling within the scope of the disclosed examples.

[0039] (Mixed reality environment)

[0040] Like all people, users of mixed reality systems exist within a real environment, i.e., within the three-dimensional portion of the "real world" and all of its contents that are perceptible to the user. For example, users perceive the real environment using their normal human senses, i.e., sight, hearing, touch, taste, and smell, and interact with the real environment by moving their body within the real environment. Locations within the real environment can be described as coordinates in a coordinate space; for example, coordinates can include latitude, longitude, and altitude relative to sea level, distance in three orthogonal dimensions from a reference point, or other suitable values. Similarly, vectors can describe qualities that have direction and magnitude in the coordinate space.

[0041] A computing device may maintain a representation of a virtual environment, for example, in a memory associated with the device. As used herein, a virtual environment is a computer representation of a three-dimensional space. A virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other properties associated with that space. In some examples, the circuitry (e.g., a processor) of a computing device may maintain and update the state of the virtual environment; i.e., the processor may determine the state of the virtual environment at a second time based on data associated with the virtual environment and / or input provided by a user at a first time. For example, if an object in the virtual environment is located at a first coordinate at a time and has certain programmed physical parameters (e.g., mass, coefficient of friction), and input received from a user indicates that a force should be applied to the object in a certain directional vector, the processor may apply the laws of kinematics and use basic mechanics to determine the location of the object at that time. The processor may use any suitable known information about the virtual environment and / or any suitable input to determine the state of the virtual environment at a certain time. In maintaining and updating the state of the virtual environment, the processor may execute any suitable software, including software related to creating and deleting virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling input and output, software for implementing network operations, software for applying asset data (e.g., animation data for moving a virtual object over time), or many other possibilities.

[0042] An output device, such as a display or speaker, can present any or all aspects of the virtual environment to the user. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, lights, etc.) that can be presented to the user. A processor can determine a representation of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, viewing axis, and frustum) and render on the display a viewable scene of the virtual environment corresponding to that representation. Any suitable rendering technique can be used for this purpose. In some examples, the viewable scene may include only some virtual objects in the virtual environment and exclude certain other virtual objects. Similarly, the virtual environment may include audio aspects that can be presented to the user as one or more audio signals. For example, a virtual object in the virtual environment may generate sounds originating from the object's location coordinates (e.g., a virtual character may speak or trigger a sound effect); or the virtual environment may be associated with musical cues or ambient sounds that may or may not be associated with a specific location. A processor can determine audio signals corresponding to the "listener" coordinates (e.g., audio signals corresponding to a composite of sounds in the virtual environment and mixed and processed to simulate the audio signals that would be heard by a listener at the listener coordinates) and present the audio signals to the user via one or more speakers.

[0043] Because the virtual environment exists only as a computer construct, the user cannot directly perceive the virtual environment using their normal senses. Instead, the user can only indirectly perceive the virtual environment as presented to the user, for example, by a display, speakers, haptic output device, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data via input devices or sensors to a processor, which can use the device or sensor data to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use that data to cause the object to respond accordingly in the virtual environment.

[0044] (Digital reverberation and ambient audio processing)

[0045] An XR system can present a user with audio signals that appear to originate at a sound source with origin coordinates and travel in the direction of an orientation vector in the system. The user can perceive these audio signals as if they were real audio signals originating from the origin coordinates of the sound source and traveling along the orientation vector.

[0046] In some cases, audio signals may be considered virtual in that they correspond to computer signals in a virtual environment and not necessarily to real sounds in a real environment. However, virtual audio signals can be presented to a user as real audio signals detectable by the human ear, for example, as generated through speakers 120A and 120B of wearable head device 100 in FIG. 1 .

[0047] Advantages of the embodiments disclosed below include reduced network bandwidth, reduced power consumption, reduced computational complexity, and reduced computational delay, which may be particularly noticeable in mobile systems, including wearable systems, where processing resources, networking resources, battery capacity, and physical size and weight are often limited.

[0048] In an environment as dynamic as AR, the system may continuously render the audio signal. Rendering the audio signal using all of the virtual speakers may particularly lead to high computational power, extensive processing, high network bandwidth, high power consumption, etc. Therefore, it may be desirable to use modified virtual speaker panning to dynamically select and use some of the fixed virtual speakers based on one or more factors.

[0049] (Example Spatial Audio System)

[0050] 5A illustrates a block diagram of an exemplary spatial audio system according to some embodiments, and FIG. 5B illustrates a flow of an exemplary method for operating the system of FIG. 5A.

[0051] Spatial audio system 500 may include a spatial modeler 510, an internal spatial representation 530, and a decoder / virtualizer 540A. Spatial modeler 510 may include a direct path portion 512, one or more reflection portions 520 (optional), and a spatial encoder 526. Spatial modeler 510 may be configured to model a virtual environment. Direct path portion 512 may include a direct source 514 and, optionally, a Doppler 516. Direct source 514 may be configured to provide an audio signal (step 552 of process 550). Doppler 516 may receive a signal from direct source 514 and may be configured to introduce a Doppler effect into its input signal (step 554). For example, Doppler 516 may vary the pitch of a sound source (e.g., pitch shift) to vary with movement of the sound source, a user of the system, or both.

[0052] The reflection portion 520 may include a sound reflector 522, an optional Doppler 516, and a delay 524. The sound reflector 522 may be configured to introduce reflections into its signal (step 556). The introduced reflections may represent one or more characteristics of the environment. The Doppler 516 in the reflection portion 520 may receive a signal from the sound reflector 522 and may be configured to introduce a Doppler effect into its input signal (step 558). The delay 524 may receive a signal from the Doppler 516 and may be configured to introduce a delay (step 560).

[0053] Spatial encoder 526 may receive signals from direct path portion 512 and reflected portion 520. In some embodiments, the signal from direct path portion 512 to spatial encoder 526 may be the output signal from Doppler 516 of direct path portion 512. In some embodiments, the signal from reflected portion 520 to spatial encoder 526 may be the output signal from delay 524 of reflected portion 520.

[0054] The spatial encoder 526 may include one or more M-way pans 528. In some embodiments, each input received by the spatial encoder 526 may be associated with a unique 528. "Panning" may refer to distributing a signal across multiple speakers, multiple locations, or both. The M-way pan 528 may be configured to distribute its input signal across multiple numbers of virtual speakers (step 562). For example, the M-way pan 528 may distribute its input signal across all M virtual speakers. For example, as shown in FIG. 5A, M may be equal to 4, and each M-way pan 528 may be configured to distribute its input signal across four virtual speakers. Although the figure illustrates a system with four virtual speakers, examples of the present disclosure may include any number of virtual speakers.

[0055] As an example, an automobile system may include left and right speakers. Sound in such a system may be panned between the left and right speakers in the automobile by splitting the sound into two, one for each speaker. The scaling volume of each speaker may be set according to the configuration of the two speakers, and the results may be sent to the left and right speakers.

[0056] As another example, a surround sound system may include multiple speakers, such as six speakers. Sound in such a system may be panned as stereo between the six speakers. The sound may be split into six (instead of two as in the example automobile system), the scaling volume of each speaker may be set according to the configuration of the six speakers, and the result may be sent to the six speakers.

[0057] For example, a first M-way pan 528 may receive the output of Doppler 516 of direct path 512, and another M-way pan 528 may receive the output of reflected portion 520. Each M-way pan 528 may split its input signal so that it may be distributed across multiple outputs. Thus, each M-way pan 528 may have a greater number of outputs than inputs.

[0058] The spatial modeler 510 may output signals to the internal spatial representation 530 (step 564). In some embodiments, the output from the spatial modeler 510 may include an output for each M-way pan 528. The internal spatial representation 530 may be configured to represent the spatial configuration of the virtual environment (step 566). One exemplary representation may include representing the relative locations of the user, sound sources, and virtual speakers. In some embodiments, the internal spatial representation 530 may output one or more signals representing the user's head pose rotation, head pose translation, sound field decoding, one or more head-related transfer functions (HRTFs), or a combination thereof, of the system 500. In some embodiments, the internal spatial representation 530 may be a representation of a non-Ambisonics multi-channel-based system, an Ambisonics / wave field-based system, or the like. One exemplary Ambisonics / wave field-based system may be Higher Order Ambisonics (HOA).

[0059] Internal spatial representation 530 may output its signal 552 to decoder / virtualizer 540A (step 568). Decoder / virtualizer 540 may decode its input signal and introduce virtual sounds into the signal (step 570). Step 570 may include multiple substeps and is discussed in more detail below. The system then outputs the signal from decoder / virtualizer 540 as left signal 502L, which may be output to the left speaker, and as right signal 502R, which may be output to the right speaker (step 580).

[0060] System 500 may include any number of different types of decoder / virtualizers 540. One exemplary decoder / virtualizer 540A is shown in Figure 5A. Other exemplary decoder / virtualizers 540 are discussed below.

[0061] The decoder / virtualizer 540A may include a rotation / translation representation 542, a sound field decoder 544, one or more HRTFs 546, and one or more combiners 548. FIG. 5C illustrates the flow of an example method for operating an example decoder / virtualizer, which may be referred to as step 570-1. The rotation / translation representation 542 may receive a signal from the interior spatial representation 530 and may be configured to introduce a representation of a movement associated with the audio signal. For example, the movement may be that of a sound source, a user, or both (step 572). The rotation / translation representation 542 may output a signal to the sound field decoder 544. The sound field decoder 544 may receive a signal from the rotation / translation representation 542 and may be configured to decode the signal (step 574). Each HRTF 546 may receive a signal from the sound field decoder 544. Each HRTF 546 may be configured to determine an HRTF corresponding to its input signal and apply it to the signal (step 576). One or more HRTFs 546 may be collectively referred to as a speaker transformer. In some embodiments, the HRTFs 546 may be configured for finite impulse response (FIR) filtering. Each combiner 548 may receive and combine signals from the HRTFs 546 (step 578).

[0062] In some embodiments, the decoder / virtualizer 540A may represent a "baseline" processing overhead, which may be complex and involve matrix calculations and long FIR filters to apply HRTF processing for each virtual speaker.

[0063] The output from combiner 548 may be an output signal from system 500. In some embodiments, output signal 502 from system 500 may be audio signals for left and right speakers (e.g., speakers 120A and 120B in FIG. 1).

[0064] In some instances, the spatial audio system of Figure 5A may be beneficial when the number of sound sources for reproduction is large. However, in some instances, the spatial audio system of Figure 5A may not be beneficial when the number of sound sources for reproduction is small. It may be desirable to utilize the efficiency of a non-Ambisonics multi-channel based spatial audio system or an Ambisonics-based spatial audio system, such as system 500 of Figure 5A, in an efficient manner for situations when the number of sound sources for reproduction is small.

[0065] There may be ways to improve the efficiency of spatialization using sound field synthesis and decoding. The first way may be through low-energy speaker detection and culling. In low-energy speaker detection and culling, if the energy output of a virtual speaker channel of a non-Ambisonics multi-channel based spatial audio system or an Ambisonics / sound field channel of an Ambisonics-based spatial audio system is less than a predetermined threshold, processing of the signal from the virtual speaker channel is not performed. In some embodiments, the system may determine whether the output of a given virtual speaker is greater than a predetermined threshold, for example, before sound field decoding is performed on the signal from that given virtual speaker. Low-energy speaker detection and culling will be discussed in more detail below.

[0066] A second method for improving the efficiency of spatialization using sound field synthesis and decoding can be source geometry-based virtual speaker culling. In source geometry-based virtual speaker culling, decoder / virtualizer processing can be selectively disabled. Selective disabling (or selective enabling) can be based on the location of the sound source relative to the user / listener. Source geometry-based virtual speaker culling will be discussed in more detail below.

[0067] A third method may be to combine low-energy speaker detection and culling techniques with source-virtual speaker coupling techniques.

[0068] The spatial modeler 510 may have a computational complexity that may represent the number of operations required to process an audio signal. The computational complexity may be proportional to M multiplied by N, where M may be equal to the number of sound sources (including direct sources and optional reflections) and N may be equal to the number of channels required to represent the Ambisonic sound field. In some embodiments, N is (O+1) 2 where O is the Ambisonics order used.

[0069] The decoder / virtualizer 540 may have a computational complexity proportional to nVS, where nVS is the number of virtual speakers. The computational power of each speaker may be high and may generally consist of a pair of FIR filters, which are typically implemented using a fast Fourier transform (FFT) or an inverse FFT (IFFT), both of which may be computationally expensive processes.

[0070] Exemplary Low Energy Output Detection and Culling Method

[0071] In some embodiments, some virtual speakers may have little or no signal input energy: for example, when a spatial audio system has a small number of sound sources. Speaker virtualization processing can be a computationally expensive (e.g., CPU-intensive) process. For example, if a sound source is located at zero degree azimuth (e.g., directly in front of the user), there may be little or no energy in the signal from a virtual speaker located at 90-270 degree azimuth (e.g., behind the user). Low-energy signals may not have a significant effect on the perceived location of the sound source, and therefore, performing speaker virtualization processing on low-energy signals and / or determining the characteristics of the corresponding virtual speakers may be computationally inefficient.

[0072] To reduce required computational resources, a system employing a low-energy output detection and culling method may include a detector located between the sound field decoder and the HRTF. Alternatively, the detector may be located between the multi-channel output and the HRTF. The detector may be configured to detect one or more energy levels associated with one or more audio signals from one or more virtual speakers.

[0073] If the energy level of the signal emanating from the virtual speaker Vn is below the energy threshold α, the signal may be considered a low-energy signal. According to which the detected energy level associated with the audio signal is below the energy threshold α, the HRTF block and its processing of the low-energy signal may be bypassed.

[0074] Determining the energy level of a signal may use any number of techniques. For example, an RMS algorithm may be applied to the signal routed to the virtual speaker to measure its energy. "Attack" and "release" times similar to those used by traditional audio compressors may be used to prevent the speaker's signal from suddenly "popping in" and "popping out."

[0075] FIG. 6 illustrates an example configuration of sound sources and speakers according to some embodiments. System 600 may include a sound source 620 and multiple speakers. The multiple speakers 622 may include one or more active virtual speakers 622A and one or more inactive virtual speakers 622B. An active virtual speaker 622A may be one whose signal is processed by the HRTF 546 at a given time. An inactive virtual speaker 622B may be one whose signal does not need to be processed by the HRTF 546, for example, because the signal has already been processed at a previous time or because the system has determined that the signal from virtual speaker 622B does not require processing. M may refer to the number of sound sources being played, and N may refer to the number of virtual speakers in the system. While the figure illustrates a single sound source, examples of the present disclosure may include any number of sound sources. Although the figure illustrates eight sound sources, examples of the present disclosure can include any number of sources, such as sixteen (N=16).

[0076] As an example, system 600 may include a single (M=1) sound source 620 and eight virtual speakers 622, as shown in the figure. At a given instance, the majority of the energy may be output across only three virtual speakers. That is, system 600 may have three active virtual speakers at a first time. For example, virtual speakers 622A-1, 622A-2, and 622A-3 may be active virtual speakers. In some embodiments, active virtual speaker 622A may be those closest to sound source 620. In addition, system 600 may include five inactive virtual speakers 622B. System 600 may determine that the energy level from each of the five inactive virtual speakers is less than an energy threshold and, in accordance with such determination, may bypass HRTF processing of signals from the five inactive virtual speakers 622B.

[0077] The system 600 may also determine that the energy level from each of the active virtual speakers is not less than an energy threshold, and may perform HRTF processing of the signals from the three active virtual speakers 622A according to such determination.

[0078] System 600 may output two signals, one for the right speaker and one for the left speaker (such as right signal 502R and left signal 502L), as shown in FIG. 5A. The reduction in the number of HRTF operations by bypassing HRTF processing may be equal to the number of inactive virtual speakers multiplied by the number of signals output from the system. In the example of FIG. 6, HRTF processing for five signals is bypassed, so 10 HRTF operations (5 inactive virtual speakers x 2 output signals) may be saved.

[0079] As another example, if a system includes 16 virtual speakers, of which 13 are inactive virtual speakers, the number of HRTF operations saved may be equal to 26 (16 virtual speakers x 2 output signals).

[0080] 7A illustrates a block diagram of an exemplary decoder / virtualizer including multiple detectors, according to some embodiments. 7B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 7A, according to some embodiments. In some embodiments, as discussed below, decoder / virtualizer 540B may be included in system 500 instead of decoder / virtualizer 540A (shown in FIG. 5A). Instead of step 570-1 (shown in FIG. 5C), step 570-2 may be included in process 550.

[0081] The decoder / virtualizer 540B may include a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more switches 712, one or more HRTFs 546, and one or more combiners 548. The decoder / virtualizer 540B may receive a signal 552 from the internal spatial representation 530 (as shown in FIG. 5A). The rotation / translation representation 542 may receive a signal from the internal spatial representation 530 and may be configured to introduce a representation of the movement of the sound source, the user, or both (step 772). The rotation / translation representation 542 may output a signal to the sound field decoder 544. The sound field decoder 544 may receive a signal from the rotation / translation representation 542 and may be configured to decode the signal (step 774). The sound field decoder 544 may output a signal to the detector 710.

[0082] The detectors 710 may receive signals from the sound field decoder 544 and may be configured to determine the energy level of their input signals (step 776). Each detector 710 may be coupled to its own switch 712. If the energy level of the input signal (from the sound field decoder 544) is greater than or equal to an energy threshold (step 778), the switch 712 may close the loop, thereby routing the input signal (from the detector 710) to the HRTF 546 to which the switch is coupled (step 780). Each HRTF determines a corresponding HRTF and applies it to the signal (step 782).

[0083] If the energy level of the input signal is less than the energy threshold, the switch 712 may be open so that the input signal (from the detector 710) is not coupled to the corresponding HRTF 546. Thus, the corresponding HRTF 546 may be bypassed (step 784).

[0084] The signals from the HRTFs 546 may be output to the combiner 548 (step 786). The combiner 548 may be configured to combine (e.g., add, sum, etc.) the signals from the HRTFs 546. Those signals that bypass the HRTFs 546 are not combined by the combiner 548. The output from the combiner 548 may be an output signal from the system 500. In some embodiments, the output signal 502 from the system 500 may be audio signals for left and right speakers (e.g., speakers 120A and 120B of FIG. 1).

[0085] In some embodiments, each detector 710 can be coupled to a unique signal corresponding to a virtual speaker. In this way, processing for each virtual speaker 622 can be performed independently (i.e., processing for one speaker, such as 622A-1, can occur without affecting the processing of another speaker, such as 622B).

[0086] In some embodiments, the type of decoder / virtualizer 540 may depend on the number of sound sources. For example, if the number of sound sources is less than or equal to a predetermined sound source threshold, decoder / virtualizer 540B of FIG. 7A may be included in system 500. In such an instance, the signal from sound field decoder 544 may be input to detector 710.

[0087] If the number of sound sources is greater than a predetermined sound source threshold, the decoder / virtualizer 540A of Figure 5A may be included in the system. In such an instance, the signal from the sound field decoder 544 may be input to the HRTF 546.

[0088] In some embodiments, the system may include a decoder / virtualizer 540 that can select whether to perform or bypass the detector and its energy level detection. FIG. 8A illustrates a block diagram of an exemplary decoder / virtualizer according to some embodiments. FIG. 8B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 8A according to some embodiments. In some embodiments, instead of decoder / virtualizer 540A (shown in FIG. 5A) and decoder / virtualizer 540B (shown in FIG. 7A), decoder / virtualizer 540C may be included in system 500. Instead of step 570-1 (shown in FIG. 5C), step 570-3 may be included in process 550.

[0089] Decoder / virtualizer 540C, like decoder / virtualizer 540B discussed above, may include a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more first switches 712, one or more HRTFs 546, and one or more combiners 548. Steps 872, 874, and 882 may correspond to and be similar to steps 772, 774, and 782 discussed above.

[0090] Decoder / virtualizer 540C may also include a second switch 814. The second switch 814 may be configured to open or close a first loop from sound field decoder 544 to detector 710 and first switch 712. Additionally or alternatively, the second switch 814 may be configured to open or close a second loop from system 500 that bypasses detector 710 and first switch 712. In some embodiments, second switch 814 may be a bidirectional switch configured to select between passing the signal directly to detector 710 (first loop) or directly to HRTF 546 (second loop).

[0091] For example, the system may determine whether the number of sound sources is greater than or equal to a predetermined sound source threshold (step 876). If the number of sound sources is greater than or equal to the predetermined sound source threshold, the second switch 814 may close the second loop and pass the signal from the sound field decoder 544 directly to the HRTFs 546 (step 878). Each HRTF 546 then determines a corresponding HRTF and applies it to the signal (step 880). When the number of sound sources exceeds in number, the likelihood that the signal will have a low energy level may be reduced.

[0092] On the other hand, if the number of sound sources is less than a predetermined sound source threshold, the signal is likely to have a low energy level, and therefore the second switch 814 may close the first loop and pass the signal from the sound field decoder 544 directly to the detector 710 (step 882). The detector 710 may receive the signal from the sound field decoder 544 and may be configured to determine the energy level of that input signal (step 884). If the energy level of the input signal (from the sound field decoder 544) is equal to or greater than the energy threshold (step 886), the switch 712 closes the loop, thereby routing that input signal (from the detector 710) to the HRTF 546 to which the switch is coupled (step 888). If the energy level of the input signal is less than the energy threshold, the switch 712 may open so that the input signal (from the detector 710) is not coupled to the corresponding HRTF 546, causing the HRTF 546 to be bypassed (step 890).

[0093] The signal from the HRTF 546 may be output to the combiner 548 (step 892).

[0094] In some embodiments, one or more energy threshold detections may be active in response to energy, hi some embodiments, one or more energy threshold detections may be active in response to amplitude, and may be subject to conventional attack, release times, etc.

[0095] Exemplary Source Geometry-Based Speaker Culling Method

[0096] Source-geometry-based virtual speaker culling may be another method for reducing CPU consumption. In some embodiments, source-geometry-based virtual speaker culling may include selectively disabling decoder / virtualizer processing (e.g., decoder / virtualizer 540A of FIG. 5A , decoder / virtualizer 540B of FIG. 7A , decoder / virtualizer 540C of FIG. 8A , etc.). In some embodiments, the selective disabling (or selective enabling) may be based on the location of the sound source relative to the user / listener. In some embodiments, the selective disabling of decoder / virtualizer processing may include bypassing all of the decoder / virtualizer processing blocks.

[0097] In source geometry-based virtual speaker culling, Ambisonic outputs can be calculated. If the Ambisonic outputs require a significant amount of energy to be decoded, it may be beneficial to use simpler methods (requiring less CPU consumption), such as real-time energy detection methods. Additionally, in some embodiments, real-time energy detection methods can perform calculations less frequently.

[0098] FIG. 9 illustrates an example configuration of sound sources and speakers according to some embodiments. System 900 may include a sound source 920 and multiple speakers. Compared to system 600 of FIG. 6, sound source 920 may be located at a second position that may differ from the first position of sound source 620 of FIG. 6. Multiple speakers 922 may include one or more active virtual speakers 922A, one or more inactive virtual speakers 922B, and one or more inactive virtual speakers 922C. Active virtual speakers 922A and inactive virtual speakers 922B may correspond to and be similar to active virtual speakers 622A and inactive virtual speakers 622B of FIG. 6, respectively.

[0099] Inactive virtual speaker 922C may differ from inactive virtual speaker 922B in that virtual speaker 922C is active at a first time but its signal is processed at a second time (e.g., a ring-out period). In the example of FIG. 9, sound source 920 may have moved from a first position (e.g., close to virtual speaker 922C) to a second position (e.g., not close to virtual speaker 922B). Due to the movement of the sound source, the two virtual speakers may no longer have sound sources to mix within them at the second time. Due to the filtering of the two virtual speakers, the two virtual speakers may need to be active for a subsequent frame (e.g., a second time) to properly complete the filtering.

[0100] In some embodiments, a system may include decoder / virtualizer 540 in a system that uses active virtual speakers. FIG. 10A illustrates a block diagram of an exemplary decoder / virtualizer used in a system that includes active speakers, according to some embodiments. FIG. 10B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 10A, according to some embodiments. In some embodiments, decoder / virtualizer 540D may be included in system 500 in place of decoder / virtualizer 540A (shown in FIG. 5A), decoder / virtualizer 540B (shown in FIG. 7A), and decoder / virtualizer 540C (shown in FIG. 8A). Step 570-4 may be included in process 550 in place of step 570-1 (shown in FIG. 5C), step 570-2 (shown in FIG. 7B), and step 570-3 (shown in FIG. 8B).

[0101] Decoder / virtualizer 540C, like decoder / virtualizer 540B and decoder / virtualizer 540C discussed above, may include a sound field decoder 544, one or more HRTFs 546, and one or more combiners 548. Steps 1072, 1076, 1078, and 1080 may correspond to and be similar to steps 872, 874, and 782 discussed above.

[0102] The decoder / virtualizer 540D may also include a rotation / translation representation 1042 and a sound field decode decision 1044. The rotation / translation representation 1042 may receive a signal from the interior space representation 530 and may be configured to introduce a representation of the movement of the sound source, the user, or both (step 1072). The representation of movement may also take into account the orientation / altitude of the sound source 920. The rotation / translation representation 542 may output a signal to the sound field decoder decision 1044.

[0103] The sound field decoder determination 1044 may receive signals from the rotation / translation representation 1042 and may be configured to determine signals that have "salient" outputs and pass those signals to the sound field decoder 544 (step 1074). A salient output may be an output that will affect the perceived sound. For example, a salient output may be an audio signal that has an amplitude that is above a predetermined amplitude threshold. The sound field decoder 544 may receive signals from the sound field decoder determination 1044 that have salient outputs and may be configured to decode the signals (step 1076). In some embodiments, the sound field decoder 1044 may receive signals from the sound field decoder determination 1044 that have salient outputs. Each HRTF 546 may receive a signal from the sound field decoder 544. Each HRTF 546 may be configured to determine an HRTF that corresponds to its input signal and apply it to the signal (step 1078). One or more HRTFs 546 may be collectively referred to as a speaker transformer. Each combiner 548 may receive and combine signals from an HRTF 546 (step 1080).

[0104] In some embodiments, those audio signals that do not have significant output (e.g., have amplitudes below a predetermined amplitude threshold) may not be passed to the sound field decoder 544. Thus, the sound field decoder 544 and HRTF 546 on audio signals that do not have significant output may be bypassed.

[0105] An exemplary source geometry-based speaker culling method can designate virtual speakers to be active virtual speakers based on the location (e.g., X, Y, Z location) of the sound source. The location of the sound source may represent the location of the source object. The system may determine the location of each sound source and determine a virtual speaker located proximate to each sound source. In some embodiments, the determination of the virtual speaker located proximate to the sound source may be performed, for example, at the beginning of every video frame (in a video frame rate-based approach). A video frame rate-based approach may require fewer computations than other approaches, such as a sample rate-based approach.

[0106] A sound source may contribute significantly to a particular virtual loudspeaker, for example, based on the video frame rate-based approach calculation and the Ambisonic decoding formula. As discussed above, a virtual loudspeaker that contributes little or no energy when decoded may be bypassed from the corresponding Ambisonic decoding and HRTF processing of the decoded Ambisonics channel. In some embodiments, the system may disable any processing blocks that are bypassed.

[0107] An example pseudocode for performing the specified method may be as follows: For each sound source, S and decode channel n Enable[n] |= f(sourcePosition Vector3, sourceOrientation Vector3, ListenerPosition Vector3, ListenerOrientation Vector3, VirtualSpeakerPosition[n] Vector3). (Ambisonic / sound field example) For each Ambisonic Decode Channel If (Enable[n]) { AmbisonicDecode(n) Virtualize(n) } (Multi-channel example) For each channel If (Enable[n]) { Virtualize(n) }

[0108] With respect to the above pseudocode, the variable sourcePosition may refer to the position of the sound source, sourceOrientation may refer to the orientation of the sound source, ListenerPosition may refer to the position of the user / listener, ListenerOrientation may refer to the orientation of the user / listener, VirtualSpeakerPosition may refer to the position of the virtual speaker, AmbisonicDecode may refer to a function that performs Ambisonic decoding, and Virtualize may refer to a function that performs virtualization.

[0109] With respect to the pseudocode above, for each sound source S and decode channel n, the decode channel n may be enabled based on one or more factors such as the position of the sound source S, the orientation of the sound source S, the position of the user / listener, the orientation of the user / listener, and the position of the virtual speaker. Still referring to the pseudocode above, for each Ambisonic decode channel, if the channel is enabled, the system may execute an AmbisonicDecode function and a Virtualize function.

[0110] The pseudocode can be enhanced by providing a "ring-out" period for each virtual speaker. For example, if a source moves in position during a video frame, it can be determined that a virtual speaker may no longer have any sound sources mixed into it. However, due to the filtering of the virtual speaker, that virtual speaker may need to be the active speaker for the following frame in order to properly complete the filtering.

[0111] Examples of the present disclosure may include using all active sound sources to determine decoded sound field outputs that have "significant" outputs (e.g., outputs that will affect the perceived sound field). Ambisonics or non-Ambisonics multi-channel outputs that will affect the perceived sound field may be decoded. Furthermore, in some embodiments, only the HRTFs 546 corresponding to those detected outputs are processed. When the number of sound sources is small or large but close to each other, there may be significant CPU savings for synthetically generated Ambisonic sound fields or non-Ambisonic multi-channel rendering.

[0112] Example Method Combination of Source Geometry-Based Virtual Speaker Culling Method and Low Energy Output Detection and Culling Method

[0113] In some embodiments, both source geometry-based virtual speaker culling and low energy output detection and culling may be used consecutively to further reduce CPU consumption. As described above, source geometry-based virtual speaker culling may include, for example, selectively disabling virtual speaker processing based on, for example, the location of the sound source relative to the user / listener. Low energy output detection and culling may include, for example, placing a signal energy / level detector between the sound field decoding or multi-channel output and the HRTF processing. The output / results of source geometry-based virtual speaker culling may be input to low energy output detection and culling.

[0114] With respect to the systems and methods described above, elements of the systems and methods can be implemented by one or more computer processors (e.g., CPUs or DSPs), as appropriate. The present disclosure is not limited to any particular configuration of computer hardware, including computer processors, used to implement these elements. In some cases, multiple computer systems can be employed to implement the systems and methods described above. For example, a first computer processor (e.g., a processor of a wearable device coupled to a microphone) can be utilized to receive input microphone signals and perform initial processing of those signals (e.g., signal conditioning and / or segmentation, such as those described above). A second (perhaps more computationally powerful) processor can then be utilized to perform more computationally intensive processing, such as determining probability values ​​associated with speech segments of those signals. Another computer device, such as a cloud server, can host a speech recognition engine, to which the input signals are ultimately provided. Other suitable configurations will become apparent and are within the scope of the present disclosure.

[0115] Although the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are to be understood as being included within the scope of the disclosed examples as defined by the appended claims.

Claims

1. A method for spatially rendering an audio signal, said method comprising: determining a model of the virtual environment; determining a spatial configuration of the virtual environment; determining a plurality of signals associated with the spatial configuration and the model; determining whether an energy level of the plurality of signals exceeds a predetermined threshold; decoding one or more signals of the plurality of signals in accordance with a determination that the energy level exceeds the predetermined threshold, the energy level corresponding to the one or more signals; rendering the audio signal based on the one or more decoded signals; A method comprising:

2. applying the one or more signals to a head-related transfer function (HRTF) in accordance with a determination that the energy level exceeds the predetermined threshold; and refraining from applying the one or more signals to the HRTF in accordance with a determination that the energy level does not exceed the predetermined threshold. The method of claim 1 further comprising:

3. Decoding the one or more signals includes applying a first set of processing blocks to the one or more signals, the method further comprising: bypassing a second set of processing blocks in accordance with a determination that the energy level does not exceed the predetermined threshold, the second set of processing blocks being associated with one or more inactive virtual speakers; The method of claim 1 , comprising:

4. The method described in claim 3, wherein the bypassing of the second set of processing blocks includes refraining from transmitting the one or more signals to a decoder, the decoder comprising the second set of processing blocks.

5. In accordance with a determination that the energy level exceeds the predetermined threshold, transmitting the one or more signals to the decoder, the decoder comprising the first set of processing blocks. The method of claim 3 further comprising:

6. The method of claim 1, further comprising: applying a HRTF to one or more signals in accordance with a determination that the energy level exceeds the predetermined threshold, the one or more signals being received from the decoder. The method of claim 5 further comprising:

7. The method of claim 6, wherein determining the model of the virtual environment comprises: receiving one or more input sound signals from a direct sound source and from a reflected sound source; modifying the one or more input sound signals to simulate a Doppler effect; applying a delay to the one or more input sound signals; panning the one or more sound signals across a plurality of virtual speakers; Including, Decoding the one or more signals includes: The method of claim 1 , comprising determining one or more virtual sounds associated with movement of one or more of the direct sound source, the reflected sound source, and a user.

8. A system for spatially rendering an audio signal, the system comprising: a wearable head device configured to provide the audio signal to a user; one or more processors configured to perform the method; Equipped with The method comprises: determining a model of the virtual environment; determining a spatial configuration of the virtual environment; determining a plurality of signals associated with the spatial configuration and the model; determining whether an energy level of the plurality of signals exceeds a predetermined threshold; decoding one or more signals of the plurality of signals in accordance with a determination that the energy level exceeds the predetermined threshold, the energy level corresponding to the one or more signals; rendering the audio signal based on the one or more decoded signals; Including, the system.

9. The method comprises: applying the one or more signals to a head-related transfer function (HRTF) in accordance with a determination that the energy level exceeds the predetermined threshold; and refraining from applying the one or more signals to the HRTF in accordance with a determination that the energy level does not exceed the predetermined threshold. The system of claim 8 further comprising:

10. The decoding of the one or more signals comprises applying a first set of processing blocks to the one or more signals, the method further comprising: bypassing a second set of processing blocks in accordance with a determination that the energy level does not exceed the predetermined threshold, the second set of processing blocks being associated with one or more inactive virtual speakers; The system of claim 8 , comprising:

11. The system described in claim 10, wherein the bypassing of the second set of processing blocks includes refraining from transmitting the one or more signals to a decoder, the decoder comprising the second set of processing blocks.

12. The method comprising: transmitting the one or more signals to the decoder in accordance with a determination that the energy level exceeds the predetermined threshold, the decoder comprising the first set of processing blocks; The system of claim 11 further comprising:

13. The method comprising: applying a HRTF to the one or more signals in accordance with a determination that the energy level exceeds the predetermined threshold, the one or more signals being received from the decoder; The system of claim 12 further comprising:

14. The method of claim 13, wherein determining the model of the virtual environment comprises: receiving one or more input sound signals from a direct sound source and from a reflected sound source; modifying the one or more input sound signals to simulate a Doppler effect; applying a delay to the one or more input sound signals; panning the one or more sound signals across a plurality of virtual speakers; Including, Decoding the one or more signals includes: The system of claim 8 , further comprising determining one or more virtual sounds associated with movement of one or more of the direct sound source, the reflected sound source, and a user.

Citation Information

Patent Citations

  • Speech spatialization and environmental simulation

    JP2010520671A