Efficient Rendering of Virtual Sound Fields

By implementing modified virtual speaker panning and dynamically selecting active speakers in the audio system, the inefficiencies in processing multiple sound sources are addressed, resulting in improved computational efficiency and resource utilization.

JP7699632B2Active Publication Date: 2025-06-27MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023150990
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-12
Filing Date
2023-09-19
Publication Date
2025-06-27
Estimated Expiration
2039-06-12

AI Technical Summary

Technical Problem

Existing spatial audio rendering systems face inefficiencies in processing large numbers of sound sources, leading to high computational requirements and resource utilization.

Method used

The proposed audio system employs modified virtual speaker panning, dynamically selecting a subset of active virtual speakers based on proximity to sound sources, and using multiple FIR filters for each active speaker to enhance rendering efficiency.

Benefits of technology

This approach reduces the number of FIR filters required, thereby decreasing computational resources and improving processing efficiency, especially when handling a large number of sound sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699632000001
    Figure 0007699632000001
  • Figure 0007699632000002
    Figure 0007699632000002
  • Figure 0007699632000003
    Figure 0007699632000003
Patent Text Reader

Abstract

To provide efficient rendering of virtual soundfields.SOLUTION: An audio system and method of spatially rendering audio signals that uses modified virtual speaker panning is disclosed. The audio system may include a fixed number F of virtual speakers, and the modified virtual speaker panning may dynamically select and use a subset P of the fixed virtual speakers. The subset P of virtual speakers may be selected using a low energy speaker detection and culling method, a source geometry-based culling method, or both. One or more processing blocks in a decoder / virtualizer may each be bypassed based on an energy level of an associated audio signal or a location of a sound source relative to a user / listener. In some embodiments, a virtual speaker that is designated as an active virtual speaker at a first time may also be designated as an active virtual speaker at a second time to ensure the processing completion.SELECTED DRAWING: Figure 5A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 684,093, filed on Jun. 12, 2018, which is hereby incorporated by reference in its entirety. (Technical Field)

[0002] The present disclosure generally relates to spatial audio rendering and associated systems. More specifically, the present disclosure relates to systems and methods for enhancing the efficiency of virtual - speaker - based spatial audio systems.

Background Art

[0003] Virtual environments are prevalent in computing environments and find use in video games (where the virtual environment can represent a game world), maps (where the virtual environment can represent terrain to be navigated), simulations (where the virtual environment can simulate a real - world environment), digital storytelling (where virtual characters can interact with each other within the virtual environment), and many other applications. Modern computer users generally perceive and interact with virtual environments comfortably. However, a user's virtual - environment experience can be limited by the technology used to present the virtual environment. For example, conventional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may not be able to realize a virtual environment so as to attract people and create a realistic and immersive experience.

[0004] Virtual reality ("VR"), augmented reality ("AR"), mixed reality ("MR"), and related technologies (collectively, "XR") share the ability to present sensory information corresponding to a virtual environment represented by data within a computer system to a user of an XR system. Such systems can provide a uniquely enhanced sense of immersion and presence by combining virtual visual and audio cues with the sights and sounds of reality. Thus, it may be desirable to present digital sound to a user of an XR system such that the sound appears to occur naturally within the user's physical environment and is consistent with the sound that the user expects. Generally speaking, a user expects that virtual sounds will have the acoustic characteristics of the physical environment in which they are heard. For example, a user of an XR system within a large concert hall expects the virtual sound of the XR system to have a sound quality similar to that of a large cave, and conversely, a user within a small apartment will expect the sound to be more attenuated, closer, and direct. Additionally, a user expects that virtual sounds will be provided without latency.

[0005] Among other techniques, ambisonics and non-ambisonics can be used to generate spatial audio. For a number of sound source objects, ambisonics and non-ambisonics can be an efficient way to render spatial audio by virtue of their design and architecture. This can be particularly true when reflections are modeled. Ambisonics and non-ambisonics multi-channel based spatial audio systems can render an audio signal through a number of steps. Exemplary steps can include an encoding step per source, a fixed overhead sound field decoding step, and / or a fixed speaker virtualization step. One or more hardware components can perform the steps.

[0006] In a first method for rendering an audio signal, each sound source can have a pair of its own finite impulse response (FIR) filters. In such a system, the perceived position of the sound can be changed by changing the filter coefficients of the FIR filters. In some embodiments, each sound can use a plurality (e.g., two pairs) of FIR filters. Each pair can use two filters (i.e., four FIR filters). When the sound moves around the virtual environment, the FIR filters can be cross-faded. In some embodiments, four FIR filters can be used for each sound.

[0007] In a second method for rendering an audio signal, virtual speaker panning can be implemented using a fixed number of virtual speakers. Each sound source can be panned across the fixed virtual speakers. In some embodiments, a plurality (e.g., two) of FIR filters can be used for each virtual speaker. Virtual speaker panning can be efficient for certain applications and can use very few computational resources.

[0008] In some embodiments, one method can have high efficiency compared to other methods depending on the number of sounds being played simultaneously. For example, 30 sounds can be played simultaneously. If four FIR filters are used for each sound source, 120 FIR filters (30 sound sources × 4 FIR filters per sound source = 120 FIR filters) can be required for the first method. If two FIR filters are used for each virtual speaker, only 32 FIR filters can be required for the second method (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters).

[0009] As another example, only one sound may be reproduced. The first method may require only four FIR filters (one sound source × four FIR filters per sound source = four FIR filters), while the second method may require 32 FIR filters (16 virtual speakers × two FIR filters per virtual speaker = 32 FIR filters).

[0010] As illustrated through the above examples, the first method may be beneficial for a small number of sounds, and the second method may be beneficial for a large number of sounds. Therefore, an audio system and method that enhance efficiency may be desired based on the number of sound sources at a given time.

Summary of the Invention

Means for Solving the Problems

[0011] An audio system and method for rendering an audio signal are disclosed, and the system uses modified virtual speaker panning. The audio system may include a fixed number F of virtual speakers, and the modified virtual speaker panning may dynamically select and use a subset P of the fixed virtual speakers. Each sound source may be panned across the subset P of virtual speakers. In some embodiments, multiple (e.g., two) FIR filters may be used for each virtual speaker in the subset P. The subset P of virtual speakers may be selected based on one or more factors such as proximity to the sound source. The subset P of virtual speakers may be referred to as active speakers.

[0012] The modified virtual speaker panning method can be compared, as an example, with the first and second methods disclosed above. When three sounds are played simultaneously and the audio system has 16 fixed virtual speakers, the first method may require 12 FIR filters (3 sound sources × 4 FIR filters per sound source = 12 FIR filters), and the second method may require 32 FIR filters (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters). On the other hand, the modified virtual speaker panning method may dynamically select three virtual speakers such that they are active virtual speakers as part of subset P. The modified virtual speaker panning method may require six FIR filters, i.e., two FIR filters for each active virtual speaker (3 virtual speakers × 2 FIR filters = 6 FIR filters). The present invention provides, for example, the following. (Item 1) A method for spatially rendering an audio signal, the method comprising: modeling a virtual environment using a spatial model; distributing signals from the spatial model over a plurality of virtual speakers using a spatial encoder; representing a spatial configuration of the virtual environment using an internal spatial representation; decoding signals from the internal spatial representation using a decoder / virtualizer; introducing virtual sounds into the decoded signals using a decoder / virtualizer; selectively bypassing one or more processing blocks associated with non-active virtual speakers within the decoder / virtualizer; combining signals from the decoder / virtualizer; outputting the combined signals as the audio signal and a method comprising the steps of: (Item 2) determining an energy level associated with the signals from the sound field decoder; determining whether each of the detected energy levels is less than an energy threshold and further comprising wherein the selective bypass of the one or more processing blocks comprises bypassing head-related transfer function (HRTF) processing of the corresponding signal from the sound field decoder according to a determination that the detected energy level of at least one of the virtual speakers is less than the energy threshold wherein the sound field decoder is the method according to item 1, included within the decoder / virtualizer (Item 3) The method according to item 2, further comprising performing HRTF processing of the corresponding signal from the sound field decoder according to a determination that the detected energy level of at least one of the virtual speakers is not less than the energy threshold (Item 4) further comprising determining whether the number of sound sources is equal to or greater than a predetermined sound source threshold wherein the selective bypass of the one or more processing blocks comprises bypassing a plurality of detectors and passing a signal from the sound field decoder directly to a plurality of HRTF blocks when the number of sound sources is equal to or greater than the predetermined sound source threshold wherein the plurality of detectors and the plurality of HRTF blocks are included within the decoder / virtualizer, the method according to item 1 (Item 5) The method according to item 4, further comprising passing a signal from the sound field decoder directly to the plurality of detectors according to a determination that the number of sound sources is less than the predetermined sound source threshold (Item 6) determining the location of each sound source and determining which of the plurality of virtual speakers are located in proximity to each respective sound source The method according to item 1, further comprising (Item 7) The determination of which of the plurality of virtual speakers is located close to each of the sound sources is performed in all video frames, the method according to item 6. (Item 8) The selective bypass of the one or more processing blocks in the decoder / virtualizer includes bypassing all of the one or more processing blocks associated with at least one speaker that is not located close to each of the sound sources in the decoder / virtualizer, the method according to item 6. (Item 9) Introducing an expression of movement associated with the audio signal using a rotation / translation expression, Determining whether the amplitude of the signal from the rotation / translation expression is equal to or greater than a predetermined amplitude threshold and further including, The selective bypass of the one or more processing blocks in the decoder / virtualizer includes bypassing the sound field decoder and a plurality of HRTF blocks when the amplitude of the signal from the rotation / translation expression is not equal to or greater than the predetermined amplitude threshold, The sound field decoder and the plurality of HRTF blocks are included in the decoder / virtualizer, the method according to item 1. (Item 10) According to the determination that the amplitude of the signal from the rotation / translation expression is equal to or greater than the predetermined amplitude threshold, decoding the signal from the rotation / translation expression, determining a head-related transfer function (HRTF) and applying it to the decoded signal and further including, the method according to item 9. (Item 11) The plurality of virtual speakers includes inactive virtual speakers and active virtual speakers at a first time, and at least one of the active virtual speakers at the first time is designated as inactive at a second time while the signal is being processed, the method according to item 1. (Item 12) A system, the system comprising: A wearable head device configured to provide an audio signal to a user; A circuit configured to spatially render the audio signal; The circuit comprising: A spatial model configured to model a virtual environment; A spatial encoder configured to distribute signals from the spatial model across a plurality of virtual speakers; An internal spatial representation configured to represent a spatial configuration of the virtual environment; A decoder / virtualizer configured to decode signals from the internal spatial representation and introduce virtual sound into the decoded signals; The decoder / virtualizer comprising: A rotation / translation representation configured to introduce a representation of movement associated with the audio signal; An acoustic decoder configured to be configurable to decode signals from the rotation / translation representation; A plurality of head-related transfer function (HRTF) blocks configured to determine an HRTF corresponding to an input signal thereof and apply the corresponding HRTF to the input signal thereof; A plurality of combiners configured to combine signals from the plurality of HRTF blocks and output the audio signal; The system further comprising: A plurality of detectors configured to receive signals from the acoustic decoder and determine an energy level associated with the signals from the acoustic decoder; A plurality of first switches configured to pass the signals from the acoustic decoder to the plurality of HRTF blocks when the determined energy level is not less than an energy threshold; (Item 13) The system according to item 12, further comprising: The system according to item 12, further comprising: The system according to item 12, further comprising: (Item 14) Further comprising a second switch, said second switch receives the signal from the sound field decoder, and selectively passes the signal from the sound field decoder directly to the plurality of detectors or the plurality of HRTF blocks. The system according to item 13, which is configured to perform the above operations. (Item 15) Further comprising a sound field decoding decision, said sound field decoding decision determines whether the amplitude of the signal from the rotation / translation representation is greater than a predetermined amplitude threshold, and passes the signal from the rotation / translation representation to the sound field decoder according to the determination that the amplitude of the signal from the rotation / translation representation is greater than the predetermined amplitude threshold. The system according to item 12, which is configured to perform the above operations.

Brief Description of the Drawings

[0013]

Figure 1

[0014]

Figure 2

[0015]

Figure 3

[0016]

Figure 4

[0017]

Figure 5A

[0018]

Figure 5B

[0019]

Figure 5C

[0020]

Figure 6

[0021]

Figure 7A

[0022]

Figure 7B

[0023]

Figure 8A

[0024]

Figure 8B

[0025]

Figure 9

[0026]

Figure 10A

[0027]

Figure 10B

DETAILED DESCRIPTION OF THE INVENTION

[0028] In the following description of examples, reference is made to the accompanying drawings that form a part hereof and in which are shown, by way of illustration, specific examples that may be practiced. It is to be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.

[0029] (Exemplary Wearable System)

[0030] FIG. 1 illustrates an exemplary wearable head device 100 configured to be worn on a user's head. The wearable head device 100 can be part of a broader wearable system comprising one or more components such as a head device (e.g., the wearable head device 100), a handheld controller (e.g., the handheld controller 200 described below), and / or an auxiliary unit (e.g., the auxiliary unit 300 described below). In some examples, the wearable head device 100 can be used for virtual reality, augmented reality, or mixed reality systems or applications. The wearable head device 100 can include one or more displays such as displays 110A and 110B (left and right transmissive displays, with associated components for coupling light from the displays to the user's eyes such as orthogonal pupil expansion (OPE) grating sets 112A / 112B and exit pupil expansion (EPE) grating sets 114A / 114B), left and right acoustic structures such as speakers 120A and 120B (mounted on respective gooseneck arms 122A and 122B and positioned adjacent to the user's left and right ears), one or more sensors such as an infrared sensor, an accelerometer, a GPS unit, an inertial measurement unit (IMU) (e.g., IMU 126), an acoustic sensor (e.g., microphone 150), an orthogonal coil electromagnetic receiver (e.g., receiver 127 shown mounted on the left gooseneck arm 122A), left and right cameras directed away from the user (e.g., depth (time-of-flight) cameras 130A and 130B), and left and right eye cameras directed towards the user (e.g., for detecting the user's eye movements) (e.g., eye cameras 128 and 128B). However, the wearable head device 100 can incorporate any suitable display technology and any suitable number, type, or combination of sensors or other components without departing from the scope of the invention.In some examples, the wearable head device 100 may include one or more microphones 150 configured to detect an audio signal generated by the user's voice, such microphones may be positioned within the wearable head device adjacent to the user's mouth. In some examples, the wearable head device 100 may incorporate networking features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other wearable systems. The wearable head device 100 may further include components such as a battery, a processor, a memory, a storage unit, or various input devices (e.g., buttons, touch pads), or may be coupled to a handheld controller (e.g., handheld controller 200) or an auxiliary unit (e.g., auxiliary unit 300) that includes one or more such components. In some examples, the sensor may be configured to output a set of coordinates of the head-mounted unit relative to the user's environment, provide the input to a processor, and perform a simultaneous localization and mapping (SLAM) procedure and / or a visual odometry algorithm. In some examples, the wearable head device 100 may be coupled to the handheld controller 200 and / or the auxiliary unit 300, as further described below.

[0031] Figure 2 illustrates an exemplary mobile handheld controller component 200 of an exemplary wearable system. In some examples, the handheld controller 200 may communicate wired or wirelessly with the wearable head device 100 and / or the auxiliary unit 300 described below. In some examples, the handheld controller 200 includes a handle portion 220 to be held by a user and one or more buttons 240 disposed along the upper surface 210. In some examples, the handheld controller 200 may be configured for use as an optical tracking target; for example, sensors (e.g., cameras or other optical sensors) of the wearable head device 100 may be configured to detect the position and / or orientation of the handheld controller 200, which in turn may indicate the position and / or orientation of the hand of the user holding the handheld controller 200. In some examples, the handheld controller 200 may include one or more input devices such as a processor, memory, storage unit, display, or those described above. In some examples, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 100). In some examples, the sensors can detect the position or orientation of the handheld controller 200 relative to the wearable head device 100 or another component of the wearable system. In some examples, the sensors may be positioned within the handle portion 220 of the handheld controller 200 and / or may be mechanically coupled to the handheld controller. The handheld controller 200 may be configured to provide one or more output signals corresponding to, for example, the pressed state of the button 240 or the position, orientation, and / or movement (e.g., via an IMU) of the handheld controller 200. Such output signals may be used as inputs to the processor of the wearable head device 100, inputs to the auxiliary unit 300, or inputs to another component of the wearable system.In some examples, the handheld controller 200 can include one or more microphones to detect sound (e.g., the user's speech, ambient sound) and, in some cases, provide a signal corresponding to the detected sound to a processor (e.g., the processor of the wearable head device 100).

[0032] FIG. 3 illustrates an exemplary auxiliary unit 300 of an exemplary wearable system. In some examples, the auxiliary unit 300 can communicate wired or wirelessly with the wearable head device 100 and / or the handheld controller 200. The auxiliary unit 300 can include a battery to provide energy for operating one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including a display, sensors, acoustic structures, processors, microphones, and / or other components of the wearable head device 100 or the handheld controller 200). In some examples, the auxiliary unit 300 can include one or more sensors such as a processor, memory, storage unit, display, one or more input devices, and / or those described above. In some examples, the auxiliary unit 300 includes a clip 310 (e.g., a belt worn by the user) for attaching the auxiliary unit to the user. The advantage of using the auxiliary unit 300 to store one or more components of the wearable system is that doing so allows relatively large or heavy components to be carried on the user's waist, chest, or back, which is relatively well-suited for supporting large and heavy objects, rather than being mounted on the user's head (e.g., if stored within the wearable head device 100) or carried by the user's hand (e.g., if stored within the handheld controller 200). This can be particularly advantageous for relatively heavy or bulky components such as a battery.

[0033] Figure 4 shows an exemplary functional block diagram that may correspond to an exemplary wearable system 400, which may include, for example, the wearable head device 100 described above, the handheld controller 200, and the auxiliary unit 300. In some examples, the wearable system 400 may be used for virtual reality, augmented reality, or mixed reality applications. As shown in FIG. 4, the wearable system 400 may include an exemplary handheld controller 400B, herein referred to as a "totem" (and may correspond to the handheld controller 200 described above), and the handheld controller 400B may include a totem / headgear 6-degree-of-freedom (6DOF) totem subsystem 404A. The wearable system 400 may also include an exemplary wearable head device 400A (which may correspond to the wearable headgear device 100 described above), and the wearable head device 400A may include a totem / headgear 6DOF headgear subsystem 404B. In an example, the 6DOF totem subsystem 404A and the 6DOF headgear subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translational directions and rotations along three axes). The six degrees of freedom may be represented relative to the coordinate system of the wearable head device 400A. The three translational offsets may be represented as X, Y, and Z offsets, a translation matrix, or some other representation within such a coordinate system. The rotational degrees of freedom may be represented as a column, vector, rotation matrix, quaternion, or some other representation of yaw, pitch, and roll rotations. In some examples, one or more depth cameras 444 (and / or one or more non-depth cameras) included within the wearable head device 400A and / or one or more optical targets (e.g., buttons 240 of the handheld controller 200 as described above or dedicated optical targets included within the handheld controller) may be used for 6DOF tracking.In some examples, the handheld controller 400B can include a camera as described above, and the headgear 400A can include an optical target for optical tracking in conjunction with the camera. In some examples, each of the wearable head device 400A and the handheld controller 400B includes a set of three orthogonally oriented solenoids that are used to wirelessly transmit and receive three distinguishable signals. By measuring the relative magnitudes of the three distinguishable signals received at each of the coils used for reception, the 6DOF of the handheld controller 400B relative to the wearable head device 400A can be determined. In some examples, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that is useful for providing improved accuracy and / or more timely information regarding fast movement of the handheld controller 400B.

[0034] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space fixed with respect to the wearable head device 400A) to an inertial coordinate space or to an environmental coordinate space. For example, such a transformation may be necessary so that the display of the wearable head device 400A presents virtual objects at their expected positions and orientations with respect to the real environment, rather than at fixed positions and orientations on the display (e.g., at the same position on the display of the wearable head device 400A). For example, a virtual person sitting on a real chair facing forward, regardless of the position and orientation of the wearable head device 400A. This can maintain the illusion that the virtual objects are present within the real environment (and, for example, do not appear unnaturally positioned within the real environment as the wearable head device 400A shifts and rotates). In some examples, a compensating transformation between coordinate spaces can be determined by processing images from a depth camera 444 (e.g., using a Simultaneous Localization and Mapping (SLAM) and / or Visual Odometry procedure) to determine the transformation of the wearable head device 400A with respect to the inertial or environmental coordinate system. In the example shown in FIG. 4, the depth camera 444 can be coupled to a SLAM / Visual Odometry block 406 and provide the image to the block 406. The SLAM / Visual Odometry block 406 implementation can include a processor configured to process this image and then determine the position and orientation of the user's head, which can be used to identify the transformation between the head coordinate space and the real coordinate space. Similarly, in some examples, an additional source of information regarding the user's head pose and location is obtained from the IMU 409 of the wearable head device 400A. Information from the IMU 409 can be integrated with information from the SLAM / Visual Odometry block 406 to provide more timely information regarding improved accuracy and / or faster adjustment of the user's head pose and position.

[0035] In some examples, the depth camera 444 can supply a 3D image to a hand gesture tracker 411 that can be implemented within the processor of the wearable head device 400A. The hand gesture tracker 411 can identify a user's hand gesture, for example, by matching the 3D image received from the depth camera 444 to a stored pattern representing the hand gesture. Other suitable techniques for identifying a user's hand gesture will also be apparent.

[0036] In some examples, one or more processors 416 may be configured to receive data from the headgear subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, microphone (not shown), and / or hand gesture tracker 411. The processor 416 can also send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A in examples where the handheld controller 400B is not connected. The processor 416 may further communicate with additional components such as the audiovisual content memory 418, graphical processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to the head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to the left source of the light 424 modulated per image and a right channel output coupled to the right source of the light 426 modulated per image. The GPU 420 can output stereoscopic image data to the sources of the light 424, 426 modulated per image. The DSP audio spatializer 422 can output audio to the left speaker 412 and / or right speaker 414. The DSP audio spatializer 422 can receive from the processor 416 an input indicating a direction vector from the user to a virtual sound source (e.g., movable by the user via the handheld controller 400B). Based on the direction vector, the DSP audio spatializer 422 can determine the corresponding HRTF (e.g., by accessing the HRTF or interpolating multiple HRTFs). The DSP audio spatializer 422 can then apply the determined HRTF to an audio signal such as an audio signal corresponding to the virtual sound generated by the virtual object.This can improve the credibility and realism of virtual sounds by incorporating the user's relative position and orientation with respect to the virtual sounds in the mixed reality environment, i.e., by presenting virtual sounds that match the user's expectations of what the virtual sounds would sound like if they were real sounds in the real environment.

[0037] In some examples, such as those shown in FIG. 4, one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 may be included within an auxiliary unit 400C (which may correspond to the auxiliary unit 320 described above). The auxiliary unit 400C includes a battery 427, which can power its components and / or which can supply power to the wearable head device 400A and / or the handheld controller 400B. Incorporating such components within an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which can in turn reduce fatigue in the user's head and neck.

[0038] FIG. 4 presents elements corresponding to various components of the exemplary wearable system 400, although various other suitable arrangements of these components will be apparent to those skilled in the art. For example, the elements presented in FIG. 4 as being associated with the auxiliary unit 400C may instead be associated with the wearable head device 400A or the handheld controller 400B. Further, some wearable systems may eliminate the handheld controller 400B or the auxiliary unit 400C entirely. Such changes and modifications should be understood to be included within the scope of the disclosed examples.

[0039] (Mixed Reality Environment)

[0040] Like all people, users of a composite reality system exist within a physical environment, i.e., within the three-dimensional portion of the "real world" that is perceptible by the user and all of its contents. For example, the user perceives the physical environment using its normal human senses, i.e., vision, hearing, touch, taste, and smell, and interacts with the physical environment by moving its own body within the physical environment. Locations within the physical environment can be described as coordinates within a coordinate space, e.g., the coordinates can include latitude, longitude, and altitude relative to sea level, distances in three orthogonal dimensions from a reference point, or other suitable values. Similarly, a vector can describe a quality having a direction and magnitude within the coordinate space.

[0041] A computing device can maintain a representation of a virtual environment, for example, in a memory associated with the device. As used herein, a virtual environment is a computer representation of a three-dimensional space. The virtual environment can include representations of any object, action, signal, parameter, coordinate, vector, or other characteristic associated with that space. In some examples, a circuit (e.g., a processor) of the computing device can maintain and update the state of the virtual environment, i.e., the processor can determine the state of the virtual environment at a second time based on data associated with the virtual environment and / or input provided by a user at a first time. For example, if an object within the virtual environment is located at a first coordinate at a certain time, has certain programmed physical parameters (e.g., mass, coefficient of friction), and input received from a user indicates that a force should be applied to the object in a certain direction vector, the processor can apply the laws of kinematics and use basic mechanics to determine the location of the object at that time. The processor can use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at a given time. In maintaining and updating the state of the virtual environment, the processor can execute any suitable software, and any suitable software can include software related to the creation and deletion of virtual objects within the virtual environment, software (e.g., scripts) for defining the behavior of virtual objects or characters within the virtual environment, software for defining the behavior of signals (e.g., audio signals) within the virtual environment, software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling input and output, software for implementing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.

[0042] Output devices such as displays or speakers can present any or all aspects of a virtual environment to the user. For example, the virtual environment can include virtual objects (which can include representations of inanimate objects, people, animals, light, etc.) that can be presented to the user. The processor can determine the display of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, line of sight, and frustum), and render a visible scene of the virtual environment corresponding to that display on the display. Any suitable rendering technique can be used for this purpose. In some examples, the visible scene can include only some of the virtual objects within the virtual environment and exclude other virtual objects. Similarly, the virtual environment can include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects within the virtual environment can generate sounds arising from the location coordinates of the object (e.g., a virtual character can speak or cause sound effects); or the virtual environment can be associated with musical cues or ambient sounds that may or may not be associated with a particular location. The processor can determine an audio signal corresponding to "listener" coordinates (e.g., an audio signal that simulates the audio signal that would be heard by a listener at the listener coordinates, mixed and processed to correspond to the composite of the sounds within the virtual environment), and present the audio signal to the user via one or more speakers.

[0043] Since the virtual environment only exists as a computer construct, a user cannot directly perceive the virtual environment using their normal senses. Instead, the user can only indirectly perceive the virtual environment as presented to the user, for example, by way of a display, speakers, a tactile output device, and the like. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment, but can provide input data to a processor that can use the device or sensor data to update the virtual environment via an input device or sensor. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object within the virtual environment, and the processor can use that data to cause the object to respond accordingly within the virtual environment.

[0044] (Digital Reverberation and Environmental Audio Processing)

[0045] An XR system can present an audio signal to a user that appears to originate from a sound source with origin coordinates and travel in the direction of a direction vector in the system. The user can perceive these audio signals as if they were real audio signals originating from the origin coordinates of the sound source and traveling along the direction vector.

[0046] In some cases, the audio signals can be considered virtual in that they correspond to computer signals within the virtual environment and do not necessarily correspond to real sounds in the real environment. However, the virtual audio signals can be presented to the user as real audio signals detectable by a human ear, for example, as being generated via speakers 120A and 120B of the wearable head device 100 in FIG. 1.

[0047] The advantages of the embodiments disclosed below include reduced network bandwidth, reduced power consumption, reduced computational complexity, and reduced computational latency. These advantages can be particularly prominent in mobile systems, including wearable systems, where processing resources, networking resources, battery capacity, and physical size and weight are often limited.

[0048] In an environment as dynamic as AR, the system can continuously render an audio signal. Rendering an audio signal using all of the virtual speakers can particularly lead to high computational power, a large amount of processing, high network bandwidth, high power consumption, etc. Therefore, it may be desirable to use modified virtual speaker panning to dynamically select and use a portion of the fixed virtual speakers based on one or more factors.

[0049] (Exemplary Spatial Audio System)

[0050] FIG. 5A illustrates a block diagram of an exemplary spatial audio system according to some embodiments. FIG. 5B illustrates a flow of an exemplary method for operating the system of FIG. 5A.

[0051] The spatial audio system 500 may include a spatial model 510, an internal space representation 530, and a decoder / virtualizer 540A. The spatial model 510 may include a direct path portion 512, one or more reflection portions 520 (optional), and a spatial encoder 526. The spatial model 510 may be configured to model a virtual environment. The direct path portion 512 may include a direct source 514 and optionally a Doppler 516. The direct source 514 may be configured to provide an audio signal (step 552 of process 550). The Doppler 516 may receive a signal from the direct source 514 and may be configured to introduce a Doppler effect into the input signal (step 554). For example, the Doppler 516 may change the pitch of the sound source (e.g., pitch shift) to vary with the movement of the sound source, the user of the system, or both.

[0052] The reflection portion 520 may include a sound reflector 522, an optional Doppler 516, and a delay 524. The sound reflector 522 may be configured to introduce a reflection into its signal (step 556). The introduced reflection may represent one or more characteristics of the environment. The Doppler 516 within the reflection portion 520 may receive a signal from the sound reflector 522 and may be configured to introduce a Doppler effect into its input signal (step 558). The delay 524 may receive a signal from the Doppler 516 and may be configured to introduce a delay (step 560).

[0053] The spatial encoder 526 may receive signals from the direct path portion 512 and the reflection portion 520. In some embodiments, the signal from the direct path portion 512 to the spatial encoder 526 may be the output signal from the Doppler 516 of the direct path portion 512. In some embodiments, the signal from the reflection portion 520 to the spatial encoder 526 may be the output signal from the delay 524 of the reflection portion 520.

[0054] The spatial encoder 526 may include one or more M-direction pans 528. In some embodiments, each input received by the spatial encoder 526 may be associated with a unique 528. "Panning" may refer to distributing a signal across multiple speakers, multiple locations, or both. The M-direction pan 528 may be configured to distribute its input signal across a plurality of virtual speakers (step 562). For example, the M-direction pan 528 may be able to distribute its input signal across all M virtual speakers. For example, as shown in FIG. 5A, M may be equal to 4, and each M-direction pan 528 may be configured to distribute its input signal across four virtual speakers. The figure illustrates a system having four virtual speakers, but the examples of the present disclosure may include any number of virtual speakers.

[0055] As an example, an automotive system may include left and right speakers. Sound in such a system can be panned between the left and right speakers in the automobile by splitting the sound into one, two for each speaker. The scaling volume of each speaker can be set according to the configuration of the two speakers, and the result can be sent to the left and right speakers.

[0056] As another example, a surround sound system may include a plurality of speakers, such as six speakers. Sound in such a system can be panned as stereo between the six speakers. The sound can be split into six (instead of two as in the automotive system example), the scaling volume of each speaker can be set according to the configuration of the six speakers, and the result can be sent to the six speakers.

[0057] For example, a first M-direction pan 528 may receive the output of the Doppler 516 of the direct path 512, and other M-direction pans 528 may receive the output of the reflected portion 520. Each M-direction pan 528 can split its input signal so that it can be distributed across multiple outputs. Thus, each M-direction pan 528 can have a number of outputs greater than the input.

[0058] The spatial modeler 510 may output a signal to the internal space representation 530 (step 564). In some embodiments, the output from the spatial modeler 510 can include the output of each M-direction pan 528. The internal space representation 530 may be configured to represent the spatial configuration of the virtual environment (step 566). An exemplary representation can include representing the relative locations of the user, sound sources, and virtual speakers. In some embodiments, the internal space representation 530 may output one or more signals representing the head pose rotation, head pose translation, sound field decoding, one or more head-related transfer functions (HRTFs), or combinations thereof of a user of the system 500. In some embodiments, the internal space representation 530 can be a representation such as a non-ambisonic multi-channel-based system, an ambisonics / wave field-based system, etc. An exemplary ambisonics / wave field-based system can be higher-order ambisonics (HOA).

[0059] The internal space representation 530 may output that signal 552 to the decoder / virtualizer 540A (step 568). The decoder / virtualizer 540 may decode the input signal and introduce virtual sounds into the signal (step 570). Step 570 can include a plurality of sub-steps, which will be discussed in more detail below. The system may then output the signal from the decoder / virtualizer 540 as a left signal 502L that can be output to the left speaker and a right signal 502R that can be output to the right speaker (step 580).

[0060] The system 500 may include any number of different types of decoders / virtualizers 540. An exemplary decoder / virtualizer 540A is shown in FIG. 5A. Other exemplary decoders / virtualizers 540 will be discussed below.

[0061] The decoder / virtualizer 540A may include a rotation / translation representation 542, a sound field decoder 544, one or more HRTFs 546, and one or more combiners 548. FIG. 5C illustrates a flow of an exemplary method for operating an exemplary decoder / virtualizer, which may be referred to as step 570-1. The rotation / translation representation 542 may receive a signal from the internal space representation 530 and may be configured to introduce a representation of movement associated with the audio signal. For example, the movement may be the sound source, the user, or both (step 572). The rotation / translation representation 542 can output the signal to the sound field decoder 544. The sound field decoder 544 may receive a signal from the rotation / translation representation 542 and may be configured to decode the signal (step 574). Each HRTF 546 may receive a signal from the sound field decoder 544. Each HRTF 546 may be configured to determine the HRTF corresponding to its input signal and apply it to the signal (step 576). One or more HRTFs 546 may be collectively referred to as a speaker virtualizer. In some embodiments, the HRTF 546 may be configured for finite impulse response (FIR) filtering. Each combiner 548 may receive and combine signals from the HRTF 546 (step 578).

[0062] In some embodiments, the decoder / virtualizer 540A may represent a "baseline" processing overhead. The baseline processing overhead is complex and may involve matrix calculations and long FIR filters for applying HRTF processing for each virtual speaker.

[0063] The output from the combiner 548 may be the output signal from the system 500. In some embodiments, the output signal 502 from the system 500 may be an audio signal for the left and right speakers (e.g., speakers 120A and 120B of FIG. 1).

[0064] In some instances, when the number of sound sources for playback is large, the spatial audio system of FIG. 5A can be beneficial. However, in some instances, when the number of sound sources for playback is small, the spatial audio system of FIG. 5A may not be beneficial. It may be desirable to utilize the efficiency of a non-ambisonic multi-channel based spatial audio system such as system 500 of FIG. 5A or an ambisonics-based spatial audio system in an efficient way for situations when the number of sound sources for playback is small.

[0065] There may exist a method for improving the efficiency of spatialization using sound field synthesis and decoding. A first method may be through low energy speaker detection and killing. In low energy speaker detection and killing, if the energy output of a virtual speaker channel of a non-ambisonic multi-channel based spatial audio system or an ambisonics / sound field channel of an ambisonics-based spatial audio system is less than a predetermined threshold, the processing of the signal from the virtual speaker channel is not performed. In some embodiments, the system may determine, for example, whether the output of a given virtual speaker is greater than a predetermined threshold before sound field decoding is performed on the signal from that given virtual speaker. Low energy speaker detection and killing is discussed in more detail below.

[0066] A second method for improving the efficiency of spatialization using sound field synthesis and decoding may be source geometry-based virtual speaker killing. In source geometry-based virtual speaker killing, decoder / virtualizer processing can be selectively disabled. The selective disabling (or selective enabling) can be based on the location of the sound source relative to the user / listener. Source geometry-based virtual speaker killing is discussed in more detail below.

[0067] A third method may be to combine low energy speaker detection and killing techniques with source-virtual speaker coupling techniques.

[0068] The spatial model 510 may have a computational complexity that can represent the number of operations required to process an audio signal. The computational complexity may be proportional to M multiplied by N, where M may be equal to the number of sound sources (including direct sources and optional reflections), and N may be equal to the number of channels required to represent an ambisonic sound field. In some embodiments, N may be equal to (O + 1) 2 where O is the order of the ambisonics used.

[0069] The decoder / virtualizer 540 may have a computational complexity proportional to nVS, where nVS is the number of virtual speakers. The computational ability of each speaker may be high, and it may generally consist of a pair of FIR filters, which are typically implemented using a fast Fourier transform (FFT) or an inverse FFT (IFFT), and both of which may be computationally expensive processes.

[0070] (Exemplary Low-Energy Output Detection and Culling Method)

[0071] In some embodiments, some virtual speakers may have little or no signal input energy: for example, when a spatial audio system has a small number of sound sources. The speaker virtualization process can be a computationally expensive (e.g., CPU-intensive) process. For example, when a sound source is located at the 0-degree azimuth (e.g., directly in front of the user), there may be little or no energy in the signals from virtual speakers located at 90 degrees to 270 degrees azimuth (e.g., behind the user). Low-energy signals may also have no significant effect on the perceived location of the sound source, and thus, performing the speaker virtualization process on low-energy signals and / or determining the characteristics of the corresponding virtual speakers can be computationally inefficient.

[0072] To reduce the required computational resources, a system that employs a low-energy output detection and curving method can include a detector positioned between an acoustic field decoder and an HRTF. Alternatively, the detector can be positioned between a multi-channel output and the HRTF. The detector can be configured to detect one or more energy levels associated with one or more audio signals from one or more virtual speakers.

[0073] If the energy level of a signal emitted from virtual speaker Vn is less than the energy threshold α, the signal can be considered a low-energy signal. In accordance with the detected energy level associated with the audio signal being less than the energy threshold α, the HRTF block and its processing of the low-energy signal can be bypassed.

[0074] Determination of the energy level of a signal can use any number of techniques. For example, an RMS algorithm can be applied to the signal routed to the virtual speaker to measure its energy. Similar "attack" and "release" times used by conventional audio compressors can be used to prevent the signal of the speaker from suddenly "popping in" and "popping out".

[0075] FIG. 6 illustrates an exemplary configuration of a sound source and speakers according to some embodiments. System 600 may include a sound source 620 and a plurality of speakers. The plurality of speakers 622 may include one or more active virtual speakers 622A and one or more non-active virtual speakers 622B. The signal of the active virtual speaker 622A may be processed by the HRTF 546 at a given time. The non-active virtual speaker 622B may be such that its signal does not need to be processed by the HRTF 546, for example, because its signal has already been processed at a previous time or because the system has determined that the signal from the virtual speaker 622B does not require processing. M may refer to the number of sound sources to be reproduced, and N may refer to the number of virtual speakers in the system. The figure illustrates a single sound source, but the examples of the present disclosure can include any number of sound sources. The figure illustrates eight sound sources, but the examples of the present disclosure can include any number of sources such as 16 (N = 16).

[0076] As an example, system 600 can include a single (M = 1) sound source 620 and eight virtual speakers 622, as shown in the figure. In a given instance, most of the energy can be output over only three virtual speakers. That is, system 600 can have three active virtual speakers at a first time. For example, virtual speakers 622A-1, 622A-2, and 622-3 can be active virtual speakers. In some embodiments, the active virtual speakers 622A can be those closest to the sound source 620. Additionally, system 600 can include five non-active virtual speakers 622B. System 600 can determine that the energy level from each of the five non-active virtual speakers is less than an energy threshold, and in accordance with such a determination, can bypass the HRTF processing of the signals from the five non-active virtual speakers 622B.

[0077] System 600 may also determine that the energy level from each of the active virtual speakers is not less than the energy threshold, and in accordance with such determination, may perform HRTF processing on the signals from the three active virtual speakers 622A.

[0078] As shown in FIG. 5A, system 600 may output two signals, i.e., one for the right speaker (such as right signal 502R and left signal 502L) and one for the left speaker. The reduction in the number of HRTF operations by bypassing the HRTF processing may be equal to the number of non-active virtual speakers multiplied by the number of signals output from the system. In the example of FIG. 6, since the HRTF processing of five signals is bypassed, 10 times (5 non-active virtual speakers × 2 output signals) of HRTF operations can be saved.

[0079] As another example, if the system includes 16 virtual speakers where 13 are non-active virtual speakers, the number of HRTF operations saved may be equal to 26 times (16 virtual speakers × 2 output signals).

[0080] FIG. 7A illustrates a block diagram of an exemplary decoder / virtualizer including a plurality of detectors according to some embodiments. FIG. 7B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 7A according to some embodiments. In some embodiments, as discussed below, instead of decoder / virtualizer 540A (shown in FIG. 5A), decoder / virtualizer 540B may be included within system 500. Instead of step 570-1 (shown in FIG. 5C), step 570-2 may be included within process 550.

[0081] The decoder / virtualizer 540B can include a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more switches 712, one or more HRTFs 546, and one or more combiners 548. The decoder / virtualizer 540B can receive a signal 552 from an internal space representation 530 (such as shown in FIG. 5A). The rotation / translation representation 542 can receive a signal from the internal space representation 530 and can be configured to introduce a representation of the movement of the sound source, the user, or both (step 772). The rotation / translation representation 542 can output the signal to the sound field decoder 544. The sound field decoder 544 can receive a signal from the rotation / translation representation 542 and can be configured to decode the signal (step 774). The sound field decoder 544 can output the signal to the detector 710.

[0082] The detector 710 can receive a signal from the sound field decoder 544 and can be configured to determine the energy level of the input signal (step 776). Each detector 710 can be coupled to a respective switch 712. If the energy level of the input signal (from the sound field decoder 544) is greater than or equal to an energy threshold (step 778), the switch 712 can close the loop, thereby routing the input signal (from the detector 710) to the HRTF 546 to which the switch is coupled (step 780). Each HRTF determines the corresponding HRTF and applies it to the signal (step 782).

[0083] If the energy level of the input signal is less than the energy threshold, the switch 712 can open so that the input signal (from the detector 710) is not coupled to the corresponding HRTF 546. Thus, the corresponding HRTF 546 can be bypassed (step 784).

[0084] The signal from HRTF546 can be output to combiner 548 (step 786). Combiner 548 can be configured to combine (e.g., add, aggregate, etc.) the signals from HRTF546. Those signals that bypass HRTF546 cannot be combined by combiner 548. The output from combiner 548 can be the output signal from system 500. In some embodiments, the output signal 502 from system 500 can be an audio signal for the left and right speakers (e.g., speakers 120A and 120B of FIG. 1).

[0085] In some embodiments, each detector 710 can be coupled to a unique signal corresponding to a virtual speaker. In this way, the processing of each virtual speaker 622 can be performed independently (i.e., the processing of one speaker such as 622A-1 can be performed without affecting the processing of another speaker such as 622B).

[0086] In some embodiments, the type of decoder / virtualizer 540 can depend on the number of sound sources. For example, if the number of sound sources is less than or equal to a predetermined sound source threshold, the decoder / virtualizer 540B of FIG. 7A can be included within system 500. In such an instance, the signal from sound field decoder 544 can be input to detector 710.

[0087] If the number of sound sources is greater than a predetermined sound source threshold, the decoder / virtualizer 540A of FIG. 5A can be included within the system. In such an instance, the signal from sound field decoder 544 can be input to HRTF546.

[0088] In some embodiments, the system may include a decoder / virtualizer 540 that can select whether to perform detector and its energy level detection or bypass it. FIG. 8A illustrates a block diagram of an exemplary decoder / virtualizer according to some embodiments. FIG. 8B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 8A according to some embodiments. In some embodiments, instead of decoder / virtualizer 540A (shown in FIG. 5A) and decoder / virtualizer 540B (shown in FIG. 7A), decoder / virtualizer 540C may be included within system 500. Instead of step 570-1 (shown in FIG. 5C), step 570-3 may be included within process 550.

[0089] Similar to decoder / virtualizer 540B discussed above, decoder / virtualizer 540C can include a rotation / translation representation 542, an acoustic field decoder 544, one or more detectors 710, one or more first switches 712, one or more HRTFs 546, and one or more combiners 548. Steps 872, 874, and 882 may be similar and corresponding to steps 772, 774, and 782 discussed above.

[0090] Decoder / virtualizer 540C may also include a second switch 814. The second switch 814 can be configured to open or close a first loop from the acoustic field decoder 544 to the detector 710 and the first switch 712. Additionally, or alternatively, the second switch 814 can be configured to open or close a second loop from system 500 that bypasses the detector 710 and the first switch 712. In some embodiments, the second switch 814 can be a bidirectional switch configured to select between passing the signal directly to the detector 710 (first loop) or directly to the HRTF 546 (second loop).

[0091] For example, the system can determine whether the number of sound sources is greater than or equal to a predetermined sound source threshold (step 876). If the number of sound sources is greater than or equal to the predetermined sound source threshold, the second switch 814 can close the second loop and directly pass the signal from the sound field decoder 544 to the HRTF 546 (step 878). Each HRTF 546 then determines the corresponding HRTF and applies it to the signal (step 880). When the number of sound sources exceeds a certain number, the possibility that the signal has a low energy level can be reduced.

[0092] On the other hand, if the number of sound sources is less than the predetermined sound source threshold, the signal is likely to have a low energy level. Therefore, the second switch 814 can close the first loop and directly pass the signal from the sound field decoder 544 to the detector 710 (step 882). The detector 710 can receive the signal from the sound field decoder 544 and can be configured to determine the energy level of the input signal (step 884). If the energy level of the input signal (from the sound field decoder 544) is greater than or equal to the energy threshold (step 886), the switch 712 can close the loop, thereby routing the input signal (from the detector 710) to the HRTF 546 to which the switch is connected (step 888). If the energy level of the input signal is less than the energy threshold, the switch 712 can open so that the input signal (from the detector 710) is not connected to the corresponding HRTF 546, allowing the HRTF 546 to be bypassed (step 890).

[0093] The signal from the HRTF 546 can be output to the combiner 548 (step 892).

[0094] In some embodiments, one or more energy threshold detections can be active in response to energy. In some embodiments, one or more energy threshold detections can be active in response to amplitude and can receive conventional attack, release times, etc.

[0095] (Example source geometry-based speaker culling method)

[0096] Source geometry-based virtual speaker culling can be another way to reduce CPU consumption. In some embodiments, source geometry-based virtual speaker culling can include selectively disabling decoder / virtualizer processing (e.g., decoder / virtualizer 540A of FIG. 5A, decoder / virtualizer 540B of FIG. 7A, decoder / virtualizer 540C of FIG. 8A, etc.). In some embodiments, the selective disabling (or selective enabling) can be based on the location of the sound source relative to the user / listener. In some embodiments, the selective disabling of decoder / virtualizer processing can include the step of bypassing all of the processing blocks of the decoder / virtualizer.

[0097] In source geometry-based virtual speaker culling, an ambisonic output can be calculated. If the ambisonic output requires a significant amount of energy to be decoded, it may be beneficial to use a simpler method (requiring less CPU consumption) such as a real-time energy detection method. Additionally, in some embodiments, the real-time energy detection method can perform calculations at a lower frequency.

[0098] FIG. 9 illustrates an exemplary configuration of a sound source and speakers according to some embodiments. System 900 can include a sound source 920 and a plurality of speakers. Compared to system 600 of FIG. 6, the sound source 920 can be located at a second position different from the first position of the sound source 620 of FIG. 6. The plurality of speakers 922 can include one or more active virtual speakers 922A, one or more inactive virtual speakers 922B, and one or more inactive virtual speakers 922C. The active virtual speaker 922A and the inactive virtual speaker 922B can be similar corresponding to the active virtual speaker 622A and the inactive virtual speaker 622B of FIG. 6, respectively.

[0099] The inactive virtual speaker 922C may differ from the inactive virtual speaker 922B in that the virtual speaker 922C is active at a first time, but its signal is being processed at a second time (e.g., a ringout period). In the example of FIG. 9, the sound source 920 may be moving from a first position (e.g., close to the virtual speaker 922C) to a second position (e.g., not close to the virtual speaker 922). Due to the movement of the sound source, the two virtual speakers may no longer have a sound source mixing into them at the second time. Due to the filtering of the two virtual speakers, the two virtual speakers may need to be active for a subsequent frame (e.g., the second time) in order to properly complete the filtering process.

[0100] In some embodiments, the system may include a decoder / virtualizer 540 within a system that uses active virtual speakers. FIG. 10A illustrates a block diagram of an exemplary decoder / virtualizer used in a system that includes active speakers, according to some embodiments. FIG. 10B illustrates a flow of an exemplary method for operating the decoder / virtualizer of FIG. 10A, according to some embodiments. In some embodiments, instead of the decoder / virtualizer 540A (shown in FIG. 5A), the decoder / virtualizer 540B (shown in FIG. 7A), and the decoder / virtualizer 540C (shown in FIG. 8A), a decoder / virtualizer 540D may be included within the system 500. Instead of step 570-1 (shown in FIG. 5C), step 570-2 (shown in FIG. 7B), and step 570-3 (shown in FIG. 8B), a step 570-4 may be included within the process 550.

[0101] Similar to the decoder / virtualizer 540B and decoder / virtualizer 540C discussed above, the decoder / virtualizer 540C can include a sound field decoder 544, one or more HRTFs 546, and one or more combiners 548. Steps 1072, 1076, 1078, and 1080 may be similar and correspond to steps 872, 874, and 782 discussed above.

[0102] The decoder / virtualizer 540D may also include a rotation / translation representation 1042 and a sound field decoding decision 1044. The rotation / translation representation 1042 may receive a signal from the internal space representation 530 and may be configured to introduce a representation of the movement of the sound source, the user, or both (step 1072). The representation of the movement may also take into account the azimuth / elevation of the sound source 920. The rotation / translation representation 542 can output the signal to the sound field decoder decision 1044.

[0103] The sound field decoder decision 1044 may receive a signal from the rotation / translation representation 1042, determine signals with "significant" outputs, and may be configured to pass those signals to the sound field decoder 544 (step 1074). A significant output may be an output that will affect the perceived sound. For example, a significant output may be an audio signal having an amplitude above a predetermined amplitude threshold. The sound field decoder 544 may receive a signal from the sound field decoder decision 1044 having a significant output and may be configured to decode the signal (step 1076). In some embodiments, the sound field decoder 1044 may receive a signal from the sound field decoder decision 1044 having a significant output. Each HRTF 546 may receive a signal from the sound field decoder 544. Each HRTF 546 may be configured to determine the HRTF corresponding to its input signal and apply it to the signal (step 1078). The one or more HRTFs 546 may be collectively referred to as speaker virtualizers. Each combiner 548 may receive signals from the HRTFs 546 and combine them (step 1080).

[0104] In some embodiments, those audio signals that do not have a significant output (e.g., have an amplitude less than a predetermined amplitude threshold) may not be passed to the sound field decoder 544. Thus, the sound field decoder 544 and the HRTF 546 on audio signals that do not have a significant output can be bypassed.

[0105] An exemplary source geometry-based speaker culling method can specify virtual speakers as if they were active virtual speakers based on the position of the sound source (e.g., X, Y, Z location). The location of the sound source can represent the location of the source object. The system can determine the location of each sound source and determine virtual speakers located proximate to each sound source. In some embodiments, the determination of virtual speakers located proximate to the sound source can be performed, for example, at the start of each video frame (in a video frame rate-based approach). The video frame rate-based approach may require fewer calculations than other approaches such as a sample rate-based approach.

[0106] The sound source can contribute significantly to a particular virtual speaker, for example, based on a video frame rate-based approach calculation and an ambisonic decoding scheme. As discussed above, virtual speakers that contribute little or no energy when decoded can have the corresponding ambisonic decoding and HRTF processing of the decoded ambisonic channels bypassed. In some embodiments, the system can disable any processing blocks that are bypassed.

[0107] Exemplary pseudocode for implementing the specification method can be as follows: For each sound source, S and decode channel n Enable[n] |= f(sourcePosition Vector3, sourceOrientation Vector3, ListenerPosition Vector3, ListenerOrientation Vector3, VirtualSpeakerPosition[n] Vector3). (Ambisonic / Sound Field Example) For each Ambisonic Decode Channel If (Enable[n]) { AmbisonicDecode(n) Virtualize(n) } (Multi-Channel Example) For each Channel If (Enable[n]) { Virtualize(n) }

[0108] Regarding the above pseudo-code, the variable sourcePosition may refer to the position of the sound source, sourceOrientation may refer to the orientation of the sound source, ListenerPosition may refer to the position of the user / listener, ListenerOrientation may refer to the orientation of the user / listener, VirtualSpeakerPosition may refer to the position of the virtual speaker, AmbisonicDecode may refer to a function that performs Ambisonic decoding, and Virtualize may refer to a function that performs virtualization.

[0109] Regarding the above pseudo-code, for each sound source S and decode channel n, the decode channel n may be enabled based on one or more factors such as the position of the sound source S, the orientation of the sound source S, the position of the user / listener, the orientation of the user / listener, and the position of the virtual speaker. Still referring to the above pseudo-code, for each Ambisonic decode channel, if the channel is enabled, the system may execute the AmbisonicDecode function and the Virtualize function.

[0110] The pseudo-code can be enhanced by providing a "ring-out" period for each virtual speaker. For example, if the source has moved in position within a video frame, it may be determined that the virtual speaker no longer has any sound sources to mix into it. However, due to the filtering of the virtual speaker, that virtual speaker may need to be an active speaker for subsequent frames in order to properly complete the filtering process.

[0111] Examples of the present disclosure can include determining a decoded sound field output that uses all active sound sources and has a "significant" output (e.g., an output that will affect the perceived sound field). Ambisonics or non-ambisonics multi-channel outputs that will affect the perceived sound field can be decoded. Further, in some embodiments, only the HRTF546 corresponding to those detected outputs is processed. When the number of sound sources is small or, although numerous, they are close to each other, significant CPU savings can be achieved for synthetically generated ambisonic sound fields or non-ambisonics multi-channel renderings.

[0112] (Exemplary method combination of source geometry-based virtual speaker culling method and low-energy output detection and culling method)

[0113] In some embodiments, both source geometry-based virtual speaker culling and low-energy output detection and culling can be used continuously to further reduce CPU consumption. As described above, source geometry-based virtual speaker culling can include, for example, selectively disabling virtual speaker processing based on the location of the sound source relative to the user / listener. Low-energy output detection and culling can include, for example, installing a signal energy / level detector between sound field decoding or multi-channel output and HRTF processing. The output / result of source geometry-based virtual speaker culling can be input into low-energy output detection and culling.

[0114] Regarding the systems and methods described above, the elements of the systems and methods can be implemented, as appropriate, by one or more computer processors (e.g., a CPU or DSP). The present disclosure is not limited to any particular configuration of computer hardware that includes computer processors used to implement these elements. In some cases, multiple computer systems can be employed to implement the systems and methods described above. For example, a first computer processor (e.g., the processor of a wearable device coupled to a microphone) can be utilized to receive input microphone signals and perform initial processing of those signals (e.g., signal conditioning and / or segmentation such as that described above). A second (presumably more computationally powerful) processor can then be utilized to perform more computationally intensive processing such as determining probability values associated with the utterance segments of those signals. Another computer device such as a cloud server can host the speech recognition engine and the input signals are ultimately provided thereto. Other suitable configurations will also be apparent and are within the scope of the present disclosure.

[0115] While the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. For example, the elements of one or more implementations can be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications should be understood to be included within the scope of the disclosed examples as defined by the appended claims.

Claims

1. A method for spatially rendering an audio signal, the method comprising: determining a model of a virtual environment; determining a spatial configuration of the virtual environment, the spatial configuration comprising at least a user location, a sound source location, and a virtual speaker location; determining one or more signals associated with the spatial configuration and further associated with the user location, the sound source location, or the virtual speaker location; determining whether an amplitude of one or more signals corresponding to the sound source in the virtual environment exceeds a predetermined amplitude threshold; decoding the one or more signals according to the determination that the amplitude of the one or more signals exceeds the predetermined amplitude threshold; rendering the audio signal based on the one or more signals A method comprising.

2. A method for spatially rendering an audio signal, the method comprising: determining a model of a virtual environment; determining a spatial configuration of the virtual environment, the spatial configuration comprising at least a user location, a sound source location, and a virtual speaker location; determining one or more signals associated with the spatial configuration and further associated with the user location, the sound source location, or the virtual speaker location; determining whether one or more signals corresponding to the sound source in the virtual environment exceed a predetermined threshold; decoding the one or more signals according to the determination that the one or more signals exceed the predetermined threshold, wherein decoding the one or more signals includes performing a first set of one or more processing blocks; selectively bypassing a second set of one or more processing blocks, the second set of one or more processing blocks being associated with one or more inactive virtual speakers; rendering the audio signal based on the one or more signals A method comprising.

3. A method for spatially rendering an audio signal, the method comprising: determining a model of a virtual environment; determining a spatial configuration of the virtual environment, the spatial configuration comprising at least a user location, a sound source location, and a virtual speaker location; Determining one or more signals associated with the spatial configuration and further associated with the user location, the sound source location, or the virtual speaker location; Determining whether one or more signals corresponding to the sound source in the virtual environment exceed a predetermined threshold; Decoding the one or more signals according to the determination that the one or more signals exceed the predetermined threshold; Determining an energy level associated with the one or more signals; Determining whether the energy level is less than an energy threshold; Performing head-related transfer function (HRTF) processing on the one or more signals according to the determination that the energy level is not less than the energy threshold; Suspending the performance of the HRTF processing on the one or more signals according to the determination that the energy level is less than the energy threshold; Rendering the audio signal based on the one or more signals A method comprising.

4. Further comprising determining whether the number of sound sources in the virtual environment exceeds a predetermined sound source threshold, The method according to claim 2, wherein the selective bypass of the one or more processing blocks of the second set includes bypassing a plurality of detectors according to the determination that the number of sound sources exceeds the predetermined sound source threshold.

5. The method according to claim 4, further comprising detecting an energy level of the one or more signals using the plurality of detectors according to the determination that the number of sound sources does not exceed the predetermined sound source threshold.

6. Determining whether the energy level is less than an energy threshold; Performing head-related transfer function (HRTF) processing on the one or more signals according to the determination that the energy level is not less than the energy threshold; Suspending the performance of the HRTF processing on the one or more signals according to the determination that the energy level is less than the energy threshold The method according to claim 5, further comprising.

7. Determining the model of the virtual environment includes: Receiving one or more sound signals from at least a direct sound source and a reflected sound source; Modifying the one or more sound signals to simulate a Doppler effect; Adding a delay to the one or more sound signals; Panning the one or more sound signals across a plurality of virtual speakers comprising, decoding the one or more signals further comprises determining one or more virtual sounds associated with movement of the sound source, the user, or both, the method of claim 1. **Claim 8** An audio signal spatially rendering system, the system comprising: A wearable head device configured to provide the audio signal to a user; and One or more processors configured to execute a method, The method comprising: Determining a model of a virtual environment; Determining a spatial configuration of the virtual environment, the spatial configuration comprising at least a user location, a sound source location, and a virtual speaker location; Determining one or more signals associated with the spatial configuration and further associated with one or more of the user location, the sound source location, or the virtual speaker location; Determining whether an amplitude of one or more signals corresponding to the sound source in the virtual environment exceeds a predetermined amplitude threshold; Decoding the one or more signals in accordance with the determination that the amplitude of the one or more signals exceeds the predetermined amplitude threshold; and Rendering the audio signal based on the one or more signals. **Claim 9** An audio signal spatially rendering system, the system comprising: A wearable head device configured to provide the audio signal to a user; and One or more processors configured to execute a method, The method comprising: Determining a model of a virtual environment; Determining a spatial configuration of the virtual environment, the spatial configuration comprising at least a user location, a sound source location, and a virtual speaker location; Determining one or more signals associated with the spatial configuration and further associated with one or more of the user location, the sound source location, or the virtual speaker location; Determining whether one or more signals corresponding to the sound source in the virtual environment exceed a predetermined threshold; Decoding the one or more signals in accordance with the determination that the one or more signals exceed the predetermined threshold, wherein decoding the one or more signals comprises performing a first set of one or more processing blocks. ​ ​ ​ ​ Selectively bypassing one or more processing blocks of the second set, wherein the one or more processing blocks of the second set are associated with one or more inactive virtual speakers, and Rendering the audio signal based on the one or more signals A system comprising. **Claim 10**: A system for spatially rendering an audio signal, the system comprising A wearable head device configured to provide the audio signal to a user, and One or more processors configured to execute a method Comprising The method comprising Determining a model of a virtual environment, and Determining a spatial configuration of the virtual environment, the spatial configuration comprising at least a user location, a sound source location, and a virtual speaker location, and Determining one or more signals associated with the spatial configuration and further associated with one or more of the user location, the sound source location, or the virtual speaker location, and Determining whether one or more signals corresponding to the sound source within the virtual environment exceed a predetermined threshold, and Decoding the one or more signals according to the determination that the one or more signals exceed the predetermined threshold, and Determining an energy level associated with the one or more signals, and Determining whether the energy level is less than an energy threshold, and Performing head-related transfer function (HRTF) processing on the one or more signals according to the determination that the energy level is not less than the energy threshold, and Suspending the performance of the HRTF processing on the one or more signals according to the determination that the energy level is less than the energy threshold, and Rendering the audio signal based on the one or more signals A system comprising. **Claim 11** The method further comprises Determining whether the number of sound sources within the virtual environment exceeds a predetermined threshold, and The selective bypassing of the one or more processing blocks of the second set comprises bypassing a plurality of detectors according to the determination that the number of sound sources exceeds the predetermined threshold. The system according to claim 9. **Claim 12** The method comprises The system of claim 11, further comprising detecting an energy level of the one or more signals using the plurality of detectors in accordance with a determination that a number of the sound sources does not exceed the predetermined threshold.

13. The method comprises: determining whether the energy level is less than an energy threshold; performing head-related transfer function (HRTF) processing of the one or more signals in accordance with a determination that the energy level is not less than the energy threshold; and deferring performance of the HRTF processing of the one or more signals in accordance with a determination that the energy level is less than the energy threshold. The system of claim 12, further comprising.

14. Determining the model of the virtual environment comprises: receiving one or more sound signals from at least a direct sound source and a reflected sound source; modifying the one or more sound signals to simulate a Doppler effect; adding a delay to the one or more sound signals; and panning the one or more sound signals across a plurality of virtual speakers. Including Decoding the one or more signals further comprises: determining one or more virtual sounds associated with movement of a sound source, a user, or both. The system of claim 8.

Citation Information

Patent Citations

  • Speech processing apparatus and method, encoder, and program

    JP2017055149A

  • Audio Response Based on User Worn Microphones to Direct or Adapt Program Responses System and Method

    US20180014140A1