Efficient rendering of virtual sound fields
By dynamically selecting a combination of virtual speaker subsets and finite impulse response filters, the virtual sound field rendering is optimized, solving the high computational and bandwidth problems of audio systems in virtual environments, and improving rendering efficiency and immersion.
Patent Information
- Application Number
- CN202510808983.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-12
- Filing Date
- 2019-06-12
- Publication Date
- 2025-09-16
AI Technical Summary
Existing virtual environment audio systems require excessive computing resources and network bandwidth when rendering virtual sound fields, especially when processing a large number of sound sources, resulting in low efficiency and inability to provide immersive and instant audio experience.
A modified virtual speaker placement method is adopted to dynamically select and use a subset of fixed virtual speakers, combined with finite impulse response filters to optimize the rendering process of audio signals.
It reduces computational complexity and network bandwidth requirements, lowers power consumption, improves audio rendering efficiency and immersion, and is suitable for mobile devices.
Smart Images

Figure CN120659006A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application “Efficient Rendering of Virtual Sound Field” with application number 201980048983.7 (filing date is June 12, 2019).
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 684,093, filed June 12, 2018, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0004] The present disclosure relates generally to spatial audio rendering and associated systems. More specifically, the present disclosure relates to systems and methods for improving the efficiency of virtual speaker-based spatial audio systems. Background Art
[0005] Virtual environments are ubiquitous in computing environments and can be found in video games (where a virtual environment can represent a game world); maps (where a virtual environment can represent a terrain to be navigated); simulations (where a virtual environment can simulate a real environment); digital storytelling (where virtual characters can interact with each other in a virtual environment); and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the technology used to present virtual environments can limit the user's experience in the virtual environment. For example, traditional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may not be able to implement a virtual environment in a way that creates a compelling, realistic, and immersive experience.
[0006] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present sensory information to a user of an XR system that corresponds to a virtual environment represented by data in a computer system. By combining virtual visual and audio cues with real sight and sound, such systems can provide a uniquely striking sense of immersion and realism. Therefore, it may be desirable to present digital sound to a user of an XR system in such a way that the sound appears to occur naturally in the user’s real environment and conforms to the user’s expectations of sound. Generally speaking, users expect that virtual sound will have the acoustic characteristics of the real environment in which the sound is heard. For example, a user of an XR system in a large concert hall will expect the XR system’s virtual sound to have a large, hollow sound quality; in contrast, a user in a small apartment will expect the sound to be softer, closer, and immediate. Additionally, users expect that virtual sound will be presented without delay.
[0007] Among other techniques, ambisonics and non-ambisonics can be used to generate spatial audio. For a large number of sound source objects, ambisonics or non-ambisonics can be an effective way to render spatial audio due to their design and architecture. This is especially true when modeling reflections. A spatial audio system based on ambisonics and non-ambisonic multi-channels can render audio signals in several steps. Example steps can include an encoding step for each source, a fixed overhead sound field decoding step, and / or a fixed speaker virtualization step. One or more hardware components can perform these steps.
[0008] In a first method for rendering an audio signal, each sound source can have its own pair of finite impulse response (FIR) filters. In such a system, the perceived position of the sound is changed by changing the filter coefficients of the FIR filter. In some embodiments, each sound can use multiple (e.g., two pairs) of FIR filters. Each pair can use two filters (i.e., four FIR filters). When the sound moves in the virtual environment, the FIR filters can crossfade. In some embodiments, each sound can use four FIR filters.
[0009] In a second method for rendering audio signals, virtual speaker panning can be implemented using a fixed number of virtual speakers. Each sound source can be panned on a fixed virtual speaker. In some embodiments, each virtual speaker can use multiple (e.g., two) FIR filters. Virtual speaker panning can be efficient for specific applications and can use negligible computing resources.
[0010] In some embodiments, depending on the number of sounds played simultaneously, a particular method may have improved efficiency compared to another method. For example, 30 sounds may be played simultaneously. If four FIR filters are used for each sound source, the first method may require 120 FIR filters (30 sound sources × 4 FIR filters per sound source = 120 FIR filters). If 2 FIR filters are used for each virtual speaker, the second method may only require 32 FIR filters (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters).
[0011] As another example, there may be only one sound being played. The first approach may only require four FIR filters (1 sound source x 4 FIR filters per sound source = 4 FIR filters), while the second approach may require 32 FIR filters (16 virtual speakers x 2 FIR filters per virtual speaker = 32 FIR filters).
[0012] As shown in the above examples, the first method can be beneficial for a small number of sounds, and the second method can be beneficial for a large number of sounds. Therefore, there may be a need for an audio system and method that improves efficiency based on the number of sound sources at a given time. Summary of the Invention
[0013] An audio system and method for rendering an audio signal using modified virtual speaker placement is disclosed. The audio system can include a fixed number F of virtual speakers, and the modified virtual speaker placement can dynamically select and use a subset P of the fixed virtual speakers. Each sound source can be placed on the subset P of virtual speakers. In some embodiments, multiple (e.g., two) FIR filters can be used for each virtual speaker in the subset P. The subset P of virtual speakers can be selected based on one or more factors, such as proximity to the sound source. The subset P of virtual speakers can be referred to as active speakers.
[0014] As an example, the modified virtual speaker positioning method can be compared with the first and second methods disclosed above. If three sounds are played simultaneously and the audio system has 16 fixed virtual speakers, the first method may require 12 FIR filters (3 sound sources × 4 FIR filters per sound source = 12 FIR filters), and the second method may require 32 FIR filters (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters). On the other hand, the modified virtual speaker positioning method can dynamically select three virtual speakers as active virtual speakers as part of the subset P. The modified virtual speaker positioning method may require six FIR filters, with each active virtual speaker requiring two FIR filters (3 virtual speakers × 2 FIR filters = 6 FIR filters). BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 An example wearable system is shown in accordance with some embodiments.
[0016] Figure 2 An example handheld controller is shown that may be used in conjunction with an example wearable system in accordance with some embodiments.
[0017] Figure 3 An example auxiliary unit that may be used in conjunction with an example wearable system is shown in accordance with some embodiments.
[0018] Figure 4 An example functional block diagram for an example wearable system is shown in accordance with some embodiments.
[0019] Figure 5AA block diagram of an example spatial audio system is shown, in accordance with some embodiments.
[0020] Figure 5B A method for operating according to some embodiments is shown. Figure 5A Flow of an example method of a system.
[0021] Figure 5C A flow diagram of an example method for operating an example decoder / virtualizer is shown in accordance with some embodiments.
[0022] Figure 6 Shown are example configurations of sound sources and speakers according to some embodiments.
[0023] Figure 7A A block diagram of an example decoder / virtualizer including multiple detectors is shown in accordance with some embodiments.
[0024] Figure 7B A method for operating according to some embodiments is shown. Figure 7A Flow of an example method of a decoder / virtualizer.
[0025] Figure 8A A block diagram of an example decoder / virtualizer is shown in accordance with some embodiments.
[0026] Figure 8B A method for operating according to some embodiments is shown. Figure 8A Flow of an example method of a decoder / virtualizer.
[0027] Figure 9 Shown are example configurations of sound sources and speakers according to some embodiments.
[0028] Figure 10A A block diagram is shown of an example decoder / virtualizer for use in a system including active speakers, according to some embodiments.
[0029] Figure 10B A method for operating according to some embodiments is shown. Figure 10A Flow of an example method of a decoder / virtualizer. DETAILED DESCRIPTION
[0030] In the following description of the examples, reference is made to the accompanying drawings which form a part hereof, and in which is shown by way of illustration specific examples that may be practiced. It should be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.
[0031] Example Wearable System
[0032] Figure 1An example wearable head device 100 configured to be worn on a user's head is shown. The wearable head device 100 can be part of a broader wearable system that includes one or more components, such as a head device (e.g., wearable head device 100), a handheld controller (e.g., handheld controller 200 described below), and / or an auxiliary unit (e.g., auxiliary unit 300 described below). In some examples, the wearable head device 100 can be used in virtual reality, augmented reality, or mixed reality systems or applications. The wearable head device 100 may include one or more displays, such as displays 110A and 110B (which may include left and right transmissive displays, and associated components for coupling light from the displays to the user's eyes, such as orthogonal pupil expansion (OPE) grating sets 112A / 112B and exit pupil expansion (EPE) grating sets 114A / 114B); left and right acoustic structures, such as speakers 120A and 120B (which may be mounted on temples 122A and 122B, respectively, and positioned adjacent to the user's left and right ears); one or more sensors, such as an infrared sensor, an accelerometer, a GPS unit, an inertial measurement unit (IMU) (e.g., an IMU 126), an acoustic sensor (e.g., microphone 150); an orthogonal coil electromagnetic receiver (e.g., receiver 127 shown mounted to the left temple arm 122A); left and right cameras oriented away from the user (e.g., depth (time of flight) cameras 130A and 130B); and left and right eye cameras oriented toward the user (e.g., for detecting the user's eye movements) (e.g., eye cameras 128 and 128B). However, the wearable head device 100 can incorporate any suitable display technology, as well as any suitable number, type, or combination of sensors, or other components without departing from the scope of the present invention. In some examples, the wearable head device 100 can incorporate one or more microphones 150 configured to detect audio signals generated by the user's voice; such microphones can be positioned in the wearable head device adjacent to the user's mouth. In some examples, the wearable head device 100 can incorporate networking features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other wearable systems. The wearable head device 100 may further include components such as a battery, a processor, a memory, a storage unit, or various input devices (e.g., buttons, touchpads); or may be coupled to a handheld controller (e.g., handheld controller 200) or an auxiliary unit (e.g., auxiliary unit 300) that includes one or more such components. In some examples, the sensor can be configured to output a set of coordinates of the head-mounted unit relative to the user's environment and can provide input to a processor that performs a simultaneous localization and mapping (SLAM) process and / or a visual odometry algorithm.In some examples, as described further below, the wearable head device 100 can be coupled to a handheld controller 200 and / or an auxiliary unit 300.
[0033] Figure 2 An example mobile handheld controller assembly 200 of an example wearable system is shown. In some examples, the handheld controller 200 can communicate wired or wirelessly with the wearable head device 100 and / or auxiliary unit 300 described below. In some examples, the handheld controller 200 includes a handle portion 220 to be held by a user, and one or more buttons 240 disposed along a top surface 210. In some examples, the handheld controller 200 can be configured to serve as an optical tracking target; for example, a sensor (e.g., a camera or other optical sensor) of the wearable head device 100 can be configured to detect the position and / or orientation of the handheld controller 200, and thereby, by extension, indicate the position and / or orientation of the hand of the user holding the handheld controller 200. In some examples, such as those described above, the handheld controller 200 can include a processor, memory, a storage unit, a display, or one or more input devices. In some examples, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 100). In some examples, the sensor can detect the position or orientation of the handheld controller 200 relative to the wearable head device 100 or relative to another component of the wearable system. In some examples, the sensor can be positioned in the handle portion 220 of the handheld controller 200 and / or can be mechanically coupled to the handheld controller. The handheld controller 200 can be configured to provide one or more output signals, for example, corresponding to the pressed state of the button 240; or, the position, orientation and / or movement of the handheld controller 200 (for example, via an IMU). Such output signals can be used as input to a processor of the wearable head device 100, the auxiliary unit 300, or another component of the wearable system. In some examples, the handheld controller 200 can include one or more microphones to detect sounds (for example, the user's voice, ambient sounds), and in some cases, provide signals corresponding to the detected sounds to a processor (for example, a processor of the wearable head device 100).
[0034] Figure 3An example auxiliary unit 300 of an example wearable system is shown. In some examples, the auxiliary unit 300 can communicate with the wearable head device 100 and / or the handheld controller 200 in wired or wireless communication. The auxiliary unit 300 may include a battery to provide energy to operate one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including a display, sensors, acoustic structures, processors, microphones, and / or other components of the wearable head device 100 or the handheld controller 200). In some examples, as described above, the auxiliary unit 300 may include a processor, a memory, a storage unit, a display, one or more input devices, and / or one or more sensors. In some examples, the auxiliary unit 300 includes a clip 310 for attaching the auxiliary unit to a user (e.g., a belt worn by the user). An advantage of using auxiliary unit 300 to house one or more components of the wearable system is that doing so allows large or heavy components to be carried on the user's waist, chest, or back (which are relatively well-suited for supporting large and heavy objects), rather than being mounted on the user's head (e.g., if housed in wearable head device 100) or carried by the user's hand (e.g., if housed in handheld controller 200). This may be particularly advantageous for relatively heavy or bulky components (such as batteries).
[0035] Figure 4 1 shows an example functional block diagram that may correspond to an example wearable system 400 (such as may include the above-described example wearable head device 100, handheld controller 200, and auxiliary unit 300). In some examples, the wearable system 400 may be used for virtual reality, augmented reality, or mixed reality applications. Figure 4As shown in , the wearable system 400 may include an example handheld controller 400B, referred to herein as a "totem" (and may correspond to the handheld controller 200 described above); the handheld controller 400B may include a totem-to-headgear six degrees of freedom (6DOF) totem subsystem 404A. The wearable system 400 may also include an example wearable head device 400A (which may correspond to the wearable headgear device 100 described above); the wearable head device 400A includes a totem-to-helmet 6DOF helmet subsystem 404B. In this example, the 6DOF totem subsystem 404A and the 6DOF helmet subsystem 404B jointly determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translation directions and rotations along three axes). The six degrees of freedom can be expressed relative to the coordinate system of the wearable head device 400A. In such a coordinate system, the three translation offsets can be expressed as X, Y, and Z offsets, which can be expressed as a translation matrix, or some other representation. The rotational degrees of freedom can be expressed as a sequence of yaw, pitch, and roll rotations; as a vector; as a rotation matrix; as a quaternion; or as some other representation. In some examples, one or more depth cameras 444 (and / or one or more non-depth cameras) included in the wearable head device 400A; and / or one or more optical sights (e.g., the button 240 of the handheld controller 200 described above, or a dedicated optical sight included in the handheld controller) can be used for 6DOF tracking. In some examples, as described above, the handheld controller 400B can include a camera; and the helmet 400A can include an optical sight for optical tracking in conjunction with the camera. In some examples, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids for wirelessly transmitting and receiving three distinguishable signals. By measuring the relative amplitudes of the three distinguishable signals received in each coil used for reception, the 6DOF of the handheld controller 400B relative to the wearable head device 400A can be determined. In some examples, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that can be used to provide improved accuracy and / or more timely information about rapidly moving handheld controller 400B.
[0036] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space that is fixed relative to the wearable head device 400A) to an inertial coordinate space or an ambient coordinate space. For example, such a transformation may be necessary for the display of the wearable head device 400A to present a virtual object (e.g., a virtual person sitting in a real chair, facing forward, regardless of the position and orientation of the wearable head device 400A) at an expected position and orientation relative to the real environment rather than a fixed position and orientation on the display (e.g., at the same position in the display of the wearable head device 400A). This can maintain the illusion that the virtual object exists in the real environment (and, for example, moves and rotates as the wearable head device 400A does not appear to be unnaturally positioned in the real environment). In some examples, a compensating transformation on the coordinate space can be determined by processing images from the depth camera 444 (e.g., using simultaneous localization and mapping (SLAM) and / or visual odometry processes) to determine the transformation of the wearable head device 400A relative to the inertial or ambient coordinate system. Figure 4 In the example shown in , the depth camera 444 can be coupled to the SLAM / visual odometry module 406 and can provide images to the module 406. An implementation of the SLAM / visual odometry module 406 may include a processor configured to process the image and determine the position and orientation of the user's head, which can then be used to identify a transformation between the head coordinate space and the actual coordinate space. Similarly, in some examples, an additional source of information about the user's head pose and position is obtained from the IMU 409 of the wearable head device 400A. Information from the IMU 409 can be integrated with information from the SLAM / visual odometry module 406 to provide improved accuracy and / or more timely information for rapid adjustment of the user's head pose and position.
[0037] In some examples, depth camera 444 can provide 3D images to gesture tracker 411, which can be implemented in a processor of wearable head device 400A. Gesture tracker 411 can recognize a user's gesture, for example, by matching the 3D images received from depth camera 444 with stored patterns representing gestures. Other suitable techniques for recognizing user gestures will be apparent.
[0038] In some examples, one or more processors 416 can be configured to receive data from the headset subsystem 404B, the IMU 409, the SLAM / visual odometry module 406, the depth camera 444, a microphone (not shown), and / or the gesture tracker 411. The processor 416 can also send and receive control signals from the 6DOF totem system 404A. In examples where the handheld controller 400B is untethered, the processor 416 can be wirelessly coupled to the 6DOF totem system 404A. The processor 416 can further communicate with additional components such as an audiovisual content storage 418, a graphics processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 can be coupled to a head-related transfer function (HRTF) storage 425. The GPU 420 can include a left channel output coupled to a left source 424 of image-forming modulated light and a right channel output coupled to a right source 426 of image-forming modulated light. The GPU 420 can output stereo image data to a source of image modulated light 424, 426. The DSP audio sound field locator 422 can output audio to the left speaker 412 and / or the right speaker 414. The DSP audio sound field locator 422 can receive an input indicating a direction vector from the user to the virtual sound source from the processor 416 (the virtual sound source can be moved by the user, for example, via a handheld controller 400B). Based on the direction vector, the DSP audio sound field locator 422 can determine the corresponding HRTF (for example, by accessing the HRTF, or by interpolating multiple HRTFs). The DSP audio sound field locator 422 can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. By combining the relative position and orientation of the user relative to the virtual sound in a mixed reality environment, that is, by presenting a virtual sound that matches the user's expectation that the virtual sound sounds like a real sound in a real environment, the credibility and authenticity of the virtual sound can be enhanced.
[0039] In some examples, such as Figure 4 As shown in , one or more of the processor 416, GPU 420, DSP audio sound field localizer 422, HRTF memory 425, and audio / video content memory 418 may be included in an auxiliary unit 400C (which may correspond to the auxiliary unit 300 described above). The auxiliary unit 400C may include a battery 427 to power its components and / or to power the wearable head device 400A and / or the handheld controller 400B. Including such components in an auxiliary unit that can be mounted to the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue on the user's head and neck.
[0040] although Figure 4 Elements corresponding to the various components of the example wearable system 400 are presented, but various other suitable arrangements of these components will become apparent to those skilled in the art. For example, the auxiliary unit 400C is associated with Figure 4 The elements presented in the examples may be associated with the wearable head device 400A or the handheld controller 400B instead. In addition, some wearable systems may completely abandon the handheld controller 400B or the auxiliary unit 400C. Such changes and modifications should be understood to be included within the scope of the disclosed examples.
[0041] Mixed reality environment
[0042] Like all people, users of mixed reality systems exist in a real environment—that is, the three-dimensional portion of the "real world" and all its contents that the user can perceive. For example, users perceive the real environment using ordinary human senses (sight, sound, touch, taste, smell) and interact with the real environment by moving their bodies within the real environment. Positions in the real environment can be described as coordinates in a coordinate space; for example, coordinates can include latitude, longitude, and altitude relative to sea level; distances in three orthogonal dimensions from a reference point; or other suitable values. Similarly, vectors can describe quantities that have direction and magnitude in a coordinate space.
[0043] A computing device may maintain a representation of a virtual environment, for example, in a memory associated with the device. As used herein, a virtual environment is a computational representation of a three-dimensional space. A virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other features associated with the space. In some examples, circuitry (e.g., a processor) of a computing device may maintain and update the state of the virtual environment; that is, the processor may determine the state of the virtual environment at a second time based on data associated with the virtual environment and / or input provided by a user at a first time. For example, if an object in the virtual environment is located at a first coordinate at that time and has certain programmed physical parameters (e.g., mass, coefficient of friction); and an input is received from the user indicating that a force should be applied to the object in a direction vector; the processor may apply the laws of kinematics to determine the position of the object at that time using basic mechanics. The processor may use any known appropriate information about the virtual environment and / or any appropriate input to determine the state of the virtual environment at that time. In maintaining and updating the state of the virtual environment, the processor may execute any appropriate software, including software relating to creating and deleting virtual objects in the virtual environment; software (e.g., scripts) for defining the behavior of virtual objects or characters in the virtual environment; software for defining the behavior of signals (e.g., audio signals) in the virtual environment; software for creating and updating parameters associated with the virtual environment; software for generating audio signals in the virtual environment; software for processing input and output; software for implementing network operations; software for applying asset data (e.g., animation data to move virtual objects over time); or many other possibilities.
[0044] An output device (such as a display or speakers) can present any or all aspects of the virtual environment to the user. For example, the virtual environment can include virtual objects (which can include representations of inanimate objects, people, animals, lights, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" having origin coordinates, viewing axes, and a frustum) and render the visible scene of the virtual environment corresponding to that view to the display. Any suitable rendering technique can be used for this purpose. In some examples, the visible scene may include only some virtual objects in the virtual environment and exclude certain other virtual objects. Similarly, the virtual environment can include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects in the virtual environment can generate sounds originating from the object's location coordinates (e.g., a virtual character can speak or cause a sound effect); or the virtual environment can be associated with music cues or ambient sounds that may or may not be associated with a specific location. The processor may determine audio signals corresponding to "listener" coordinates, e.g., audio signals corresponding to a synthesis of sounds in a virtual environment, and mix and process to simulate the audio signals that would be heard by the listener at the listener coordinates, and present the audio signals to the user via one or more speakers.
[0045] Because the virtual environment exists only as a computational construct, the user cannot directly perceive the virtual environment using their ordinary senses. Instead, the user can only perceive the virtual environment indirectly, for example, through a display, speakers, tactile output devices, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment; however, input data can be provided via input devices or sensors to a processor that can use the device or sensor data to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use this data to cause the object in the virtual environment to respond accordingly.
[0046] Digital reverb and ambient audio processing
[0047] The XR system can present audio signals to the user, which originate from a sound source with an origin coordinate and propagate in the system in a direction with an orientation vector. The user can perceive these audio signals as if they were real audio signals originating from the sound source's origin coordinate and propagating along the orientation vector.
[0048] In some cases, audio signals may be considered virtual because they correspond to computed signals in a virtual environment and do not necessarily correspond to real sounds in a real environment. However, virtual audio signals may be used, for example, as Figure 1The real audio signals detectable by the human ear generated by the speakers 120A and 120B of the wearable head device 100 are presented to the user.
[0049] Advantages of the embodiments disclosed below include reduced network bandwidth, reduced power consumption, reduced computational complexity, and reduced computational latency. These advantages are particularly important for mobile systems (including wearable systems) where processing resources, network resources, battery capacity, and physical size and weight are often at a premium.
[0050] In a dynamic environment such as AR, the system may be constantly rendering audio signals. Using all virtual speakers to render audio signals can be computationally intensive, require extensive processing, require high network bandwidth, consume high power, and so on. Therefore, it may be desirable to use modified virtual speaker placements, dynamically selecting and using a subset of fixed virtual speakers based on one or more factors.
[0051] Example spatial audio system
[0052] Figure 5A A block diagram of an example spatial audio system is shown, in accordance with some embodiments. Figure 5B Shown for operation Figure 5A Flow of an example method of a system.
[0053] The spatial audio system 500 may include a spatial modeler 510, an internal spatial representation 530, and a decoder / virtualizer 540A. The spatial modeler 510 may include a direct path section 512, one or more reflection sections 520 (optional), and a spatial encoder 526. The spatial modeler 510 may be configured to model a virtual environment. The direct path section 512 may include a direct source 514 and, optionally, a Doppler 516. The direct source 514 may be configured to provide an audio signal (step 552 of process 550). The Doppler 516 may receive a signal from the direct source 514 and may be configured to introduce a Doppler effect into its input signal (step 554). For example, the Doppler 516 may change the pitch of a sound source (e.g., pitch shift) to change relative to the motion of the sound source, the user of the system, or both.
[0054] Reflection unit 520 may include an acoustic reflector 522, an optional Doppler 516, and a delay 524. Acoustic reflector 522 may be configured to introduce reflections into its signal (step 556). The introduced reflections may represent one or more properties of the environment. Doppler 516 in reflection unit 520 may receive a signal from acoustic reflector 522 and may be configured to introduce a Doppler effect into its input signal (step 558). Delay 524 may receive a signal from Doppler 516 and may be configured to introduce a delay (step 560).
[0055] Spatial encoder 526 may receive signals from direct path section 512 and reflection section(s) 520. In some embodiments, the signal from direct path section 512 to spatial encoder 526 may be an output signal of Doppler 516 from direct path section 512. In some embodiments, the signal(s) from reflection section(s) 520 to spatial encoder 526 may be an output signal(s) of delay(s) 524 from reflection section(s) 520.
[0056] The spatial encoder 526 may include one or more M-way panners (pans) 528. In some embodiments, each input received by the spatial encoder 526 may be associated with a unique M-way panner 528. "Panning" may refer to distributing a signal across multiple speakers, multiple locations, or both. The M-way panner 528 may be configured to distribute its input signal across multiple virtual speakers (step 562). For example, the M-way panner 528 may distribute its input signal across all M virtual speakers. For example, Figure 5A As shown, M can be equal to four, and each M-way positioner 528 can be configured to distribute its input signal across four virtual speakers. Although the figure shows a system with four virtual speakers, examples of the present disclosure can include any number of virtual speakers.
[0057] As an example, a car system may include a left speaker and a right speaker. The sound in such a system can be positioned on the left and right speakers of the car by splitting the sound into two (one for each speaker). The scaling volume of each speaker can be set according to the configuration of the two speakers, and the result can be sent to the left and right speakers.
[0058] As another example, a surround sound system may include multiple speakers, such as six speakers. The sound in such a system can be positioned as stereo across the six speakers. The sound can be split into six (instead of two in the example car system), and the scaled volume of each speaker can be set based on the configuration of the six speakers, with the result being sent to all six speakers.
[0059] For example, the first M-way positioner 528 may receive the output of the Doppler 516 of the direct path 512, while the other M-way positioners 528 may receive the output of the reflection unit 520. Each M-way positioner 528 may split its input signal so that it can be distributed among multiple outputs. In this way, each M-way positioner 528 may have a greater number of outputs than inputs.
[0060] The spatial modeler 510 may output a signal to the interior spatial representation 530 (step 564). In some embodiments, the output(s) from the spatial modeler 510 may include the output of each of the M-way positioners 528. The interior spatial representation 530 may be configured to represent the spatial configuration of the virtual environment (step 566). An example representation may include representing the relative positions of the user, the sound source(s), and the virtual speaker(s). In some embodiments, the interior spatial representation 530 may output one or more signals representing a head pose rotation, a head pose translation, a sound field decode, one or more head-related transfer functions (HRTFs), or a combination thereof, of a user of the system 500. In some embodiments, the interior spatial representation 530 may be a representation of a non-ambisonic multi-channel based system, an Ambisonics / Wavefield based system, or the like. An exemplary Ambisonics / Wavefield based system may be High Order Ambisonics (HOA).
[0061] The interior spatial representation 530 may output its signal 552 to the decoder / virtualizer 540A (step 568). The decoder / virtualizer 540 may decode its input signal and introduce virtualized sound into the signal (step 570). Step 570 may include multiple sub-steps and will be discussed in more detail below. The system then outputs the signal from the decoder / virtualizer 540 as the left signal 502L to the left speaker and as the right signal 502R to the right speaker (step 580).
[0062] System 500 may include any number of different types of decoders / virtualizers 540 . Figure 5A One example decoder / virtualizer 540A is shown in FIG. Other example decoders / virtualizers 540 are discussed below.
[0063] The decoder / virtualizer 540A may include a rotation / translation representation 542 , a sound field decoder 544 , one or more HRTFs 546 , and one or more combiners 548 . Figure 5CA process for an example method of operating an example decoder / virtualizer is shown, which may be referred to as step 570-1. The rotation / translation representation 542 may receive (multiple) signals from the internal space representation 530 and may be configured to introduce a representation of motion associated with the audio signal. For example, the motion may be the motion of a sound source, a user, or both (step 572). The rotation / translation representation 542 may output (multiple) signals to a sound field decoder 544. The sound field decoder 544 may receive (multiple) signals from the rotation / translation representation 542 and may be configured to decode the signals (step 574). Each HRTF 546 may receive (multiple) signals from the sound field decoder 544. Each HRTF 546 may be configured to determine an HRTF corresponding to its input signal and apply it to the signal (step 576). One or more HRTFs 546 may be collectively referred to as a speaker virtualizer. In some embodiments, the HRTF 546 may be configured for finite impulse response (FIR) filtering. Each combiner 548 may receive and combine the signal(s) from the HRTF(s) 546 (step 578).
[0064] In some embodiments, decoder / virtualizer 540A may represent a "baseline" processing overhead. The baseline processing overhead may be complex, involving matrix calculations and long FIR filters to apply the HRTF processing to each virtual speaker.
[0065] The output from the combiner 548 may be the output signal that forms the system 500. In some embodiments, the output signal 502 from the system 500 may be the output signal for the left and right speakers (e.g., Figure 1 audio signals to speakers 120A and 120B).
[0066] In some cases, when the number of sound sources used for playback is large, Figure 5A However, in some cases, when the number of sound sources used for playback is small, Figure 5A It may be desirable to utilize a non-ambisonic multi-channel based spatial audio system or an Ambisonics based spatial audio system (such as a 4K ... Figure 5A The efficiency of the system 500).
[0067] There can be multiple ways to improve spatialization efficiency using sound field synthesis and decoding. A first way can be through low-energy speaker detection and cull. In low-energy speaker detection and cull, if the energy output of a virtual speaker channel of a non-fidelity stereo multi-channel based spatial audio system or a high-fidelity stereo / sound field channel of a high-fidelity stereo based spatial audio system is less than a predetermined threshold, no processing of the signal from the virtual speaker channel is performed. In some embodiments, for example, before performing sound field decoding on the signal from that given virtual speaker, the system can determine whether the output of a given virtual speaker is above a predetermined threshold. Low-energy speaker detection and cull will be discussed in more detail below.
[0068] A second approach to improving spatialization efficiency using sound field synthesis and decoding can be source geometry-based virtual speaker selection. In source geometry-based virtual speaker selection, decoder / virtualizer processing can be selectively disabled. Selective disabling (or selective enabling) can be based on the position of the sound source(s) relative to the user / listener. Source geometry-based virtual speaker selection is discussed in detail below.
[0069] A third approach could be to combine low-energy loudspeaker detection and selection techniques with source-virtual loudspeaker coupling techniques.
[0070] The spatial modeler 510 may have a computational complexity, which may represent the number of operations required to process the audio signal. The computational complexity may be proportional to M times N, where M may be equal to the number of sound sources (including direct sound sources and optional reflections) and N may be equal to the number of channels required to represent an ambisonics field. In some embodiments, N may be equal to (O+1). 2 , where O is the order of the ambisonics used.
[0071] The decoder / virtualizer 540 can have a computational complexity proportional to nVS, where nVS is the number of virtual speakers. The computational power of each speaker can be high and typically consists of a pair of FIR filters, typically implemented using a Fast Fourier Transform (FFT) or Inverse FFT (IFFT), both of which are computationally expensive processes.
[0072] Example Low Energy Output Detection and Sorting Method
[0073] In some embodiments, some virtual speakers may have little or no signal input energy; for example, when the number of sound sources in the spatial audio system is small. Speaker virtualization processing can be a computationally expensive (e.g., CPU intensive) process. For example, if there is a sound source located at zero degrees azimuth (e.g., directly in front of the user), there may be little or no energy in the signal from virtual speakers located between 90 degrees and 270 degrees azimuth (e.g., behind the user). Low-energy signals may not have a significant impact on the perceived location of the sound source, and therefore performing speaker virtualization processing and / or determining the characteristics of the corresponding virtual speakers on low-energy signals may be computationally inefficient.
[0074] To reduce the required computational resources, a system employing a low-energy output detection and selection method can include a detector located between the sound field decoder and the HRTF. Alternatively, the detector can be located between the multi-channel output and the HRTF. The detector can be configured to detect one or more energy levels associated with one or more audio signals from one or more virtual speakers.
[0075] If the energy level of the signal from the virtual speaker Vn is less than the energy threshold α, the signal can be considered a low energy signal. Based on the detected energy level associated with the audio signal being less than the energy threshold α, the HRTF block and its processing of the low energy signal can be bypassed.
[0076] A variety of techniques can be used to determine the energy level of a signal. For example, an RMS algorithm can be applied to the signal routed to the virtual speaker to measure its energy. Attack and release times, similar to those used by traditional audio compressors, can be used to prevent the speaker's signal from suddenly "jumping" in and out.
[0077] Figure 6An example configuration of sound sources and speakers according to some embodiments is shown. The system 600 may include a sound source 620 and a plurality of speakers. The plurality of speakers 622 may include one or more active virtual speakers 622A and one or more inactive virtual speakers 622B. The active virtual speaker 622A may be a virtual speaker whose signal is processed by the HRTF 546 at a given time. The inactive virtual speaker 622B may be a speaker whose signal does not need to be processed by the HRTF 546 because, for example, its signal has already been processed at a previous time, or because the system determines that the signal from the virtual speaker 622B does not need to be processed. M may represent the number of sound sources being played, and N may represent the number of virtual speakers in the system. Although the figure shows a single sound source, examples of the present disclosure may include any number of sound sources. Although the figure shows eight sound sources, examples of the present disclosure may include any number of sound sources, such as 16 (N=16).
[0078] As an example, as shown, the system 600 may include a single (M=1) sound source 620 and eight virtual speakers 622. In a given situation, most of the energy is output on only three virtual speakers. That is, the system 600 may have three active virtual speakers at the first time. For example, virtual speakers 622A-1, 622A-2, and 622-3 may be active virtual speakers. In some embodiments, the active virtual speaker 622A may be the speaker closest to the sound source 620. In addition, the system 600 may include five inactive virtual speakers 622B. The system 600 may determine that the energy level from each of the five inactive virtual speakers is less than an energy threshold, and based on this determination, the HRTF processing of the signals from the five inactive virtual speakers 622B may be bypassed.
[0079] The system 600 may also determine that the energy level from each active virtual speaker is not less than an energy threshold, and based on this determination, may perform HRTF processing on the signals from the three active virtual speakers 622A.
[0080] System 600 may output two signals, one for the right speaker and one for the left speaker, e.g. Figure 5A The right signal 502R and the left signal 502L are shown. The reduction in the number of HRTF operations due to bypassing HRTF processing can be equal to the number of inactive virtual speakers multiplied by the number of signals output from the system. Figure 6 In the example of , since HRTF processing of five signals is bypassed, 10 (five inactive virtual speakers × two output signals) HRTF operations can be saved.
[0081] As another example, if the system includes 16 virtual speakers, 13 of which are inactive virtual speakers, the number of saved HRTF operations may be equal to 26 (16 virtual speakers x two output signals).
[0082] Figure 7A A block diagram of an example decoder / virtualizer including multiple detectors is shown in accordance with some embodiments. Figure 7B A method for operating according to some embodiments is shown. Figure 7A In some embodiments, the decoder / virtualizer 540A ( Figure 5A ), the decoder / virtualizer 540B may be included in the system 500 as described below. Figure 5C As shown), step 570-2 may be included in process 550.
[0083] The decoder / virtualizer 540B can include a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more switches 712, one or more HRTFs 546, and one or more combiners 548. The decoder / virtualizer 540B can be configured to generate a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more switches 712, one or more HRTFs 546, and one or more combiners 548. Figure 5A The rotation / translation representation 542 can receive the signal(s) 552 from the interior space representation 530 and can be configured to introduce a representation of the motion of the sound source(s), the user, or both (step 772). The rotation / translation representation 542 can output the signal(s) to the sound field decoder 544. The sound field decoder 544 can receive the signal(s) from the rotation / translation representation 542 and can be configured to decode the signal (step 774). The sound field decoder 544 can output the signal(s) to the detector(s) 710.
[0084] The detector(s) 710 may receive signals from the sound field decoder 544 and may be configured to determine the energy level of its input signal (step 776). Each detector 710 may be coupled to a unique switch 712. If the energy level of the input signal (from the sound field decoder 544) is greater than or equal to an energy threshold (step 778), the switch 712 may close the loop, thereby routing its input signal (from the detector 710) to the HRTF 546 coupled to the switch (step 780). Each HRTF determines a corresponding HRTF and applies it to the signal (step 782).
[0085] If the energy level of the input signal is less than the energy threshold, the switch 712 can be opened so that its input signal (from the detector 710) is not coupled to the corresponding HRTF 546. Therefore, the corresponding HRTF 546 can be bypassed (step 784).
[0086] The signals from the HRTF(s) 546 can be output to the combiner 548 (step 786). The combiner 548 can be configured to combine (e.g., add, sum, etc.) the signals from the HRTF(s) 546. Signals that bypass the HRTF 546 may not be combined by the combiner 548. The output from the combiner 548 may form the output signal of the system 500. In some embodiments, the output signal 502 from the system 500 may be for the left and right speakers (e.g., Figure 1 audio signals to speakers 120A and 120B).
[0087] In some embodiments, each detector 710 can be coupled to a unique signal corresponding to a virtual speaker. In this way, processing for each virtual speaker 622 can be performed independently (i.e., processing for one speaker (e.g., 622A-1) can occur without affecting processing for another speaker (e.g., 622B)).
[0088] In some embodiments, the type of decoder / virtualizer 540 may depend on the number of sound sources. For example, if the number of sound sources is less than or equal to a predetermined sound source threshold, then Figure 7A The decoder / virtualizer 540B may be included in the system 500. In this case, the signal from the sound field decoder 544 may be input to the detector(s) 710.
[0089] If the number of sound sources is greater than the predetermined sound source threshold, then Figure 5A A decoder / virtualizer 540A may be included in the system. In this case, the signal from the sound field decoder 544 may be input to the HRTF 546.
[0090] In some embodiments, the system may include a decoder / virtualizer 540 that can choose whether to execute or bypass the detector and its energy level detection. Figure 8A A block diagram of an example decoder / virtualizer is shown in accordance with some embodiments. Figure 8B A method for operating according to some embodiments is shown. Figure 8A In some embodiments, the decoder / virtualizer 540A ( Figure 5A ) and decoder / virtualizer 540B ( Figure 7A As shown), decoder / virtualizer 540C may be included in system 500. Alternatively, step 570-1 ( Figure 5C As shown), step 570-3 may be included in process 550.
[0091] The decoder / virtualizer 540C can include a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more first switches 712, one or more HRTFs 546, and one or more combiners 548, similar to the decoder / virtualizer 540B discussed above. Steps 872, 874, and 882 can be similar to steps 772, 774, and 782 discussed above, respectively.
[0092] The decoder / virtualizer 540C may also include a second switch 814. The second switch 814 can be configured to open or close a first loop from the sound field decoder 544 to the detector(s) 710 and the first switch(es) 712. Additionally or alternatively, the second switch 814 can be configured to open or close a second loop from the system 500 that bypasses the detector(s) 710 and the first switch(es) 712. In some embodiments, the second switch 814 can be a bidirectional switch configured to select between passing the signal directly to the detector 710 (first loop) or directly to the HRTF 546 (second loop).
[0093] For example, the system can determine whether the number of sound sources is greater than or equal to a predetermined sound source threshold (step 876). If the number of sound sources is greater than or equal to the predetermined sound source threshold, the second switch 814 can close the second loop and pass the signal from the sound field decoder 544 directly to the HRTF 546 (step 878). Each HRTF 546 then determines a corresponding HRTF and applies it to the signal (step 880). When the number of sound sources is greater, the likelihood of a signal with a low energy level may decrease.
[0094] On the other hand, if the number of sound sources is less than the predetermined sound source threshold, the signal is more likely to have a low energy level, so the second switch 814 can close the first loop and pass the signal from the sound field decoder 544 directly to the (multiple) detectors 710 (step 882). The (multiple) detectors 710 can receive signals from the sound field decoder 544 and can be configured to determine the energy level of their input signals (step 884). If the energy level of the input signal (from the sound field decoder 544) is greater than or equal to the energy threshold (step 886), the switch 712 can close the loop, thereby routing its input signal (from the detector 710) to the HRTF 546 to which the switch is coupled (step 888). If the energy level of the input signal is less than the energy threshold, the switch 712 can be opened so that its input signal (from the detector 710) is not coupled to the corresponding HRTF 546, and the HRTF 546 is bypassed (step 890).
[0095] The signals from the HRTF(s) 546 may be output to the combiner 548 (step 892).
[0096] In some embodiments, one or more energy threshold detections can be activated in response to energy. In some embodiments, one or more energy threshold detections can be activated in response to amplitude, and can be subject to traditional attack, release times, etc.
[0097] Example of a loudspeaker selection method based on source geometry
[0098] Virtual speaker selection based on source geometry can be another way to reduce CPU consumption. In some embodiments, virtual speaker selection based on source geometry can include selectively disabling decoder / virtualizer processing (e.g., Figure 5A Decoder / virtualizer 540A, Figure 7A Decoder / virtualizer 540B, Figure 8A In some embodiments, selective disabling (or selective enabling) of decoder / virtualizer processing can include bypassing all of the processing blocks of the decoder / virtualizer.
[0099] Using virtual speaker selection based on source geometry, high-fidelity stereo output can be calculated. If the high-fidelity stereo output requires a lot of energy to decode, it may be beneficial to use a simpler method (that requires less CPU consumption), such as a real-time energy detection method. In addition, in some embodiments, the real-time energy detection method can perform calculations less frequently.
[0100] Figure 9 9 shows an example configuration of a sound source and a speaker according to some embodiments. System 900 may include a sound source 920 and a plurality of speakers. Figure 6 Compared to the system 600, the sound source 920 can be located at a second location, which can be different from Figure 6 The plurality of speakers 922 may include one or more active virtual speakers 922A, one or more inactive virtual speakers 922B, and one or more inactive virtual speakers 922C. The active virtual speakers 922A and the inactive virtual speakers 922B may be respectively similar to Figure 6 The active virtual speaker 622A and the inactive virtual speaker 622B.
[0101] The inactive virtual speaker 922C may differ from the inactive virtual speaker 922B in that the virtual speaker 922C may be active at a first time, but its signal is processed at a second time (eg, a ring output period). Figure 9 In the example shown in FIG, sound source 920 may have moved from a first position (e.g., near virtual speaker 922C) to a second position (e.g., not near virtual speaker 922). Due to the movement of the sound source, the two virtual speakers may no longer have the sound source mixed into them at the second time. Due to the filtering process of the two virtual speakers, the two virtual speakers may need to be active in the subsequent frame (e.g., the second time) to correctly complete the filtering process.
[0102] In some embodiments, the system may include a decoder / virtualizer 540 in a system that uses active virtual speakers. Figure 10A A block diagram is shown of an example decoder / virtualizer for use in a system including active speakers, according to some embodiments. Figure 10B A method for operating according to some embodiments is shown. Figure 10A In some embodiments, instead of ( Figure 5A ) decoder / virtualizer 540A, ( Figure 7A ) decoder / virtualizer 540B and (in Figure 8A ) decoder / virtualizer 540C, decoder / virtualizer 540D may be included in system 500. Instead of ( Figure 5C Step 570-1, ( Figure 7B ) step 570-2, and ( Figure 8B As shown) steps 570-3 and 570-4 may be included in process 550.
[0103] The decoder / virtualizer 540C can include a sound field decoder 544, one or more HRTFs 546, and one or more combiners 548, similar to the decoder / virtualizer 540B and the decoder / virtualizer 540C discussed above. Steps 1072, 1076, 1078, and 1080 can be similar to steps 872, 874, and 782 discussed above, respectively.
[0104] The decoder / virtualizer 540D may also include a rotation / translation representation 1042 and a sound field decoder determiner 1044. The rotation / translation representation 1042 may receive signal(s) from the interior space representation 530 and may be configured to introduce a representation of the motion of the sound source(s), the user, or both (step 1072). The representation of motion may also take into account the azimuth / pitch of the sound source 920. The rotation / translation representation 542 may output the signal(s) to the sound field decoder determiner 1044.
[0105] The sound field decoder determiner 1044 can receive (multiple) signals from the rotation / translation representation 1042 and can be configured to determine which signal has a "significant" output and pass those signals to the sound field decoder 544 (step 1074). A significant output can be an output that affects the perceived sound. For example, a significant output can be an audio signal with an amplitude greater than or equal to a predetermined amplitude threshold. The sound field decoder 544 can receive (multiple) signals from the sound field decoder determiner 1044 with a significant output and can be configured to decode the signal (step 1076). In some embodiments, the sound field decoder 1044 can receive a signal with a significant output from the sound field decoder determiner 1044. Each HRTF 546 can receive (multiple) signals from the sound field decoder 544. Each HRTF 546 can be configured to determine an HRTF corresponding to its input signal and apply the HRTF to the signal (step 1078). One or more HRTFs 546 can be collectively referred to as a speaker virtualizer. Each combiner 548 may receive and combine the signal(s) from the HRTF(s) 546 (step 1080).
[0106] In some embodiments, those audio signals that do not have a significant output (e.g., have an amplitude less than a predetermined amplitude threshold) may not be passed to the sound field decoder 544. Thus, for audio signals that do not have a significant output, the sound field decoder 544 and the HRTF 546 may be bypassed.
[0107] An example source geometry-based speaker selection method can designate a virtual speaker as an active virtual speaker based on the location (e.g., X, Y, Z location) of a sound source. The location of the sound source can represent the location of a sound source object. The system can determine the location of each sound source and determine which virtual speaker(s) are located near the corresponding sound source. In some embodiments, the determination of which virtual speaker is located near the sound source can be performed, for example, at the beginning of each video frame (video frame rate-based method). Compared to other methods such as sampling rate-based methods, video frame rate-based methods may require less computation.
[0108] Based on methods such as those calculated based on the video frame rate and the Ambisonics decoding formula, a sound source may contribute significantly to a particular virtual speaker. As described above, if decoded, a virtual speaker that contributes little energy can cause the corresponding Ambisonics decoding and HRTF processing of the decoded Ambisonics channel to be bypassed. In some embodiments, the system can disable any bypassed processing blocks.
[0109] Example pseudocode for executing the specified method could be:
[0110]
[0111]
[0112] Regarding the above pseudocode, the variable sourcePosition may refer to the position of the sound source, sourceOrientation may refer to the orientation of the sound source, ListenerPosition may refer to the position of the user / listener, ListenerOrientation may refer to the orientation of the user / listener, VirtualSpeakerPosition may relate to the position of the virtual speaker, AmbisonicDecode may refer to a function that performs high-fidelity stereo decoding, and Virtualize may refer to a function that performs virtualization.
[0113] With respect to the above pseudo code, for each sound source S and decoding channel n, decoding channel n can be enabled based on one or more factors, such as the location of the sound source S, the orientation of the sound source S, the location of the user / listener, the orientation of the user / listener, and the location of the virtual speaker. Still referring to the above pseudo code, for each Ambisonics decoding channel, if the channel is enabled, the system can perform Ambisonics decoding functions and virtualization functions.
[0114] The pseudocode can be improved by providing a "ringing" period for each virtual speaker. For example, if a source moves in position during a video frame, it can be determined that a virtual speaker may no longer have any sound sources mixed into it. However, due to the filtering process of the virtual speaker, this virtual speaker may need to be the active speaker for the next frame in order for the filtering process to complete correctly.
[0115] Examples of the present disclosure may include using all active sound sources to determine which decoded sound field outputs have "significant" outputs (e.g., outputs that will affect the perceived sound field). Ambisonics or non-Ambisonics multi-channel outputs that will affect the perceived sound field may be decoded. Furthermore, in some embodiments, only the HRTFs 546 corresponding to those detected outputs are processed. For synthetically generated Ambisonics sound fields or non-Ambisonics multi-channel renderings, where the number of sound sources is small, or where the number of sound sources is large but close together, significant CPU savings may be possible.
[0116] Example of a source geometry-based virtual loudspeaker selection method and a low-energy output detection and selection method French combination
[0117] In some embodiments, source geometry-based virtual speaker selection and low energy output detection and selection can be used sequentially to further reduce CPU consumption. As described above, source geometry-based virtual speaker selection can include, for example, selectively disabling virtual speaker processing based on, for example, the position of the sound source relative to the user / listener. Low energy output detection and selection can include, for example, placing a signal energy / level detector between the sound field decoding or multi-channel output and the HRTF processing. The output / results of the source geometry-based virtual speaker selection can be input to the low energy output detection and selection.
[0118] With respect to the above-described systems and methods, the elements of the systems and methods can be suitably implemented by one or more computer processors (e.g., a CPU or DSP). The present disclosure is not limited to any particular configuration of computer hardware, including computer processors, for implementing these elements. In some cases, multiple computer systems can be used to implement the above-described systems and methods. For example, a first computer processor (e.g., a processor of a wearable device coupled to a microphone) can be used to receive input microphone signals and perform initial processing of those signals (e.g., signal conditioning and / or segmentation, such as described above). A second (and perhaps more powerful) processor can then be used to perform more computationally intensive processing, such as determining probability values associated with speech segments of those signals. Another computer device, such as a cloud server, can host a speech recognition engine, ultimately providing it with input signals. Other suitable configurations will be apparent and within the scope of the present disclosure.
[0119] Although the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are understood to be within the scope of the disclosed examples as defined by the appended claims.
Claims
1. A method for spatially rendering an audio signal, the method comprising: determining a spatial configuration of the virtual environment, wherein the spatial configuration includes sound source locations and virtual speaker locations; determining one or more signals associated with the sound source location; decoding the one or more signals based on determining that the distance between the sound source position and the virtual speaker position is less than a predetermined distance; and The audio signal is rendered based on the decoded one or more signals.
2. The method according to claim 1, further comprising: further applying the one or more signals to a head-related transfer function (HRTF) based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance; Based on determining that the distance between the sound source position and the virtual speaker position is not less than the predetermined distance, applying the one or more signals to the HRTF is abandoned.
3. The method according to claim 1, wherein Decoding the one or more signals comprises applying a first set of processing blocks to the one or more signals, and wherein the method further comprises: Further based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance, a second group of processing blocks is bypassed, the second group of processing blocks being associated with one or more inactive virtual speakers.
4. The method according to claim 3, wherein: The bypassing of the second set of processing blocks includes forgoing transmitting the one or more signals to a decoder including the second set of processing blocks.
5. The method according to claim 3, further comprising: Further based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance, the one or more signals are transmitted to a decoder including the first set of processing blocks.
6. The method according to claim 5, further comprising: Further based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance, the output of the decoder is applied to the HRTF.
7. The method according to claim 1, wherein Determining the spatial configuration of the virtual environment includes: receiving one or more input sound signals, the one or more input sound signals comprising a first input sound signal from a direct sound source and a second input sound signal from a reflected sound source; modifying the one or more input sound signals to simulate a Doppler effect; applying a delay to the one or more input sound signals; and Positioning the one or more input sound signals on a plurality of virtual speakers, wherein decoding the one or more signals comprises: One or more virtual sounds are determined, wherein the one or more virtual sounds are associated with one or more of the direct sound source, the reflected sound source, and a movement of a user.
8. The method according to claim 1, further comprising: determining whether an energy level of the one or more signals exceeds a predetermined energy level threshold; as well as Based on determining that the energy level exceeds the predetermined energy level threshold, the one or more signals are decoded.
9. The method according to claim 1, wherein: The virtual environment includes a plurality of sound source locations, and wherein the method further comprises: determining whether the number of sound source locations in the virtual environment exceeds a predetermined sound source threshold; and The one or more signals are decoded based on determining that the number of sound source locations does not exceed the predetermined sound source threshold.
10. A system comprising: a wearable head device configured to provide an audio signal to a user; as well as One or more processors configured to perform a method comprising: determining a spatial configuration of the virtual environment, wherein the spatial configuration includes sound source locations and virtual speaker locations; determining one or more signals associated with the sound source location; decoding the one or more signals based on determining that the distance between the sound source position and the virtual speaker position is less than a predetermined distance; and The audio signal is rendered based on the decoded one or more signals.
11. The system according to claim 10, wherein: The method further comprises: further applying the one or more signals to a head-related transfer function (HRTF) based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance; Based on determining that the distance between the sound source position and the virtual speaker position is not less than the predetermined distance, applying the one or more signals to the HRTF is abandoned.
12. The system according to claim 10, wherein: Decoding the one or more signals comprises applying a first set of processing blocks to the one or more signals, and wherein the method further comprises: Further based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance, a second group of processing blocks is bypassed, the second group of processing blocks being associated with one or more inactive virtual speakers.
13. The system according to claim 12, wherein: The bypassing of the second set of processing blocks includes forgoing transmitting the one or more signals to a decoder including the second set of processing blocks.
14. The system according to claim 12, wherein: The method further comprises: Further based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance, the one or more signals are transmitted to a decoder including the first set of processing blocks.
15. The system according to claim 14, wherein: The method further comprises: Further based on determining that the distance between the sound source position and the virtual speaker position is less than the predetermined distance, the output of the decoder is applied to the HRTF.
16. The system according to claim 10, wherein: Determining the spatial configuration of the virtual environment includes: receiving one or more input sound signals, the one or more input sound signals comprising a first input sound signal from a direct sound source and a second input sound signal from a reflected sound source; modifying the one or more input sound signals to simulate a Doppler effect; applying a delay to the one or more input sound signals; and Positioning the one or more input sound signals on a plurality of virtual speakers, wherein decoding the one or more signals comprises: One or more virtual sounds are determined, wherein the one or more virtual sounds are associated with one or more of the direct sound source, the reflected sound source, and a movement of a user.
17. The system according to claim 10, wherein: The method further comprises: determining whether an energy level of the one or more signals exceeds a predetermined energy level threshold; and Based on determining that the energy level exceeds the predetermined energy level threshold, the one or more signals are decoded.
18. The system according to claim 10, wherein: The virtual environment includes a plurality of sound source locations, and wherein the method further comprises: determining whether the number of sound source locations in the virtual environment exceeds a predetermined sound source threshold; and The one or more signals are decoded based on determining that the number of sound source locations does not exceed the predetermined sound source threshold.