Efficient Rendering of Virtual Sound Fields
By dynamically selecting a subset of virtual speakers, the problem of unbalanced computing resources under different sound sources is solved, and the rendering efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN201980048983.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-12
- Filing Date
- 2019-06-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2039-06-12
AI Technical Summary
The existing virtual speaker positioning method uses uneven computing resources and efficiency when processing different numbers of sound sources, resulting in waste of resources in the case of a small number of sound sources or excessive computational complexity in the case of a large number of sound sources.
Using the modified virtual speaker positioning method, a subset of fixed virtual speakers is dynamically selected and used, based on sound source proximity and other factors, unnecessary computing resource consumption is reduced.
By dynamically selecting a virtual speaker subset, the computing complexity and resource consumption are reduced, and the system's rendering efficiency under different sound sources is improved.
Smart Images

Figure CN112470102B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 684,093, filed on Jun. 12, 2018, the content of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to spatial audio rendering and associated systems. More specifically, the present disclosure relates to systems and methods for improving the efficiency of virtual - speaker - based spatial audio systems. Background Art
[0004] Virtual environments are ubiquitous in computing environments and can be found in use in video games (where the virtual environment can represent the game world); maps (where the virtual environment can represent the terrain to be navigated); simulations (where the virtual environment can simulate a real - world environment); digital storytelling (where virtual characters can interact with each other in the virtual environment); and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the technologies used to present virtual environments may limit the user experience in the virtual environment. For example, traditional displays (e.g., 2D displays) and audio systems (e.g., fixed speakers) may not be able to implement the virtual environment in a way that creates a compelling, realistic, and immersive experience.
[0005] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively referred to as “XR”) share the ability to present sensory information corresponding to a virtual environment represented by data in a computer system to a user of an XR system. By combining virtual visual and audio cues with real vision and sound, such systems can provide a uniquely prominent sense of immersion and realism. Thus, it may be desirable to present digital sound to a user of an XR system in such a way that the sound appears to occur naturally in the user's real environment and conforms to the user's expectations of sound. Generally, users expect that virtual sound will have the acoustic characteristics of the real environment in which the sound is heard. For example, a user of an XR system in a large concert hall will expect the virtual sound of the XR system to have a huge, hollow sound quality; in contrast, a user in a small apartment will expect the sound to be softer, closer, and more immediate. Additionally, users expect that virtual sound will be presented without latency.
[0006] Among other techniques, high-fidelity stereo (ambisonics) and non-high-fidelity stereo can be used to generate spatial audio. For a large number of sound source objects, high-fidelity stereo or non-high-fidelity stereo can be an effective method of rendering spatial audio due to their design and architecture. This is especially true when modeling reflections. A spatial audio system based on high-fidelity stereo and non-high-fidelity stereo multi-channels can render an audio signal through several steps. Example steps can include an encoding step for each source, a fixed overhead sound field decoding step, and / or a fixed speaker virtualization step. One or more hardware components can perform these steps.
[0007] In a first method for rendering an audio signal, each sound source can have its own pair of finite impulse response (FIR) filters. In such a system, the perceived position of the sound is changed by varying the filter coefficients of the FIR filters. In some embodiments, each sound can use multiple (e.g., two pairs) of FIR filters. Each pair can use two filters (i.e., four FIR filters). The FIR filters are capable of crossfading as the sound moves in the virtual environment. In some embodiments, each sound can use four FIR filters.
[0008] In a second method for rendering an audio signal, a fixed number of virtual speakers can be used to achieve virtual speaker panning. Each sound source can be panned across the fixed virtual speakers. In some embodiments, each virtual speaker can use multiple (e.g., two) FIR filters. Virtual speaker panning can be effective for certain applications and can use a negligible amount of computational resources.
[0009] In some embodiments, depending on the number of sounds being played simultaneously, a particular method can have improved efficiency compared to another method. For example, 30 sounds can be played simultaneously. If each sound source uses four FIR filters, the first method may require 120 FIR filters (30 sound sources × 4 FIR filters per sound source = 120 FIR filters). If each virtual speaker uses 2 FIR filters, the second method may only require 32 FIR filters (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters).
[0010] As another example, only one sound may be being played. The first method may only require four FIR filters (1 sound source × 4 FIR filters per sound source = 4 FIR filters), while the second method may require 32 FIR filters (16 virtual speakers × 2 FIR filters per virtual speaker = 32 FIR filters).
[0011] As shown in the above examples, the first method can be beneficial for a small number of sounds, and the second method can be beneficial for a large number of sounds. Therefore, an audio system and method that can improve efficiency based on the number of sound sources at a given time may be needed. Summary of the Invention
[0012] An audio system and method for rendering an audio signal using a modified virtual loudspeaker layout are disclosed. The audio system may include a fixed number F of virtual loudspeakers, and the modified virtual loudspeaker layout may dynamically select and use a subset P of the fixed virtual loudspeakers. Each sound source may be positioned on the subset P of virtual loudspeakers. In some embodiments, multiple (e.g., two) FIR filters may be used for each virtual loudspeaker in the subset P. The subset P of virtual loudspeakers may be selected based on one or more factors, such as proximity to the sound source. The subset P of virtual loudspeakers may be referred to as active loudspeakers.
[0013] As an example, the modified virtual loudspeaker layout method may be compared with the first and second methods disclosed above. If three sounds are played simultaneously and the audio system has 16 fixed virtual loudspeakers, the first method may require 12 FIR filters (3 sound sources × 4 FIR filters per sound source = 12 FIR filters), and the second method may require 32 FIR filters (16 virtual loudspeakers × 2 FIR filters per virtual loudspeaker = 32 FIR filters). On the other hand, the modified virtual loudspeaker layout method can dynamically select three virtual loudspeakers as part of the subset P of active virtual loudspeakers. The modified virtual loudspeaker layout method may require six FIR filters, two FIR filters per active virtual loudspeaker (3 virtual loudspeakers × 2 FIR filters = 6 FIR filters). Brief Description of the Drawings
[0014] Figure 1 An example wearable system according to some embodiments is shown.
[0015] Figure 2 An example handheld controller that can be used in conjunction with the example wearable system according to some embodiments is shown.
[0016] Figure 3 An example auxiliary unit that can be used in conjunction with the example wearable system according to some embodiments is shown.
[0017] Figure 4 An example functional block diagram for the example wearable system according to some embodiments is shown.
[0018] Figure 5AShows a block diagram of an example spatial audio system according to some embodiments.
[0019] Figure 5B Shows an example method flow for operating a Figure 5A system according to some embodiments.
[0020] Figure 5C Shows an example method flow for operating an example decoder / virtualizer according to some embodiments.
[0021] Figure 6 Shows an example configuration of sound sources and speakers according to some embodiments.
[0022] Figure 7A Shows a block diagram of an example decoder / virtualizer including multiple detectors according to some embodiments.
[0023] Figure 7B Shows an example method flow for operating a Figure 7A decoder / virtualizer according to some embodiments.
[0024] Figure 8A Shows a block diagram of an example decoder / virtualizer according to some embodiments.
[0025] Figure 8B Shows an example method flow for operating a Figure 8A decoder / virtualizer according to some embodiments.
[0026] Figure 9 Shows an example configuration of sound sources and speakers according to some embodiments.
[0027] Figure 10A Shows a block diagram of an example decoder / virtualizer used in a system including active speakers according to some embodiments.
[0028] Figure 10B Shows an example method flow for operating a Figure 10A decoder / virtualizer according to some embodiments. Detailed Description
[0029] In the following description of the examples, reference is made to the accompanying drawings which form a part hereof, and in which are shown by way of illustration specific examples that may be practiced. It is to be understood that other examples may be used and structural changes may be made without departing from the scope of the disclosed examples.
[0030] Example Wearable System
[0031] Figure 1FIG. 0 shows an example wearable head-mounted device 100 configured to be worn on a user's head. The wearable head-mounted device 100 can be part of a broader wearable system that includes one or more components such as a head-mounted device (e.g., the wearable head-mounted device 100), a handheld controller (e.g., the handheld controller 200 described below), and / or an auxiliary unit (e.g., the auxiliary unit 300 described below). In some examples, the wearable head-mounted device 100 can be used in virtual reality, augmented reality, or mixed reality systems or applications. The wearable head-mounted device 100 can include one or more displays such as displays 110A and 110B (which can include left and right transmissive displays and associated components for coupling light from the displays to the user's eyes such as ortho-pupil expansion (OPE) grating sets 112A / 112B and exit-pupil expansion (EPE) grating sets 114A / 114B); left and right acoustic structures such as speakers 120A and 120B (which can be mounted on temple arms 122A and 122B respectively and positioned adjacent to the user's left and right ears); one or more sensors such as infrared sensors, accelerometers, GPS units, inertial measurement units (IMUs) (e.g., IMU 126), acoustic sensors (e.g., microphone 150); an ortho-coil electromagnetic receiver (e.g., the receiver 127 shown mounted to the left temple arm 122A); left and right cameras oriented away from the user (e.g., depth (time-of-flight) cameras 130A and 130B); and left and right eye cameras oriented towards the user (e.g., for detecting the user's eye movements) (e.g., eye cameras 128 and 128B). However, the wearable head-mounted device 100 can incorporate any suitable display technology, as well as any suitable number, type, or combination of sensors, or other components without departing from the scope of the present invention. In some examples, the wearable head-mounted device 100 can incorporate one or more microphones 150 configured to detect audio signals generated by the user's voice; such microphones can be positioned in the wearable head-mounted device adjacent to the user's mouth. In some examples, the wearable head-mounted device 100 can incorporate networking features (e.g., Wi-Fi capability) to communicate with other devices and systems including other wearable systems. The wearable head-mounted device 100 can further include components such as a battery, a processor, a memory, a storage unit, or various input devices (e.g., buttons, touchpads); or can be coupled to a handheld controller (e.g., the handheld controller 200) or an auxiliary unit (e.g., the auxiliary unit 300) that includes one or more such components. In some examples, the sensors can be configured to output a set of coordinates of the head-mounted unit relative to the user's environment and can provide input to a processor that performs a simultaneous localization and mapping (SLAM) process and / or a visual odometry algorithm.In some examples, as further described below, the wearable head-mounted device 100 can be coupled to the handheld controller 200 and / or the auxiliary unit 300.
[0032] Figure 2 FIG. shows an example mobile handheld controller assembly 200 of an example wearable system. In some examples, the handheld controller 200 can communicate with the wearable head-mounted device 100 and / or the auxiliary unit 300 described below, either wired or wirelessly. In some examples, the handheld controller 200 includes a handle portion 220 to be held by a user, and one or more buttons 240 disposed along the top surface 210. In some examples, the handheld controller 200 can be configured to serve as an optical tracking target; for example, sensors (e.g., cameras or other optical sensors) of the wearable head-mounted device 100 can be configured to detect the position and / or orientation of the handheld controller 200, and thus, by extension, can indicate the position and / or orientation of the hand of the user holding the handheld controller 200. In some examples, as described above, the handheld controller 200 can include a processor, a memory, a storage unit, a display, or one or more input devices. In some examples, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head-mounted device 100). In some examples, the sensors can detect the position or orientation of the handheld controller 200 relative to the wearable head-mounted device 100 or relative to another component of the wearable system. In some examples, the sensors can be located in the handle portion 220 of the handheld controller 200, and / or can be mechanically coupled to the handheld controller. The handheld controller 200 can be configured to provide one or more output signals, e.g., corresponding to the pressed state of the button 240; or, the position, orientation, and / or motion of the handheld controller 200 (e.g., via an IMU). Such output signals can be used as inputs to the processor of the wearable head-mounted device 100, the auxiliary unit 300, or another component of the wearable system. In some examples, the handheld controller 200 can include one or more microphones to detect sounds (e.g., the user's voice, ambient sounds), and in some cases, provide signals corresponding to the detected sounds to a processor (e.g., the processor of the wearable head-mounted device 100).
[0033] Figure 3An example auxiliary unit 300 of an example wearable system is shown. In some examples, the auxiliary unit 300 may communicate with the wearable head-mounted device 100 and / or the handheld controller 200 in a wired or wireless manner. The auxiliary unit 300 may include a battery to provide energy to operate one or more components of the wearable system, such as the wearable head-mounted device 100 and / or the handheld controller 200 (including a display, sensors, acoustic structures, processors, microphones, and / or other components of the wearable head-mounted device 100 or the handheld controller 200). In some examples, as described above, the auxiliary unit 300 may include a processor, a memory, a storage unit, a display, one or more input devices, and / or one or more sensors. In some examples, the auxiliary unit 300 includes a clip 310 for attaching the auxiliary unit to a user (e.g., a belt worn by the user). The advantage of using the auxiliary unit 300 to house one or more components of the wearable system is that doing so allows large or heavy components to be carried on the user's waist, chest, or back (which are relatively well-suited for supporting larger and heavier objects), rather than being mounted on the user's head (e.g., if housed in the wearable head-mounted device 100) or carried by the user's hand (e.g., if housed in the handheld controller 200). This may be particularly advantageous for relatively heavy or bulky components, such as batteries.
[0034] Figure 4 An example functional block diagram is shown that may correspond to an example wearable system 400 (such as may include the example wearable head-mounted device 100, the handheld controller 200, and the auxiliary unit 300 described above). In some examples, the wearable system 400 may be used for virtual reality, augmented reality, or mixed reality applications. As Figure 4As shown, the wearable system 400 may include an example handheld controller 400B, herein referred to as a "totem" (and may correspond to the handheld controller 200 described above); the handheld controller 400B may include a six degrees of freedom (6DOF) totem-to-headgear subsystem 404A. The wearable system 400 may also include an example wearable head device 400A (which may correspond to the wearable helmet device 100 described above); the wearable head device 400A includes a 6DOF headgear-to-headgear subsystem 404B. In this example, the 6DOF totem subsystem 404A and the 6DOF headgear subsystem 404B together determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offsets in three translational directions and rotations about three axes). The six degrees of freedom may be expressed relative to the coordinate system of the wearable head device 400A. In such a coordinate system, the three translational offsets may be expressed as X, Y, and Z offsets, may be expressed as a translation matrix, or some other representation. The rotational degrees of freedom may be expressed as a sequence of yaw, pitch, and roll rotations; expressed as a vector; expressed as a rotation matrix; expressed as a quaternion; or expressed as some other representation. In some examples, one or more depth cameras 444 (and / or one or more non-depth cameras) included in the wearable head device 400A; and / or one or more optical sights (e.g., the button 240 of the handheld controller 200 described above, or a dedicated optical sight included in the handheld controller) may be used for 6DOF tracking. In some examples, as described above, the handheld controller 400B may include a camera; and the helmet 400A may include an optical sight for use with the camera for optical tracking. In some examples, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids for wirelessly transmitting and receiving three distinguishable signals. By measuring the relative amplitudes of the three distinguishable signals received in each coil for reception, the 6DOF of the handheld controller 400B relative to the wearable head device 400A can be determined. In some examples, the 6DOF totem subsystem 404A may include an inertial measurement unit (IMU), which may be used to provide improved accuracy and / or more timely information regarding the fast-moving handheld controller 400B.
[0035] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head-mounted device 400A) to an inertial coordinate space or an environmental coordinate space. For example, such a transformation may be necessary for the display of the wearable head-mounted device 400A to present virtual objects (e.g., a virtual person sitting on a real chair and facing forward, regardless of the position and orientation of the wearable head-mounted device 400A) at an expected position and orientation relative to the real environment rather than at a fixed position and orientation on the display (e.g., at the same position in the display of the wearable head-mounted device 400A). This can maintain the illusion that the virtual object exists in the real environment (and, for example, does not appear unnaturally positioned in the real environment as the wearable head-mounted device 400A moves and rotates). In some examples, a compensating transformation on the coordinate space can be determined by processing images from the depth camera 444 (e.g., using Simultaneous Localization and Mapping (SLAM) and / or visual odometry processes) to determine the transformation of the wearable head-mounted device 400A relative to an inertial or environmental coordinate system. In Figure 4 the example shown in, the depth camera 444 can be coupled to the SLAM / visual odometry module 406 and can provide images to the module 406. The SLAM / visual odometry module 406 can be implemented to include a processor that is configured to process the images and determine the position and orientation of the user's head, which can then be used to identify the transformation between the head coordinate space and the actual coordinate space. Similarly, in some examples, additional information sources regarding the user's head pose and position are obtained from the IMU 409 of the wearable head-mounted device 400A. The information from the IMU 409 can be integrated with the information from the SLAM / visual odometry module 406 to provide improved accuracy and / or more timely information for a quick adjustment of the user's head pose and position.
[0036] In some examples, the depth camera 444 can provide 3D images to a gesture tracker 411, which can be implemented in the processor of the wearable head-mounted device 400A. The gesture tracker 411 can identify the user's gestures, for example, by matching the 3D images received from the depth camera 444 with stored patterns representing gestures. Other suitable techniques for identifying user gestures will be apparent.
[0037] In some examples, one or more processors 416 may be configured to receive data from the helmet subsystem 404B, IMU 409, SLAM / visual odometry module 406, depth camera 444, microphone (not shown), and / or gesture tracker 411. The processor 416 may also send and receive control signals from the 6DOF totem system 404A. In examples such as where the handheld controller 400B is untethered, the processor 416 may be wirelessly coupled to the 6DOF totem system 404A. The processor 416 may further communicate with additional components such as the audio-visual content memory 418, graphics processing unit (GPU) 420, and / or digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to the head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to the left source 424 of the imaging modulated light and a right channel output coupled to the right source 426 of the imaging modulated light. The GPU 420 may output stereoscopic image data to the sources of the imaging modulated light 424, 426. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or right speaker 414. The DSP audio spatializer 422 may receive an input from the processor 416 indicating a direction vector from the user to a virtual sound source (which may be moved by the user, for example, via the handheld controller 400B). Based on the direction vector, the DSP audio spatializer 422 may determine a corresponding HRTF (e.g., by accessing the HRTF, or by interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. By incorporating the user's relative position and orientation with respect to the virtual sound in the mixed reality environment, that is, by presenting a virtual sound that matches the user's expectation of a real sound that would sound like it is in the real environment, the believability and realism of the virtual sound can be enhanced.
[0038] In some examples, such as Figure 4 as shown, one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / video content memory 418 may be included in the auxiliary unit 400C (which may correspond to the auxiliary unit 300 described above). The auxiliary unit 400C may include a battery 427 to power its components and / or to power the wearable head device 400A and / or the handheld controller 400B. Including such components in an auxiliary unit that can be mounted to the user's waist may limit the size and weight of the wearable head device 400A, which in turn may reduce fatigue in the user's head and neck.
[0039] AlthoughFigure 4 shows elements corresponding to the various components of the example wearable system 400, but various other suitable arrangements of these components will be apparent to those skilled in the art. For example, the elements presented in Figure 4 may alternatively be associated with the wearable head device 400A or the handheld controller 400B. Additionally, some wearable systems may entirely dispense with the handheld controller 400B or the auxiliary unit 400C. Such changes and modifications should be understood to be included within the scope of the disclosed examples.
[0040] Mixed Reality Environment
[0041] Like all people, the users of a mixed reality system also exist in the real environment, that is, the three-dimensional part of the "real world" that the user can perceive and all its contents. For example, the user perceives the real environment using ordinary human senses (vision, sound, touch, taste, smell) and interacts with the real environment by moving their body within the real environment. A location in the real environment can be described as coordinates in a coordinate space; for example, the coordinates can include latitude, longitude, and altitude relative to sea level; distances in three orthogonal dimensions from a reference point; or other suitable values. Similarly, a vector can describe a quantity that has a direction and a magnitude in the coordinate space.
[0042] A computing device may maintain a representation of a virtual environment in a memory associated with the device, for example. As used herein, a virtual environment is a computational representation of a three-dimensional space. The virtual environment may include representations of any objects, actions, signals, parameters, coordinates, vectors, or other characteristics associated with that space. In some examples, circuitry (e.g., a processor) of the computing device may maintain and update the state of the virtual environment; that is, the processor may determine the state of the virtual environment at a second time based on data associated with the virtual environment and / or user-provided input at a first time. For example, if an object in the virtual environment is located at a first coordinate at a given time and has certain programmed physical parameters (e.g., mass, coefficient of friction); and an input is received from a user indicating that a force should be applied to the object in a direction vector; then the processor may apply the laws of kinematics to determine the position of the object at that time using basic mechanics. The processor may use any known appropriate information about the virtual environment and / or any appropriate input to determine the state of the virtual environment at that time. When maintaining and updating the state of the virtual environment, the processor may execute any appropriate software, including software related to creating and deleting virtual objects in the virtual environment; software for defining the behavior of virtual objects or characters in the virtual environment (e.g., scripts); software for defining the behavior of signals (e.g., audio signals) in the virtual environment; software for creating and updating parameters associated with the virtual environment; software for generating audio signals in the virtual environment; software for processing input and output; software for implementing network operations; software for applying asset data (e.g., animation data to move virtual objects over time); or many other possibilities.
[0043] An output device (such as a display or a speaker) can present any or all aspects of the virtual environment to the user. For example, the virtual environment can include virtual objects that can be presented to the user (which can include representations of inanimate objects, people, animals, lights, etc.). The processor can determine a view of the virtual environment (e.g., corresponding to a “camera” with origin coordinates, view axes, and a frustum); and render the visible scene of the virtual environment corresponding to that view to the display. Any suitable rendering technique can be used for this purpose. In some examples, the visible scene can include only some of the virtual objects in the virtual environment and exclude certain other virtual objects. Similarly, the virtual environment can include audio aspects that can be presented to the user as one or more audio signals. For example, a virtual object in the virtual environment can generate a sound originating from the object's position coordinates (e.g., a virtual character can speak or cause a sound effect); or the virtual environment can be associated with music cues or ambient sounds that may or may not be associated with a specific location. The processor can determine the audio signals corresponding to the “listener” coordinates, e.g., the audio signals corresponding to sound synthesis in the virtual environment, and mix and process to simulate the audio signals that would be heard by a listener at the listener coordinates, and present the audio signals to the user via one or more speakers.
[0044] Since the virtual environment exists only as a computational construct, the user cannot directly perceive the virtual environment using ordinary senses. Instead, the user can only indirectly perceive the virtual environment presented to the user, for example, through a display, a speaker, a haptic output device, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment; but can provide input data to the processor via an input device or a sensor, and the processor can use the device or sensor data to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use this data to make the object respond accordingly in the virtual environment.
[0045] Digital Reverb and Ambient Audio Processing
[0046] The XR system can present an audio signal to the user that originates from a sound source with origin coordinates and propagates in the system in a direction with an orientation vector. The user can perceive these audio signals as if they were real audio signals originating from the origin coordinates of the sound source and propagating along the orientation vector.
[0047] In some cases, the audio signals can be considered virtual because they correspond to computational signals in the virtual environment and do not necessarily correspond to real sounds in the real environment. However, virtual audio signals can be, for example, as via Figure 1The real audio signals detectable by the human ear generated by the speakers 120A and 120B of the wearable head device 100 in [0] are presented to the user.
[0048] Advantages of the embodiments disclosed below include reduced network bandwidth, reduced power consumption, reduced computational complexity, and reduced computational latency. These advantages are particularly important for mobile systems, including wearable systems, where processing resources, network resources, battery capacity, and physical size and weight are typically very valuable.
[0049] In a dynamic environment such as AR, the system can continuously render audio signals. Rendering audio signals using all virtual speakers can particularly lead to high computational power, a large amount of processing, high network bandwidth, high power consumption, etc. Therefore, it may be necessary to use a modified virtual speaker placement to dynamically select and use a subset of fixed virtual speakers based on one or more factors.
[0050] Example Spatial Audio System
[0051] Figure 5A A block diagram of an example spatial audio system according to some embodiments is shown. Figure 5B Shows an example method for operating Figure 5A the system.
[0052] The spatial audio system 500 can include a spatial modeler 510, an internal spatial representation 530, and a decoder / virtualizer 540A. The spatial modeler 510 can include a direct path section 512, one or more reflection sections 520 (optional), and a spatial encoder 526. The spatial modeler 510 can be configured to model a virtual environment. The direct path section 512 can include a direct source 514, and optionally, a Doppler 516. The direct source 514 can be configured to provide an audio signal (step 552 of process 550). The Doppler 516 can receive a signal from the direct source 514 and can be configured to introduce a Doppler effect in its input signal (step 554). For example, the Doppler 516 can change the pitch of the sound source (e.g., pitch shift) to change with respect to the motion of the sound source, the user of the system, or both.
[0053] The reflection unit 520 may include a sound reflector 522, an optional Doppler 516, and a delay unit 524. The sound reflector 522 may be configured to introduce reflections in its signal (step 556). The introduced reflections may represent one or more properties of the environment. The Doppler 516 in the reflection unit 520 may receive the signal from the sound reflector 522 and may be configured to introduce the Doppler effect into its input signal (step 558). The delay unit 524 may receive the signal from the Doppler 516 and may be configured to introduce a delay (step 560).
[0054] The spatial encoder 526 may receive signals from the direct path unit 512 and the (multiple) reflection units 520. In some embodiments, the signal from the direct path unit 512 to the spatial encoder 526 may be the output signal of the Doppler 516 from the direct path unit 512. In some embodiments, the (multiple) signals from the (multiple) reflection units 520 to the spatial encoder 526 may be the (multiple) output signals of the (multiple) delay units 524 from the (multiple) reflection units 520.
[0055] The spatial encoder 526 may include one or more M-way pans 528. In some embodiments, each input received by the spatial encoder 526 may be associated with a unique M-way pan 528. "Panning" may refer to distributing a signal across multiple speakers, multiple locations, or both. The M-way pan 528 may be configured to distribute its input signal across multiple virtual speakers (step 562). For example, the M-way pan 528 may distribute its input signal across all M virtual speakers. For example, as Figure 5A shown, M may be equal to four, and each M-way pan 528 may be configured to distribute its input signal across four virtual speakers. Although the figure shows a system with four virtual speakers, examples of the present disclosure may include any number of virtual speakers.
[0056] As an example, an automotive system may include a left speaker and a right speaker. Sound in such a system may be panned across the left and right speakers of the vehicle by splitting the sound into two (one for each speaker). The scaling volume of each speaker may be set according to the configuration of the two speakers, and the result may be sent to the left and right speakers.
[0057] As another example, a surround sound system may include multiple speakers, such as six speakers. Sound in such a system may be panned across the six speakers as stereo. The sound may be split into six (as opposed to two in the automotive system example), the scaling volume of each speaker may be set according to the configuration of the six speakers, and the result may be sent to the six speakers.
[0058] For example, the first M-way positioner 528 can receive the output of the Doppler 516 of the direct path 512, while the other M-way positioners 528 can receive the output of the reflector 520. Each M-way positioner 528 can split its input signal such that it can be distributed over multiple outputs. In this way, each M-way positioner 528 can have a greater number of outputs than inputs.
[0059] The spatial modeler 510 can output a signal to the internal spatial representation 530 (step 564). In some embodiments, the output(s) from the spatial modeler 510 can include the output of each M-way positioner 528. The internal spatial representation 530 can be configured to represent the spatial configuration of the virtual environment (step 566). One example representation can include representing the relative positions of the user, the sound source(s), and the virtual speaker(s). In some embodiments, the internal spatial representation 530 can output one or more signals representing the rotational head pose, translational head pose, sound field decoding, one or more head-related transfer functions (HRTFs), or a combination thereof of the user of the system 500. In some embodiments, the internal spatial representation 530 can be a representation of a non-high-fidelity stereo multi-channel based system, a high-fidelity stereo / wave field based system, etc. An exemplary high-fidelity stereo / wave field based system can be higher-order ambisonics (HOA).
[0060] The internal spatial representation 530 can output its signal 552 to the decoder / virtualizer 540A (step 568). The decoder / virtualizer 540 can decode its input signal and introduce virtualized sound into the signal (step 570). Step 570 can include multiple sub-steps and will be discussed in more detail below. Then, the system outputs the signal from the decoder / virtualizer 540 as the left signal 502L to the left speaker and as the right signal 502R to the right speaker (step 580).
[0061] The system 500 can include any number of different types of decoder / virtualizers 540. Figure 5A An example decoder / virtualizer 540A is shown. Other example decoder / virtualizers 540 will be discussed below.
[0062] The decoder / virtualizer 540A can include a rotation / translation representation 542, a sound field decoder 544, one or more HRTFs 546, and one or more combiners 548. Figure 5CShows a flow of an example method for operating an example decoder / virtualizer, which may be referred to as step 570-1. The rotation / translation representation 542 may receive signals from the inner space representation 530 and may be configured to introduce a representation of motion associated with the audio signal. For example, the motion may be the motion of the sound source, the user, or both (step 572). The rotation / translation representation 542 may output the signals to the sound field decoder 544. The sound field decoder 544 may receive signals from the rotation / translation representation 542 and may be configured to decode the signals (step 574). Each HRTF 546 may receive signals from the sound field decoder 544. Each HRTF 546 may be configured to determine the HRTF corresponding to its input signal and apply it to the signal (step 576). One or more HRTFs 546 may be collectively referred to as a speaker virtualizer. In some embodiments, the HRTF 546 may be configured for finite impulse response (FIR) filtering. Each combiner 548 may receive and combine signals from the HRTFs 546 (step 578).
[0063] In some embodiments, the decoder / virtualizer 540A may represent the "baseline" processing overhead. The baseline processing overhead may be complex, involving matrix calculations and long FIR filters to apply HRTF processing to each virtual speaker.
[0064] The output from the combiner 548 may be the output signal of the system 500. In some embodiments, the output signal 502 from the system 500 may be an audio signal for the left and right speakers (e.g., Figure 1 speakers 120A and 120B).
[0065] In some cases, when the number of sound sources for playback is large, Figure 5A a spatial audio system may be beneficial. However, in some cases, when the number of sound sources for playback is small, Figure 5A a spatial audio system may not be beneficial. It may be desirable to utilize the efficiency of a spatial audio system based on non-high-fidelity stereo multi-channel or a spatial audio system based on high-fidelity stereo (such as Figure 5A system 500) in a manner that is effective for the case when the number of sound sources for playback is small.
[0066] There can be multiple ways to use sound field synthesis and decoding to improve spatialization efficiency. The first method can be low-energy speaker detection and culling. In low-energy speaker detection and culling, if the energy output of the virtual speaker channel of a spatial audio system based on non-fidelity stereo multi-channel or the high-fidelity stereo / sound field channel of a spatial audio system based on high-fidelity stereo is less than a predetermined threshold, then the processing of the signal from the virtual speaker channel is not performed. In some embodiments, for example, before performing sound field decoding on the signal from a given virtual speaker, the system can determine whether the output of the given virtual speaker is higher than the predetermined threshold. Low-energy speaker detection and culling will be discussed in more detail below.
[0067] The second method of using sound field synthesis and decoding to improve spatialization efficiency can be source geometry-based virtual speaker selection. In source geometry-based virtual speaker selection, decoder / virtualizer processing can be selectively disabled. The selective disabling (or selective enabling) can be based on the position of the sound source(s) relative to the user / listener. Source geometry-based virtual speaker selection will be discussed in detail below.
[0068] The third approach can be to combine low-energy speaker detection and culling techniques with source-virtual speaker coupling techniques.
[0069] The spatial modeler 510 can have a computational complexity, which can represent the number of operations required to process an audio signal. The computational complexity can be proportional to M multiplied by N, where M can be equal to the number of sound sources (including direct sound sources and optional reflections), and N can be equal to the number of channels required to represent the high-fidelity sound field. In some embodiments, N can be equal to (O + 1) 2 , where O is the order of the high-fidelity stereo used.
[0070] The decoder / virtualizer 540 can have a computational complexity proportional to nVS, where nVS is the number of virtual speakers. The computational power of each speaker can be high and is typically composed of a pair of FIR filters, usually implemented using the fast Fourier transform (FFT) or inverse FFT (IFFT), both of which are computationally expensive processes.
[0071] Example Low Energy Output Detection and Selection Method
[0072] In some embodiments, some virtual speakers may have little or no signal input energy; for example, when the number of sound sources in a spatial audio system is small. Speaker virtualization processing can be a computationally expensive (e.g., CPU-intensive) process. For example, if there is a sound source located at zero azimuth (e.g., directly in front of the user), there may be little or no energy in the signals from virtual speakers located between 90 degrees and 270 degrees azimuth (e.g., behind the user). Low-energy signals may not have a significant impact on the perceived location of the sound source, and thus performing speaker virtualization processing on low-energy signals and / or determining the characteristics of the corresponding virtual speakers may be computationally inefficient.
[0073] To reduce the computational resources required, a system employing a low-energy output detection and selection method can include a detector located between the sound field decoder and the HRTF. Alternatively, the detector can be located between the multi-channel output and the HRTF. The detector can be configured to detect one or more energy levels associated with one or more audio signals from one or more virtual speakers.
[0074] If the energy level of the signal from virtual speaker Vn is less than the energy threshold α, the signal can be considered a low-energy signal. Based on the detected energy level of the audio signal being less than the energy threshold α, the HRTF block and its processing of the low-energy signal can be bypassed.
[0075] The determination of the energy level of a signal can use a variety of techniques. For example, the RMS algorithm can be applied to the signal routed to the virtual speaker to measure its energy. Similar to the "attack" and "release" times used in traditional audio compressors, these can be used to prevent the signal of the speaker from suddenly "jumping" in and out.
[0076] Figure 6Shows an example configuration of a sound source and speakers according to some embodiments. System 600 may include a sound source 620 and a plurality of speakers. The plurality of speakers 622 may include one or more active virtual speakers 622A and one or more inactive virtual speakers 622B. An active virtual speaker 622A may be a virtual speaker whose signal is processed by the HRTF 546 at a given time. An inactive virtual speaker 622B may be a speaker whose signal does not need to be processed by the HRTF 546, for example because its signal has been processed at a previous time, or because the system determines that the signal from the virtual speaker 622B does not need to be processed. M may represent the number of sound sources being played, and N may represent the number of virtual speakers in the system. Although the figure shows a single sound source, examples of the present disclosure may include any number of sound sources. Although the figure shows eight sound sources, examples of the present disclosure may include any number of sound sources, such as 16 (N = 16).
[0077] As an example, as shown, system 600 may include a single (M = 1) sound source 620 and 8 virtual speakers 622. In a given situation, most of the energy is output only on three virtual speakers. That is, system 600 may have three active virtual speakers at a first time. For example, virtual speakers 622A-1, 622A-2, and 622-3 may be active virtual speakers. In some embodiments, the active virtual speaker 622A may be the speaker closest to the sound source 620. Additionally, system 600 may include five inactive virtual speakers 622B. System 600 may determine that the energy level from each of the five inactive virtual speakers is less than an energy threshold, and based on this determination, may bypass the HRTF processing of the signals from the five inactive virtual speakers 622B.
[0078] System 600 may also determine that the energy level from each active virtual speaker is not less than the energy threshold, and based on this determination, may perform HRTF processing on the signals from the three active virtual speakers 622A.
[0079] System 600 may output two signals, one for the right speaker and one for the left speaker, for example, the right signal 502R and the left signal 502L as shown. The reduction in the number of HRTF operations due to bypassing the HRTF processing may be equal to the number of inactive virtual speakers multiplied by the number of signals output from the system. In the Figure 5A example, since the HRTF processing of five signals is bypassed, 10 (five inactive virtual speakers × two output signals) HRTF operations may be saved. Figure 6 example, since the HRTF processing of five signals is bypassed, 10 (five inactive virtual speakers × two output signals) HRTF operations may be saved.
[0080] As another example, if the system includes 16 virtual speakers, where 13 are inactive virtual speakers, the number of HRTF operations saved can be equal to 26 (16 virtual speakers × two output signals).
[0081] Figure 7A A block diagram of an example decoder / virtualizer including multiple detectors is shown in accordance with some embodiments. Figure 7B An example method for operating a decoder / virtualizer in accordance with some embodiments is shown. Figure 7A of the decoder / virtualizer 540A (shown in Figure 5A ), the decoder / virtualizer 540B may be included in the system 500, as described below. Instead of step 570-1 (as shown in Figure 5C ), step 570-2 may be included in the process 550.
[0082] The decoder / virtualizer 540B can include a rotation / translation representation 542, an acoustic field decoder 544, one or more detectors 710, one or more switches 712, one or more HRTFs 546, and one or more combiners 548. The decoder / virtualizer 540B can receive (a) signal(s) 552 from an internal spatial representation 530 (shown in Figure 5A ). The rotation / translation representation 542 can receive a signal from the internal spatial representation 530 and can be configured to introduce a representation of the motion of (a) sound source(s), the user, or both (step 772). The rotation / translation representation 542 can output (a) signal(s) to the acoustic field decoder 544. The acoustic field decoder 544 can receive a signal from the rotation / translation representation 542 and can be configured to decode the signal (step 774). The acoustic field decoder 544 can output the signal to (a) detector(s) 710.
[0083] (A) detector(s) 710 can receive a signal from the acoustic field decoder 544 and can be configured to determine the energy level of its input signal (step 776). Each detector 710 can be coupled to a unique switch 712. If the energy level of the input signal (from the acoustic field decoder 544) is greater than or equal to an energy threshold (step 778), the switch 712 can close the loop, thereby routing its input signal (from the detector 710) to the HRTF 546 coupled to the switch (step 780). Each HRTF determines the corresponding HRTF and applies it to the signal (step 782).
[0084] If the energy level of the input signal is less than the energy threshold, the switch 712 can open such that its input signal (from the detector 710) is not coupled to the corresponding HRTF 546. Accordingly, the corresponding HRTF 546 can be bypassed (step 784).
[0085] Signals from the (multiple) HRTFs 546 can be output to the combiner 548 (step 786). The combiner 548 can be configured to combine (e.g., add, sum, etc.) the signals from the (multiple) HRTFs 546. Those signals that bypass the HRTF 546 may not be combined by the combiner 548. The output from the combiner 548 can be the output signal that forms the output of the system 500. In some embodiments, the output signal 502 from the system 500 can be an audio signal for the left and right speakers (e.g., Figure 1 speakers 120A and 120B).
[0086] In some embodiments, each detector 710 can be coupled to a unique signal corresponding to a virtual speaker. In this way, the processing for each virtual speaker 622 can be performed independently (i.e., the processing for one speaker (e.g., 622A-1) can occur without affecting the processing for another speaker (e.g., 622B)).
[0087] In some embodiments, the type of decoder / virtualizer 540 can depend on the number of sound sources. For example, if the number of sound sources is less than or equal to a predetermined sound source threshold, then Figure 7A the decoder / virtualizer 540B can be included in the system 500. In this case, the signal from the sound field decoder 544 can be input to the (multiple) detectors 710.
[0088] If the number of sound sources is greater than the predetermined sound source threshold, then Figure 5A the decoder / virtualizer 540A can be included in the system. In this case, the signal from the sound field decoder 544 can be input to the HRTF 546.
[0089] In some embodiments, the system can include a decoder / virtualizer 540 that can select whether to perform or bypass the detector and its energy level detection. Figure 8A A block diagram of an example decoder / virtualizer according to some embodiments is shown. Figure 8B An example method flow for operating a Figure 8A decoder / virtualizer according to some embodiments is shown. In some embodiments, instead of the decoder / virtualizer 540A ( Figure 5A shown in Figure 7A ) and the decoder / virtualizer 540B ( Figure 5C shown in
[0090] The decoder / virtualizer 540C can include a rotation / translation representation 542, a sound field decoder 544, one or more detectors 710, one or more first switches 712, one or more HRTFs 546, and one or more combiners 548, similar to the decoder / virtualizer 540B discussed above. Steps 872, 874, and 882 can be correspondingly similar to steps 772, 774, and 782 discussed above.
[0091] The decoder / virtualizer 540C can also include a second switch 814. The second switch 814 can be configured to open or close a first loop from the sound field decoder 544 to the detector(s) 710 and the first switch(es) 712. Additionally or alternatively, the second switch 814 can be configured to open or close a second loop that bypasses the detector(s) 710 and the first switch(es) 712 from the system 500. In some embodiments, the second switch 814 can be a two-way switch configured to select between directly delivering a signal to the detector 710 (the first loop) or directly delivering the signal to the HRTF 546 (the second loop).
[0092] For example, the system can determine whether the number of sound sources is greater than or equal to a predetermined sound source threshold (step 876). If the number of sound sources is greater than or equal to the predetermined sound source threshold, the second switch 814 can close the second loop and cause the signal from the sound field decoder 544 to be directly delivered to the HRTF 546 (step 878). Then, each HRTF 546 determines the corresponding HRTF and applies it to the signal (step 880). When there are more sound sources, the likelihood of signals with a low energy level may decrease.
[0093] On the other hand, if the number of sound sources is less than the predetermined sound source threshold, the signal is more likely to have a low energy level. Thus, the second switch 814 can close the first loop and cause the signal from the sound field decoder 544 to be directly delivered to the detector(s) 710 (step 882). The detector(s) 710 can receive the signal from the sound field decoder 544 and can be configured to determine the energy level of its input signal (step 884). If the energy level of the input signal (from the sound field decoder 544) is greater than or equal to the energy threshold (step 886), the switch 712 can close the loop, thereby routing its input signal (from the detector 710) to the HRTF 546 coupled to it by the switch (step 888). If the energy level of the input signal is less than the energy threshold, the switch 712 can open so that its input signal (from the detector 710) is not coupled to the corresponding HRTF 546, thereby bypassing the HRTF 546 (step 890).
[0094] Signals from the (multiple) HRTFs 546 can be output to the combiner 548 (step 892).
[0095] In some embodiments, one or more energy threshold detections can be activated in response to energy. In some embodiments, one or more energy threshold detections can be activated in response to amplitude and can be subject to traditional attacks, release times, etc.
[0096] Example Source-Geometry-Based Loudspeaker Selection Method
[0097] Virtual speaker selection based on source geometry can be another way to reduce CPU consumption. In some embodiments, virtual speaker selection based on source geometry can include selectively disabling decoder / virtualizer processing (e.g., Figure 5A decoder / virtualizer 540A of Figure 7A decoder / virtualizer 540B of Figure 8A decoder / virtualizer 540C of, etc.). In some embodiments, selective disabling (or selective enabling) can be based on the (multiple) positions of the (multiple) sound sources relative to the user / listener. In some embodiments, selective disabling of decoder / virtualizer processing can include bypassing all of the processing blocks of the decoder / virtualizer.
[0098] Using virtual speaker selection based on source geometry, a high-fidelity stereo output can be calculated. If the high-fidelity stereo output requires a large amount of energy for decoding, it may be beneficial to use a simpler method (requiring less CPU consumption), such as a real-time energy detection method. Additionally, in some embodiments, the real-time energy detection method can perform calculations less frequently.
[0099] Figure 9 An example configuration of a sound source and speakers according to some embodiments is shown. System 900 can include a sound source 920 and multiple speakers. Compared with Figure 6 system 600 of Figure 6 the sound source 920 can be located at a second position, which can be different from Figure 6 the first position of the sound source 620 of
[0100] The inactive virtual speaker 922C may differ from the inactive virtual speaker 922B in that the virtual speaker 922C may be active at a first time, but its signal is processed at a second time (e.g., the ringing output period). In Figure 9 the example of, the sound source 920 may have moved from a first position (e.g., close to the virtual speaker 922C) to a second position (e.g., not close to the virtual speaker 922). Due to the movement of the sound source, the two virtual speakers may no longer have the sound sources that are mixed into them at the second time. Due to the filtering process of the two virtual speakers, the two virtual speakers may need to be active in a subsequent frame (e.g., the second time) to correctly complete the filtering process.
[0101] In some embodiments, the system may include a decoder / virtualizer 540 in a system using active virtual speakers. Figure 10A A block diagram of an example decoder / virtualizer used in a system including active speakers according to some embodiments is shown. Figure 10B An example method for operating Figure 10A the decoder / virtualizer is shown in a flow. In some embodiments, instead of the decoder / virtualizer 540A (shown in Figure 5A ), the decoder / virtualizer 540B (shown in Figure 7A ), and the decoder / virtualizer 540C (shown in Figure 8A ), the decoder / virtualizer 540D may be included in the system 500. Instead of the steps 570-1 (shown in Figure 5C ), the step 570-2 (shown in Figure 7B ), and the step 570-3 (shown in Figure 8B ), the step 570-4 may be included in the process 550.
[0102] The decoder / virtualizer 540C can include an acoustic field decoder 544, one or more HRTFs 546, and one or more combiners 548, similar to the decoder / virtualizer 540B and the decoder / virtualizer 540C discussed above. Steps 1072, 1076, 1078, and 1080 can be correspondingly similar to steps 872, 874, and 782 discussed above.
[0103] The decoder / virtualizer 540D may further include a rotation / translation representation 1042 and an acoustic field decoding determiner 1044. The rotation / translation representation 1042 may receive signals from the internal space representation 530 and may be configured to introduce a representation of the movement of the sound source(s), the user, or both (step 1072). The representation of the movement may also consider the azimuth / pitch of the sound source 920. The rotation / translation representation 542 can output the signals to the acoustic field decoding determiner 1044.
[0104] The sound field decoder determiner 1044 may receive signals from the rotation / translation representation 1042 and may be configured to determine which signals have a "significant" output and pass those signals to the sound field decoder 544 (step 1074). A significant output may be an output that affects the perceived sound. For example, a significant output may be an audio signal having an amplitude greater than or equal to a predetermined amplitude threshold. The sound field decoder 544 may receive signals from the sound field decoder determiner 1044 having a significant output and may be configured to decode the signals (step 1076). In some embodiments, the sound field decoder 1044 may receive signals having a significant output from the sound field decoder determiner 1044. Each HRTF 546 may receive signals from the sound field decoder 544. Each HRTF 546 may be configured to determine the HRTF corresponding to its input signal and apply that HRTF to the signal (step 1078). One or more HRTFs 546 may be collectively referred to as a speaker virtualizer. Each combiner 548 may receive and combine signals from the HRTFs 546 (step 1080).
[0105] In some embodiments, those audio signals that do not have a significant output (e.g., have an amplitude less than a predetermined amplitude threshold) may not be passed to the sound field decoder 544. Thus, for audio signals that do not have a significant output, the sound field decoder 544 and the HRTF 546 may be bypassed.
[0106] An example source geometry-based speaker selection method can designate virtual speakers as active virtual speakers based on the position of the sound source (e.g., X, Y, Z position). The position of the sound source may represent the position of the sound source object. The system may determine the position of each sound source and determine which virtual speaker(s) is / are located near the corresponding sound source. In some embodiments, the determination of which virtual speaker is located near the sound source may be performed, for example, at the start of each video frame (video frame rate-based method). Compared to other methods such as a sampling rate-based method, the video frame rate-based method may require less computation.
[0107] Based on, for example, a video frame rate-based method calculation and a high-fidelity stereo decoding formula, a sound source may contribute significantly to a particular virtual speaker. As described above, virtual speakers that contribute little energy, if decoded, may cause the corresponding high-fidelity stereo decoding and HRTF processing of the decoded high-fidelity stereo channels to be bypassed. In some embodiments, the system may disable any processing blocks that are bypassed.
[0108] An example pseudocode for performing the specified method may be:
[0109] For each sound source, S, and decoding channel, n
[0110] Enabled[n] = f(sourcePosition vector3, sourceOrientation vector3, ListenerPosition vector3, ListenerOrientation vector3, VirtualSpeakerPosition[n] vector3).
[0111] High-fidelity stereo / sound field example
[0112] For each high-fidelity stereo decoding channel
[0113] if (Enabled[n]) {
[0114] AmbisonicDecode(n)
[0115] Virtualize(n)
[0116] }
[0117] Multi-channel example
[0118] For each channel
[0119] if (Enabled[n]) {
[0120] Virtualize(n)
[0121] }
[0122] Regarding the above pseudocode, the variable sourcePosition can refer to the position of the sound source, sourceOrientation can refer to the orientation of the sound source, ListenerPosition can refer to the position of the user / listener, ListenerOrientation can refer to the orientation of the user / listener, VirtualSpeakerPosition can relate to the position of the virtual speaker, AmbisonicDecode can refer to the function of performing high-fidelity stereo decoding, and Virtualize can refer to the function of performing virtualization.
[0123] Regarding the above pseudocode, for each sound source S and decoding channel n, the decoding channel n can be enabled based on one or more factors such as the position of the sound source S, the orientation of the sound source S, the position of the user / listener, the orientation of the user / listener, and the position of the virtual speaker. Still referring to the above pseudocode, for each high-fidelity stereo decoding channel, if the channel is enabled, the system can perform the high-fidelity stereo decoding function and the virtualization function.
[0124] The pseudocode can be improved by providing a "ringing" period for each virtual speaker. For example, if the source moves in position during a video frame, it can be determined that a virtual speaker may no longer have any sound source mixed into it. However, due to the filtering process of the virtual speaker, that virtual speaker may need to be an active speaker for the next frame to correctly complete the filtering process.
[0125] Examples of the present disclosure may include using all active sound sources to determine which decoded sound field outputs have "significant" outputs (e.g., outputs that will affect the perceived sound field). High-fidelity stereo or non-high-fidelity stereo multi-channel outputs that will affect the perceived sound field can be decoded. Additionally, in some embodiments, only the HRTFs546 corresponding to those detected outputs are processed. For synthetically generated high-fidelity stereo sound fields or non-high-fidelity stereo multi-channel renderings where the number of sound sources is small, or where the number of sound sources is large but they are close to each other, significant CPU savings may be achieved.
[0126] Example Method Combinations of Source-Geometry-Based Virtual Loudspeaker Selection Method and Low Energy Output Detection and Selection Method Method Combinations
[0127] In some embodiments, both virtual speaker selection based on source geometry and low-energy output detection and selection can be used sequentially to further reduce CPU consumption. As described above, virtual speaker selection based on source geometry can include, for example, selectively disabling virtual speaker processing based on, for example, the position of the sound source relative to the user / listener. Low-energy output detection and selection can include, for example, placing a signal energy / level detector between sound field decoding or multi-channel output and HRTF processing. The output / result of virtual speaker selection based on source geometry can be input into low-energy output detection and selection.
[0128] Regarding the above systems and methods, the elements of the systems and methods can be suitably implemented by one or more computer processors (e.g., a CPU or DSP). The present disclosure is not limited to any particular configuration of computer hardware for implementing these elements, including computer processors. In some cases, multiple computer systems can be employed to implement the above systems and methods. For example, a first computer processor (e.g., the processor of a wearable device coupled to a microphone) can be employed to receive input microphone signals and perform initial processing of those signals (e.g., signal conditioning and / or segmentation, such as described above). Then a second (and perhaps more computationally powerful) processor can be employed to perform more computationally intensive processing, such as determining probability values associated with speech segments of those signals. Another computer device, such as a cloud server, can host a speech recognition engine and ultimately provide input signals to it. Other suitable configurations will be apparent and are within the scope of the present disclosure.
[0129] Although the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications should be understood to be included within the scope of the disclosed examples as defined by the appended claims.
Claims
1. A method for spatially rendering an audio signal, the method comprising: Determining a model of a virtual environment; Determining a spatial configuration of the virtual environment, wherein the spatial configuration at least includes a user position, a sound source position, and a virtual speaker position; Receiving one or more input signals associated with the virtual environment; Determining whether the number of sound sources in the virtual environment exceeds a predetermined threshold; Based on the determination that the number of sound sources is less than the predetermined threshold, directly passing a signal from a sound field decoder to a plurality of detectors, the plurality of detectors receiving the signal from the sound field decoder and determining an energy level associated with the one or more input signals; Based on the determination that the number of sound sources is greater than or equal to the predetermined threshold, bypassing all of one or more processing blocks associated with at least one speaker not close to a corresponding sound source in a decoder / virtualizer; Decoding the one or more input signals; Applying the decoded one or more input signals to a speaker virtualizer; and Rendering an audio signal based on the decoded one or more input signals.
2. The method according to claim 1, further comprising: Determining whether the energy level is less than an energy threshold; Based on the determination that the energy level is not less than the energy threshold, performing head-related transfer function (HRTF) processing on the one or more input signals; And Based on the determination that the energy level is less than the energy threshold, foregoing performing the HRTF processing on the one or more input signals.
3. The method according to claim 1, further comprising: Determining whether the energy level is less than an energy threshold; Based on the determination that the energy level is not less than the energy threshold, performing head-related transfer function (HRTF) processing on the one or more input signals; And Based on the determination that the energy level is less than the energy threshold, foregoing performing the HRTF processing on the one or more input signals.
4. The method according to claim 1, further comprising: Determining the position of each sound source; And Determining which virtual speaker among a plurality of virtual speakers associated with the virtual environment is close to a corresponding sound source.
5. The method according to claim 4, wherein Performing the determination of which virtual speaker among the plurality of virtual speakers is close to a corresponding sound source for each video frame.
6. The method according to claim 4, wherein The plurality of virtual speakers include inactive virtual speakers and active virtual speakers at a first time, wherein at least one of the active virtual speakers at the first time is designated as inactive at a second time, but the signal is processed.
7. The method according to claim 1, wherein Determining the model of the virtual environment includes: Receiving one or more sound signals from one or more of a direct sound source and a reflected sound source; Modifying the one or more sound signals to simulate a Doppler effect; Adding a delay to the one or more sound signals; and Positioning the one or more sound signals on a plurality of virtual speakers, and Wherein decoding the one or more input signals includes: Determining one or more virtualized sounds associated with a movement of a direct sound source, a reflected sound source, or a user.
8. A system for spatially rendering an audio signal, the system comprising: A wearable head device configured to provide an audio signal to a user; and one or more processors configured to execute a method, the method comprising: determining a model of a virtual environment; determining a spatial configuration of the virtual environment, wherein the spatial configuration at least includes a user location, a sound source location, and a virtual speaker location; receiving one or more input signals associated with the virtual environment; determining whether a number of sound sources in the virtual environment exceeds a predetermined threshold; in accordance with a determination that the number of sound sources is less than the predetermined threshold, directly passing a signal from a sound field decoder to a plurality of detectors that receive the signal from the sound field decoder and determine an energy level associated with the one or more input signals; in accordance with a determination that the number of sound sources is greater than or equal to the predetermined threshold, bypassing all of one or more processing blocks associated with at least one speaker not close to a corresponding sound source in a decoder / virtualizer; decoding the one or more input signals; applying the decoded one or more input signals to a speaker virtualizer; and rendering an audio signal based on the decoded one or more input signals.
9. The system according to claim 8, wherein The method further comprises: determining whether the energy level is less than an energy threshold; in accordance with a determination that the energy level is not less than the energy threshold, performing head-related transfer function (HRTF) processing on the one or more input signals; and in accordance with a determination that the energy level is less than the energy threshold, foregoing performing the HRTF processing on the one or more input signals.
10. The system according to claim 8, wherein, The method further comprises: determining whether the energy level is less than an energy threshold; in accordance with a determination that the energy level is not less than the energy threshold, performing head-related transfer function (HRTF) processing on the one or more input signals; and in accordance with a determination that the energy level is less than the energy threshold, foregoing performing the HRTF processing on the one or more input signals.
11. The system according to claim 8, wherein Determining the model of the virtual environment includes: receiving one or more sound signals from a sound source; modifying the one or more sound signals to simulate a Doppler effect; adding a delay to the one or more sound signals; and positioning the one or more sound signals on a plurality of virtual speakers, and wherein decoding the one or more input signals includes: determining one or more virtualized sounds associated with movement of the sound source or movement of the user.
Citation Information
Patent Citations
Method, computer readable storage medium, and apparatus for determining a target sound scene at a target position from two or more source sound scenes
US20170245089A1