Mapping Virtual Sound Sources to Physical Speakers in Extended Reality Applications

By determining the impact of the virtual sound source in the XR environment and mapping the sound components to the loudspeaker, the problem that the XR system cannot accurately locate the sound in the loudspeaker system is solved, and more realistic audio scene presentation is achieved.

CN111459444BActive Publication Date: 2025-06-10HARMAN INT IND INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010069777.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-22
Filing Date
2020-01-21
Publication Date
2025-06-10
Estimated Expiration
2040-07-04

AI Technical Summary

Technical Problem

Existing XR systems are unable to accurately locate and orient sound in the loudspeaker system, resulting in the inability to present realistic audio scenes in the same way as headphone systems.

Method used

By determining the influence of the virtual sound source associated with the XR environment, the sound components associated with the virtual sound source are generated and mapped to speakers included in a plurality of loudspeakers, and the sound components are output for playback on the speakers.

Benefits of technology

It realizes improved realism and immersion quality in the amplifier system, and provides a more realistic audio experience through dynamic spatial presentation of virtual sound sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111459444B_ABST
    Figure CN111459444B_ABST
Patent Text Reader

Abstract

One or more embodiments include an audio processing system for generating an audio scene for an extended reality (XR) environment. The audio processing system determines that a first virtual sound source associated with the XR environment affects a sound in the audio scene. The audio processing system generates a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene. The audio processing system maps the sound component to a first loudspeaker included in a plurality of loudspeakers. The audio processing system outputs at least a first portion of the component for playback on the first loudspeaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to audio signal processing, and more particularly to mapping virtual sound sources to physical speakers in extended reality applications. Background Art

[0002] Extended reality (XR) systems, such as augmented reality (AR) systems and virtual reality (VR) systems, are increasingly popular methods for experiencing immersive computer-generated and prerecorded audio-visual environments. In an AR system, virtual computer-generated objects are projected relative to the real-world environment. In one type of AR system, a user wears a special transparent device (such as an AR headset), through which the user views physical objects in the real world as well as computer-generated virtual objects presented on the display surface of the AR headset. In other types of AR systems, when the user views the physical real-world environment, the device projects an image of the virtual object directly onto the user's eyes. In other types of AR systems, the user holds a mobile device such as a smartphone or a tablet computer. A camera associated with the mobile device captures an image of the physical real-world environment. Then, a processor associated with the mobile device presents one or more virtual objects and overlays the presented virtual objects on the display screen of the mobile device. For any of these types of AR systems, the virtual objects are displayed as objects in the physical real-world environment.

[0003] Similarly, in a VR system, virtual computer-generated objects are projected into a virtual computer-generated environment. In a typical VR system, a user wears a special device (such as a VR headset), through which the user views virtual objects in the virtual environment.

[0004] In addition, XR systems typically include a pair of headphones for delivering spatial audio directly to the user's ears. Spatial audio in XR systems involves the presentation of virtual sound sources (also referred to herein as "virtual sound artifacts") as well as environmental effects (such as echo or reverb), depending on the characteristics of the virtual space that the XR user is viewing. The complete set of virtual sound sources and associated environmental effects is referred to herein as an "audio scene" or a "sound scene". The various virtual sound sources in the environment can be fixed or mobile. A fixed virtual sound source is a sound source that appears to remain in a fixed position from the user's perception. In contrast, a mobile virtual sound source is a sound source that appears to move from one position to another from the user's perception.

[0005] Because the positions of the left and right headphone speakers are known relative to the user's ears, the XR system can accurately generate a realistic audio scene that includes all stationary and moving virtual sound sources. Generally speaking, the XR system presents the virtual sound sources so that there is the best possible correlation between the virtual sound sources heard by the user and the corresponding VR objects seen by the user on the display of the XR headset ( For example , based on the auditory angle, perceived distance, and / or perceived loudness). In this way, the VR objects and the corresponding virtual sound sources are considered to be realistic relative to how objects are seen and heard in a real-world environment.

[0006] One problem with the above method is that, compared to the same audio scene experienced via one or more loudspeakers placed in a physical environment, the audio scene experienced via headphones is generally less realistic. For example, the sound wave pressure of the sound waves generated by loudspeakers is generally much greater than that of the sound waves generated by headphones. Therefore, loudspeakers can generate a sound pressure level (SPL) that causes a physical sensation in the user, while headphones generally cannot generate such an SPL. In addition, compared to headphones, loudspeakers can generally generate audio signals with greater directivity and locality. Therefore, compared to the audio from virtual sound sources and environmental effects that only come out of headphones, the audio from the same sources and effects that come out of physical loudspeakers may sound and feel more realistic. Generally speaking, compared to the virtual sound sources and environmental effects heard via headphones, the same virtual sound sources and environmental effects heard via loudspeakers seem more realistic. In addition, the increased sound wave pressure generated by loudspeakers can provide an instinctive effect that is usually not obtainable from the sound produced by headphones. Therefore, in the case of a loudspeaker system, the user may be able to hear and also feel the audio scene generated by the loudspeaker system more realistically compared to the audio scene generated by headphones.

[0007] However, one disadvantage of a loudspeaker-based system is that the XR system generally cannot accurately position and direct sound between two or more loudspeakers in the loudspeaker system. Therefore, current XR systems cannot accurately implement the dynamic positioning of virtual sound sources in a loudspeaker system in the same way as a headphone-based XR system. Therefore, current XR systems generally cannot present a realistic audio scene via a loudspeaker system.

[0008] As mentioned above, improved techniques for generating an audio scene for an XR environment would be useful. Summary of the Invention

[0009] Various embodiments of the present disclosure describe a computer - implemented method for generating an audio scene for an extended reality (XR) environment. The method includes determining that a first virtual sound source associated with the XR environment affects a sound in the audio scene. The method further includes generating a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene. The method further includes mapping the sound component to a first loudspeaker included in a plurality of loudspeakers. The method further includes outputting at least a first portion of the component for playback on the first loudspeaker.

[0010] Other embodiments include, but are not limited to, an audio processing system that implements one or more aspects of the disclosed technology, and a computer - readable medium that includes instructions for performing one or more aspects of the disclosed technology, and a method for performing one or more aspects of the disclosed technology.

[0011] At least one technical advantage of the disclosed technology over the prior art is that, relative to existing methods, an audio scene for an XR environment is generated with improved realism and immersive quality. Via the disclosed technology, virtual sound sources are presented with increased realism through the dynamic spatialization of XR virtual audio sources relative to the user's position, orientation, and / or direction. Additionally, due to the physical characteristics of loudspeakers in terms of directivity and physical sound pressure, users may experience better audio quality and a more realistic experience compared to what they might experience with headphones. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to understand the manner in which the above - described features of one or more embodiments can be made more specific, one or more embodiments briefly summarized above may be described in more detail by reference to certain specific embodiments, some of which are illustrated in the drawings. However, it should be noted that the drawings only show typical embodiments and should not be considered in any way to limit their scope, as the scope of the present disclosure also encompasses other embodiments.

[0013] Figure 1 A system configured to implement one or more aspects of the present disclosure is shown;

[0014] Figure 2 is a more detailed illustration of an Figure 1 audio processing system according to various embodiments;

[0015] Figure 3 is a conceptual diagram showing how audio associated with a Figure 1 system according to various embodiments is mapped to a set of loudspeakers

[0016] Figures 4A - 4B shows, according to various embodiments, by Figure 1Exemplary arrangements of system-generated virtual sound sources relative to a set of loudspeakers;

[0017] Figures 5A - 5C Shows exemplary arrangements of audio panoramas generated by the Figure 1 system relative to a set of loudspeakers according to various embodiments;

[0018] Figure 6 Shows exemplary arrangements of virtual sound sources generated by the Figure 1 system relative to a set of loudspeakers and a set of head-mounted speakers according to various embodiments; and

[0019] Figures 7A - 7C Illustrates a flowchart of method steps for generating an audio scene for an XR environment according to various embodiments. Detailed Description

[0020] In the following description, numerous specific details are set forth to provide a more thorough understanding of certain specific embodiments. However, those skilled in the art will appreciate that other embodiments may be practiced without one or more of these specific details or with additional specific details.

[0021] As further described herein, an audio processing system optimizes the reproduction of an XR sound scene using available speakers (including stand-alone loudspeakers and head-mounted speaker systems) in an XR environment. The disclosed audio processing system optimizes the mapping of virtual sound sources in an XR system to physical speakers in the XR environment. In this way, the audio processing system provides a high-fidelity sound reproduction that closely represents the XR environment using the available speakers and speaker arrangements in the XR environment.

[0022] System Overview

[0023] Figure 1 Shows a system 100 configured to implement one or more aspects of the present disclosure. As shown, system 100 includes, but is not limited to, an XR system 102, an audio processing system 104, loudspeakers 120, and head-mounted speakers 130 that communicate with each other via a communication network 110. The communication network 110 can be any suitable environment that enables communication between remote or local computer systems and computing devices, including but not limited to Bluetooth communication channels, wireless and wired LANs (local area networks), and Internet-based WANs (wide area networks). Additionally or alternatively, any technically feasible combination of the XR system 102, the audio processing system 104, the loudspeakers 120, and the head-mounted speakers 130 can communicate with each other via one or more point-to-point communication links (such as exemplary communication links 124 and 134).

[0024] The XR system 102 includes, but is not limited to, computing devices, which may be stand-alone servers, clusters or "farms" of servers, one or more network devices, or any other device suitable for implementing one or more aspects of the present disclosure. Illustratively, the XR system 102 communicates over the communication network 110 via the communication link 112.

[0025] In operation, the XR system 102 generates an XR environment that replicates a virtual scene, overlays a physical reality scene with virtual content, and / or plays panoramic ( For example , 360°) immersive video and / or audio content. The audio content typically takes the form of virtual sound sources, which may include, but are not limited to, virtual sound emitters, virtual sound absorbers, and virtual sound reflectors. A virtual sound emitter is a virtual sound source at a specific location and having a specific orientation and / or direction that generates one or more sounds and / or other audio signals. The virtual sound emitter may be ambient (non-localized relative to the user), localized (at a fixed position in the XR environment), or mobile (moving within the XR environment).

[0026] In particular, for ambient virtual sound sources, an ambient virtual sound source is a virtual sound source that has no apparent location, direction, or orientation. Thus, an ambient virtual sound source appears to come from anywhere in the XR environment rather than from a specific location, direction, and / or orientation. Such ambient virtual sound sources may be presented to all speakers simultaneously. Additionally or alternatively, the ambient virtual sound source may be presented to non-directional speakers such as subwoofers. Generally, an ambient virtual sound source is an artificial construct for representing a virtual sound source perceived by the human ear as a non-localized sound source.

[0027] In a first example, the sound of rain is generated by a large number of raindrops falling, where theoretically each individual raindrop is a localized or mobile sound source contributing to the sound of rain. Generally, the human ear does not separately perceive the sound of each raindrop as a localized or mobile virtual sound source from a specific location, direction, and / or orientation. Instead, the human ear perceives the sound of rain as coming from anywhere within the XR environment. Thus, the XR system 102 may, without loss of generality, generate the sound of rain as an ambient virtual sound source.

[0028] In a second example, the sound of applause is generated by many people clapping, where theoretically each individual clap is a localized or mobile sound source contributing to the sound of applause. Generally, the human ear does not separately perceive the sound of each clap as a localized or mobile virtual sound source from a specific location, direction, and / or orientation. Instead, the human ear perceives the sound of applause as coming from anywhere within the XR environment. Thus, the XR system 102 may, without loss of generality, generate the sound of applause as an ambient virtual sound source.

[0029] In a third example, a single positioned or moving virtual sound source can generate sound in a room with many hard surfaces. Thus, when sound waves emanating from the positioned or moving virtual sound source interact with the hard surfaces, the positioned or moving virtual sound source may generate many sound reflections or echoes. As a specific example, a coin dropped in a church or lecture hall may generate so many sound reflections or echoes that the human ear cannot perceive the specific location, direction, and / or orientation of either the coin or any of the individual sound reflections or echoes. Instead, the human ear perceives the sound of the coin dropping as well as sound reflections or echoes from various locations within the XR environment. Thus, the XR system 102 can, without loss of generality, generate the sound of the coin dropping and the resulting sound reflections or echoes as ambient virtual sound sources.

[0030] Additionally or alternatively, the XR system 102 can generate a positioned or moving virtual sound source for each individual sound source contributing to the ambient virtual sound sources. The XR system 102 can generate a separate positioned or virtual sound source for each drop of rain in a rainfall, each clap of a clapping audience, and each sound reflection or echo when a coin drops in a church. The audio processing system 104 then presents a separate audio signal for each positioned or virtual sound source and maps each audio signal of each positioned or virtual sound source to one or more speakers. In this case, the XR system 102 and the audio processing system 104 do not necessarily generate and present an ambient virtual sound source for the rainfall, the clapping, or the sound reflections or echoes caused by the coin drop.

[0031] A virtual sound absorber is a virtual sound source at a specific location and having a specific orientation and / or direction that absorbs at least a portion of the sound and / or other audio signals contacting the virtual sound absorber. Similarly, a virtual sound reflector is a virtual sound source at a specific location and having a specific orientation and / or direction that reflects at least a portion of the sound and / or other audio signals contacting the virtual sound reflector.

[0032] In addition, the XR system 102 generally includes sensing hardware to track a user's head pose and spatial location for video and audio spatialization purposes. The XR system 102 transmits the user's head pose and spatial location as well as data corresponding to one or more virtual sound sources to the audio processing system 104.

[0033] The audio processing system 104 includes, but is not limited to, a computing device, which can be a stand-alone server, a cluster or "farm" of servers, one or more network devices, or any other device suitable for implementing one or more aspects of the present disclosure. Illustratively, the audio processing system 104 communicates over the communication network 110 via the communication link 114.

[0034] In operation, the audio processing system 104 maps virtual sound sources in the XR environment to physical speakers in the physical viewing environment of the XR user in a manner that optimally outputs or "presents" the audio associated with the XR environment. Given the characteristics of the XR environment, virtual sound source objects, physical speakers, and the user's physical environment, the audio processing system 104 optimizes the assignment of virtual sound sources to speakers. The audio processing system 104 then transmits the optimized audio signals to each of the loudspeakers 120 and, if present, to each of the head-mounted speakers 130.

[0035] In addition, the audio processing system 104 incorporates a speaker system including one or more loudspeakers 120 and optionally one or more head-mounted speakers 130. The loudspeakers 120 can include one or more speakers at fixed locations within a physical environment such as inside a room or a vehicle. To accurately generate realistic audio from the XR environment regardless of the specific physical environment, the audio processing system 104 can compensate for the acoustic characteristics of the physical environment. For example, the audio processing system 104 can compensate for inappropriate echo or reverberation effects caused by the physical environment. The audio processing system 104 will measure the frequency response characteristics at various locations in the physical environment. Then, when generating audio for the loudspeakers 120, the audio processing system 104 will include an audio signal that can invert or otherwise compensate for the acoustic characteristics of the physical environment.

[0036] When measuring the acoustic characteristics of the physical environment, the audio processing system 104 can consider the known properties of the loudspeakers 120 and the physical environment. These known properties can include but are not limited to speaker directivity, speaker frequency response characteristics, the spatial location of the speakers, and the physical environment frequency response characteristics. In this regard, speaker directivity can include the property that sounds emitted by low-frequency loudspeakers 120 (such as subwoofers) are generally not perceived as originating from a specific location but rather from the entire environment. In contrast, high-frequency loudspeakers 120 emit sound waves that are more strongly perceived as originating from a specific location. Speaker directivity has implications for audio analysis and audio mapping, as further described herein. Speaker frequency response characteristics include consideration of the optimal playback frequency bands for a particular loudspeaker 120 and / or individual drivers within a particular loudspeaker 120. The spatial location of the speakers can include consideration of the three-dimensional spatial location of each loudspeaker 120 relative to other speakers and relative to the user's physical environment. The spatial location of each speaker can include the position of the speaker in the physical space in the horizontal dimension and the height of the speaker in the vertical dimension. The physical environment frequency response characteristics include consideration of the frequency response characteristics or transfer functions of the physical environment at each loudspeaker 120 location and at each user's location.

[0037] The loudspeaker 120 converts one or more electrical signals into sound waves and directs the sound waves into the physical environment. Illustratively, the loudspeaker 120 may communicate over the communication network 110 via the communication link 122. Additionally or alternatively, the loudspeaker 120 may communicate with the audio processing system 104 via the point-to-point communication link 124.

[0038] The head-mounted speaker 130 converts one or more electrical signals into sound waves and directs the sound waves into one or both of the user's left and right ears. The head-mounted speaker 130 may be any technically feasible configuration, including but not limited to headphones, earbuds, and speakers integrated into an XR kit. Illustratively, the head-mounted speaker 130 may communicate over the communication network 110 via the communication link 132. Additionally or alternatively, the head-mounted speaker 130 may communicate with the audio processing system 104 via the point-to-point communication link 134.

[0039] It should be understood that the systems shown herein are illustrative and may be subject to variations and modifications. For example, the system 100 may include any technically feasible number of loudspeakers 120. Additionally, in an XR environment with one user, the user may receive audio only from the loudspeaker 120, or may receive audio from both the loudspeaker 120 and the head-mounted speaker 130. Similarly, in a multi-user XR environment with two or more users, each of the users may receive audio only from the loudspeaker 120, or may receive audio from both the loudspeaker 120 and the head-mounted speaker 130. In some embodiments, some users may receive audio only from the loudspeaker 120, while other users may receive audio from both the loudspeaker 120 and the head-mounted speaker 130.

[0040] Operation of the Audio Processing System

[0041] As further described herein, the audio processing system 104 presents an XR sound scene that includes a plurality of ambient virtual sound sources, positional virtual sound sources, and moving virtual sound sources. The audio processing system 104 presents the XR sound scene onto a fixed speaker system within the XR environment. The XR environment may be an indoor location, such as a room, the passenger compartment of a car or other vehicle, or any other technically feasible environment. Via the techniques disclosed herein, the audio processing system 104 replicates the XR sound scene with high fidelity via the available speaker arrangement and speaker frequency response characteristics of the loudspeaker 120 system. To appropriately replicate the XR sound scene, the audio processing system 104 maps a set of virtual sound sources associated with the XR system 102 to a set of physical speakers. Generally, the physical speakers are located at fixed positions within the XR environment.

[0042] The audio processing system 104 dynamically places virtual sound sources within the XR environment so that they appear to emanate from the correct position, direction, and / or orientation within the XR environment. Additionally, as the virtual sound sources move within the XR environment and as the user reference frame within the XR environment changes, the audio processing system 104 dynamically adjusts the relative position, direction, and / or orientation of the virtual sound sources. For example, when a user drives a virtual vehicle within the XR environment, performs turns, accelerations, and decelerations, the audio processing system 104 can dynamically adjust the relative position, direction, and / or orientation of the virtual sound sources. In some embodiments, the audio processing system 104 replicates the XR sound scene via a system of loudspeakers 120 in conjunction with one or more head-mounted speakers 130, where the head-mounted speakers 130 move within the physical environment as the associated user moves.

[0043] Figure 2 is of an audio processing system 104 according to various embodiments Figure 1 is a more detailed illustration of the audio processing system 104. As shown, the audio processing system 104 includes, but is not limited to, a processor 202, a storage device 204, an input / output (I / O) device interface 206, a network interface 208, an interconnect 210, and a system memory 212.

[0044] The processor 202 retrieves and executes programming instructions stored in the system memory 212. Similarly, the processor 202 stores and retrieves application data resident in the system memory 212. The interconnect 210 facilitates the transfer of, such as, programming instructions and application data between the processor 202, the input / output (I / O) device interface 206, the storage device 204, the network interface 208, and the system memory 212. The I / O device interface 206 is configured to receive input data from a user I / O device 222. Examples of the user I / O device 222 can include one or more buttons, a keyboard, and a mouse or other pointing device. The I / O device interface 206 can also include an audio output unit configured to generate an electrical audio output signal, and the user I / O device 222 can also include a speaker configured to generate an acoustic output in response to the electrical audio output signal. Another example of the user I / O device 222 is a display device, which generally represents any technically feasible means for generating an image for display. For example, the display device can be a liquid crystal display (LCD) monitor, an organic light emitting diode (OLED) monitor, or a digital light processing (DLP) monitor. The display device can be a TV including a broadcast or cable tuner for receiving digital or analog TV signals. The display device can be included in a VR / AR head-mounted kit. Additionally, the display device can project an image onto one or more surfaces such as a wall or a projection screen, or can project an image directly onto the user's eyes.

[0045] Includes a processor 202 to represent a single central processing unit (CPU), multiple CPUs, a single CPU with multiple processing cores, etc. And, typically includes a system memory 212 to represent random access memory. The storage device 204 can be a disk drive storage device. Although shown as a single unit, the storage device 204 can be a combination of fixed and / or removable storage devices, such as a fixed disk drive, a floppy disk drive, a tape drive, a removable memory card, or an optical storage device, a network-attached storage (NAS), or a storage area network (SAN). The processor 202 communicates with other computing devices and systems via a network interface 208, where the network interface 208 is configured to transmit and receive data via a communication network.

[0046] The system memory 212 includes, but is not limited to, an audio analysis and preprocessing application 232, an audio mapping application 234, and a data storage area 242. When executed by the processor 202, the audio analysis and preprocessing application 232 and the audio mapping application 234 will perform one or more operations associated with Figure 1 the audio processing system 104 as further described herein. When performing operations associated with the audio processing system 104, the audio analysis and preprocessing application 232 and the audio mapping application 234 can store data in and retrieve data from the data storage area 242.

[0047] In operation, the audio analysis and preprocessing application 232 determines the audio properties of the sound components of the virtual sound source regarding presenting audio data related to the virtual sound source via one or more loudspeakers 120 and / or head-mounted speakers 130. Some virtual sound sources may correspond to visual objects that generate sound in the XR environment. Additionally or alternatively, some virtual sound sources can correspond to specific audio generation locations in the XR environment scene that do not have a corresponding visual object. Additionally or alternatively, some virtual sound sources can correspond to environmental or background audio tracks that have no location or corresponding visual object in the XR environment.

[0048] In some embodiments, certain virtual sound sources can be associated with a developer override for the reproduction of ambient sound, positional sound, or moving sound of the virtual sound source. A developer override is a rule by which, when a virtual sound source or virtual sound source category meets certain criteria, the corresponding virtual sound source is assigned to an available speaker in a predetermined manner via a specific mapping. If a virtual sound source is subject to a developer override, the audio analysis and preprocessing application 232 will not analyze or preprocess the virtual sound source before transmitting it to the audio mapping application 234 for mapping.

[0049] The audio analysis and preprocessing application 232 can perform frequency analysis to determine suitability for spatialization. If the virtual sound source includes low frequencies, the audio analysis and preprocessing application 232 can present the virtual sound source in a non-spatialized manner. The audio mapping application 234 then maps the virtual sound source to one or more subwoofers and / or equivalently to all speakers 120. If the virtual sound source includes mid to high frequencies, the audio analysis and preprocessing application 232 presents the virtual sound source in a spatialized manner. The audio mapping application 234 then maps the virtual sound source to one or more speakers that are closest to the location, direction, and / or orientation corresponding to the virtual sound source. Generally speaking, low frequencies, mid frequencies, and high frequencies can be defined as non-overlapping and / or overlapping frequency ranges in any technically feasible manner. In one non-limiting example, low frequencies can be defined as frequencies in the range of 20 Hertz (Hz) to 200 Hz, mid frequencies can be defined as frequencies in the range of 200 Hz to 5,000 Hz, and high frequencies can be defined as frequencies in the range of 5,000 Hz to 20,000 Hz.

[0050] In some embodiments, the audio analysis and preprocessing application 232 can generate a priority list for sound source mapping. In such an embodiment, the audio analysis and preprocessing application 232 can prioritize the mapping or assignment of certain virtual sound sources or certain passbands associated with virtual sound sources before performing the mapping or assignment of lower priority virtual sound sources or passbands.

[0051] In some embodiments, the audio analysis and preprocessing application 232 can map separate multiple overlapping sounds present in a single audio stream. Additionally or alternatively, the audio analysis and preprocessing application 232 can map the overlapping components as a single sound in a single audio stream before analyzing the several overlapping components separately.

[0052] In addition, the audio analysis and preprocessing application 232 can analyze additional properties of the virtual sound source that affect how the audio analysis and preprocessing application 232 and the audio mapping application 234 present the virtual sound source. For example, the audio analysis and preprocessing application 232 can analyze the position of the virtual sound source in the XR environment. Additionally or alternatively, the audio analysis and preprocessing application 232 can analyze the distance between the virtual sound source and an acoustic reflection surface (such as a virtual sound reflector) and / or an acoustic absorption surface (such as a virtual sound absorber) within the XR environment. Additionally or alternatively, the audio analysis and preprocessing application 232 can analyze the amplitude or volume of the sound generated by the virtual sound source. Additionally or alternatively, when the virtual sound source and the user are represented in the XR environment, the audio analysis and preprocessing application 232 can analyze the shortest straight-line path from the virtual sound source to the user. Additionally or alternatively, the audio analysis and preprocessing application 232 can analyze the reverberation properties of the virtual surface located near the virtual sound source in the XR environment. Additionally or alternatively, the audio analysis and preprocessing application 232 can analyze the masking properties of the virtual objects nearby in the XR environment.

[0053] In some embodiments, the audio analysis and preprocessing application 232 can analyze the audio interaction of virtual sound sources that are adjacent to each other in the audio scene. For example, the audio analysis and preprocessing application 232 can determine that virtual sound sources located near each other can mask each other. In such an embodiment, the audio analysis and preprocessing application 232 can suppress the virtual sound source that would otherwise be masked, rather than forwarding the sound that may be masked to the audio mapping application 234. In this way, the audio mapping application 234 does not consume processing resources for mapping virtual sound sources that are subsequently masked by other virtual sound sources. To account for such audio interaction between virtual sound sources in the analysis, the audio analysis and preprocessing application 232 can additionally but not limited to analyze the distance from one virtual sound source to other virtual sound sources, the amplitude or volume of one virtual sound source relative to other virtual sound sources, and the spectral properties of the audio generated by the virtual sound source.

[0054] In operation, the audio mapping application 234 analyzes the virtual sound sources received from the audio analysis and preprocessing application 232 to determine the optimal assignment of the virtual sound sources to physical speakers (including the loudspeaker 120 and the head-mounted speaker 130). In doing so, the audio mapping application 234 performs two different processes, namely the cost function process and the optimization process, as now described.

[0055] First, the audio mapping application 234 performs a cost function process to calculate the cost of assigning virtual sound sources to physical speakers such as loudspeaker 120 and head-mounted speaker 130. During the performance of the cost function process, the audio mapping application 234 analyzes the user's ability to localize a particular sound based on the sound quality properties of the corresponding virtual sound source and the overall sound pressure level contributed by the particular virtual sound source rather than other virtual sound sources.

[0056] The audio mapping application 234 calculates a cost function based on the sound quality properties of the virtual sound source that support the user's successful spatial localization of the virtual sound source, and the properties include but are not limited to the frequency of the virtual sound source ( For example for monaural spectral cues), the amplitude or volume of the virtual sound source ( For example for interaural level differences), and the sound propagation model associated with the virtual sound source.

[0057] In some embodiments, when analyzing the ability to localize a given source, the audio mapping application 234 may additionally analyze the presence of other virtual sound sources in the virtual space, including but not limited to the overlap of the frequency distributions of multiple virtual sound sources, interfering noise, background noise, and multi-source simplification (such as the W-disjoint orthogonality method (WDO)). Additionally, the audio mapping application 234 analyzes other properties of the virtual sound source, and the other properties may depend on other virtual geospatial and acoustic variables, including but not limited to the angle of the virtual sound source with respect to the user in the XR environment, the distance of the virtual sound source from the user in the XR environment, the amplitude or volume of the virtual sound source to the user in the XR environment, and the type of the virtual sound source (i.e., ambient sound source, localization sound source, or moving sound source). In some embodiments, the cost function may also be based on the frequency response and sensitivity of one or more of the loudspeakers 120 and / or the head-mounted speakers 130.

[0058] In some embodiments, the audio mapping application 234 can map 'k' virtual sound sources to 'l' physical speakers by generating a vector's' of speaker assignments, where the length of's' corresponds to 'k'. The index's i ' corresponds to the index of the virtual sound source 'i', where 1 ≤ i ≤ k. The value of's i ' corresponds to the assignment of the virtual sound source to a speaker of the speaker system, where 1 ≤ s i i ≤ l.

[0059] In some embodiments, the audio mapping application 234 can calculate a cost function The cost function quantifies the cost of reproducing the virtual sound source 'i' on the speaker 'j' according to the following equation 1:

[0060]

[0061] Where A(i, j) is the absolute distance relative to the user between the angle of virtual sound source 'i' and physical speaker 'j', and F(i) is a frequency deviation function that preferentially processes the spatialization of sound sources with higher frequencies.

[0062] In addition, the audio mapping application 234 can calculate A(i, j) according to Equation 2 below:

[0063] A(i, j) = α|γ i -δ j | (2)

[0064] Where γ is a vector including the angular offsets in the virtual space for all virtual sound sources, and δ is a vector including the angular offsets of all physical speakers.

[0065] As disclosed above, F(i) is a frequency deviation function that preferentially processes the spatialization of sound sources with higher frequencies because, relative to lower-frequency sound sources, higher-frequency sound sources are generally perceived as more directional. Therefore, calculating F(i) ensures that sound sources dominated by high frequencies are weighted more highly than those dominated by low frequencies. The audio mapping application 234 can calculate F(i) according to Equation 3 below:

[0066]

[0067] Where ω i is the dominant audio frequency of virtual sound source i. For example, the audio mapping application 234 can determine the value of ω i based on the maximum energy analysis of the Fourier spectrum of sound source i.

[0068] The above cost function process is an exemplary technique for determining the relative cost of assigning virtual sound sources to one or more physical speakers. Any other technically feasible method for determining the relative cost of assigning virtual sound sources to one or more physical speakers may be considered within the scope of the present disclosure.

[0069] Second, after performing the cost function process, the audio mapping application 234 performs an optimization process. The audio mapping application 234 performs the optimization process by the following method: using the cost function to determine the best mapping of virtual sound sources to physical speakers (including loudspeaker 120 and head-mounted speaker 130). By performing the optimization process, the audio mapping application 234 determines the assignment of virtual sound sources to physical speakers such that the cost function is minimized under this assignment.

[0070] The audio mapping application 234 can perform the optimization process via any technically feasible technique based on the nature of the cost function, including but not limited to least squares optimization, convex optimization, and simulated annealing optimization. Additionally, since the main goal of the audio processing system 104 is to assign a set of virtual sound sources to a set of fixed speakers, a combinatorial optimization method may be particularly applicable. Given the formulation of the cost function as described in connection with equations 1 - 3, the Hungarian algorithm is an applicable technique for determining the optimal assignment of virtual sound sources to physical speakers.

[0071] The audio mapping application 234 performs the Hungarian algorithm by generating a cost matrix using the cost function defined by equations 1 - 3 above, the cost matrix assigning the cost of playing back each virtual sound source to each of the physical speakers in the XR environment. Via the Hungarian algorithm, the audio mapping application 234 calculates the optimal assignment of the 'k' virtual sound sources across the 'l' speakers and the cost of each assignment. A possible cost matrix can be constructed, including the 'k' virtual sound sources and the 'l' physical speakers, as shown in Table 1 below:

[0072]

[0073] The optimization process described above is an exemplary technique for optimally assigning virtual sound sources to one or more physical speakers. Any other technically feasible method for optimally assigning virtual sound sources to one or more physical speakers may be considered within the scope of the present disclosure.

[0074] Alternative embodiments are now described where a loudspeaker 120 is employed in combination with a common head - mounted speaker 130, and where a loudspeaker 120 is employed in combination with a head - mounted speaker 130 having an audio - transparent function.

[0075] In some embodiments, the audio processing system 104 may utilize the loudspeaker 120 together with a common head-mounted speaker 130 to generate an audio scene. Such an audio scene may be optimal and more realistic compared to an audio scene generated only for the loudspeaker 120. Generally, the XR system 102 may track the position of the head-mounted speaker 130 when the user moves within a physical environment with the head-mounted speaker 130. The XR system 102 may transmit the position of the head-mounted speaker 130 to the audio processing system 104. In these embodiments, the audio processing system 104 may prioritize processing the mapping to the head-mounted speaker 130 as the main speaker. The audio processing system 104 may utilize the loudspeaker 120 to create a more immersive and realistic audio scene by mapping relatively ambient, atmospheric, and distant sounds to the loudspeaker 120. In a multi-user XR environment, the audio processing system 104 may deliver audio content specific to a single user through the head-mounted speaker 130 associated with that single user. For example, in a game where two users are competing against each other and receiving different instructions, the audio processing system 104 may deliver user-specific commentary only to the head-mounted speaker 130 associated with the target user. The audio processing system 104 may deliver environmental and audio sounds to the loudspeaker 130.

[0076] Additionally, the audio processing system 104 may further optimize the mapping based on the type of the head-mounted speaker 130, such as an open-headphone, a closed-headphone, etc. For example, if the user wears a pair of open headphones, the audio processing system 104 may map more audio to the loudspeaker 130 relative to closed headphones because the user can hear more audio generated by the loudspeaker 130.

[0077] In some embodiments, the audio processing system 104 may utilize the loudspeaker 120 together with a head-mounted speaker 130 having an audio transparency function to generate an audio scene. Such an audio scene may be optimal and more realistic compared to an audio scene generated solely for the loudspeaker 120 or for the loudspeaker 120 together with a common head-mounted speaker 130. The head-mounted speaker 130 equipped with the audio transparency function includes an external microphone mounted on the head-mounted speaker 130. The external microphone samples audio near the head-mounted speaker 130 and converts the sampled audio into an audio signal. The head-mounted speaker 130 includes a mixer that mixes the audio signal from the microphone with the audio signal received from the audio processing system 104. In this way, the audio processing system 104 may modify the audio signal transmitted to the head-mounted speaker 130 to account for the audio signal from the microphone. Thus, the audio processing system 104 may generate a more immersive and realistic audio scene compared to a system that employs a head-mounted speaker 130 without the audio transparency function.

[0078] In addition, the audio processing system 104 can control the amount of audio passed from the microphone to the head-mounted speaker 130. In doing so, the audio processing system 104 can control the level of the audio signal from the microphone relative to the audio signal received from the audio processing system 104. In this way, the audio processing system 104 can adjust the relative levels of the two audio signals to generate an audio scene with improved depth and realism. In one example, when another human user or a computer-generated player fires a virtual gun and the virtual bullet fired from the virtual gun is advancing towards the user, the user may be playing a first-person shooter game. The user then enables the in-game time warp function to dodge the virtual bullet. When the gun is first fired, the audio processing system 104 maps the sound of the virtual bullet so that the sound is emitted from the combination of two of the speakers 120 in front of the physical room where the user is playing the game. In addition, the audio processing system 104 adjusts the audio mix on the head-mounted speaker 130 so that audio transparency is fully turned on. Thus, the user hears the sound from the speaker 130. When the virtual bullet approaches the user, the audio processing system 104 adjusts the audio mix on the head-mounted speaker 130 to reduce the sound transmitted by the audio transparency function and increase the sound received from the audio processing system 104. In this way, the user hears the sound of the virtual bullet from the audio processing system 104, thus providing a more realistic and immersive sound when the virtual bullet travels near the user.

[0079] Figure 3 is a conceptual diagram 300 showing how audio associated with Figure 1 the system 100 according to various embodiments is mapped to a set of speakers. As shown, the conceptual diagram 300 includes a user 310, an XR system 102, virtual sound sources 320(0)-320(4), an audio analysis and preprocessing application 232, an audio mapping application 234, and speakers 330(0)-330(2). The XR system 102, the audio analysis and preprocessing application 232, the audio mapping application 234, and the speakers 330(0)-330(2) operate in substantially the same manner as described in connection with Figures 1 - 2 the above, with differences further described below.

[0080] As shown, user 310 wears an XR headset coupled to XR system 102. XR system 102 may be embedded within the XR headset or may be a separate system communicatively coupled to the XR headset via a wired or wireless communication link. XR system 102 generates one or more virtual sound sources 320(0)-320(4) based on certain virtual objects and audio effects generated by XR system 102. Additionally, virtual sound sources 320(0)-320(4) are based on actions performed by user 310, including but not limited to moving within the XR environment, moving within the physical real-world environment, and manipulating various controls on a controller device (not explicitly shown). Virtual sound sources 320(0)-320(4) include any technically feasible combination of virtual sound emitters, virtual sound absorbers, and virtual sound reflectors. For example, virtual sound sources 320(2)-320(4) may be virtual sound emitters that generate one or more sounds. Each of virtual sound sources 320(2)-320(4) may be an ambient, localized, or moving virtual sound emitter in any technically feasible combination. Each of virtual sound sources 320(0)-320(1) may be: a virtual sound absorber that absorbs incoming sound, a virtual sound reflector that reflects incoming sound, or a virtual sound source that absorbs a first portion of incoming sound and reflects a second portion of incoming sound.

[0081] Information about virtual sound sources 320(0)-320(4) is transmitted to audio analysis and preprocessing application 232. Audio analysis and preprocessing application 232 performs frequency analysis to determine suitability for spatialization. If virtual sound source 320 includes low frequencies, audio analysis and preprocessing application 232 presents virtual sound source 320 in a non-spatialized manner. If virtual sound source 320 includes mid to high frequencies, audio analysis and preprocessing application 232 presents virtual sound source 320 in a spatialized manner. After performing frequency analysis, the preprocessed virtual sound source 320 reaches audio mapping application 234.

[0082] Audio mapping application maps each of the preprocessed virtual sound sources 320 to one or more speakers 330(0)-330(2) based on the frequency analysis information generated by audio analysis and preprocessing application 232. If the virtual sound source includes low frequencies, audio mapping application 234 maps the virtual sound source to one or more subwoofers and / or equivalently to multiple speakers 330(0)-330(2). If the virtual sound source includes mid to high frequencies, audio mapping application 234 then maps the virtual sound source to one or more speakers that are closest to the position, direction, and / or orientation corresponding to the virtual sound source. Audio mapping application 234 transmits the sound associated with each of virtual sound sources 320(0)-320(4) to the appropriate speakers 330(0)-330(2).

[0083] Various scenarios for arranging and mapping virtual sound sources to speakers are now described.

[0084] Figures 4A - 4B An exemplary arrangement of virtual sound sources 420(0)-420(4) generated by Figure 1 system 100 according to various embodiments relative to a set of loudspeakers 430(0)-430(8) is shown. The virtual sound sources 420(0)-420(4) and the loudspeakers 430(0)-430(8) function in substantially the same manner as those described in connection with Figures 1 - 3 which, with differences further described below.

[0085] As Figure 4A shown in exemplary arrangement 400 of Figure 1 , user 410 is surrounded by five virtual sound sources 420(0)-420(4) generated by XR system 102 of

[0086] Virtual sound source 420(0) is a virtual sound emitter that generally points towards user 410. Thus, user 410 hears all or most of the sound generated by virtual sound source 420(0). Virtual sound sources 420(1) and 420(2) are virtual sound emitters that are angled towards user 410. Thus, user 410 hears a small portion of the sound generated by virtual sound sources 420(1) and 420(2). Virtual sound source 420(3) is a virtual sound emitter that generally points away from user 410. Thus, user 410 directly hears little or no sound generated by virtual sound source 420(3). Additionally, virtual sound source 420(3) generally points towards virtual sound source 420(4).

[0087] In some embodiments, all of the virtual sound sources 420(0)-420(4) may be perceived as being located relative to the direction of the user 410's head. In such an embodiment, when the user 410 rotates his or her head and / or body rather than remaining in a static position, orientation, and / or alignment, the virtual sound sources may appear to rotate with the user 410. Thus, the virtual sound sources 420(0)-420(3) may be positioned virtual sound emitters that are perceived by the user 410 as being more or less static sound sources that remain in a fixed position, orientation, and / or alignment relative to the user 410. Similarly, the virtual sound source 420(4) may be a virtual sound absorber and / or virtual sound reflector that remains in a fixed position, orientation, and / or alignment relative to the user 410.

[0088] As Figure 4B shown in the exemplary arrangement 450 of Figure 4A FIG. 8, the user 410 is surrounded by the same five virtual sound sources 420(0)-420(4) shown in Figure 4B FIG. 9. Even when the user 410 looks in a different direction relative to Figure 4A FIG. 10, the virtual sound sources 420(0)-420(4) maintain the same relative position, orientation, and / or alignment relative to the user 410. The audio processing system 104 maps each of the virtual sound sources 420(0)-420(4) to one or more of the loudspeakers 430(0)-430(8). The loudspeakers 430(0)-430(7) are directional speakers. The audio processing system 104 may map mid-frequency to high-frequency sounds from the positioned and / or moving virtual sound sources 420 to the loudspeakers 430(0)-430(7). The loudspeaker 430(8) is an omnidirectional speaker, such as a subwoofer. The audio processing system 104 may map the ambient virtual sound sources 420 as well as low-frequency sounds from the positioned and / or moving virtual sound sources 420 to the loudspeaker 430(8). Additionally or alternatively, the audio processing system 104 may map the ambient virtual sound sources 420 as well as low-frequency sounds from the positioned and / or moving virtual sound sources 420 more or less equivalently to the loudspeakers 430(0)-430(7).

[0089] The audio processing system 104 maps sound to speakers 430(0)-430(7) based on the relative position, direction, and / or orientation of corresponding virtual sound sources 420(0)-420(4). In this regard, the audio processing system 104 mainly maps the sound generated by the virtual sound source 420(0) to the speaker 430(3). Additionally, the audio processing system 104 may map a portion of the sound generated by the virtual sound source 420(0) to one or more additional speakers, such as the speaker 430(5). The audio processing system 104 mainly maps the sound generated by the virtual sound source 420(1) to the speaker 430(7). Additionally, the audio processing system 104 may map a portion of the sound generated by the virtual sound source 420(1) to one or more additional speakers, such as the speakers 430(4) and 430(6). The audio processing system 104 mainly maps the sound generated by the virtual sound source 420(2) to the speaker 430(4). Additionally, the audio processing system 104 may map a portion of the sound generated by the virtual sound source 420(2) to one or more additional speakers, such as the speaker 430(2). The audio processing system 104 mainly maps the sound generated by the virtual sound source 420(3) and the absorption and reflection effects of the virtual sound source 420(4) to the speaker 430(1). Additionally, the audio processing system 104 may map a portion of the sound generated by the virtual sound source 420(3) and the absorption and / or reflection effects of the virtual sound source 420(4) to one or more additional speakers, such as the speakers 430(0) and 430(2).

[0090] In one possible scenario, the user 410 is a player of a VR action game. In contrast to playing game audio using headphones, the user 410 plays the game audio via a surround sound system exemplified by the speakers 430(0)-430(8). The audio processing system 104 presents the game audio to the speakers 430(0)-430(8) in a manner that enhances the level of immersion experienced by the user 410 during the game. The audio processing system 104 automatically detects when the orientation of the user 410 changes in the XR environment (for example, when the user 410 is driving a car and turns left), and reassigns the virtual sound sources 420(0)-420(4) to different speakers 430(0)-430(8) in the XR environment. To maintain the audio presentation in the XR environment as realistic as possible. The audio processing system 104 tracks moving virtual sound sources ( For example , attack helicopters, projectiles, vehicles, etc.), and automatically assigns the speakers 430(0)-430(8) to play these virtual sound sources consistent with the virtual sound sources 420(0)-420(4) generated by the XR system 102.

[0091] Figures 5A - 5C Shows, according to various embodiments, byFigure 1 An exemplary arrangement of the audio panorama 520 generated by the system 100 relative to a set of loudspeakers 530(0)-530(8). The loudspeakers 530(0)-530(8) operate in substantially the same manner as those described in conjunction with Figures 1 - 4B and are described further below.

[0092] As shown in the exemplary arrangement 500 of Figure 5A , the user 510 is surrounded by the virtual audio panorama 520. Generally speaking, the audio panorama 520 can be used as a virtual sound source that forms a 360° sound circle around the user 510. In addition, the audio panorama 530 can have a vertical dimension such that the audio panorama forms a virtual dome (or sphere) of sound around the user 510 in both the horizontal and vertical dimensions. The audio panorama 520 has the effect of generating sound continuously around and above the user 510. As shown, the audio panorama 520 includes a focus indicator 540 that identifies the direction the user is facing. Depending on one or more characteristics of the audio panorama 520, the audio panorama 520 can rotate or remain fixed as the user rotates his or her head to face different directions.

[0093] In one possible scenario, the user 510 is a mountain bike. The user 510 records his or her mountain bike ride via a helmet equipped with a 360° camera mounted on his or her helmet. The camera includes a directional microphone system that records the spatial audio panorama 520 based on the direction the user 510 is looking at any given time. Later, the user 510 views the previously recorded mountain bike ride on the XR system 102 and the audio processing system 104.

[0094] When viewing panoramic video content via the XR system 102 associated with the XR headset of the user 510, the audio processing system 104 tracks the position, direction, and / or orientation of the user 510 in the physical environment relative to the focus indicator 540 of the original audio panorama 520. Based on the position, direction, and / or orientation of the user 510 in the physical environment relative to the focus indicator 540, the audio processing system 104 automatically adjusts the position, direction, and / or orientation of the audio panorama 510 as needed according to the movement of the user 510 and the movement recorded in the original mountain bike ride, thereby bringing a more realistic and vivid XR experience when viewing the previously recorded mountain bike ride.

[0095] In a first example, the sound represented by the audio panorama 520 can be an ambient virtual sound source. One such ambient virtual sound source can be rainfall. As further described herein, an ambient virtual sound source is a virtual sound source that does not have an apparent position, direction, or orientation. Thus, the ambient virtual sound source appears to come from anywhere within the audio panorama 520 rather than from a specific position, direction, and / or orientation. In such a case, when the user 510 rotates his or her head, the audio panorama 520 will not rotate.

[0096] In a second example, the user 510 can replay a virtual mountain bike ride through a forest that the user 510 recorded during an actual mountain bike ride. The user 510 can ride his or her bike in a straight line when recording the original mountain bike ride. Subsequently, when replaying the previously recorded virtual mountain bike ride, the user can turn his or her head to the left while the virtual bike continues to travel straight in the same direction. In such a case, if the ambient sound represented by the audio panorama 520 is replayed via a physical loudspeaker rather than via a head-mounted speaker, the audio panorama 520 will not rotate. Thus, the audio processing system 104 will not adjust the presented sound that is transmitted to the physical loudspeaker. In a specific example, the rustling of leaves on a single tree and the pecking of a woodpecker appear to come from a virtual tree directly to the left of the user 510. If, during the replay of the mountain bike ride, the user 510 turns his or her head to the left to face the virtual tree, the audio panorama 520 will not rotate, and the sound associated with the virtual tree will continue to be presented to the same physical loudspeaker. Because the user 510 is now facing left, the sound associated with the virtual tree appears to be in front of him or her.

[0097] In a third example, user 510 can replay a virtual mountain bike ride through a forest that user 510 recorded during an actual mountain bike ride. User 510 may keep his or her head stationary during the replay, but both user 510 and the bike can change direction within the audio panorama 520 based on the previously recorded mountain bike ride. In this case, XR system 102 will track the direction of the bike, so the ambient sounds represented by the audio panorama 520 will rotate in the opposite direction by approximately the same amount. In a specific example, when user 510 is riding the bike and initially records the scene as audio panorama 520, user 510 can turn his or her bike to the left. During the left turn, user 510's head will remain aligned with the bike. Subsequently, when user 510 replays the previously recorded virtual mountain bike ride, user 510 can experience the virtual mountain bike ride while standing or sitting still without turning his or her head. During the above-mentioned left turn, the extended reality environment as represented by the video appears to rotate to the right. In other words, the extended reality environment as represented by the video will rotate in the opposite direction to the right by approximately the same amount as the original left turn. Thus, if user 510 turns 90° to the left, the extended reality environment as represented by the video will rotate 90° in the opposite direction to the right. Similarly, the audio panorama 520 will also rotate 90° in the opposite direction to the right to maintain the correct orientation of the virtual sound sources represented by the audio panorama 520. Thus, the rustling of the leaves of a single tree directly to the left of user 510 and the pecking of a woodpecker before the left turn seem to be in front of user 510 after the left turn.

[0098] As shown in Figure 5B the exemplary arrangement 550 of Figure 5A user 510 is surrounded by the same audio panorama 520 as shown in Figure 5B . Even when user 510 looks in Figure 5A a different direction relative to Figure 5A , the audio panorama 520 does not rotate based on the movement of user 510. In this regard, the focus indicator 540 of the audio panorama 520 remains in the same direction as shown in Figure 5B . The audio processing system 104 continues to map the audio panorama 520 to the speakers 530(0)-530(8) based on frequency and based on the position, direction, and / or orientation of the focus indicator 540 of the audio panorama 520 relative to the speakers 530(0)-530(8). Figure 5B The scene shown in Figure 5B corresponds to the first example above, where the virtual sound sources represented by the audio panorama 520 are ambient sound sources. Figure 5B The scene shown in also corresponds to the second example above, where user 510 rides his or her bike in a straight line while recording the original mountain bike ride, but turns his or her head when subsequently replaying the previously recorded virtual mountain bike ride.

[0099] As shown in the exemplary arrangement 560 of Figure 5C , the user 510 is Figure 5A surrounded by the same audio panorama 520 as shown in. As shown, the user 510 looks towards Figure 5C in a different direction relative to Figure 5A . The audio panorama 520 rotates in the opposite direction by an amount approximately equal to the movement of the user 510. In this regard, the focus indicator 540 of the audio panorama 520 moves in the direction opposite to the movement of the user 510. The audio processing system 104 maps the audio panorama 520 to the loudspeakers 530(0)-530(8) based on frequency and based on the new position, orientation, and / or alignment of the focus indicator 540 of the audio panorama 520 relative to the loudspeakers 530(0)-530(8). Figure 5C The scene shown in corresponds to the third example above, where the user 510 keeps his or her head stationary during playback, but both the user 510 and the bicycle change direction within the audio panorama 520 based on a previously recorded mountain bike ride.

[0100] Figure 6 shows an exemplary arrangement 600 of virtual sound sources 620 generated by the Figure 1 system 100 according to various embodiments relative to a set of loudspeakers 630(0)-630(8) and a set of head-mounted speakers 615. The virtual sound sources 620 and the loudspeakers 630(0)-630(8) operate in substantially the same manner as described in connection with Figures 1 - 5C , with differences further described below.

[0101] As shown, the user 610 faces the virtual sound source 620. The user 610 is surrounded by the loudspeakers 630(0)-630(8) and wears a set of head-mounted speakers 615. Since the user 610 is near the virtual sound source 620, the audio processing system 104 can map most or all of the mid-frequency to high-frequency sounds emitted by the virtual sound source 620 to the headphones 615. The audio processing system 104 can map most or all of the low-frequency sounds emitted by the virtual sound source 620 to the loudspeaker 630(8). Additionally or alternatively, the audio processing system 104 can map most or all of the low-frequency sounds emitted by the virtual sound source 620 more or less equivalently to the loudspeakers 630(0)-630(7). Further, the audio processing system 104 can map additional ambient virtual sound sources, localized virtual sound sources, and moving virtual sound sources (not explicitly shown) to one or more of the loudspeakers 630(0)-630(8), as further described herein.

[0102] In one possible scenario, user 610 uses XR system 102 to complete a virtual training exercise on how to diagnose and repair an engine. XR system 102 generates an avatar of an instructor in the XR environment. XR system 102 also generates various virtual sound sources such as virtual sound source 620, which generates sounds related to the virtual engine that is running. Audio processing system 104 generates the sounds that user 610 hears via one or more of loudspeakers 630(0)-630(8) and head-mounted speaker 615. When user 610 moves around the virtual engine in the XR environment, audio processing system 104 tracks user 610. When user 610 moves his or her head closer to the virtual engine, audio processing system 104 routes the instructor's voice to head-mounted speaker 615. In this way, audio processing system 104 can compensate for the fact that the instructor's voice may be masked by the noise generated by the virtual sound source related to the running virtual engine. Additionally or alternatively, audio processing system 104 can map the low-frequency rumble of the virtual engine to loudspeaker 630(8) or more or less equivalently to loudspeakers 630(0)-630(7). In this way, user 610 experiences these low-frequency rumbles as ambient non-spatial sounds. Further, the user may experience a physical sensation from loudspeakers 630(0)-630(8) that approximates the physical vibration of the running engine.

[0103] If the instructor wants to draw user 610's attention to an engine component that generates a high-pitched whooshing sound, audio processing system 104 can map the high-pitched whooshing sound to one of the directional loudspeakers 630(0)-630(7) in the room. In this way, user 610 can experience the high-pitched sound as a directional sound and can more easily identify the virtual engine component that is generating the sound.

[0104] Figures 7A - 7C A flowchart of method steps for generating an audio scene for an XR environment in accordance with various embodiments is presented. Although the method steps are described in connection with Figures 1 - 6 the system, those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.

[0105] As shown, method 700 begins at step 702, where audio processing system 104 receives the acoustic characteristics of loudspeakers 120 and the physical environment. These acoustic characteristics can include, but are not limited to, speaker directivity, speaker frequency response characteristics, the three-dimensional spatial location of the speakers, and the physical environment frequency response characteristics.

[0106] At step 704, the audio processing system 104 receives parameters regarding the virtual sound source. These parameters include the position, amplitude or volume, direction, and / or orientation of the virtual sound source. These parameters also include whether the virtual sound source is an ambient virtual sound source, a localized virtual sound source, or a moving virtual sound source. These parameters also include whether the virtual sound source is a virtual sound emitter, a virtual sound absorber, or a virtual sound reflector. These parameters also include any other information describing how the virtual sound source generates or affects sound in the XR environment.

[0107] At step 706, the audio processing system 104 determines whether the virtual sound source generates sound in the XR environment. If the virtual sound source generates sound in the XR environment, the method 700 proceeds to step 712, where the audio processing system 104 generates one or more preprocessed virtual sound sources based on the incoming virtual sound source. The preprocessed virtual sound sources include information regarding the spectrum of the virtual sound source. For example, the preprocessed virtual sound sources include information regarding whether the virtual sound source includes any one or more of low-frequency, mid-frequency, and high-frequency sound components.

[0108] At step 714, the audio processing system 104 determines whether the preprocessed virtual sound source is an ambient sound source. If the preprocessed virtual sound source is an ambient sound source, the method 700 proceeds to step 716, where the audio processing system 104 generates ambient audio data based on the preprocessed virtual sound source and the stored metadata. The stored metadata includes information related to the loudspeakers 120 and the acoustic characteristics of the physical environment. The stored metadata also includes information related to virtual sound sources (such as virtual sound absorbers and virtual sound reflectors) that affect audio in the XR environment. At step 718, the audio processing system 104 outputs or presents the ambient sound components of the audio data via the ambient speaker system. When performing this step, the audio processing system 104 may map the ambient audio data to one or more subwoofers. Additionally or alternatively, the audio processing system 104 may map the ambient audio data to all of the directional loudspeakers 120, equivalently mapping the virtual sound source to all of the directional loudspeakers 120.

[0109] At step 720, the audio processing system 104 determines whether there are additional virtual sound sources to process. If there are additional virtual sound sources to process, the method 700 proceeds to step 704 as described above. On the other hand, if there are no additional virtual sound sources to process, the method 700 terminates.

[0110] Return to step 714. If the preprocessed virtual sound source is not an ambient sound source, the preprocessed virtual sound source is a localization virtual sound source or a moving virtual sound source. In this case, method 700 proceeds to step 722, where audio processing system 104 generates a speaker mapping for the virtual sound source. Audio processing system 104 generates the mapping based on the frequency components of the virtual sound source. Low-frequency sound components can be mapped to the ambient speaker system. Mid-frequency sound components and high-frequency sound components can be mapped to one or more directional speakers in the spatial speaker system. At step 724, audio processing system 104 generates ambient sound components based on the low-frequency sound components of the virtual sound source and based on the stored metadata. The stored metadata includes information related to the acoustic characteristics of the loudspeakers 120 and the physical environment. The stored metadata also includes information related to virtual sound sources (such as virtual sound absorbers and virtual sound reflectors) that affect audio in the XR environment. At step 726, audio processing system 104 outputs or presents the low-frequency sound components of the virtual sound source via the ambient speaker system. When performing this step, audio processing system 104 can map the ambient sound components of the audio data to one or more subwoofers. Additionally or alternatively, audio processing system 104 can equivalently map the ambient sound components of the audio data related to the virtual sound source to all directional loudspeakers 120.

[0111] At step 728, audio processing system 104 generates speaker-specific sound components based on the mid-frequency sound components and high-frequency sound components of the virtual sound source and based on the stored metadata. The stored metadata includes information related to the acoustic characteristics of the loudspeakers 120 and the physical environment. The stored metadata also includes information related to virtual sound sources (such as virtual sound absorbers and virtual sound reflectors) that affect audio in the XR environment. At step 730, audio processing system 104 outputs or presents the mid-frequency sound components and high-frequency sound components of the audio data related to the virtual sound source via one or more speakers in the spatial speaker system. Then, the method proceeds to step 720 as described above.

[0112] Return to step 706. If the virtual sound source does not generate sound in the XR environment, method 700 proceeds to step 712, where audio processing system 104 determines whether the virtual sound source affects the sound in the XR environment. If the virtual sound source does not affect the sound in the XR environment, method 700 proceeds to step 720 as described above. On the other hand, if the virtual sound source does affect the sound in the XR environment, the virtual sound source is a virtual sound absorber and / or a virtual sound reflector. Method 700 proceeds to step 710, where audio processing system 104 calculates and stores metadata related to the virtual sound source. The metadata includes, but is not limited to, the position of the virtual sound source, the orientation of the virtual sound source, and data regarding how the virtual sound source absorbs and / or reflects audio of various frequencies. Then, the method proceeds to step 720 as described above.

[0113] In summary, the audio processing system presents an XR audio scene to the loudspeaker system. In some embodiments, the audio processing system presents an XR audio scene to the loudspeaker system together with one or more sets of headphones. The audio processing system includes an audio analysis and preprocessing application that receives ambient parameters of the physical environment. The audio analysis and preprocessing application also receives data related to one or more virtual objects generated by the XR system. For each virtual object, the audio analysis and preprocessing application may also determine whether the virtual object affects one or more sounds generated by other virtual objects within the audio scene. If the virtual object (such as by absorbing or reflecting certain sounds) affects the sounds related to other virtual objects, the analysis and preprocessing application generates and stores metadata that defines how the virtual object affects the other sounds.

[0114] In addition, the audio analysis and preprocessing application determines whether the virtual object generates sound. If the virtual object generates sound, the analysis and preprocessing application generates a virtual sound source corresponding to the virtual object. Then, the analysis and preprocessing application determines whether the virtual sound source is an ambient sound source, a localized sound source, or a moving sound source. If the virtual sound source is an ambient sound source, the audio mapping application included in the audio processing system generates ambient audio data. The ambient audio data is based on the virtual object and the stored metadata related to other virtual objects and the physical environment. The audio mapping application presents the ambient sound component of the audio data via the ambient loudspeaker system. If the virtual sound source is a localized sound source or a moving sound source, the audio mapping application determines the current position of the virtual sound source and generates speaker-specific audio data. The speaker-specific audio data is based on the virtual object and the stored metadata related to other virtual objects and the physical environment. The audio mapping application presents the speaker-specific sound component of the audio data via the spatial loudspeaker system.

[0115] At least one technical advantage of the disclosed technology over the prior art is that, relative to existing methods, an audio scene for an XR environment is generated with improved realism and immersive quality. Via the disclosed technology, virtual sound sources are presented with increased realism through the dynamic spatialization of XR virtual audio sources relative to the user's position, orientation, and / or direction. Additionally, due to the physical characteristics of speakers in terms of directivity and physical sound pressure, the user may experience better audio quality and a more realistic experience compared to what may be experienced using headphones.

[0116] 1. In some embodiments, a computer-implemented method for generating an audio scene for an extended reality (XR) environment includes: determining that a first virtual sound source associated with the XR environment affects a sound in the audio scene; generating a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene; mapping the sound component to a first loudspeaker included in a plurality of loudspeakers; and outputting at least a first portion of the sound component for playback on the first loudspeaker.

[0117] 2. The computer-implemented method of clause 1, wherein the first virtual sound source includes a localized virtual sound source, and further includes: determining a virtual position associated with the first virtual sound source; and determining that the first loudspeaker is closer to the virtual position than a second loudspeaker included in the plurality of loudspeakers.

[0118] 3. The computer-implemented method of clause 1 or clause 2, wherein the first virtual sound source includes a localized virtual sound source, and further includes: determining that the first loudspeaker is included in a spatial loudspeaker system, the spatial loudspeaker system including a subset of loudspeakers within the plurality of loudspeakers; determining a virtual position associated with the first virtual sound source; determining that each of a first loudspeaker and a second loudspeaker included in the subset of loudspeakers is closer to the virtual position than a third loudspeaker included in the subset of loudspeakers; mapping the sound component to the second loudspeaker; and outputting at least a second portion of the sound component for playback on the second loudspeaker.

[0119] 4. The computer-implemented method of any one of clauses 1-3, wherein the first virtual sound source includes a moving virtual sound source, and further includes: determining that the first virtual sound source has moved from a first virtual position to a second virtual position; and determining that the first loudspeaker is closer to the second virtual position than a second loudspeaker included in the plurality of loudspeakers.

[0120] 5. The computer-implemented method according to any one of clauses 1-4, the computer-implemented method further comprising: determining that the first virtual sound source has moved from the second virtual position to a third virtual position; determining that a second loudspeaker is closer to the third virtual position than a first loudspeaker; removing at least a first portion of the sound component that is not output to the first loudspeaker; mapping the sound component to the second loudspeaker; and outputting at least a second portion of the sound component for playback on the second loudspeaker.

[0121] 6. The computer-implemented method according to any one of clauses 1-5, the computer-implemented method further comprising: determining that a second virtual sound source associated with the XR environment affects the sound in the audio scene; determining that the second virtual sound source includes a virtual sound absorber that absorbs at least a portion of the sound component associated with the first virtual sound source; determining an absorption value based on at least a portion of the sound component associated with the first virtual sound source; and reducing at least a portion of the sound component associated with the first virtual sound source based on the absorption value.

[0122] 7. The computer-implemented method according to any one of clauses 1-6, the computer-implemented method further comprising: determining that a second virtual sound source associated with the XR environment affects the sound in the audio scene; determining that the second virtual sound source includes a virtual sound reflector that reflects at least a portion of the sound component associated with the first virtual sound source; determining a reflection value based on at least a portion of the sound component associated with the first virtual sound source; and increasing at least a portion of the sound component associated with the first virtual sound source based on the reflection value.

[0123] 8. The computer-implemented method according to any one of clauses 1-7, wherein the first virtual sound source includes an ambient virtual sound source, and the first loudspeaker includes a subwoofer.

[0124] 9. The computer-implemented method according to any one of clauses 1-8, wherein the first virtual sound source includes an ambient virtual sound source, and further comprising: determining that the first loudspeaker is included in a spatial loudspeaker system that includes a subset of loudspeakers within a plurality of loudspeakers; mapping the sound component to each of the loudspeakers included in the plurality of loudspeakers in addition to the first loudspeaker; and outputting at least a portion of the sound component for playback on each of the loudspeakers included in the plurality of loudspeakers in addition to the first loudspeaker.

[0125] 10. In some embodiments, a computer-readable storage medium includes instructions that, when executed by a processor, cause the processor to generate an audio scene for an extended reality (XR) environment by performing the following steps: determining that a first virtual sound source associated with the XR environment affects a sound in the audio scene; generating a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene; mapping the sound component to a first speaker included in a plurality of speakers based on an audio frequency present in the sound component; and outputting the sound component for playback on the first speaker.

[0126] 11. The computer-readable storage medium according to clause 10, wherein the first virtual sound source includes an ambient virtual sound source, and the first speaker includes a subwoofer.

[0127] 12. The computer-readable storage medium according to clause 10 or clause 11, further comprising: determining that the first virtual sound source is placed at a fixed virtual location; classifying the sound component associated with the first virtual sound source as a localized virtual sound source; and determining that the first speaker is closer to the fixed virtual location than a second speaker included in the plurality of speakers.

[0128] 13. The computer-readable storage medium according to any one of clauses 10-12, further comprising: determining that the first virtual sound source is placed at a fixed virtual location; classifying the sound component associated with the first virtual sound source as a localized virtual sound source; determining that each of the first speaker and a second speaker included in the plurality of speakers is closer to the fixed virtual location than a third speaker included in the plurality of speakers; mapping the sound component to the second speaker; and outputting at least a second portion of the sound component for playback on the second speaker.

[0129] 14. The computer-readable storage medium according to any one of clauses 10-13, further comprising: determining that the first virtual sound source has moved from a first virtual location to a second virtual location; classifying the sound component associated with the first virtual sound source as a moving virtual sound source; and determining that the first speaker is closer to the second virtual location than a second speaker included in the plurality of speakers.

[0130] 15. The computer-readable storage medium according to any one of clauses 10-14, further comprising: determining that the first virtual sound source has moved from the second virtual location to a third virtual location; determining that the second speaker is closer to the third virtual location than the first speaker; removing at least a first portion of the sound component that is not output to the first speaker; mapping the sound component to the second speaker; and outputting at least a second portion of the sound component for playback on the second speaker.

[0131] 16. The computer-readable storage medium according to any one of clauses 10-15 further includes: determining that the first virtual sound source includes a sound component below a specified frequency; classifying the sound component as a surrounding virtual sound source; mapping the sound component to each of a plurality of speakers in addition to the first speaker; and outputting at least a portion of the sound component for playback on each of the plurality of speakers in addition to the first speaker.

[0132] 17. The computer-readable storage medium according to any one of clauses 10-16, wherein the first virtual sound source includes a low-frequency sound component and the first speaker includes a subwoofer.

[0133] 18. The computer-readable storage medium according to any one of clauses 10-17, wherein the first virtual sound source includes at least one of a mid-frequency sound component and a high-frequency sound component, and wherein the first speaker is included in a spatial speaker system that includes a subset of speakers within a plurality of speakers.

[0134] 19. The computer-readable storage medium according to any one of clauses 10-18, wherein the first speaker is within a threshold distance of the first virtual sound source and the first speaker includes a head-mounted speaker.

[0135] 20. In some embodiments, a system includes: a plurality of speakers; and an audio processing system coupled to the plurality of speakers and configured to: determine that a first virtual object included in an extended reality (XR) environment is associated with a first virtual sound source; determine that the first virtual sound source affects sound in an audio scene associated with the XR environment; generate a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene; map the sound component to a first speaker included in a plurality of speakers; and output the sound component for playback on the first speaker.

[0136] Any one of the claim elements recited in any of the claims in any manner and / or any and all combinations of any of the elements described in this application fall within the intended scope of the present disclosure and protection.

[0137] The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0138] Aspects of the present implementation can be embodied as a system, a method, or a computer program product. Thus, aspects of the present disclosure may take the form of an entirely hardware implementation, an entirely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software aspects and hardware aspects, which may all be generally referred to herein as a "module" or a "system". In addition, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0139] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0140] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device support the implementation of the functions / acts specified in one or more blocks of the flowchart and / or block diagram. Such a processor may be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable processor.

[0141] The flow charts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of possible implementations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each frame in the flow chart or block diagram may represent a code module, a code segment or a code portion, which includes one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative implementations, the functions indicated in the frame may appear in an order different from the order indicated in the accompanying drawings. For example, the two frames shown in succession can actually be executed roughly in parallel, or the frames can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each frame in the block diagram and / or the flow chart diagram and the frame combination in the block diagram and / or the flow chart diagram can be implemented by a system based on special-purpose hardware, and the system based on special-purpose hardware performs a specified function or action, or is implemented by a combination of special-purpose hardware and computer instructions.

[0142] While the foregoing is directed to various embodiments of the present disclosure, other and further embodiments of the present disclosure may be envisaged without departing from the basic scope of the present disclosure, and the scope of the present disclosure is determined by the following claims.

Claims

1. A computer-implemented method for generating an audio scene for an extended reality (XR) environment, the method comprises: determining that a first virtual sound source associated with the XR environment affects a sound in the audio scene; generating a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene; mapping the sound component to a first loudspeaker included in a plurality of loudspeakers by calculating a cost function to determine the contribution of the first virtual sound source within the audio scene to one or more of the plurality of loudspeakers, wherein the cost function of the first loudspeaker is based on: a frequency deviation function that prioritizes the spatialization of sound sources with higher frequencies, and the absolute distance between the angle of the first virtual sound source and the first loudspeaker relative to the user; and outputting at least a first portion of the sound component for playback on the first loudspeaker.

2. The computer-implemented method according to claim 1, wherein the first virtual sound source includes a positioned virtual sound source, and wherein calculating the cost function comprises: determining a virtual position associated with the first virtual sound source; and determining that a first cost of the first loudspeaker for the virtual position is lower than a second cost of a second loudspeaker included in the plurality of loudspeakers.

3. The computer-implemented method according to claim 2, wherein the first virtual sound source includes a positioned virtual sound source, and wherein calculating the cost function comprises: determining that the first loudspeaker is included in a spatial loudspeaker system that includes a subset of the loudspeakers within the plurality of loudspeakers, determining a virtual position associated with the first virtual sound source, and determining that a first cost of the first loudspeaker for the virtual position and a second cost of a second loudspeaker included in the subset of the loudspeakers are each lower than a third cost of a third loudspeaker included in the subset of the loudspeakers; and further comprises: mapping the sound component to the second loudspeaker; and outputting at least a second portion of the sound component for playback on the second loudspeaker.

4. The computer-implemented method according to claim 2, wherein the first virtual sound source includes a moving virtual sound source, and wherein calculating the cost function comprises: determining that the first virtual sound source has moved from a first virtual position to a second virtual position; and determining that a first cost of the first loudspeaker for the second virtual position is lower than a second cost of a second loudspeaker included in the plurality of loudspeakers.

5. The computer-implemented method according to claim 4, wherein calculating the cost function further comprises: determining that the first virtual sound source has moved from the second virtual position to a third virtual position, and determining that a second cost of the second loudspeaker for the third virtual position is lower than the first cost of the first loudspeaker; and further comprises: removing at least a first portion of the sound component that is not output to the first loudspeaker; mapping the sound component to the second loudspeaker; and Output at least a second portion of the sound component for playback on the second loudspeaker.

6. The computer-implemented method according to claim 2, wherein a second virtual sound source associated with the XR environment affects the sound in the audio scene, and wherein calculating the cost function further comprises: determining that the second virtual sound source includes a virtual sound absorber that absorbs at least a portion of the sound component associated with the first virtual sound source, and determining an absorption value based on at least the portion of the sound component associated with the first virtual sound source; and further comprises reducing at least the portion of the sound component associated with the first virtual sound source based on the absorption value.

7. The computer-implemented method according to claim 2, wherein a second virtual sound source associated with the XR environment affects the sound in the audio scene, and wherein calculating the cost function further comprises: determining that the second virtual sound source includes a virtual sound reflector that reflects at least a portion of the sound component associated with the first virtual sound source, and determining a reflection value based on at least the portion of the sound component associated with the first virtual sound source; and further comprises increasing at least the portion of the sound component associated with the first virtual sound source based on the reflection value.

8. The computer-implemented method according to claim 2, wherein the first virtual sound source includes an ambient virtual sound source, and wherein calculating the cost function further comprises: determining that the first loudspeaker is included in a spatial loudspeaker system that includes a subset of the loudspeakers within the plurality of loudspeakers; and further comprises: mapping the sound component to each of the loudspeakers included in the plurality of loudspeakers in addition to the first loudspeaker; and outputting at least a portion of the sound component for playback on each of the loudspeakers included in the plurality of loudspeakers in addition to the first loudspeaker.

9. The computer-implemented method according to claim 1, wherein the first virtual sound source includes an ambient virtual sound source, and the first loudspeaker includes a subwoofer.

10. A computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to generate an audio scene for an extended reality (XR) environment by performing the following steps: determining that a first virtual sound source associated with the XR environment affects the sound in the audio scene; generating a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene; mapping the sound component to a first loudspeaker included in a plurality of loudspeakers by calculating a cost function to determine the contribution of the first virtual sound source to one or more of the loudspeakers included in the plurality of loudspeakers within the audio scene, wherein the cost function of the first loudspeaker is based on the following: a frequency deviation function that prioritizes the spatialization of sound sources with higher frequencies, and the absolute distance between the angle of the first virtual sound source and the first loudspeaker relative to the user; and Output at least a first portion of the sound component for playback on the first loudspeaker.

11. The computer-readable storage medium according to claim 10, wherein the sound component is mapped based on audio frequencies present in the sound component.

12. The computer-readable storage medium according to claim 10, the computer-readable storage medium further comprises: Determine that the first virtual sound source is placed at a fixed virtual position; Classify the sound component associated with the first virtual sound source as a localized virtual sound source; and Determine that the first loudspeaker is closer to the fixed virtual position than a second loudspeaker included in the plurality of loudspeakers.

13. The computer-readable storage medium according to claim 10, the computer-readable storage medium further comprises: Determine that the first virtual sound source is placed at a fixed virtual position; Classify the sound component associated with the first virtual sound source as a localized virtual sound source; Determine that each of the first loudspeaker and a second loudspeaker included in the plurality of loudspeakers is closer to the fixed virtual position than a third loudspeaker included in the plurality of loudspeakers; Map the sound component to the second loudspeaker; and Output at least a second portion of the sound component for playback on the second loudspeaker.

14. The computer-readable storage medium according to claim 10, the computer-readable storage medium further comprises: Determine that the first virtual sound source has moved from a first virtual position to a second virtual position; Classify the sound component associated with the first virtual sound source as a moving virtual sound source; and Determine that the first loudspeaker is closer to the second virtual position than a second loudspeaker included in the plurality of loudspeakers.

15. The computer-readable storage medium according to claim 14, the computer-readable storage medium further comprises: Determine that the first virtual sound source has moved from the second virtual position to a third virtual position; Determine that the second loudspeaker is closer to the third virtual position than the first loudspeaker; Remove the at least first portion of the sound component that is not output to the first loudspeaker; Map the sound component to the second loudspeaker; and Output at least a second portion of the sound component for playback on the second loudspeaker.

16. The computer-readable storage medium according to claim 10, the computer-readable storage medium further comprises: Determine that the first virtual sound source includes a sound component below a specified frequency; Classify the sound component as a surrounding virtual sound source; In addition to the first loudspeaker, map the sound component to each loudspeaker included in the plurality of loudspeakers; and Output at least a portion of the sound component for playback on each loudspeaker included in the plurality of loudspeakers in addition to the first loudspeaker.

17. The computer-readable storage medium according to claim 10, wherein the first virtual sound source includes a low-frequency sound component, and the first loudspeaker includes a subwoofer.

18. The computer-readable storage medium according to claim 10, wherein the first virtual sound source includes at least one of a mid-frequency sound component and a high-frequency sound component, and wherein the first loudspeaker is included in a spatial loudspeaker system that includes a subset of the loudspeakers within the plurality of loudspeakers.

19. The computer-readable storage medium according to claim 10, wherein the first loudspeaker is within a threshold distance from the first virtual sound source, and the first loudspeaker includes a head-mounted loudspeaker.

20. A system, the system comprising: a plurality of loudspeakers; and an audio processing system coupled to the plurality of loudspeakers and configured to: determine that a first virtual object included in an extended reality (XR) environment is associated with a first virtual sound source; determine that the first virtual sound source affects sound in an audio scene associated with the XR environment; generate a sound component associated with the first virtual sound source based on the contribution of the first virtual sound source to the audio scene; map the sound component to a first loudspeaker included in the plurality of loudspeakers by calculating a cost function to determine the contribution of the first virtual sound source within the audio scene to one or more of the loudspeakers included in the plurality of loudspeakers, wherein the cost function of the first loudspeaker is based on: a frequency deviation function that prioritizes the spatialization of sound sources with higher frequencies, and the absolute distance between the angle of the first virtual sound source and the first loudspeaker relative to the user; and output the sound component for playback on the first loudspeaker.

Citation Information

Patent Citations

  • System for dynamically creating and rendering audio objects

    US20120232910A1

  • Augmented reality (AR) audio with position and action triggered virtual sound effects

    US20130236040A1

  • Spatial audio with remote speakers

    US20160212538A1

  • Panning of Audio Objects to Arbitrary Speaker Layouts

    US20160212559A1