Controlling rendering of audio data
By controlling audio rendering in the audio playback system and generating audio renderers for internal and external regions using metadata and listener location, the problem of insufficient auditory experience in existing technologies is solved, achieving more efficient processor and bandwidth utilization and a more realistic XR experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2026-03-20
AI Technical Summary
Existing computer-mediated reality systems struggle to provide a realistic and immersive audio experience, especially as video experiences improve, resulting in insufficient auditory input that makes it difficult for users to identify the source of audio content.
By controlling audio rendering in the audio playback system, audio renderers for internal and external regions are generated using metadata and listener location. Processor cycles, memory, and bandwidth consumption can be flexibly adjusted to separate the boundaries between internal and external regions, providing a more realistic XR experience.
It improves the operational efficiency of the audio playback system, reduces processor cycles and bandwidth consumption, and provides a more immersive XR experience, allowing users to be more realistically placed in the XR environment.
Smart Images

Figure CN116195276B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Application Serial No. 17 / 469,421, filed September 8, 2021, entitled “CONTROLLING RENDERING OF AUDIO DATA,” and U.S. Provisional Application Serial No. 63 / 085,437, filed September 30, 2020, entitled “CONTROLLING RENDERING OF AUDIO DATA,” the entire contents of each of which are incorporated herein by reference. U.S. Application Serial No. 17 / 469,421, filed September 8, 2021, claims the benefit of U.S. Provisional Application Serial No. 63 / 085,437, filed September 30, 2020. In addition, this application is related to U.S. Patent Application Serial No. 17 / 038,618, filed September 30, 2020, entitled “CONTROLLING RENDERING OF AUDIO DATA,” and U.S. Provisional Application Serial No. 62 / 909,104, filed October 1, 2019, entitled “CONTROLLING RENDERING OF AUDIO DATA.” TECHNICAL FIELD
[0002] The present disclosure relates to processing of audio data. BACKGROUND
[0003] Computer-mediated reality systems are under development to allow a computing device to augment or add to, subtract or subtract from, or generally modify an existing reality experienced by a user. Computer-mediated reality systems (which can also be referred to as “extended reality systems” or “XR systems”) can include, for example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceptual success of a computer-mediated reality system often involves the ability of such a computer-mediated reality system to provide a realistic, immersive experience in both the video and audio experience in a way that aligns with the way a user expects. While the human visual system is more sensitive than the human auditory system (e.g., in terms of perceived localization of various objects within a scene), ensuring a sufficient auditory experience is an increasingly important factor in ensuring a realistic, immersive experience, particularly as the video experience improves to allow for better localization of video objects, which enables a user to better identify the source of audio content. SUMMARY
[0004] The present disclosure generally relates to techniques for controlling audio rendering at an audio playback system. The techniques can enable an audio playback system to perform flexible rendering in terms of complexity (as defined by processor cycles, memory, and / or consumed bandwidth) while also allowing for internal and external rendering for an XR experience, as defined by a boundary separating an internal region and an external region. Moreover, the audio playback system can utilize metadata or other indications specified in a bitstream representing audio data to configure an audio Tenderer while also referencing a listener position relative to the boundary to generate an audio Tenderer for accounting for the internal region or the external region. The boundary can also be referred to herein as a range or a spatial range.
[0005] Accordingly, the techniques can improve operation of an audio playback system as the audio playback system can reduce the amount of processor cycles, memory, and / or consumed bandwidth when configured to perform a low complexity rendering. When performing a high complexity rendering, the audio playback system can provide a more immersive XR experience, which can result in a user of the audio playback system being more realistically placed in the XR experience.
[0006] In one example, the techniques are directed to a device configured to process audio data, the device comprising: a memory configured to store one or more speaker feeds; and one or more processors implemented in circuitry and communicatively coupled to the memory, the one or more processors configured to: determine whether a boundary separating an internal region and an external region is present; determine, based on a determination that the boundary is present, a transition distance value, the transition distance value indicating a size of a transition zone; obtain a listener position, the listener position indicating a virtual position of the device relative to the internal region; obtain, based at least in part on the boundary and the listener position, a current Tenderer; apply the current Tenderer to the audio data to obtain the one or more speaker feeds.
[0007] In another example, the techniques are directed to a method for processing audio data, the method comprising: determining whether a boundary separating an internal region and an external region is present; determining, based on a determination that the boundary is present, a transition distance value, the transition distance value indicating a size of a transition zone; obtaining a listener position, the listener position indicating a virtual position of the device relative to the internal region; obtaining, based at least in part on the boundary and the listener position, a current Tenderer; applying the current Tenderer to the audio data to obtain the one or more speaker feeds; and storing the one or more speaker feeds.
[0008] In another example, the technology is directed to a non-transitory computer- readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine whether a boundary separating an interior region and an exterior region exists; determine, based on a determination that the boundary exists, a transition distance value that indicates a size of a transition zone; obtain a listener position that indicates a virtual position of a device relative to the interior region; obtain, based at least in part on the boundary and the listener position, a current renderer; apply the current renderer to audio data to obtain one or more speaker feeds; and store the one or more speaker feeds.
[0009] In one example, the technology is directed to a device configured to process one or more audio streams, the device comprising: one or more processors configured to: determine whether a boundary separating an interior region and an exterior region exists; determine, based on the boundary existing, a transition distance value that indicates a size of a transition zone, wherein the transition distance value is 0; obtain a listener position that indicates a position of the device relative to the interior region; obtain, based on the boundary, the listener position, and the transition distance value being 0, a current renderer that is either an interior renderer configured to render audio data for the interior region or an exterior renderer configured to render audio data for the exterior region; apply the current renderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0010] In one example, the technology is directed to a device configured to process one or more audio streams, the device comprising: one or more processors configured to: determine whether a boundary separating an interior region and an exterior region exists; determine, based on the boundary existing, a transition distance value that indicates a size of a transition zone, wherein the transition distance value is greater than 0; obtain a listener position that indicates a position of the device relative to the interior region; obtain, based on the boundary, the listener position, and the transition distance value being greater than 0, a current renderer that is either an interior renderer configured to render audio data for the interior region or an exterior renderer configured to render audio data for the exterior region or both the interior renderer and the exterior renderer; apply the current renderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0011] In another example, the technology is directed to a method for processing one or more audio streams, the method comprising: determining whether a boundary separating an interior region and an exterior region exists; determining a transition distance value based on the boundary existing, the transition distance value indicating a size of a transition zone, wherein the transition distance value is 0; obtaining a listener position, the listener position indicating a position of a device relative to the interior region; obtaining a current renderer based on the boundary, the listener position, and the transition distance value being 0, the current renderer being either as an interior renderer configured to render audio data for the interior region or as an exterior renderer configured to render audio data for the exterior region; applying the current renderer to the audio data to obtain one or more speaker feeds; and storing the one or more speaker feeds.
[0012] In another example, the technology is directed to a method for processing one or more audio streams, the method comprising: determining whether a boundary separating an interior region and an exterior region exists; determining a transition distance value based on the boundary existing, the transition distance value indicating a size of a transition zone, wherein the transition distance value is greater than 0; obtaining a listener position, the listener position indicating a position of a device relative to the interior region; obtaining a current renderer based on the boundary, the listener position, and the transition distance value being greater than 0, the current renderer being either as an interior renderer configured to render audio data for the interior region or as an exterior renderer configured to render audio data for the exterior region or as both the interior renderer and the exterior renderer; applying the current renderer to the audio data to obtain one or more speaker feeds; and storing the one or more speaker feeds.
[0013] In another example, the technology is directed to a device configured to process one or more audio streams, the device comprising: means for determining whether a boundary separating an interior region and an exterior region exists; means for determining a transition distance value based on the boundary existing, the transition distance value indicating a size of a transition zone, wherein the transition distance value is 0; means for obtaining a listener position, the listener position indicating a position of a device relative to the interior region; means for obtaining a current renderer based on the boundary, the listener position, and the transition distance value being 0, the current renderer being either as an interior renderer configured to render audio data for the interior region or as an exterior renderer configured to render audio data for the exterior region; means for applying the current renderer to the audio data to obtain one or more speaker feeds; and means for storing the one or more speaker feeds.
[0014] In another example, the technique pertains to a device configured to process one or more audio streams, the device comprising: determining whether a boundary exists separating an inner region and an outer region; determining a transition distance value based on the existence of the boundary, the transition distance value indicating the size of a transition zone, wherein the transition distance value is greater than 0; obtaining a listener position indicating the position of the device relative to the inner region; obtaining a current renderer based on the boundary, the listener position, and the transition distance value being greater than 0, the current renderer being either an inner renderer configured to render audio data for the inner region, or an outer renderer configured to render audio data for the outer region, or both an inner and outer renderer; applying the current renderer to the audio data to obtain one or more speaker feeds; and storing the one or more speaker feeds.
[0015] In another example, the technique targets a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: determine whether a boundary exists separating an inner region and an outer region; determine a transition distance value based on the existence of the boundary, the transition distance value indicating the size of a transition zone, wherein the transition distance value is 0; obtain a listener position indicating the position of a device relative to the inner region; obtain a current renderer based on the boundary, the listener position, and the transition distance value of 0, the current renderer being either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; apply the current renderer to the audio data to obtain one or more speaker feeds; and store one or more speaker feeds.
[0016] In another example, the technique targets a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: determine whether a boundary exists separating an inner region and an outer region; determine a transition distance value based on the existence of the boundary, the transition distance value indicating the size of a transition zone, wherein the transition distance value is greater than 0; obtain a listener position indicating the position of a device relative to the inner region; obtain a current renderer based on the boundary, the listener position, and the transition distance value being greater than 0, the current renderer being either an inner renderer configured to render audio data for the inner region, or an outer renderer configured to render audio data for the outer region, or both an inner and outer renderer; apply the current renderer to the audio data to obtain one or more speaker feeds; and store one or more speaker feeds.
[0017] Details of one or more examples of this disclosure are set forth in the accompanying drawings and the description below. Further features, objects, and advantages of various aspects of the technology will become apparent from the description, the drawings, and the claims. Attached Figure Description
[0018] FIG. 1A and FIG. 1B This is a diagram illustrating a system that can perform various aspects of the techniques described in this disclosure.
[0019] FIG. 2 This is a diagram illustrating examples of low-complexity rendering of extended reality (XR) scenes based on various aspects of the techniques described in this disclosure.
[0020] FIG. 3 This is a diagram illustrating examples of highly complex rendering of XR scenes, including transition distances, based on various aspects of the techniques described in this disclosure.
[0021] FIG. 4A and FIG. 4B This is a diagram showing an example of a VR device.
[0022] FIG. 5A and FIG. 5B This is a diagram illustrating example systems that can perform various aspects of the techniques described in this disclosure.
[0023] FIG. 6A - FIG. 6G yes FIG. 1A and FIG. 1B The example shown is a block diagram of an example audio playback system performing various aspects of the techniques described in this disclosure.
[0024] FIG. 7 This is a diagram illustrating examples of rendering extended reality (XR) scenes using various aspects of the techniques described in this disclosure.
[0025] FIG. 8 This is another example of rendering an extended reality (XR) scene, illustrating various aspects of the techniques described in this disclosure.
[0026] FIG. 9 This is another example of rendering an extended reality (XR) scene, illustrating various aspects of the techniques described in this disclosure.
[0027] FIG. 10 This is another example of rendering an extended reality (XR) scene, illustrating various aspects of the techniques described in this disclosure.
[0028] FIG. 11 This is a flowchart illustrating the example rendering techniques disclosed herein.
[0029] FIG. 12 An example of a wireless communications system that supports audio streaming is shown in accordance with aspects of the present disclosure.
[0030] FIG. 13 is a flow diagram illustrating FIG. 1A example operations of a source device shown in
[0031] FIG. 14 is a flow diagram illustrating FIG. 1A example operations of a content consumer device shown in
[0032] FIG. 15 is a flow diagram illustrating example audio processing techniques in accordance with various aspects of the present disclosure. DETAILED DESCRIPTION
[0033] There are many different ways to represent a soundfield. Example formats include a channel-based audio format, an object-based audio format, and a scene-based audio format. A channel-based audio format refers to a 5.1 surround sound format, a 7.1 surround sound format, a 22.2 surround sound format, or any other channel-based format that positions audio channels to specific locations around a listener in order to recreate a soundfield.
[0034] An object-based audio format can refer to a format in which audio objects, often encoded using pulse code modulation (PCM) and referred to as PCM audio objects, are specified in order to represent a soundfield. Such audio objects can include metadata that identifies a position of the audio object relative to a listener or other reference point in a soundfield, such that the audio objects can be rendered to one or more speaker channels for playback in an effort to recreate the soundfield. The techniques described in this disclosure can be applied to any of the aforementioned formats, including a scene-based audio format, a channel-based audio format, an object-based audio format, or any combination thereof.
[0035] A scene-based audio format can include a hierarchical set of elements that define a soundfield in three dimensions. One example of a hierarchical set of elements is a set of spherical harmonic coefficients (SHCs). The following expression demonstrates a description or representation of a soundfield using SHCs:
[0036]
[0037] The expression shows that at a point of the soundfield at time t, i can be uniquely represented by SHCs, Here, c is the speed of sound (~343 m / s), is a reference point (or observation point), j n (·) is the n-th order spherical Bessel function, and is an n-th order and m-th sub-order spherical harmonic basis function (which can also be referred to as a spherical basis function). It can be appreciated that the term in brackets is a frequency domain representation of the signal (i.e., It can be approximated with various time-frequency transforms (e.g., a discrete Fourier transform (DFT), a discrete cosine transform (DCT), or a wavelet transform). Other examples of a set of levels include a set of wavelet transform coefficients and other sets of coefficients of multi-resolution basis functions.
[0038] The SHCs (which can also be referred to as ambisonic coefficients) can be physically acquired (e.g., recorded) by various microphone array configurations, or alternatively, they can be derived from a channel-based or object-based description of the soundfield. The SHCs represent a scene-based audio, where the SHCs can be input to an audio encoder to obtain encoded SHCs, which can facilitate more efficient transmission or storage. For example, a fourth-order representation involving (1 + 4) 2 (25, so fourth order) coefficients.
[0039] As noted above, the SHCs can be derived from microphone recordings using a microphone array. Various examples of how to physically acquire SHCs from a microphone array are described in Poletti, M, “Three-Dimensional Surround Sound Systems Based on Spherical Harmonics,” J. Audio Eng. Soc, Vol. 53, No. 11, 2005 November, pp. 1004-1025.
[0040] The following equation can illustrate how to derive SHCs from an object-based description. The coefficients of the soundfield corresponding to individual audio objects can be represented as:
[0041]
[0042] where i is is an n-th order (second kind) spherical Hankel function, and is the position of the object. Knowing the object source energy g(co) as a function of frequency (e.g., using time-frequency analysis techniques, such as performing a fast Fourier transform on a pulse-code modulated (PCM) stream) can enable converting each PCM object and corresponding position to Furthermore, (since the above is a linear and orthogonal decomposition) it is possible to see that each object The coefficients are additive. In this way, multiple PCM objects can be generated by... The coefficients are represented (e.g., as the sum of coefficient vectors of individual objects). These coefficients can contain information about the sound field (based on pressure in 3D coordinates), and are shown above at the viewpoint. The transformation from a single object to the representation of the entire sound field in the vicinity.
[0043] Computer-mediated reality systems (also known as “extended reality systems” or “XR systems”) are under development to take advantage of the many potential benefits offered by surround sound coefficients. For example, surround sound coefficients can represent a sound field in three dimensions in a way that can potentially enable accurate three-dimensional (3D) localization of sound sources within the sound field. Therefore, XR devices can render surround sound coefficients to a speaker feed that, when played back through one or more speakers, accurately reproduces the sound field.
[0044] The use of surround sound coefficients in XR enables the development of numerous use cases that rely on a more immersive sound field provided by surround sound coefficients, particularly for computer gaming applications and real-time video streaming applications. In these highly dynamic use cases that depend on low-latency reproduction of the sound field, XR devices may prefer surround sound coefficients over other representations that are more difficult to manipulate or involve complex rendering. More information on these use cases will be discussed below. FIG. 1A and FIG. 1B supply.
[0045] While this disclosure describes VR devices, various aspects of the technology can be performed in the context of other devices, such as mobile devices. In this example, a mobile device (e.g., a so-called smartphone) can present the displayed world via a screen that can be mounted on a user's head or viewed as in normal mobile device use. Therefore, any information on the screen can be part of the mobile device. The mobile device is capable of providing tracking information 41, allowing both the VR experience (when worn) and the normal experience to view the displayed world, where the normal experience can still allow the user to view the displayed world, thus displaying VR-lite type experiences (e.g., raising the device and rotating or panning it to view different parts of the displayed world).
[0046] This disclosure provides various combinations of opacity and interpolation distance properties for rendering internal surround sound fields in 6DoF (and other) use cases. Furthermore, this disclosure discusses examples of low-complexity and high-complexity rendering solutions for internal surround sound fields that can be specified by a single bit. In an example encoder input format (EIF), there may be properties indicating whether the surround sound field description is an internal or external field. In an internal sound field, the sound source is within a specified boundary described by a mesh or simple geometric object, while for an external sound field, the sound source is described as being outside the boundary. The opacity property used for the internal sound field can specify whether the contribution of a sound field that does not have a direct line of sight to the listener contributes to rendering the sound field for the listener when the listener is outside the boundary. Additionally, the distance property can specify a buffer region around the boundary, where interpolation is used between rendering the internal field for the external listener and the internal listener. As used herein, the buffer region may also be referred to as a transition distance.
[0047] Therefore, various aspects of the technology described herein enable the determination of a user's listener position during navigation in VR or other XR settings, whether the listener position is within a geometric boundary (where all sound sources radiate unobstructed toward the listener within the geometric boundary), and whether the listener position is outside the geometric boundary. When the listener position is determined to be outside the geometric boundary, these aspects also enable the assignment of opacity attributes to each sound source obstructed relative to the listener, and, when the listener position indicates that the listener is outside the geometric boundary, the interpolation of the sound field within the geometric boundary based on the opacity attributes, and the rendering of the interpolated sound field.
[0048] FIG. 1A and FIG. 1B This is a diagram illustrating various aspects of a system capable of performing the techniques described in this disclosure. For example... FIG. 1A As illustrated in the example, system 10 includes a source device 12A and a content consumer device 14A. While described in the context of source device 12A and content consumer device 14A, these techniques can be implemented in any context in which any hierarchical representation of the sound field is encoded to form a bitstream representing audio data. Furthermore, source device 12A can represent any form of computing device capable of generating hierarchical representations of the sound field, and is generally described herein in the context of a VR content creator device. Similarly, content consumer device 14A can represent any form of computing device capable of implementing the audio stream interpolation techniques and audio playback described in this disclosure, and is generally described herein in the context of a VR client device.
[0049] Source device 12A can be operated by an entertainment company or other entity that can generate multi-channel audio content for consumption by an operator of a content consumer device (e.g., content consumer device 14A). In many VR scenarios, source device 12A generates audio content in conjunction with video content. Source device 12A includes content capture device 300 and content soundfield representation generator 302.
[0050] Content capture device 300 can be configured to interface or otherwise communicate with one or more microphones 5A-5N (“microphones 5”). Microphones 5 can represent or other types of 3D audio microphones that are capable of capturing a soundfield and representing it as corresponding scene-based audio data 11A-11N (which can also be referred to as ambisonic coefficients 11A-11N or “ambisonic coefficients 11”). In the context of scene-based audio data 11 (which is another way of referring to ambisonic coefficients 11), each of microphones 5 can represent a cluster of microphones arranged within a single housing according to a set geometry that facilitates generation of ambisonic coefficients 11. Thus, the term “microphone” can refer to a cluster of microphones (which are effectively geometrically arranged transducers) or a single microphone (which can be referred to as a spot microphone).
[0051] Ambisonic coefficients 11 can represent one example of an audio stream. Thus, ambisonic coefficients 11 can also be referred to as audio stream 11. Although primarily described with respect to ambisonic coefficients 11, these techniques can be performed with respect to other types of audio streams, including pulse code modulated (PCM) audio streams, channel-based audio streams, object-based audio streams, etc.
[0052] In some examples, content capture device 300 can include integrated microphones that are integrated into a housing of content capture device 300. Content capture device 300 can interface with microphones 5 wirelessly or via a wired connection. In contrast to capturing audio data via microphones 5 or in conjunction with capturing audio data, content capture device 300 can process ambisonic coefficients 11 after inputting ambisonic coefficients 11 via some type of removable storage device, wirelessly, and / or via a wired input process, or alternatively or in conjunction with the foregoing, generate or otherwise create ambisonic coefficients 11 (from stored sound samples, for example, as is common in game applications). Thus, various combinations of content capture device 300 and microphones 5 are possible.
[0053] The content capture device 300 can also be configured to interface with or otherwise be in communication with a soundfield representation generator 302. The soundfield representation generator 302 can comprise any type of hardware device capable of interfacing with the content capture device 300. The soundfield representation generator 302 can use the ambisonic coefficients 11 provided by the content capture device 300 to generate various representations of the same soundfield represented by the ambisonic coefficients 11.
[0054] For example, to generate different representations of a soundfield using stereo coefficients (again, this is one example of an audio stream), the soundfield representation generator 302 can use a coding scheme for a stereo representation of a soundfield known as mixed-order ambisonics (MOA), as discussed in greater detail in U.S. Application Serial No. 15 / 672,058, filed August 8, 2017, entitled “MIXED-ORDER AMBISONICS (MOA) AUDIO DATA FOR COMPUTER-MEDIATED REALITY SYSTEMS,” which was published as U.S. Patent Publication No. 20190007781 on January 3, 2019.
[0055] To generate a particular MOA representation of a soundfield, the soundfield representation generator 302 can generate a partial subset of the entire set of stereo coefficients. For example, each MOA representation generated by the soundfield representation generator 302 can provide a certain degree of accuracy with respect to some regions of the soundfield, but less accuracy in other regions. In one example, an MOA representation of a soundfield can include eight (8) uncompressed stereo coefficients, while a third-order stereo representation of the same soundfield can include sixteen (16) uncompressed stereo coefficients. Thus, each MOA representation of a soundfield generated as a partial subset of stereo coefficients can be less storage intensive and less bandwidth intensive (if and when transmitted as part of a bitstream 27 over a transmission channel shown) than a corresponding third-order ambisonic representation of the same soundfield generated from ambisonic coefficients.
[0056] Although described with respect to MOA representations, the techniques of the present disclosure can also be performed with respect to first-order ambisonic (FOA) representations, in which all ambisonic coefficients associated with first- and zero-order spherical basis functions are used to represent a soundfield. In other words, as opposed to using a partial non-zero subset of ambisonic coefficients to represent a soundfield, the soundfield representation generator 302 can use all ambisonic coefficients of a given order N to represent a soundfield, resulting in a total number of ambisonic coefficients equal to (N + 1) 2 .
[0057] In this regard, the surround sound audio data (which is another way of referring to the surround sound coefficients in the MOA representation or the full rank representation (e.g., the first rank representation noted above)) can include surround sound coefficients associated with spherical basis functions having a first rank or less (which can be referred to as “first rank surround sound audio data”), surround sound coefficients associated with spherical basis functions having a hybrid rank and sub-rank (which can be referred to as the “MOA representation” discussed above), or surround sound coefficients associated with spherical basis functions having a rank greater than one (referred to above as the “full rank representation”).
[0058] In some examples, the content capture device 300 can be configured to communicate wirelessly with the sound field representation generator 302. In some examples, the content capture device 300 can communicate with the sound field representation generator 302 via one or both of a wireless connection or a wired connection. Via the connection between the content capture device 300 and the sound field representation generator 302, the content capture device 300 can provide content in various forms, which are described herein for purposes of discussion as portions of the surround sound coefficients 11.
[0059] In some examples, the content capture device 300 can utilize various aspects of the sound field representation generator 302 (in terms of hardware or software capabilities of the sound field representation generator 302). For example, the sound field representation generator 302 can include specialized hardware configured to perform psychoacoustic audio encoding (or specialized software that, when executed, causes one or more processors to perform psychoacoustic audio encoding), such as the Unified Speech and Audio Coder denoted as “USAC” set forth by the Moving Picture Experts Group (MPEG), the MPEG-H 3D Audio Coding standard, the MPEG-I Immersive Audio standard, or a proprietary standard, such as AptX TM (including various versions of AptX, such as Enhanced AptX-E-AptX, AptX live, AptX stereo, and AptX High Definition-AptX-HD), Advanced Audio Coding (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey’s Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Windows Media Audio (WMA).
[0060] The content capture device 300 can not include psychoacoustic audio encoder- specific hardware or specific software, but rather provide the audio aspects of the content 301 in a non-psychoacoustic audio coded form. The soundfield representation generator 302 can assist in capturing the content 301 by at least partially performing psychoacoustic audio encoding with respect to the audio aspects of the content 301.
[0061] The soundfield representation generator 302 can also assist in content capture and transmission by generating one or more bitstreams 21 based at least in part on the audio content generated from the ambisonic coefficients 11 (e.g., the MOA representation, the third-order ambisonic representation, and / or the first-order ambisonic representation). The bitstreams 21 can represent compressed versions of the ambisonic coefficients 11 (and / or a partial subset thereof of the MOA representation used to form the soundfield) and any other different types of content 301 (e.g., compressed versions of spherical video data, image data, or textual data).
[0062] The soundfield representation generator 302 can generate the bitstreams 21 for transmission, e.g., across a transmission channel, which can be a wired or wireless channel, a data storage device, etc. The bitstreams 21 can represent encoded versions of the ambisonic coefficients 11 (and / or a partial subset thereof of the MOA representation used to form the soundfield) and can include a primary bitstream and another side bitstream, which can be referred to as side channel information. In some instances, the bitstreams 21 representing compressed versions of the ambisonic coefficients 11 can conform to bitstreams generated according to the MPEG-H 3D Audio Coding standard.
[0063] The content consumer device 14A can be operated by an individual and can represent a VR client device. While described with respect to a VR client device, the content consumer device 14A can represent other types of devices, e.g., an augmented reality (AR) client device, a mixed reality (MR) client device (or any other type of head-mounted display device or extended reality (XR) device), a standard computer, a headset, an earpiece, or any other device capable of tracking head movement and / or general translational movement of the individual operating the content consumer device 14A. As FIG. 1A As shown in the example of FIG. 1, the content consumer device 14A includes an audio playback system 16A, which can refer to any form of audio playback system capable of rendering ambisonic coefficients (whether in the form of first-, second-, and / or third-order ambisonic representations and / or MOA representations) for playback as multi-channel audio content.
[0064] The content consumer device 14A can retrieve the bitstreams 21 directly from the source device 12A. In some examples, the content consumer device 14A can interface with a network including a fifth generation (5G) cellular network to retrieve the bitstreams 21 or otherwise cause the source device 12A to transmit the bitstreams 21 to the content consumer device 14A.
[0065] Although shown as being transmitted directly to the content consumer device 14A in FIG. 1A The source device 12A can output the bitstream 21 to an intermediate device located between the source device 12A and the content consumer device 14A. The intermediate device can store the bitstream 21 for later delivery to the content consumer device 14A, which can request the bitstream. The intermediate device can include a file server, a web server, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a smart phone, or any other device capable of storing the bitstream 21 for later retrieval by an audio decoder. The intermediate device can reside in a content delivery network capable of streaming the bitstream 21 (and possibly in combination with sending a corresponding video data bitstream) to subscribers (e.g., the content consumer device 14A) that request the bitstream 21.
[0066] Alternatively, the source device 12A can store the bitstream 21 to a storage medium, e.g., a compact disc, a digital video disc, a high definition video disc, or other storage medium, where most are capable of being read by a computer and thus can be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, the transmission channel can refer to the channel through which the content stored to the medium is transmitted (and can include retail stores and other store-based delivery mechanisms). In any case, the techniques of this disclosure should therefore not be limited in this regard to FIG. 1A examples.
[0067] As noted above, the content consumer device 14A includes an audio playback system 16A. The audio playback system 16A can represent any system capable of playing back multi-channel audio data. The audio playback system 16A can include a plurality of different audio renderers 22. The renderers 22 can each provide different forms of audio rendering, where the different forms of rendering can include one or more of various ways of performing vector-based amplitude panning (VBAP) and / or one or more of various ways of performing soundfield synthesis. As used herein, “A and / or B” means “A or B” or “A and B” both.
[0068] The audio playback system 16A can also include an audio decoding device 24. The audio decoding device 24 can represent a device configured to decode the bitstream 21 to output reconstructed surround coefficients 11A’-11N’ (which can form a full first-, second-, and / or third-order ambisonic representation or a subset thereof or a decomposition thereof of the same soundfield, e.g., a dominant audio signal, ambient surround coefficients, and vector-based signals described in the MPEG-H 3D Audio Coding standard and / or the MPEG-I Immersive Audio standard).
[0069] Accordingly, the surround coefficients 11A’-11N’ (“surround coefficients 11’”) can be similar to the full set or a partial subset of the surround coefficients 11, but differ due to lossy operations (e.g., quantization) and / or transmission via a transmission channel. The audio playback system 16A can obtain the surround audio data 15 from the different streams of surround coefficients 11’ after decoding the bitstream 21 to obtain the surround coefficients 11’, and render the surround audio data 15 to the output speaker feeds 25. The speaker feeds 25 can drive one or more speakers (for ease of illustration, not shown in the example). The surround representation of the soundfield can be normalized in a variety of ways, including N3D, SN3D, FuMa, N2D, or SN2D. FIG. 1A
[0070] To select an appropriate renderer or, in some instances, generate an appropriate renderer, the audio playback system 16A can obtain speaker information 13 indicating a number of speakers and / or a spatial geometry of the speakers. In some instances, the audio playback system 16A can obtain the speaker information 13 using reference microphones and output signals to activate (or, in other words, drive) the speakers in a manner that dynamically determines the speaker information 13 via the reference microphones. In other instances, or in combination with the dynamic determination of the speaker information 13, the audio playback system 16A can prompt a user to interact with the audio playback system 16A and input the speaker information 13.
[0071] The audio playback system 16A can select one of the one or more audio renderers 22 based on the speaker information 13. In some instances, the audio playback system 16A can generate one of the one or more audio renderers 22 based on the speaker information 13 when none of the one or more audio renderers 22 are within a threshold similarity measure (in terms of speaker geometry) to the speaker geometry specified in the speaker information 13. In some instances, the audio playback system 16A can generate one of the one or more audio renderers 22 based on the speaker information 13 without first attempting to select an existing one of the one or more audio renderers 22.
[0072] When outputting the speaker feeds 25 to headphones, the audio playback system 16A can utilize one of the renderers 22 that provides binaural rendering using head-related transfer functions (HRTFs) or other functionality capable of rendering to left and right speaker feeds 25 for headphone speaker playback. The term “speaker” or “transducer” generally refers to any speaker, including a loudspeaker, a headphone speaker, etc. The rendered speaker feeds 25 can then be played back by one or more speakers.
[0073] While described as rendering a loudspeaker feed 25 from ambisonic audio data 15, references to rendering of the loudspeaker feed 25 can refer to other types of rendering, e.g., rendering that is directly incorporated into the decoding of ambisonic audio data 15 from bitstream 21. Examples of alternative rendering can be found in Appendix G of the MPEG-H 3D Audio Coding standard, where rendering occurs during the dominant signal formation and ambient signal formation prior to soundfield synthesis. Thus, references to rendering of ambisonic audio data 15 should be understood to refer to both the actual rendering of ambisonic audio data 15 or the decomposition of ambisonic audio data 15 or its representations (e.g., the dominant audio signals, ambient ambisonic coefficients, and / or vector-based signals, which can also be referred to as V-vectors, noted above).
[0074] As described above, content consumer device 14A can represent a VR device in which a human wearable display is mounted in front of the eyes of a user operating the VR device. FIG. 4A and FIG. 4B are diagrams illustrating examples of VR devices 400A and 400B. In the examples of FIG. 4A VR device 400A is coupled to or otherwise includes headphones 404 that can reproduce a soundfield represented by ambisonic audio data 15 (which is another way of referring to ambisonic coefficients) through playback of a loudspeaker feed 25. Loudspeaker feed 25 can represent an analog or digital signal that is capable of causing a diaphragm within a transducer of headphones 404 to vibrate at various frequencies. Such a process is often referred to as driving headphones 404.
[0075] Video, audio, and other sensory data can play an important role in a VR experience. To participate in a VR experience, user 402 can wear VR device 400A (which can also be referred to as VR headset 400A) or other wearable electronic device. A VR client device (e.g., VR headset 400A) can track head movements of user 402 and adjust video data shown via VR headset 400A to account for the head movements, providing an immersive experience in which user 402 can experience a virtual world shown in the video data in visual three-dimensions.
[0076] While VR (and other forms of AR and / or MR, which can generally be referred to as computer-mediated reality devices) can allow user 402 to visually reside in a virtual world, often VR headset 400A can lack the ability to aurally place the user in the virtual world. In other words, a VR system (which can include a computer (not shown in the examples of FIG. 4A to render video data and audio data) and VR headset 400A) can be unable to support full three-dimensional immersion aurally.
[0077] FIG. 4Bis a diagram illustrating an example of a wearable device 400B that can operate in accordance with various aspects of the techniques described in this disclosure. In various examples, the wearable device 400B can represent a VR headset (e.g., the VR headset 400A described above), an AR headset, an MR headset, or any other type of XR headset. Augmented reality “AR” can refer to computer-rendered images or data that are overlaid on the real world in which a user is actually located. Mixed reality “MR” can refer to computer-rendered images or data that are world-locked to a particular location in the real world, or can refer to a variant on VR in which partially computer-rendered 3D elements and partially photographed real elements are combined into an immersive experience that simulates the physical presence of the user in the environment. Extended reality “XR” can represent a catch-all term for VR, AR, and MR. More information on the XR terminology can be found in the document by Jason Peterson, titled “Virtual Reality, Augmented Reality, and Mixed Reality Definitions,” dated July 7, 2017.
[0078] The wearable device 400B can represent other types of devices, such as a watch (including a so-called “smart watch”), glasses (including a so-called “smart glasses”), earphones (including a so-called “wireless earphones” and “smart earphones”), smart clothing, smart jewelry, and the like. Whether the VR device represents a watch, glasses, and / or earphones, the wearable device 400B can communicate with a computing device that supports the wearable device 400B via a wired connection or a wireless connection.
[0079] In some instances, the computing device that supports the wearable device 400B can be integrated within the wearable device 400B, and thus, the wearable device 400B can be considered the same device as the computing device that supports the wearable device 400B. In other instances, the wearable device 400B can communicate with a separate computing device that can support the wearable device 400B. In this regard, the term “support” should not be understood to require a separate, dedicated device, but rather one or more processors configured to perform various aspects of the techniques described in this disclosure can be integrated within the wearable device 400B or within a computing device that is separate from the wearable device 400B.
[0080] For example, when the wearable device 400B represents an example of a VR device 400B, a separate, dedicated computing device (e.g., a personal computer including one or more processors) can render audio and visual content, while the wearable device 400B can determine, in accordance with various aspects of the techniques described in this disclosure, panning head movements based on which the dedicated computing device can render audio content (as a speaker feed) in accordance with the determined panning head movements. As another example, when the wearable device 400B represents smart glasses, the wearable device 400B can include one or more processors that determine panning head movements (by engaging within one or more sensors of the wearable device 400B) and render a speaker feed based on the determined panning head movements.
[0081] As shown, the wearable device 400B includes one or more directional speakers, as well as one or more tracking and / or recording cameras. Further, the wearable device 400B includes one or more inertial, haptics, and / or health sensors, one or more eye tracking cameras, one or more high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of the wearable device 400B can include persistent see-through display technology and hardware.
[0082] The wearable device 400B also includes connectivity hardware, which can represent one or more network interfaces that support multi-mode connectivity (e.g., 4G communication, 5G communication, Bluetooth, etc.). The wearable device 400B also includes one or more ambient light sensors and bone conduction transducers. In some instances, the wearable device 400B can also include one or more passive and / or active cameras with fisheye lenses and / or long lenses. While not shown in FIG. 4B While not shown in the example of FIG. 4B, the wearable device 400B can also include one or more light-emitting diode (LED) lights. In some examples, the LED lights can be referred to as “super-bright” LED lights. In some implementations, the wearable device 400B can also include one or more rear-facing cameras. It will be appreciated that the wearable device 400B can exhibit a variety of different form factors.
[0083] Further, the tracking and recording cameras and other sensors can facilitate determining panning distances. While not shown in the example of FIG. 4B, the wearable device 400B can include other types of sensors for detecting panning distances. FIG. 4B
[0084] While described with respect to particular examples of wearable devices (e.g., the VR device 400B and the smart glasses discussed above with respect to the examples of FIGS. 4A and 4B, respectively), and other devices set forth in the examples of FIGS. 5-7, one of ordinary skill in the art will appreciate that the techniques described herein can be implemented with respect to other devices. FIG. 4B FIG. 1A and FIG. 1B While described with respect to particular examples of wearable devices (e.g., the VR device 400B and the smart glasses discussed above with respect to the examples of FIGS. 4A and 4B, respectively), and other devices set forth in the examples of FIGS. 5-7, one of ordinary skill in the art will appreciate that the techniques described herein can be implemented with respect to other devices. FIG. 1A - FIG. 4B The related descriptions can apply to other examples of wearable devices. For example, other wearable devices (e.g., smart glasses) can include sensors through which translational head movement is obtained. As another example, other wearable devices such as smart watches can include sensors through which translational movement is obtained. Thus, the techniques described in this disclosure should not be limited to a particular type of wearable device, but any wearable device can be configured to perform the techniques described in this disclosure.
[0085] In any event, the audio aspect of VR has been classified into three separate categories of immersion. The first category provides the lowest level of immersion and is referred to as three degrees of freedom (3DOF). 3DOF refers to audio rendering that takes into account head movement in three degrees of freedom (yaw, pitch, and roll), thereby allowing the user to look freely in any direction around them. However, 3DOF does not take into account translational head movement, in which the head is not centered around the optical and acoustic center of the sound field.
[0086] The second category is referred to as 3DOF plus (3DOF+), which provides three degrees of freedom (yaw, pitch, and roll) in addition to limited spatial translational movement (due to the head moving away from the optical and acoustic center within the sound field). 3DOF+ can provide support for perceptual effects such as motion parallax, which can enhance the sense of immersion.
[0087] The third category is referred to as six degrees of freedom (6DOF), which renders audio data in a manner that takes into account three degrees of freedom (yaw, pitch, and roll) of head movement and also takes into account translational movement (x, y, and z translation) of the user in space. Spatial translation can be tracked by sensors that track the user’s position in the physical world or by input controllers.
[0088] 3DOF rendering is the current state of the art for the audio aspect of VR. Thus, the audio aspect of VR is less immersive compared to the video aspect, thereby potentially reducing the overall sense of immersion for the user experience and introducing localization errors (e.g., when the aural playback does not match or is not fully correlated with the visual scene).
[0089] While 3DOF rendering is the current state, more immersive audio rendering (e.g., 3DOF+ and 6DOF rendering) can result in higher complexity in terms of processor cycles consumed, memory and bandwidth consumed, etc. To reduce complexity, audio playback system 16A can include an interpolation device 30 (“INT device 30”) that can select a subset of surround sound coefficients 11’ as surround audio data 15. Then, interpolation device 30 can interpolate the selected subset of surround sound coefficients 11’, apply various weighting (as defined by a measure of importance to the auditory scene, e.g., according to a gain analysis or other analysis (e.g., a directional analysis, etc.)), and then sum the weighted surround sound coefficients 11’ to form surround audio data 15. Interpolation device 30 can select a subset of surround sound coefficients, thereby reducing the number of operations performed when rendering surround audio data 15 (as increasing the number of surround sound coefficients 11’ likewise increases the number of operations performed to render speaker feeds 25 from surround audio data 15).
[0090] Thus, there can be instances in which high complexity audio rendering can be important in providing an immersive experience, and other instances in which low complexity audio rendering can be sufficient to provide the same immersive experience. Moreover, having the capability to provide high complexity audio rendering while also supporting low complexity audio rendering can enable devices with different processing capabilities to perform audio rendering, thereby potentially accelerating adoption of XR devices as lower cost devices (which can have lower processing capabilities as compared to higher host devices) can allow more people to purchase and experience XR.
[0091] According to techniques described in this disclosure, various ways are described by which to implement low complexity audio rendering while providing an option for high complexity audio rendering with additional metadata or other indications for controlling audio rendering at audio playback system 16A. These techniques can enable audio playback system 16A to perform flexible rendering in terms of complexity (as defined by processor cycles, memory, and / or bandwidth consumed) while also allowing for internal and external rendering for an XR experience, as defined by a boundary separating an internal region and an external region. As used herein, a “region” can refer to a two-dimensional space, a three-dimensional space, or a volume. Moreover, audio playback system 16A can utilize metadata or other indications specified in a bitstream representing audio data to configure one or more audio renderers 22 while also referencing a listener position 17 relative to the boundary to generate one or more audio renderers 22 to account for the internal region or the external region.
[0092] Accordingly, these techniques can improve the operation of an audio playback system because, when configured to perform low complexity rendering, the audio playback system 16A can reduce the number of processor cycles, memory, and / or bandwidth consumed. When performing high complexity rendering, the audio playback system 16A can provide a more immersive XR experience, which can result in the user of the audio playback system 16A being more realistically placed in the XR experience.
[0093] As FIG. 1A As shown by the example of FIG. 3, the audio playback system 16A can include a renderer generation unit 32, which represents a unit configured to generate or otherwise obtain one or more of the audio renderers 22 in accordance with various aspects of the techniques described in this disclosure. In some examples, the renderer generation unit 32 can perform the process described above to generate one or more audio renderers 22 based on the listener position 17 and the speaker geometry in the speaker information 13.
[0094] However, in addition, the renderer generation unit 32 can obtain various indications 31 (e.g., syntax elements or other types of metadata) from the bitstream 21 (which can be parsed by the audio decoding device 24). Thus, the soundfield representation generator 302 can specify the indications 31 in the bitstream 21 prior to transmission of the bitstream 21 to the audio playback system 16A. As one example, the soundfield representation generator 302 can receive the indications 31 from the content capture device 300. An operator, editor, or other individual can specify the indications 31 through interaction with the content capture device 300 or some other device such as a content editing device.
[0095] The one or more indications 31 can include an indication indicating a complexity of rendering performed by the audio playback system 16A, an indication of an opacity for rendering a secondary source present in the surround sound coefficients, and / or an indication of a transition distance around an interior region in which rendering is interpolated between interior rendering and exterior rendering. The indication of complexity can indicate the complexity as either low complexity or high complexity (as a Boolean value, where true indicates low complexity and false indicates high complexity). The indication of opacity can either indicate opaque or transparent (as a Boolean value, where true indicates opaque and false indicates transparent, although opacity can be defined as a floating point number with a value of 0 to 1). The indication of transition distance can indicate the distance as a value.
[0096] The soundfield representation generator 302 can also specify in the bitstream 21 a boundary separating the inner region from the outer region. As noted above, the soundfield representation generator 302 can also specify one or more indications 31 for controlling rendering of the surround sound coefficients 11 for the inner region and the outer region. The soundfield representation generator 302 can output the bitstream 21 for delivery (near real-time delivery via network streaming, etc., or for later delivery as described above).
[0097] The audio playback system 16A can obtain the bitstream 21 and invoke an audio decoding device 24 to decompress the bitstream to obtain the surround sound audio coefficients 11’ and to parse the indications 31 from the bitstream 21. The audio decoding device 24 can output the indications 31 to the renderer generation unit 32 along with an indication of the boundary. The audio playback system 16A can also interface with the tracking device 306 to obtain the listener positions 17, with the boundary, the listener positions 17, and the indications 31 being provided to the renderer generation unit 32.
[0098] The renderer generation unit 32 can thus obtain an indication of the boundary separating the inner region and the outer region. The renderer generation unit 32 can also obtain the listener positions 17 indicating the virtual positions of the content consumer device 14A relative to the inner region.
[0099] The renderer generation unit 32 can then obtain, based on the boundary and the listener positions 17, a current renderer of the one or more audio renderers 22 to use when rendering the surround sound audio data 15 to the one or more loudspeaker feeds 25. The current renderer can be configured to render the surround sound audio data 25 for the inner region (and thereby operate as an inner renderer) or configured to render the audio data for the outer region (and thereby operate as an outer renderer).
[0100] Determining whether to configure the current renderer as an inner or outer rendering (or as an interpolation or crossfading therebetween) can depend on where the content consumer device 14A resides relative to the boundary in the XR scene. For example, when the content consumer device 14A is in the XR scene and outside the inner region defined by the boundary for each listener position 17, the renderer generation unit 32 can configure the current renderer to operate as an outer renderer. When the content consumer device 14A is in the XR scene and inside the inner region defined by the boundary for each listener position 17, the renderer generation unit 32 can configure the current renderer to operate as an inner renderer. The renderer generation unit 32 can output the current renderer, with the audio playback system 16A can apply the current renderer to the surround sound audio data 15 to obtain the loudspeaker feeds 25.
[0101] More information on the indication of complexity, the indication of opacity, and the indication of transition distance will be described below with respect to FIG. 2 and FIG. 3 examples.
[0102] FIG. 1B is a block diagram illustrating another example system 100 configured to perform various aspects of the techniques described in this disclosure. The system 100 is similar to the system 10 shown in FIG. 1A except that in the audio playback system 16B of the content consumer device 14B, FIG. 1A one or more audio renderers 22 shown are replaced with a binaural renderer 102 capable of performing binaural rendering using one or more HRTFs or other functionality capable of rendering to left and right speaker feeds 103. Thus, in some examples, the current renderer can be a binaural renderer.
[0103] The audio playback system 16B can output the left and right speaker feeds 103 to headphones 104, which can represent another example of a wearable device and can be coupled to an additional wearable device (e.g., a watch, the VR headset noted above, smart glasses, smart clothing, a smart ring, a smart bracelet, or any other type of smart jewelry (including a smart necklace), etc.) to facilitate reproduction of the sound field. The headphones 104 can be coupled to the additional wearable device wirelessly or via a wired connection.
[0104] Further, the headphones 104 can be coupled to the audio playback system 16 via a wired connection (e.g., a standard 3.5 mm audio jack, a universal system bus (USB) connection, an optical audio jack, or other form of wired connection) or wirelessly (e.g., by way of a Bluetooth TM connection, a wireless network connection, etc.). The headphones 104 can recreate the sound field represented by the surround sound coefficients 11 based on the left and right speaker feeds 103. The headphones 104 can include left and right headphone speakers that are powered by (or, in other words, driven by) the corresponding left and right speaker feeds 103.
[0105] FIG. 2 is a diagram illustrating an example of a low complexity rendering for an extended reality (XR) scene in accordance with various aspects of the techniques described in this disclosure. As FIG. 2 shown in the example of the XR scene 200, the operator 202 is operating a content consumer device 14A (not shown for ease of illustration). The XR scene 200 also includes a boundary 204 separating an interior region 206 and an exterior region 208.
[0106] While in the example of the XR scene 200, FIG. 2A single boundary 204 is shown in the example, but the XR scene 200 can include multiple boundaries separating different interior regions from the exterior region 208. Moreover, while shown as a single boundary 204, a boundary can exist within other boundaries, overlap with other boundaries, and so on. When a boundary exists within other boundaries, the interior region defined by the larger boundary can operate as an exterior boundary (for rendering purposes) with respect to rendering for the interior region defined by the boundary within the outer boundary.
[0107] In any case, first assuming that the operator 202 is in the exterior region 208 relative to the boundary 204, the renderer generation unit 32 (of the content consumer device 14A) can first determine whether the indication of complexity indicates a high complexity or a low complexity. For purposes of illustration, assume that the indication of complexity indicates a low complexity, the renderer generation unit 32 can determine a first distance between the listener position 17 and a center 210 of the interior region 206 (as one example, the first distance is computed based on the boundary 204, which can be represented as a shape, a list of points, a spline, or any other geometric representation). The renderer generation unit 32 can next determine a second distance between the boundary 204 and the center 210.
[0108] The renderer generation unit 32 can then compare the first distance to the second distance to determine that the operator 202 resides outside of the boundary 204. That is, when the first distance is greater than the second distance, the renderer generation unit 32 can determine that the operator 202 is located outside of the boundary 204. For a low complexity configuration, the renderer generation unit 32 can generate a current renderer to render the surround sound audio data 15 for the interior region 206 such that the sound field represented by the surround sound audio data 15 originates from the center 210 of the interior region 206. The renderer generation unit 32 can render the surround sound audio data 15 to be located theta (0) degrees from a direction that the operator 202 is facing.
[0109] Having the sound field appear to originate from a single point (e.g., the center 210 in this example) can reduce complexity in terms of processing cycles, memory, and bandwidth consumption, as it can result in fewer loudspeaker feeds to represent the sound field (and potentially reduce panning, mixing, and other audio operations), while also potentially maintaining an immersive experience. Further reductions in processor cycles, memory, and bandwidth consumption can occur when the renderer generation unit 32 utilizes only a single surround sound coefficient of the surround sound audio data 15 (e.g., a surround sound coefficient corresponding to a spherical basis function with zero order, a surround sound coefficient corresponding to a spherical basis function with zero order represents a gain of the sound field and does not provide much spatial information, thus not requiring complex rendering) instead of processing multiple surround sound coefficients from the surround sound audio data 15.
[0110] Next, assume that the operator 202 moves into the interior region 206. The Tenderer Generation Unit 32 can receive the updated listener position 17 and perform the same process as described above to determine that the operator 202 is located in the interior region 206 (because the first distance is less than the second distance). For the low complexity indication and in response to determining that the operator 202 resides in the interior region, the Tenderer Generation Unit 206 can output an updated current Tenderer configured to render the surround sound audio data 15 such that the sound field represented by the surround sound audio data 15 appears throughout the interior region 206 (this can be referred to as full or normal rendering because all of the surround sound audio data 15 can be rendered such that the audio sources within the sound field are placed precisely around the operator 202).
[0111] In this way, when an interior venue is designated to be rendered using a low complexity Tenderer for low latency applications or artistic purposes, the properties for the buffer region distance or opacity are not utilized in generating the current Tenderer. In this case, when the listener 202 (which is another way of referring to the operator 202) is outside the interior venue region 206 (which is another way of referring to the interior region 206), the W surround sound channels (corresponding to the audio data for the zeroth and sub-order of spherical harmonics, a 00 (t)). When the listener 202 is within the interior venue region 206, the surround sound sound field is normally played back from all directions.
[0112] FIG. 3 is a diagram showing an example of high complexity rendering for an XR scene including a transition distance according to various aspects of the techniques described in this disclosure. The XR scene 220 is similar to the XR scene 200 shown in the example of FIG. 2 the example of FIG. 2, except that its indication of complexity indicates high complexity. In response to the indication indicating high complexity, the Tenderer Generation Unit 32 can utilize an indication of a transition distance 222 that results in a transition zone 224 (which can also be referred to as an “interpolation zone 224”). In some examples, the transition distance 222 can be a configurable threshold, or can be defined as a small value relative to the exterior region 208 or the interior region 206, e.g., 20% of a percentage of the distance to the center of the interior region 206.
[0113] First assume that the operator 202 resides in the exterior region 208, the Tenderer Generation Unit 32 can perform the same process as described above with respect to FIG. 2The example describes determining that the operator 202 resides in the outer region 208. In response to determining that the operator 202 resides in the outer region 208, the Tenderer Generation Unit 32 can next determine whether the indication of complexity indicates high complexity or low complexity. For purposes of illustration, assume that the indication of complexity indicates high complexity, the Tenderer Generation Unit 32 can determine that the indication of opacity indicates whether to indicate opaque or transparent.
[0114] When the indication of opacity indicates opaque, the Tenderer Generation Unit 32 can configure the current Tenderer to discard the secondary audio sources that exist in the soundfield represented by the surround sound audio data 15 that are not directly within the line of sight of the operator 202. In other words, the Tenderer Generation Unit 32 can configure the current Tenderer based on the listener position 17 and the boundary 204 to exclude adding secondary sources to the listener position 17 that indicate a position that is not directly within the line of sight. When the indication of opacity indicates transparent, the Tenderer Generation Unit 32 reverts to normal rendering that considers all secondary sources.
[0115] When the current Tenderer is configured for outer rendering using high complexity, the Tenderer Generation Unit 32 can configure the current Tenderer to render the surround sound audio data 15 such that the soundfield represented by the surround sound audio data 15 expands outward depending on the distance between the listener position 17 and the boundary 204 in all cases (e.g., opaque or transparent). In other words, the Tenderer Generation Unit 32 can configure the current Tenderer to render the surround sound audio data 15 such that the soundfield represented by the surround sound audio data 15 expands outward depending on the distance between the listener position 17 and the boundary 204. FIG. 3 In the example, the outward expansion is represented as theta (0) degrees. The distance is illustrated by the two dashed lines 226A and 226B, resulting in an outward expansion of theta.
[0116] When the operator 202 moves into the transition zone 224, the Tenderer Generation Unit 32 can update the current Tenderer to interpolate or cross-fade between the outer Tenderer and the inner Tenderer in response to determining that the listener position 17 is within the transition distance 222 of the boundary 204. An example of interpolation can be (1-a)*internal_rendering + a*external_rendering, where a is a fraction based on how close the operator 202 is to the soundfield boundary 204 (this is another way of referring to the boundary 204). If the outer Tenderer and the inner Tenderer include different orders of surround sound, for example, if the outer Tenderer renders 1storder surround sound and the inner Tenderer renders 4thorder surround sound, the current Tenderer can interpolate or cross-fade to 2ndorder surround sound and 3rdorder surround sound as the operator 202 moves from the outer region 208 through the transition zone 224 toward the inner region 206. Similarly, the current Tenderer can interpolate or cross-fade to 3rdorder surround sound and 2ndorder surround sound as the operator 202 moves from the inner region 206 through the transition zone 224 back to the outer region 208.
[0117] The audio playback system 16A can then apply the updated current renderer to obtain one or more updated loudspeaker feeds 25. For example, the current renderer can cross-fade between the external renderer and the internal renderer.
[0118] When the operator 202 is fully moved within the internal region 206, the renderer generation unit 32 can generate the current renderer to render normally. That is, the renderer generation unit 32 can generate the current renderer using all of the surround sound coefficients of the surround sound audio data 15 and in a manner that correctly places each of the audio sources in the sound field (e.g., not all sources are positioned at the same location, such as the center 210 of the internal region 206 shown in the example), to fully render the surround sound audio data 15 that resides within the internal region 206. FIG. 2
[0119] In other words, the internal sound field of the internal field region 206 represented in surround sound format allows for secondary sources on the boundary, and according to Huygens’ principle, these secondary sources contribute to the sound at the listener 202. When the opacity property is true, the renderer generation unit 32 can not add the contribution of secondary sources for which the listener 202 does not have a direct line of sight.
[0120] In a high complexity renderer, the listener 202 can hear the internal sound field of the internal field region 206 as an outwardly panned source depending on how far they are from the internal sound field 206. When the listener 202 moves from the external sound field 208 (which is another way of referring to the external region 208) to the internal sound field 206 (which is another way of referring to the internal region 206), the rendering can change and this shift can be done smoothly. The Buffer_Distance property specifies the distance when performing interpolation between the rendering for external to internal listeners 202. One example interpolation scheme includes (1 - a) * internal_rendering + a * external_rendering. This variable can represent a fraction based on how close the listener is to the sound field boundary 204.
[0121] For example, a symphony orchestra can be represented as an internal surround sound sound field. In this case, the listener 202 should have contributions from all the instruments, so the opacity is false. In the case where the internal field represents a crowd and the intent is to change the listening experience as the listener moves outside the boundary 204, the opacity can be set to true.
[0122] Thus, the addition of opacity properties, interpolation buffer distance properties, and complexity properties (which is another way of referring to the indication 31) can be specified to support rendering of interior ambisonic sound fields into the MPEG-I encoder input format. Several use cases can illustrate the usefulness of these properties. These properties can facilitate control of the rendering of interior sound fields at the listener position for 6DoF (and other) use cases.
[0123] While described with respect to VR devices as shown in the examples of FIG. 4A and FIG. 4B , these techniques can be performed by other types of wearable devices, including watches (e.g., so-called “smart watches”), glasses (e.g., so-called “smart glasses”), earphones (including wireless earphones coupled via a wireless connection, or smart earphones coupled via a wired or wireless connection), and any other type of wearable device. Thus, these techniques can be performed by any type of wearable device through which a user can interact with the wearable device while the wearable device is being worn by the user.
[0124] FIG. 5A and FIG. 5B are diagrams showing example systems that can perform various aspects of the techniques described in this disclosure. FIG. 5A An example is shown in which the source device 12B further includes a camera 201. The camera 201 can be configured to capture video data and provide the captured raw video data to the content capture device 300. The content capture device 300 can provide the video data to another component of the source device 12B for further processing into viewport partitioned portions.
[0125] In the example of FIG. 5A , the content consumer device 14C further includes a wearable device 800. It will be understood that, in various implementations, the wearable device 800 can be included in the content consumer device 14C or coupled externally to the content consumer device 14C. As discussed above with respect to FIG. 4A and FIG. 4B , the wearable device 800 includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewports) and for rendering audio data.
[0126] FIG. 5B An example is shown that is similar to the example shown in FIG. 5A , except that one or more of the audio renderers 22 shown in FIG. 5A are replaced with a binaural renderer 102 that is capable of performing binaural rendering using one or more HRTFs or other functionality capable of rendering to left and right speaker feeds 103. The audio playback system 16 can output the left and right speaker feeds 103 to headphones 104.
[0127] Headphones 104 can be coupled to audio playback system 16 via a wired connection (e.g., a standard 3.5 mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other form of wired connection) or wirelessly (e.g., through Bluetooth TM connection, a wireless network connection, etc.). Headphones 104 can recreate the sound field represented by surround sound coefficients 11 based on left and right speaker feeds 103. Headphones 104 can include left and right headphone speakers that are powered (or, in other words, driven) by corresponding left and right speaker feeds 103.
[0128] FIG. 6A is FIG. 1A and FIG. 1B A block diagram of an audio playback system shown in the example of FIG. 1 1 A, in performance of various aspects of the techniques described in this disclosure. Audio playback system 16C can represent an example of audio playback system 16A and / or audio playback system 16B. Audio playback system 16 can include audio decoding device 24 in combination with 6DOF audio Tenderer 22A, which can represent one example of one or more audio renderers 22 shown in the example of FIG. 1 1 B. FIG. 1A
[0129] Audio decoding device 24 can include low-latency decoder 900A, audio decoder 900B, and local audio buffer 902. Low-latency decoder 900A can process XR audio bitstream 21 A to obtain audio stream 901 A, where low-latency decoder 900A can perform relatively low complexity decoding (as compared to audio decoder 900B) to facilitate low-latency reconstruction of audio stream 901 A. Audio decoder 900B can perform relatively higher complexity decoding (as compared to audio decoder 900A) with respect to audio bitstream 21 B to obtain audio stream 901 B. Audio decoder 900B can perform audio decoding in compliance with the MPEG-H 3D Audio Coding standard. Local audio buffer 902 can represent a unit configured to buffer local audio content, which local audio buffer 902 can output as audio stream 903.
[0130] Bitstream 21 (including one or more of XR audio bitstream 21 A and / or audio bitstream 21 B) can also include XR metadata 905 A (which can include the above-identified microphone position information) and 6DOF metadata 905B (which can specify various parameters related to 6DOF audio rendering). 6DOF audio renderer 22A can obtain audio streams 901 A and / or 901 B and XR metadata 905 A, 6DOF metadata 905B, listener position 17, and HRTF 23 from buffers 910 and / or 903, and render speaker feeds 25 and / or 103 based on the listener position and microphone positions. In FIG. 6A In examples of FIG. 6A In examples of
[0131] FIG. 6B is FIG. 1A and FIG. 1B A block diagram of an audio playback system shown in examples of FIG. 6B Example audio playback system 16D of FIG. 6AThe audio playback system 16D is similar to the audio playback system 16C, but the audio playback system 16D also includes an audio object Tenderer 912 and a 3DOF audio Tenderer 914. Each of the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22B can receive the listener position 17 and the HRTF 23. In this example, the output of the audio object Tenderer 912, the 3DOF audio Tenderer 914, or the output of the 6DOF audio Tenderer can be sent to a binauralizer 916, which can perform binaural rendering. In some examples, each of the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22B can output surround sound. The output of the binauralizer 916 can be sent to the interpolation device 30B. The interpolation device 30B can include a controller 918. Although a single output from the audio decoding device 24 is shown, in some examples, the low-latency decoder 900A, the audio decoder 900B, and the local audio buffer 902 can each have separate connections to each of the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22A. In FIG. 6B In the example shown, the interpolation device 30B can interpolate the binauralized audio from the binauralizer 916. In FIG. 6B In the example shown, the interpolation device 30B also includes a controller 918, which can control the functionality of the interpolation device 30B. Although shown as part of the interpolation device 30, in some examples, the controller 918 can be located elsewhere in the audio playback system 16D. In some examples, any of the low-latency decoder 900A, the audio decoder 900B, the local audio buffer 902, the buffer 910, the audio object Tenderer 912, the 3DOF audio Tenderer 914, the 6DOF audio Tenderer 22B, the binauralizer 916, and the interpolation device 30B can be implemented in one or more processors.
[0132] FIG. 6C is FIG. 1A and FIG. 1B A block diagram of an audio playback system shown in the example of FIG. 10 when performing various aspects of the techniques described in this disclosure. FIG. 6C The example audio playback system 16E is similar to the audio playback system 16D of FIG. 6Baudio playback system 16D, however, instead of sending their outputs to binauralizer 916, audio object Tenderer 912, 3DOF audio Tenderer 914, or 6DOF audio Tenderer 22A sends their outputs to interpolation device 30B, which in turn sends outputs to binauralizer 916. In some examples, each of audio object Tenderer 912, 3DOF audio Tenderer 914, and 6DOF audio Tenderer 22B can output surround sound. In FIG. 6C In examples where audio playback system 16E includes interpolation device 30B, interpolation device 30B can interpolate surround sound coefficients from two or more of audio object Tenderer 912, 3DOF audio Tenderer 914, or 6DOF audio Tenderer 22A. In FIG. 6C In examples where audio playback system 16E includes interpolation device 30B, interpolation device 30B can interpolate surround sound coefficients from two or more of audio object Tenderer 912, 3DOF audio Tenderer 914, or 6DOF audio Tenderer 22A. In
[0133] FIG. 6D In examples where audio playback system 16E includes interpolation device 30B, interpolation device 30B can interpolate surround sound coefficients from two or more of audio object Tenderer 912, 3DOF audio Tenderer 914, or 6DOF audio Tenderer 22A. In FIG. 1A and FIG. 1B A block diagram of an audio playback system shown in FIG. 16F as it performs various aspects of the techniques described in this disclosure. FIG. 6D Example audio playback system 16F of FIG. 16F is similar to audio playback system 16E of FIG. 16E FIG. 6C audio playback system 16E of FIG. 16E, however, audio playback system 16F does not include binauralizer 916. In some examples, each of audio object Tenderer 912, 3DOF audio Tenderer 914, and 6DOF audio Tenderer 22B can output surround sound. In FIG. 6D In examples where audio playback system 16E includes interpolation device 30B, interpolation device 30B can interpolate surround sound coefficients from two or more of audio object Tenderer 912, 3DOF audio Tenderer 914, or 6DOF audio Tenderer 22A. In FIG. 6DIn the example shown in FIG. 9, the audio playback system includes a low latency decoder 900A, an audio decoder 900B, a local audio buffer 902, a buffer 910, an audio object Tenderer 912, a 3DOF audio Tenderer 914, a 6DOF audio Tenderer 22B, a binauralizer 916, and an interpolation device 30B. In some examples, the low latency decoder 900A, the audio decoder 900B, the local audio buffer 902, the buffer 910, the audio object Tenderer 912, the 3DOF audio Tenderer 914, the 6DOF audio Tenderer 22B, the binauralizer 916, and the interpolation device 30B can be implemented in one or more processors.
[0134] FIG. 6E In the example shown in FIG. 9, the audio playback system includes a low latency decoder 900A, an audio decoder 900B, a local audio buffer 902, a buffer 910, an audio object Tenderer 912, a 3DOF audio Tenderer 914, a 6DOF audio Tenderer 22B, a binauralizer 916, and an interpolation device 30B. In some examples, the low latency decoder 900A, the audio decoder 900B, the local audio buffer 902, the buffer 910, the audio object Tenderer 912, the 3DOF audio Tenderer 914, the 6DOF audio Tenderer 22B, the binauralizer 916, and the interpolation device 30B can be implemented in one or more processors. FIG. 1A FIG. 1B A block diagram of an audio playback system shown in the example of FIG. 9 when performing various aspects of the techniques described in this disclosure. FIG. 6E An example audio playback system 16G is similar to the audio playback system 16F of FIG. 6D However, the 3DOF audio Tenderer 914 is part of a 6DOF audio Tenderer 22C, rather than a separate device. In some examples, each of the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22C can output surround sound. In some examples, the 3DOF audio Tenderer 914 and the 6DOF audio Tenderer 22C can output surround sound. FIG. 6E In the example shown in FIG. 9, the interpolation device 30B can interpolate surround sound coefficients from two or more of the audio object Tenderer 912, the 3DOF audio Tenderer 914, or the 6DOF audio Tenderer 22B, or interpolate binauralized audio from the audio object Tenderer 912, the 3DOF audio Tenderer 914, and / or the 6DOF audio Tenderer 22B. In some examples, the interpolation device 30B can interpolate surround sound coefficients from two or more of the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22B, or interpolate binauralized audio from the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22B. FIG. 6E In the example shown in FIG. 9, the interpolation device 30B can interpolate surround sound coefficients from two or more of the audio object Tenderer 912, the 3DOF audio Tenderer 914, or the 6DOF audio Tenderer 22B, or interpolate binauralized audio from the audio object Tenderer 912, the 3DOF audio Tenderer 914, and / or the 6DOF audio Tenderer 22B. In some examples, the interpolation device 30B can interpolate surround sound coefficients from two or more of the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22B, or interpolate binauralized audio from the audio object Tenderer 912, the 3DOF audio Tenderer 914, and the 6DOF audio Tenderer 22B.
[0135] FIG. 6F In the example shown in FIG. 9, the audio playback system includes a low latency decoder 900A, an audio decoder 900B, a local audio buffer 902, a buffer 910, an audio object Tenderer 912, a 3DOF audio Tenderer 914, a 6DOF audio Tenderer 22B, a binauralizer 916, and an interpolation device 30B. In some examples, the low latency decoder 900A, the audio decoder 900B, the local audio buffer 902, the buffer 910, the audio object Tenderer 912, the 3DOF audio Tenderer 914, the 6DOF audio Tenderer 22B, the binauralizer 916, and the interpolation device 30B can be implemented in one or more processors. FIG. 1A FIG. 1B A block diagram of an audio playback system shown in the example of FIG. 9 when performing various aspects of the techniques described in this disclosure. FIG. 6F An example audio playback system 16G is similar to the audio playback system 16F of FIG. 6A audio playback system 16C, however, the audio decoder 900C includes an audio object renderer 912, a higher order ambisonics (HOA) renderer 922, and a binauralizer 916, and the interpolation device 30C is a separate device and includes a controller 918 and a 6DOF audio Tenderer 22A. As referred to herein, HOA can include surround sound of order greater than 1. In FIG. 6F In the example of FIG. 9C, the interpolation device 30C can interpolate surround sound coefficients from two or more sources in the buffer 910, or interpolate binauralized audio from the binauralizer 916. In FIG. 6F In the example of FIG. 9C, the interpolation device 30C also includes a controller 918 that can control the functions of the interpolation device 30C. Although shown as part of the interpolation device 30C, in some examples the controller 918 can be located elsewhere in the audio playback system 16H. While several examples of audio playback systems have been set forth in FIG. 6A - FIG. 6F FIG. 9A, other examples including other combinations of the various elements of FIG. 6A - FIG. 6F may fall within the scope of the present disclosure. In some examples, any of the low latency decoder 900A, the audio decoder 900C, the local audio buffer 902, the buffer 910, and the interpolation device 30B can be implemented in one or more processors.
[0136] FIG. 6G is FIG. 1A and FIG. 1B A block diagram of an audio playback system shown in the example of FIG. 9D when performing various aspects of the techniques described in this disclosure. FIG. 6G The example audio playback system 16I of FIG. 9D is similar to the audio playback system 16H of FIG. 6F FIG. 9C, however, the audio decoder 900D includes a FOA / MOA Tenderer 924. The FOA / MOA Tenderer 924 can be a FOA Tenderer and / or a MOA Tenderer. In FIG. 6G In the example of FIG. 9D, the interpolation device 30C can interpolate surround sound coefficients from two or more sources in the buffer 910 (e.g., a 4th order surround sound signal from the HOA Tenderer 922 and a 1st order surround sound signal from the FOA / MOA Tenderer 924), or interpolate binauralized audio from the binauralizer 916. In FIG. 6G In the example of FIG. 9D, the interpolation device 30C also includes a controller 918 that can control the functions of the interpolation device 30C. Although shown as part of the interpolation device 30C, in some examples the controller 918 can be located elsewhere in the audio playback system 16I. While several examples of audio playback systems have been set forth in FIG. 6A - FIG. 6G FIG. 9A, other examples including other combinations of the various elements of FIG. 6A - FIG. 6G or lacking FIG. 6A - FIG. 6GOther examples of various components may fall within the scope of this disclosure. In some examples, any one of the low-latency decoder 900A, audio decoder 900C, local audio buffer 902, buffer 910, and interpolation device 30B may be implemented in one or more processors.
[0137] HOA signals are currently used for playback of ambient sound sources in 6DOF scenes. The origin of the sound source defines the spatial extent of the HOA signal (also referred to as the boundary or range in this paper). When the listener moves outside the spatial extent, the old HOA rendering is no longer effective.
[0138] This disclosure describes the transition between HOA rendering within a spatial range and object-based rendering outside the range. In EIF (N19211, MPEG-I 6DoF Audio Encoder Input Format, online, 2020 (hereinafter referred to as N19211)), HOA sources use high-order surround sound signals with both directionality and location to declare the sound emission source. Most (if not all) HOA renderers do not use the location of the HOA source for rendering. Therefore, it is unclear in EIF how to define a HOA source to render a HOA signal in such a 3DoF manner. Furthermore, it may be desirable to render a HOA signal in 3DoF within the range and in 6DoF outside the range. To achieve this and clarify the old HOA rendering, new attributes can be introduced into the HOA source definition to serve as a flag indicating the complexity of the audio signal to be rendered, such as an indication of whether the HOA signal is rendered in 3DoF or 6DoF.
[0139] As shown in Table 1, the attribute is6DoF can be introduced to indicate the audio playback system (e.g., FIG. 6C The audio playback system 16E switches between 3DoF and 6DoF rendering of the HOA signal. In this example, the default value is "false," meaning that the audio playback system 16E should render the HOA signal using conventional 3DoF HOA rendering (as in an MPEG-H decoder), where only the listener's orientation is used, and position is ignored. For example, the audio playback system 16E can obtain an indication of the current renderer's complexity (e.g., is6DoF) from the bitstream representing the audio data (e.g., bitstream 27). The audio playback system 16E can obtain the current renderer based on the boundary, listener position, and complexity indication. In the example where the complexity indication is false, the audio playback system 16E can render the HOA signal inside the boundary (also referred to herein as the range), and the renderer is free to choose how to render the signal outside the range.
[0140] When is6DoF is set to "true", the audio playback system 16E can render the HOA signal in 6DoF within the extent. If the extent is not defined, the audio playback system 16E can render the HOA signal in 6DoF anywhere in the scene. Setting is6DoF to "true" also enables the group and refDistance attributes, which would otherwise be ignored.
[0141] The extentTransform and transitionDistance attributes depend on the extent of the HOA source. For clarity, the text indicating the dependency is added in the following descriptions of these attributes in Table 1. Table 2 provides a summary of the behavior with different combinations of the is6DoF and extent attributes. Table 3 provides a summary of the behavior with different combinations of the extent and extent attributes. For example, the audio playback system 16A can render the audio signal based on the behavior set forth in Tables 1-3.
[0142] The addition of the table to the HOASource definition in N19211 is shown between <add>With< / add> It should be noted that the references to figures in Table 1 are references to figures in N19211, not to the figures of the present disclosure.
[0143]
[0144]
[0145] Table 1
[0146]
[0147] Table 2
[0148]
[0149] Table 3
[0150]
[0151]
[0152]
[0153] If the HOA signal is music content to be used as background music following the listener, cspace would be set to "user".
[0154] In this case, the orientation of the HOA source is always aligned with the listener.
[0155]
[0156] The following is an example with a 3DOF HOA with a range. In this example, the listener can explore the interior and exterior of a birdhouse. The HOASource includes an HOA signal with bird sounds and a range that is identical to the birdhouse geometry. When the listener is inside the birdhouse, the HOASource will be rendered as legacy 3DoF HOA signals. When the listener travels 2 meters of transitionDistance beyond the range, the rendering will transition from 3DoF to unspecified rendering by the renderer. When the listener is outside the birdhouse (outside the range), the bird sounds are expected to be heard as if originating from inside the birdhouse.
[0157]
[0158] By adding an is6DoF attribute to the HOASource, the HOA content can be explicitly specified to be rendered as either a legacy 3DoF HOA source or a 6DoF source. The provided examples also show how to render 3DoF HOA with and without a range.
[0159] For discussion with respect to the remaining figures, either of the source devices 12A( FIG. 1A - FIG. 1B ) or 12B( FIG. 5A - FIG. 5B ) can be referred to as the source device 12, either of the content consumer devices 14A( FIG. 1A and FIG. 5A ) or 14B( FIG. 1B and FIG. 5B ) can be referred to as the content consumer device 14, and either of the audio playback systems 16A( FIG. 1A and FIG. 5A ), 16B( FIG. 1B and FIG. 5B ), 16C( FIG. 6A ), 16D( FIG. 6B ), 16E( FIG. 6C ), 16F( FIG. 6D ), 16G( FIG. 6E ), or 16H( FIG. 6F ) can be referred to as the audio playback system 16.
[0160] FIG. 7is a diagram illustrating an example of rendering for an extended reality (XR) scene in accordance with various aspects of the techniques described in this disclosure. In this example, the interior region 706 represents a birdhouse. For example, HOA content can be recorded by a person in the birdhouse, e.g., using a microphone on a wearable device (or other source device, e.g., source device 12) within the interior region 706. There can be birds 704A-704C inside the birdhouse that the person hears as if coming from the surroundings. For example, another person (“the listener”) wishes to listen to the content captured by the person in the birdhouse. The listener 702 can receive the HOA stream and have a wearable device with a 3DoF HOA renderer, e.g., audio playback system 16. In this example, the listener 702 hears the 3DoF HOA rendering when the listener 702 is inside the boundary that defines the interior region 706 (also referred to herein as a spatial extent or extent). In this case, the spatial extent can be defined by a sphere, but the spatial extent can be any three-dimensional shape.
[0161] FIG. 8 is a diagram illustrating another example of rendering for an extended reality (XR) scene in accordance with various aspects of the techniques described in this disclosure. If a transition distance 710 is defined, there is an enveloping sphere 712 (or other three-dimensional shape) that encloses the spatial extent or interior region 706. When the listener 702 is within the transition zone defined by the boundary of the interior region 706 (e.g., the spatial extent) and the transition distance 710, the audio playback system (e.g., audio playback system 16) can interpolate or cross-fade between the interior renderer (e.g., 3DOoF audio renderer 914 or 6DoF audio renderer 22A) and the exterior renderer (e.g., object-based audio renderer 912), or can interpolate or cross-fade between different orders of ambisonics (e.g., using HOA renderer 922), as discussed above with respect to FIG. 3 In some examples, the transition distance 710 can be a configurable threshold, or can be defined as a small value relative to the interior region 706 or the exterior region 716, e.g., 20% of a percentage of the distance to the center of the interior region 706.
[0162] FIG. 9is a diagram illustrating another example of rendering for an extended reality (XR) scene in accordance with various aspects of the techniques described in this disclosure. In some examples, for object-based rendering of HOA signals that can occur when the listener 702 is outside the spatial extent (e.g., the interior region 706), there can be two options. In one option, the audio playback system 16 can take the first HOA channel of the HOA signal and use it as the audio for object rendering via the audio object Tenderer 912. In the case that the HOA signal includes a defined position, the audio playback system 16 can use that particular position as the position of the audio object, e.g., the bird at position 714. In the case that the HOA signal does not include a defined position, the audio playback system 16 can compute the geometric center of the spatial extent (e.g., the interior region 706) and use that geometric center as the position of the audio object.
[0163] FIG. 10 is a diagram illustrating another example of rendering for an extended reality (XR) scene in accordance with various aspects of the techniques described in this disclosure. In this example, the audio playback system 16 can render the HOA signal into virtual loudspeakers 720A-720I, which can be located at points around a sphere (or other three-dimensional shape represented by the spatial extent (e.g., the interior region 706)). The audio playback system 16 can use these virtual loudspeakers 720A-720I as the multiple objects that the audio object Tenderer 912 can render. In some examples, the HOA rendering to virtual loudspeakers can be pre-rendered and encoded, e.g., by the source device 12, to facilitate object rendering. For example, if the listener 702 is outside the extent, only the virtual loudspeakers can be transmitted. This can reduce the need for cumbersome HOA rendering and object rendering when the listener is outside the extent.
[0164] FIG. 11 is a flowchart of an example rendering technique in accordance with the present disclosure. The audio playback system 16 can obtain a listener position (720). For example, the audio playback system 16 can receive the listener position 17, e.g., by the tracking device 306. The audio playback system 16 can determine whether the listener position 17 is within the extent (e.g., the interior region 706 (722)). If the listener position 17 is within the extent (“yes” path from block 722), the audio playback system 16 can render the HOA in 3DoF or 6DoF (e.g., by the 3DoF Tenderer 914 or the 6DoF Tenderer 22A (724)). In some examples, whether the audio playback system 16 renders in 3DoF or 6DoF can be based on an indication of complexity of rendering, e.g., a 6DoF flag.
[0165] If the listener position 17 is not inside the range (the “No” path from block 722), the audio playback system 16 can determine whether range transformation is enabled (726). For example, the audio playback system 16 can parse a flag in the bitstream 27 to determine whether range transformation is enabled. If range transformation is not enabled (the “No” path from block 726), the audio playback system 16 can stop rendering (728). For example, the audio playback system 16 can not have obtained any renderers as external renderers at all. If range transformation is enabled (the “Yes” path from block 726), the audio playback system 16 can determine whether the transition distance is greater than zero (730). For example, the audio playback system 16 can determine the value of the transition distance in the bitstream 27. If the transition distance is greater than zero (the “Yes” path from block 730), the audio playback system 16 can optionally cross-fade the internal and external renderers (if the listener position 17 is within the transition distance from the range) (732). Since cross-fading is optional, this block is shown in dashed lines. For example, the audio playback system 16 can render objects and HOA as in block 724, and as in the example of FIG. 9 or FIG. 10 the techniques shown in FIG. 3 , interpolate or cross-fade the rendered signals. In some examples, the audio playback system 16 can render multiple ambisonic orders (e.g., using the HOA renderer 724) and interpolate or cross-fade between different ambisonic orders. For example, the updated current renderer of the audio playback system 16 can cross-fade between different ambisonic orders of the external and internal renderers. In some examples, as the listener position moves from the internal area through the transition area toward the external area, the updated current renderer cross-fades from a higher ambisonic order to a lower ambisonic order. If the transition distance is zero (the “No” path from block 730) (or if the listener position 17 is outside the transition distance from the range), the audio playback system 16 can render objects (734), e.g., using the techniques shown in FIG. 9 or FIG. 10 .
[0166] FIG. 12 An example of a wireless communications system 100 that supports audio streaming, in accordance with aspects of the present disclosure, is shown. The wireless communications system 100 includes base stations 105, UEs 115, and a core network 130. In some examples, the wireless communications system 100 can be a Long Term Evolution (LTE) network, an LTE-Advanced (LTE-A) network, an LTE-A Pro network, or a New Radio (NR) network. In some cases, wireless communications system 100 can support enhanced broadband communications, ultra-reliable (e.g., mission critical) communications, low latency communications, or communications with low-cost and low-complexity devices.
[0167] The base stations 105 can wirelessly communicate with the UEs 115 via one or more base station antennas. The base stations 105 described herein can include or can be referred to by those skilled in the art as a base transceiver station, a radio base station, an access point, a radio transceiver, a NodeB, an eNodeB (eNB), a next-generation NodeB or giga-NodeB (either of which can be referred to as a gNB), a Home NodeB, or a Home eNodeB. The wireless communications system 100 can include base stations 105 of different types (e.g., macro or small cell base stations). The UEs 115 described herein can be able to communicate with various types of base stations 105 and network equipment including macro eNBs, small cell eNBs, gNBs, relay base stations, and the like.
[0168] Each base station 105 can be associated with a particular geographic coverage area 110 in which communication with various UEs 115 is supported. Each base station 105 can provide communication coverage for a respective geographic coverage area 110 via communication links 125, and communication links 125 between a base station 105 and a UE 115 can utilize one or more carriers. Communication links 125 shown in wireless communications system 100 can include uplink transmissions from a UE 115 to a base station 105, or downlink transmissions from a base station 105 to a UE 115. Downlink transmissions can also be called forward link transmissions while uplink transmissions can also be called reverse link transmissions.
[0169] The geographic coverage area 110 for a base station 105 can be divided into sectors making up a portion of the geographic coverage area 110, and each sector can be associated with a small cell or other type of cell or various combinations of cells. For example, each base station 105 can provide communication coverage for a macro cell, a small cell, a hot spot, or other types of cells or various combinations of cells. In some examples, base stations 105 can be movable and therefore provide communication coverage for a moving geographic coverage area 110. In some examples, different geographic coverage areas 110 associated with different technologies can overlap, and overlapping geographic coverage areas 110 associated with different technologies can be supported by the same base station 105 or by different base stations 105. The wireless communications system 100 can include, for example, a heterogeneous LTE / LTE-A / LTE-A Pro or NR network in which different types of base stations 105 provide coverage for various geographic coverage areas 110.
[0170] The UEs 115 can be dispersed throughout the wireless communication system 100, and each UE 115 can be stationary or mobile. A UE 115 can also be referred to as a mobile device, a wireless device, a remote device, a handheld device, or a subscriber device, or some other suitable terminology, where the“device” can also be referred to as a unit, a station, a terminal, or a client. A UE 115 can also be a personal electronic device such as a cellular phone, a personal digital assistant (PDA), a tablet computer, a laptop computer, or a personal computer. In examples of the present disclosure, a UE 115 can be any one of the audio sources described in the present disclosure, including a VR headset, an XR headset, an AR headset, a vehicle, a smartphone, a microphone, a microphone array, or any other device that includes a microphone or that is capable of transmitting a captured and / or synthesized audio stream. In some examples, the synthesized audio stream can be an audio stream stored in memory or previously created or synthesized. In some examples, a UE 115 can also refer to a wireless local loop (WLL) station, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, or an MTC device, among others, which can be implemented in various articles such as appliances, vehicles, meters, or the like.
[0171] Some UEs 115, such as MTC or IoT devices, can be low cost or low complexity devices, and can provide for automated communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC can refer to data communication technologies that allow devices to communicate with one another or a base station 105 without human intervention. In some examples, M2M communication or MTC can include communications from devices that exchange and / or utilize audio metadata and / or privacy data based on cryptography to indicate privacy restrictions to switch, mask, and / or null various audio streams and / or audio sources, as will be described in more detail below.
[0172] In some cases, a UE 115 can also be able to communicate directly with other UEs 115 (e.g., using a peer-to-peer (P2P) or device-to-device (D2D) protocol). One or more of a group of UEs 115 utilizing D2D communications can be within the geographic coverage area 110 of a base station 105. Other UEs 115 in such a group can be outside the geographic coverage area 110 of a base station 105, or be otherwise unable to receive transmissions from a base station 105. In some cases, groups of UEs 115 communicating via D2D communications can utilize a one-to-many (1:M) system in which each UE 115 transmits to every other UE 115 in the group. In some cases, a base station 105 facilitates the scheduling of resources for D2D communications. In other cases, D2D communications are carried out between UEs 115 without the involvement of a base station 105.
[0173] The base stations 105 can communicate with the core network 130 and with one another. For example, the base stations 105 can interface with the core network 130 through backhaul links 132 (e.g., via an SI, N2, N3, or other interface). The base stations 105 can communicate with one another over backhaul links 134 (e.g., via an X2, Xn, or other interface) either directly (e.g., direct communication between base stations 105) or indirectly (e.g., via the core network 130).
[0174] In some cases, the wireless communications system 100 can utilize both licensed and unlicensed radio frequency spectrum bands. For example, the wireless communications system 100 can employ License Assisted Access (LAA), LTE-Unlicensed (LTE-U) radio access technology, or NR technology in an unlicensed
[0175] FIG. 13 is a flowchart illustrating example operations of a source device in performing various aspects of the techniques described in this disclosure. The source device 12 can obtain a bitstream 21 representing scene-based audio data 11 in the manner described above (800). The soundfield representation generator 302 of the source device 12 can specify a boundary separating an interior region and an exterior region in the bitstream 21 (802). FIG. 1A
[0176] As noted above, the soundfield representation generator 302 can also specify one or more indications 31 that control rendering of ambisonic coefficients 11 for the interior region and the exterior region (804). The soundfield representation generator 302 can output the bitstream 21 for delivery (near real-time delivery via network streaming, etc., or for later delivery as noted above) (806).
[0177] FIG. 14 is a flowchart illustrating example operations of a source device in performing various aspects of the techniques described in this disclosure. The source device 12 can obtain a bitstream 21 representing scene-based audio data 11 in the manner described above (800). The soundfield representation generator 302 of the source device 12 can specify a boundary separating an interior region and an exterior region in the bitstream 21 (802). FIG. 1A A flow diagram illustrating example operations of a content consumer device in performing various aspects of the techniques described in this disclosure. The audio playback system 16 can obtain the bitstream 21 and invoke the audio decoding device 24 to decompress the bitstream to obtain the surround sound audio coefficients 11' and the indications 31 from the bitstream 21. The audio decoding device 24 can output the indications 31 to the renderer generation unit 32 along with an indication of the boundary. The audio playback system 16A can also interface with the tracking device 306 to obtain the listener position 17, where the boundary, the listener position 17, and the indications 31 are provided to the renderer generation unit 32.
[0178] Thus, the renderer generation unit 32 can obtain an indication of a boundary separating an interior region and an exterior region (1000). The renderer generation unit 32 can also obtain a listener position 17 indicating a virtual position of the content consumer device 14 relative to the interior region (1002).
[0179] The renderer generation unit 32 can then obtain a current renderer to use when rendering the surround sound audio data 15 to one or more speaker feeds 25 based on the boundary and the listener position 17. The current renderer can be configured to render the surround sound audio data 25 for the interior region (and thereby operate as an interior renderer) or configured to render the audio data for the exterior region (and thereby operate as an exterior renderer) (1004). The renderer generation unit 32 can output the current renderer, where the audio playback system 16 can apply the current renderer to the surround sound audio data 15 to obtain the speaker feeds 25 (1006).
[0180] FIG. 15 is a flow diagram illustrating example audio processing techniques in accordance with various aspects of the disclosure. The audio playback system 16 can determine whether a boundary separating an interior region and an exterior region exists (1100). For example, the audio playback system 16 can obtain an indication of a boundary separating an interior region and an exterior region from the bitstream 21. Based on determining that the boundary exists, the audio playback system 16 can determine a transition distance value indicating a size of a transition zone (1102). For example, the audio playback system 16 can obtain an indication of the transition distance value from the bitstream 21.
[0181] The Tenderer Generation Unit 32 can obtain a listener position, the listener position indicating a virtual position of a device (e.g., the content consumer device 14) relative to the interior region (1104). For example, the audio playback system 16 can interface with the tracking device 306 to obtain the listener position 17 and provide the listener position 17 to the Tenderer Generation Unit 32. The Tenderer Generation Unit 32 can obtain a current Tenderer based at least in part on the boundary and the listener position (1106). For example, the Tenderer Generation Unit can generate the current Tenderer based at least in part on the boundary and the listener position. The one or more audio renderers 22 can apply the current Tenderer to the audio data to obtain one or more speaker feeds 25 (1108). For example, the one or more audio renderers 22 can apply the current Tenderer to the surround sound audio data 15 to obtain the one or more speaker feeds 25. The content consumer device 14 can store the one or more speaker feeds 25 (1110). For example, the content consumer device 14 can store the one or more speaker feeds 25 in memory.
[0182] In some examples, the transition distance value is 0, the current Tenderer includes either an interior Tenderer configured to render the audio data for the interior region or an exterior Tenderer configured to render the audio data for the exterior region, and wherein the obtaining of the current Tenderer is further based on the transition distance value being 0.
[0183] In some examples, the transition distance value is greater than 0, the current Tenderer includes either an interior Tenderer configured to render the audio data for the interior region or an exterior Tenderer configured to render the audio data for the exterior region, or both the interior Tenderer and the exterior Tenderer, and wherein the obtaining of the current Tenderer is further based on the transition distance being greater than 0.
[0184] In this regard, various aspects of the techniques described in this disclosure can implement the following clauses.
[0185] Clause 1A. A device configured to process one or more audio streams, the device comprising: one or more processors configured to: obtain an indication of a boundary separating an interior region and an exterior region; obtain a listener position, the listener position indicating a position of the device relative to the interior region; obtain a current Tenderer based on the boundary and the listener position, the current Tenderer either as an interior Tenderer configured to render audio data for the interior region or as an exterior Tenderer configured to render audio data for the exterior region; apply the current Tenderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0186] Clause 2A. The device of clause 1A, wherein the one or more processors are configured to: determine a first distance between the listener position and a center of the interior region; determine a second distance between the boundary and the center of the interior region; and obtain the current renderer based on the first distance and the second distance.
[0187] Clause 3A. The device of any combination of clauses 1A and 2A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the exterior renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data originates from a center of the interior region.
[0188] Clause 4A. The device of any combination of clauses 1A and 2A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the interior renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0189] Clause 5A. The device of any combination of clauses 1A and 2A, wherein the audio data includes surround sound audio data representing a primary audio source and a secondary audio source, wherein the one or more processors are further configured to: obtain an indication of an opacity of the secondary audio source, and wherein the one or more processors are configured to: obtain the current renderer based on the listener position, the boundary, and the indication.
[0190] Clause 6A. The device of clause 5A, wherein the one or more processors are configured to: obtain the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0191] Clause 7A. The device of any combination of clauses 5A and 6A, wherein the one or more processors are further configured to: obtain the current renderer when the indication of the opacity is enabled and based on the listener position and the boundary, the current renderer excluding adding the secondary source to a location for which the listener position indicates is not directly in line of sight.
[0192] Clause 8A. The device of any combination of clauses 5A-7A, wherein the exterior renderer is configured to: render the audio data such that a sound field represented by the audio data expands outward depending on a distance between the listener position and the boundary.
[0193] Clause 9A. The device of any combination of clauses 5A-8A, wherein the one or more processors are further configured to: update the current renderer to interpolate between the exterior renderer and the interior renderer in response to determining that the listener position is within a buffer distance from the boundary, to obtain an updated current renderer; and apply the current renderer to the audio data to obtain one or more updated loudspeaker feeds.
[0194] Clause 10A. The device of clause 9A, wherein the one or more processors are further configured to obtain an indication of the buffer distance from a bitstream representing the audio data.
[0195] Clause 11A. The device of any combination of clauses 1A-10A, wherein the one or more processors are further configured to obtain an indication of a complexity of the current renderer from a bitstream representing the audio data, and wherein the one or more processors are configured to obtain the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0196] Clause 12A. The device of clause 11A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the one or more processors are configured to obtain an outer renderer when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, such that the outer renderer is configured to render the surround sound audio data such that a soundfield represented by the surround sound audio data originates from a center of the inner region.
[0197] Clause 13A. The device of clause 11A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the one or more processors are configured to obtain an outer renderer when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, such that the outer renderer is configured to render the audio data such that a soundfield represented by the audio data expands outward depending on a distance between the listener position and the boundary.
[0198] Clause 14A. A method for processing one or more audio streams, the method comprising: obtaining, by one or more processors, an indication of a boundary separating an inner region and an outer region; obtaining, by the one or more processors, a listener position indicating a position of a device relative to the inner region; obtaining, by the one or more processors, a current renderer based on the boundary and the listener position, the current renderer being either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; and applying, by the one or more processors, the current renderer to the audio data to obtain one or more loudspeaker feeds.
[0199] Clause 15A. The method of clause 14A, wherein obtaining the current renderer comprises: determining a first distance between the listener position and a center of the inner region; determining a second distance between the boundary and the center of the inner region; and obtaining the current renderer based on the first distance and the second distance.
[0200] Clause 16A. The method of any combination of clauses 14A and 15A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the outer renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data originates from a center of the interior region.
[0201] Clause 17A. The method of any combination of clauses 14A and 15A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the inner renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0202] Clause 18A. The method of any combination of clauses 14A and 15A, wherein the audio data includes surround sound audio data representing a primary audio source and a secondary audio source, wherein the method further comprises: obtaining an indication of an opacity of the secondary audio source, and wherein obtaining the current renderer comprises obtaining the current renderer based on the listener position, the boundary, and the indication.
[0203] Clause 19A. The method of clause 18A, wherein obtaining the indication of the opacity comprises obtaining the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0204] Clause 20A. The method of any combination of clauses 18A and 19A, further comprising: when the indication of the opacity is enabled, and based on the listener position and the boundary, obtaining the current renderer that excludes adding the secondary source to a location indicated by the listener position that is not directly in line of sight.
[0205] Clause 21A. The method of any combination of clauses 18A-20A, wherein the outer renderer is configured to: render the audio data such that a sound field represented by the audio data expands outward depending on a distance between the listener position and the boundary.
[0206] Clause 22A. The method of any combination of clauses 18A-21A, further comprising: in response to determining that the listener position is within a buffer distance from the boundary, updating the current renderer to interpolate between the outer renderer and the inner renderer in order to obtain an updated current renderer; and applying the current renderer to the audio data to obtain one or more updated loudspeaker feeds.
[0207] Clause 23A. The method of clause 22A, further comprising: obtaining an indication of the buffer distance from a bitstream representing the audio data.
[0208] Clause 24A. The method of any combination of clauses 14A-23A, further comprising: obtaining, from a bitstream representing the audio data, an indication of a complexity of the current renderer, and wherein obtaining the current renderer comprises: obtaining the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0209] Clause 25A. The method of clause 24A, wherein the audio data comprises surround sound audio data associated with spherical basis functions having a zeroth order, and wherein obtaining the current renderer comprises: obtaining, when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, an outside renderer, such that the outside renderer is configured to render the surround sound audio data so that a soundfield represented by the surround sound audio data originates from a center of the interior region.
[0210] Clause 26A. The method of clause 24A, wherein the audio data comprises surround sound audio data associated with spherical basis functions having a zeroth order, and wherein obtaining the current renderer comprises: obtaining, when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, an outside renderer, such that the outside renderer is configured to render the audio data so that a soundfield represented by the audio data expands outward depending on a distance between the listener position and the boundary.
[0211] Clause 27A. A device configured to process one or more audio streams, the device comprising: means for obtaining an indication of a boundary separating an interior region and an exterior region; means for obtaining a listener position indicating a position of the device relative to the interior region; means for obtaining, based on the boundary and the listener position, a current renderer that is either an interior renderer configured to render audio data for the interior region or an exterior renderer configured to render audio data for the exterior region; and means for applying the current renderer to the audio data to obtain one or more loudspeaker feeds.
[0212] Clause 28A. The device of clause 27A, wherein the means for obtaining the current renderer comprises: means for determining a first distance between the listener position and a center of the interior region; means for determining a second distance between the boundary and the center of the interior region; and means for obtaining the current renderer based on the first distance and the second distance.
[0213] Clause 29A. The device of any combination of clauses 27A and 28A, wherein the audio data comprises surround sound audio data associated with spherical basis functions having a zeroth order, and wherein the exterior renderer is configured to render the surround sound audio data so that a soundfield represented by the surround sound audio data originates from a center of the interior region.
[0214] Clause 30A. The device of any combination of clauses 27A and 28A, wherein the audio data comprises surround sound audio data associated with a spherical basis function having a zeroth order, and wherein the interior Tenderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0215] Clause 31A. The device of any combination of clauses 27A and 28A, wherein the audio data comprises surround sound audio data representing a primary audio source and a secondary audio source, wherein the device further comprises: means for obtaining an indication of an opacity of the secondary audio source, and wherein the means for obtaining the current Tenderer comprises: means for obtaining the current Tenderer based on the listener position, the boundaries, and the indication.
[0216] Clause 32A. The device of clause 31A, wherein the means for obtaining the indication of the opacity comprises: means for obtaining the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0217] Clause 33A. The device of any combination of clauses 31A and 32A, further comprising: means for obtaining the current Tenderer based on the listener position and the boundaries when the indication of the opacity is enabled, the current Tenderer excluding adding the secondary source to a location indicated by the listener position as not being directly in line of sight.
[0218] Clause 34A. The device of any combination of clauses 31A-33A, wherein the exterior Tenderer is configured to: render the audio data such that a sound field represented by the audio data unfolds outwardly depending on a distance between the listener position and the boundaries.
[0219] Clause 35A. The device of any combination of clauses 31A-34A, further comprising: means for updating the current Tenderer to interpolate between the exterior Tenderer and the interior Tenderer in response to determining that the listener position is within a buffer distance from the boundaries, to obtain an updated current Tenderer; and means for applying the current Tenderer to the audio data to obtain one or more updated loudspeaker feeds.
[0220] Clause 36A. The device of clause 35A, further comprising: means for obtaining an indication of the buffer distance from a bitstream representing the audio data.
[0221] Clause 37A. The device of any combination of clauses 27A-36A, further comprising: means for obtaining an indication of a complexity of the current Tenderer from a bitstream representing the audio data, and wherein the means for obtaining the current Tenderer comprises means for obtaining the current Tenderer based on the boundaries, the listener position, and the indication of the complexity.
[0222] Clause 38A. The device of clause 37A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zero order, and wherein the means for obtaining a current renderer includes: means for obtaining an outer renderer when the listener position is outside the boundary and when the indication of complexity indicates a low complexity, such that the outer renderer is configured to render the surround sound audio data, such that a sound field represented by the surround sound audio data emanates from a center of the inner region.
[0223] Clause 39A. The device of clause 37A, wherein the audio data includes surround sound audio data associated with spherical basis functions having a zero order, and wherein the means for obtaining a current renderer includes: means for obtaining an outer renderer when the listener position is outside the boundary and when the indication of complexity indicates a low complexity, such that the outer renderer is configured to render the audio data, such that a sound field represented by the audio data expands outward depending on a distance between the listener position and the boundary.
[0224] Clause 40A. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: obtain an indication of a boundary separating an inner region and an outer region; obtain a listener position indicating a position of a device relative to the inner region; obtain, based on the boundary and the listener position, a current renderer that is either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; and apply the current renderer to the audio data to obtain one or more loudspeaker feeds.
[0225] Clause 1B. A device configured to generate a bitstream representing audio data, the device comprising: a memory configured to store the audio data; and one or more processors coupled to the memory and configured to: obtain, based on the audio data, the bitstream representing the audio data; specify, in the bitstream, a boundary separating an inner region and an outer region; specify, in the bitstream, one or more indications that control rendering of the audio data for the inner region or the outer region; and output the bitstream.
[0226] Clause 2B. The device of clause 1B, wherein the one or more indications include an indication indicating a complexity of the rendering.
[0227] Clause 3B. The device of clause 2B, wherein the indication of the complexity indicates either a low complexity or a high complexity.
[0228] Clause 4B. The device of any combination of clauses 1B-3B, wherein the one or more indications include an indication indicating an opacity for rendering a secondary source present in the audio data.
[0229] Clause 5B. The device of clause 4B, wherein the indication of opacity indicates that the opacity is either opaque or transparent.
[0230] Clause 6B. The device of any combination of clauses 1B-5B, wherein the one or more indications include an indication of a buffer distance around the interior region in which rendering is interpolated between interior rendering and exterior rendering.
[0231] Clause 7B. The device of any combination of clauses 1B-6B, wherein the audio data comprises surround sound audio data.
[0232] Clause 8B. A method for generating a bitstream representing audio data, the method comprising: obtaining, based on the audio data, a bitstream representing the audio data; specifying, in the bitstream, a boundary separating an interior region and an exterior region; specifying, in the bitstream, one or more indications that control rendering of audio data for the interior region or the exterior region; and outputting the bitstream.
[0233] Clause 9B. The method of clause 8B, wherein the one or more indications include an indication of a complexity of rendering.
[0234] Clause 10B. The method of clause 9B, wherein the indication of complexity indicates either a low complexity or a high complexity.
[0235] Clause 11B. The method of any combination of clauses 8B-10B, wherein the one or more indications include an indication of an opacity for rendering of a secondary source present in the audio data.
[0236] Clause 12B. The method of clause 11B, wherein the indication of opacity indicates that the opacity is either opaque or transparent.
[0237] Clause 13B. The device of any combination of clauses 8B-12B, wherein the one or more indications include an indication of a buffer distance around the interior region in which rendering is interpolated between interior rendering and exterior rendering.
[0238] Clause 14B. The method of any combination of clauses 8B-13B, wherein the audio data comprises surround sound audio data.
[0239] Clause 15B. A device configured to generate a bitstream representing audio data, the device comprising: means for obtaining, based on the audio data, the bitstream representing the audio data; means for specifying, in the bitstream, a boundary separating an inner region and an outer region; means for specifying, in the bitstream, one or more indications that control rendering of the audio data for the inner region or the outer region; and means for outputting the bitstream.
[0240] Clause 16B. The device of clause 15B, wherein the one or more indications include an indication that indicates a complexity of the rendering.
[0241] Clause 17B. The device of clause 16B, wherein the indication of the complexity indicates either a low complexity or a high complexity.
[0242] Clause 18B. The device of any combination of clauses 15B-17B, wherein the one or more indications include an indication that indicates an opacity for rendering a secondary source present in the audio data.
[0243] Clause 19B. The device of clause 18B, wherein the indication of the opacity indicates that the opacity is either opaque or transparent.
[0244] Clause 20B. The device of any combination of clauses 15B-19B, wherein the one or more indications include an indication that indicates a buffer distance around the inner region in which rendering is interpolated between the inner rendering and the outer rendering.
[0245] Clause 21B. The device of any combination of clauses 15B-20B, wherein the audio data includes surround sound audio data.
[0246] Clause 22B. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: obtain, based on audio data, a bitstream representing the audio data; specify, in the bitstream, a boundary separating an inner region and an outer region; specify, in the bitstream, one or more indications that control rendering of the audio data for the inner region or the outer region; and output the bitstream.
[0247] Clause 1C. A device configured to process one or more audio streams, the device comprising: one or more processors configured to: determine whether a boundary separating an interior region and an exterior region is present; determine, based on the boundary being present, a transition distance value that indicates a size of a transition zone, wherein the transition distance value is 0; obtain a listener position that indicates a position of the device relative to the interior region; obtain, based on the boundary, the listener position, and the transition distance value being 0, a current renderer that is either an interior renderer configured to render audio data for the interior region or an exterior renderer configured to render audio data for the exterior region; apply the current renderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0248] Clause 2C. The device of clause 1C, wherein the one or more processors are configured to: determine a first distance between the listener position and a center of the interior region; determine a second distance between the boundary and the center of the interior region; and obtain, based on the first distance and the second distance, the current renderer.
[0249] Clause 3C. The device of clause 1C or clause 2C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the exterior renderer is configured to render an audio object that includes only a first channel of the surround sound audio data.
[0250] Clause 4C. The device of clause 1C or clause 2C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the exterior renderer is configured to render the HOA to a plurality of virtual loudspeakers.
[0251] Clause 5C. The device of clauses 1C-4C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the interior renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0252] Clause 6C. The device of any of clauses 1C-5C, wherein the audio data comprises surround sound audio data representing a primary audio source and a secondary audio source, wherein the one or more processors are further configured to: obtain an indication of an opacity of the secondary audio source, and wherein the one or more processors are configured to: obtain, based on the listener position, the boundary, and the indication, the current renderer.
[0253] Clause 7C. The device of clause 6C, wherein the one or more processors are configured to: obtain the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0254] Clause 8C. The device of clause 6C or clause 7C, wherein the one or more processors are further configured to: obtain, when the indication of opacity is enabled, and based on the listener position and the boundary, a current renderer that excludes adding the secondary source to a position for which the listener position indicates is not directly in line of sight.
[0255] Clause 9C. The device of clause 6C, wherein the external renderer is configured to: render the audio data such that the sound field represented by the audio data is panned out depending on a distance between the listener position and the boundary.
[0256] Clause 10C. The device of any of clauses 1C-9C, wherein the one or more processors are further configured to: obtain, from the bitstream representing the audio data, an indication of a transition distance.
[0257] Clause 11C. The device of any of clauses 1C-10C, wherein the one or more processors are further configured to: obtain, from the bitstream representing the audio data, an indication of a complexity of the current renderer; and wherein the one or more processors are configured to: obtain the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0258] Clause 12C. The device of clause 11C, wherein the indication of the complexity comprises a 6DOF flag.
[0259] Clause 13C. The device of clause 12C, wherein the 6DOF flag is false.
[0260] Clause 14C. The device of clause 13C, wherein the one or more processors are configured to: obtain, based at least in part on the 6DOF flag being false, the internal renderer as a 3DOF renderer.
[0261] Clause 15C. The device of clause 14C, wherein the one or more processors are further configured to: determine a range transform that indicates whether the device renders audio sources outside the boundary.
[0262] Clause 16C. The device of clause 15C, wherein the range transform is true and the listener position is outside the boundary, wherein the one or more processors are configured to: obtain, based on the range transform being true and the listener position being outside the boundary, the external renderer as the current renderer.
[0263] Clause 17C. The device of clause 15C, wherein the range transform is false and the listener position is outside the boundary, wherein the one or more processors are configured to: obtain, based on the range transform being false and the listener position being outside the boundary, the external renderer as the current renderer, wherein the external renderer is a no-renderer.
[0264] Clause 18C. The device of clause 12C, wherein the 6DOF flag is true.
[0265] Clause 19C. The device of clause 18C, wherein the one or more processors are configured to obtain the internal Tenderer as the 6DOF Tenderer based at least in part on the 6DOF flag being true.
[0266] Clause 20C. The device of clause 19C, wherein the one or more processors are further configured to determine a range transform, the range transform indicating whether the device renders audio sources outside the boundary.
[0267] Clause 21C. The device of clause 20C, wherein the range transform is true and the listener position is outside the boundary, wherein the one or more processors are configured to obtain the external Tenderer as the current Tenderer based on the range transform being true and the listener position being outside the boundary.
[0268] Clause 22C. The device of clause 20C, wherein the range transform is false and the listener position is outside the boundary, wherein the one or more processors are configured to obtain the external Tenderer as the current Tenderer based on the range transform being false and the listener position being outside the boundary, wherein the external Tenderer is a no Tenderer.
[0269] Clause 23C. A device configured to process one or more audio streams, the device comprising: one or more processors configured to: determine whether a boundary separating an inner region and an outer region exists; determine a transition distance value based on the boundary existing, the transition distance value indicating a size of a transition zone, wherein the transition distance value is greater than 0; obtain a listener position, the listener position indicating a position of the device relative to the inner region; obtain a current Tenderer based on the boundary, the listener position, and the transition distance value being greater than 0, the current Tenderer either as an internal Tenderer configured to render audio data for the inner region, or as an external Tenderer configured to render audio data for the outer region, or as both the internal Tenderer and the external Tenderer; apply the current Tenderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0270] Clause 24C. The device of clause 23C, wherein the one or more processors are configured to: determine a first distance between the listener position and a center of the inner region; determine a second distance between the boundary and the center of the inner region; and obtain the current Tenderer based on the first distance and the second distance.
[0271] Clause 25C. The device of clause 23C or clause 24C, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the external Tenderer is configured to render audio objects that include only a first channel of the surround sound audio data.
[0272] Clause 26C. The device of clause 23C or clause 24C, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the external Tenderer is configured to render the HOA to a plurality of virtual loudspeakers.
[0273] Clause 27C. The device of any of clauses 23C-26C, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the internal Tenderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0274] Clause 28C. The device of any of clauses 23C-27C, wherein the audio data includes surround sound audio data representing a primary audio source and a secondary audio source, wherein the one or more processors are further configured to obtain an indication of an opacity of the secondary audio source, and wherein the one or more processors are configured to obtain the current Tenderer based on the listener position, the boundaries, and the indication.
[0275] Clause 29C. The device of clause 28C, wherein the one or more processors are configured to obtain the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0276] Clause 30C. The device of clause 28C or clause 29C, wherein the one or more processors are further configured to obtain the current Tenderer based on the listener position and the boundaries when the indication of the opacity is enabled, the current Tenderer excluding adding the secondary source to a position for which the listener position indicates is not directly in line of sight.
[0277] Clause 31C. The device of clause 28C, wherein the external Tenderer is configured to render the audio data such that a sound field represented by the audio data expands outward depending on a distance between the listener position and the boundaries.
[0278] Clause 32C. The device of any of clauses 23C-31C, wherein the one or more processors are further configured to obtain an indication of a transition distance from a bitstream representing the audio data.
[0279] Clause 33C. The device of clause 32C, wherein the one or more processors are further configured to: update the current renderer to interpolate between the outer renderer and the inner renderer to obtain an updated current renderer in response to determining that the listener position is within a transition distance from the boundary; and apply the current renderer to the audio data to obtain the one or more updated speaker feeds.
[0280] Clause 34C. The device of any of clauses 23C-33C, wherein the one or more processors are further configured to: obtain an indication of a complexity of the current renderer from a bitstream representing the audio data; and wherein the one or more processors are configured to obtain the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0281] Clause 35C. The device of clause 34C, wherein the indication of the complexity comprises a 6DOF flag.
[0282] Clause 36C. The device of clause 35C, wherein the 6DOF flag is false.
[0283] Clause 37C. The device of clause 36C, wherein the one or more processors are configured to obtain the inner renderer as a 3DOF renderer based at least in part on the 6DOF flag being false.
[0284] Clause 38C. The device of clause 37C, wherein the one or more processors are further configured to determine a range transform that indicates whether the device renders audio sources outside the boundary.
[0285] Clause 39C. The device of clause 38C, wherein the range transform is true and the listener position is outside the boundary, wherein the one or more processors are configured to obtain the outer renderer as the current renderer based on the range transform being true and the listener position being outside the boundary.
[0286] Clause 40C. The device of clause 38C, wherein the range transform is false and the listener position is outside the boundary, wherein the one or more processors are configured to obtain the outer renderer as the current renderer based on the range transform being false and the listener position being outside the boundary, wherein the outer renderer is a no-renderer.
[0287] Clause 41C. The device of clause 35C, wherein the 6DOF flag is true.
[0288] Clause 42C. The device of clause 41C, wherein the one or more processors are configured to obtain the inner renderer as a 6DOF renderer based at least in part on the 6DOF flag being true.
[0289] Clause 43C. The device of clause 42C, wherein the one or more processors are further configured to: determine a range transform, the range transform indicating whether the device renders audio sources outside the boundary.
[0290] Clause 44C. The device of clause 43C, wherein the range transform is true and the listener position is outside the boundary, wherein the one or more processors are configured to: obtain the outside renderer as the current renderer based on the range transform being true and the listener position being outside the boundary.
[0291] Clause 45C. The device of clause 43C, wherein the range transform is false and the listener position is outside the boundary, wherein the one or more processors are configured to: obtain the outside renderer as the current renderer based on the range transform being false and the listener position being outside the boundary, wherein the outside renderer is a no- renderer.
[0292] Clause 46C. A method for processing one or more audio streams, the method comprising: determining whether a boundary separating an inner region and an outer region exists; determining a transition distance value based on the boundary existing, the transition distance value indicating a size of a transition zone, wherein the transition distance value is 0; obtaining a listener position, the listener position indicating a position of a device relative to the inner region; obtaining a current renderer based on the boundary, the listener position, and the transition distance value being 0, the current renderer being either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; applying the current renderer to the audio data to obtain one or more loudspeaker feeds; and storing the one or more loudspeaker feeds.
[0293] Clause 47C. The method of clause 46C, further comprising: determining a first distance between the listener position and a center of the inner region; determining a second distance between the boundary and the center of the inner region; and obtaining the current renderer based on the first distance and the second distance.
[0294] Clause 48C. The method of clause 46C or clause 47C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the outer renderer is configured to render an audio object comprising only a first channel of the surround sound audio data.
[0295] Clause 49C. The device of clause 46C or clause 47C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the outer renderer is configured to render the HOA to a plurality of virtual loudspeakers.
[0296] Clause 50C. The method of any of clauses 46C-49C, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the interior Tenderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0297] Clause 51C. The method of any of clauses 46C-50C, wherein the audio data includes surround sound audio data representing a primary audio source and a secondary audio source, further comprising: obtaining an indication of an opacity of the secondary audio source, and wherein obtaining the current Tenderer is based on the listener position, the boundary, and the indication.
[0298] Clause 52C. The method of clause 51C, wherein the indication of the opacity is within a bitstream.
[0299] Clause 53C. The method of clause 51C or clause 52C, further comprising: when the indication of the opacity is enabled, and based on the listener position and the boundary, obtaining the current Tenderer excludes adding the secondary source to a location for which the listener position indicates is not directly in line of sight.
[0300] Clause 54C. The method of clause 51C, wherein the exterior Tenderer is configured to: render the audio data such that a sound field represented by the audio data unfolds outwardly depending on a distance between the listener position and the boundary.
[0301] Clause 55C. The method of any of clauses 46C-51C, further comprising: obtaining an indication of a transition distance from a bitstream representing the audio data.
[0302] Clause 56C. The method of any of clauses 46C-52C, further comprising: obtaining an indication of a complexity of the current Tenderer from a bitstream representing the audio data; and wherein obtaining the current Tenderer is based on the boundary, the listener position, and the indication of the complexity.
[0303] Clause 57C. The method of clause 56C, wherein the complexity indication includes a 6DOF flag.
[0304] Clause 58C. The method of clause 57C, wherein the 6DOF flag is false.
[0305] Clause 59C. The method of clause 58C, further comprising: obtaining the interior Tenderer as a 3DOF Tenderer based at least in part on the 6DOF flag being false.
[0306] Clause 60C. The method of clause 59C, further comprising: determining a range transform that indicates whether the device renders audio sources outside the boundary.
[0307] The method of Clause 61C and Clause 60C, wherein the range transformation is true and the listener position is outside the boundary, further includes: obtaining an external renderer as the current renderer based on the range transformation being true and the listener position being outside the boundary.
[0308] The method of Clause 62C and Clause 60C, wherein the range transformation is false and the listener position is outside the boundary, further includes: obtaining an external renderer as the current renderer based on the range transformation being true and the listener position being outside the boundary, wherein the external renderer is no renderer.
[0309] The method of Clause 63C and Clause 57C, wherein the 6DOF flag is true.
[0310] The methods of Clause 64C and Clause 63C further include: obtaining an internal renderer as a 6DOF renderer based at least in part on the 6DOF flag being true.
[0311] The method of Clause 65C and Clause 64C further includes: determining a range transformation that indicates whether the device renders an audio source outside the boundary.
[0312] The method of Clause 66C and Clause 65C, wherein the range transformation is true and the listener position is outside the boundary, further includes: obtaining an external renderer as the current renderer based on the range transformation being true and the listener position being outside the boundary.
[0313] The method of Clause 67C and Clause 65C, wherein the range transformation is false and the listener position is outside the boundary, further includes: obtaining an external renderer as the current renderer based on the range transformation being false and the listener position being outside the boundary, wherein the external renderer is no renderer.
[0314] Item 68C. A method for processing one or more audio streams, the method comprising: determining whether a boundary exists separating an inner region and an outer region; determining a transition distance value based on the existence of the boundary, the transition distance value indicating the size of a transition area, wherein the transition distance value is greater than 0; obtaining a listener position indicating the position of a device relative to the inner region; obtaining a current renderer based on the boundary, the listener position, and the transition distance value being greater than 0, the current renderer being either an inner renderer configured to render audio data for the inner region, or an outer renderer configured to render audio data for the outer region, or both an inner and outer renderer; applying the current renderer to the audio data to obtain one or more speaker feeds; and storing the one or more speaker feeds.
[0315] Clause 69C. The method of clause 68C, further comprising: determining a first distance between the listener position and a center of the interior region; determining a second distance between the boundary and the center of the interior region; and obtaining the current renderer based on the first distance and the second distance.
[0316] Clause 70C. The device of clause 68C or clause 69C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the exterior renderer is configured to render an audio object comprising only a first channel of the surround sound audio data.
[0317] Clause 71C. The method of clause 68C or clause 69C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the exterior renderer is configured to render the HOA to a plurality of virtual loudspeakers.
[0318] Clause 72C. The method of any of clauses 68C-71C, wherein the audio data comprises surround sound audio data associated with spherical basis functions, and wherein the interior renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0319] Clause 73C. The method of any of clauses 68C-72C, wherein the audio data comprises surround sound audio data representing a primary audio source and a secondary audio source, further comprising: obtaining an indication of an opacity of the secondary audio source, and wherein obtaining the current renderer is based on the listener position, the boundary, and the indication.
[0320] Clause 74C. The method of clause 73C, further comprising: obtaining the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0321] Clause 75C. The method of clause 73C or clause 74C, further comprising: when the indication of the opacity is enabled, and based on the listener position and the boundary, obtaining the current renderer that excludes adding the secondary source to a location indicated by the listener position that is not directly in line of sight.
[0322] Clause 76C. The method of clause 73C, wherein the exterior renderer is configured to: render the audio data such that a sound field represented by the audio data expands outward depending on a distance between the listener position and the boundary.
[0323] Clause 77C. The method of any of clauses 68C-76C, further comprising: obtaining an indication of a transition distance from a bitstream representing the audio data.
[0324] Clause 78C. The method of clause 77C, further comprising: responsive to determining that the listener position is within the transition distance from the boundary, updating the current renderer to interpolate between the outer renderer and the inner renderer to obtain an updated current renderer; and applying the current renderer to the audio data to obtain the one or more updated speaker feeds.
[0325] Clause 79C. The method of any of clauses 68C-78C, further comprising: obtaining an indication of a complexity of the current renderer from a bitstream representing the audio data, and obtaining the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0326] Clause 80C. The method of clause 79C, wherein the indication of the complexity comprises a 6DOF flag.
[0327] Clause 81C. The method of clause 80C, wherein the 6DOF flag is false.
[0328] Clause 82C. The method of clause 81C, further comprising: obtaining the inner renderer as a 3DOF renderer based at least in part on the 6DOF flag being false.
[0329] Clause 83C. The method of clause 82C, further comprising: determining a range transform that indicates whether the device renders audio sources outside the boundary.
[0330] Clause 84C. The method of clause 83C, wherein the range transform is true and the listener position is outside the boundary, further comprising: obtaining the outer renderer as the current renderer based on the range transform being true and the listener position being outside the boundary.
[0331] Clause 85C. The method of clause 83C, wherein the range transform is false and the listener position is outside the boundary, further comprising: obtaining the outer renderer as the current renderer based on the range transform being false and the listener position being outside the boundary, wherein the outer renderer is a no-renderer.
[0332] Clause 86C. The method of clause 80C, wherein the 6DOF flag is true.
[0333] Clause 87C. The method of clause 86C, further comprising: obtaining the inner renderer as a 6DOF renderer based at least in part on the 6DOF flag being true.
[0334] Clause 88C. The method of clause 87C, further comprising: determining a range transform that indicates whether the device renders audio sources outside the boundary.
[0335] Clause 89C. The method of clause 88C, wherein the range transform is true and the listener position is outside the boundary, further comprising: obtaining, as the current renderer, an outside renderer based on the range transform being true and the listener position being outside the boundary.
[0336] Clause 90C. The method of clause 88C, wherein the range transform is false and the listener position is outside the boundary, further comprising: obtaining, as the current renderer, an outside renderer based on the range transform being false and the listener position being outside the boundary, wherein the outside renderer is a no-renderer.
[0337] Clause 91C. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determine whether a boundary separating an inner region and an outer region exists; determine, based on the boundary existing, a transition distance value that indicates a size of a transition zone, wherein the transition distance value is 0; obtain a listener position that indicates a position of a device relative to the inner region; obtain, based on the boundary, the listener position, and the transition distance value being 0, a current renderer that is either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; apply the current renderer to the audio data to obtain one or more speaker feeds; and store the one or more speaker feeds.
[0338] Clause 92C. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determine whether a boundary separating an inner region and an outer region exists; determine, based on the boundary existing, a transition distance value that indicates a size of a transition zone, wherein the transition distance value is greater than 0; obtain a listener position that indicates a position of a device relative to the inner region; obtain, based on the boundary, the listener position, and the transition distance value being greater than 0, a current renderer that is either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region or both the inner renderer and the outer renderer; apply the current renderer to the audio data to obtain one or more speaker feeds; and store the one or more speaker feeds.
[0339] Clause 93C. A device for processing one or more audio streams, the device comprising: means for determining whether a boundary separating an inner region and an outer region is present; means for determining a transition distance value based on the boundary being present, the transition distance value indicating a size of a transition zone, wherein the transition distance value is 0; means for obtaining a listener position indicating a position of the device relative to the inner region; means for obtaining a current renderer based on the boundary, the listener position, and the transition distance value being 0, the current renderer being either as an inner renderer configured to render audio data for the inner region or as an outer renderer configured to render audio data for the outer region; means for applying the current renderer to the audio data to obtain one or more speaker feeds; and means for storing the one or more speaker feeds.
[0340] Clause 94C. A device for processing one or more audio streams, the device comprising: means for determining whether a boundary separating an inner region and an outer region is present; means for determining a transition distance value based on the boundary being present, the transition distance value indicating a size of a transition zone, wherein the transition distance value is greater than 0; means for obtaining a listener position indicating a position of the device relative to the inner region; means for obtaining a current renderer based on the boundary, the listener position, and the transition distance value being greater than 0, the current renderer being either as an inner renderer configured to render audio data for the inner region or as an outer renderer configured to render audio data for the outer region or as both the inner renderer and the outer renderer; means for applying the current renderer to the audio data to obtain one or more speaker feeds; and means for storing the one or more speaker feeds.
[0341] Clause 1D. A device configured to process audio data, the device comprising: a memory configured to store one or more speaker feeds; and one or more processors implemented in circuitry and communicatively coupled to the memory, the one or more processors configured to: determine whether a boundary separating an inner region and an outer region is present; determine a transition distance value based on determining that the boundary is present, the transition distance value indicating a size of a transition zone; obtain a listener position indicating a virtual position of the device relative to the inner region; obtain a current renderer based at least in part on the boundary and the listener position; and apply the current renderer to the audio data to obtain the one or more speaker feeds.
[0342] Clause 2D. The device of clause 1D, wherein the one or more processors are further configured to: determine a first distance between the listener position and a center of the inner region; determine a second distance between the boundary and the center of the inner region; and obtain the current renderer based on the first distance and the second distance.
[0343] Clause 3D. The device of clause 1D or clause 2D, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the external renderer is configured to render audio objects that include only a first channel of the surround sound audio data.
[0344] Clause 4D. The device of clause 1D or clause 2D, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the external renderer is configured to render the HOA to a plurality of virtual loudspeakers.
[0345] Clause 5D. The device of any of clauses 1D-4D, wherein the audio data includes surround sound audio data associated with spherical basis functions, and wherein the internal renderer is configured to render the surround sound audio data such that a sound field represented by the surround sound audio data appears throughout the interior region.
[0346] Clause 6D. The device of any of clauses 1D-5D, wherein the audio data includes surround sound audio data representing a primary audio source and a secondary audio source, wherein the one or more processors are further configured to: obtain an indication of an opacity of the secondary audio source, and obtain the current renderer based on the listener position, the boundary, and the indication.
[0347] Clause 7D. The device of clause 6D, wherein the one or more processors are configured to: obtain the indication of the opacity of the secondary source from a bitstream representing the audio data.
[0348] Clause 8D. The device of clause 6D or clause 7D, wherein the one or more processors are further configured to: obtain the current renderer based on the listener position and the boundary when the indication of the opacity is enabled, the current renderer excluding adding the secondary source to a location for which the listener position indicates is not directly in line of sight.
[0349] Clause 9D. The device of clause 6D, wherein the external renderer is configured to: render the audio data such that a sound field represented by the audio data unfolds outwardly depending on a distance between the listener position and the boundary.
[0350] Clause 10D. The device of any of clauses 1D-9D, wherein the one or more processors are further configured to: obtain an indication of a transition distance value from a bitstream representing the audio data.
[0351] Clause 11D. The device of any of clauses 1D-10D, wherein the one or more processors are further configured to: obtain an indication of a complexity of the current renderer from a bitstream representing the audio data; and obtain the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0352] Clause 12D. The device of clause 11D, wherein the indication of complexity comprises a 6DOF flag.
[0353] Clause 13D. The device of clause 12D, wherein the 6DOF flag is false.
[0354] Clause 14D. The device of clause 13D, wherein the one or more processors are configured to obtain the internal Tenderer as a 3DOF Tenderer based at least in part on the 6DOF flag being false.
[0355] Clause 15D. The device of clause 14D, wherein the one or more processors are further configured to determine a range transform that indicates whether the device renders audio sources outside the boundary.
[0356] Clause 16D. The device of clause 15D, wherein the range transform is true and the listener position is outside the boundary, wherein the one or more processors are further configured to obtain an external Tenderer as the current Tenderer based on the range transform being true and the listener position being outside the boundary.
[0357] Clause 17D. The device of clause 15D, wherein the range transform is false and the listener position is outside the boundary, wherein the one or more processors are configured to obtain an external Tenderer as the current Tenderer based on the range transform being false and the listener position being outside the boundary, wherein the external Tenderer is a no Tenderer.
[0358] Clause 18D. The device of clause 12D, wherein the 6DOF flag is true.
[0359] Clause 19D. The device of clause 18D, wherein the one or more processors are configured to obtain the internal Tenderer as a 6DOF Tenderer based at least in part on the 6DOF flag being true.
[0360] Clause 20D. The device of clause 19D, wherein the one or more processors are further configured to determine a range transform that indicates whether the device renders audio sources outside the boundary.
[0361] Clause 21D. The device of clause 20D, wherein the range transform is true and the listener position is outside the boundary, wherein the one or more processors are further configured to obtain an external Tenderer as the current Tenderer based on the range transform being true and the listener position being outside the boundary.
[0362] Clause 22D. The device of clause 20D, wherein the range transform is false and the listener position is outside the boundary, wherein the one or more processors are configured to obtain an external Tenderer as the current Tenderer based on the range transform being false and the listener position being outside the boundary, wherein the external Tenderer is a no Tenderer.
[0363] Clause 23D. The device of any of clauses 1D-22D, wherein the transition distance value is 0, wherein the current renderer comprises either an interior renderer configured to render the audio data for the interior region or an exterior renderer configured to render the audio data for the exterior region, and wherein the one or more processors are configured to obtain the current renderer further based on the transition distance value being 0.
[0364] Clause 24D. The device of any of clauses 1D-22D, wherein the transition distance value is greater than 0, wherein the current renderer comprises either an interior renderer configured to render the audio data for the interior region or an exterior renderer configured to render the audio data for the exterior region, or both the interior renderer and the exterior renderer, and wherein the one or more processors are configured to obtain the current renderer further based on the transition distance value being greater than 0.
[0365] Clause 25D. The device of clause 24D, wherein the one or more processors are further configured to: in response to determining that the listener position is within the transition distance value from the boundary, update the current renderer to interpolate between the exterior renderer and the interior renderer to obtain an updated current renderer; and apply the updated current renderer to the audio data to obtain one or more updated speaker feeds.
[0366] Clause 26D. The device of clause 25D, wherein the updated current renderer cross-fades between the exterior renderer and the interior renderer.
[0367] Clause 27D. The device of clause 26D, wherein the updated current renderer cross-fades between different ambisonic orders of the exterior renderer and the interior renderer.
[0368] Clause 28D. The device of clause 27D, wherein the updated current renderer cross-fades from a higher ambisonic order to a lower ambisonic order as the listener position moves from the interior region through the transition zone toward the exterior region.
[0369] Clause 29D. A method for processing audio data, the method comprising: determining whether a boundary separating an interior region and an exterior region exists; determining a transition distance value based on determining that the boundary exists, the transition distance value indicating a size of a transition zone; obtaining a listener position, the listener position indicating a virtual position of a device relative to the interior region; obtaining a current renderer based at least in part on the boundary and the listener position; applying the current renderer to the audio data to obtain one or more speaker feeds; and storing the one or more speaker feeds.
[0370] Clause 30D. The method of clause 29D, wherein the transition distance value is 0, wherein the current renderer comprises either an interior renderer configured to render audio data for the interior region or an exterior renderer configured to render audio data for the exterior region, and wherein obtaining the current renderer is further based on the transition distance value being 0.
[0371] Clause 31D. The method of clause 29D, wherein the transition distance value is greater than 0, wherein the current renderer comprises either an interior renderer configured to render audio data for the interior region or an exterior renderer configured to render audio data for the exterior region, or both the interior renderer and the exterior renderer, and wherein obtaining the current renderer is further based on the transition distance value being greater than 0.
[0372] Clause 32D. A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determine whether a boundary separating an interior region and an exterior region exists; determine, based on a determination that the boundary exists, a transition distance value that indicates a size of a transition zone; obtain a listener position that indicates a virtual position of a device relative to the interior region; obtain, based at least in part on the boundary and the listener position, a current renderer; apply the current renderer to audio data to obtain one or more speaker feeds; and store the one or more speaker feeds.
[0373] Clause 33D. A device configured to process one or more audio streams, the device comprising: means for determining whether a boundary separating an interior region and an exterior region exists; means for determining, based on a determination that the boundary exists, a transition distance value that indicates a size of a transition zone; means for obtaining a listener position that indicates a virtual position of a device relative to the interior region; means for obtaining, based at least in part on the boundary and the listener position, a current renderer; means for applying the current renderer to audio data to obtain one or more speaker feeds; and means for storing the one or more speaker feeds.
[0374] It is recognized that depending on the example, certain acts or events of any of the techniques described herein can be performed in a different sequence, can be added, merged, or omitted altogether (e.g., all described acts or events can not be necessary to practice the techniques). Moreover, in certain examples, acts or events can be performed concurrently, e.g., through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.
[0375] In some examples, the VR device (or streaming device) can use a network interface coupled to a memory of the VR / streaming device to communicate exchange messages to an external device, where the exchange messages are associated with the plurality of available representations of the soundfield. In some examples, the VR device can use an antenna coupled to the network interface to receive a wireless signal that includes data packets, audio packets, video packets, or transport protocol data associated with the plurality of available representations of the soundfield. In some examples, one or more microphone arrays can capture the soundfield.
[0376] In some examples, the plurality of available representations of the soundfield stored to the memory device can include a plurality of object-based representations of the soundfield, a higher order ambisonic representation of the soundfield, a hybrid order ambisonic representation of the soundfield, a combination of an object-based representation of the soundfield and a higher order ambisonic representation of the soundfield, a combination of an object-based representation of the soundfield and a hybrid order ambisonic representation of the soundfield, or a combination of a hybrid order representation of the soundfield and a higher order ambisonic representation of the soundfield.
[0377] In some examples, one or more of the soundfield representations of the plurality of available representations of the soundfield can include at least one high resolution region and at least one lower resolution region, and where the selected representation based on the steering angle provides greater spatial accuracy with respect to the at least one high resolution region and lesser spatial accuracy with respect to the lower resolution region.
[0378] In one or more examples, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer- readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0379] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any
[0380] Instructions can be executed by one or more processors, including fixed function processing circuitry and / or programmable processing circuitry, e.g., one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein the term “processor” can refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0381] The techniques of this disclosure can be implemented in a variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functionality of devices configured to perform the disclosed techniques, but do not necessarily require separate physical components. As described above, various units can be combined in a codec hardware unit or provided by a set of inter-operating hardware units including one or more processors as described above, with suitable software and / or firmware.
[0382] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A device configured to process audio data, the device comprising: A memory configured to store one or more speaker feeds; as well as One or more processors, implemented in a circuit system and communicatively coupled to the memory, are configured to: Determine if a boundary exists that separates the internal and external regions; The transition distance value is determined based on the existence of the boundary, and the transition distance value indicates the size of the transition zone; Obtain the listener's location, which indicates the virtual position of the device relative to the internal area; The current renderer is obtained at least in part based on the boundary and the listener's position, wherein the current renderer either includes an internal renderer configured to render audio data for the inner region, or an external renderer configured to render the audio data for the outer region, or both the internal and external renderers; and Apply the current renderer to the audio data to obtain the one or more speaker feeds. The one or more processors are further configured to: When the transition distance value is greater than 0: In response to determining that the listener's position is within the transition distance value from the boundary, the current renderer is updated to interpolate between the external renderer and the internal renderer to obtain an updated current renderer; and The updated current renderer is applied to the audio data to obtain one or more updated speaker feeds.
2. The device according to claim 1, wherein, The one or more processors are further configured to: Determine a first distance between the listener's location and the center of the internal region; Determine a second distance between the boundary and the center of the interior region; and The current renderer is obtained based on the first distance and the second distance.
3. The device according to claim 1, in, The audio data includes surround sound audio data associated with spherical basis functions. The external renderer is configured to render an audio object that includes only the first channel of the surround sound audio data.
4. The device according to claim 1, in, The audio data includes surround sound audio data associated with spherical basis functions. The external renderer is configured to render the HOA to multiple virtual speakers.
5. The device according to claim 1, in, The audio data includes surround sound audio data associated with spherical basis functions. The internal renderer is configured to render the surround sound audio data such that the sound field represented by the surround sound audio data appears throughout the internal region.
6. The device according to claim 1, in, The audio data includes surround sound audio data representing primary and secondary audio sources. The one or more processors are further configured to: Obtain an indication of the opacity of the secondary audio source, and The current renderer is obtained based on the listener's location, the boundary, and the indication.
7. The device according to claim 6, wherein, The one or more processors are configured to obtain an indication of the opacity of the secondary audio source from a bitstream representing the audio data.
8. The device according to claim 6, wherein, The one or more processors are further configured to: when the opacity indication is enabled, and to obtain the current renderer based on the listener position and the boundary, the current renderer not to add the secondary audio source to a location indicated by the listener position as not directly in line of sight.
9. The device according to claim 6, wherein, An external renderer is configured to render the audio data such that the sound field represented by the audio data expands outward depending on the distance between the listener's position and the boundary.
10. The device according to claim 1, wherein, The one or more processors are further configured to obtain an indication of the transition distance value from a bitstream representing the audio data.
11. The device according to claim 1, wherein, The one or more processors are further configured to: Obtain an indication of the complexity of the current renderer from the bitstream representing the audio data; and The current renderer is obtained based on the indications of the boundary, the listener's location, and the complexity.
12. The device according to claim 11, wherein, The complexity indicator includes the 6DOF flag.
13. The device according to claim 12, wherein, The 6DOF flag is false.
14. The device according to claim 13, wherein, One or more processors are configured to obtain an internal renderer as a 3DOF renderer, at least in part based on the 6DOF flag being false.
15. The device according to claim 14, wherein, The one or more processors are further configured to: determine a range transformation, the range transformation indicating whether the device renders audio sources outside the boundary.
16. The device according to claim 15, wherein, The range transformation is true and the listener's position is outside the boundary, wherein the one or more processors are further configured to: obtain an external renderer as the current renderer based on the range transformation being true and the listener's position being outside the boundary.
17. The device according to claim 15, wherein, The range transformation is false and the listener's position is outside the boundary, wherein the one or more processors are configured to: obtain an external renderer as the current renderer based on the range transformation being false and the listener's position being outside the boundary, wherein the external renderer is no renderer.
18. The device according to claim 12, wherein, The 6DOF flag is true.
19. The device according to claim 18, wherein, One or more processors are configured to obtain an internal renderer as a 6DOF renderer, at least in part based on the 6DOF flag being true.
20. The device according to claim 19, wherein, The one or more processors are further configured to: determine a range transformation, the range transformation indicating whether the device renders audio sources outside the boundary.
21. The device according to claim 20, wherein, The range transformation is true and the listener's position is outside the boundary, wherein the one or more processors are further configured to: obtain an external renderer as the current renderer based on the range transformation being true and the listener's position being outside the boundary.
22. The device according to claim 20, wherein, The range transformation is false and the listener's position is outside the boundary, wherein the one or more processors are configured to: obtain an external renderer as the current renderer based on the range transformation being false and the listener's position being outside the boundary, wherein the external renderer is no renderer.
23. The device according to claim 1, wherein, The interpolation uses (1-a)*internal_rendering+a*external_rendering, where internal_rendering is the internal renderer, external_rendering is the external renderer, and a is a score based on how close the listener is to the boundary.
24. A method for processing audio data, the method comprising: Determine if a boundary exists that separates the internal and external regions; The transition distance value is determined based on the existence of the boundary, and the transition distance value indicates the size of the transition zone; Obtain the listener's location, which indicates the virtual position of the device relative to the internal area; The current renderer is obtained at least in part based on the boundary and the listener's position, wherein the current renderer either includes an internal renderer configured to render audio data for the inner region, or an external renderer configured to render the audio data for the outer region, or both the internal renderer and the external renderer. Apply the current renderer to the audio data to obtain one or more speaker feeds; and Store the feed from the one or more speakers. The method further includes: When the transition distance value is greater than 0: In response to determining that the listener's position is within the transition distance value from the boundary, the current renderer is updated to interpolate between the external renderer and the internal renderer to obtain an updated current renderer; and The updated current renderer is applied to the audio data to obtain one or more updated speaker feeds.
25. A computer-readable storage medium having instructions stored thereon, said instructions, when executed, causing one or more processors to perform the following operations: Determine if a boundary exists that separates the internal and external regions; The transition distance value is determined based on the existence of the boundary, and the transition distance value indicates the size of the transition zone; Obtain the listener's location, which indicates the virtual position of the device relative to the internal area; The current renderer is obtained at least in part based on the boundary and the listener's position, wherein the current renderer either includes an internal renderer configured to render audio data for the inner region, or an external renderer configured to render the audio data for the outer region, or both the internal renderer and the external renderer. Apply the current renderer to the audio data to obtain one or more speaker feeds; and Store the feed from the one or more speakers. When executed, the instructions also cause the one or more processors to perform the following operations: When the transition distance value is greater than 0: In response to determining that the listener's position is within the transition distance value from the boundary, the current renderer is updated to interpolate between the external renderer and the internal renderer to obtain an updated current renderer; and The updated current renderer is applied to the audio data to obtain one or more updated speaker feeds.
Citation Information
Patent Citations
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US10405126B2
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US20190007781A1
Controlling rendering of audio data
US20210099825A1
Determining renderers for spherical harmonic coefficients
CN104956695A
Switching rendering mode based on location data
EP3410747A1