Controlling the presentation of audio data
By dynamically adjusting the audio renderer in a real system with computers as media, and optimizing audio presentation based on boundaries and listener locations, the existing system solves the problem of balancing complexity and immersive experience in the audio experience, achieving efficient resource utilization and more realistic auditory immersion.
Patent Information
- Application Number
- CN202080062647.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2020-10-01
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2040-10-01
AI Technical Summary
The existing computer-mediated reality systems have difficulty providing a realistic immersive experience in audio experience, especially as video experience improves, the importance of auditory experience gradually increases, and existing audio playback systems are difficult to balance between processing complexity and immersive experience.
By obtaining boundary indications of the internal and external areas and listener locations, dynamically adjust the configuration of the audio renderer, generate internal or external audio renderings, optimize processor cycles, memory and bandwidth usage, and achieve flexible audio rendering.
Reduce processing resource consumption at low complexity, while providing a more realistic immersive XR experience at high complexity, improving users' auditory immersion.
Smart Images

Figure CN114424587B_ABST
Abstract
Description
[0001] This application claims priority to U.S. patent application Ser. No. 17 / 038,618, filed on September 30, 2020, entitled “CONTROLLING RENDERING OF AUDIO DATA,” which claims the benefit of U.S. Provisional Application Serial No. 62 / 909,104, filed on October 1, 2019, entitled “CONTROLLING RENDERING OF AUDIO DATA,” the entire contents of each of which are incorporated herein by reference as if fully set forth herein. Technical Field
[0002] The present disclosure relates to the processing of audio data. Background Art
[0003] Computer-mediated reality systems are being developed to allow computing devices to augment or add to, remove or subtract from, or more generally modify, the existing reality that a user experiences. Computer-mediated reality systems (which may also be referred to as "extended reality systems" or "XR systems") may include, for example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceived success of a computer-mediated reality system generally relates to the ability of such a computer-mediated reality system to provide a realistic immersive experience in terms of both the video and audio experience, where the video and audio experience conform to the manner expected by the user. Although the human visual system is more sensitive than the human auditory system (e.g., in terms of perceived positioning of various objects within a scene), ensuring an adequate auditory experience is an increasingly important factor in ensuring a realistic immersive experience, particularly as the video experience improves to allow better positioning of video objects, which enables the user to better identify the source of the audio content. Summary of the Invention
[0004] The present disclosure generally relates to techniques for controlling audio rendering at an audio playback system. The techniques can enable the audio playback system to perform flexible rendering in terms of complexity (such as complexity defined by processor cycles, memory, and / or bandwidth consumed), while also allowing for internal and external rendering of an XR experience as defined by a boundary that separates an internal area from an external area. In addition, the audio playback system can configure the audio renderer using metadata or other indications specified in a bitstream representing the audio data, while also generating the audio renderer to account for the internal area or the external area with reference to the listener's position relative to the boundary.
[0005] Thus, the technology can improve the operation of an audio playback system because, when the audio playback system is configured to perform low-complexity rendering, the audio playback system can reduce the amount of processor cycles, memory, and / or bandwidth consumed. When performing high-complexity rendering, the audio playback system can provide a more immersive XR experience, which can result in a user of the audio playback system being more realistically placed in the XR experience.
[0006] In one example, the technology is directed to a device configured to process one or more audio streams, the device comprising: one or more processors configured to: obtain an indication of a boundary separating an interior area from an exterior area; obtain a listener position indicating a position of the device relative to the interior area; based on the boundary and the listener position, obtain a current renderer as either an interior renderer configured to render audio data for the interior area or an external renderer configured to render audio data for the exterior area; apply the current renderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0007] In another example, the technology is directed to a method of processing one or more audio streams, the method comprising: obtaining, by one or more processors, an indication of a boundary separating an inner region from an outer region; obtaining, by the one or more processors, a listener position indicating a position of a device relative to the inner region; obtaining, by the one or more processors, a current renderer as an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region based on the boundary and the listener position; and applying, by the one or more processors, the current renderer to the audio data to obtain one or more speaker feeds.
[0008] In another example, the technology is directed to a device configured to process one or more audio streams, the device including: a unit for obtaining an indication of a boundary separating an inner region from an outer region; a unit for obtaining a listener position indicating a position of the device relative to the inner region; a unit for obtaining, based on the boundary and the listener position, a current renderer that is either an inner renderer configured to render audio data for the inner region or an external renderer configured to render audio data for the outer region; and a unit for applying the current renderer to the audio data to obtain one or more speaker feeds.
[0009] In another example, the technology is directed to a non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: obtain an indication of a boundary separating an interior area from an exterior area; obtain a listener position indicating a position of a device relative to the interior area; based on the boundary and the listener position, obtain a current renderer that is either an interior renderer configured to render audio data for the interior area or an external renderer configured to render audio data for the exterior area; and apply the current renderer to the audio data to obtain one or more speaker feeds.
[0010] In another example, the technology is directed to a device configured to generate a bitstream representing audio data, the device including: a memory configured to store the audio data; and one or more processors coupled to the memory and configured to: obtain a bitstream representing the audio data based on the audio data; specify in the bitstream a boundary separating an inner region from an outer region; specify in the bitstream one or more indications for controlling presentation of the audio data for the inner region or the outer region; and output the bitstream.
[0011] In another example, the technology is directed to a method of generating a bitstream representing audio data, the method comprising: obtaining a bitstream representing the audio data based on the audio data; specifying a boundary separating an inner region from an outer region in the bitstream; specifying one or more indications in the bitstream that control presentation of the audio data for the inner region or the outer region; and outputting the bitstream.
[0012] In another example, the technology is directed to a device configured to generate a bitstream representing audio data, the device including: a unit for obtaining a bitstream representing the audio data based on the audio data; a unit for specifying a boundary separating an inner region from an outer region in the bitstream; a unit for specifying one or more indications in the bitstream that control presentation of the audio data for the inner region or the outer region; and a unit for outputting the bitstream.
[0013] In another example, the technology is directed to a non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: obtain a bitstream representing the audio data based on audio data; specify in the bitstream a boundary that separates an inner region from an outer region; specify in the bitstream one or more indications for controlling presentation of the audio data by the inner region or the outer region; and output the bitstream.
[0014] The details of one or more examples of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of various aspects of the technology will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1A and 1B is a diagram illustrating a system that can perform various aspects of the techniques described in this disclosure.
[0016] Figure 2 is a diagram illustrating an example of low-complexity rendering for an extended reality (XR) scene, in accordance with various aspects of the techniques described in this disclosure.
[0017] Figure 3 is a diagram illustrating an example of high complexity rendering including a distance buffer for an XR scene, in accordance with various aspects of the techniques described in this disclosure.
[0018] Figure 4A and 4B is a diagram illustrating an example of a VR device.
[0019] Figure 5A and 5B is a diagram illustrating an example system that can perform various aspects of the techniques described in this disclosure.
[0020] Figures 6A-6F It shows Figure 1A and 1B Block diagrams of various examples of an audio playback device shown in the examples of FIG. 1 in performing various aspects of the techniques described in this disclosure.
[0021] Figure 7 An example of a wireless communication system supporting audio streaming according to aspects of the present disclosure is shown.
[0022] Figure 8 It shows Figure 1A Flowchart of example operations of a source device shown in FIG. 1 in performing various aspects of the techniques described in this disclosure.
[0023] Figure 9 It shows Figure 1A Flowchart of example operations of a content consumer device shown in FIG. 1 in performing various aspects of the techniques described in this disclosure. DETAILED DESCRIPTION
[0024] There are many different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound formats, 7.1 surround sound formats, 22.2 surround sound formats, or any other channel-based format that positions audio channels to specific locations around the listener in order to recreate the sound field.
[0025] An object-based audio format may refer to a format in which audio objects are specified to represent a sound field, and the audio objects are often encoded using pulse code modulation (PCM) and are referred to as PCM audio objects. Such audio objects may include metadata that identifies the position of the audio objects in the sound field relative to the listener or other reference points, so that the audio objects can be rendered to one or more speaker channels for playback in an effort to recreate the sound field. The techniques described in this disclosure may be applied to any of the aforementioned formats, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.
[0026] A scene-based audio format may include a hierarchical collection of elements that define a sound field in three dimensions. An example of a hierarchical collection of elements is a collection of spherical harmonic coefficients (SHCs). The following expression demonstrates the description or representation of a sound field using SHCs:
[0027]
[0028] This expression shows that at any point in the sound field The pressure at p i At time t, SHC Uniquely represents. Here, c is the speed of sound (~343 m / s), is the reference point (or observation point), j n (·) is a spherical Bessel function of order n, and are the spherical harmonic basis functions of order n and suborder m (also called spherical basis functions). It can be recognized that the terms in the square brackets are signals (i.e., ), which can be approximated by various time-frequency transforms, such as discrete Fourier transform (DFT), discrete cosine transform (DCT), or wavelet transform. Other examples of hierarchical sets include sets of wavelet transform coefficients and other sets of coefficients of multiresolution basis functions.
[0029] SHC It can be physically acquired (e.g., recorded) by various microphone array configurations, or alternatively, it can be derived from a channel-based or object-based description of the sound field. SHC (which can also be called ambisonic coefficients) represents scene-based audio, where the SHC can be input to an audio encoder to obtain encoded SHC that can facilitate more efficient transmission or storage. For example, a method involving (1+4) 2 (25, hence fourth order) The fourth order representation of the coefficients.
[0030] As described above, SHC can be derived from microphone recordings using a microphone array. Various examples of how to physically obtain SHC from a microphone array are described in Poletti, M., "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics" (J. Audio Eng. Soc., Vol. 53, No. 11, November 2005, pp. 1004-1025).
[0031] The following equation illustrates how the SHC is derived from an object-based description. The coefficients for the sound field corresponding to a single audio object are It can be expressed as:
[0032]
[0033] where i is is a spherical Hankel function of order n (of the second kind), and is the position of the object. Knowing the object source energy g(ω) as a function of frequency (e.g., using time-frequency analysis techniques such as performing a Fast Fourier Transform on a Pulse Code Modulation - PCM - stream) allows converting each PCM object and corresponding position into an SHC Furthermore, it can be shown (due to the linear and orthogonal decomposition above) that for each object The coefficients are additive. In this way, multiple PCM objects can be represented by The coefficients are represented (e.g. as the sum of the coefficient vectors of the individual objects). The coefficients may contain information about the sound field (pressure as a function of 3D coordinates) and the above representation is at the observation point The transformation from individual objects to a representation of the entire sound field.
[0034] Computer-mediated reality systems (which may also be referred to as "extended reality systems" or "XR systems") are being developed to exploit the many potential benefits provided by ambisonics coefficients. For example, ambisonics coefficients can represent a sound field in three dimensions, potentially enabling accurate three-dimensional (3D) positioning of sound sources within the sound field. Thus, an XR device can render the ambisonics coefficients to a speaker feed that accurately reproduces the sound field when played through one or more speakers.
[0035] Using ambisonics coefficients for XR can enable the development of many use cases that rely on the more immersive sound field provided by ambisonics coefficients, especially for computer gaming applications and live video streaming applications. In these highly dynamic use cases that rely on low-latency reproduction of the sound field, XR devices can prefer ambisonics coefficients over other representations that are more difficult to manipulate or involve complex rendering. More information about these use cases is available below. Figure 1A and 1B supply.
[0036] While described in this disclosure with respect to VR devices, various aspects of the technology may be performed in the context of other devices, such as mobile devices. In this case, a mobile device (such as a so-called smartphone) may present the displayed world via a screen that may be mounted to the head of the user 102 or viewed as would be done during normal use of the mobile device. In this way, any information on the screen may be part of the mobile device. The mobile device may be able to provide tracking information 41, thereby allowing the displayed world to be viewed in both a VR experience (when head-mounted) and a normal experience, wherein the normal experience may still allow the user to view the displayed world, thereby presenting a VR-lite type of experience (e.g., picking up the device and rotating or translating the device to view different parts of the displayed world).
[0037] The present disclosure may provide various combinations of opacity attributes and interpolation distance attributes to render interior ambisonic sound fields for 6 DoF (and other) use cases. In addition, the present disclosure discusses examples of low-complexity and high-complexity rendering solutions for interior ambisonic sound fields, which may be specified by a single binary bit. In an example encoder input format, there may be an attribute indicating whether the ambisonic sound field description is an interior field or an exterior field. In an interior sound field, sound sources are within a specified boundary described by a mesh or simple geometric object, while for an exterior sound field, sound sources are described as being outside the boundary. The opacity attribute of the interior sound field may specify whether contributions that do not have a direct line of sight to the listener contribute to the rendering of the sound field for the listener when the listener is outside the boundary. In addition, the distance attribute may specify a buffer area around the boundary in which interpolation between the rendering of the interior field for an exterior listener to the interior listener is used.
[0038] Thus, various aspects of the technology described herein may enable determining a user's listener position while navigating in a VR or other XR setting, determining whether the listener position is within a geometric boundary, wherein all sound sources radiate toward the listener without obstructions within the geometric boundary, and determining whether the listener position is outside the geometric boundary. Various aspects of the technology may also enable assigning an opacity attribute to each sound source that is obstructed relative to the listener when the listener position is determined to be outside the geometric boundary, and performing interpolation of a sound field within the geometric boundary based on the opacity attribute when the listener position indicates that the listener is outside the geometric boundary, and presenting the interpolated sound field.
[0039] Figure 1A and 1B is a diagram illustrating a system that can perform various aspects of the techniques described in this disclosure. Figure 1A As shown in the example of , system 10 includes a source device 12 and a content consumer device 14. Although described in the context of source device 12 and content consumer device 14, the techniques may be implemented in any context in which any layered representation of a sound field is encoded to form a bitstream representing audio data. Furthermore, source device 12 may represent any form of computing device capable of generating a layered representation of a sound field, and is generally described herein in the context of being a VR content creator device. Likewise, content consumer device 14 may represent any form of computing device capable of implementing the audio stream interpolation techniques and audio playback described in this disclosure, and is generally described herein in the context of being a VR client device.
[0040] Source device 12 may be operated by an entertainment company or other entity that generates multi-channel audio content for consumption by an operator of a content consumer device, such as content consumer device 14. In many VR scenarios, source device 12 generates audio content in conjunction with video content. Source device 12 includes a content capture device 300 and a content sound field representation generator 302.
[0041] The content capture device 300 may be configured to interface with or otherwise communicate with one or more microphones 5A-5N ("microphones 5"). Microphones 5 may represent Or other types of 3D audio microphones capable of capturing a sound field and representing the sound field as corresponding scene-based audio data 11A to 11N (which are also referred to as ambisonic sound coefficients 11A to 11N or “ambisonic sound coefficients 11”). In the context of scene-based audio data 11 (which is another way of referring to ambisonic sound coefficients 11), each of the microphones 5 may represent a cluster of microphones arranged within a single housing according to a collective geometry that is conducive to generating ambisonic sound coefficients 11. As such, the term microphone may refer to a cluster of microphones (which is actually geometrically arranged transducers) or a single microphone (which may be referred to as a point microphone).
[0042] The ambisonic coefficients 11 may represent an example of an audio stream. Therefore, the ambisonic coefficients 11 may also be referred to as the audio stream 11. Although primarily described with respect to the ambisonic coefficients 11, the techniques may be performed with respect to other types of audio streams, including pulse code modulation (PCM) audio streams, channel-based audio streams, object-based audio streams, and the like.
[0043] In some examples, the content capture device 300 may include an integrated microphone integrated into the housing of the content capture device 300. The content capture device 300 may interface with the microphone 5 wirelessly or via a wired connection. Instead of or in conjunction with capturing audio data via the microphone 5, the content capture device 300 may process the ambisonic sound coefficients 11 after inputting the ambisonic sound coefficients 11 via some type of removable storage device, wirelessly, and / or via a wired input process, or alternatively or in conjunction with the foregoing, generating the ambisonic sound coefficients 11, or otherwise creating the ambisonic sound coefficients 11 (from stored sound samples, such as is common in gaming applications, etc.). In this way, various combinations of the content capture device 300 and the microphone 5 are possible.
[0044] The content capture device 300 may also be configured to interface with or otherwise communicate with a sound field representation generator 302. The sound field representation generator 302 may include any type of hardware device capable of interfacing with the content capture device 300. The sound field representation generator 302 may use the ambisonics coefficients 11 provided by the content capture device 300 to generate various representations of the same sound field represented by the ambisonics coefficients 11.
[0045] For example, to generate different representations of the sound field using ambisonic coefficients (which is also an example of an audio stream), the sound field representation generator 24 may use a coding scheme for ambisonic representation of the sound field known as Mix Order Ambisonics (MOA), as discussed in more detail in U.S. application serial number 15 / 672,058, entitled “MIXED-ORDER AMBISONICS (MOA) AUDIO DATAFO COMPUTER-MEDIATED REALITY SYSTEMS,” filed on August 8, 2017 and published on January 3, 2019 as U.S. Patent Publication No. 20190007781.
[0046] To generate a particular MOA representation of a sound field, the sound field representation generator 24 may generate a partial subset of the full set of ambisonic coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide accuracy for some areas of the sound field, but less accuracy in other areas. In one example, the MOA representation of a sound field may include eight (8) uncompressed ambisonic coefficients, while a third-order ambisonic representation of the same sound field may include sixteen (16) uncompressed ambisonic coefficients. Thus, each MOA representation of a sound field generated as a partial subset of the ambisonic coefficients may be less storage-intensive and less bandwidth-intensive (if and when transmitted as part of the bitstream 27 over the transmission channel shown) than a corresponding third-order ambisonic representation of the same sound field generated from the ambisonic coefficients.
[0047] Although described with respect to an MOA representation, the techniques of the present disclosure may also be performed with respect to a first-order ambisonic (FOA) representation, in which all ambisonic coefficients associated with the first-order spherical basis functions and the zeroth-order spherical basis functions are used to represent the sound field. That is, the sound field representation generator 302 may use all ambisonic coefficients of a given order N to represent the sound field, rather than using a partial non-zero subset of the ambisonic coefficients to represent the sound field, resulting in the sum of the ambisonic coefficients being equal to (N+1). 2 .
[0048] In this regard, ambisonic audio data (which is another way of referring to ambisonic coefficients in an MOA representation or a full-order representation (e.g., the first-order representation described above)) may include ambisonic coefficients associated with spherical basis functions having an order of one or less (which may be referred to as "first-order ambisonic audio data"), ambisonic coefficients associated with spherical basis functions having mixed and sub-orders (which may be referred to as the "MOA representation" discussed above), or ambisonic coefficients associated with spherical basis functions having an order greater than one (which is referred to as the "full-order representation" above).
[0049] In some examples, the content capture device 300 can be configured to communicate wirelessly with the sound field representation generator 302. In some examples, the content capture device 300 can communicate with the sound field representation generator 302 via one or both of a wireless connection or a wired connection. Via the connection between the content capture device 300 and the sound field representation generator 302, the content capture device 300 can provide content in various forms of content, which for purposes of discussion is described herein as part of the ambisonics coefficients 11.
[0050] In some examples, the content capture device 300 may utilize various aspects of the sound field representation generator 302 (in terms of hardware or software capabilities of the sound field representation generator 302). For example, the sound field representation generator 302 may include dedicated hardware (or dedicated software that, when executed, causes one or more processors to perform psychoacoustic audio coding) configured to perform psychoacoustic audio coding (psychoacoustic audio coding such as the Unified Speech and Audio Codec denoted as “USAC” as set forth by the Moving Picture Experts Group (MPEG), the MPEG-H 3D Audio coding standard, the MPEG-I Immersive Audio standard, or a proprietary standard such as AptX). TM (Including various versions of AptX, such as Enhanced AptX (E-AptX), AptX live, AptXstereo and AptX high definition (AptX-HD), Advanced Audio Coding (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey's Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus and Windows Media Audio (WMA)).
[0051] The content capture device 300 may not include dedicated hardware or dedicated software for a psychoacoustic audio encoder, but rather provide the audio aspects of the content 301 in a non-psychoacoustic audio coded form. The sound field representation generator 302 may assist in capturing the content 301 at least in part by performing psychoacoustic audio encoding on the audio aspects of the content 301.
[0052] The sound field representation generator 302 may also assist in content capture and transmission by generating one or more bitstreams 21 based at least in part on audio content (e.g., an MOA representation, a third-order ambisonic representation, and / or a first-order ambisonic representation) generated from the ambisonic coefficients 11. The bitstream 21 may represent a compressed version of the ambisonic coefficients 11 (and / or a partial subset thereof used to form the MOA representation of the sound field) and any other different types of content 301 (e.g., a compressed version of spherical video data, image data, or text data).
[0053] The sound field representation generator 302 may generate a bitstream 21 for transmission across a transmission channel, such as a wired or wireless channel, a data storage device, or the like (as an example). The bitstream 21 may represent an encoded version of the ambisonics coefficients 11 (and / or a partial subset thereof used to form an MOA representation of the sound field) and may include a main bitstream and another side bitstream, which may be referred to as side channel information. In some cases, the bitstream 21 representing the compressed version of the ambisonics coefficients 11 may conform to a bitstream generated according to the MPEG-H 3D Audio coding standard.
[0054] The content consumer device 14 may be operated by an individual and may represent a VR client device. Although described with respect to a VR client device, the content consumer device 14 may represent other types of devices, such as an augmented reality (AR) client device, a mixed reality (MR) client device (or any other type of head-mounted display device or extended reality - XR- device), a standard computer, a headset, headphones, or any other device capable of tracking the head movements and / or general translational movements of the individual operating the client consumer device 14. Figure 1A As shown in the example of , content consumer device 14 includes audio playback system 16A, which may refer to any form of audio playback system capable of rendering ambisonic coefficients (whether in the form of first-order, second-order, and / or third-order ambisonic representations and / or MOA representations) for playback as multi-channel audio content.
[0055] Content consumer device 14 may extract bitstream 21 directly from source device 12. In some examples, content consumer device 12 may interface with a network, including a fifth generation (5G) cellular network, to extract bitstream 21 or otherwise cause source device 12 to transmit bitstream 21 to content consumer device 14.
[0056] Although Figure 1A 14. Although shown in FIG. 14 as being sent directly to content consumer device 14, source device 12 may output bitstream 21 to an intermediary device located between source device 12 and content consumer device 14. The intermediary device may store bitstream 21 for later delivery to content consumer device 14, which may request the bitstream. The intermediary device may include a file server, a network server, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a smartphone, or any other device capable of storing bitstream 21 for later extraction by an audio decoder. The intermediary device may reside in a content delivery network that is capable of streaming bitstream 21 (and possibly along with sending a corresponding video data bitstream) to a subscriber (such as content consumer device 14) requesting the bitstream 21.
[0057] Alternatively, source device 12 may store bitstream 21 to a storage medium such as a compact disc, digital video disc, high-definition video disc, or other storage medium, most of which are capable of being read by a computer and thus may be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this case, the transmission channel may refer to the channel through which the content stored to the medium is sent (and may include retail stores and other store-based delivery mechanisms). In any case, the techniques of this disclosure should not be limited to Figure 1A .
[0058] As described above, the content consumer device 14 includes an audio playback system 16. The audio playback system 16 can represent any system capable of playing back multi-channel audio data. The audio playback system 16A can include a plurality of different audio renderers 22. The renderers 22 can each provide a different form of audio presentation, where the different forms of presentation can include one or more of various ways of performing Vector Basis Amplitude Panning (VBAP) and / or one or more of various ways of performing sound field synthesis. As used herein, "A and / or B" means "A or B", or "both A and B".
[0059] The audio playback system 16A may further include an audio decoding device 24. The audio decoding device 24 may represent a device configured to decode the bitstream 21 to output reconstructed ambisonic coefficients 11A′-11N′ (which may form a complete first-order, second-order, and / or third-order ambisonic representation, or a subset thereof, forming an MOA representation of the same sound field, or a decomposition thereof, such as the primary audio signal, ambisonic coefficients, and vector-based signals described in the MPEG-H 3D Audio coding standard and / or the MPEG-I Immersive Audio standard).
[0060] Thus, the ambisonic coefficients 11A'-11N' ("ambisonic coefficients 11") may be similar to the full set or a partial subset of the ambisonic coefficients 11, but may differ due to lossy operations (e.g., quantization) and / or transmission via a transmission channel. The audio playback system 16 may, after decoding the bitstream 21 to obtain the ambisonic coefficients 11', obtain the ambisonic audio data 15 from the different streams of ambisonic coefficients 11' and render the ambisonic audio data 15 to output a speaker feed 25. The speaker feed 25 may drive one or more speakers (for ease of illustration, in FIG. Figure 1A (not shown in the example). The ambisonic representation of the sound field can be normalized in a variety of ways, including N3D, SN3D, FuMa, N2D, or SN2D.
[0061] To select an appropriate renderer, or in some cases generate an appropriate renderer, the audio playback system 16A may obtain loudspeaker information 13 indicating the number of loudspeakers and / or the spatial geometry of the loudspeakers. In some examples, the audio playback system 16A may dynamically determine the loudspeaker information 13 via the reference microphone, using a reference microphone and outputting a signal to activate (or in other words, drive) the loudspeaker to obtain the loudspeaker information 13. In other examples, or in conjunction with the dynamic determination of the loudspeaker information 13, the audio playback system 16A may prompt the user to interface with the audio playback system 16A and input the loudspeaker information 13.
[0062] The audio playback system 16A may select one of the audio renderers 22 based on the microphone information 13. In some examples, the audio playback system 16A may generate one of the audio renderers 22 based on the microphone information 13 when none of the audio renderers 22 are within a certain threshold similarity metric (in terms of microphone geometry) to the microphone geometry specified in the microphone information 13. In some cases, the audio playback system 16A may generate one of the audio renderers 22 based on the microphone information 13 without first attempting to select an existing one of the audio renderers 22.
[0063] When outputting the speaker feeds 25 to the headphones, the audio playback system 16A can utilize one of the renderers 22 that uses a head-related transfer function (HRTF) or other function that can be rendered to the left and right speaker feeds 25 for headphone speaker playback to provide binaural rendering. The terms "speaker" or "transducer" can generally refer to any speaker, including a loudspeaker, headphone speaker, etc. The one or more speakers can then play back the rendered speaker feeds 25.
[0064] Although described as rendering the speaker feeds 25 from the ambisonic audio data 15, reference to rendering the speaker feeds 25 may refer to other types of rendering, such as rendering that is directly incorporated into the decoding of the ambisonic audio data 15 from the bitstream 21. An example of an alternative rendering can be found in Annex G of the MPEG-H 3D Audio coding standard, where rendering occurs during the formation of the main and background signals prior to synthesis of the sound field. Therefore, reference to rendering the ambisonic audio data 15 should be understood to refer to both rendering the actual ambisonic audio data 15 or a decomposition or representation of the ambisonic audio data 15 (such as the main audio signal, ambisonic coefficients, and / or vector-based signals, which may also be referred to as V-vectors, mentioned above).
[0065] As described above, the content consumer device 14 may represent a VR device in which a human wearable display is mounted in front of the eyes of a user operating the VR device. Figure 4A and 4B is a diagram showing examples of VR devices 400A and 400B. Figure 4A In the example of FIG4 , VR device 400A is coupled to or otherwise includes headphones 404, which can reproduce a sound field represented by ambisonic audio data 15 (which is another way to refer to ambisonic coefficients 15) by playing back speaker feeds 25. Speaker feeds 25 can represent analog or digital signals that can cause a membrane within a transducer of headphones 404 to vibrate at various frequencies. This process is generally referred to as driving headphones 404.
[0066] Video, audio, and other sensory data may play an important role in a VR experience. To participate in a VR experience, a user 402 may wear a VR device 400A (which may also be referred to as a VR headset 400A) or other wearable electronic device. A VR client device (e.g., VR headset 400A) may track the head movements of the user 402 and adjust the video data displayed via the VR headset 400A to account for the head movements, thereby providing an immersive experience in which the user 402 may experience a virtual world displayed in the video data in visual three dimensions.
[0067] While VR (and other forms of AR and / or MR, which may generally be referred to as computer-mediated reality devices) may allow a user 402 to visually be present in a virtual world, a VR headset 400A may often lack the ability to auditorily place the user in the virtual world. That is, a VR system (which may include a computer responsible for presenting video data and audio data, for ease of illustration, may be used in conjunction with a computer). Figure 4A Not shown in the example, and VR headset 400A) may not be able to support full three-dimensional immersion in hearing.
[0068] Figure 4B is a diagram illustrating an example of a wearable device 400B that can operate according to various aspects of the technology described in this disclosure. In various examples, wearable device 400B can represent a VR headset (such as VR headset 400A described above), an AR headset, an MR headset, or any other type of XR headset. Augmented reality "AR" can refer to computer-rendered images or data overlaid on the real world where the user is actually located. Mixed reality "MR" can refer to computer-rendered images or data that are world-locked to a specific location in the real world, or can refer to a variant of VR in which some computer-rendered 3D elements and some photographed reality elements are combined into an immersive experience that simulates the user's physical presence in the environment. Extended reality "XR" can represent a general term for VR, AR, and MR. More information about XR terminology can be found in Jason Peterson's document entitled "Virtual Reality, Augmented Reality, and Mixed Reality Definitions" dated July 7, 2017.
[0069] Wearable device 400B may represent other types of devices, such as watches (including so-called “smart watches”), glasses (including so-called “smart glasses”), headphones (including so-called “wireless headphones” and “smart headphones”), smart clothing, smart jewelry, etc. Regardless of whether it represents a VR device, watch, glasses and / or headphones, wearable device 400B can communicate with a computing device that supports wearable device 400B via a wired connection or a wireless connection.
[0070] In some examples, the computing device supporting wearable device 400B can be integrated within wearable device 400B, and thus, wearable device 400B can be considered to be the same device as the computing device supporting wearable device 400B. In other examples, wearable device 400B can communicate with a separate computing device that can support wearable device 400B. In this regard, the term "support" should not be understood as requiring a separate dedicated device, but rather one or more processors configured to perform various aspects of the techniques described in this disclosure can be integrated within wearable device 400B or integrated within a computing device separate from wearable device 400B.
[0071] For example, when wearable device 400B represents an example of a VR device 400B, in accordance with various aspects of the technology described in this disclosure, a separate dedicated computing device (such as a personal computer including the one or more processors) can render audio and visual content, and wearable device 400B can determine, based on the translational head movement, the translational head movement at which the dedicated computing device can render the audio content (as a speaker feed). As another example, when wearable device 400B represents smart glasses, wearable device 400B can include the one or more processors that both determine the translational head movement (by interfacing with one or more sensors on wearable device 400B) and render the speaker feed based on the determined translational head movement.
[0072] As shown, wearable device 400B includes one or more directional speakers and one or more tracking and / or recording cameras. In addition, wearable device 400B includes one or more inertial, tactile and / or health sensors, one or more eye tracking cameras, one or more high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of wearable device 400B may include durable translucent display technology and hardware.
[0073] Wearable device 400B also includes connectivity hardware, which may represent one or more network interfaces that support multi-mode connectivity, such as 4G communication, 5G communication, Bluetooth, etc. Wearable device 400B also includes one or more ambient light sensors and bone conduction transducers. In some examples, wearable device 400B may also include one or more passive and / or active cameras with wide-angle lenses and / or telephoto lenses. Although in Figure 4B Although not shown, wearable device 400B may also include one or more light-emitting diode (LED) lights. In some examples, the LED light(s) may be referred to as "super bright" LED light(s). In some embodiments, wearable device 400B may also include one or more rear-facing cameras. It will be appreciated that wearable device 400B may take on a variety of different form factors.
[0074] Additionally, tracking and recording cameras and other sensors can facilitate determination of translation distance. Figure 4B Not shown in the example of , the wearable device 400B may include other types of sensors for detecting translation distance.
[0075] Although described with respect to specific examples of wearable devices, such as those described above with respect to Figure 4B The VR device 400B discussed in the example Figure 1A and Figure 1B Other devices described in the examples, but those skilled in the art will understand that Figures 1A-4B The related description can be applied to other examples of wearable devices. For example, other wearable devices such as smart glasses can include sensors for obtaining translational head movement. As another example, other wearable devices such as smart watches can include sensors for obtaining translational movement. Therefore, the techniques described in this disclosure should not be limited to a specific type of wearable device, but rather any wearable device can be configured to perform the techniques described in this disclosure.
[0076] In any case, the audio aspects of VR can be divided into three separate categories of immersion. The first category provides the lowest level of immersion and is called three degrees of freedom (3DOF). 3DOF refers to audio presentation that handles head movement in three degrees of freedom (yaw, pitch, and roll), thereby allowing the user to look around freely in any direction. However, 3DOF cannot handle translational head movement, where the head is not centered on the optical and acoustic center of the sound field.
[0077] The second category, called 3DOF plus (3DOF+), provides three degrees of freedom (yaw, pitch, and roll) plus limited spatial translational movement due to head movement away from the optical and acoustic centers within the sound field. 3DOF+ can provide support for perceptual effects such as motion parallax, which can enhance immersion.
[0078] The third category, called six degrees of freedom (6DOF), presents audio data in a way that accounts for three degrees of freedom in terms of head movement (yaw, pitch, and roll) and also accounts for the user's translation in space (x, y, and z translation). Spatial translation can be caused by sensors that track the user's position in the physical world or through input controllers.
[0079] 3DOF rendering is the current state of the art for audio in VR. As such, audio in VR is less immersive than video, potentially reducing the overall immersion experienced by the user and introducing positioning errors (e.g., such as when auditory playback does not match or accurately correlates with the visual scene).
[0080] While 3DOF rendering is the current state of the art, more immersive audio rendering, such as 3DOF+ and 6DOF rendering, may result in increased complexity in terms of extended processor cycles, consumed memory, bandwidth, etc. In an effort to reduce complexity, the audio playback system 16A may include an interpolation device 30 ("INT device 30") that may select a subset of ambisonic coefficients 11' as ambisonic audio data 15. The interpolation device 30 may then interpolate the selected subset of ambisonic coefficients 11', apply various weights (such as weights defined by measured importance to the auditory scene—e.g., weights defined based on gain analysis or other analysis (such as directivity analysis), etc.), and then sum the weighted ambisonic coefficients 11' to form the ambisonic audio data 15. The interpolation device 30 may select a subset of the ambisonic coefficients, thereby reducing the number of operations performed when rendering the ambisonic audio data 15 (because increasing the number of ambisonic coefficients 11′ also increases the number of operations performed for rendering the loudspeaker feeds 25 from the ambisonic audio data 15).
[0081] Thus, there may be instances where high-complexity audio rendering may be important in providing an immersive experience, and other instances where low-complexity audio rendering may be sufficient to provide the same immersive experience. Furthermore, having the ability to provide high-complexity audio rendering while also supporting low-complexity audio rendering may enable devices with different processing capabilities to perform audio rendering, potentially accelerating the adoption of XR devices, as low-cost devices (which may have lower processing power than higher-cost devices) may allow more people to purchase and experience XR.
[0082] In accordance with the techniques described in this disclosure, various methods are described by which low-complexity audio rendering can be achieved while providing options for high-complexity audio rendering with additional metadata or other indications for controlling the audio rendering at the audio playback system 16A. These techniques can enable the audio playback system 16A to perform flexible rendering in terms of complexity (such as complexity defined by processor cycles, consumed memory, and / or bandwidth) while also allowing for internal and external rendering of the XR experience as defined by a boundary separating the internal area from the external area. In addition, the audio playback system 16A can configure the audio renderer 22 using metadata or other indications specified in the bitstream representing the audio data, while also generating the audio renderer 22 to address the internal area or the external area with reference to the listener position 17 relative to the boundary.
[0083] Thus, this technique can improve the operation of the audio playback system because, when configured to perform low-complexity rendering, the audio playback system 16A can reduce the amount of processor cycles, memory consumed, and / or bandwidth consumed. When performing high-complexity rendering, the audio playback system 16A can provide a more immersive XR experience, which can more realistically place the user of the audio playback system 16A in the XR experience.
[0084] like Figure 1A As shown in the example of , the audio playback system 16A can include a renderer generation unit 32, which represents a unit configured to generate or otherwise obtain one or more of the audio renderers 22 in accordance with various aspects of the techniques described in this disclosure. In some examples, the renderer generation unit 32 can perform the above-described process to generate the audio renderer 22 based on the listener position 17 and the speaker geometry 13.
[0085] However, in addition, the renderer generation unit 32 can obtain various indications 31 (e.g., syntax elements or other types of metadata) from the bitstream 21 (which can be parsed by the audio decoding device 24). As such, the sound field representation generator 302 can specify the indications 31 in the bitstream 21 before sending the bitstream 21 to the audio playback device 16A. As an example, the sound field representation generator 302 can receive the indications 31 from the content capture device 300. An operator, editor, or other individual can specify the indications 31 by interacting with the content capture device 300 or some other device such as a content editing device.
[0086] The one or more indications 31 may include an indication of the complexity of the rendering to be performed by the audio playback system 16A, an indication of the opacity for rendering secondary sources present in the ambisonic coefficients s, and / or an indication of a buffer distance around an inner region, where the rendering is interpolated between the inner rendering and the outer rendering. The indication indicating complexity may indicate complexity as low complexity or high complexity (as a Boolean value, where "true" is used for low complexity and "false" is used for high complexity). The indication indicating opacity may indicate opacity or non-opaqueness (as a Boolean value, where "true" indicates opacity and "false" indicates non-opaqueness, although opacity may be defined as a floating point number with a value between zero and one). The indication indicating a distance buffer may indicate a distance as a value.
[0087] The sound field representation generator 302 may also specify a boundary separating the inner region from the outer region in the bitstream 21. As mentioned above, the sound field representation generator 302 may also specify one or more indications 31 for controlling the rendering of the ambisonic coefficients 11 for the inner region and the outer region. The sound field representation generator 302 may output the bitstream 21 for delivery (either in near real time via network streaming, etc., or at a later time as described above).
[0088] The audio playback system 16A may obtain the bitstream 21 and call the audio decoding device 24 to decompress the bitstream to obtain the ambisonic audio coefficients 11′, and parse the indication 31 from the bitstream 21. The audio decoding device 24 may output the indication 31 along with the indication of the boundary to the renderer generation unit 32. The audio playback system 16A may also interface with the tracking device 306 to obtain the listener position 17, wherein the boundary, the listener position 17, and the indication 31 are provided to the renderer generation unit 32.
[0089] In this way, the renderer generation unit 32 may obtain an indication of a boundary that separates the inner area from the outer area.The renderer generation unit 32 may also obtain a listener position 17 that indicates the position of the content consumer device 14 relative to the inner area.
[0090] The renderer generation unit 32 may then, based on the boundaries and the listener position 17, obtain a current renderer 22 to use when rendering the ambisonics audio data 15 to the one or more speaker feeds 25. The current renderer 22 may be configured to render the ambisonics audio data 25 for the inner zone (and thereby operate as an inner renderer) or to render the audio data for the outer zone (and thereby operate as an outer renderer).
[0091] Determining whether to configure the current renderer 22 as an internal renderer or an external renderer may depend on where the content consumer device 14 is located relative to the boundary in the XR scene. For example, when the content consumer device 14 is in the XR scene and outside the internal area defined by the boundary according to the listener position 17, the renderer generation unit 32 may configure the current renderer 22 to operate as an external renderer. When the content consumer device 14 is in the XR scene and within the internal area defined by the boundary according to the listener position 17, the renderer generation unit 32 may configure the current renderer 22 to operate as an internal renderer. The renderer generation unit 32 may output the current renderer 22, where the audio playback system 16A may apply the current renderer 22 to the ambisonic audio data 15 to obtain the speaker feed 25.
[0092] The following will be Figure 2 and 3More information about the indication regarding complexity, the indication regarding opacity, and the indication regarding the distance buffer is described with reference to examples of FIG.
[0093] Figure 2 is a diagram illustrating an example of low-complexity rendering for an extended reality (XR) scene according to various aspects of the technology described in this disclosure. Figure 2 As shown in the example of , XR scene 200 includes an operator 202 operating content consumer device 14A (not shown for ease of illustration). XR scene 200 also includes a boundary 204 that separates an inner area 206 from an outer area 208.
[0094] Despite Figure 2 While a single boundary 204 is shown in the example of , the XR scene 200 may include multiple boundaries that separate different interior regions from an exterior region 208. Furthermore, although shown as a single boundary 204, boundaries may exist within other boundaries, overlap other boundaries, etc. When boundaries exist within other boundaries, the interior region defined by the larger boundary may operate as an outer boundary (for presentation purposes) relative to the presentation of the interior region defined by the boundary within the outer boundary.
[0095] In any case, first assuming that the operator 202 is in the outer region 208 relative to the boundary 204, the renderer generation unit 32 (of the content consumer device 14A) may first determine whether the indication of complexity indicates high complexity or low complexity. For purposes of illustration, assuming that the indication of complexity indicates low complexity, the renderer generation unit 32 may determine a first distance between the listener position 17 and the center 210 of the inner region 206 (as an example, calculated based on the boundary 204, which may be represented as a shape, a list of points, a spline, or any other geometric representation). The renderer generation unit 32 may then determine a second distance between the boundary 204 and the center 210.
[0096] The renderer generation unit 32 may then compare the first distance with the second distance to determine that the operator 202 exists outside the boundary 204. That is, when the first distance is greater than the second distance, the renderer generation unit 32 may determine that the operator 202 is located outside the boundary 204. For the low complexity configuration, the renderer generation unit 32 may generate the current renderer 22 to render the ambisonics audio data 15 for the interior area 206 such that the sound field represented by the ambisonics audio data 15 originates from the center 210 of the interior area 206. The renderer generation unit 32 may render the ambisonics audio data 15 to be positioned theta (θ) degrees from the direction the operator 202 is facing.
[0097] Making the sound field appear to originate from a single point, i.e., the center 210 in this example, can reduce complexity in terms of processing cycles, memory, and bandwidth consumption, as this can result in fewer speaker feeds being used to represent the sound field (and potentially reducing panning, mixing, and other audio operations), while also potentially maintaining an immersive experience. Further reductions in processor cycles, memory, and bandwidth consumption can be achieved when the renderer generation unit 32 utilizes only a single ambisonic coefficient of the ambisonic audio data 15 (e.g., an ambisonic coefficient corresponding to a spherical basis function with order zero, which represents the gain of the sound field and does not provide much spatial information (if any) and therefore does not require complex rendering) instead of processing multiple ambisonic coefficients from the ambisonic audio data 15.
[0098] Next, assume that the operator 202 moves into the interior area 206. The renderer generation unit 32 may receive the updated listener position 17 and perform the same processing described above to determine that (because the first distance is less than the second distance) the operator 202 is located in the interior area 206. For a low complexity indication and in response to determining that the operator 202 is present in the interior area, the renderer generation unit 206 may output an updated current renderer 22 that is configured to render the ambisonic audio data 15 such that the sound field represented by the ambisonic audio data 15 appears in the entire interior area 206 (which may be referred to as a full or normal rendering because all of the ambisonic audio data 15 may be rendered such that audio sources within the sound field are accurately placed around the operator 202).
[0099] Thus, when the interior field is specified to be rendered using a low complexity renderer for low latency applications or for artistic intent, the buffer distance or opacity properties are not used when generating the current renderer 22. In this case, when the listener 202 (which is another way of referring to the operator 202) is outside the interior field area 206 (which is another way of referring to the interior field 206), the W ambisonic reverb channels (corresponding to the spherical harmonic function α) are played from the center 210 of the interior field area 206 toward the listener 202. 00 When the listener 202 is located within the infield area 206, the ambisonic sound field is typically played back from all directions.
[0100] Figure 3 is a diagram illustrating an example of high complexity rendering of an XR scene including a distance buffer according to various aspects of the techniques described herein. The XR scene 220 is similar to Figure 2In the example of the XR scene 200 shown in FIG, it is assumed that the indication of complexity indicates high complexity. In response to the indication indicating high complexity, the renderer generation unit 32 may utilize the indication of the distance buffer (which is shown as “distance buffer 222”) to generate a transition region 224 (which may also be referred to as “interpolation region 224”).
[0101] First, it is assumed that the operator 202 exists in the external area 208. The renderer generation unit 32 can use the above Figure 2 In the manner described in the example of FIG. , the renderer generation unit 32 determines that the operator 202 is present in the outside area 208. In response to determining that the operator 202 is present in the outside area 208, the renderer generation unit 32 may next determine whether the indication of complexity indicates high complexity or low complexity. For the purpose of illustration, assuming that the indication of complexity indicates high complexity, the renderer generation unit 32 may determine whether the indication of opacity indicates opacity or non-opaqueness.
[0102] When the indication of opacity indicates opacity, the renderer generation unit 32 may configure the current renderer 22 to discard secondary audio sources present in the sound field represented by the ambisonic audio data 15 that are not directly in the line of sight of the operator 202. That is, the renderer generation unit 32 may configure the current renderer 22 based on the listener position 17 and the boundary 204 to exclude the addition of secondary audio sources indicated as not directly in the line of sight by the listener position 17. When the indication of opacity indicates non-opaqueness, the renderer generation unit 32 returns to normal rendering that takes into account all secondary sources.
[0103] When the current renderer 22 is configured for external presentation using high complexity, the renderer generation unit 32 may configure the current renderer 22 to render the ambisonic audio data 15 in all cases (e.g., occlusive or non-occlusive) such that the sound field represented by the ambisonic audio data 15 is spread out according to the distance between the listener position 17 and the boundary 204. In the example of FIG1 , the spread angle is expressed as theta (θ) degrees. This distance is illustrated by two dashed lines 226A and 226B, thereby generating a spread of θ degrees.
[0104] When the operator 202 moves into the transition region 224, in response to determining that the listener position 17 is within the distance buffer 222 of the boundary 204, the renderer generation unit 32 can update the current renderer 22 to interpolate between the external renderer and the internal renderer. An example of interpolation can be (1-a)*internal_rendering+a*external_rendering, where a is a fraction based on the proximity of the listener 202 to the sound field boundary 204 (which is another way of referring to the boundary 204). The audio playback system 16A can then apply the updated current renderer 22 to obtain one or more updated speaker feeds 25.
[0105] When the operator 202 moves completely into the inner area 206, the renderer generation unit 32 may generate the current renderer 22 for normal rendering. That is, the renderer generation unit 32 may generate the current renderer 22 to use all ambisonic coefficients of the ambisonic audio data 15 and in a manner that correctly places each audio source in the sound field (e.g., not positioning all sources at the same location, e.g., Figure 2 ) to fully present the ambisonic audio data 15 existing within the inner area 206 .
[0106] That is, the interior sound field 206 represented in the ambisonic format allows for secondary sources on the boundaries, and these secondary sources contribute to the sound at the listener 202 according to Huygen's principle. When the opacity property is true, the renderer generation unit 32 may not add the contribution of secondary sources to which the listener 202 does not have a direct line of sight.
[0107] In a high complexity renderer, the listener 202 can hear the interior sound field 206 as an expanded source based on the listener's distance from the interior sound field 206. When the listener 202 moves from the exterior sound field 208 (which is another way of referring to the exterior area 208) to the interior sound field 206 (which is another way of referring to the interior area 206), the rendering can change and the movement can be smooth. The Buffer_Distance attribute specifies the distance when interpolation between the renderings for the exterior and interior listeners 202 is performed. An example interpolation scheme includes (1-a)*internal_rendering+a*external_rendering. This variable can represent a score based on the listener's proximity to the sound field boundary 204.
[0108] For example, an orchestra can be represented as an interior ambisonic sound field. In this case, listener 202 should hear contributions from all instruments, so opacity is set to false. If the interior field represents a crowd and the intent is to change the listening experience as the listener moves around the outside of boundary 204, opacity can be set to true.
[0109] Thus, an opacity attribute, an interpolation buffer distance attribute, and a complexity attribute (which is another way of referring to indication 31) may be specified to support rendering of interior ambisonic sound fields to the MPEG-I encoder input format. Several usage scenarios may illustrate the usefulness of these attributes. These attributes may facilitate control of the rendering of interior sound fields at the listener's position for 6DOF (and other) use cases.
[0110] Figure 1B is a block diagram illustrating another example system 100 configured to perform various aspects of the techniques described in this disclosure. System 100 is similar to Figure 1A The system 10 shown in FIG. Figure 1A The audio renderer 22 shown in FIG. 1 is replaced by a binaural renderer 102 which is capable of performing binaural rendering using one or more HRTFs or other functions which can be rendered to left and right speaker feeds 103 .
[0111] The audio playback system 16B may output the left and right speaker feeds 103 to headphones 104, which may represent another example of a wearable device and which may be coupled to additional wearable devices to facilitate reproduction of the sound field, such as a watch, the VR headsets mentioned above, smart glasses, smart clothing, a smart ring, a smart bracelet, or any other type of smart jewelry (including a smart necklace), etc. The headphones 104 may be coupled to the additional wearable devices wirelessly or via a wired connection.
[0112] Additionally, the headset 104 may be connected via a wired connection (such as a standard 3.5 mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other form of wired connection) or wirelessly (such as via Bluetooth TM The headset 104 may be coupled to the audio playback system 16 via a wireless connection, a wireless network connection, etc. The headset 104 may recreate the sound field represented by the ambisonics coefficients 11 based on the left and right speaker feeds 103. The headset 104 may include a left earphone speaker and a right earphone speaker powered (or, in other words, driven) by the corresponding left and right speaker feeds 103.
[0113] Although for Figure 4A and 4BThe techniques are described with reference to a VR device shown in the example of FIG, but these techniques can be performed by other types of wearable devices, including watches (such as so-called "smart watches"), glasses (such as so-called "smart glasses"), headphones (including wireless headphones coupled via a wireless connection, or smart headphones coupled via a wired or wireless connection), and any other type of wearable device. Thus, these techniques can be performed by any type of wearable device through which a user can interact with the wearable device while the user is wearing the wearable device.
[0114] Figure 5A and 5B is a diagram illustrating an example system that can perform various aspects of the techniques described in this disclosure. Figure 5A An example is described in which source device 12 further includes camera 200. Camera 200 may be configured to capture video data and provide the captured raw video data to content capture device 300. Content capture device 300 may provide the video data to another component of source device 12 for further processing into viewport-divided portions.
[0115] exist Figure 5A In the example of , the content consumer device 14 also includes a wearable device 800. It will be understood that in various embodiments, the wearable device 800 can be included in the content consumer device 14 or externally coupled to the content consumer device 14. Figure 4A and 4B As discussed, wearable device 800 includes display hardware and speaker hardware for outputting video data (eg, as associated with various viewports) and for presenting audio data.
[0116] Figure 5B Shown with Figure 5A The example shown is similar to the example shown, except that Figure 5A The audio renderer 22 is shown replaced with a binaural renderer 102 that can perform binaural rendering using one or more HRTFs or other functions that can be rendered to left and right speaker feeds 103. The audio playback system 16 can output the left and right speaker feeds 103 to headphones 104.
[0117] The headset 104 may be connected via a wired connection (such as a standard 3.5 mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other form of wired connection) or wirelessly (such as via Bluetooth TMThe headset 104 may be coupled to the audio playback system 16 via a wireless connection, a wireless network connection, etc. The headset 104 may recreate the sound field represented by the ambisonics coefficients 11 based on the left and right speaker feeds 103. The headset 104 may include a left earphone speaker and a right earphone speaker powered (or, in other words, driven) by the corresponding left and right speaker feeds 103.
[0118] Figure 6A When performing various aspects of the technology described in this disclosure Figure 1A and 1B . The audio playback device 16C may represent an example of the audio playback device 16A and / or the audio playback device 16B. The audio playback system 16 may include an audio decoding device 24 in combination with a 6DOF audio renderer 22A, which may represent an example of the audio playback device 16C. Figure 1A An example of an audio renderer 22 is shown in the example of .
[0119] The audio decoding device 24 may include a low-latency decoder 900A, an audio decoder 900B, and a local audio buffer 902. The low-latency decoder 900A may process the XR audio bitstream 21A to obtain an audio stream 901A. The low-latency decoder 900A may perform relatively low-complexity decoding (compared to the audio decoder 900B) to facilitate low-latency reconstruction of the audio stream 901A. The audio decoder 900B may perform relatively high-complexity decoding (compared to the audio decoder 900A) on the audio bitstream 21B to obtain the audio stream 901B. The audio decoder 900B may perform audio decoding compliant with the MPEG-H 3D Audio coding standard. The local audio buffer 902 may represent a unit configured to buffer local audio content. The local audio buffer 902 may output the local audio content as an audio stream 903.
[0120] The bitstream 21 (including one or more of the XR audio bitstream 21A and / or the audio bitstream 21B) may also include XR metadata 905A (which may include the microphone position information described above) and 6DOF metadata 905B (which may specify various parameters related to 6DOF audio rendering). The 6DOF audio renderer 22A may obtain the audio streams 901A and / or 901B and / or the audio stream 903 from the buffer 910, along with the XR metadata 905A, the 6DOF metadata 905B, the listener position 17, and the HRTF 23, and render the speaker feeds 25 and / or 103 based on the listener position and the microphone position. Figure 6A In the example of , the 6DOF audio renderer 22A includes an interpolation device 30A, which can perform various aspects of the audio stream selection and / or interpolation techniques described in more detail above to facilitate 6DOF audio rendering. Figure 6AIn the example of , the 6DOF audio renderer 22A also includes a controller 920, which can pass appropriate metadata and audio signals to the interpolation device 30A. The interpolation device 30A can interpolate the ambisonic coefficients from two or more sources in the buffer 910, or interpolate the binauralized audio from the audio object renderer and / or the 6DOF audio renderer 22A. Although shown as part of 6DOF, in some examples, the controller 920 can be located elsewhere in the audio playback device 16C. In some examples, any one of the low-latency decoder 900A, audio decoder 900B, local audio buffer 902, buffer 910, and 6DOF audio renderer 22A can be implemented in one or more processors.
[0121] Figure 6B When performing various aspects of the technology described in this disclosure Figure 1A and 1B A block diagram of an audio playback device is shown in the example. Figure 6B An example audio playback device 16D is similar to Figure 6A 16C, however, the audio playback device 16D also includes an audio object renderer 912 and a 3DOF audio renderer 914. Each of the audio object renderer 912, the 3DOF audio renderer 914, and the 6DOF audio renderer 22B can receive the listener position 17 and the HRTF 23. In this example, the output of the audio object renderer 912, the 3DOF audio renderer 914, or the output of the 6DOF audio renderer can be sent to a binauralizer 916, which can perform binaural rendering. In some examples, each of the audio object renderer 912, the 3DOF audio renderer 914, and the 6DOF audio renderer 22B can output ambisonic sound. The output of the binauralizer 916 can be sent to the interpolation device 30B. The interpolation device 30B may include a controller 918. Although a single output from the audio decoding device 24 is shown, in some examples, the low-latency decoder 900A, the audio decoder 900B, and the local audio buffer 902 may each have separate connections to each of the audio object renderer 912, the 3DOF audio renderer 914, and the 6DOF audio renderer 22A. Figure 6B In the example of , the interpolation device 30B can interpolate the binauralized audio from the binauralizer 916. Figure 6BIn the example shown, the interpolation device 30B also includes a controller 918 that can control the functions of the interpolation device 30B. Although shown as part of the interpolation device 30, in some examples, the controller 918 can be located elsewhere in the audio playback device 16D. In some examples, any of the low-latency decoder 900A, audio decoder 900B, local audio buffer 902, buffer 910, audio object renderer 912, 3DOF audio renderer 914, 6DOF audio renderer 22B, binauralizer 916, and interpolation device 30B can be implemented in one or more processors.
[0122] Figure 6C When performing various aspects of the technology described in this disclosure Figure 1A and 1B A block diagram of an audio playback device is shown in the example. Figure 6C An example audio playback device 16E is similar to Figure 6B 916, however, instead of the audio object renderer 912, the 3DOF audio renderer 914, or the 6DOF audio renderer 22A sending their output to the binauralizer 916, the audio object renderer 912, the 3DOF audio renderer 914, or the 6DOF audio renderer 22A sends their output to the interpolation device 30B, which in turn sends its output to the binauralizer 916. In some examples, each of the audio object renderer 912, the 3DOF audio renderer 914, and the 6DOF audio renderer 22B can output ambisonic sound. Figure 6C In the example of , the interpolation device 30B may interpolate ambisonic coefficients from two or more of the audio object renderer 912, the 3DOF audio renderer 914, or the 6DOF audio renderer 22B. Figure 6C In the example shown, the interpolation device 30B also includes a controller 918 that can control the functions of the interpolation device 30B. Although shown as part of the interpolation device 30B, in some examples, the controller 918 can be located elsewhere in the audio playback device 16E. In some examples, any of the low-latency decoder 900A, audio decoder 900B, local audio buffer 902, buffer 910, audio object renderer 912, 3DOF audio renderer 914, 6DOF audio renderer 22B, binauralizer 916, and interpolation device 30B can be implemented in one or more processors.
[0123] Figure 6D When performing various aspects of the technology described in this disclosure Figure 1A and 1B A block diagram of an audio playback device is shown in the example. Figure 6DAn example audio playback device 16F is similar to Figure 6C However, the audio playback device 16G does not include the binauralizer 916. In some examples, each of the audio object renderer 912, the 3DOF audio renderer 914, and the 6DOF audio renderer 22B can output ambisonic sound. Figure 6D In the example of , the interpolation device 30B may interpolate ambisonic coefficients from two or more of the audio object renderer 912, the 3DOF audio renderer 914, or the 6DOF audio renderer 22B, or interpolate binauralized audio from the audio object renderer 912, the 3DOF audio renderer 914, and / or the 6DOF audio renderer 22B. Figure 6D In the example shown, the interpolation device 30B also includes a controller 918 that can control the functions of the interpolation device 30B. Although shown as part of the interpolation device 30, in some examples, the controller 918 can be located elsewhere in the audio playback device 16F. In some examples, any of the low-latency decoder 900A, the audio decoder 900B, the local audio buffer 902, the buffer 910, the audio object renderer 912, the 3DOF audio renderer 914, the 6DOF audio renderer 22B, the binauralizer 916, and the interpolation device 30B can be implemented in one or more processors.
[0124] Figure 6E When performing various aspects of the technology described in this disclosure Figure 1A and 1B A block diagram of an audio playback device is shown in the example. Figure 6E An example audio playback device 16G is similar to Figure 6D In some examples, each of the audio object renderer 912, the 3DOF audio renderer 914, and the 6DOF audio renderer 22C can output ambisonic sound. Figure 6E In the example of , the interpolation device 30B may interpolate ambisonic coefficients from two or more of the audio object renderer 912, the 3DOF audio renderer 914, or the 6DOF audio renderer 22B, or interpolate binauralized audio from the audio object renderer 912, the 3DOF audio renderer 914, and / or the 6DOF audio renderer 22B. Figure 6EIn the example shown, the interpolation device 30B also includes a controller 918 that can control the functions of the interpolation device 30B. Although shown as part of the interpolation device 30B, in some examples, the controller 918 can be located elsewhere in the audio playback device 16G. In some examples, any of the low-latency decoder 900A, audio decoder 900B, local audio buffer 902, buffer 910, audio object renderer 912, 6DOF audio renderer 22C, and interpolation device 30B can be implemented in one or more processors.
[0125] Figure 6F When performing various aspects of the technology described in this disclosure Figure 1A and 1B A block diagram of an audio playback device is shown in the example. Figure 6F An example audio playback device 16H is similar to Figure 6A The audio playback device 16C of FIG. 1 is a block diagram of an audio decoder 900C, however, the audio decoder 900C includes an audio object renderer 912, an HOA renderer 922, and a binauralizer 916, and the interpolation device 30C is an independent device and includes a controller 918 and a 6DOF audio renderer 22A. Figure 6F In the example of , the interpolation device 30C may interpolate ambisonic coefficients from two or more sources in the buffer 910 or interpolate binauralized audio from the binauralizer 916. Figure 6F In the example of , the interposer device 30C also includes a controller 918 that can control the functions of the interposer device 30C. Although shown as part of the interposer device 30C, in some examples, the controller 918 can be located elsewhere in the audio playback device 16H. Figures 6A-6F Several examples of audio playback devices have been described in , but these include Figures 6A-6F Other examples of other combinations of various elements may fall within the scope of this disclosure. In some examples, any one of the low-delay decoder 900A, audio decoder 900C, local audio buffer 902, buffer 910, and interpolation device 30B may be implemented in one or more processors.
[0126] Figure 7An example of a wireless communication system 100 supporting audio streaming according to aspects of the present disclosure is shown. The wireless communication system 100 includes a base station 105, a UE 115, and a core network 130. In some examples, the wireless communication system 100 can be a Long Term Evolution (LTE) network, an Advanced LTE (LTE-A) network, an LTE-A Pro network, or a New Radio (NR) network. In some cases, the wireless communication system 100 can support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low-latency communication, or communication with low-cost and low-complexity devices.
[0127] The base station 105 can communicate wirelessly with the UE 115 via one or more base station antennas. The base station 105 described herein may include or may be referred to by those skilled in the art as a base transceiver station, a radio base station, an access point, a radio transceiver, a Node B, an eNodeB (eNB), a next generation Node B or a giga Node B (any of which may be referred to as a gNB), a Home NodeB, a Home eNodeB, or some other appropriate terminology. The wireless communication system 100 may include different types of base stations 105 (e.g., macro cell base stations or small cell base stations). The UE 115 described herein is capable of communicating with various types of base stations 105 and network devices, including macro eNBs, small cell eNBs, gNBs, relay base stations, etc.
[0128] Each base station 105 may be associated with a particular geographic coverage area 110 in which it supports communications with various UEs 115. Each base station 105 may provide communication coverage for the respective geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 may utilize one or more carriers. The communication link 125 shown in the wireless communication system 100 may include an uplink transmission from the UE 115 to the base station 105, or a downlink transmission from the base station 105 to the UE 115. Downlink transmissions may also be referred to as forward link transmissions, while uplink transmissions may also be referred to as reverse link transmissions.
[0129] The geographic coverage area 110 of a base station 105 can be divided into sectors that constitute a portion of the geographic coverage area 110, and each sector can be associated with a cell. For example, each base station 105 can provide communication coverage for a macrocell, a small cell, a hotspot, or other types of cells, or various combinations thereof. In some examples, the base stations 105 can be mobile and, therefore, provide communication coverage for mobile geographic coverage areas 110. In some examples, different geographic coverage areas 110 associated with different technologies can overlap, and overlapping geographic coverage areas 110 associated with different technologies can be supported by the same base station 105 or different base stations 105. The wireless communication system 100 can include, for example, a heterogeneous LTE / LTE-A / LTE-A Pro or NR network, in which different types of base stations 105 provide coverage for various geographic coverage areas 110.
[0130] UE 115 can be dispersed throughout the wireless communication system 100, and each UE 115 can be stationary or mobile. UE 115 can also be referred to as a mobile device, a wireless device, a remote device, a handheld device, or a subscriber device, or some other appropriate term, where "device" can also be referred to as a unit, a station, a terminal, or a client. UE 115 can also be a personal electronic device, such as a cellular phone, a personal digital assistant (PDA), a tablet computer, a laptop computer, or a personal computer. In the examples of the present disclosure, UE 115 can be any one of the audio sources described in the present disclosure, including a VR headset, an XR headset, an AR headset, a vehicle, a smart phone, a microphone, a microphone array, or any other device including a microphone or capable of sending a captured and / or synthesized audio stream. In some examples, the synthesized audio stream can be an audio stream stored in a memory or previously created or synthesized. In some examples, UE 115 may also refer to a wireless local loop (WLL) station, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, or an MTC device, etc., which may be implemented in various articles of manufacture such as appliances, vehicles, meters, etc.
[0131] Some UEs 115 (such as MTC or IoT devices) may be low-cost or low-complexity devices and may provide automatic communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC may refer to data communication technology that allows devices to communicate with each other or with a base station 105 without human intervention. In some examples, M2M communication or MTC may include communications from devices that exchange and / or use audio metadata indicating privacy restrictions and / or password-based privacy data to switch, block, and / or disable various audio streams and / or audio sources, as will be described in more detail below.
[0132] In some cases, UE 115 can also communicate directly with other UEs 115 (e.g., using a peer-to-peer (P2P) or device-to-device (D2D) protocol). One or more UEs 115 in a group of UEs 115 utilizing D2D communication may be within the geographic coverage area 110 of base station 105. Other UEs 115 in the group may be outside the geographic coverage area 110 of base station 105 or unable to receive transmissions from base station 105 for other reasons. In some cases, a group of UEs 115 communicating via D2D communication may utilize a one-to-many (1:M) system, in which each UE 115 transmits to each other UE 115 in the group. In some cases, base station 105 facilitates the scheduling of resources for D2D communication. In other cases, D2D communication is performed between UEs 115 without involving base station 105.
[0133] The base stations 105 can communicate with the core network 130 and with each other. For example, the base stations 105 can interface with the core network 130 via a backhaul link 132 (e.g., via an S1, N2, N3, or other interface). The base stations 105 can communicate with each other directly (e.g., directly between the base stations 105) or indirectly (e.g., via the core network 130) on a backhaul link 134 (e.g., via an X2, Xn, or other interface).
[0134] In some cases, the wireless communication system 100 can utilize both licensed and unlicensed radio spectrum bands. For example, the wireless communication system 100 can employ licensed assisted access (LAA), LTE unlicensed (LTE-U) radio access technology, or NR technology in an unlicensed band such as the 5 GHz ISM band. When operating in an unlicensed radio spectrum band, wireless devices such as base stations 105 and UEs 115 can employ a listen before talk (LBT) process to ensure that the frequency channel is idle before sending data. In some cases, operations in the unlicensed band can be based on a carrier aggregation configuration in combination with component carriers operating in the licensed band (e.g., LAA). Operations in the unlicensed spectrum can include downlink transmissions, uplink transmissions, peer to peer transmissions, or a combination of these. Duplexing in the unlicensed spectrum can be based on frequency division duplexing (FDD), time division duplexing (TDD), or a combination of both.
[0135] Figure 8 It shows Figure 1A , a flow chart of example operations of a source device as shown in FIG. 1 in performing various aspects of the techniques described in this disclosure. Source device 12 may obtain a bitstream 21 representing scene-based audio data 11 in the manner described above (800). Sound field representation generator 302 of source device 12 may specify a boundary separating an interior region from an exterior region in bitstream 21 (802).
[0136] As mentioned above, the sound field representation generator 302 may also specify one or more indications 31 for controlling the rendering of the ambisonic coefficients 11 for the inner and outer zones (804). The sound field representation generator 302 may output the bitstream 21 for delivery (either in near real time via network streaming, etc., or later as described above) (806).
[0137] Figure 9 It shows Figure 1A , in performing various aspects of the techniques described in this disclosure. The audio playback system 16A may obtain a bitstream 21 and invoke an audio decoding device 24 to decompress the bitstream to obtain ambisonic audio coefficients 11′, and parse indications 31 from the bitstream 21. The audio decoding device 24 may output the indications 31 along with indications of boundaries to a renderer generation unit 32. The audio playback system 16A may also interface with a tracking device 306 to obtain a listener position 17, wherein the boundaries, listener position 17, and indications 31 are provided to the renderer generation unit 32.
[0138] In this manner, renderer generation unit 32 may obtain an indication of a boundary separating the inner region from the outer region 950. Renderer generation unit 32 may also obtain listener position 17 indicating a position of content consumer device 14 relative to the inner region 952.
[0139] The renderer generation unit 32 may then obtain a current renderer 22 to use when rendering the ambisonics audio data 15 to one or more speaker feeds 25 based on the boundaries and the listener position 17. The current renderer 22 may be configured to render the ambisonics audio data 25 for the inner zone (thereby operating as an inner renderer) or to render the audio data for the outer zone (thereby operating as an outer renderer) (954). The renderer generation unit 32 may output the current renderer 22, where the audio playback system 16A may apply the current renderer 22 to the ambisonics audio data 15 to obtain the speaker feeds 25 (956).
[0140] In this regard, various aspects of the technology described in this disclosure may implement the following provisions.
[0141] Item 1A. A device for processing one or more audio streams, the device comprising: one or more processors configured to: obtain an indication of a boundary separating an inner region from an outer region; obtain a listener position indicating a position of the device relative to the inner region; based on the boundary and the listener position, obtain a current renderer that is either an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; apply the current renderer to the audio data to obtain one or more speaker feeds; and a memory coupled to the one or more processors and configured to store the one or more speaker feeds.
[0142] Clause 2A. A device according to clause 1A, wherein the one or more processors are configured to: determine a first distance between the listener position and the center of the interior area; determine a second distance between the boundary and the center of the interior area; and obtain the current renderer based on the first distance and the second distance.
[0143] Clause 3A. The apparatus of any combination of clauses 1A and 2A, wherein the audio data comprises ambisonic audio data associated with spherical basis functions having order zero, and wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the interior region.
[0144] Clause 4A. The apparatus of any combination of clauses 1A and 2A, wherein the audio data comprises ambisonic audio data associated with spherical basis functions having order zero, and wherein the interior renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data appears throughout the interior region.
[0145] Clause 5A. The apparatus of any combination of clauses 1A and 2A, wherein the audio data comprises ambisonic audio data representing a primary audio source and a secondary audio source, wherein the one or more processors are further configured to obtain an indication of an opacity of the secondary audio source, and wherein the one or more processors are configured to obtain the current renderer based on the listener position, the boundary, and the indication.
[0146] Clause 6A. The apparatus of Clause 5A, wherein the one or more processors are configured to obtain the indication of acoustic opacity of the secondary audio source from a bitstream representing the audio data.
[0147] Clause 7A. A device according to any combination of clauses 5A and 6A, wherein the one or more processors are further configured to: when the indication of the opacity is enabled and based on the listener position and the boundary, obtain the current renderer, which excludes the addition of the secondary audio sources that the listener position indicates are not directly in line of sight.
[0148] Clause 8A. The apparatus of any combination of clauses 5A-7A, wherein the external renderer is configured to render the audio data such that a sound field represented by the audio data is spread out according to a distance between the listener position and the boundary.
[0149] Clause 9A. An apparatus according to any combination of clauses 5A-8A, wherein the one or more processors are further configured to: in response to determining that the listener position is within a buffer distance from the boundary, update the current renderer to interpolate between the external renderer and the internal renderer to obtain an updated current renderer; and apply the current renderer to the audio data to obtain one or more updated speaker feeds.
[0150] Clause 10A. The device of Clause 9A, wherein the one or more processors are further configured to obtain the indication of the buffer distance from a bitstream representing the audio data.
[0151] Clause 11A. An apparatus according to any combination of clauses 1A to 10A, wherein the one or more processors are further configured to: obtain an indication of the complexity of the current renderer from a bitstream representing the audio data, and wherein the one or more processors are configured to: obtain the current renderer based on the boundary, the listener position and the indication of the complexity.
[0152] Clause 12A. The apparatus of clause 11A, wherein the audio data comprises ambisonic audio data associated with spherical basis functions having order zero, and wherein the one or more processors are configured to: when the listener position is outside the boundary and when the indication of the complexity indicates low complexity, obtain the external renderer, such that the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the interior region.
[0153] Clause 13A. An apparatus according to clause 11A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having a zero order, and wherein the one or more processors are configured to: when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, obtain the external renderer, such that the external renderer is configured to render the audio data such that the sound field represented by the audio data is expanded according to the distance between the listener position and the boundary.
[0154] Clause 14A. A method of processing one or more audio streams, the method comprising: obtaining, by one or more processors, an indication of a boundary separating an inner region from an outer region; obtaining, by one or more processors, a listener position indicating a position of the device relative to the inner region; obtaining, by the one or more processors, a current renderer that is an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region based on the boundary and the listener position; and applying, by the one or more processors, the current renderer to the audio data to obtain one or more speaker feeds.
[0155] Clause 15A. A method according to Clause 14A, wherein obtaining the current renderer includes: determining a first distance between the listener position and the center of the inner area; determining a second distance between the boundary and the center of the inner area; and obtaining the current renderer based on the first distance and the second distance.
[0156] Clause 16A. The method of any combination of clauses 14A and 15A, wherein the audio data comprises ambisonic audio data associated with spherical basis functions having order zero, and wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the interior region.
[0157] Clause 17A. A method according to any combination of clauses 14A and 15A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having order zero, and wherein the interior renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data appears throughout the interior region.
[0158] Clause 18A. A method according to any combination of clauses 14A and 15A, wherein the audio data includes ambisonic audio data representing a primary audio source and a secondary audio source, wherein the method further comprises: obtaining an indication of an opacity of the secondary audio source, and wherein obtaining the current renderer includes: obtaining the current renderer based on the listener position, the boundary, and the indication.
[0159] Clause 19A. The method of Clause 18A, wherein obtaining the indication of the acoustic opacity comprises obtaining the indication of the acoustic opacity of the secondary audio source from a bitstream representing the audio data.
[0160] Clause 20A. The method of any combination of clauses 18A and 19A, further comprising: when the indication of the opacity is enabled and based on the listener position and the boundary, obtaining the current renderer, the current renderer excluding the addition of the secondary audio sources that the listener position indicates are not directly in line of sight.
[0161] Clause 21A. The method of any combination of clauses 18A-20A, wherein the external renderer is configured to render the audio data such that a sound field represented by the audio data is spread out according to a distance between the listener position and the boundary.
[0162] Clause 22A. The method of any combination of clauses 18A-21A, further comprising: in response to determining that the listener position is within a buffer distance from the boundary, updating the current renderer to interpolate between the external renderer and the internal renderer to obtain an updated current renderer; and applying the current renderer to audio data to obtain one or more updated speaker feeds.
[0163] Clause 23A. The method of Clause 22A, further comprising obtaining an indication of the buffer distance from a bitstream representing the audio data.
[0164] Clause 24A. The method of any combination of clauses 14A to 23A, further comprising obtaining an indication of the complexity of the current renderer from a bitstream representing the audio data, and wherein obtaining the current renderer comprises obtaining the current renderer based on the boundary, the listener position, and the indication of the complexity.
[0165] Clause 25A. A method according to clause 24A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having a zeroth order, and wherein obtaining the current renderer includes: when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, obtaining the outer renderer, such that the outer renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner region.
[0166] Clause 26A. A method according to clause 24A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having a zero order, and wherein obtaining the current renderer includes: when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, obtaining the external renderer, such that the external renderer is configured to render the audio data such that the sound field represented by the audio data is expanded according to the distance between the listener position and the boundary.
[0167] Clause 27A. A device configured to process one or more audio streams, the device comprising: a unit for obtaining an indication of a boundary separating an inner region from an outer region; a unit for obtaining a listener position indicating a position of the device relative to the inner region; a unit for obtaining, based on the boundary and the listener position, a current renderer that is an inner renderer configured to render audio data for the inner region or an outer renderer configured to render audio data for the outer region; and a unit for applying the current renderer to the audio data to obtain one or more speaker feeds.
[0168] Clause 28A. An apparatus according to Clause 27A, wherein the unit for obtaining the current renderer includes: a unit for determining a first distance between the listener position and the center of the inner area; a unit for determining a second distance between the boundary and the center of the inner area; and a unit for obtaining the current renderer based on the first distance and the second distance.
[0169] Clause 29A. An apparatus according to any combination of clauses 27A and 28A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having order zero, and wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the interior region.
[0170] Clause 30A. An apparatus according to any combination of clauses 27A and 28A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having order zero, and wherein the interior renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data appears throughout the interior region.
[0171] Clause 31A. An apparatus according to any combination of clauses 27A and 28A, wherein the audio data includes ambisonic audio data representing a primary audio source and a secondary audio source, wherein the apparatus further comprises: means for obtaining an indication of an opacity of the secondary audio source, and wherein the means for obtaining the current renderer comprises: means for obtaining the current renderer based on the listener position, the boundary, and the indication.
[0172] Clause 32A. The apparatus of Clause 31A, wherein means for obtaining the indication of the acoustic opacity comprises means for obtaining the indication of the acoustic opacity of the secondary audio source from a bitstream representing the audio data.
[0173] Clause 33A. The apparatus of any combination of clauses 31A and 32A, further comprising: a unit for obtaining the current renderer when the indication indicating the opacity is enabled and based on the listener position and the boundary, the current renderer excluding the addition of the secondary audio sources indicated by the listener position as not being directly in line of sight.
[0174] Clause 34A. The apparatus of any combination of clauses 31A-33A, wherein the external renderer is configured to render the audio data such that a sound field represented by the audio data is spread out according to a distance between the listener position and the boundary.
[0175] Clause 35A. The apparatus of any combination of clauses 31A-34A, further comprising: a unit for updating the current renderer to interpolate between the external renderer and the internal renderer to obtain an updated current renderer in response to determining that the listener position is within a buffer distance from the boundary; and a unit for applying the current renderer to audio data to obtain one or more updated speaker feeds.
[0176] Clause 36A. The apparatus of Clause 35A, further comprising means for obtaining an indication of the buffer distance from a bitstream representing the audio data.
[0177] Clause 37A. An apparatus according to any combination of clauses 27A to 36A, further comprising: a unit for obtaining an indication of the complexity of the current renderer from a bitstream representing the audio data, and wherein the unit for obtaining the current renderer comprises: a unit for obtaining the current renderer based on the boundary, the listener position and the indication of the complexity.
[0178] Clause 38A. An apparatus according to clause 37A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having a zeroth order, and wherein the means for obtaining the current renderer includes: means for obtaining the outer renderer when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, such that the outer renderer is configured to render the ambisonic audio data such that the sound field represented by the ambisonic audio data originates from the center of the inner region.
[0179] Clause 39A. An apparatus according to clause 37A, wherein the audio data includes ambisonic audio data associated with spherical basis functions having a zero order, and wherein the means for obtaining the current renderer includes: means for obtaining the external renderer when the listener position is outside the boundary and when the indication of the complexity indicates a low complexity, such that the external renderer is configured to render the audio data such that the sound field represented by the audio data is expanded according to the distance between the listener position and the boundary.
[0180] Item 40A. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: obtain an indication of a boundary separating an interior area from an exterior area; obtain a listener position indicating a position of the device relative to the interior area; based on the boundary and the listener position, obtain a current renderer that is either an internal renderer configured to render audio data for the interior area or an external renderer configured to render audio data for the exterior area; and apply the current renderer to the audio data to obtain one or more speaker feeds.
[0181] Item 1B. A device configured to generate a bitstream representing audio data, the device comprising: a memory configured to store the audio data; and one or more processors coupled to the memory and configured to: obtain a bitstream representing the audio data based on the audio data; specify in the bitstream a boundary separating an inner region from an outer region; specify in the bitstream one or more indications for controlling presentation of the audio data for the inner region or the outer region; and output the bitstream.
[0182] Clause 2B. The apparatus of Clause 1B, wherein the one or more indications include an indication indicating a complexity of the presentation.
[0183] Clause 3B. The apparatus of clause 2B, wherein the indication indicating complexity indicates low complexity or high complexity.
[0184] Clause 4B. The apparatus of any combination of clauses 1B-3B, wherein the one or more indications include an indication indicating an opacity of presentation for a secondary source present in the audio data.
[0185] Clause 5B. The apparatus of Clause 4B, wherein the indication for indicating acoustic opacity indicates acoustic opacity as either acoustic opaque or non-acoustic opaque.
[0186] Clause 6B. An apparatus as described in any combination of clauses 1B-5B, wherein the one or more indications include an indication for indicating a buffer distance around the inner area in which the presentation is interpolated between the inner presentation and the outer presentation.
[0187] Clause 7B. The apparatus of any combination of clauses 1B-6B, wherein the audio data comprises ambisonic audio data.
[0188] Item 8B. A method for generating a bitstream representing audio data, the method comprising: obtaining the bitstream representing the audio data based on the audio data; specifying a boundary separating an inner region from an outer region in the bitstream; specifying one or more indications in the bitstream for controlling the presentation of the audio data for the inner region or the outer region; and outputting the bitstream.
[0189] Clause 9B. The method of Clause 8B, wherein the one or more indications include an indication indicating a complexity of the presentation.
[0190] Clause 10B. The method of clause 9B, wherein the indication indicating complexity indicates low complexity or high complexity.
[0191] Clause 11B. The method of any combination of clauses 8B-10B, wherein the one or more indications include an indication indicating an opacity of presentation for a secondary source present in the audio data.
[0192] Clause 12B. The method of Clause 11B, wherein the indication for indicating acoustic opacity indicates acoustic opacity as either acoustic opaque or non-acoustic opaque.
[0193] Clause 13B. The apparatus of any combination of clauses 8B-12B, wherein the one or more indications include an indication of a buffer distance around the interior region in which the presentation is interpolated between the interior presentation and the exterior presentation.
[0194] Clause 14B. The method of any combination of clauses 8B-13B, wherein the audio data comprises ambisonic audio data.
[0195] Item 15B. A device configured to generate a bitstream representing audio data, the device comprising: a unit for obtaining a bitstream representing the audio data based on the audio data; a unit for specifying a boundary separating an inner region from an outer region in the bitstream; a unit for specifying one or more indications in the bitstream for controlling the presentation of the audio data for the inner region or the outer region; and a unit for outputting the bitstream.
[0196] Clause 16B. The apparatus of Clause 15B, wherein the one or more indications include an indication indicating a complexity of the presentation.
[0197] Clause 17B. The apparatus of clause 16B, wherein the indication indicating complexity indicates low complexity or high complexity.
[0198] Clause 18B. The apparatus of any combination of clauses 15B-17B, wherein the one or more indications include an indication indicating an opacity of presentation for a secondary source present in the audio data.
[0199] Clause 19B. The apparatus of clause 18B, wherein the indication for indicating acoustic opacity indicates acoustic opacity as either acoustic opaque or non-acoustic opaque.
[0200] Clause 20B. The apparatus of any combination of clauses 15B-19B, wherein the one or more indications include an indication of a buffer distance around the interior region in which the presentation is interpolated between the interior presentation and the exterior presentation.
[0201] Clause 21B. The apparatus of any combination of clauses 15B through 20B, wherein the audio data comprises ambisonic audio data.
[0202] Item 22B. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: obtain a bitstream representing the audio data based on the audio data; specify in the bitstream a boundary that separates an inner region from an outer region; specify in the bitstream one or more indications for controlling the presentation of the audio data for the inner region or the outer region; and output the bitstream.
[0203] It should be appreciated that, according to examples, certain operations or events of any of the techniques described herein may be performed in a different sequence, may be added, combined, or omitted entirely (e.g., not all described operations or events are necessary to practice the techniques). Furthermore, in some examples, operations or events may be performed concurrently, such as through multithreading, interrupt handling, or multiple processors, rather than sequentially.
[0204] In some examples, a VR device (or streaming device) may use a network interface coupled to a memory of the VR / streaming device to transmit exchange messages to an external device, wherein the exchange messages are associated with multiple available representations of the sound field. In some examples, the VR device may use an antenna coupled to the network interface to receive a wireless signal including data packets, audio packets, video packets, or transport protocol data associated with multiple available representations of the sound field. In some examples, one or more microphone arrays may capture the sound field.
[0205] In some examples, the multiple available representations of the sound field stored to the memory device may include: multiple object-based representations of the sound field, a higher-order ambisonic representation of the sound field, a mixed-order ambisonic representation of the sound field, a combination of the object-based representation of the sound field and the higher-order ambisonic representation of the sound field, a combination of the object-based representation of the sound field and the mixed-order ambisonic representation of the sound field, or a combination of the mixed-order representation of the sound field and the higher-order ambisonic representation of the sound field.
[0206] In some examples, one or more of the sound field representations in multiple available representations of the sound field may include at least one high-resolution region and at least one lower-resolution region, and wherein the selected presentation based on the steering angle provides greater spatial accuracy relative to the at least one high-resolution region and less spatial accuracy relative to the lower-resolution region.
[0207] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or sent via a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to physical media such as data storage media, or communication media including any media that facilitates, for example, the transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) a non-transitory physical computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include computer-readable media.
[0208] As an example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory or any other medium that can be used to store the required program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technologies such as infrared, radio and microwaves are used to send instructions from a website, server or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, radio and microwaves are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals or other temporary media, but are instead directed to non-temporary physical storage media. As used herein, disks and optical disks include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically with lasers. The above combinations should also be included within the scope of computer-readable media.
[0209] Instructions may be executed by one or more processors comprising fixed-function processing circuits and / or programmable processing circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be implemented entirely in one or more circuits or logic elements.
[0210] The techniques of the present disclosure can be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of an apparatus configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units may be combined in a codec hardware unit in conjunction with appropriate software and / or firmware or provided by a collection of interoperable hardware units, including one or more processors as described above.
[0211] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A device configured to process one or more audio streams, the device comprising: One or more processors configured to: Obtaining an indication of the boundary separating the inner area from the outer area; obtaining a listener position indicating a position of the device relative to the interior area; determining a first distance between the listener position and a center of the interior area; determining a second distance between the boundary and the center of the interior region; obtaining, based on the first distance and the second distance, a current renderer that is an internal renderer configured to render audio data for the internal area or an external renderer configured to render the audio data for the external area; applying the current renderer to the audio data to obtain one or more speaker feeds; as well as a memory coupled to the one or more processors and configured to store the one or more speaker feeds, wherein the audio data comprises ambisonic audio data associated with a spherical basis function having a zeroth order, and Wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner area.
2. The device according to claim 1, wherein Said indication about the boundaries is obtained from the bitstream.
3. The device according to claim 1, in, The one or more processors are further configured to obtain an indication of complexity of the current renderer from a bitstream representing the audio data, the indication of complexity being used to indicate high complexity or low complexity.
4. The device according to claim 1, in, The interior renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data appears throughout the interior area.
5. The device according to claim 1, in, The audio data includes ambisonic audio data representing a primary audio source and a secondary audio source, wherein the one or more processors are further configured to: obtain an indication of acoustic opacity of the secondary audio source, and The one or more processors are configured to obtain the current renderer based on the first distance, the second distance, and the indication of acoustic opacity.
6. The device according to claim 5, wherein The one or more processors are configured to obtain the indication of the acoustic opacity of the secondary audio source from a bitstream representing the audio data.
7. The apparatus according to claim 5, wherein The one or more processors are further configured to, when the indication of the acoustic opacity is enabled and based on the first distance and the second distance, obtain the current renderer that excludes adding the secondary audio sources indicated by the listener position as not being directly in line of sight.
8. The apparatus according to claim 5, wherein The external renderer is configured to render the audio data such that a sound field represented by the audio data is spread out according to a distance between the listener position and the boundary.
9. The apparatus according to claim 5, wherein The one or more processors are further configured to: In response to determining that the listener position is within the buffer distance from the boundary, updating the current renderer to interpolate between the outer renderer and the inner renderer to obtain an updated current renderer; as well as The current renderer is applied to the audio data to obtain one or more updated speaker feeds.
10. The apparatus according to claim 9, wherein The one or more processors are further configured to obtain an indication of the buffer distance from a bitstream representing the audio data.
11. The device according to claim 1, in, The one or more processors are further configured to obtain an indication of the complexity of the current renderer from a bitstream representing the audio data, and Wherein the one or more processors are configured to obtain the current renderer based on the boundary, the listener position and the indication of the complexity.
12. The device according to claim 11, in, The one or more processors are configured to, when the listener position is outside the boundary and when the indication of the complexity indicates low complexity, obtain the external renderer such that the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner region.
13. The device according to claim 11, in, The one or more processors are configured to: when the listener position is outside the boundary and when the indication about the complexity indicates low complexity, obtain the external renderer, so that the external renderer is configured to render the audio data so that the sound field represented by the audio data is expanded according to the distance between the listener position and the boundary.
14. A method of processing one or more audio streams, the method comprising: obtaining, by one or more processors, an indication of a boundary separating an inner region from an outer region; obtaining, by the one or more processors, a listener position indicating a position of a device relative to the interior area; determining, by the one or more processors, a first distance between the listener position and a center of the interior area; determining, by the one or more processors, a second distance between the boundary and the center of the interior area; Obtaining, by the one or more processors, a current renderer that is an internal renderer configured to render audio data for the internal area or an external renderer configured to render the audio data for the external area based on the first distance and the second distance; as well as applying, by the one or more processors, the current renderer to the audio data to obtain one or more speaker feeds, wherein the audio data comprises ambisonic audio data associated with a spherical basis function having a zeroth order, and Wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner area.
15. The method according to claim 14, wherein Said indication about the boundaries is obtained from the bitstream.
16. The method according to claim 14, further comprising: An indication of complexity of the current renderer is obtained from a bitstream representing the audio data, the indication of complexity being used to indicate high complexity or low complexity.
17. The method according to claim 14, in, The interior renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data appears throughout the interior area.
18. The method according to claim 14, in, The audio data includes ambisonic audio data representing a primary audio source and a secondary audio source, Wherein the method further comprises: obtaining an indication of the acoustic opacity of the secondary audio source, and The obtaining of the current renderer comprises obtaining the current renderer based on the first distance, the second distance, and the indication of acoustic opacity.
19. The method according to claim 18, wherein Obtaining the indication of the opacity comprises obtaining the indication of the opacity of the secondary audio source from a bitstream representing the audio data.
20. The method of claim 18, further comprising: When the indication of the acoustic opacity is enabled and based on the first distance and the second distance, the current renderer is obtained, the current renderer excluding the addition of the secondary audio sources indicated by the listener position as not being directly in line of sight.
21. The method according to claim 18, wherein The external renderer is configured to render the audio data such that a sound field represented by the audio data is spread out according to a distance between the listener position and the boundary.
22. The method of claim 18, further comprising: In response to determining that the listener position is within the buffer distance from the boundary, updating the current renderer to interpolate between the outer renderer and the inner renderer to obtain an updated current renderer; as well as The current renderer is applied to the audio data to obtain one or more updated speaker feeds.
23. The method according to claim 22, further comprising: An indication of the buffer distance is obtained from a bitstream representing the audio data.
24. The method of claim 14, further comprising: obtaining an indication of the complexity of the current renderer from a bitstream representing the audio data, and Wherein obtaining the current renderer comprises: obtaining the current renderer based on the boundary, the listener position and the indication of the complexity.
25. The method according to claim 24, in, Obtaining the current renderer comprises, when the listener position is outside the boundary and when the indication of the complexity indicates low complexity, obtaining the outer renderer, such that the outer renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner area.
26. The method according to claim 24, in, Obtaining the current renderer includes: when the listener position is outside the boundary and when the indication about the complexity indicates low complexity, obtaining the external renderer, so that the external renderer is configured to render the audio data so that the sound field represented by the audio data is expanded according to the distance between the listener position and the boundary.
27. A device configured to generate a bitstream representing audio data, the device comprising: a memory configured to store the audio data; as well as one or more processors coupled to the memory and configured to: obtaining the bitstream representing the audio data based on the audio data; specifying a boundary separating an inner region from an outer region in the bitstream; specifying in the bitstream one or more indications that control rendering of the audio data for the inner region or the outer region, and enabling determination of a first distance between a listener position indicating a position of a device relative to the inner region and a center of the inner region, and a second distance between the boundary and the center of the inner region, to obtain a current renderer as either an inner renderer configured to render the audio data for the inner region or an outer renderer configured to render the audio data for the outer region; as well as outputting the bit stream, wherein the audio data comprises ambisonic audio data associated with a spherical basis function having a zeroth order, and Wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner area.
28. The apparatus of claim 27, wherein The one or more indications include an indication indicating a complexity of the presentation.
29. The apparatus of claim 28, wherein The indication for indicating the complexity indicates low complexity or high complexity.
30. A method of generating a bitstream representing audio data, the method comprising: obtaining the bitstream representing the audio data based on the audio data; specifying a boundary separating an inner region from an outer region in the bitstream; specifying in the bitstream one or more indications that control rendering of the audio data for the inner region or the outer region, and enabling determination of a first distance between a listener position indicating a position of a device relative to the inner region and a center of the inner region, and a second distance between the boundary and the center of the inner region, to obtain a current renderer as either an inner renderer configured to render the audio data for the inner region or an outer renderer configured to render the audio data for the outer region; as well as outputting the bit stream, wherein the audio data comprises ambisonic audio data associated with a spherical basis function having a zeroth order, and Wherein the external renderer is configured to render the ambisonic audio data such that a sound field represented by the ambisonic audio data originates from a center of the inner area.
Citation Information
Patent Citations
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US10405126B2
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US20190007781A1
An apparatus and associated methods in the field of virtual reality
CN110121695A
Audio parallax for virtual reality, augmented reality, and mixed reality
CN110168638A
Audio rendering for augmented reality
EP3506082A1