Timer-based access for audio streaming and rendering

By storing timing information and multiple audio streams in a computer-mediated real system, and controlling the access of the audio stream based on timing information, the problem of unrealistic audio experience in the prior art is solved, and higher immersion and audio source positioning accuracy are achieved.

CN114051736BActive Publication Date: 2025-07-01QUALCOMM INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202080047109.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-01
Filing Date
2020-07-02
Publication Date
2025-07-01
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

Existing computer-mediated reality systems have difficulty providing realistic immersion in audio experience, especially in case of improved video experience, ensuring that the source positioning and immersion of audio content are reduced.

Method used

Through devices that store timing information and multiple audio streams, access to the audio stream is controlled based on timing information, adaptive audio capture, synthesis and rendering are realized, and multiple audio formats such as audio formats based on channels, objects and scenes are supported.

Benefits of technology

Improves the audio experience in computer-mediated real-life systems, enhances immersion, ensures accurate positioning and rendering of audio sources, and adapts to different acoustic environments and audio needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114051736B_ABST
    Figure CN114051736B_ABST
Patent Text Reader

Abstract

Example devices and methods for timer-based access for audio streaming and rendering are presented. For example, a device configured to play one or more of a plurality of audio streams includes a memory configured to store timing information and the plurality of audio streams. The device further includes one or more processors coupled to the memory. The one or more processors are configured to control access to at least one of the plurality of audio streams based on the timing information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to Related Applications

[0002] This application claims priority to U.S. Patent Application No. 16 / 918,465, filed on July 1, 2020, and U.S. Provisional Application No. 62 / 870,599, filed on July 3, 2019, the entire contents of both of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the processing of media data, such as audio data. Background Art

[0004] Computer - mediated reality systems are being developed to allow computing devices to enhance or add, remove or subtract, or generally modify the existing reality of a user experience. Computer - mediated reality systems (which may also be referred to as "extended reality systems" or "XR systems") can include, for example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceived success of computer - mediated reality systems is generally related to the ability of such computer - mediated reality systems to provide a realistic immersive experience in terms of video and audio experiences, where the video and audio experiences are adjusted in the way the user expects. Although the human visual system is more sensitive than the human auditory system (e.g., in terms of the perceptual localization of various objects in a scene), ensuring an adequate auditory experience is an increasingly important factor in ensuring a realistic immersive experience, especially as video experiences are improved to allow better localization of video objects, enabling the user to better identify the source of audio content. Summary of the Invention

[0005] This disclosure generally relates to the auditory aspects of the user experience of computer - mediated reality systems, including virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. Various aspects of the technology can provide adaptive audio capture, synthesis, and rendering for extended reality systems. As used herein, an acoustic environment is represented as an indoor environment or an outdoor environment, or both an indoor environment and an outdoor environment. An acoustic environment can include one or more sub - acoustic spaces, which can include various acoustic elements. Examples of outdoor environments can include cars, buildings, walls, forests, etc. An acoustic space can be an example of an acoustic environment and can be an indoor space or an outdoor space. As used herein, an audio element is a sound captured by a microphone (e.g., directly captured from a near - field source or reflected from either a real or synthetic far - field source), or a previously synthesized sound field, or a mono sound synthesized from text - to - speech, or a reflection of a virtual sound from an object in the acoustic environment.

[0006] In one example, aspects of the technology are directed to an apparatus configured to store timing information and a plurality of audio streams in a memory; and one or more processors coupled to the memory and configured to control access to at least one of the plurality of audio streams based on the timing information.

[0007] In another example, aspects of the technology are directed to a method of playing one or more of a plurality of audio streams, including: storing, by a memory, timing information and a plurality of audio streams; and controlling access to at least one of the plurality of audio streams based on the timing information.

[0008] In another example, aspects of the technology are directed to an apparatus configured to play one or more of a plurality of audio streams, the apparatus including: means for storing the plurality of audio streams and means for controlling access to at least one of the plurality of audio streams based on the timing information.

[0009] In another example, aspects of the technology are directed to a non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: store timing information and a plurality of audio streams; and control access to at least one of the plurality of audio streams based on the timing information.

[0010] Details of one or more examples of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the aspects of the technology will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1A-1C is a diagram illustrating a system that can perform aspects of the technology described in the disclosure.

[0012] Figure 2 is a diagram illustrating an example of a VR device worn by a user.

[0013] Figure 3A-3E is a diagram more particularly illustrating Figure 1A-1C an example operation of a stream selection unit shown in the example of

[0014] Figure 4A-4C is a diagram illustrating Figure 1A-1C an example operation of a stream selection unit controlling access to at least one of a plurality of audio streams based on timing information in the example of

[0015] Figure 4D and 4E are diagrams further illustrating the use of timing information, such as timing metadata, according to aspects of the technology described in the disclosure.

[0016] Figure 4F and 4G is a diagram illustrating the use of a temporary request for more access in accordance with various aspects of the technology described in the present disclosure.

[0017] Figure 4H and 4I is a diagram illustrating an example of a privacy area provided in accordance with various aspects of the technology described in the present disclosure.

[0018] Figure 4J and 4K is a diagram illustrating the use of a layer of an audio rendering service in accordance with various aspects of the technology described in the present disclosure.

[0019] Figure 4L is a state transition diagram illustrating state transitions in accordance with various aspects of the technology described in the present disclosure.

[0020] Figure 4M is a diagram of a vehicle in accordance with various aspects of the technology described in the present disclosure.

[0021] Figure 4N is a diagram of a moving vehicle in accordance with various aspects of the technology described in the present disclosure.

[0022] Figure 4O is a flowchart illustrating an example technique for controlling access to at least one of a plurality of audio streams based on timing information using an authorization level.

[0023] Figure 4P is a flowchart illustrating an example technique for controlling access to at least one of a plurality of audio streams based on timing information using a trigger and a delay.

[0024] Figure 5 is a diagram illustrating an example of a wearable device that can operate in accordance with various aspects of the technology described in the present disclosure.

[0025] Figure 6A and 6B is a diagram illustrating other example systems that can execute various aspects of the technology described in the present disclosure.

[0026] Figure 7 is a block diagram illustrating example components of one or more of the source device and the content consumer device shown in the example of FIG. 1.

[0027] Figure 8A-8C is a diagram illustrating Figure 1A-1C example operations of a stream selection unit performing stream selection techniques in the example shown.

[0028] Figure 9 is a conceptual diagram illustrating an example of a wireless communication system in accordance with aspects of the present disclosure. Detailed Implementation Modes

[0029] Currently, rendering an XR scene with many audio sources (which can be obtained, for example, from audio capture devices in a live scene) may render an audio source containing sensitive information, and such sensitive information will be better restricted, or if access is allowed, such access should not be permanent. According to the technology of the present disclosure, a single audio stream can be restricted from rendering, or can be temporarily rendered based on timing information such as time or duration. For better audio interpolation, certain individual audio streams or clusters of audio streams can be enabled or disabled within a fixed duration. Therefore, the technology of the present disclosure provides a flexible way to control access to audio streams based on time.

[0030] There are many different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound formats, 7.1 surround sound formats, 22.2 surround sound formats, or any other channel-based format that localizes audio channels to specific positions around a listener to reconstruct the sound field.

[0031] An object-based audio format may refer to a format in which audio objects that are typically encoded using pulse code modulation (PCM) and are called PCM audio objects are specified to represent the sound field. Such audio objects may include position information, such as position metadata, which identifies the position of the audio object relative to a listener or other reference point in the sound field, so that the audio object can be rendered into one or more speaker channels for playback in an attempt to reconstruct the sound field. The technology described in the present disclosure can be applied to any of the following formats, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.

[0032] A scene-based audio format may include a hierarchical collection of elements that define the sound field in three dimensions. An example of a hierarchical collection of elements is a Spherical Harmonic Coefficient (SHC) collection. The following expression shows the description or representation of the sound field using SHC:

[0033]

[0034] This expression shows that at time t, any point in the sound field i has a pressure p that can be uniquely represented by SHC Here, c is the speed of sound (~343 m / s), n (·) is the spherical Bessel function of the nth order, and are spherical harmonic basis functions of order n and sub-order m (also referred to as spherical basis functions). It can be recognized that the terms in the square brackets are the frequency-domain representation of the signal (e.g., ), which can be estimated by various time-frequency transforms such as the discrete Fourier transform (DFT), discrete cosine transform (DCT), or wavelet transform. Other examples of hierarchical sets include sets of wavelet transform coefficients and other coefficient sets of multi-resolution basis functions.

[0035] The SHC can be physically acquired (e.g., recorded) through various microphone array configurations Alternatively, they can be derived from channel-based or object-based descriptions of the sound field. The SHC (also referred to as stereophonic reverberation coefficients) represents scene-based audio, where the SHC can be input into an audio encoder to obtain an encoded SHC that can facilitate more efficient transmission or storage. For example, a fourth-order representation involving (1 + 4) 2 (25 and thus fourth-order) coefficients can be used.

[0036] As described above, the SHC can be derived from microphone recordings using a microphone array. Poletti, M. described various examples of how to physically acquire the SHC from a microphone array in "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics" in J. Audio Eng. Soc., Vol. 53, No. 11, pp. 1004 - 1025, November 2005.

[0037] The following equation can illustrate how to derive the SHC from an object-based description. The coefficients corresponding to a single audio object in the sound field can be expressed as:

[0038]

[0039] where i is is the spherical Hankel function of the second kind of order n, and is the position of the object. Knowing the object source energy g(ω) as a function of frequency (e.g., using time-frequency analysis techniques such as performing a fast Fourier transform on a pulse code modulation PCM stream) enables each PCM object and its corresponding position to be converted into SHC Furthermore, (since the above is a linear and orthogonal decomposition) it can be seen that the coefficients of each object are additive. In this way, multiple PCM objects can be represented by Coefficient representation (e.g., as a sum of coefficient vectors for individual objects). The coefficients can contain information about the acoustic field (pressure as a function of three-dimensional (3D) coordinates), and the above representation enables the transformation of the representation from individual objects to the entire acoustic field near the observation point nearby.

[0040] Computer-mediated reality systems (which may also be referred to as "extended reality systems" or "XR systems") are being developed to take advantage of the many potential benefits offered by stereophonic reverberation coefficients. For example, stereophonic reverberation coefficients can represent the acoustic field in three dimensions in a way that potentially enables accurate 3D localization of sound sources within the acoustic field. In this way, an XR device can render the stereophonic reverberation coefficients to a speaker feed, which, when played through one or more speakers, can accurately reproduce the acoustic field.

[0041] As another example, the stereophonic reverberation coefficients can be transformed or rotated to account for user movement without overly complex mathematical operations, thereby potentially meeting the low-latency requirements of XR devices. Additionally, the stereophonic reverberation coefficients are hierarchical, thus naturally accommodating scalability through order reduction (which can eliminate the stereophonic reverberation coefficients associated with higher orders), thereby potentially enabling dynamic adaptation of the acoustic field to meet the latency and / or battery requirements of XR devices.

[0042] Using stereophonic reverberation coefficients for XR devices can develop multiple use cases that rely on a more immersive acoustic field provided by the stereophonic reverberation coefficients, particularly for computer game applications and live video streaming applications. In these high-dynamic use cases that rely on low-latency acoustic field reproduction, XR devices may prefer stereophonic reverberation coefficients over other representations that are more difficult to manipulate or involve complex rendering. More information about these use cases is provided below regarding Figure 1A-1C provided.

[0043] Although described in the context of VR devices in this disclosure, various aspects of the technology can be performed in the context of other devices such as mobile devices. In such a case, a mobile device (such as a so-called smartphone) can present the acoustic space via a screen, which can be mounted to the head of user 102 or viewed as is done during normal use of the mobile device. In this way, any information on the screen can be part of the mobile device. The mobile device may be able to provide tracking information, thereby allowing for VR experiences (when worn on the head) and normal experiences to view the acoustic space, where the normal experience can still allow the user to view the acoustic space providing a VR-lite experience (e.g., holding up the device and rotating or translating the device to view different parts of the acoustic space).

[0044] Figure 1A-1C is a diagram of a system that can perform various aspects of the technology described in this disclosure. As in Figure 1AAs shown in the example of, system 10 includes source device 12A and content consumer device 14A. Although described in the context of source device 12A and content consumer device 14A, these techniques may also be implemented in any context where any representation of a sound field is encoded to form a bitstream representing audio data. Additionally, source device 12A may represent any form of computing device capable of generating a representation of a sound field and is generally described herein in the context of being a VR content creator device. Similarly, content consumer device 14A may represent any form of computing device capable of implementing the rendering techniques described in the present disclosure as well as audio playback and is generally described herein in the context of being a VR client device.

[0045] Source device 12A may be operated by an entertainment company or other entity that can generate mono and / or multi-channel audio content for consumption by an operator of a content consumer device such as content consumer device 14A. In some VR scenarios, source device 12A generates audio content in combination with video content. Source device 12A includes content capture device 20, content editing device 22, and sound field representation generator 24. Content capture device 20 may be configured to connect to or otherwise communicate with microphone 18.

[0046] Microphone 18 may represent or other types of 3D audio microphones that are capable of capturing a sound field and representing it as audio data 19, which may refer to one or more of the above-described scene-based audio data (such as stereo reverberation coefficients), object-based audio data, and channel-based audio data. Although described as a 3D audio microphone, microphone 18 may also represent other types of microphones (such as omnidirectional microphones, point microphones, unidirectional microphones, etc.) configured to capture audio data 19. Audio data 19 may represent an audio stream or include an audio stream.

[0047] In some examples, content capture device 20 may include integrated microphone 18 integrated into the housing of content capture device 20. Content capture device 20 may be connected to microphone 18 wirelessly or via a wired connection. Instead of capturing audio data 19 via microphone 18 or in combination with capturing audio data 19 via microphone 18, content capture device 20 may process audio data 19 after processing input audio data 19 wirelessly and / or via a wired input from some type of removable storage. Thus, according to the present disclosure, various combinations of content capture device 20 and microphone 18 are possible.

[0048] The content capture device 20 may also be configured to connect to or otherwise communicate with the content editing device 22. In some cases, the content capture device 20 may include the content editing device 22 (which, in some cases, may represent software or a combination of software and hardware, including software executed by the content capture device 20 to configure the content capture device 20 to perform a particular form of content editing). The content editing device 22 may represent a unit configured to edit or otherwise alter the content 21 (including the audio data 19) received from the content capture device 20. The content editing device 22 may output the edited content 23 and associated metadata 25 to the sound field representation generator 24.

[0049] The sound field representation generator 24 may include any type of hardware device capable of connecting to the content editing device 22 (or the content capture device 20). Although not shown in the Figure 1A example, the sound field representation generator 24 may use the edited content 23 including the audio data 19 and the metadata 25 provided by the content editing device 22 to generate one or more bitstreams 27. In an example focusing on the audio data 19 Figure 1A , the sound field representation generator 24 may generate one or more representations of the same sound field represented by the audio data 19 to obtain a bitstream 27 including a representation of the sound field and the audio metadata 25.

[0050] For example, to generate different representations of a sound field using a stereo reverberation coefficient (which is also an example of the audio data 19), the sound field representation generator 24 may use an encoding scheme for a stereo reverberation representation of the sound field, called Mixed Order Ambisonics (MOA), as discussed in more detail in U.S. Application Serial No. 15 / 672,058, filed August 8, 2017, and titled "MIXED-ORDER AMBISONICS (MOA) AUDIO DATA FOR COMPUTER-MEDIATED REALITY SYSTEMS", and published as U.S. Patent Publication No. 20190007781 on January 3, 2019.

[0051] To generate a particular MOA representation of the sound field, the sound field representation generator 24 may generate a partial subset of the complete set of stereophonic reverberation coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide accuracy for some regions of the sound field, but lower accuracy in other regions. In one example, the MOA representation of the sound field may include eight (8) uncompressed stereophonic reverberation coefficients, while the third-order stereophonic reverberation representation of the same sound field may include sixteen (16) uncompressed stereophonic reverberation coefficients. Thus, each MOA representation of the sound field generated as a partial subset of the stereophonic reverberation coefficients may have lower storage intensity and bandwidth intensity than the corresponding third-order stereophonic reverberation representation of the same sound field generated from the stereophonic reverberation coefficients (if and when transmitted as part of the bit stream 27 through the illustrated transmission channel).

[0052] Although described with respect to the MOA representation, the techniques of the present disclosure may also be performed with respect to a first-order stereophonic reverberation (FOA) representation, where all stereophonic reverberation coefficients associated with first-order spherical basis functions and zero-order spherical basis functions are used to represent the sound field. In other words, the sound field representation generator 302 may use all stereophonic reverberation coefficients of a given order N to represent the sound field, rather than using a partial non-zero subset of the stereophonic reverberation coefficients to represent the sound field, such that the total stereophonic reverberation coefficients equal (N + 1) 2 。

[0053] In this regard, stereophonic reverberation audio data (which is another way of referring to stereophonic reverberation coefficients in the MOA representation or full-order representation, such as the first-order representation mentioned above) may include stereophonic reverberation coefficients associated with spherical basis functions of one or fewer orders (which may be referred to as "first-order stereophonic reverberation audio data"), stereophonic reverberation coefficients associated with spherical basis functions having mixed orders and sub-orders (which may be referred to as the "MOA representation" above), or stereophonic reverberation coefficients associated with spherical basis functions having more than one order (referred to above as the "full-order representation").

[0054] In some examples, the content capture device 20 or the content editing device 22 may be configured to communicate wirelessly with the sound field representation generator 24. In some examples, the content capture device 20 or the content editing device 22 may communicate with the sound field representation generator 24 via one or both of a wireless connection and a wired connection. Via the connection between the content capture device 20 or the content editing device 22 and the sound field representation generator 24, the content capture device 20 or the content editing device 22 may provide content in various content forms, which for purposes of discussion is described herein as part of the audio data 19.

[0055] In some examples, the content capture device 20 may utilize aspects of the sound field representation generator 24 (in terms of the hardware or software capabilities of the sound field representation generator 24). For example, the sound field representation generator 24 may include dedicated hardware that is configured to (or dedicated software that, when executed, causes one or more processors to) perform psychoacoustic audio coding (such as the Unified Speech and Audio Coding represented as "USAC" proposed by the Moving Picture Experts Group (MPEG), the MPEG-H 3D audio coding standard, the MPEG-I immersive audio standard, or proprietary standards such as AptXTM (including various versions of AptX such as enhanced AptX - E - AptX, live AptX, AptX stereo, and AptX high definition - AptX - HD), Advanced Audio Coding (AAC), Audio Codec 3 (AC - 3), Apple Lossless Audio Codec (ALAC), MPEG - 4 Audio Lossless Streaming (ALS), Enhanced AC - 3, Free Lossless Audio Codec (FLAC), Monkey's Audio, MPEG - 1 Audio Layer II (MP2), MPEG - 1 Audio Layer III (MP3), Opus, and Windows Media Audio (WMA).

[0056] The content capture device 20 may not include dedicated hardware or specialized software for psychoacoustic audio coding, but rather may provide the audio aspects of the content 21 in the form of non - psychoacoustic audio coding. The sound field representation generator 24 may assist in the capture of the content 21 by at least partially performing psychoacoustic audio coding on the audio aspects of the content 21.

[0057] The sound field representation generator 24 may also assist in content capture and transmission by generating one or more bitstreams 27, at least partially based on audio content (e.g., MOA representation and / or first - order stereophonic reverberation representation) generated from the audio data 19 (in the case where the audio data 19 includes scene - based audio data). The bitstream 27 may represent a compressed version of the audio data 19 and any other different types of content 21 (such as a compressed version of spherical video data, image data, or text data).

[0058] As an example, the sound field representation generator 24 may generate a bitstream 27 for transmission across a transmission channel, which may be a wired or wireless channel, a data storage device, etc. The bitstream 27 may represent an encoded version of the audio data 19 and may include a main bitstream and an additional bitstream, which may be referred to as side channel information or metadata. In some cases, the bitstream 27 representing a compressed version of the audio data 19 (which may also represent scene-based audio data, object-based audio data, channel-based audio data, or a combination thereof) may conform to a bitstream generated according to the MPEG-H 3D audio coding standard and / or the MPEG-I immersive audio standard.

[0059] The content consumer device 14 may be operated by an individual and may represent a VR client device. Although described with respect to a VR client device, the content consumer device 14 may represent other types of devices, such as an augmented reality (AR) client device, a mixed reality (MR) client device (or other XR client device), a standard computer, a headset, headphones, a mobile device (including so-called smartphones), or any other device capable of tracking the head movement and / or general translational movement of an individual operating the content consumer device 14. As Figure 1A shown in the example of, the content consumer device 14 includes an audio playback system 16A, which may refer to any form of audio playback system capable of rendering audio data for playback as single-channel and / or multi-channel audio content.

[0060] Although shown in Figure 1A as being sent directly to the content consumer device 14, the source device 12A may output the bitstream 27 to an intermediate device located between the source device 12A and the content consumer device 14A. The intermediate device may store the bitstream 27 for later delivery to a content consumer device 14A that requests the bitstream 27. The intermediate device may include a file server, a web server, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a smartphone, or any other device capable of storing the bitstream 27 for later retrieval by an audio decoder. The intermediate device may reside in a content delivery network that is capable of streaming the bitstream 27 to subscribers such as the content consumer device 14 that request the bitstream 27 (and possibly in combination with sending a corresponding video data bitstream).

[0061] Alternatively, the source device 12A can store the bitstream 27 to a storage medium such as a compact disc, a digital video disc, a high definition video disc, or other storage media, most of which are readable by a computer and can thus be referred to as computer-readable storage media or non-transitory computer-readable storage media. In this context, the transmission channel can refer to the channel through which the content stored on the medium (e.g., in the form of one or more bitstreams 27) is sent (and can include retail stores and other store-based delivery mechanisms). Thus, in any case, the techniques of the present disclosure should not be limited in this regard to Figure 1A examples of

[0062] As described above, the content consumer device 14 includes an audio playback system 16A. The audio playback system 16A can represent any system capable of playing back mono-channel and / or multi-channel audio data. The audio playback system 16A can include a plurality of different renderers 32. Each audio renderer 32 can provide a different form of rendering, where the rendering of different audio forms can include one or more of various ways of performing Vector-Base Amplitude Panning (VBAP), and / or one or more of various ways of performing sound field synthesis. As used herein, "A and / or B" means "A or B", or both "A and B".

[0063] The audio playback system 16A can also include an audio decoding device 34. The audio decoding device 34 can represent a device configured to decode the bitstream 27 to output audio data 19' (where the prime notation can indicate that the audio data 19' is different from the audio data 19 due to lossy compression of the audio data 19, such as quantization). Again, the audio data 19' can include scene-based audio data, which in some examples can form a complete first (or higher) order stereo reverberation representation or a subset of an MOA representation of the same sound field, its decomposition (such as the primary audio signal, the ambient stereo reverberation coefficient, and the vector-based signal described in the MPEG-H 3D audio coding standard), or other forms of scene-based audio data.

[0064] Other forms of scene-based audio data include audio data defined according to the Higher Order Ambisonics (HOA) Transport Format (HTF). More information about the HTF can be found in the Technical Specification (TS) titled "Higher Order Ambisonics (HOA) Transport Format" in ETSI TS 103589 V1.1.1, dated June 2018 (2018-06) of the European Telecommunications Standards Institute (ETSI), and in U.S. Patent Publication No. 2019 / 0918028, titled "PRIORITY INFORMATION FOR HIGHER ORDER AMBISONIC AUDIO DATA", filed on December 20, 2018. In any case, this audio data 19' may be similar to the complete set or a partial subset of the audio data 19', but may differ due to lossy operations (e.g., quantization) and / or transmission via the transmission channel.

[0065] This audio data 19' may include channel-based audio data as an alternative to or in combination with scene-based audio data. This audio data 19' may include object-based audio data or channel-based audio data as an alternative to or in combination with scene-based audio data. Thus, this audio data 19' may include any combination of scene-based audio data, object-based audio data, and channel-based audio data.

[0066] The audio renderer 32 of the audio playback system 16A may render the audio data 19' to output a speaker feed 35 after the audio decoding device 34 has decoded the bitstream 27 to obtain the audio data 19'. The speaker feed 35 may drive one or more speakers (not illustrated in the example for the purpose of illustration). Various audio representations, including scene-based audio data of the sound field (and possibly channel-based audio data and / or object-based audio data), may be normalized in various ways, including N3D, SN3D, FuMa, N2D, or SN2D. Figure 1A The example of

[0067] To select a suitable renderer, or in some cases, generate a suitable renderer, the audio playback system 16A can obtain speaker information 37 indicative of a plurality of speakers (e.g., loudspeakers or headphone speakers) and / or the spatial geometry of the speakers. In some cases, the audio playback system 16A can obtain the speaker information 37 using a reference microphone and can drive the speakers in a manner that dynamically determines the speaker information 37 (which can refer to the output of an electrical signal to cause a transducer to vibrate). In other cases, or in combination with the dynamic determination of the speaker information 37, the audio playback system 16A can prompt the user to interact with the audio playback system 16A and input the speaker information 37.

[0068] The audio playback system 16A can select one of the audio renderers 32 based on the speaker information 37. In some cases, when none of the audio renderers 32 is within a certain threshold similarity metric (in terms of speaker geometry) of the speaker geometry specified in the speaker information 37, the audio playback system 16A can generate one of the audio renderers 32 based on the speaker information 37. In some cases, the audio playback system 16A can generate one of the audio renderers 32 based on the speaker information 37 without first attempting to select an existing one of the audio renderers 32.

[0069] When outputting the speaker feed 35 to headphones, the audio playback system 16A can utilize one of the renderers 32 that uses a head-related transfer function (HRTF) or other function capable of rendering to left and right speaker feeds 35 for headphone speaker playback to provide binaural rendering, such as a binaural room impulse response renderer. The term "speaker" or "transducer" can generally refer to any speaker, including loudspeakers, headphone speakers, bone conduction speakers, earbud speakers, wireless headphone speakers, etc. One or more speakers can then play back the rendered speaker feed 35 to reproduce the sound field.

[0070] Although described as rendering the speaker feed 35 from the audio data 19', the reference to the rendering of the speaker feed 35 can refer to other types of rendering, such as rendering directly incorporated into the decoding of the audio data from the bitstream 27. Examples of alternative rendering can be found in Appendix G of the MPEG-H 3D Audio standard, where rendering occurs during primary signal formation and background signal formation prior to sound field synthesis. Thus, the reference to the rendering of the audio data 19' should be understood to refer to the actual rendering of the audio data 19' or the decomposition or representation of the audio data 19' (such as the primary audio signal, the environmental stereo reverberation coefficient, and / or the vector-based signal - which can also be referred to as the V vector or the multi-dimensional stereo reverberation space vector) described above.

[0071] The audio playback system 16A can also adjust the audio renderer 32 based on the tracking information 41. That is, the audio playback system 16A can be connected to a tracking device 40, which is configured to track the head movement and possible translational movement of a user of the VR device. The tracking device 40 can represent one or more sensors (e.g., cameras - including depth cameras, gyroscopes, magnetometers, accelerometers, light emitting diodes - LEDs, etc.), which are configured to track the head movement and possible translational movement of a user of the VR device. The audio playback system 16A can adjust the audio renderer 32 based on the tracking information 41 such that the speaker feeds 35 reflect the changes and possible translational movement of the user's head to correctly reproduce the sound field in response to such movement.

[0072] Figure 1B is a block diagram of another example system 50 configured to perform various aspects of the techniques described in the present disclosure. The system 50 is similar to Figure 1A the system 10 shown, except that Figure 1A the audio renderer 32 shown in

[0073] is replaced by a binaural renderer 42 capable of performing binaural rendering using one or more head-related transfer functions (HRTFs) or other functions capable of rendering to left and right speaker feeds 43.

[0074] TM The audio playback system 16B can output the left and right speaker feeds 43 to headphones 48, which can represent another example of a wearable device and can be coupled to additional wearable devices to facilitate the reproduction of the sound field, such as a watch, the VR headset mentioned above, smart glasses, smart clothing, smart rings, smart bracelets, or any other type of smart jewelry (including smart necklaces), etc. The headphones 48 can be coupled to the additional wearable devices wirelessly or via a wired connection.

[0075] Figure 1C is a block diagram of another example system 60. The example system 60 is similar to Figure 1AExample system 10, but the source device 12B of system 60 does not include a content capture device. The source device 12B includes a synthesis device 29. A content developer can use the synthesis device 29 to generate a synthesized audio source. The synthesized audio source can have location information associated therewith, which can identify the location of the audio source relative to a listener or other reference point in the sound field, such that the audio source can be rendered into one or more speaker channels for playback in an attempt to reconstruct the sound field.

[0076] For example, a content developer can generate a synthesized audio stream for a video game. Although Figure 1C the example of Figure 1A is shown with the content consumer device 14 of the example of Figure 1C the source device 12B of the example of Figure 1B can be used with the content consumer device 14B of Figure 1C In some examples, the source device 12B of

[0077] can also include a content capture device, such that the bitstream 27 can include a captured audio stream and a synthesized audio stream.

[0077] As described above, the content consumer device 14A or 14B (either of which may hereinafter be referred to as the content consumer device 14) can represent a VR device, where a human wearable display (which may also be referred to as a "head-mounted display") is mounted in front of the eyes of a user operating the VR device. Figure 2 is a diagram illustrating an example of a VR device 1100 worn by a user 1102. The VR device 1100 is coupled to or otherwise includes headphones 1104, and the headphones 1104 can reproduce the sound field represented by the audio data 19' through playback of the speaker feed 35. The speaker feed 35 can represent an analog or digital signal capable of causing a diaphragm within a transducer of the headphones 104 to vibrate at various frequencies, where this process is generally referred to as driving the headphones 1104.

[0078] Video, audio, and other sensory data may play an important role in a VR experience. To participate in a VR experience, a user 1102 can wear a VR device 1100 (which may also be referred to as a VR client device 1100) or other wearable electronic device. The VR client device (such as the VR device 1100) can include a tracking device (e.g., the tracking device 40), which is configured to track the head movement of the user 1102 and adjust the video data displayed via the VR device 1100 to account for the head movement, thereby providing an immersive experience in which the user 1102 can experience an acoustic space displayed in three visible dimensions in the video data. The acoustic space can refer to a virtual world (where all worlds are simulated), an augmented world (where a portion of the world is enhanced by virtual objects), or a physical world (where real-world images are virtually navigated).

[0079] Although VR (and other forms of AR and / or MR) may allow user 1102 to visually reside in a virtual world, VR device 1100 may often lack the ability to auditorily place the user in an acoustic space. In other words, a VR system (which may include a computer responsible for rendering video data and audio data (not shown for purposes of illustration in the Figure 2 example of ) and VR device 1100) may not support full three-dimensional auditory immersion (and in some cases, actually in a manner that reflects the displayed scene presented to the user via VR device 1100).

[0080] Although described in the context of a VR device in this disclosure, various aspects of the technology may be implemented in the context of other devices such as a mobile device. In this case, a mobile device (such as a so-called smartphone) may present an acoustic space via a screen that may be mounted to the head of user 1102 or viewed as done during normal use of the mobile device. Thus, any information on the screen may be part of the mobile device. The mobile device may be able to provide tracking information 41, thereby allowing for VR experiences (when worn on the head) and normal experiences to view the acoustic space, where the normal experience may still allow the user to view the acoustic space providing a VR-lite experience (e.g., lifting the device and rotating or translating the device to view different parts of the acoustic space).

[0081] In any case, returning to the VR device context, the audio aspect of VR has been divided into three different immersion categories. The first category provides the lowest level of immersion, called three degrees of freedom (3DOF). 3DOF refers to audio rendering that takes into account the movement of the head in three degrees of freedom (yaw, pitch, and roll), thereby allowing the user to freely look around in any direction. However, 3DOF cannot account for translational head movements where the head is not centered on the optical and acoustic center of the sound field.

[0082] The second category (called 3DOF plus (3DOF+)) provides three degrees of freedom (yaw, pitch, and roll) in addition to limited spatial translational movement due to the head moving away from the acoustic and optical center of the sound field. 3DOF+ can provide support for perceptual effects such as motion parallax, which can enhance the sense of immersion.

[0083] The third category (called six degrees of freedom (6DOF)) renders audio data in a manner that takes into account three degrees of freedom in terms of head movement (yaw, pitch, and roll), while also considering the user's translation in space (x, y, and z translations). The spatial translation can be introduced via sensors that track the user's position in the physical world or via an input controller.

[0084] 3DOF rendering is the latest technology in the audio aspect of VR. As such, the immersion in the audio aspect of VR is not as good as that in the video aspect, which may reduce the overall immersion of the user experience. However, VR is evolving rapidly and may quickly develop to support both 3DOF+ and 6DOF, which may provide opportunities for other use cases.

[0085] For example, interactive game applications can utilize 6DOF to facilitate fully immersive games where the user moves themselves within the VR world and can interact with virtual objects by walking up to them. Additionally, interactive live stream applications can utilize 6DOF to allow VR client devices to experience a live stream of a concert or a sports event as if they were there themselves, allowing users to move around during the concert or sports event.

[0086] There are multiple difficulties associated with these use cases. In the case of fully immersive games, latency may need to be kept at a low level to enable gameplay that does not cause nausea or motion sickness. Additionally, from an audio perspective, audio playback latency that causes desynchronization with video data may reduce immersion. Moreover, for certain types of game applications, spatial accuracy may be important for allowing accurate responses, including regarding how users perceive sound, as this allows users to anticipate actions that are not currently in view.

[0087] In the context of live stream applications, a large number of source devices 12A or 12B (either of which may hereinafter be referred to as source device 12) can stream content 21, where the source devices 12 can have widely different capabilities. For example, one source device can be a smartphone with a digital fixed lens camera and one or more microphones, while another source device can be a production-grade television device capable of obtaining video with much higher resolution and quality than a smartphone. However, in the context of live stream applications, all source devices can provide streams of different qualities, and VR devices can attempt to select a suitable stream from them to provide the expected experience.

[0088] Furthermore, similar to game applications, latency in audio data that causes desynchronization with video data may result in reduced immersion. Additionally, spatial accuracy may also be important so that users can better understand the environment or location of different audio sources. Moreover, privacy may become an issue when users use cameras and microphones for live streaming, as users may not want the live stream to be completely open to the public.

[0089] In the context of a streaming application (live or recorded), there may be a large number of audio streams associated with different levels of quality and / or content. The audio streams can represent any type of audio data, including scene-based audio data (e.g., stereo reverberant audio data, including FOA audio data, MOA audio data, and / or HOA audio data), channel-based audio data, and object-based audio data. Selecting only one of the potentially large number of audio streams to reconstruct the sound field may not provide an experience that ensures a sufficient level of immersion. However, due to the different spatial localizations between multiple audio streams, selecting multiple audio streams may create interference, potentially reducing the immersion.

[0090] According to the techniques described in this disclosure, the audio decoding device 34 can adaptively select between the audio streams available via the bitstream 27 (which is represented by the bitstream 27 and thus the bitstream 27 can also be referred to as "audio stream 27"). The audio decoding device 34 can select between different audio streams of the audio stream 27 based on audio location information (ALI) (e.g., Figure 1A-1C 45A in). In some examples, the audio location information can be included as metadata accompanying the audio stream 27, where the audio location information can define coordinates in the acoustic space for the microphones that captured the corresponding audio stream 27 or virtual coordinates for the synthesized audio stream. The ALI 45A can represent the capture location where the corresponding one of the audio stream 27 in the acoustic space was captured or the virtual coordinates where the corresponding one in the audio stream was synthesized. The audio decoding device 34 can select a subset of the audio stream 27 based on the ALI 45A, where the subset of the audio stream 27 excludes at least one of the audio stream 27. The audio decoding device 34 can output the subset of the audio stream 27 as audio data 19' (which can also be referred to as "audio data 19'").

[0091] In addition, the audio decoding device 34 can obtain tracking information 41, which the content consumer device 14 can convert to device location information (DLI) (e.g., Figure 1A-1C 45B in). The DLI 45B can represent the virtual or actual location of the content consumer device 14 in the acoustic space, which can be defined as one or more device coordinates in the acoustic space. The content consumer device 14 can provide the DLI 45B to the audio decoding device 34. The audio decoding device 34 can then select the audio data 19' from the audio stream 27 based on the ALI 45A and the DLI 45B. The audio playback system 16A can then reproduce the corresponding sound field based on the audio data 19'.

[0092] In this regard, the audio decoding device 34 can adaptively select a subset of the audio stream 27 to obtain audio data 19' that can result in a more immersive experience (compared to selecting a single audio stream or all of the audio data 19'). In this way, various aspects of the techniques described in the present disclosure can improve the operation of the audio decoding device 34 (as well as the audio playback system 16A or 16B and the content consumer device 14) itself by potentially enabling the audio decoding device 34 to better spatialize sound sources within the sound field, thereby enhancing the sense of immersion.

[0093] In operation, the audio decoding device 34 can interact with one or more source devices 12 to determine the ALI 45A of each of the audio streams 27. As Figure 1A shown in the example of, the audio decoding device 34 can include a stream selection unit 44, which can represent a unit configured to perform various aspects of the audio stream selection techniques described in the present disclosure.

[0094] The stream selection unit 44 can generate a constellation map (CM) 47 based on the ALI 45A. The CM 47 can define the ALI 45A for each of the audio streams 27. The stream selection unit 44 can also perform an energy analysis for each of the audio streams 27 to determine an energy map for each audio stream 27, and store the energy map together with the ALI 45A in the CM 47. The energy map can collectively define the energy of the common sound field represented by the audio stream 27.

[0095] The stream selection unit 44 can then determine the (multiple) distances between the device location represented by the DLI 45B and the (multiple) capture locations or (multiple) synthesis locations represented by the ALI 45A, which are associated with at least one and possibly each of the audio streams 27. The stream selection unit 44 can then select the audio data 19' from the audio stream 27 based on the (multiple) distances, as discussed in more detail below with respect to Figure 3A -3F.

[0096] In addition, in some examples, the stream selection unit 44 may also select the audio data 19' from the audio stream 27 based on the energy map stored in the CM 47, the ALI 45A, and the DLI 45B (where the ALI 45A and the DLI 45B are presented together in the form of the above-mentioned distance (which may also be referred to as "relative distance")). For example, the stream selection unit 44 may analyze the energy map presented in the CM 47 to determine the audio source location (ASL) 49 of the audio source of the emitted sound captured by the microphone (such as the microphone 18) in the common sound field and represented by the audio stream 27. The stream selection unit 44 may then determine the audio data 19' from the audio stream 27 based on the ALI 45A, the DLI 45B, and the ASL 49. More information on how the stream selection unit 44 may select the stream is discussed below with respect to Figure 3A -3F.

[0097] Figure 3A -3F is a diagram that more specifically illustrates Figure 1A-1C an example operation of the stream selection unit 44 shown in the example. As Figure 3A shown in the example, the stream selection unit 44 may determine that the DLI 45B indicates that the content consumer device 14 (illustrated as the VR device 1100) is at the virtual location 300A. The stream selection unit 44 may then determine the ALI45A of one or more of the audio elements 302A - 302J (collectively referred to as the audio elements 302), where the audio elements 302A - 302J may represent not only the microphone (such as Figure 1A the microphone 18 shown in), but also other types of capture devices, including other XR devices, mobile phones (including so-called smartphones), etc., or synthetic sound fields, etc.

[0098] As described above, the stream selection unit 44 may obtain the audio stream 27. The stream selection unit 44 may be connected to the audio elements 302A - 302J to obtain the audio stream 27. In some examples, the stream selection unit 44 may interact with an interface (such as a receiver, a transmitter, and / or a transceiver) to obtain the audio stream 27 according to the fifth-generation (5G) cellular standard, a personal area network (PAN) (such as BluetoothTM), or some other open-source, proprietary, or standardized communication protocol. The wireless communication of the audio stream is represented as a lightning ball in Figure 3A-3E the example, where the selected audio data 19' is shown as communication from one or more of the selected audio elements 302 to the VR device 1100.

[0099] In any case, the stream selection unit 44 may then obtain the energy map in the above-described manner, analyze the energy map to determine the audio source location 304, which may represent Figure 1AAn example of the ASL 49 as shown in the example. This energy map can represent the audio source location 304 because the energy at this audio source location 304 may be higher than the surrounding area. Assuming that each in the energy map can represent this higher energy, the stream selection unit 44 can triangulate the audio source location 304 based on the higher energy in the energy map.

[0100] Next, the stream selection unit 44 can determine the audio source distance 306A as the distance between the audio source location 304 and the virtual location 300A of the VR device 1100. The stream selection unit 44 can compare the audio source distance 306A with an audio source distance threshold. In some examples, the stream selection unit 44 can derive the audio source distance threshold based on the energy of the audio source 308. That is, when the audio source 308 has a high energy (or in other words, when the sound of the audio source 308 is louder), the stream selection unit 44 can increase the audio source distance threshold. When the audio source 308 has a low energy (or in other words, when the sound of the audio source 308 is quieter), the stream selection unit 44 can decrease the audio source distance threshold. In other examples, the stream selection unit 44 can obtain a statically defined audio source distance threshold, which can be statically defined or specified by the user 1102.

[0101] In any case, when the audio source distance 306A is greater than the audio source distance threshold (assumed for illustrative purposes in this example), the stream selection unit 44 can select a single audio stream captured by the audio elements 302A - 302J (“audio elements 302”) in the audio stream 27. The stream selection unit 44 can output the corresponding one in the audio stream 27, and the audio decoding device 34 can decode and output it as the audio data 19'.

[0102] Assuming that the user 1102 moves from the virtual location 300A to the virtual location 300B, the stream selection unit 44 can determine the audio source distance 306B as the distance between the audio source location 304 and the virtual location 300B. In some examples, the stream selection unit 44 can be updated only after a certain configurable release time, which can refer to the time from when the listener stops moving until the receiver area increases.

[0103] In any case, the stream selection unit 44 can compare the audio source distance 306A with the audio source distance threshold again. The stream selection unit 44 can select multiple audio streams captured by the audio elements 302A - 302J ("audio elements 302") in the audio stream 27 when the audio source distance 306B is less than or equal to the audio source distance threshold (assumed for illustrative purposes in this example). The stream selection unit 44 can output corresponding ones in the audio stream 27, and the audio decoding device 34 can decode and output them as the audio data 19'.

[0104] The stream selection unit 44 can also determine one or more proximity distances between the virtual location 300B and one or more (and possibly each) of the capture locations represented by the ALI 45A. The stream selection unit 44 can then compare the one or more proximity distances with a threshold proximity distance. When one or more of the proximity distances are greater than the threshold proximity distance, the stream selection unit 44 can select a smaller number of audio streams 27 compared to when one or more of the proximity distances are less than or equal to the threshold proximity distance to obtain the audio data 19'. However, when one or more of the proximity distances are less than or equal to the threshold proximity distance, the stream selection unit 44 can select a larger number of audio streams 27 compared to when one or more of the proximity distances are less than or equal to the threshold proximity distance to obtain the audio data 19'.

[0105] In other words, the stream selection unit 44 can attempt to select those audio streams in the audio stream 27 such that the audio data 19' is most closely aligned with and around the virtual location 300B. The proximity distance threshold can define such a threshold where the user 1102 of the VR device 1100 can set the threshold or the stream selection unit 44 can dynamically determine the threshold again based on the quality of the audio elements 302F - 302J, the gain or loudness of the audio source 308, the tracking information 41 (e.g., for determining whether the user 1102 is facing the audio source 308), or any other factor.

[0106] In this regard, the stream selection unit 44 can increase the audio spatialization accuracy when the listener is at the location 300B. Additionally, when the listener is at the location 300A, the stream selection unit 44 can reduce the bit rate because only the audio stream captured by the audio element 302A instead of multiple audio streams of the audio elements 302B - 302J is used to reproduce the sound field.

[0107] Refer to the following Figure 3BExample, the stream selection unit 44 may determine that the audio stream of the audio element 302A is corrupted, noisy, or unavailable. Assuming that the distance of the audio source 306A is greater than the audio source distance threshold, the stream selection unit 44 may remove the audio stream from the CM 47 according to the techniques described in more detail above and repeat in the audio stream 27 to select a single one of the audio streams 27 (e.g., in Figure 3B the example, the audio stream captured by the microphone 302B).

[0108] Refer to below Figure 3C Example, the stream selection unit 44 may obtain a new audio stream (the audio stream of the audio element 302K) and corresponding new audio information including the ALI 45A, such as metadata. The stream selection unit 44 may add the new audio stream to the CM 47 representing the audio stream 27. Assuming that the distance of the audio source 306A is greater than the audio source distance threshold, the stream selection unit 44 may then repeat in the audio stream 27 according to the techniques described in more detail above to select a single one of the audio streams 27 (e.g., in Figure 3C the example, the audio stream captured by the audio element 302B).

[0109] In Figure 3D the example, the audio element 302 is replaced by a specific example device 320A - 320J ("device 320"), where device 320A represents a dedicated microphone 320A, and devices 320B, 320C, 320D, 320G, 320H, and 320J represent smartphones. Devices 320E, 320F, and 320I may represent VR devices. Each of the devices 320 may include an audio element 302 that captures the audio stream 27 to be selected according to various aspects of the stream selection techniques described in the present disclosure.

[0110] Figure 3E is a conceptual diagram illustrating an example concert with three or more audio elements. In Figure 3EIn the example, multiple musicians are depicted on stage 323. Singer 312 is located behind audio element 310A. String section 314 is depicted behind audio element 310B. Drummer 316 is depicted behind audio element 310C. Other musicians 318 are depicted behind audio element 310D. Audio elements 310A - 301D may represent captured audio streams corresponding to the sounds received by the microphones. In some examples, microphones 310A - 310D may represent synthesized audio streams. For example, audio element 310A may represent the captured audio stream(s) mainly associated with singer 312, but the audio stream(s) may also include sounds produced by other band members such as string section 314, drummer 316, or other musicians 318, while audio element 310B may represent the captured audio stream(s) mainly associated with string section 314, but includes sounds produced by other band members. In this way, each of audio elements 310A - 310D may represent different audio stream(s).

[0111] In addition, multiple devices are depicted. These devices represent user devices located at multiple different listening positions. Headphones 321 are located near audio element 310A, but between audio element 310A and audio element 310B. Thus, according to the techniques of the present disclosure, stream selection unit 44 may select at least one of the audio streams to produce an audio experience for the user of headphones 321 similar to that of a user located at the position of headphones 321 in FIG. 3F. Similarly, VR goggles 322 are shown located behind audio element 310C and between drummer 316 and other musicians 318. The stream selection unit 44 may select at least one audio stream to produce an audio experience for the user of VR goggles 322 similar to that of a user located at the position of VR goggles 322 in FIG. 3F.

[0112] Smart glasses 324 are shown located at a relatively central position between audio elements 310A, 310C, and 310D. The stream selection unit 44 may select at least one audio stream to produce an audio experience for the user of smart glasses 324 similar to that of a user located at the position of smart glasses 324 in FIG. 3F. In addition, device 326 (which may represent any device capable of implementing the techniques of the present disclosure, such as a mobile handheld, a speaker array, headphones, VR goggles, smart glasses, etc.) is shown located in front of audio element 310B. The stream selection unit 44 may select at least one audio stream to produce an audio experience for the user of device 326 similar to that of a user located at Figure 3E the position of device 325 in. Although specific devices are discussed for specific positions, the use of any of the depicted devices may provide an indication of a different desired listening position than that depicted in Figure 3E the figure.

[0113] Figure 4A-4C is illustratedFigure 1A-1C Flowchart of an operational example of a stream selection unit 44 that controls access to at least one of a plurality of audio streams based on timing information as shown in the example of Figure 4A The example of

[0114] In many contexts, there are audio streams that may be inappropriate or offensive to some individuals. For example, during a live sports event, someone may use offensive language on the field. This may also be the case in some video games. In other live events (such as a conference), sensitive discussions may occur. By using the start time, the stream selection unit 44 of the content consumer device 14 can filter out unwanted or sensitive audio streams and exclude them from playback to the user. The timing information (such as timing metadata) can be associated with individual audio streams or privacy regions (discussed in more detail with respect to Figure 4H and 4J ).

[0115] In some cases, the source device 12 can apply the start time. For example, in a conference where a sensitive discussion will occur at a given time, the content creator or source can create and apply the start time at which the discussion will begin so that only certain individuals with appropriate permissions can hear the discussion. For others without appropriate permissions, the stream selection unit 44 can filter out or otherwise exclude the audio stream(s) for the discussion.

[0116] In other cases, such as the sports event example, the content consumer device 14 can create and apply the start time. In this way, the user can exclude offensive language during audio playback.

[0117] Now discuss the use (400) of start time information (such as start time metadata). The stream selection unit 44 can obtain the input audio stream and the metadata associated with the audio stream, including location information and start time information, and store them in the memory of the content consumer device 14 (401). The stream selection unit 44 can obtain the location information (402). As described above, the location information can be associated with capture coordinates in an acoustic space. The start time information can be associated with each stream or privacy region (will be discussed with respect to Figure 4FFor a more thorough discussion). For example, during a live event, sensitive discussions may occur, or inappropriate language may be used or topics discussed that are targeted at certain audiences. For example, if a sensitive meeting at the conference will be held at 1:00 PM GMT, the content creator or source can set a start time for the (multiple) audio streams or (multiple) privacy areas that include the audio associated with the meeting up to 1:00 PM GMT. In one example, the stream selection unit 44 can compare the start time with the current time (403), and if the start time is equal to or later than the current time, the stream selection unit 44 can filter out or otherwise exclude those audio streams or privacy areas that have an associated start time (404). In some examples, the content consumer device 14 can stop downloading the excluded audio streams.

[0118] In another example, when the stream selection unit 44 filters out or excludes an audio stream or privacy area, the content consumer device 14 can send a message to the source device 12 instructing the source device 12 to stop sending the excluded stream (405). In this way, the content consumer device does not receive the excluded stream and can save bandwidth within the transmission channel.

[0119] In one example, the audio playback system 16 (which can represent the audio playback system 16A or the audio playback system 16B for simplicity) can change the gain, enhance or attenuate the audio output, based on the start time associated with the audio stream or privacy area. In another example, the audio playback system 16 can not change the gain. The audio decoding device 34 can also combine two or more selected audio streams together (406). For example, the combination of the selected audio streams can be done by mixing or interpolation or another variation of sound field operation. The audio decoding device can output a subset of the audio streams (407).

[0120] In one example, the audio playback system 16 can allow the user to override the start time. For example, the content consumer device 14 can obtain from the user 1102 an override request, for example, to add at least one of the excluded audio streams among several audio streams (408). In the example where the content consumer device 14 sends a message to tell the source device to stop sending the excluded audio stream or privacy area (405), the content consumer device 14 will send a new message to tell the source device to start sending those audio streams or privacy areas again (409). If the start time is overridden, the audio decoding device 34 can add or combine those corresponding streams or privacy areas with the subset of the audio streams or privacy areas (410). For example, the combination of the selected audio streams can be done by mixing or interpolation or another variation of sound field operation. The audio decoding device 34 can include the selected streams in the audio output (411).

[0121] Figure 4B Is a diagram Figure 1A-1C The flowchart of the operation example in which the stream selection unit shown in the example of [] controls the access to at least one of a plurality of audio streams based on timing information. In this example, the timing information is the duration. In some examples, the timing information may be timing metadata. In some examples, the timing metadata may be included in the audio metadata. In some cases, the content creator or source may wish to provide a more complete experience during a temporary time period. For example, the content provider or source may wish to do so during an advertisement or a trial period when attempting to get users to upgrade their service level.

[0122] The stream selection unit 44 may store the input audio streams and information, such as metadata associated with them, including location information and start time metadata (421) in the memory of the content consumer device 14. The stream selection unit 44 may obtain the location information (422). The stream selection unit 44 may do so by, for example, reading the location information from the memory in the case of a single audio stream, or by calculating it in the case of a privacy area. As described above, the location information may be associated with the capture coordinates in the acoustic space. Duration metadata may be associated with each stream or privacy area and may be set to any duration. For example, in an example of providing a complete experience for a limited time period, the source device or the content consumer device may set the duration to one hour (as an example only). The stream selection unit 44 may compare the duration with a timer (423). If the timer is equal to or greater than the duration, the stream selection unit 44 may exclude the audio stream or privacy area associated with the duration, thereby selecting a subset of the audio streams (424). If the timer is less than the duration, the stream selection unit 44 will not exclude those streams or privacy areas (425).

[0123] And Figure 4A As in the example of [], if the duration is overridden (not shown for simplicity), the content consumer device 14 may send a message to the source device 12 telling it to stop sending the excluded streams and send another message to start resending the excluded streams. This can save bandwidth within the transmission channel.

[0124] In one example, the audio playback system 16 may change the gain based on the duration associated with the audio stream or privacy area, thereby enhancing or weakening the audio output. In another example, the audio playback system may not change the gain. The audio decoding device 34 may combine two or more selected audio streams together (426). For example, the combination of the selected audio streams may be done by mixing or interpolation or another variation of sound field operation. Then the audio decoding device 34 may output a subset of the audio streams (427).

[0125] By using the start time and / or duration as access control, the stream selector unit 44 can maintain access control even when there is no connection to the source device. For example, when the content consumer device 14 is offline and playing stored audio, the stream selector unit 44 can still compare the start time with the current time or the duration with a timer and implement offline access control.

[0126] Figure 4C is a flowchart (430) of an example of operations in which aspects of a stream selection technique are performed by the stream selection unit shown in the example of Figure 1A-1C The source device 12 can make different sound fields available, such as a FOA sound field, a higher-order stereophonic reverberation sound field (HOA), or a MOA sound field. A user of the content consumer device 14 can make a request (431) to change the audio experience on the content consumer device 14 via a user interface. For example, a user experiencing a FOA sound field may desire an enhanced experience and request an HOA or MOA sound field. If the content consumer device receives the necessary coefficients and is configured to change the stereophonic reverberation sound field type (432), it can then change the stereophonic reverberation sound field type (433) and the stream selector unit 44 can output an audio stream (434). If the content consumer device 14 does not receive the necessary coefficients or is not configured to change the stereophonic reverberation sound field type, the content consumer device 14 can send a request (435) to the source device 12 to make the change. The source device can make the change and send the new sound field to the content consumer device 14. The audio decoding device 34 can then receive the new sound field (436) and output an audio stream (437). The use of different types of stereophonic reverberation sound fields can also be used in conjunction with Figure 4A the start time example of Figure 4B and the duration example of

[0127] Figure 4D and 4EFIG. is a diagram further illustrating the use of timing information (such as timing metadata) for various aspects of the techniques described in the present invention. A static audio source 441 is shown, such as an activated microphone. In some examples, the static audio source 441 can be a live audio source. In other examples, the static audio source 441 can be a synthetic audio source. A dynamic audio source 442 is also shown, such as in a mobile handheld device where the user sets when to record. In some examples, the dynamic audio source can be a live audio source. In other examples, the dynamic audio source 442 can be a synthetic source. One or more of the static audio source 441 and / or the dynamic audio source 442 can capture audio information 443. A controller 444 can process the audio information 443. In Figure 4D this example, the controller 444 can be implemented in one or more processors 440 in the content consumer device 14. In Figure 4E this example, the controller 444 can be implemented in one or more processors 448 in the source device 12. The controller 444 can divide the audio information into regions, create audio streams and tag the audio streams with information (such as metadata, including location information about the audio sources 441 and 442, and region partitioning, including the boundaries of the regions, e.g., via centroid and radius data). In some examples, the controller 444 can provide the location information in a manner different from the metadata. The controller 444 can perform these functions online or offline. The controller 444 can also assign timing information (such as timing metadata) to each of the audio streams or regions, such as start time information or duration information. The controller 444 can provide burst (e.g., periodic) or fixed (e.g., continuous) audio streams and associated information, such as metadata, to the content consumer device 14. The controller 444 can also assign gains and / or clearings to be applied to the audio streams.

[0128] The stream selection unit 44 can use the timing metadata to provide burst or fixed audio streams to the user during rendering. Thus, the user's experience may change based on the timing metadata. The user can request the controller 444 to override the timing metadata and change the user's access to the audio stream or privacy region via the link 447.

[0129] Figure 4F and 4G FIG. is a diagram illustrating the use of a temporary request for more access for various aspects of the techniques described in the present disclosure. In as Figure 4FIn the example shown, content consumer device 14 renders audio streams 471, 472, and 473 represented by the depicted audio elements to user 470. Content consumer device 14 does not render audio stream 474, which is also represented by an audio element. In this case, if the user wants to temporarily enhance their experience, they can send a request via the user interface to temporarily grant them access to audio stream 474. The stream selector unit can then add audio stream 474 as shown in Figure 4G In some examples, content consumer device 14 can send a message requesting access to source device 12. In other examples, stream selection unit 44 can add audio stream 474 without sending a message to source device 12.

[0130] Figure 4H and 4I are diagrams illustrating the concept of privacy regions in accordance with various aspects of the techniques described in the present disclosure. User 480 is shown near several groups of audio elements, each group of audio elements representing an audio stream. It may be useful to authorize which streams are used to create user 480's audio experience in groups rather than individually. For example, in the example of a meeting, multiple audio elements may be receiving sensitive information. Thus, a privacy region can be created.

[0131] Source device 12 or content consumer device 14 can respectively assign an authorization level (e.g., rank) to the user and an authorization level (e.g., rank) for each privacy region. For example, controller 444 can assign a gain and clear metadata, and in this example, assign a rank for each privacy region. For example, privacy region 481 may contain audio streams 4811, 4812, and 4813. Privacy region 482 may contain audio streams 4821, 4822, and 4823. Privacy region 483 may contain audio streams 4831, 4832, and 4833. As shown in Table 1, controller 444 can label these audio streams as belonging to their respective privacy regions and can also associate a gain and clear metadata with them. As shown in Table 1, G is the gain, and N is clear or exclude. In this example, user 480's rank with respect to privacy regions 481 and 483 is 2, but the rank with respect to privacy region 482 is 3. As shown in the table, stream selection unit 44 will exclude or clear region 482 and it will not be available for rendering unless user 480 wants to override it.

[0132] The resulting rendering is shown in Figure 4H below.

[0133] Region Marker Metadata Level 461,463 4611-4613,4631-4633 G - 20dB, N = 0 2 462 4621-4623 G - N / A, N = 1 3

[0134] Table 1

[0135] Temporal information (such as temporal metadata) can be used to temporarily change the level of one or more of the privacy regions. For example, the source device 12 can assign a duration to region 462 that will raise the level to 2 for a period of time (e.g., 5 minutes). Then the stream selector unit 44 will not exclude or clear the privacy region 482 during that duration. In another example, the source device 12 can assign a start time to privacy region 461 of 12:00pm GMT, which will lower the level to 3. Then the stream selector unit 44 will exclude the privacy region 461. If the stream selector unit 44 does both, the user will receive the audio stream from privacy regions 462 and 463 (instead of 461 as Figure 4I shown).

[0136] The content consumer device 14 can use temporal information (such as temporal metadata) and comparisons as timestamps and store them in memory as a way to maintain an event record for each region.

[0137] Figure 4J and 4K are diagrams illustrating the use of layers of a service for audio rendering according to various aspects of the present disclosure. User 480 is depicted as being surrounded by audio elements. In this example, the audio elements in privacy region 482 represent a FOA sound field. The audio elements within privacy region 481 represent a HOA or MOA sound field. In Figure 4J , the content consumer device 14 is using a FOA sound field. In this example, certain individual streams or stream groups can be enabled to obtain better audio interpolation. The source device 12 may wish to make higher resolution rendering available for a temporary period, such as an advertisement or trailer for higher resolution rendering. In another example, as discussed above with respect to Figure 4C , the user can request higher resolution rendering. Then the content consumer device 14 can provide an enhanced experience, as Figure 4K shown.

[0138] Another way to utilize temporal information (such as temporal metadata) is for node modification as part of an audio scene update for a 6DOF use case as described below. Currently, audio scene updates occur instantaneously, and this is not always desirable. Figure 4L is a state transition diagram illustrating state transitions according to various aspects of the techniques described in the present disclosure. In this case, the temporal information is temporal metadata, and the temporal metadata is a delay (fireOnStartTime) and a duration (updateDuration). This temporal metadata can be included in the audio metadata.

[0139] One may wish to update the audio scene of the user experience based on a condition that occurs, but not update it immediately when the condition occurs. One may also wish to extend the time it takes for the content consumer device 14 to perform the update. In this way, the stream selection unit 44 can use the modifiable fireOnStartTime to delay the start of the update and use the updateDuration to change the time it takes to complete the update, thereby affecting the stream selection and updating the audio scene in a controlled manner. The source device 12 or the content consumer device 14 can determine or modify the fireOnStartTime and / or the updateDuration.

[0140] A condition (490) may occur, such as a nearby car starting, which may make a delayed update in the audio scene desirable. The source device 12 or the content consumer device 14 can set the delay (491) by setting the fireOnStartTime. The fireOnStartTime can be the delay time or the time after the occurrence of the condition at which the update of the audio scene starts. The stream selection unit 44 can compare the timer with the fireOnStartTime and, if the timer is equal to or greater than the fireOnStartTime, start the update of the audio scene (492). The stream selection unit 44 can update the audio scene during a transition duration (494) based on the update duration (493) and complete the update when the transition duration (494) has passed (495). The stream selection unit 44 can modify the audio scene as discussed in Table 2 below:

[0141]

[0142]

[0143] Table 2

[0144] Figure 4MDescription of a vehicle 4000 in accordance with aspects of the technology of the present disclosure. The stream selection unit 44 can sequentially update three object sources (audio sources) of the vehicle based on modifiable timing parameters fireOnStartTime and updateDuration. The content consumer device 14 or the source device 12 can set or modify these parameters. In this example, the three object sources are the engine 4001, radio 4002, and exhaust 4003 of the vehicle 4000. The source device 12 or the content consumer device 14 can assign its own local trigger time (fireOnStartTime) and duration to complete the transition (updateDuration) for each object source, the engine 4001, radio 4002, and exhaust 4003. The stream selection unit 44 can apply the fireOnStartTime regardless of the interpolation attributes mentioned in Table 2. The stream selection unit 44 can also consider the updateDuration as the effect of the interpolation attribute. For example, if the attribute is set to "true", then the stream selection unit 44 can utilize the dateDuration and update during the dateDuration, otherwise the stream selection unit 44 can immediately transition the audio scene.

[0145] The following code provides an example in accordance with aspects of the technology of the present disclosure:

[0146]

[0147]

[0148] Figure 4N Description of a moving vehicle 4100 in accordance with aspects of the technology of the present disclosure. This description represents a scenario where the stream selection unit 44 can update the audio scene in terms of location when the vehicle 4100 is navigating on a highway. In this example, there are five object sources: the engine 4101, tire 1 4102, tire 2 4103, radio 4104, and exhaust 4105. The position update after the update duration is affected is the final position after the update time. The intermediate update / interpolation between the update durations is applied as part of the audio renderer, and different interpolation schemes can be applied according to personal preference or can be situational. The following code gives an example:

[0149]

[0150]

[0151] These techniques may be particularly useful in the case of virtual teleportation. In this case, the audio signal may be perceived by the user as emanating from the direction where the virtual teleportation image is located. The virtual image may be a different passenger or driver in another vehicle or other fixed environment (e.g., school, office, or home). The virtual image (e.g., virtual passenger) may include two-dimensional avatar data or three-dimensional avatar data. When the virtual passenger speaks, it may sound as if the (multiple) virtual passengers are located at the position projected on the digital display of the head-mounted headphone device or the digital display viewed by the (multiple) cameras coupled to the head-mounted headphone device (e.g., the orientation on the screen). That is, the (multiple) virtual passengers may be coupled to a two-dimensional audio signal or a three-dimensional audio signal. The two-dimensional audio signal or three-dimensional audio signal may include one or more audio objects (e.g., human speech) spatially oriented towards the position of the virtual image relative to the screen of the digital display on the head-mounted headphone device or the digital display coupled to the head-mounted headphone device. The loudspeaker for generating the two-dimensional or three-dimensional audio signal may be installed and integrated into the head-mounted headphone device. In other embodiments, the loudspeakers may be distributed at different positions within the vehicle 4100, and may render the audio signal such that the sound from the audio stream is perceived as being located at the position where the virtual teleportation image is located. In an alternative embodiment, the "teleportation" may be the sound being teleported rather than the virtual image. Thus, a person in the vehicle or wearing the head-mounted headphone device may hear a person's sound or voice as if these people were near them, e.g., beside them, in front of them, behind them, etc.

[0152] It may be useful to include a "listener event trigger" in the audio metadata for the virtual teleportation use case, as the controller can control the listener navigation between positions through the trigger. The controller may use this listener event trigger to initiate the teleportation.

[0153] Figure 4O is a flowchart illustrating an example technique for using authorization levels to control access to at least one of a plurality of audio streams based on timing information. Now discuss the use of the authorization level (430). The stream selection unit 44 may determine the authorization level (504) of the user 1102. For example, the user 1102 may have a rank associated with them, as discussed above with respect to Figure 4H and 4I discussed. The stream selection unit 44 compares the authorization level of the user 1102 with the authorization levels of one or more privacy regions. For example, each privacy region may have an associated authorization level, as discussed above with respect to Figure 4H and 4I discussed. The stream selection unit 44 may select a subset of the plurality of audio streams based on the comparison. For example, the stream selection unit 44 may determine that the user 1102 is not authorized to access Figure 4HThe privacy area 482 and can exclude or empty the area 482. Therefore, the audio streams 4821, 4822, and 4823 will be excluded from a subset of the several audio streams.

[0154] Figure 4P is a flowchart illustrating an example technique for using a trigger and a delay to control access to at least one of several audio streams based on timing information. Now discuss the use of the trigger and the delay (510). For example, the stream selection unit 44 can detect the trigger (512). For example, the stream selection unit 44 can detect a local trigger time, such as fireOnStartTime or a listener event trigger. The stream selection unit 44 can compare the delay with a timer (514). For example, the stream selection unit 44 can compare updateDuration or other delay with the timer. If the delay is less than the timer ( Figure 4P the "no" path), then the stream selection unit 44 can continue to compare the delay with the timer. If the delay is greater than or equal to the timer, the stream selection unit can select a subset of the several audio streams (516). In this way, the stream selection unit can wait until the delay is equal to or greater than the timer to select a subset of the several audio streams.

[0155] Figure 5 is a diagram illustrating an example of a wearable device 500 that can operate according to various aspects of the techniques described in the present disclosure. In various examples, the wearable device 500 can represent a VR headset (such as the VR device 1100 described above), an AR headset, an MR headset, or any other type of extended reality (XR) headset. Augmented reality "AR" can refer to computer-rendered images or data superimposed on the real world where the user is actually located. Mixed reality "MR" can refer to computer-rendered images or data that are world locked to a specific location in the real world, or can refer to a variant of VR where some computer-rendered 3D elements and some captured real elements are combined to create an immersive experience simulating the user's physical presence in the environment. Extended reality "XR" can represent the collective term for VR, AR, and MR. More information about the XR terminology can be found in the document titled "Virtual Reality, Augmented Reality, and Mixed Reality Definitions" published by Jason Peterson on July 7, 2017.

[0156] The wearable device 500 can represent other types of devices, such as watches (including so-called "smart watches"), glasses (including so-called "smart glasses"), headphones (including so-called "wireless headphones" and "smart headphones"), smart clothing, smart jewelry, etc. Whether representing a VR device, a watch, glasses, and / or headphones, the wearable device 500 can communicate with a computing device that supports the wearable device 500 via a wired connection or a wireless connection.

[0157] In some cases, the computing device that supports the wearable device 500 can be integrated within the wearable device 500, such that the wearable device 500 can be regarded as the same device as the computing device that supports the wearable device 500. In other cases, the wearable device 500 can communicate with a separate computing device that can support the wearable device 500. In this regard, the term "support" should not be construed as requiring a separate dedicated device, but rather one or more processors configured to perform various aspects of the techniques described in this disclosure can be integrated within the wearable device 500 or within a computing device separate from the wearable device 500.

[0158] For example, when the wearable device 500 represents the VR device 1100, a separate dedicated computing device (such as a personal computer including one or more processors) can render audio and video content, and the wearable device 500 can determine translational head movements. Once the translational head movements are determined, the dedicated computing device can render audio content (as a speaker feed) based on the translational head movements according to various aspects of the techniques described in this disclosure. As another example, when the wearable device 500 represents smart glasses, the wearable device 500 can include one or more processors that both determine translational head movements (by connecting to one or more sensors of the wearable device 500) and render a speaker feed based on the determined translational head movements.

[0159] As shown, the wearable device 500 includes a rear camera, one or more directional speakers, one or more tracking and / or recording cameras, and can include one or more light-emitting diode (LED) lights. In some examples, the (multiple) LED lights can be referred to as (multiple) "ultra-bright" LED lights. Additionally, the wearable device 500 includes one or more eye-tracking cameras, high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of the wearable device 500 can include durable semi-transparent display technology and hardware.

[0160] The wearable device 500 further includes connectivity hardware, which may represent one or more network interfaces that support multi-mode connectivity such as 4G communication, 5G communication, etc. The wearable device 500 also includes an ambient light sensor, one or more cameras and night vision sensors, and one or more bone conduction transducers. In some cases, the wearable device 500 may also include one or more passive and / or active cameras with a fish-eye lens and / or a telephoto lens. It will be understood that the wearable device 500 may exhibit a variety of different form factors.

[0161] In addition, the tracking and recording cameras and other sensors can facilitate the determination of the translation distance. Although not illustrated in the examples of Figure 5 the wearable device 500 may include other types of sensors for detecting the translation distance.

[0162] Although described with respect to specific examples of wearable devices such as the VR device 1100 discussed above in the examples regarding Figure 2 and other devices illustrated in the examples of Figure 1A-1C those of ordinary skill in the art will understand that the descriptions related to Figure 1A-1C and 2 may be applicable to other examples of wearable devices. For example, other wearable devices such as smart glasses may include sensors through which translational head movements are obtained. As another example, other wearable devices such as smart watches may include sensors through which translational head movements are obtained. Thus, the techniques described in this disclosure should not be limited to a particular type of wearable device, but any wearable device may be configured to perform the techniques described in this disclosure.

[0163] Figure 6A and 6B are diagrams of example systems that illustrate various aspects of the techniques described in this disclosure. Figure 6A An example is illustrated in which the source device 12C further includes a camera 600. The camera 600 may be configured to capture video data and provide the captured raw video data to the content capture device 20. The content capture device 20 may provide the video data to another component of the source device 12C for further processing into viewport-divided portions.

[0164] In Figure 6A the example, the content consumer device 14C further includes a VR device 1100. It will be understood that in various embodiments, the VR device 1100 may be included in the content consumer device 14C or externally coupled to the content consumer device 14C. The VR device 1100 includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewports) and for rendering audio data.

[0165] Figure 6B illustrates an example in which Figure 6A the audio renderer 32 shown therein is replaced by a binaural renderer 42 that is capable of performing binaural rendering using one or more HRTFs or other functions capable of rendering to the left and right speaker feeds 43. The audio playback system 16C of the content consumer device 14D can output the left and right speaker feeds 43 to the headphones 48.

[0166] The headphones 48 can be coupled to the audio playback system 16C via a wired connection (such as a standard 3.5 mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other forms of wired connection) or wirelessly (such as via Bluetooth TM connection, a wireless network connection, etc.). The headphones 48 can reconstruct the sound field 19' represented by the audio data based on the left and right speaker feeds 43. The headphones 48 can include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the respective left and right speaker feeds 43.

[0167] Figure 7 is a block diagram of example components of one or more of the source device 12 and the content consumer device 14 shown in the example that Figure 1A-1C illustrates. In the example of Figure 7 the device 710 includes a processor 712 (which may be referred to as "one or more processors" or "(multiple) processors"), a Graphics Processing Unit (GPU) 714, a system memory 716, a display processor 718, one or more integrated speakers 740, a display 703, a user interface 720, an antenna 721, and a transceiver module 722. In an example where the device 710 is a mobile device, the display processor 718 is a Mobile Display Processor (MDP). In some examples, such as an example where the device 710 is a mobile device, the processor 712, the GPU 714, and the display processor 718 can be formed as an integrated circuit (IC).

[0168] For example, an IC can be considered a processing chip within a chip package and can be a System on Chip (SoC). In some examples, two of the processor 712, the GPU 714, and the display processor 718 can be housed together in the same IC, while the other can be housed in a different integrated circuit (e.g., a different chip package), or all three can be housed in different ICs or on the same IC. However, in an example where the device 710 is a mobile device, the processor 712, the GPU 714, and the display processor 718 may all be housed in different integrated circuits.

[0169] Examples of the processor 712, GPU 714, and display processor 718 include, but are not limited to, one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The processor 712 can be the central processing unit (CPU) of the device 710. In some examples, the GPU 714 can be specialized hardware that includes integrated and / or discrete logic circuits that provide the GPU 714 with massively parallel processing capabilities suitable for graphics processing. In some cases, the GPU 714 can also include general purpose processing capabilities and can be referred to as a general purpose GPU (GPGPU) when implementing general purpose processing tasks (e.g., non-graphics related tasks). The display processor 718 can also be specialized integrated circuit hardware that is designed to retrieve image content from the system memory 716, combine the image content into image frames, and output the image frames to the display 703.

[0170] The processor 712 can execute various types of applications. Examples of applications include web browsers, email applications, spreadsheets, video games, other applications that generate visual objects for display, or any of the application types listed in more detail above. The system memory 716 can store instructions for executing the applications. Execution of one of the applications on the processor 712 causes the processor 712 to generate graphic data for the image content to be displayed and (possibly via the integrated speaker 740) audio data 19 to be played. The processor 712 can send the graphic data of the image content to the GPU 714 for further processing based on instructions or commands sent by the processor 712 to the GPU 714.

[0171] The processor 712 can communicate with the GPU 714 according to a specific application programming interface (API). Examples of such APIs include of API, the of the Khronos Group, or OpenGL TM and OpenCL; however, aspects of the present disclosure are not limited to DirectX, OpenGL, or OpenCL APIs and can be extended to other types of APIs. Additionally, the techniques described in the present disclosure do not need to run according to an API, and the processor 712 and the GPU 714 can communicate using any processing.

[0172] The system memory 716 can be the memory of the device 710. The system memory 716 can include one or more computer-readable storage media. Examples of the system memory 716 include, but are not limited to, random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other media that can be used to carry or store desired program code in the form of instructions and / or data structures and that can be accessed by a computer or a processor.

[0173] In some examples, the system memory 716 can include instructions that cause the processor 712, the GPU 714, and / or the display processor 718 to perform the functions ascribed to the processor 712, the GPU 714, and / or the display processor 718 in this disclosure. Thus, the system memory 716 can be a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors (e.g., the processor 712, the GPU 714, and / or the display processor 718) to perform various functions.

[0174] The system memory 716 can include non-transitory storage media. The term "non-transitory" indicates that the storage media are not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be construed to mean that the system memory 716 is non-removable or that its contents are static. As an example, the system memory 716 can be removed from the device 710 and moved to another device. As another example, a memory substantially similar to the system memory 716 can be inserted into the device 710. In certain examples, the non-transitory storage media can store data that can change over time (e.g., in RAM).

[0175] The user interface 720 can represent one or more hardware or virtual (meaning a combination of hardware and software) user interfaces through which a user can connect to the device 710. The user interface 720 can include physical buttons, switches, toggle switches, lights, or their virtual versions. The user interface 720 can also include a physical or virtual keyboard, a touch interface - such as a touch screen, haptic feedback, etc.

[0176] The processor 712 may include one or more hardware units (including so-called “processing cores”) configured to perform all or part of the operations discussed above with respect to any one or more of the modules, units, or other functional components of the content creator device and / or the content consumer device. The antenna 721 and the transceiver module 722 may represent units configured to establish and maintain a connection between the source device 12 and the content consumer device 14. The antenna 721 and the transceiver module 722 may represent one or more receivers and / or one or more transmitters capable of wireless communication according to one or more wireless communication protocols, such as the fifth generation (5G) cellular standard, personal area network (PAN) protocols (such as BluetoothTM), or other open source, proprietary, or other communication standards. For example, the transceiver module 722 may receive and / or transmit wireless signals. The transceiver module 722 may represent a separate transmitter, a separate receiver, both a separate transmitter and a separate receiver, or a combined transmitter and receiver. The antenna 721 and the transceiver module 722 may be configured to receive encoded audio data. Similarly, the antenna 721 and the transceiver module 722 may be configured to transmit encoded audio data.

[0177] Figure 8A-8C illustrates Figure 1A-1C a flowchart of example operations in which the stream selection unit 44 of the example shown performs various aspects of the stream selection technique. First, referring to Figure 8A the example of, the stream selection unit 44 may obtain the audio streams 27 from all enabled audio elements, where the audio streams 27 may include corresponding audio information, such as metadata, such as ALI 45A(800). The stream selection unit 44 may perform an energy analysis for each of the audio streams 27 to calculate respective energy maps (802).

[0178] Next, the stream selection unit 44 may iterate through different combinations (804) of audio elements (defined in CM 47) based on the proximity to the audio source 308 (as defined by the audio source distances 306A and / or 306B) and the audio elements (as defined by the proximity distance discussed above). As Figure 8A shown, the audio elements may be sorted or otherwise associated with different access rights. The stream selection unit 44 may iterate in the manner described above to identify whether a larger subset or a reduced subset of the audio streams 27 is required based on the listener position represented by the DLI 45B (which is another way of referring to a “virtual position” or “device position”) and the audio element position represented by the ALI 45A (806, 808).

[0179] When a larger subset of the audio stream 27 is desired, the stream selection unit 44 may add (a)udio element(s) to the audio data 19′, or in other words, additional (a)udio stream(s) (such as when the user is closer to the audio source in the example of Figure 3A ). When a reduced subset of the audio stream 27 is desired, the stream selection unit 44 may remove (a)udio element(s) from the audio data 19′, or in other words, existing (a)udio stream(s) (such as when the user is farther from the audio source in the example of Figure 3A ).

[0180] In some examples, the stream selection unit 44 may determine that the current constellation of audio elements is the optimal set (or, in other words, the existing audio data 19′ will remain the same as the selection process described herein resulting in the same audio data 19′) (804), and the process may return to 802. However, when an audio stream is added to or removed from the audio data 19′, the stream selection unit 44 may update the CM 47 (814), generating a constellation history (815) (including position, energy map, etc.).

[0181] In addition, the stream selection unit 44 may determine whether the privacy setting enables or disables the addition of audio elements (where the privacy setting may refer to digital access rights that restrict access to one or more of the audio streams 27, e.g., by password, authorization level or rank, time, etc.) (816, 818). When the privacy setting enables the addition of audio elements, the stream selection unit 44 may add (a)udio element(s) to the updated CM 47 (which means adding (a)udio stream(s) to the audio data 19′) (820). When the privacy setting disables the addition of audio elements, the stream selection unit 44 may remove (a)udio element(s) from the updated CM 47 (which means removing (a)udio stream(s) from the audio data 19′) (822). In this way, the stream selection unit 44 can identify the set of newly enabled audio elements (824).

[0182] The stream selection unit 44 may iterate in this way and update various inputs at any given frequency. For example, the stream selection unit 44 may update the privacy setting at the user interface rate (meaning the update is driven by an update input via the user interface). As another example, the stream selection unit 44 may update the position at the sensor rate (meaning the position changes as the audio elements move). The stream selection unit 44 may also update the energy map at the audio frame rate (meaning the energy map is updated per frame).

[0183] Referring to the example of Figure 8B below, except that the stream selection unit 44 may not determine the CM 47 based on the energy map, the stream selection unit 44 may operate as described above with respect to Figure 8AOperate in the described manner. In this way, the stream selection unit 44 can obtain the audio stream 27 from all enabled audio elements, where the audio stream 27 can include corresponding audio information, such as metadata, such as ALI 45A(840). The stream selection unit 44 can determine whether the privacy setting enables or disables the addition of audio elements (where the privacy setting can refer to digital access rights that restrict access to one or more of the audio streams 27, e.g., via a password, authorization level or rank, time, etc.)(842, 844).

[0184] When the privacy setting enables the addition of audio elements, the stream selection unit 44 can add the (multiple) audio elements to the updated CM 47 (which means adding the (multiple) audio streams to the audio data 19′)(846). When the privacy setting disables the addition of audio elements, the stream selection unit 44 can remove the (multiple) audio elements from the updated CM 47 (which means removing the (multiple) audio streams from the audio data 19′)(848). In this way, the stream selection unit 44 can identify the set of newly enabled audio elements(850). The stream selection unit 44 can iterate through different combinations of audio elements in the CM 47(852) to determine the constellation history(854), which represents the audio data 19′.

[0185] The stream selection unit 44 can iterate in this way and update various inputs according to any given frequency. For example, the stream selection unit 44 can update the privacy setting at the user interface rate (meaning the update is driven by an update input via the user interface). As another example, the stream selection unit 44 can update the location at the sensor rate (meaning the location changes as the audio elements move).

[0186] Refer to the following Figure 8C example. Except that the stream selection unit 44 may not determine the CM 47 based on the audio elements enabled by the privacy setting, the stream selection unit 44 can operate in the manner described above regarding Figure 8A description. In this way, the stream selection unit 44 can obtain the audio stream 27 from all enabled audio elements, where the audio stream 27 can include corresponding audio information, such as metadata, such as ALI 45A(860). The stream selection unit 44 can perform an energy analysis for each of the audio streams 27 to calculate their respective energy maps(862).

[0187] Next, the stream selection unit 44 can iterate through different combinations of audio elements (defined in the CM 47)(864) based on the proximity to the audio source 308 (as defined by the audio source distance 306A and / or 306B) and the audio elements (as defined by the proximity distance discussed above). As Figure 8CAs shown, audio elements can be sorted or otherwise associated with different access rights. The stream selection unit 44 can iterate (866, 868) in the manner described above to identify whether a larger subset or a reduced subset of the audio stream 27 is needed based on the listener position represented by the DLI 45B (which again refers to another way of the above-mentioned "virtual position" or "device position") and the audio element position represented by the ALI 45A.

[0188] When a larger subset of the audio stream 27 is needed, the stream selection unit 44 can add (one or more) audio elements to the audio data 19', or in other words, additional (one or more) audio streams (such as when the user is closer to the audio source in the example) (870). When a reduced subset of the audio stream 27 is needed, the stream selection unit 44 can remove (one or more) audio elements from the audio data 19', or in other words, existing (one or more) audio streams (such as when the user is farther from the audio source in the example) (872). Figure 3A When a larger subset of the audio stream 27 is needed, the stream selection unit 44 can add (one or more) audio elements to the audio data 19', or in other words, additional (one or more) audio streams (such as when the user is closer to the audio source in the example) (870). When a reduced subset of the audio stream 27 is needed, the stream selection unit 44 can remove (one or more) audio elements from the audio data 19', or in other words, existing (one or more) audio streams (such as when the user is farther from the audio source in the example) (872). Figure 3A When a larger subset of the audio stream 27 is needed, the stream selection unit 44 can add (one or more) audio elements to the audio data 19', or in other words, additional (one or more) audio streams (such as when the user is closer to the audio source in the example) (870). When a reduced subset of the audio stream 27 is needed, the stream selection unit 44 can remove (one or more) audio elements from the audio data 19', or in other words, existing (one or more) audio streams (such as when the user is farther from the audio source in the example) (872).

[0189] In some examples, the stream selection unit 44 can determine that the current constellation of audio elements is the optimal set (or, in other words, the existing audio data 19' will remain the same as the selection process described herein resulting in the same audio data 19') (864), and the process can return to 862. However, when an audio stream is added to or removed from the audio data 19', the stream selection unit 44 can update the CM 47 (874) and generate a constellation history (875).

[0190] The stream selection unit 44 can iterate in this way and update various inputs at any given frequency. For example, the stream selection unit 44 can update the position at the sensor rate (meaning the position changes due to the movement of the audio elements). The stream selection unit 44 can also update the energy map at the audio frame rate (meaning the energy map is updated per frame).

[0191] Figure 9 An example of a wireless communication system 100 according to aspects of the present disclosure is illustrated. The wireless communication system 100 includes a base station 105, a UE 115, and a core network 130. In some examples, the wireless communication system 100 can be a Long Term Evolution (LTE) network, an Advanced LTE (LTE-A) network, an LTE-A Pro network, a 5th generation cellular network, or a New Radio (NR) network. In some cases, the wireless communication system 100 can support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low-latency communication, or communication with low-cost and low-complexity devices.

[0192] Base station 105 may communicate wirelessly with UE 115 via one or more base station antennas. The base station 105 described herein may include or may be referred to by those skilled in the art as a base station transceiver, radio base station, access point, radio transceiver, NodeB, eNodeB (eNB), next-generation NodeB, or gigabit NodeB (any of which may be referred to as a gNB), home NodeB, home eNodeB, or some other suitable term. The wireless communication system 100 may include different types of base stations 105 (e.g., macro cell base stations or small cell base stations). The UE 115 described herein may be capable of communicating with various types of base stations 105 and network devices, including macro eNBs, small cell eNBs, gNBs, relay base stations, and the like.

[0193] Each base station 105 may be associated with a specific geographic coverage area 110 in which communication with respective UEs 115 is supported. Each base station 105 may provide communication coverage for the corresponding geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 may utilize one or more carriers. The communication link 125 shown in the wireless communication system 100 may include an uplink transmission from the UE 115 to the base station 105, or a downlink transmission from the base station 105 to the UE 115. The downlink transmission may also be referred to as a forward link transmission, and the uplink transmission may also be referred to as a reverse link transmission.

[0194] The geographic coverage area 110 of the base station 105 may be divided into sectors that form part of the geographic coverage area 110, and each sector may be associated with a cell. For example, each base station 105 may provide communication coverage for a macro cell, a small cell, a hotspot, or other types of cells or various combinations thereof. In some examples, the base station 105 may be mobile and thus provide communication coverage for a mobile geographic coverage area 110. In some examples, different geographic coverage areas 110 associated with different technologies may overlap, and the overlapping geographic coverage areas 110 associated with different technologies may be supported by the same base station 105 or different base stations 105. The wireless communication system 100 may include, for example, a heterogeneous LTE / LTE-A / LTE-A Pro, fifth-generation, or NR network, where different types of base stations 105 provide coverage for various geographic coverage areas 110.

[0195] The UEs 115 can be dispersed throughout the wireless communication system 100, and each UE 115 can be stationary or mobile. The UE 115 can also be referred to as a mobile device, a wireless device, a remote device, a handheld device, or a subscriber device, or some other suitable term, where "device" can also be referred to as a unit, a station, a terminal, or a client. The UE 115 can also be a personal electronic device, such as a cellular phone, a personal digital assistant (PDA), a tablet computer, a laptop computer, or a personal computer. In an example of the present disclosure, the UE 115 can be any audio source among the audio sources described in the present disclosure, including a VR headset, an XR headset, an AR headset, a vehicle, a smartphone, a microphone, a microphone array, or any other device including a microphone or capable of transmitting a captured and / or synthesized audio stream. In some examples, the UE 115 can also refer to a wireless local loop (WLL) station, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, or a machine type communication (MTC) device, etc., which can be implemented in various articles, such as appliances, vehicles, meters, etc.

[0196] Some UEs 115 (such as MTC or IoT devices) can be low-cost or low-complexity devices and can provide automatic communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC can refer to a data communication technology that allows devices to communicate with each other or with the base station 105 without human intervention. In some examples, M2M communication or MTC can include communication from devices that exchange and / or use audio metadata, which can include timing metadata for affecting an audio stream and / or an audio source.

[0197] In some cases, the UE 115 can also be capable of directly communicating with other UEs 115 (e.g., using a peer-to-peer (P2P) or device-to-device (D2D) protocol). One or more UEs in a group of UEs 115 that utilize D2D communication can be within the geographical coverage area 110 of the base station 105. Other UEs 115 in the group can be outside the geographical coverage area 110 of the base station 105, or in other cases cannot receive transmissions from the base station 105. In some cases, multiple groups of UEs 115 that communicate via D2D communication can utilize a one-to-many (1:M) system, in which each UE 115 transmits to each other UE 115 in the group. In some cases, the base station 105 facilitates resource scheduling for D2D communication. In other cases, D2D communication is performed between UEs 115 without the participation of the base station 105.

[0198] Base station 105 can communicate with the core network 130 and can communicate with each other. For example, base station 105 can interface with core network 130 via a backhaul link 132 (e.g., via S1, N2, N3, or other interfaces). Base station 105 can communicate with each other directly (e.g., directly between base stations 105) or indirectly (e.g., via core network 130) via a backhaul link 134 (e.g., via X2, Xn, or other interfaces).

[0199] In some cases, wireless communication system 100 can utilize both licensed and unlicensed radio frequency bands. For example, wireless communication system 100 can employ licensed-assisted access (LAA), unlicensed LTE (LTE-U) radio access technology, or NR technology in an unlicensed band such as the 5 GHz industrial, scientific, and medical (ISM) band. When operating in an unlicensed radio frequency band, wireless devices such as base station 105 and UE 115 can employ a listen-before-talk (LBT) procedure to ensure that the frequency channel is clear before transmitting data. In some cases, operation in the unlicensed band can be based on a carrier aggregation configuration together with a component carrier operating in a licensed band (e.g., LAA). Operation in the unlicensed spectrum can include downlink transmission, uplink transmission, peer-to-peer transmission, or a combination of these transmissions. Duplexing in the unlicensed spectrum can be based on frequency-division duplexing (FDD), time-division duplexing (TDD), or a combination of both.

[0200] According to the techniques of the present disclosure, a single audio stream can be restricted from rendering or can be temporarily rendered based on timing information such as time or duration. For better audio interpolation, certain single audio streams or clusters of audio streams can be enabled or disabled for a fixed duration. Thus, the techniques of the present disclosure provide a flexible way to control access to audio streams based on time.

[0201] It should be noted that the methods described herein describe possible embodiments, and the operations and steps can be rearranged or otherwise modified, and other embodiments are possible. Additionally, aspects from two or more of these methods can be combined.

[0202] It should be recognized that, according to examples, certain actions or events of any of the techniques described herein can be performed in a different order, can be added, combined, or entirely omitted (e.g., not all described actions or events are necessary to practice the technique). Additionally, in certain examples, actions or events can be performed concurrently rather than sequentially, such as by multithreading, interrupt processing, or multiple processors.

[0203] In some examples, a VR device (or streaming device) may communicate an exchange message to an external device using a network interface coupled to a memory of the VR / streaming device, where the exchange message is associated with multiple available representations of an acoustic field. In some examples, the VR device may use an antenna coupled to the network interface to receive a wireless signal that includes a data packet, an audio packet, a video protocol, or a transport protocol data associated with multiple available representations of the acoustic field. In some examples, one or more microphone arrays may capture the acoustic field.

[0204] In some examples, multiple available representations of the acoustic field stored to a memory device may include multiple object-based representations of the acoustic field, higher-order stereophonic reverberation representations of the acoustic field, hybrid-order stereophonic reverberation representations of the acoustic field, a combination of an object-based representation of the acoustic field and a higher-order stereophonic reverberation representation of the acoustic field, a combination of an object-based representation of the acoustic field and a hybrid-order stereophonic reverberation representation of the acoustic field, or a combination of a hybrid-order representation of the acoustic field and a higher-order stereophonic reverberation representation of the acoustic field.

[0205] In some examples, one or more of the acoustic field representations in the multiple available representations of the acoustic field may include at least one high-resolution region and at least one lower-resolution region, and where a selected representation based on a steering angle provides higher spatial accuracy for at least one high-resolution region and lower spatial accuracy for the lower-resolution region.

[0206] The present disclosure includes the following examples.

[0207] Example 1 A device configured to play one or more of a plurality of audio streams, comprising: a memory configured to store timing metadata, the plurality of audio streams and corresponding audio metadata, and location information associated with coordinates of an acoustic space, where a respective one of the plurality of audio streams is captured in the acoustic space; and one or more processors coupled to the memory and configured to: select a subset of the plurality of audio streams based on the timing metadata and the location information, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0208] Example 2 The device according to Example 1, wherein the one or more processors are further configured to obtain the location information.

[0209] Example 3 The device according to Example 2, wherein the excluded stream is associated with one or more privacy regions, and the one or more processors obtain the location information by determining the location information.

[0210] Example 4 The device according to Example 2, wherein the one or more processors obtain the location information by reading the location information from the memory.

[0211] Example 5. A device according to any combination of Examples 1-4, wherein one or more processors are further configured to combine at least two in a subset of a plurality of audio streams.

[0212] Example 6. A device according to Example 5, wherein one or more processors combine at least two in a subset of a plurality of audio streams by at least one of mixing or interpolation.

[0213] Example 7. A device according to any combination of Examples 1-6, wherein one or more processors are further configured to change the gain of one or more in a subset of a plurality of audio streams.

[0214] Example 8. A device according to any combination of Examples 1-7, wherein the timing metadata includes a start time when at least one of a plurality of audio streams includes audio content.

[0215] Example 9. A device according to Example 8, wherein one or more processors are configured to: compare the start time with the current time; and when the start time is equal to or greater than the current time, select a subset of a plurality of audio streams.

[0216] Example 10. A device according to any combination of Examples 1-9, wherein the timing metadata includes a duration of at least one of a plurality of audio streams.

[0217] Example 11. A device according to Example 10, wherein one or more processors are configured to: compare the duration with a timer; and when the duration is equal to or greater than the timer, select a subset of a plurality of audio streams.

[0218] Example 12. A device according to Example 10, wherein one or more processors are further configured to: select a second subset of a plurality of audio streams based on location information, the second subset of a plurality of audio streams excluding at least one of a plurality of audio streams; and interpolate between the subset of a plurality of audio streams and the second subset of a plurality of audio streams by the duration.

[0219] Example 13. A device according to any combination of Examples 1-12, wherein one or more processors are further configured to: obtain a request from a user to select a subset of a plurality of audio streams; and select a subset of a plurality of audio streams based on the user request, location information, and timing metadata.

[0220] Example 14. A device according to any combination of Examples 1-13, wherein the timing metadata is received from a source device.

[0221] Example 15. A device according to Examples 1-13, wherein one or more processors are further configured to generate timing metadata.

[0222] Example 16. The apparatus according to Examples 1-15, wherein one or more processors are configured to: obtain a request for one of a plurality of stereo reverberation sound field types from a user; and reproduce a corresponding sound field based on the request for one of the plurality of stereo reverberation sound field types and a plurality of audio streams or a subset of the plurality of audio streams.

[0223] Example 17. The apparatus according to Example 16, wherein the plurality of stereo reverberation sound field types includes at least two of a first-order stereo reverberation sound field (FOA), a higher-order stereo reverberation sound field (HOA), and a mixed-order stereo reverberation sound field (MOA).

[0224] Example 18. The apparatus according to any combination of Examples 1-17, further comprising a display device.

[0225] Example 19. The apparatus according to Example 18, further comprising a microphone, wherein the one or more processors are further configured to receive a voice command from the microphone and control the display device based on the voice command.

[0226] Example 20. The apparatus according to any combination of Examples 1-19, further comprising one or more speakers.

[0227] Example 21. The apparatus according to any combination of Examples 1-20, wherein the apparatus comprises an extended reality headset, and wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0228] Example 22. The apparatus according to any combination of Examples 1-20, wherein the apparatus comprises an extended reality headset, and wherein the acoustic space comprises a virtual world.

[0229] Example 23. The apparatus according to any combination of Examples 1-22, further comprising a head-mounted display configured to present the acoustic space.

[0230] Example 24. The apparatus according to any combination of Examples 1-20, wherein the apparatus comprises a mobile handheld device.

[0231] Example 25. The apparatus according to any combination of Examples 1-24, further comprising a wireless transceiver coupled to one or more processors and configured to receive a wireless signal.

[0232] Example 26. The apparatus according to Example 25, wherein the wireless signal is Bluetooth.

[0233] Example 27. The apparatus according to Example 25, wherein the wireless signal is 5G.

[0234] Example 28. The apparatus according to any combination of Examples 1-27, wherein the apparatus comprises a vehicle.

[0235] Example 29. An apparatus according to any combination of Examples 1-25, wherein the timing metadata includes a latency, and wherein one or more processors are further configured to: detect a trigger; compare the latency with a timer; and wait until the latency is equal to or greater than the timer to select a subset of a plurality of audio streams.

[0236] Example 30. A method of playing one or more of a plurality of audio streams, comprising: storing, by a memory, timing metadata, a plurality of audio streams and corresponding audio metadata, and location information associated with coordinates of an acoustic space, wherein a respective one of the plurality of audio streams is captured in the acoustic space; and selecting, by the one or more processors, a subset of the plurality of audio streams based on the timing metadata and the location information, wherein the subset of the plurality of audio streams excludes at least one of the plurality of audio streams.

[0237] Example 31. The method according to Example 30, further comprising obtaining, by the one or more processors, the location information.

[0238] Example 32. The method according to Example 31, wherein the excluded stream is associated with one or more privacy regions, and the location information is obtained by determining the location information.

[0239] Example 33. The method according to Example 31, wherein the location information is obtained by reading the location information from the memory.

[0240] Example 34. The method according to any combination of Examples 31-33, further comprising combining, by the one or more processors, at least two of the subset of the plurality of audio streams.

[0241] Example 35. The method according to Example 34, wherein at least two of the subset of the plurality of audio streams are combined by at least one of mixing or interpolation.

[0242] Example 36. The method according to any combination of Examples 30-35, further comprising changing, by the one or more processors, a gain of one or more of the subset of the plurality of audio streams.

[0243] Example 37. The method according to any combination of Examples 30-36, wherein the timing metadata includes a start time when at least one of the plurality of audio streams includes audio content.

[0244] Example 38. The method according to Example 37, further comprising: comparing, by the one or more processors, the start time with the current time; and selecting, by the one or more processors, a subset of the plurality of audio streams when the start time is equal to or greater than the current time.

[0245] Example 39. The method according to any combination of Examples 30-38, wherein the timing metadata includes a duration of at least one of the plurality of audio streams.

[0246] Example 40. The method according to Example 39 further includes: comparing, by one or more processors, a duration with a timer; and when the duration is equal to or greater than the timer, selecting, by the one or more processors, a subset of a plurality of audio streams.

[0247] Example 41. The method according to Example 39 further includes: selecting, by the one or more processors, a second subset of a plurality of audio streams based on location information, the second subset of the plurality of audio streams excluding at least one of the plurality of audio streams; and interpolating, by the one or more processors, between the subset of the plurality of audio streams and the second subset of the plurality of audio streams by the duration.

[0248] Example 42. The method according to any combination of Examples 30-41 further includes: obtaining, from a user, a request to select a subset of a plurality of audio streams; and selecting, by the one or more processors, a subset of a plurality of audio streams based on the user request, location information, and timing metadata.

[0249] Example 43. The method according to any combination of Examples 30-42, wherein the timing metadata is received from a source device.

[0250] Example 44. The method according to any combination of Examples 30-42 further includes generating, by the one or more processors, the timing metadata.

[0251] Example 45. The method according to any combination of Examples 30-44 further includes: obtaining, from a user, a request for one of a plurality of stereo reverberation field types; and reproducing, by the one or more processors, a corresponding sound field based on the request for one of a plurality of stereo reverberation field types and a plurality of audio streams or a subset of a plurality of audio streams.

[0252] Example 46. The method according to Example 45, wherein the plurality of stereo reverberation field types includes at least two of a first-order stereo reverberation field (FOA), a higher-order stereo reverberation field (HOA), and a mixed-order stereo reverberation field (MOA).

[0253] Example 47. The method according to any combination of Examples 30-46 further includes a microphone that receives a voice command and controls, by one or more processors, a display device based on the voice command.

[0254] Example 48. The method according to any combination of Examples 30-47 further includes outputting a subset of a plurality of audio streams to the one or more speakers.

[0255] Example 49. The method according to any combination of Examples 30-48, wherein the acoustic space includes a scene represented by video data captured by a camera.

[0256] Example 50. A method according to any combination of Examples 30-48, wherein the acoustic space includes a virtual world.

[0257] Example 51. A method according to any combination of Examples 30-50, further comprising presenting, by the one or more processors, the acoustic space on a head-mounted device.

[0258] Example 52. A method according to any combination of Examples 30-51, further comprising presenting, by the one or more processors, the acoustic space on a mobile handheld device.

[0259] Example 53. A method according to any combination of Examples 30-52, further comprising receiving a wireless signal.

[0260] Example 54. A method according to Example 53, wherein the wireless signal is Bluetooth.

[0261] Example 55. A method according to Example 53, wherein the wireless signal is 5G.

[0262] Example 56. A method according to any combination of Examples 30-55, further comprising presenting, by the one or more processors, the acoustic space inside a vehicle.

[0263] Example 57. A method according to any combination of Examples 30-56, wherein the timing metadata includes a delay, and wherein the method further comprises: detecting, by the one or more processors, a trigger; comparing, by the one or more processors, the delay with a timer; and waiting until the delay is equal to or greater than the timer to select a subset of the plurality of audio streams.

[0264] Example 58. A device configured to play one or more of a plurality of audio streams, the device comprising: means for storing timing metadata, the plurality of audio streams and corresponding audio metadata, and location information associated with coordinates of an acoustic space, wherein a respective one of the plurality of audio streams is captured in the acoustic space; and means for selecting a subset of the plurality of audio streams based on the timing metadata and the location information, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0265] Example 59. The device according to Example 58, further comprising means for obtaining the location information.

[0266] Example 60. The device according to Example 59, wherein the excluded stream is associated with one or more privacy regions, and the location information is obtained by determining the location information.

[0267] Example 61. The device according to Example 59, wherein the location information is obtained by reading the location information from the memory.

[0268] The apparatus according to any combination of Examples 58 - 60 further includes components for combining at least two of a plurality of subsets of audio streams.

[0269] The apparatus according to Example 62, wherein at least two of a plurality of subsets of audio streams are combined by at least one of mixing or interpolation.

[0270] The apparatus according to any combination of Examples 58 - 63 further includes components for changing the gain of one or more of a plurality of subsets of audio streams.

[0271] The apparatus according to any combination of Examples 58 - 64, wherein the timing metadata includes a start time when at least one of a plurality of audio streams includes audio content.

[0272] The apparatus according to Example 65 further includes components for comparing the start time with the current time; and components for selecting a subset of a plurality of audio streams when the start time is equal to or greater than the current time.

[0273] The apparatus according to any combination of Examples 58 - 66, wherein the timing metadata includes a duration of at least one of a plurality of audio streams.

[0274] The apparatus according to Example 67 further includes components for comparing the duration with a timer; and components for selecting a subset of a plurality of audio streams when the duration is equal to or greater than the timer.

[0275] The apparatus according to Example 67 further includes: components for selecting a second subset of a plurality of audio streams based on location information, the second subset of a plurality of audio streams excluding at least one of a plurality of audio streams; and components for interpolating between a subset of a plurality of audio streams and the second subset of a plurality of audio streams by the duration.

[0276] The apparatus according to any combination of Examples 58 - 69 further includes: components for obtaining a request from a user to select a subset of a plurality of audio streams; and components for selecting a subset of a plurality of audio streams based on the user request, location information, and timing metadata.

[0277] The apparatus according to any combination of Examples 58 - 70, wherein the timing metadata is received from a source device.

[0278] The apparatus according to any combination of Examples 58 - 70 further includes components for generating the timing metadata.

[0279] Apparatus according to any combination of Examples 58 - 72, further comprising: means for obtaining from a user a request for one of a plurality of stereophonic reverberant field types; and means for reproducing a corresponding sound field based on the request for one of the plurality of stereophonic reverberant field types and a plurality of audio streams or a subset of the plurality of audio streams.

[0280] Apparatus according to Example 73, wherein the plurality of stereophonic reverberant field types includes at least two of a first - order stereophonic reverberant field (FOA), a higher - order stereophonic reverberant field (HOA), and a mixed - order stereophonic reverberant field (MOA).

[0281] Apparatus according to any combination of Examples 58 - 74, further comprising means for receiving a voice command and means for controlling a display device based on the voice command.

[0282] Apparatus according to any combination of Examples 58 - 75, further comprising means for outputting a subset of the plurality of audio streams to the one or more loudspeakers.

[0283] Apparatus according to any combination of Examples 58 - 76, wherein the acoustic space includes a scene represented by video data captured by a camera.

[0284] Apparatus according to any combination of Examples 58 - 76, wherein the acoustic space includes a virtual world.

[0285] Apparatus according to any combination of Examples 58 - 78, further comprising means for presenting the acoustic space on a head - mounted device.

[0286] Apparatus according to any combination of Examples 58 - 78, further comprising means for presenting the acoustic space on a mobile handheld device.

[0287] Apparatus according to any combination of Examples 58 - 80, further comprising means for receiving a wireless signal.

[0288] Apparatus according to Example 81, wherein the wireless signal is Bluetooth.

[0289] Apparatus according to Example 81, wherein the wireless signal is 5G.

[0290] Apparatus according to any combination of Examples 58 - 83, further comprising means for presenting the acoustic space inside a vehicle.

[0291] Apparatus according to any combination of Examples 58 - 84, wherein the timing metadata includes a latency, and wherein the apparatus further comprises: means for detecting a trigger; means for comparing the latency with a timer; and means for waiting until the latency is equal to or greater than the timer to select a subset of the plurality of audio streams.

[0292] Example 86 A non - transitory computer - readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: store timing metadata, the plurality of audio streams and corresponding audio metadata, and location information associated with coordinates of an acoustic space, wherein a respective one of the plurality of audio streams is captured in the acoustic space; and select a subset of the plurality of audio streams based on the timing metadata and the location information, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0293] Example 87 The non - transitory computer - readable storage medium according to Example 86, further comprising instructions that, when executed, cause one or more processors to obtain the location information.

[0294] Example 88 The non - transitory computer - readable storage medium according to Example 87, wherein the excluded stream is associated with one or more privacy regions, and the one or more processors obtain the location information by determining the location information.

[0295] Example 89 The non - transitory computer - readable storage medium according to Example 87, wherein the one or more processors obtain the location information by reading the location information from the memory.

[0296] Example 90 The non - transitory computer - readable storage medium according to any combination of Examples 86 - 89, further comprising instructions that, when executed, cause one or more processors to combine at least two of the subset of the plurality of audio streams.

[0297] Example 91 The non - transitory computer - readable storage medium according to Example 90, wherein at least two of the subset of the plurality of audio streams are combined by at least one of mixing or interpolation.

[0298] Example 92 The non - transitory computer - readable storage medium according to any combination of Examples 86 - 91, further comprising instructions that, when executed, cause one or more processors to change the gain of one or more of the subset of the plurality of audio streams.

[0299] Example 93 The non - transitory computer - readable storage medium according to any combination of Examples 86 - 92, wherein the timing metadata includes a start time when at least one of the plurality of audio streams includes audio content.

[0300] Example 94. The non-transitory computer-readable storage medium according to Example 93 further includes instructions that, when executed, cause one or more processors to perform the following operations: compare a start time with a current time; and when the start time is equal to or greater than the current time, select a subset of a plurality of audio streams.

[0301] Example 95. The non-transitory computer-readable storage medium according to any combination of Examples 86-94, wherein the timing metadata includes the duration of at least one of a plurality of audio streams.

[0302] Example 96. The non-transitory computer-readable storage medium according to Example 95 further includes instructions that, when executed, cause one or more processors to perform the following operations: compare the duration with a timer; and when the duration is equal to or greater than the timer, select a subset of a plurality of audio streams.

[0303] Example 97. The non-transitory computer-readable storage medium according to Example 95 further includes instructions that, when executed, cause one or more processors to perform the following operations: select a second subset of a plurality of audio streams based on location information, the second subset of the plurality of audio streams excluding at least one of the plurality of audio streams; and interpolate between the subset of the plurality of audio streams and the second subset of the plurality of audio streams by the duration.

[0304] Example 98. The non-transitory computer-readable storage medium according to any combination of Examples 86-97 further includes instructions that, when executed, cause one or more processors to perform the following operations: obtain a request from a user to select a subset of a plurality of audio streams; and select a subset of the plurality of audio streams based on the user request, location information, and timing metadata.

[0305] Example 99. The non-transitory computer-readable storage medium according to any combination of Examples 86-98, wherein the timing metadata is received from a source device.

[0306] Example 100. The non-transitory computer-readable storage medium according to Examples 86-99 further includes instructions that, when executed, cause one or more processors to generate timing metadata.

[0307] Example 101. The non-transitory computer-readable storage medium according to Examples 86-100 further includes instructions that, when executed, cause one or more processors to perform the following operations:

[0308] obtain a request from a user for one of a plurality of stereo reverberation field types; and

[0309] reproduce a corresponding sound field based on the request for one of the plurality of stereo reverberation field types and the plurality of audio streams or the subset of the plurality of audio streams.

[0310] Example 102 The non-transitory computer-readable storage medium according to Example 101, wherein the plurality of stereo reverberation field types includes at least two of a first-order stereo reverberation field (FOA), a higher-order stereo reverberation field (HOA), and a mixed-order stereo reverberation field (MOA).

[0311] Example 103 The non-transitory computer-readable storage medium according to any combination of Examples 86-102, further comprising instructions that, when executed, cause one or more processors to receive a voice command from a microphone and control a display device based on the voice command.

[0312] Example 104 The non-transitory computer-readable storage medium according to any combination of Examples 86-103, further comprising instructions that, when executed, cause one or more processors to output a subset of the plurality of audio streams to one or more speakers.

[0313] Example 105 The non-transitory computer-readable storage medium according to any combination of Examples 86-104, wherein the acoustic space includes a scene represented by video data captured by a camera.

[0314] Example 106 The non-transitory computer-readable storage medium according to any combination of Examples 86-104, wherein the acoustic space includes a virtual world.

[0315] Example 107 The non-transitory computer-readable storage medium according to any combination of Examples 86-106, further comprising instructions that, when executed, cause one or more processors to present the acoustic space on a head-mounted device.

[0316] Example 108 The non-transitory computer-readable storage medium according to any combination of Examples 86-107, further comprising instructions that, when executed, cause one or more processors to present the acoustic space on a mobile handheld device.

[0317] Example 109 The non-transitory computer-readable storage medium according to any combination of Examples 86-108, further comprising instructions that, when executed, cause one or more processors to receive a wireless signal.

[0318] Example 110 The non-transitory computer-readable storage medium according to Example 109, wherein the wireless signal is Bluetooth.

[0319] Example 111 The non-transitory computer-readable storage medium according to Example 109, wherein the wireless signal is 5G.

[0320] Example 112 The non-transitory computer-readable storage medium according to any combination of Examples 86-111, further comprising instructions that, when executed, cause one or more processors to present the acoustic space inside a vehicle.

[0321] Example 113. A non-transitory computer-readable storage medium according to any combination of Examples 86-112, wherein the timing metadata includes a latency, and the non-transitory computer-readable storage medium further includes instructions that, when executed, cause one or more processors to perform the following operations: detect a trigger; compare the latency with a timer; and wait until the latency is equal to or greater than the timer to select a subset of a plurality of audio streams.

[0322] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to tangible media such as data storage media, or communication media including, for example, any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0323] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media exclude connections, carrier waves, signals, or other transitory media, but rather are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media. Exclusions

[0324] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Similarly, the techniques can be fully implemented in one or more circuits or logic elements.

[0325] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses including wireless handsets, integrated circuits (ICs) or IC sets (e.g., chip sets). Various components, modules, or units are described in this invention to emphasize functional aspects of a device configured to perform the disclosed techniques, but need not necessarily be implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units in conjunction with suitable software and / or firmware, the collection of interoperating hardware units including one or more processors as described above.

[0326] A variety of examples have been described. These examples as well as other examples are within the scope of the appended claims.

Claims

1. A device configured to play one or more of a plurality of audio streams, the device comprising: a memory configured to store timing information and the plurality of audio streams, wherein the timing information includes a delay and an update duration; and one or more processors coupled to the memory and configured to: control access to at least one of the plurality of audio streams based on the timing information; detect a trigger configured to trigger an audio scene update; compare the delay with a timer; and wait until the delay is equal to or greater than the timer, initiate a change to a plurality of audio sources for the plurality of audio streams, and the plurality of audio sources are changed during the update duration to select a subset of the plurality of audio streams.

2. The device according to claim 1, wherein the memory is further configured to store position information associated with coordinates of an acoustic space, wherein a respective one of the plurality of audio streams is captured or synthesized in the acoustic space.

3. The device according to claim 1, wherein the one or more processors are configured to control access to at least one of the plurality of audio streams by selecting a subset of the plurality of audio streams, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

4. The device according to claim 3, wherein the excluded stream is associated with one or more privacy regions.

5. The device according to claim 4, wherein the one or more processors are further configured to: determine a user's authorization level; compare the user's authorization level with an authorization level of the one or more privacy regions; and select a subset of the plurality of audio streams based on the comparison.

6. The device according to claim 3, wherein The one or more processors are further configured to: obtain from a user an override request to add at least one excluded audio stream of the plurality of audio streams; and add the at least one excluded audio stream for a limited period of time based on the override request.

7. The device according to claim 1, wherein the one or more processors are configured to control access to at least one of the plurality of audio streams by not downloading or receiving at least one of the plurality of audio streams based on the timing information.

8. The device according to claim 1, wherein the timing information includes a start time when at least one of the plurality of audio streams includes audio content.

9. The device according to claim 8, wherein, The one or more processors are configured to: compare the start time with the current time; and select a subset of the plurality of audio streams when the start time is equal to or greater than the current time.

10. The device according to claim 1, wherein the timing information includes a duration of at least one of the plurality of audio streams.

11. The device according to claim 10, wherein, The one or more processors are configured to: compare the duration with a timer; and select a subset of the plurality of audio streams when the duration is equal to or greater than the timer.

12. The device according to claim 1, wherein The one or more processors are configured to: obtain from a user a request for one of a plurality of stereo reverberation field types; and Reproduce a corresponding sound field based on the request for one of a plurality of stereophonic reverberant sound field types and the plurality of audio streams or a subset of the plurality of audio streams, wherein the plurality of stereophonic reverberant sound field types includes at least two of a first-order stereophonic reverberant sound field (FOA), a higher-order stereophonic reverberant sound field (HOA), and a mixed-order stereophonic reverberant sound field (MOA).

13. The apparatus according to claim 1, wherein the one or more processors are further configured to combine at least two of the plurality of audio streams by at least one of mixing or interpolation or another transformation of sound field manipulation.

14. The apparatus according to claim 1, wherein the one or more processors are further configured to change the gain of one or more of the plurality of audio streams.

15. The apparatus according to claim 1, further comprising a display device.

16. The apparatus according to claim 15, further comprising a microphone, wherein the one or more processors are further configured to receive voice commands from the microphone and control the display device based on the voice commands.

17. The apparatus according to claim 1, further comprising one or more speakers.

18. The apparatus according to claim 1, wherein the apparatus comprises an extended reality headset, and wherein the acoustic space includes a scene represented by video data captured by a camera.

19. The apparatus according to claim 1, wherein the apparatus comprises an extended reality headset, and wherein the acoustic space includes a virtual world.

20. The apparatus according to claim 1, further comprising a head-mounted display configured to present the acoustic space.

21. The apparatus according to claim 1, wherein the apparatus comprises one of a mobile handheld device or a vehicle.

22. The apparatus according to claim 1, further comprising a wireless transceiver coupled to the one or more processors and configured to receive wireless signals.

23. A method of playing one or more of a plurality of audio streams, comprising: storing, by a memory, timing information and the plurality of audio streams, wherein the timing information includes a delay and an update duration; and controlling access to at least one of the plurality of audio streams based on the timing information; detecting a trigger configured to trigger an audio scene update; comparing the delay with a timer; and waiting until the delay is equal to or greater than the timer, starting a change to a plurality of audio sources for the plurality of audio streams, and changing the plurality of audio sources during the update duration to select a subset of the plurality of audio streams.

24. The method according to claim 23, further comprising storing position information associated with coordinates of an acoustic space, wherein a respective one of the plurality of audio streams is captured or synthesized in the acoustic space.

25. The method according to claim 23, wherein controlling access to at least one of the plurality of audio streams includes selecting a subset of the plurality of audio streams, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

26. The method according to claim 25, wherein the excluded stream is associated with one or more privacy regions.

27. The method according to claim 26, further comprising: determining an authorization level of a user; comparing the authorization level of the user with the authorization level of the one or more privacy regions; and selecting a subset of the plurality of audio streams based on the comparison.

28. The method according to claim 25, further comprising: obtaining a override request from the user for adding at least one excluded audio stream of the plurality of audio streams; and adding the at least one excluded audio stream for a limited period of time based on the override request.

29. The method according to claim 23, wherein controlling access to at least one of the plurality of audio streams comprises not downloading or receiving at least one of the plurality of audio streams based on the timing information.

30. The method according to claim 23, wherein the timing information comprises a start time when at least one of the plurality of audio streams comprises audio content.

31. The method according to claim 30, further comprising: comparing the start time with the current time; and selecting a subset of the plurality of audio streams when the start time is equal to or greater than the current time.

32. The method according to claim 23, wherein the timing information comprises a duration of at least one of the plurality of audio streams.

33. The method according to claim 32, further comprising: comparing the duration with a timer; and selecting a subset of the plurality of audio streams when the duration is equal to or greater than the timer.

34. The method according to claim 23, further comprising: obtaining a request from the user for one of a plurality of stereo reverberation field types; and reproducing a corresponding sound field based on the request for one of the plurality of stereo reverberation field types and the plurality of audio streams or the subset of the plurality of audio streams, wherein the plurality of stereo reverberation field types comprises at least two of a first-order stereo reverberation field (FOA), a higher-order stereo reverberation field (HOA), and a mixed-order stereo reverberation field (MOA).

35. The method according to claim 23, further comprising combining at least two of the plurality of audio streams by at least one of mixing or interpolation or another variation of sound field manipulation.

36. The method according to claim 23, further comprising changing a gain of one or more of the plurality of audio streams.

37. The method according to claim 23, further comprising receiving a voice command by a microphone and controlling a display device based on the voice command.

38. The method according to claim 23, further comprising outputting at least one of the plurality of audio streams to one or more speakers.

39. The method according to claim 23, wherein the acoustic space comprises a scene represented by video data captured by a camera.

40. The method according to claim 23, wherein the acoustic space comprises a virtual world.

41. The method according to claim 23, further comprising presenting the acoustic space on a head-mounted device.

42. The method according to claim 23, further comprising presenting an acoustic space on a mobile handset or within a vehicle.

43. The method according to claim 23, further comprising receiving a wireless signal.

44. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: store timing information and a plurality of audio streams, wherein the timing information includes a delay and an update duration; control access to at least one of the plurality of audio streams based on the timing information; detect a trigger configured to trigger an audio scene update; compare the delay with a timer; and wait until the delay is equal to or greater than the timer, initiate a change to a plurality of audio sources for the plurality of audio streams, and change the plurality of audio sources during the update duration to select a subset of the plurality of audio streams.

45. A device configured to play one or more of a plurality of audio streams, the device comprising: means for storing timing information and a plurality of audio streams, wherein the timing information includes a delay and an update duration; means for controlling access to at least one of the plurality of audio streams based on the timing information; means for detecting a trigger configured to trigger an audio scene update; means for comparing the delay with a timer; and means for waiting until the delay is equal to or greater than the timer, initiate a change to a plurality of audio sources for the plurality of audio streams, and change the plurality of audio sources during the update duration to select a subset of the plurality of audio streams.

Citation Information

Patent Citations

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US10405126B2

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US20190007781A1

  • Transporting coded audio data

    CN107925797A

  • Audio network system

    JP2010178368A

  • Method and apparatus for generating virtual or augmented reality presentations with 3D audio positioning

    US20160269712A1