Adapting an audio stream for rendering

By adaptively selecting and adapting audio streams, the rendering problem under the hardware limitations of vehicles or XR devices is solved, and audio stream rendering and audio experience are improved under hardware constraints.

CN114008707BActive Publication Date: 2025-10-28QUALCOMM INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202080047110.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-01
Filing Date
2020-07-02
Publication Date
2025-10-28
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

Vehicles or certain XR devices may be unable to render all substreams of the audio stream due to limitations in processor, memory, or other hardware, resulting in a limited audio experience.

Method used

By adaptively selecting and adapting audio streams to meet rendering thresholds, the number of substreams is reduced, and the renderer generates speaker feeds to output audio.

Benefits of technology

It enables the rendering of multiple audio streams under hardware limitations, thus improving the audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114008707B_ABST
    Figure CN114008707B_ABST
Patent Text Reader

Abstract

Typically, techniques for adapting video and audio streams for rendering are described. A device including memory and one or more processors can be configured to perform these techniques. The memory can store multiple audio streams comprising one or more substreams. The processor can determine the total number of one or more substreams of all audio streams in the multiple audio streams based on the multiple audio streams, and when the total number of substreams is greater than a rendering threshold, adapt the multiple audio streams to reduce the number of one or more substreams and obtain adapted multiple audio streams. The processor can also apply a renderer to the adapted multiple audio streams to obtain one or more speaker feeds and output the one or more speaker feeds to one or more speakers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 16 / 918,372, filed July 1, 2020, entitled “ADAPTING AUDIO STREAMS FOR RENDERING,” which claims the benefit of U.S. Provisional Application No. 62 / 870,584, filed July 3, 2019, entitled “ADAPTING AUDIO STREAMS FOR RENDERING,” the entire contents of which are incorporated herein by reference, as fully set forth herein. Technical Field

[0003] This disclosure relates to the processing of audio data. Background Technology

[0004] In many cases, rendering audio data may not be suitable for a particular audio format. For example, due to processing, memory, power, or other limitations, some vehicles or other types of devices (such as extended reality (XR) devices, which may refer to virtual reality (VR) devices, augmented reality (AR) devices, and / or mixed reality (MR) devices) may only have renderers that support certain formats. Audio streams are increasingly being offered in a variety of formats that may not be suitable for vehicles and / or XR devices, thus limiting the audio experience in these situations. Summary of the Invention

[0005] This disclosure generally relates to adapting audio streams for rendering.

[0006] In one example, aspects of the technology relate to a device configured to play one or more of a plurality of audio streams, the device comprising: a memory configured to store the plurality of audio streams, each of the plurality of audio streams representing a sound field and including one or more substreams; and one or more processors coupled to the memory and configured to: determine, based on the plurality of audio streams, the total number of one or more substreams of all audio streams in the plurality of audio streams; when the total number of one or more substreams is greater than a rendering threshold indicating the total number of substreams supported when the renderer renders the plurality of audio streams to one or more speaker feeds, adapt the plurality of audio streams to reduce the number of one or more substreams and obtain adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold; apply the renderer to the adapted plurality of audio streams to obtain one or more speaker feeds; and output the one or more speaker feeds to one or more speakers.

[0007] In another example, aspects of the technology relate to a method for playing one or more of a plurality of audio streams, the method comprising: storing a plurality of audio streams by one or more processors, each of the plurality of audio streams representing a sound field and including one or more substreams; determining, by one or more processors, the total number of one or more substreams of all audio streams in the plurality of audio streams based on the plurality of audio streams; adapting the plurality of audio streams to reduce the number of one or more substreams and obtaining adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold by one or more processors; applying a renderer to the adapted plurality of audio streams by one or more processors to obtain one or more speaker feeds; and outputting the one or more speaker feeds to one or more speakers by one or more processors.

[0008] In another example, aspects of the technology relate to a device configured to play one or more of a plurality of audio streams, the device comprising: components for storing the plurality of audio streams, each of the plurality of audio streams representing a sound field and including one or more substreams; components for determining, based on the plurality of audio streams, the total number of one or more substreams of all audio streams in the plurality of audio streams; components for adapting the plurality of audio streams to reduce the number of one or more substreams and obtaining adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold when the total number of one or more substreams is greater than a rendering threshold indicating the total number of substreams supported when the renderer renders the plurality of audio streams to one or more speaker feeds; components for applying the renderer to the adapted plurality of audio streams to obtain one or more speaker feeds; and components for outputting one or more speaker feeds to one or more speakers.

[0009] In another example, aspects of the technology relate to a non-transitory computer-readable storage medium having instructions thereon that, when executed, cause one or more processors to: store a plurality of audio streams, each of the plurality of audio streams representing a sound field and including one or more substreams; determine, based on the plurality of audio streams, the total number of one or more substreams of all audio streams in the plurality of audio streams; when the total number of one or more substreams is greater than a rendering threshold indicating the total number of substreams supported when the renderer renders the plurality of audio streams to one or more speaker feeds, adapt the plurality of audio streams to reduce the number of one or more substreams and obtain adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold; apply the renderer to the adapted plurality of audio streams to obtain one or more speaker feeds; and output the one or more speaker feeds to one or more speakers.

[0010] In another example, aspects of the technology relate to a device configured to play one or more of a plurality of audio streams, the device comprising: a memory configured to store the plurality of audio streams and corresponding audio metadata, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; and one or more processors coupled to the memory and configured to: determine the direction of arrival of each of the plurality of audio streams based on the current coordinates of the device relative to the start coordinates corresponding to the plurality of audio streams; render each of the plurality of audio streams to one or more speaker feeds based on each direction of arrival, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each direction of arrival; and output the one or more speaker feeds to reproduce the one or more sound fields represented by the plurality of audio streams.

[0011] In another example, aspects of the technology relate to a method for playing one or more of a plurality of audio streams, the device comprising: storing a plurality of audio streams and corresponding audio metadata in a memory, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; determining, by one or more processors, an arrival direction for each of the plurality of audio streams based on the current coordinates of the device relative to the start coordinates corresponding to the plurality of audio streams; rendering, by one or more processors, each of the plurality of audio streams to one or more speaker feeds based on each arrival direction, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each arrival direction; and outputting the one or more speaker feeds by one or more processors to reproduce the one or more sound fields represented by the plurality of audio streams.

[0012] In another example, aspects of the technology relate to a device configured to play one or more of a plurality of audio streams, the device comprising: components for storing the plurality of audio streams and corresponding audio metadata, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; components for determining the direction of arrival of each of the plurality of audio streams based on the current coordinates of the device relative to the start coordinates corresponding to the plurality of audio streams; components for rendering each of the plurality of audio streams to one or more speaker feeds based on each direction of arrival, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each direction of arrival; and components for outputting the one or more speaker feeds to reproduce the one or more sound fields represented by the one or more of the plurality of audio streams.

[0013] In another example, aspects of the technology relate to a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: store a plurality of audio streams and corresponding audio metadata, each of the plurality of audio streams representing a sound field, and the audio metadata including the starting coordinates of a corresponding origin for each of the plurality of audio streams; determine the direction of arrival for each of the plurality of audio streams based on the current coordinates of the device relative to the starting coordinates corresponding to the plurality of audio streams; render each of the plurality of audio streams to one or more speaker feeds based on each direction of arrival, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each direction of arrival; and output the one or more speaker feeds to reproduce the one or more sound fields represented by the plurality of audio streams.

[0014] Details of one or more examples of this disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of various aspects of these techniques will be apparent from the specification, drawings, and claims. Attached Figure Description

[0015] Figure 1A and Figure 1B This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure.

[0016] Figure 2A To show in more detail Figure 1A and Figure 1B The example shown is a block diagram of an example system.

[0017] Figure 2BThis is a flowchart illustrating example operations of the flow selection unit in performing various aspects of the techniques described in this disclosure.

[0018] Figure 2C This illustrates in more detail the various aspects of the technology described in this disclosure. Figure 2A The flowchart shows an additional example operation of the flow selection unit shown in the example.

[0019] Figure 2D-Figure 2K It is shown by Figure 1A and Figure 1B The example shown is a diagram illustrating the application of privacy settings on the source device and / or content consumer device.

[0020] Figures 3A-3F It shows in more detail Figure 1A and Figure 1B The diagram shows an example system that can perform various aspects of the techniques described in this disclosure.

[0021] Figure 4 This is an image showing an example of a VR device worn by a user.

[0022] Figure 5 This is a diagram illustrating an example of a wearable device that can operate according to various aspects of the technology described in this disclosure.

[0023] Figure 6A and Figure 6B This is a diagram illustrating other example systems that can perform various aspects of the techniques described in this disclosure.

[0024] Figure 7 This is a block diagram illustrating one or more sample components of the source device and content consumer device shown in the example of Figure 1.

[0025] Figures 8A-8C It is shown Figure 1A and Figure 1B The example shown is a flowchart illustrating the example operations of the stream selection unit when performing various aspects of the stream selection technique.

[0026] Figure 9 An example of a wireless communication system supporting privacy restrictions according to aspects of this disclosure is shown. Detailed Implementation

[0027] There are several different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound, 7.1 surround sound, 22.2 surround sound, or any other channel-based format that positions audio channels to specific locations around the listener in order to recreate the sound field.

[0028] Object-based audio formats can refer to formats in which audio objects (typically encoded using Pulse Codec Modulation (PCM) and referred to as PCM audio objects) are specified to represent a sound field. Such audio objects may include metadata identifying the position of the audio object relative to a listener or other reference points in the sound field, allowing the audio object to be rendered to one or more speaker channels for playback in an effort to recreate the sound field. The techniques described in this disclosure are applicable to any of the above formats, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.

[0029] Scene-based audio formats may consist of a hierarchical set of elements that define a sound field in three dimensions. An example of a hierarchical set of elements is the set of spherical harmonic coefficients (SHC). The following expression demonstrates a description or representation of a sound field using an SHC:

[0030]

[0031] This expression indicates that at any point in the sound field... Pressure p at the point i At time t, it can be determined by SHC. Unique representation. Here, c is the speed of sound (~343 m / s). It is the reference point (or observation point), j n (·) is an nth-order spherical Bessel function. These are the nth and mth order spherical harmonic basis functions (also called spherical basis functions). It can be seen that the terms in square brackets are the frequency domain representation of the signal (i.e.,...). This can be approximated by various time-frequency transforms, such as the Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), or Wavelet Transform. Other examples of hierarchical sets include wavelet transform coefficient sets and other coefficient sets of multi-resolution basis functions.

[0032] SHC Physical acquisition (e.g., recording) can be achieved through various microphone array configurations, or they can be derived from channel-based or object-based sound field descriptions. SHC (which can also be referred to as ambisonic reverberation coefficients) represents scene-based audio, where SHC can be input to an audio encoder to obtain an encoded SHC that facilitates more efficient transmission or storage. For example, a (1+4) approach can be used. 2 (25, therefore a fourth-order) coefficient fourth-order representation.

[0033] As mentioned above, SHC can be derived from microphone recordings using microphone arrays. Poletti, M., “Three-Dimensional Surround Sound Systems Based on Spherical Harmonics”, J. Audio Eng. Soc., Vol. 53, No. 11, November 2005, pp. 1004-1025, describes various examples of how to obtain SHC from microphone array physics.

[0034] The following equation illustrates how to derive SHC from an object-based description. The sound field coefficients correspond to a single audio object. It can be represented as:

[0035]

[0036] Where i is It is an n-order spherical Hankel function (of the second kind). It refers to the location of the object. Understanding the object's source energy g(ω) as a function of frequency (e.g., using time-frequency analysis techniques, such as performing a Fast Fourier Transform on a Pulse Coded Modulation-Pulse Stream) allows each PCM object and its corresponding location to be converted to SHC. Furthermore, each object can be shown (since the above is a linear and orthogonal decomposition). The coefficients are additive. In this way, multiple PCM objects can be generated by... The coefficients are represented (e.g., as the sum of the coefficient vectors of the individual objects). The coefficients may contain information about the sound field (pressure as a function of 3D coordinates), as shown above, representing the distance from a single object to the observation point. The transformation of the representation of the overall sound field in the vicinity.

[0037] With the growth of connectivity (e.g., cellular and other forms of wireless communication), the ability to capture and stream media content is increasing, enabling almost anyone to stream in real time or in other forms via mobile devices (or other types of devices). Thus, mobile devices can capture sound fields using one of the representations discussed above and generate audio streams, which the mobile device can then send to anyone who wants to listen. In some scenarios, audio streams may convey useful information or simply provide entertainment (e.g., music, etc.).

[0038] One area where audio streaming can benefit is in the vehicular environment. Vehicle-to-any (V2X) communication allows devices such as mobile phones to interface with a vehicle to stream audio data. In some cases, the vehicle's audio head unit receives the audio stream and reproduces the sound field represented by the audio stream through one or more speakers. In other cases, a mobile device can output a speaker feed, which the vehicle receives and uses to reproduce the sound field. In any case, V2X communication allows vehicles to communicate with mobile devices or even other vehicles to obtain audio streams.

[0039] Vehicles can communicate with each other via the V2X protocol to transmit audio streams between vehicles. In some examples, the audio stream may represent what the occupants of a first vehicle say, which can be played back by the occupants of a second vehicle so that they can hear what is being said. The spoken words may be a command representing future action by the occupants of the first vehicle (e.g., “pass from the left”). In other examples, the audio stream may represent an entertainment audio stream (e.g., streaming music) shared between the first and second vehicles.

[0040] Another area where audio streaming can bring benefits is extended reality (XR). XR devices can include virtual reality (VR) devices, augmented reality (AR) devices, and mixed reality (MR) devices. XR devices can retrieve and render audio streams to enable various operations, such as virtual attendance at events, parties, sports events, meetings, etc., teleportation (allowing users to see or experience another person's experience, such as becoming the co-driver in a vehicle), remote surgery, etc.

[0041] However, vehicles and some XR devices may only be able to render a specific number of substreams contained within an audio stream. When attempting to render multiple audio streams or a particular type of audio data represented by an audio stream (e.g., stereo reverberation audio data with a large number of coefficients per sample), the device may be unable to render all substreams of all audio streams. That is, there are processor, memory, or other physical hardware limitations (e.g., bandwidth) that may prevent existing devices from retrieving and processing all available substreams of an audio stream, especially when the audio stream may require significant bandwidth and processing resources in certain situations (e.g., stereo reverberation coefficients corresponding to higher-order spherical basis functions, such as third, fourth, fifth, and sixth orders).

[0042] Depending on various aspects of the technology, a device (e.g., a mobile handheld device, a vehicle, a vehicle audio head unit, and / or an XR device) can operate in a systematic manner to adaptively select subsets of multiple audio streams and / or substreams. The device may include any audio stream identified by the user presentation, but otherwise removes any audio stream originating from a remote location (because an audio stream may include audio metadata defining the starting location for spatial rendering purposes, as described in more detail below), any higher-order stereo reverberation coefficients (by down-order), and any streams with private specifications or other privacy settings. In this way, various substreams associated with the audio stream can be removed to suit the device's rendering constraints, thereby enabling the device to render virtually any type of audio stream and improving the operation of the device itself.

[0043] Figure 1A and Figure 1B This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure. For example... Figure 1A As illustrated in the example, system 10 includes a source device 12 and a target device 14. Although described in the context of source device 12 and target device 14, these techniques can be implemented in any context in which any representation of a sound field is encoded to form a bitstream, or in other words, an audio stream representing audio data. Furthermore, source device 12 can represent any form of computing device capable of generating a sound field representation, and is generally described herein in the context of a vehicle audio head unit. Similarly, target device 14 can represent any form of computing device capable of implementing the rendering techniques and audio playback described in this disclosure, and is generally described herein in the context of a vehicle.

[0044] Source device 12 can be an entity capable of generating audio content for use by the operator of target device 14. In some scenarios, source device 12 combines video content to generate audio content. Source device 12 includes content capture device 20, content editing device 22, and sound field representation generator 24. Content capture device 20 can be configured to interface with microphone 18 or otherwise communicate with microphone 18.

[0045] Microphone 18 can indicate Or other types of 3D audio microphones capable of capturing a sound field and representing it as audio data 19, which may refer to one or more of the scene-based audio data (e.g., stereo reverberation coefficients), object-based audio data, and channel-based audio data described above. Although described as a 3D audio microphone, microphone 18 may also represent other types of microphones configured to capture audio data 19 (e.g., omnidirectional microphones, spot microphones, unidirectional microphones, etc.).

[0046] In some examples, the content capture device 20 may include an integrated microphone 18 integrated into the housing of the content capture device 20. The content capture device 20 may interface with the microphone 18 wirelessly or via a wired connection. Instead of capturing or combining audio data 19 via the microphone 18, the content capture device 20 may process the audio data 19 after it has been input via some type of removable storage, wireless, and / or wired input process. Therefore, various combinations of the content capture device 20 and the microphone 18 are possible according to this disclosure.

[0047] Content capture device 20 may also be configured to interface with or otherwise communicate with content editing device 22. In some instances, content capture device 20 may include content editing device 22 (in some instances, content editing device 22 may represent software or a combination of software and hardware, including software executed by content capture device 20 to configure content capture device 20 to perform a particular form of content editing). Content editing device 22 may represent a unit configured to edit or otherwise modify content 21 (including audio data 19) received from content capture device 20. Content editing device 22 may output the edited content 23 and associated metadata 25 to sound field representation generator 24.

[0048] The sound field representation generator 24 may include any type of hardware device capable of interfacing with the content editing device 22 (or the content capture device 20). Although in Figure 1A Not shown in the example, the sound field representation generator 24 can use edited content 23 provided by the content editing device 22, including audio data 19 and metadata 25, to generate one or more bitstreams 25. Focusing on the audio data 19... Figure 1A In the example, the sound field representation generator 24 can generate one or more representations of the same sound field represented by the audio data 19 to obtain a bitstream 27 including the sound field representation and audio metadata 25.

[0049] For example, in order to generate different representations of the sound field using stereo reverberation coefficients (which is also an example of audio data 19), the sound field representation generator 24 can use a coding scheme for the stereo reverberation representation of the sound field, called Mixed-Order Ambisonics (MOA), as discussed in more detail in U.S. Application Serial No. 15 / 672,058, entitled “MIXED-ORDER AMBISONICS (MOA) AUDIO DATA FO COMPUTER-MEDIATED REALITY SYSTEMS”, filed August 8, 2017 and published January 3, 2019 as U.S. Patent Publication No. 2019 / 0007781.

[0050] To generate a specific MOA representation of a sound field, the sound field representation generator 24 can generate a partial subset of the complete set of stereo reverberation coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide accuracy for some regions of the sound field, but lower accuracy for others. In one example, the MOA representation of the sound field may include eight (8) uncompressed stereo reverberation coefficients, while a third-order stereo reverberation representation of the same sound field may include sixteen (16) uncompressed stereo reverberation coefficients. Therefore, each MOA representation of the sound field generated as a partial subset of stereo reverberation coefficients may have lower storage and bandwidth density than a corresponding third-order stereo reverberation representation of the same sound field generated from stereo reverberation coefficients (if and when transmitted as part of bitstream 27 on the illustrated transmission channel).

[0051] Although the MOA representation has been described, the techniques of this disclosure can also be performed with respect to the first-order stereo reverberation (FOA) representation, where all stereo reverberation coefficients associated with the first-order and zero-order spherical basis functions are used to represent the sound field. In other words, the sound field representation generator 24 can represent the sound field using all stereo reverberation coefficients of a given order N, rather than using a partial non-zero subset of the stereo reverberation coefficients, thus obtaining a total stereo reverberation coefficient equal to (N+1). 2 .

[0052] In this regard, stereo reverberant audio data (which refers to another way of representing stereo reverberant coefficients in MOA representation or full-order representation, such as the first-order representation mentioned above) may include stereo reverberant coefficients associated with spherical basis functions of first order or less (referred to as "first-order stereo reverberant audio data"), stereo reverberant coefficients associated with spherical basis functions of mixed order and sub-order (referred to as "MOA representation" discussed above), or stereo reverberant coefficients associated with spherical basis functions of order greater than 1 (referred to as "full-order representation" above).

[0053] In some examples, the content capture device 20 or the content editing device 22 may be configured to communicate wirelessly with the sound field representation generator 24. In some examples, the content capture device 20 or the content editing device 22 may communicate with the sound field representation generator 24 via one or both wireless or wired connections. Through the connection between the content capture device 20 and the sound field representation generator 24, the content capture device 20 can provide content in various forms, which, for the purposes of discussion, are described herein as part of audio data 19.

[0054] In some examples, content capture device 20 may utilize various aspects of sound field representation generator 24 (in terms of the hardware or software capabilities of sound field representation generator 24). For example, sound field representation generator 24 may include dedicated hardware configured to perform psychoacoustic audio coding (or dedicated software that, when executed, causes one or more processors to perform psychoacoustic audio coding) (e.g., by the Moving Picture Experts Group (MPEG), the MPEG-H 3D audio codec standard, the MPEG-I immersive audio standard, or proprietary standards such as AptX). TM (Including various versions of AptX, such as Enhanced AptX–E-AptX, AptX Live, AptX Stereo, and AptX High Definition–AptX HD), Advanced Audio Codec (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Windows Media Audio (WMA), which are presented as a unified speech and audio codec called “USAC”.

[0055] Content capture device 20 may not include dedicated hardware or software for psychoacoustic audio encoders, but may instead provide the audio aspects of content 21 in a non-psychoacoustic audio codec format. Sound field representation generator 24 may assist in the capture of content 21 by at least partially performing psychoacoustic audio encoding of the audio aspects of content 21.

[0056] The sound field representation generator 24 can also assist in content capture and transmission by generating one or more bitstreams 27, at least in part, based on the audio content generated from the audio data 19 (e.g., MOA representation and / or first-order stereo reverberation representation). The bitstreams 27 can represent compressed versions of the audio data 19 and any other different types of content 21 (e.g., compressed versions of spherical video data, image data, or text data).

[0057] As an example, the sound field representation generator 24 can generate a bitstream 27 for transmission across a transport channel (which may be a wired or wireless channel, a data storage device, etc.). Bitstream 27 can represent an encoded version of audio data 19 and can include a main bitstream and a side bitstream, which may be referred to as side-channel information or metadata. In some instances, bitstream 27 represents a compressed version of audio data 19 (which may also represent scene-based audio data, object-based audio data, channel-based audio data, or a combination thereof) conforming to a bitstream generated according to the MPEG-H 3D audio codec standard and / or the MPEG-I immersive audio standard.

[0058] As described above, source device 12 can represent a vehicle. Examples of vehicles include bicycles, mopeds, motorcycles, automobiles (including autonomous vehicles), aircraft (including autonomous aircraft), agricultural equipment, construction equipment, military vehicles (e.g., tanks, transport vehicles, etc.), drones or other remotely operated aircraft, helicopters, quadcopters, trains, ships, or any other means of transportation capable of transporting occupants from one location to another. In the context of a vehicle, target device 12 may not represent the vehicle as a whole, but only the vehicle's computing system, such as an audio host configured to interface with one or more audio elements (e.g., microphones) to capture the sound field represented by audio stream 27.

[0059] Although described in the context of a vehicle, source device 12 can represent a device that communicates with any of the example vehicles described above, such that source device 12 functions effectively as part of the vehicle. For example, source device 12 can represent a device that communicates via a PAN protocol (e.g., A smartphone or other mobile handheld device that communicates with a vehicle (e.g., wirelessly) using other wireless or wired communication protocols. In this case, source device 12 can represent any form of computing device configured to communicate with a vehicle, including mobile handheld devices (including so-called smartphones), laptop computers, XR devices, gaming systems (e.g., portable gaming systems), or any other computing device.

[0060] Furthermore, the target device 14 can be operated by an individual and can represent a vehicle, for example... Figures 3A-3F The example shows vehicle 14. Although described in the context of a vehicle, target device 14 could represent other types of devices, such as augmented reality (AR) client devices, mixed reality (MR) client devices (or other XR client devices), standard computers, headsets, headphones, mobile devices (including so-called smartphones), or any other device capable of reproducing a sound field based on an audio stream. Figure 1AAs shown in the example, target device 14 includes audio playback system 16A, which may refer to any form of audio playback system capable of rendering audio data for playback as mono or multi-channel audio content.

[0061] Although Figure 1A The bitstream 27 is shown as being transmitted directly to the target device 14, but the source device 12 can output the bitstream 27 to an intermediate device located between the source device 12 and the target device 14. The intermediate device can store the bitstream 27 for later transmission to the target device 14, which can request the bitstream 27. The intermediate device can include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing the bitstream 27 for later retrieval by an audio decoder. The intermediate device can reside in a content delivery network capable of streaming the bitstream 27 (and possibly in conjunction with the transmission of a corresponding video data bitstream) to a subscriber (e.g., the target device 14) that requests the bitstream 27.

[0062] Alternatively, source device 12 may store bitstream 27 to a storage medium, such as an optical disc, digital video disc, high-definition video disc, or other storage medium, most of which are computer-readable and therefore may be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, a transmission channel may refer to a channel through which the content stored to the medium (e.g., in the form of one or more bitstreams 27) is transmitted (and may include retail stores and other store-based delivery mechanisms). In any case, the technology disclosed herein should not be limited in this respect. Figure 1A Examples.

[0063] As described above, target device 14 includes audio playback system 16A. Audio playback system 16A can represent any system capable of playing back mono and / or multi-channel audio data. Audio playback system 16A can include multiple different renderers 32. Each renderer 32 can provide different forms of rendering, wherein different forms of rendering can include one or more of various methods of performing vector-based amplitude shifting (VBAP), and / or one or more of various methods of performing sound field synthesis. As used herein, “A and / or B” means “A or B”, or both “A and B”.

[0064] The audio playback system 16A may also include an audio decoding device 34. Audio decoding device 34 may refer to a device configured to decode bitstream 27 to output audio data 19' (wherein the apostrophe ' may indicate that audio data 19' differs from audio data 19 due to lossy compression (e.g., quantization) of audio data 19). Similarly, audio data 19' may include scene-based audio data, which in some examples may form a complete first-order (or higher-order) stereo reverberation representation or a subset thereof forming a MOA representation of the same sound field, its decomposition, such as the primary audio signal described in the MPEG-H 3D audio codec standard, ambient stereo reverberation coefficients, and vector-based signals (which may refer to a multidimensional spherical harmonic vector having multiple elements representing the spatial characteristics of the corresponding primary audio signal) or other forms of scene-based audio data.

[0065] Other forms of scene-based audio data include audio data defined according to the HOA (Higher Order Ambisonics (HOA) Transport Format) (HTF). More information on HTF can be found in the technical specification (TS) entitled "Higher Order Ambisonics (HOA) Transport Format" filed by the European Telecommunications Standards Institute (ETSI) in June 2018 (2018-06), ETSI TS 103 589 V1.1.1, and U.S. Patent Publication No. 2019 / 0918028 entitled "PRIORITY INFORMATION FOR HIGHER ORDER AMBISONIC AUDIO DATA" filed on December 20, 2018. In any case, audio data 19' may resemble a complete set or a subset of audio data 19', but may differ due to lossy operations (e.g., quantization) and / or transmission via a transmission channel.

[0066] Audio data 19' may include, replace, or combine scene-based audio data and channel-based audio data. Audio data 19' may include, replace, or combine scene-based audio data and object-based audio data. Therefore, audio data 19' may include any combination of scene-based audio data, object-based audio data, and channel-based audio data.

[0067] The audio renderer 32 of the audio playback system 16A can render the audio data 19' to output the speaker feed 35 after the audio decoding device 34 decodes the bitstream 27 to obtain the audio data 19'. The speaker feed 35 can drive one or more speakers (for illustrative purposes, in...). Figure 1A(Not shown in the example). Various audio representations of the sound field, including scene-based audio data (and possibly channel-based and / or object-based audio data), can be normalized in a variety of ways, including N3D, SN3D, FuMa, N2D, or SN2D.

[0068] To select an appropriate renderer, or in some instances, to generate an appropriate renderer, the audio playback system 16A may obtain speaker information 37 indicating the number of speakers (e.g., loudspeakers or headphone speakers) and / or the spatial geometry of the speakers. In some instances, the audio playback system 16A may use a reference microphone to obtain the speaker information 37 and may drive the speakers in a manner that dynamically determines the speaker information 37 (which may refer to the output of an electrical signal to cause the transducer to vibrate). In other cases, or in conjunction with the dynamic determination of the speaker information 37, the audio playback system 16A may prompt the user to interact with the audio playback system 16A and input the speaker information 37.

[0069] The audio playback system 16A may select one of the audio renderers 32 based on the speaker information 37. In some instances, the audio playback system 16A may generate one of the audio renderers 32 based on the speaker information 37 when none of the audio renderers 32 are within a certain threshold similarity metric (in terms of speaker geometry) of the speaker geometry specified in the speaker information 37. In some cases, the audio playback system 16A may generate one of the audio renderers 32 based on the speaker information 37 without first attempting to select one of the existing audio renderers 32.

[0070] When the speaker feed 35 is output to the headphones, the audio playback system 16A can utilize one of the renderers 32, which provides binaural rendering using a Head-Related Transfer Function (HRTF) or other functions (such as a binaural room impulse response renderer) capable of rendering to the left and right speaker feeds 35 for headphone speaker playback. The term "speaker" or "transducer" can generally refer to any speaker, including loudspeakers, headphone speakers, bone conduction speakers, earbud speakers, wireless headphone speakers, etc. One or more speakers can then play back the rendered speaker feed 35 to reproduce the sound field.

[0071] Although described as rendering the speaker feed 35 from audio data 19', references to rendering the speaker feed 19' can refer to other types of rendering, such as rendering directly incorporated into the audio data 35 decoded from bitstream 27. Examples of alternative rendering can be found in Appendix G of the MPEG-H3D audio standard, where rendering occurs during the formation of the main signal and background signal prior to sound field synthesis. Therefore, references to rendering audio data 19' should be understood to refer to either the rendering of the actual audio data 19' or a decomposition or representation of audio data 19' (e.g., the main audio signal, ambient stereo reverberation coefficients, and / or vector-based signals mentioned above—also referred to as V-vectors or multidimensional stereo reverberation space vectors).

[0072] The audio playback system 16A can also be adapted to the audio renderer 32 based on the tracking information 41. That is, the audio playback system 16A can interface with the tracking device 40, which is configured to determine the current coordinates of the target device 14. The tracking device 40 can represent one or more sensors (e.g., cameras—including depth cameras, gyroscopes, magnetometers, accelerometers, light-emitting diodes—LEDs, GPS units, etc.) configured to track the current coordinates of the target device 14. The audio playback system 16A can adapt the audio renderer 32 based on the tracking information 41 such that the speaker feed 35 reflects the change in the current coordinates relative to the original coordinates set in the metadata 23 of the bitstream 27 (which may represent one or more audio streams, and is therefore referred to as audio stream 27).

[0073] As described above, target device 14 can represent a vehicle. Examples of vehicles include bicycles, mopeds, motorcycles, automobiles (including autonomous vehicles), aircraft (including autonomous aircraft), agricultural equipment, construction equipment, military vehicles (e.g., tanks, transport vehicles, etc.), drones or other remotely operated aircraft, helicopters, quadcopters, trains, ships, or any other means of transportation capable of transporting occupants from one location to another. In the context of a vehicle, target device 14 may not represent the vehicle as a whole, but only the vehicle's computing system, such as an audio host configured to interface with one or more speakers to reproduce the sound field represented by audio stream 27.

[0074] Although described in the context of a vehicle, target device 14 can refer to any device that communicates with any of the example vehicles described above, such that target device 14 functions effectively as part of the vehicle. For example, target device 14 can refer to a device that communicates via a PAN protocol (e.g., A smartphone or other mobile handheld device that communicates with a vehicle (e.g., wireless communication) using other wireless or wired communication protocols. In this context, target device 14 can represent any form of computing device configured to communicate with a vehicle, including mobile handheld devices (including so-called smartphones), laptop computers, XR devices, gaming systems (e.g., portable gaming systems), or any other computing device.

[0075] With the growth of connectivity (e.g., cellular and other forms of wireless communication), the ability to capture and stream media content is increasing, enabling almost anyone to stream in real time or in other forms via mobile devices (or other types of devices). Thus, mobile devices can capture sound fields using one of the representations discussed above and generate audio streams, which the mobile device can then send to anyone who wants to listen. In some cases, audio streams may convey useful information or simply provide entertainment (e.g., music, etc.).

[0076] One area where audio streaming can benefit is in the vehicular environment. Vehicle-to-any (V2X) communication allows devices such as mobile phones to interface with a vehicle to stream audio data. In some cases, the vehicle's audio head unit receives the audio stream and reproduces the sound field represented by the audio stream through one or more speakers. In other cases, a mobile device can output a speaker feed, which the vehicle receives and uses to reproduce the sound field. In any case, V2X communication allows vehicles to communicate with mobile devices or even other vehicles to obtain audio streams.

[0077] Vehicle 12 can perform inter-vehicle communication via the V2X protocol to transmit an audio stream between vehicles 12 and 14. In some examples, the audio stream may represent what the occupants of the first vehicle 12 say, and the second vehicle 14 may play these words so that the occupants of the second vehicle 14 can hear them. The spoken words may be a command representing the future course of action of the occupants of the first vehicle 12 (e.g., "pass on the left"). In other examples, the audio stream may represent what the first vehicle 12 says via... Figures 3A-3F The example shows wireless connection 200 (including wireless connections 200A-200D) sharing entertainment audio streams (e.g., streaming music) with the second vehicle 14.

[0078] Another area where audio streaming can bring benefits is extended reality (XR). XR devices can include virtual reality (VR) devices, augmented reality (AR) devices, and mixed reality (MR) devices. XR devices can retrieve and render audio streams to enable various operations, such as virtual attendance at events, parties, sports events, meetings, etc., teleportation (allowing users to see or experience another person's experience, such as becoming the co-driver in a vehicle), remote surgery, etc.

[0079] However, vehicles and some XR devices may only be able to render a specific number of substreams contained within an audio stream. When attempting to render multiple audio streams or a particular type of audio data represented by an audio stream (e.g., stereo reverberation audio data with a large number of coefficients per sample), the device may be unable to render all substreams of all audio streams. That is, there are processor, memory, or other physical hardware limitations (e.g., bandwidth) that may prevent existing devices from retrieving and processing all available substreams of an audio stream, especially when the audio stream may require significant bandwidth and processing resources in certain situations (e.g., stereo reverberation coefficients corresponding to higher-order spherical basis functions, such as third, fourth, fifth, sixth, etc.).

[0080] Depending on various aspects of the technology, target device 14 (examples of which include mobile handheld devices, vehicles, vehicle audio heads, and / or XR devices) can operate systematically to adaptively select a subset of multiple audio streams 19'. Target device 14 may include any audio stream 27 identified by a user preset (in audio stream 19'), but additionally remove any audio streams originating from remote locations (since audio streams may include audio metadata defining the starting location for spatial rendering purposes, as described in more detail below), any higher-order stereo reverberation coefficients (by reducing the order, thereby reducing the number of inputs to audio renderer 32), and any audio stream 27 with private specifications or other privacy settings. In this way, various rendering constraints associated with the audio streams can be removed to adapt to the device, thereby enabling the device to render virtually any type of audio stream and improving the operation of the device itself.

[0081] In addition, the audio decoding device 34 can be connected wirelessly (e.g. Figure 3A The wireless connection 200 shown in the example communicates with the source device 12. The operator of the source device 12 can interface with the source device 12 to capture audio data (for illustrative purposes, it is assumed that the audio data is what the operator says).

[0082] Source device 12 may include microphone 18 or other audio capture device configured to capture audio data 19 and generate audio stream 27 based on the audio data 19. Sound field representation generator 24 can generate audio stream 27 and audio metadata, including the origin coordinates (e.g., GPS coordinates) of the corresponding audio stream 27. Sound field representation generator 24 can output audio stream 27 to target device 14 via wireless connection 200.

[0083] Audio decoding device 34 can receive audio stream 27 (which includes audio metadata) and store audio stream 27. Target device 14 can determine the direction of arrival (DoA) 212 of audio stream 27 based on its current coordinates relative to the starting coordinates corresponding to the audio stream 27, as indicated by tracking information 41. Target device 14 can determine that audio stream 27 is arriving from directly behind target device 14 and proceeding in front of target device 14 (as shown by the arrow).

[0084] The audio playback system 16A can invoke the audio renderer 32 to render the audio stream 27 as if it were arriving from the direction of arrival 212, thereby generating a speaker feed to simulate the sound field captured by the source device 12, and, for example, arriving directly behind the target device 14. In this example, the audio renderer 32 can... Figure 1A In the example, right and left rear speaker feeds are generated and output to the right and left rear speakers (not shown for illustration) of the target device 14 to reproduce the sound field 214A represented by audio stream 27.

[0085] In this example, assume that the driver of source device 12 captures audio data of the driver issuing a command so that target device 14 knows that the driver will "pass from the left". Although the following description refers to what is said, source device 12 may provide one or more audio streams, including pre-recorded audio streams, live audio streams, or any other type of audio stream.

[0086] When assuming automated operation (e.g., the computing device controls the target device 14 and issues commands causing the computing device to steer, accelerate, brake, and otherwise operate the target device 14 without human intervention), the target device 14 can analyze the audio stream 27 to extract commands indicating the action route, and operate the target device 14 based on the commands parsed from the audio stream 27 to avoid merging into the lane of the source device 12 or otherwise affecting the operation of the source device 12. In other words, the target device 14 can automatically adjust its operation based on commands or other spoken words.

[0087] Figure 2A To show in more detail Figure 1A and Figure 1B The example shown is a block diagram of an example system. Figure 2A As shown in the example, system 150 includes a local network 152A and a remote network 152B. Local network 152A can represent a network capable of operating according to local streaming protocols (including fifth-generation (5G) cellular protocols, WiFi protocols, PAN protocols, etc.). (or any other wireless protocol capable of interconnecting devices) Local streaming audio (as one or more audio streams 27, see reference) Figure 1A(Example) Local network of interconnected devices.

[0088] Remote network 152B can represent a publicly accessible, packet-based network, such as the Internet, or a private network operating according to various Layer 2, Layer 3, and other network protocols. Remote network 152B may include multiple interconnected network devices, including routers, switches, hubs, etc., for communicating packets according to network protocols to communicate audio data (such as audio streams 27, see again). Figure 1A (Example).

[0089] like Figure 2A As further shown in the example, system 150 includes local source devices 162A-162M (“local source device 162”) and remote source devices 162N-162Z (“remote source device 162”). Local source device 162 and remote source device 162 (“source device 162”) may each represent, as described above, the source device 162. Figure 1A Example of source device 12 described above. Local source device 162 can wirelessly connect to local network 152A to communicate with other devices (including target device 164) wirelessly connected to local network 152A. Similarly, remote source device 162 can wirelessly connect to local network 152A to communicate with other devices (including target device 164) wirelessly connected to local network 152A. Target device 164 can represent the above regarding... Figure 1A The example describes an example of target device 14.

[0090] In operation, target device 164 can obtain audio stream 27 from one or more of local source device 162 and remote source device 162. Target device 164 can obtain local source device 162 according to any vehicle-to-everything (V2X) protocol, such as cellular-V2X (C-V2X) protocol.

[0091] Therefore, this disclosure envisions improving the way a device allows communication or auditory experience with other people or other devices by initiating target selection to a selected target object using direct channel communication or peer-to-peer connections, V2X or C-V2X communication systems.

[0092] For example, a first device for communicating with a second device may include one or more processors configured to detect the selection of at least one target object outside the first device and to initiate a communication channel between the first device and a second device associated with the at least one target object outside the first device. Whether the selection of the at least one target object outside the first device is performed first, or whether a communication channel is initiated between the first device and the second device associated with the at least one target object outside the first device, may not be necessary. It may depend on the context or circumstances, whether a channel has been established and whether the initiation of the communication channel occurs, or whether the initiation of the communication channel is based on the detection of the selection of the at least one target object outside the first device.

[0093] For example, the communication channel between the first and second devices may have been established before the selection of at least one target object outside the devices is detected. Alternatively, the communication channel between the first and second devices may be initiated in response to the detection of the selection.

[0094] Furthermore, as a result of a communication channel between at least one target object outside the first device and the second device, one or more processors in the first device may be configured to receive audio packets from the second device. Subsequently, after receiving the audio packets, the one or more processors may be configured to decode the audio packets received from the second device to generate an audio signal and output the audio signal based on the selection of at least one target object outside the first device. It is possible that the first device and the second device can be a first vehicle and a second vehicle. This disclosure has different examples illustrating vehicles, but many of the techniques described are also applicable to other devices. That is, the two devices can be headsets, including: mixed reality headsets, head-mounted displays, virtual reality (VR) headsets, augmented reality (AR) headsets, etc.

[0095] The audio signal can be reproduced by one or more loudspeakers coupled to the first device. If the first device is a vehicle, the loudspeakers can be located in the vehicle's driver's compartment. If the first device is a headset, the loudspeakers can reproduce a binaural version of the audio signal.

[0096] Based on the selection of target objects, communication using C-V2X or V2X systems or other communication systems can be performed between one or more target objects and the first device. The second device, i.e., a headset or vehicle, can be used by one or more people to speak or play music associated with the second device. Audio / speech codecs can be used to compress speech or music emanating from inside the second vehicle or from the second headset and produce audio packets. The audio / speech codecs can be two separate codecs, such as an audio codec, or they can be a speech codec. Alternatively, a single codec can have the ability to compress both audio and speech.

[0097] Target device 164 can acquire audio stream 19' from local source device 162 to support various scenarios, such as people at parties, concerts, meetings, or other events. In some examples, audio playback device 164 can acquire audio stream 19' to support XR scenarios or experiences where users of target device 164 participate in the event through XR devices, mobile devices (including so-called smartphones), etc. The following is in conjunction with... Figures 3A-3F The example description includes additional vehicle environment information. Figure 2A-2C The rest of the discussion focuses on XR experiences, but the technology should not be limited to these XR experiences and can be extended to vehicle experiences or audio streaming in ways that some audio renderers 32 may not be able to support, or any other suitable experiences.

[0098] Assuming that target device 164 may include an audio playback device (and other functional components) similar to audio playback device 16A of target device 14, audio playback device 16A may also acquire audio stream 27 from remote source device 162 via network protocols and may use Dynamic Adaptive Streaming (DASH) based on Hypertext Transfer Protocol (HTTP). More information about DASH can be found in the international standard ISO / IEC 23009-1, entitled “Information technology – Dynamic adaptive streaming over HTTP (DASH) – Part 1: Media presentation description and segment formats”, second edition, dated May 15, 2014.

[0099] In any case, the audio playback system 16A can invoke the audio decoding device 34 to decode the audio stream 27 into an audio stream 19', some of which may have audio metadata (e.g., audio stream 19' from the local source device 162, where audio stream 19' from the remote source device 162 may not include metadata).

[0100] However, as mentioned above, the audio renderer 32 may not be able to decode all substreams of each audio stream 19'. Each substream may represent a single object of channel-based audio data, a single channel of channel-based audio data, or a single stereo reverberation coefficient of stereo reverberation audio data corresponding to a single spherical basis function (or, in other words, scene-based audio data).

[0101] In any case, to illustrate how audio renderer 32 might be unable to fully render audio stream 19', a single audio stream representing sixth-order stereo reverberation audio data can include 49 substreams, each stereo reverberation coefficient corresponding to one of the 49 spherical basis functions. In some examples, audio renderer 32 may support only 8 substreams (e.g., for rendering 7.1 channel audio data). Therefore, stream selection unit 44 (in Figure 1A (As shown in the example) the number of substreams can be reduced in a variety of different ways, as shown below regarding Figure 2B and 2C Let's discuss this in more detail.

[0102] Figure 2B This is a flowchart illustrating example operations of the stream selection unit in performing various aspects of the techniques described in this disclosure. The audio decoding device 34 may first decode the audio stream 27 to obtain a usable audio stream 19' (this refers to another way of saying audio stream 19') (170). The audio decoding device 34 may output the audio stream 19' as N audio substreams (therefore, the audio stream 19' may also be referred to as audio substream 19'), where the variable N represents the total number of audio substreams 19'. Thus, the audio decoding device 34 can obtain N audio substreams 19' (171).

[0103] The stream selection unit 44 can then determine, based on the audio sub-stream 19', the total number (e.g., N) of one or more sub-streams 19' for all the multiple audio streams 19'. The stream selection unit 44 can then compare the total number (N) with a rendering threshold (in... Figure 2B The comparison is denoted as "M" in the text (173). The rendering threshold (M) can indicate the total number of substreams supported by the audio renderer 32 when rendering the audio stream 19' to one or more speaker feeds 35. When the total number (N) is greater than the rendering threshold (M) ("Yes" 173), the stream selection unit 44 can adapt the audio stream 19' to reduce the number of substreams 19' and obtain an adapted audio stream that includes a reduced total number of substreams 19' equal to or less than the renderer threshold (M).

[0104] The stream selection unit 44 can adapt the audio stream 19' in a variety of different ways. In one example, the stream selection unit 44 can apply a user preset to the audio stream 19' (174). The audio preset can identify one or more preferred audio streams in the audio stream 19'. The stream selection unit 44 can avoid removing one or more audio streams 19' when obtaining an adapted audio stream based on the user preset.

[0105] In another example, stream selection unit 44 may apply a distance threshold (175). The distance threshold may represent a threshold defining the maximum distance (relative to target device 164) from which an audio stream can originate and is a candidate for rendering. As described above, each audio stream 19' may include audio metadata, which includes start position information identifying the starting position of the audio stream's origin. Stream selection unit 44 may adapt the audio stream 19' based on the start position information to reduce the total number of one or more sub-streams 19' and obtain an adapted audio stream.

[0106] As another example, the stream selection unit 44 can determine the type of audio data specified in the audio sub-stream 19', and adapt the audio stream 19' based on the type of audio data to reduce the total number of audio sub-streams 19', thereby obtaining an adapted audio stream. Figure 2B As shown in the example, the stream selection unit 44 can determine that the type of the audio data indicates that the audio data is stereo reverberant audio data (or in other words, a stereo reverberant stream) (176). When the type of the audio data indicates that the audio data is stereo reverberant audio data ("yes" 176), the stream selection unit 44 can apply or perform order reduction with respect to the stereo reverberant audio data to obtain a suitable audio stream (177). Since each coefficient corresponding to the same spherical basis function is represented by a separate substream, order reduction eliminates higher-order coefficients (associated with higher-order spherical basis functions), thereby reducing the number of audio substreams 19'.

[0107] When the type of audio data indicates that the audio data is not stereo reverb audio data (“No” 176), the stream selection unit 44 can determine whether the type of audio data indicates that the audio data is a multi-channel (MC) stream (178). When the type of audio data indicates that the audio data is an MC stream (“Yes” 178), the stream selection unit 44 can down-mix the MC stream 19' to reduce the number of channels of the MC stream 19' (e.g., 5.1 channel audio data down-mixed to stereo audio data or mono audio data, or, as another example, 7.1 channel audio data down-mixed to 5.1 stereo or mono audio data) (179). In this way, the stream selection unit 44 can perform down-mixing on channel-based audio data to obtain a suitable audio stream.

[0108] When the type of audio data indicates that the audio data is not an MC stream (“No” 178), the stream selection unit 44 can apply privacy settings to the audio substreams to remove any audio streams 19' marked as private or restricted (180). In some examples, the stream selection unit 44 can apply privacy settings regardless of the determination made regarding the type of audio data. Therefore, the stream selection unit 44 can adapt multiple audio streams based on privacy settings to remove one or more audio streams 19' (and all associated audio substreams 19') and obtain the adapted audio streams.

[0109] In any case, in one example, the stream selection unit 44 can apply an override to the adjusted audio substream to obtain a reduced audio substream (181). The override can indicate to the user that fewer audio streams 19' are needed, or otherwise indicate that one or more specific audio streams 19' should be selected for rendering. However, in most cases, the application of the override is optional and is therefore indicated by a dashed box. Thus, in some examples, the adjusted audio stream can be the same as the reduced audio stream.

[0110] The stream selection unit 44 can then determine whether the adjusted / reduced audio stream 19' includes a total number (N) of substreams greater than the renderer threshold (M) (173). If the total number (N) of substreams is less than the renderer threshold, the stream selection unit outputs the reduced audio substream ("No" 173), and the stream selection unit 44 can output the audio substream 19' as the adjusted audio substream, whereby the audio renderer 32 can now render one or more speaker feeds 35 based on the adjusted / reduced audio substream 19' (whereby the renderer 32 has M input constraints equal to the renderer threshold M) (182). The audio renderer 32 can then output the speaker feeds 35 to one or more speakers (183).

[0111] In some examples, the adapted audio stream includes at least one audio stream representing channel-based audio data, and the renderer includes a six-DOF renderer 32 to perform 6DOF rendering as described above. In this example, the stream selection unit 44 can obtain tracking information 41 representing the movement of the device 14, and modify the six-DOF renderer 32 to reflect the movement of the device before applying the tracking information 41.

[0112] In these and other examples, the adapted audio stream includes at least one audio stream representing stereo reverberation audio data, and the renderer 32 again includes a six-DOF renderer 32. In this example, the stream selection unit 44 can obtain tracking information 41 representing the movement of the device 14, and modify the six-DOF renderer 32 to reflect the movement of the device before applying the tracking information 41.

[0113] In the example of performing 6DOF rendering, the audio renderer 32 may have a lower rendering threshold because 6DOF rendering can consume significant resources (e.g., processor, memory, bandwidth, etc.). Therefore, the stream selection unit 44 can lower the rendering threshold to below M input constraints. For example, multichannel audio data and / or stereo reverb audio data with 6DOF rendering may not have available bandwidth, which may increase the likelihood that the stream selection unit 44 will perform down-conversion and / or down-mixing.

[0114] Figure 2C This illustrates in more detail the various aspects of the technology described in this disclosure. Figure 2A The flowchart shows an additional example operation of the flow selection unit shown in the example. Figure 2C The additional example operation of the stream selection unit shown in the example is similar to... Figure 2B The example operation of the stream selection unit shown in the example differs in that the application of overwriting is no longer optional, and the overwriting occurs in response to determining that the total number (N) of substreams 19' is greater than the rendering threshold (M).

[0115] Stream selection unit 44 can apply an overwrite to audio stream 19' to select the preferred audio stream of audio stream 19'. When the total number of audio substreams of the preferred audio stream still exceeds the threshold M ("Yes" 184), stream selection unit 44 can perform stereo reverb stream determination, and then apply down-order (177) or determine whether the audio data is an MC stream (178). When determined to be an MC stream, stream selection unit 44 can perform down-mixing (178). When determined not to be an MC stream, stream selection unit 44 only outputs the original audio stream.

[0116] In some examples, the above about Figure 2B and 2C The two example operations described are performed iteratively, either in response to a new audio stream being retrieved. Therefore, these techniques can iteratively reduce the number of substreams to satisfy the M input constraints of the audio renderer 32.

[0117] Thus, the stream selection unit 44 can process audio streams representing mono, stereo, multi-channel (e.g., 5.1 or 7.1 channels) and / or stereo reverberation audio data (e.g., fourth or sixth order).

[0118] Figure 2D-Figure 2K It is shown by Figure 1A and Figure 1B The example shown is a diagram illustrating the application of privacy settings on the source device and / or content consumer device. Figure 2D-Figure 2K The following discussion provides information as described above regarding... Figure 2BAdditional details regarding the application of privacy settings discussed. In some use cases, it may be desirable to be able to control which of the multiple audio streams generated by the source device 12 are available for playback by the content consumer device 14.

[0119] For example, audio from some capture devices of content capture device 20 may contain sensitive information, and / or audio from some capture devices of content capture device 20 may not be intended for exclusive access (e.g., unrestricted access for all users). It may be desirable to restrict access to audio from some capture devices of content capture device 20 based on the type of information captured by content capture device 20 and / or based on the location of the physical zone where content capture device 20 is located.

[0120] like Figure 2D As shown in the example, the stream selection unit 44 can determine that the VLI 45B indicates that the content consumer device 14 (shown as VR device 400) is in a virtual location 401. VR device 400 can be a listener on a 6DoF playback system. The stream selection unit 44 can then determine one or more of the audio elements 402A-402H, where CLI 45A (which may represent an audio stream captured by a microphone, including...) Figure 1A The microphone 18 shown, as well as other types of capture devices, including microphone arrays, microphone clusters, other XR devices, mobile phones (including so-called smartphones), etc., are also included. Furthermore, audio streams 402A-402H may include synthesized audio generated via a computer or other audio output device.

[0121] As described above, the stream selection unit 44 can obtain the audio stream 27. The stream selection unit 44 can interface with audio elements 402A-402H and / or with the source device 12 to obtain the audio stream 27. In some examples, the stream selection unit 44 can interact with an interface (e.g., a receiver, transmitter, and / or transceiver) to comply with fifth-generation (5G) cellular standards, such as Bluetooth. TM Audio streams are obtained through Personal Area Networks (PANs) or other open-source, proprietary, or standardized communication protocols.27 Figure 2A In the example, wireless communication of the audio stream is represented as a Lightning connection, where the selected audio stream 19' is shown communicating from one or more selected audio elements 402 and / or source device 12 to VR device 400.

[0122] exist Figure 2D In the example, VR device 400 is located at position 401, which is near audio source 408. Using the techniques described above, and in more detail below, VR device 400 can use an energy map to determine that audio source 408 is located at position 401. Figure 2DAudio elements 402D-402H are shown at location 401. Audio elements 402A-402C are not near VR device 400.

[0123] In one example of this disclosure, source device 12 can be configured to generate audio metadata that includes privacy restrictions for multiple audio streams. For example, such as Figure 2D As shown, source device 12 can be configured to generate audio metadata indicating that the audio stream associated with audio element 402H is restricted for the user of VR device 400 (or any other content consumer device). Source device 12 can transfer audio metadata to VR device 400 (or any other content consumer device).

[0124] VR device 400 can be configured to receive multiple audio streams and corresponding audio metadata, and store them in memory. Each audio stream represents a sound field, and the audio metadata includes restrictions on one or more of the multiple audio streams. VR device 400 can be configured to determine one or more audio streams from the audio metadata based on privacy restrictions. For example, VR device 400 can be configured to determine which audio streams can be played back based on privacy restrictions. VR device 400 can then generate a corresponding sound field based on one or more audio streams. Similarly, VR device 400 can be configured to determine one or more restricted audio streams (e.g., the audio stream associated with audio element 402H) from the audio metadata based on privacy restrictions, without generating a corresponding sound field for one or more restricted audio streams.

[0125] Figure 2E This is a block diagram illustrating the operation of controller 31 in one example of this disclosure. In one example, controller 31 may be implemented as processor 712. Reference is made below. Figure 7 A more detailed description of the processor 712 is provided above. Figure 1A The source device 12 can capture audio data using the content capture device 20. The content capture device 20 can capture audio data from audio element 18. Audio element 18 can include a static source, such as a static single microphone or a microphone cluster. Microphone 18 can be a real-time source. Alternatively or additionally, audio element 18 can include a dynamic audio source (e.g., dynamic in terms of usage and / or location), such as a mobile phone. In some examples, the dynamic audio source can be a synthesized audio source. The audio stream can originate from audio sources at a single physical interval or from a cluster of audio sources in a single physical location.

[0126] In some examples, grouping physically close audio sources into a cluster can be beneficial because each individual audio source in a cluster physically located can sense some or all of the audio as every other audio source in the cluster. Therefore, in some examples of this disclosure, controller 31 can be configured to switch audio from the audio source cluster ( Figure 2E The audio stream marked C in the middle, and the switching from a single audio source / element ( Figure 2E An audio stream is labeled R. In this context, a toggle can refer to marking an audio stream or group of audio streams as unrestricted (e.g., capable of decoding and / or playing) or restricted (e.g., not capable of decoding and / or playing). Turning the privacy toggle on (e.g., restricted) instructs the VR device to mute and / or generally not decode or play back the audio stream. Turning the privacy toggle off (e.g., unrestricted or public access) instructs any user to decode and play back the audio stream. In this way, audio engineers or content creators can grant exclusive access to a specific audio source to unrestricted users or based on tiered privacy settings.

[0127] like Figure 2E As shown, controller 31 can be configured to receive and / or access multiple audio streams captured by content capture device 20. Controller 31 can be configured to check for any privacy settings associated with the audio streams. That is, controller 31 can be configured to identify one or more unrestricted audio streams and one or more restricted audio streams from the multiple audio streams.

[0128] In some examples, the content creator can be configured to set privacy settings at each audio source or audio source cluster. In other examples, controller 31 can be configured to determine, for example, via explicit instructions whether privacy settings are required for a collection of multiple audio streams. In one example, controller 31 can receive a cluster graph including audio metadata 404 that indicates privacy restrictions on one or more audio sources or audio source clusters. In one example, the privacy restriction indicates whether one or more of the multiple audio streams are restricted or unrestricted. In other examples, the privacy restriction only indicates restricted audio streams. As will be explained in more detail, privacy restrictions may restrict a single audio source, a group (cluster) of audio sources, or indicate restrictions between audio sources (inter-group restrictions).

[0129] exist Figure 2E In the example, audio metadata 404 also includes individual privacy restrictions for each audio stream, indicating whether one or more of the multiple audio streams are restricted or unrestricted for each of the multiple privacy setting levels. Audio metadata 404 includes two privacy setting levels. Of course, more or fewer privacy setting levels can be used. Figure 2EThe audio metadata 404 only indicates which audio source clusters or individual audio sources are restricted for a specific privacy setting level. VR device 400 can determine that any audio stream from an audio source not listed as restricted in the metadata 404 can be unrestricted (i.e., playable). VR device 400 can determine that any audio stream from an audio source listed as restricted in the metadata 404 cannot be played. In other examples, the audio metadata 404 may indicate unrestricted and restricted audio sources / streams for each privacy setting level.

[0130] Controller 31 can be configured to embed audio metadata 404 into bitstream 27 and / or side channel 33 (see [link]). Figure 1A The audio metadata is transmitted to VR device 400 or any other content consumer device, including... Figure 1A and Figure 1B Content consumer device 14. Furthermore, in some examples, controller 31 may be configured to generate a privacy setting level among multiple privacy setting levels for VR device 400 and transmit the privacy setting level to VR device 400.

[0131] As described above, controller 31 can be part of any number of devices, including servers, network-connected servers (e.g., cloud servers), content capture devices, and / or mobile handheld devices. Controller 31 can be configured to transmit multiple audio streams via wireless links (including 5G air interfaces) and / or personal area networks (e.g., Bluetooth interfaces).

[0132] In some examples, when controller 31 is configured to send multiple audio streams to VR device 400 (e.g., in so-called online mode), controller 31 can be configured not to transmit any audio streams marked as restricted in audio metadata 404 to VR device 400. Controller 31 can still transmit audio metadata 404 to VR device 400 so that VR device 400 can determine which audio streams are being received. In other examples, controller 31 may not transmit any audio streams to VR device 400. Instead, in some examples, VR device 400 may receive audio streams directly from one or more audio sources. In these examples, VR device 400 still receives audio metadata 404 from the controller (or directly from the audio source). VR device 400 will then determine unrestricted and restricted audio streams based on audio metadata 404 and VR device 400's privacy settings level, and will avoid decoding and / or playing back any audio streams marked as restricted (e.g., based on privacy settings level).

[0133] In summary, in one example, VR device 400 can be configured to receive a privacy setting level from multiple privacy setting levels, decode audio metadata 404, and access corresponding privacy restrictions indicating whether one or more of the multiple audio streams are restricted or unrestricted corresponding to the received privacy setting level. VR device 400 can be configured to receive multiple audio streams via a wireless link (e.g., a 5G air interface and / or a Bluetooth interface). As mentioned above, VR device 400 can be an extended reality headset. In this example, VR device 400 can include a head-mounted display configured to present a displayed world. In other examples, VR device 400 can be a mobile handheld device.

[0134] In other examples of this disclosure, the VR device 400 may be configured to further perform the energy mapping technique of this disclosure in conjunction with the aforementioned audio metadata privacy restrictions. In this example, the audio metadata also includes capture location information representing a capture location in the displayed world, at which a corresponding one of a plurality of audio streams is captured. The VR device 400 may be further configured to determine location information representing the device's position in the displayed world, select a subset of the plurality of audio streams based on the location information and the capture location information, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams, and generate a corresponding sound field based on the subset of the plurality of audio streams.

[0135] Figure 2F This is a conceptual diagram illustrating an example of a single audio source (R4) being marked as restricted in audio metadata 406. In this example, controller 31 can be configured to generate audio metadata 406 to include privacy restrictions indicating whether the audio stream from a first audio capture device (e.g., R4) is restricted or unrestricted. VR device 400 receives a privacy setting level of 3. VR device 400 can then determine that audio source R4 is restricted by the privacy setting level 3 column of audio metadata 406. Therefore, VR device 400 will avoid decoding and / or playing back the audio stream from audio source R4. Figure 2F The example can be applied to situations where the physical propagation of an audio source is sufficient to enable the individual switching of a single audio source. Audio engineers or content creators can choose to switch (i.e., represented as unrestricted or restricted) certain audio elements based on physical propagation.

[0136] Figure 2GThis is a conceptual diagram illustrating an example where a cluster (C1) of audio sources is marked as restricted in audio metadata 408. In this example, controller 31 can be configured to generate audio metadata 408 to include privacy restrictions indicating whether the audio stream from a first cluster (e.g., C1) of the audio capture device is restricted or unrestricted. VR device 400 receives a privacy setting level of 3. VR device 400 can then determine that cluster C1 of the audio sources is restricted by the privacy setting level 3 column of the audio metadata 408. Therefore, VR device 400 will avoid decoding and / or playing back the audio stream from cluster C1 of the audio sources. Figure 2G Examples of this approach can be applied to situations where audio sources are densely clustered or grouped in a physical location, rendering individual switching of a single audio source ineffective. In some examples, a single audio source within a cluster can be designated as the primary audio source, and privacy restrictions on switching the primary audio source affect all other audio sources within the cluster. A proximity (e.g., distance) threshold can be used to determine which audio sources belong to the cluster.

[0137] Figure 2H This is a conceptual diagram illustrating an example where a cluster of audio sources (C1) is marked as restricted in the audio metadata 410. Furthermore, metadata 410 includes sub-columns. Any audio source in a sub-column inherits the privacy restrictions from the restricted column markings in the audio metadata 410. Therefore, a cluster of audio sources C2 is also marked as restricted. In this way, certain audio sources or clusters of audio sources can depend on each other, and controller 31 can affect multiple clusters or audio sources by simply switching a single cluster or audio source.

[0138] In this example, controller 31 can be configured to generate audio metadata 410 to include information indicating that audio streams from a second cluster (e.g., C2) of the audio capture device share the same privacy restrictions as the first cluster (e.g., C1) of the audio capture device. VR device 400 receives a privacy setting level of 3. VR device 400 can then determine that clusters C1 and C2 of the audio sources are restricted by the privacy setting level 3 column of the audio metadata 410. Therefore, VR device 400 will avoid decoding and / or playing back audio streams from clusters C1 and C2 of the audio sources.

[0139] In other examples of this disclosure, source device 12 and content consumer device 14 may use cryptographic techniques to restrict the decoding and / or playback of certain audio streams. The cryptographic techniques described below may be used alone or in conjunction with those referenced above. Figure 2D-2H The described privacy-based audio metadata technology is used in combination.

[0140] Figure 2I This is a block diagram illustrating the operation of controller 31 in one example of this disclosure. In one example, controller 31 may be implemented as processor 712. Reference is made below. Figure 7 A more detailed description of the processor 712 is provided above. Figure 1A The source device 12 may capture audio data using the content capture device 20. The content capture device 20 may capture audio data from the microphone 18. The microphone 18 may include a static source, such as a static single microphone or a microphone cluster. The microphone 18 may be a real-time source. Alternatively or additionally, the microphone 18 may include a dynamic audio source (e.g., dynamic in terms of use and / or location), such as a mobile phone. In some examples, the dynamic audio source may be a synthesized audio source. The audio stream may originate from audio sources at a single physical interval or from a cluster of audio sources at a single physical location.

[0141] In some examples, grouping physically close audio sources into clusters or partitions can be beneficial, as each individual audio source in a cluster physically located can sense some or all of the audio as every other audio source in the same physical partition. Therefore, in some examples of this disclosure, controller 31 can be configured to mask, zero, and / or toggle audio streams from audio source partitions. In this context, a mask partition may refer to adjusting the audio gain of the partition. A zero partition may refer to muting the audio from that partition (e.g., using beamforming). A toggle partition may refer to marking an audio stream or group of audio streams as unrestricted (e.g., capable of decoding and / or playing) or restricted (e.g., not capable of decoding and / or playing). Turning on a privacy toggle (e.g., restricted) instructs the VR device to mute and / or generally not decode or play back the audio stream. Turning off a privacy toggle (e.g., unrestricted or public access) instructs any user to decode and play back the audio stream. In this way, audio engineers or content creators can grant exclusive access to specific audio sources to unrestricted users or based on tiered privacy settings.

[0142] like Figure 2I As shown, controller 31 can be configured to receive and / or access multiple audio streams captured by content capture device 20. Controller 31 can be configured to partition the audio streams into certain partitions based on the physical location of the audio sources. In some examples, controller 31 can tag (e.g., generate metadata) indicating which partition a particular audio source belongs to. Controller 31 can also generate boundary metadata for the partitions, including centroid location and radius.

[0143] In some examples, the content creator can be configured to set privacy settings at each audio source or audio source cluster / partition. In other examples, the controller 31 can be configured to determine, for example, whether privacy settings are required for a collection of multiple audio streams via explicit instructions received by the controller 31. Based on the privacy settings for a partition, the controller 31 can enable a password generator to generate passwords for the partition's specific privacy settings. In some examples, the controller 31 can encrypt passwords based on an encryption type (e.g., Advanced Encryption Standard, Rivest-Shamir-Adleman (RSA) encryption, etc.).

[0144] Controller 31 can be configured to embed cryptography into bitstream 27 and / or side channel 33 (see...) Figure 1A The password is transmitted to VR device 400 or any other content consumer device, including... Figure 1A and Figure 1B The content consumer device 14. VR device 400 can be configured to receive a password from controller 31 (or from another source) and send the password back to controller 31 when an audio stream is requested. The embedding and authentication block of controller 31 embeds the individual passwords generated for the partition or audio source into the audio stream and metadata retrieved by controller 31. The embedding and authentication block of controller 31 also performs authentication based on the password provided by VR device 400. In one example, controller 31 can be configured to send only unrestricted audio streams to VR device 400 based on the authenticated password. In other examples, controller device 31 can be configured to send one or more audio streams to VR device 400 along with instructions on how the audio streams should be muted, muted, and / or toggled.

[0145] As described above, privacy settings can include blocking partitions, zeroing partitions, or switching partitions to one or more of restricted or unrestricted states. In one example, switching partitions indicates whether one or more of multiple audio streams are restricted or unrestricted. In other examples, switching privacy restrictions only indicates restricted audio streams.

[0146] In one example of this disclosure, controller 31 may be configured to store multiple audio streams, each representing a sound field, and to generate one or more of the multiple audio streams based on privacy restrictions associated with a password. In one example, controller 31 may be configured to transmit one or more of the multiple audio streams to a content consumer device.

[0147] In one example of this disclosure, the password is a master password associated with unrestricted privacy restrictions. In this example, controller 31 can be configured to generate each of multiple audio streams. The master password can be a superuser / administrator password. The master password grants unrestricted access to all audio streams.

[0148] In another example of this disclosure, the password is a permanent password associated with conditional privacy restrictions. In this example, controller 31 can be configured to generate one or more of a plurality of audio streams based on conditional privacy restrictions, wherein the conditional privacy restrictions indicate whether one or more of the plurality of audio streams are restricted or unrestricted. In one example, controller 31 can be configured to generate audio metadata (e.g., the audio metadata described above), which also includes respective conditional privacy restrictions indicating whether one or more of the plurality of audio streams are restricted or unrestricted based on the permanent password. As described below, conditional privacy restrictions may include masking (e.g., indicated by a gain value), zeroing, and / or toggling. In one example, the permanent password remains valid until reset. Controller 31 can generate permanent passwords for individual partitions and / or audio sources.

[0149] In another example of this disclosure, the password is a temporary password associated with conditional privacy restrictions. In this example, controller 31 can be configured to generate one or more of a plurality of audio streams based on conditional privacy restrictions, wherein the conditional privacy restriction permissions indicate whether one or more of the plurality of audio streams are restricted or unrestricted. In one example, controller 31 can be configured to generate audio metadata (e.g., the audio metadata described above), which also includes respective conditional privacy restrictions indicating whether one or more of the plurality of audio streams are restricted or unrestricted based on the temporary password. As described below, conditional privacy restrictions may include masking (e.g., indicated by a gain value), zeroing, and / or toggling. In one example, the permanent password remains valid until reset. In one example, the temporary password remains valid for a fixed duration and expires after the fixed duration. Controller 31 may automatically invalidate the temporary password after the duration expires.

[0150] In one example, the privacy limit includes a corresponding gain value associated with one or more of the corresponding audio streams. In another example, the privacy limit includes a corresponding zeroing indication associated with one or more of the corresponding audio streams. In yet another example, the privacy limit includes a corresponding toggle indication associated with one or more of the corresponding audio streams.

[0151] VR device 400 can be configured to store multiple audio streams, each representing a sound field, receive one or more of the multiple audio streams based on privacy restrictions associated with a password, and generate a corresponding sound field based on one or more of the multiple audio streams. In one example, VR device 400 sends a password to controller 31.

[0152] In one example, the password is a master password associated with unrestricted privacy restrictions, and the VR device 400 can be configured to receive each of multiple audio streams.

[0153] In another example, the password is a permanent password associated with conditional privacy restrictions, and the VR device 400 can be configured to receive one or more of a plurality of audio streams based on the conditional privacy restrictions, wherein the conditional privacy restrictions indicate whether one or more of the plurality of audio streams are restricted or unrestricted. The VR device 400 can be further configured to receive audio metadata, which also includes respective conditional privacy restrictions indicating whether one or more of the plurality of audio streams are restricted or unrestricted based on the permanent password.

[0154] In another example, the password is a temporary password associated with conditional privacy restrictions, and the VR device 400 can be configured to receive one or more of a plurality of audio streams based on the conditional privacy restrictions, wherein the conditional privacy restrictions indicate whether one or more of the plurality of audio streams are restricted or unrestricted. The VR device 400 can be further configured to receive audio metadata, which also includes respective conditional privacy restrictions indicating whether one or more of the plurality of audio streams are restricted or unrestricted based on the temporary password.

[0155] In one example, VR device 400 may be configured to receive a password from a host (e.g., controller 31). In another example, VR device 400 may be configured to receive a password from a source other than a host.

[0156] Figure 2J This diagram illustrates examples of masking and zeroing partitions and / or individual audio sources. In scenario 420, a password associated with the privacy restrictions of masking partition 2 (audio sources R7-R9) is issued to VR device 400. In this example, VR device 400 further receives the gain value of partition 2 to apply when playing back the audio stream from partition 2. In scenario 430, a password associated with the privacy restrictions of zeroing audio source R4 is issued to VR device 400. In this example, VR device 400 can completely mute the audio stream from audio source R4 (e.g., by beamforming or applying zero gain).

[0157] Figure 2K This diagram illustrates examples of switching partitions and / or individual audio sources. In scene 440, a password is issued to VR device 400 associated with switching partition 2 (audio sources R7-R9) to a restricted privacy setting. In this example, VR device 400 disables decoding and / or playback of the audio stream from partition 2. In scene 450, a password is issued to VR device 400 associated with switching the privacy setting of audio source R4. In this example, VR device 400 disables decoding and / or playback of the audio stream from audio source R4.

[0158] Figures 3A-3FThis is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure. First refer to... Figure 3A For example, system 10 includes a source device 12 and a target device 14. Source device 12 is shown as a bicycle 12, which communicates with target device 14 via a wireless connection 200. Target device 14 is shown as a vehicle for illustrative purposes. Although described with respect to bicycle 12 and vehicle 14, these techniques can be implemented by any type of device capable of wireless communication, including mobile handheld devices (including so-called "smartphones"), watches (including so-called "smartwatches"), laptops, audio systems (including so-called "infotainment systems"), etc.

[0159] Bicycle 12 and vehicle 14 can be operated manually or automatically. For illustrative purposes, it is assumed that bicycle 12 is manually operated by a driver (for ease of explanation, in...). Figure 3A (not shown in the example), and assume that vehicle 14 is operated automatically.

[0160] like Figure 3A As shown in the example, vehicle 14 operates automatically by traveling in the right lane 202 of road 210, which has a right lane 202 and a left lane 204. Vehicle 14 can establish a wireless connection 200 with bicycle 12 when bicycle 12 is sufficiently close (near or within a certain threshold). Alternatively, bicycle 12 can establish a wireless connection 200 when vehicle 14 is sufficiently close (near or within a threshold). Regardless of which device 12 / 14 initiates the wireless connection 200, it can be based on fifth-generation (5G) cellular standards and Personal Area Network (PAN) protocols (e.g., or any other wireless communication protocol (including WiFi) TM The wireless connection 200 can be established using protocols such as Vehicle-to-Everything (V2X) protocol, including the so-called Cellular V2X (C-V2X) protocol.

[0161] Under any circumstances, vehicle 14 can communicate with bicycle 12 via wireless connection 200. The rider of bicycle 12 (for illustrative purposes, in...) Figure 3A (Not shown again in the example) can interface with a source device (for simplicity, it is assumed that the source device is integrated into bicycle 12, and therefore bicycle 12 can be referred to as source device 12) to capture audio data (for simplicity, it is assumed that the audio data is the speech of the rider of bicycle 12). Although it is assumed that source device 12 is integrated into bicycle 12, source device 12 can be detached from bicycle and can represent the above-mentioned... Figure 1A The example describes one or more devices.

[0162] Bicycle 12 may include a microphone or other audio capture device configured to capture audio data and generate an audio stream based on the audio data. Bicycle 12 may generate an audio stream and audio metadata, including the starting coordinates (e.g., GPS coordinates) of the origin of the corresponding audio stream. Bicycle 12 may output the audio stream to vehicle 14 via wireless connection 200.

[0163] Vehicle 14 can receive audio streams and audio metadata, storing the audio stream along with the audio metadata. Vehicle 14 can determine the direction of arrival 212A of the audio stream based on its current coordinates relative to the starting coordinates corresponding to the audio stream. Figure 3A In the example, vehicle 14 can determine that the audio stream is arriving from directly behind vehicle 14 and proceeding in parallel to the front of vehicle 14 (as indicated by the arrow).

[0164] Next, vehicle 14 can render the audio stream as if it were arriving from direction 212A, thereby generating a speaker feed to simulate a sound field captured by bicycle 12 and arriving from directly behind vehicle 14. In this example, vehicle 14 can... Figure 3A The example generates right and left rear speaker feeds, which are output to the right and left rear speakers of vehicle 14 (not shown for illustration purposes) to reproduce the sound field 214A represented by the audio stream.

[0165] exist Figure 3A In the example, suppose the rider of bicycle 12 captures audio data of the rider issuing a command so that vehicle 14 knows the rider will “pass from the left.” Although the following description refers to what is said, bicycle 12A can provide one or more audio streams, including pre-recorded audio streams, live audio streams, or any other type of audio stream.

[0166] When assuming automated operation (e.g., a computing device controls vehicle 14 and issues commands causing the computing device to steer, accelerate, brake, and otherwise operate vehicle 14 without human intervention), vehicle 14 can analyze the audio stream to extract commands and operate vehicle 14 based on the commands parsed from the audio stream to avoid merging into the lane of bicycle 12 or otherwise affecting the operation of bicycle 12. That is, vehicle 14 can automatically adjust its operation based on commands or other spoken words.

[0167] Next reference Figure 3BFor example, bicycle 12 has begun to overtake vehicle 14 by moving into the left lane 204 of road 210. Assuming bicycle 12 continues to provide an audio stream to vehicle 14 via wireless connection 200, bicycle 12 can continue to update the audio metadata of the audio stream to represent updated starting coordinates indicating that the bicycle is located in the left lane 204 next to vehicle 14. Based on the updated starting coordinates and the current coordinates of vehicle 14, vehicle 14 can determine a direction of arrival 212B. Vehicle 14 can then render a speaker feed that spatializes the audio stream based on the direction of arrival 212B to appear as if it is arriving from the direction of arrival 212B. In this example, vehicle 14 can render a left rear speaker feed from the audio stream and output the left rear speaker feed to the left rear speaker to reproduce a sound field 214B that reflects the new direction of arrival 212B.

[0168] Next reference Figure 3C In the example, bicycle 12 has largely passed vehicle 14 in the left lane 204 of road 210. Assuming bicycle 12 continues to provide an audio stream to vehicle 14 via wireless connection 200, bicycle 12 can continue to update the audio metadata of the audio stream to represent updated starting coordinates indicating that the bicycle has almost passed vehicle 14 in the left lane 204. Based on the updated starting coordinates and vehicle 14's current coordinates, vehicle 14 can determine an arrival direction 212C. Vehicle 14 can then render a speaker feed that spatializes the audio stream based on the arrival direction 212C to appear as if it is arriving from the arrival direction 212C. In this example, vehicle 14 can render a left front speaker feed from the audio stream and output the left front speaker feed to the left front speaker to reproduce a sound field 214C that reflects the new arrival direction 212C.

[0169] Next reference Figure 3D For example, another vehicle may arrive, so vehicle 14 can establish another wireless connection 200B (where the original wireless connection 200 with bicycle 12 is represented as wireless connection 200A, and the original bicycle 12 is represented as bicycle 12A). An additional vehicle can act as another source device 12B; therefore, the additional vehicle can be represented as vehicle 12B, since it is assumed that source device 12B is integrated into vehicle 12B (e.g., as a vehicle audio head unit). Although it is assumed that source device 12B is integrated into vehicle 12B, source device 12B can be detached from the vehicle and can be represented as described above regarding... Figure 1A The example describes one or more devices.

[0170] In any case, vehicle 12B can provide one or more audio streams, including pre-recorded audio streams, live audio streams, or any other type of audio stream. Vehicle 12B can output the audio stream along with corresponding audio metadata, which includes the starting position (or in other words, starting coordinates) where the audio stream appears to originate. Vehicle 12B can output the audio stream along with the corresponding audio metadata via wireless connection 200B, which can be similar to the wireless connection 200 described above (and as...). Figure 3D The example now shown is a wireless connection 200A.

[0171] Vehicle 14 can receive an audio stream and corresponding audio metadata via wireless connection 200B. Vehicle 14 can display an indication of the availability of the audio stream via an integrated source device (which may include a display). The operator of vehicle 14 can select one of the indications to initiate playback of the corresponding audio stream, and then vehicle 14 can determine another direction of arrival 212D reflecting the position of vehicle 12B relative to vehicle 14 based on the corresponding starting position described in the audio metadata (e.g., GPS coordinates specifying the current position of vehicle 12B) and the current coordinates of vehicle 14.

[0172] Vehicle 14 may render one of the selected audio streams based on the direction of arrival 212D to generate one or more speaker feeds, wherein in this case, vehicle 14 may render the left rear speaker feed to reflect the position of vehicle 12B relative to the position of vehicle 14. Therefore, vehicle 14 may render the left rear speaker feed to spatialize one of the selected audio streams. Vehicle 14 may output the left rear speaker feedback to the left rear speaker to reproduce the sound field 214D represented by one of the selected audio streams.

[0173] exist Figure 3E In the example, vehicle 12C can operate as a source device (and may be referred to as "source device 12C"), providing immersive video and / or audio via a wireless connection 200C to network 220. Network 220 may include a public network (such as the Internet) or a private network in which a collection of computing devices (which may include routers, switches, hubs, etc.) are interconnected to facilitate communication of data packets between each other. Vehicle 12C can provide video and audio so that extended reality (XR) device 100 can view and / or hear the environment in which vehicle 12C operates, providing something that may be called a "racing experience," which allows the user experience of XR headset 100 (another way of referring to XR device 100) to ride in vehicle 12C.

[0174] In some cases, XR device 100 may transmit an audio stream back to vehicle 12C, thus operating vehicle 12C as the target device (hence the “12C / 14C” digital identifier for vehicle 12C / 14C). XR device 100 may include audio metadata specifying the right front passenger seat (or any other passenger seat) as the starting position. Vehicle 14C may receive the audio stream from XR device 100, determine the direction of arrival 212E based on the starting position associated with the audio stream, and render the audio stream to one or more speaker feeds based on the direction of arrival 12E. These speaker feeds spatialize the audio stream to simulate audio from XR device 100 as if the user of XR device 100 were sitting in the right front passenger seat of vehicle 14C.

[0175] This allows the operator of vehicle 14C to hear the voice of the user of XR device 100 as if the user of XR device 100 were sitting in the right front passenger seat acting as a front passenger. In this example, vehicle 14C can render the right front speaker feed and output the right front speaker feed to the right front speaker to reproduce the sound field 214E represented by the audio stream sent by XR device 100. Although described for XR device 100, it can be applied to any device (e.g., Figure 3F The mobile phone 230 shown in the example performs these techniques.

[0176] Furthermore, although the V2X protocol has been described, these technologies can retrieve audio streams from non-V2X connections, such as those described above regarding network 220, XR device 100, and mobile phone 230. Vehicle 14C can retrieve audio streams from the Internet or other public or private networks represented by network 220. Vehicle 14C can use Dynamic Adaptive Streaming (DASH) based on Hypertext Transfer Protocol (HTTP) to obtain these audio streams.

[0177] Figure 4 This is a diagram illustrating an example of a VR device 400 worn by a user 402. The VR device 400 is coupled to or otherwise includes headphones 404, which can reproduce the sound field represented by audio data 19' via playback from a speaker feed 35. The speaker feed 35 may represent an analog or digital signal capable of causing the diaphragm within the transducer of the headphones 104 to vibrate at various frequencies, a process commonly referred to as driving the headphones 104.

[0178] Video, audio, and other sensory data can play a significant role in VR experiences. To participate in a VR experience, user 402 may wear VR device 400 (also referred to as VR headset 400) or other wearable electronic devices. VR client devices (such as VR headset 400) may include tracking devices (e.g., tracking device 40) configured to track the head movements of user 402 and adapt video data displayed via VR headset 400 to take head movements into account, providing an immersive experience where user 402 can experience the displayed world shown in the video data in visual three dimensions. The displayed world may refer to a virtual world (where all worlds are simulated), an augmented world (where a portion of the world is augmented by virtual objects), or a physical world (where images of the real world are virtualized for navigation).

[0179] While VR (and other forms of AR and / or MR) allows users 402 to visually reside in a virtual world, VR headsets 400 may often lack the ability to audibly place the user in the displayed world. In other words, a VR system (which may include a computer responsible for rendering video and audio data—this computer is not shown in the example of Figure 2 for illustrative purposes, and the VR headset 400) may not support a fully immersive 3D auditory experience (and in some instances realistically reflect the displayed scene presented to the user via the VR headset 400).

[0180] Although VR devices have been described in this disclosure, various aspects of these technologies can be performed in the context of other devices, such as mobile devices. In this case, a mobile device (e.g., a so-called smartphone) can present the displayed world through a screen that can be mounted on the head of user 402 or viewed as is typically done when using a mobile device. Therefore, any information on the screen can be part of the mobile device. The mobile device may be able to provide tracking information 41, thereby allowing both a VR experience (when head-mounted) and a normal experience to view the displayed world, where the normal experience still allows the user to view the displayed world that provides a VR-lite type experience (e.g., raising the device and rotating or panning it to view different parts of the displayed world).

[0181] Figure 1B This is a block diagram illustrating another example system 50 configured to perform various aspects of the techniques described in this disclosure. System 50 is similar to... Figure 1A The system 10 shown is different in that Figure 1A The audio renderer 32 shown is replaced by a binaural renderer 42 that can perform binaural rendering using one or more head-related transfer functions (HRTF) or other functions that can render to the left and right speaker feeds 43.

[0182] The audio playback system 16B can output left and right speaker feeds 43 to headphones 44, which can represent another example of a wearable device and can be coupled to additional wearable devices to facilitate sound field reproduction, such as watches, the aforementioned VR headsets, smart glasses, smart clothing, smart rings, smart bracelets, or any other type of smart jewelry (including smart necklaces). Headphones 44 can be coupled to additional wearable devices wirelessly or via a wired connection.

[0183] In addition, the headphones 44 can be connected via wired connection (such as a standard 3.5mm audio jack, Universal System Bus (USB) connection, optical audio jack or other forms of wired connection) or wireless connection (such as via Bluetooth). TM The headphones 44 are coupled to the audio playback system 16B via a connection (such as a wireless network connection). The headphones 44 can reconstruct the sound field represented by the audio data 19' based on the left and right speaker feeds 43. The headphones 44 may include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the respective left and right speaker feeds 43.

[0184] Figure 5 This is a diagram illustrating an example of a wearable device 500 that can operate according to various aspects of the techniques described in this disclosure. In various examples, wearable device 500 may represent a VR headset (such as the VR headset 100 described above), an AR headset, a MR headset, or any other type of extended reality (XR) headset. Augmented Reality “AR” can refer to computer-rendered images or data overlaid on the real world in which the user actually resides. Mixed Reality “MR” can refer to computer-rendered images or data locked to a specific location in the real world, or can refer to a variant of VR in which a combination of computer-rendered 3D elements and filmed real elements is combined to create an immersive experience that simulates the user’s physical presence in the environment. Extended Reality “XR” may represent a general term for VR, AR, and MR. More information on the XR terminology can be found in the document entitled “Virtual Reality, Augmented Reality, and Mixed Reality Definitions” written by Jason Peterson on July 7, 2017.

[0185] Wearable device 500 can refer to other types of devices, such as watches (including so-called "smartwatches"), glasses (including so-called "smart glasses"), headphones (including so-called "wireless headphones" and "smart headphones"), smart clothing, smart jewelry, etc. Whether representing VR devices, watches, glasses, and / or headphones, wearable device 500 can communicate with computing devices that support wearable device 500 via wired or wireless connections.

[0186] In some instances, the computing device supporting the wearable device 500 may be integrated within the wearable device 500; therefore, the wearable device 500 may be considered the same device as the computing device supporting the wearable device 500. In other instances, the wearable device 500 may communicate with a separate computing device that supports the wearable device 500. In this regard, the term "support" should not be construed as requiring a separate dedicated device, but one or more processors configured to perform various aspects of the technologies described in this disclosure may be integrated within the wearable device 500 or within a computing device separate from the wearable device 500.

[0187] For example, according to various aspects of the technology described in this disclosure, when wearable device 500 represents VR device 500, a separate dedicated computing device (e.g., a personal computer including one or more processors) can render audio and video content, while wearable device 500 can determine translational head movements, and the dedicated computing device can render audio content (as a speaker feed) based on these translational head movements. As another example, when wearable device 500 represents smart glasses, wearable device 500 may include one or more processors that both determine translational head movements (via an interface within one or more sensors of wearable device 500) and render speaker feeds based on the determined translational head movements.

[0188] As shown in the figure, wearable device 500 includes one or more directional speakers and one or more tracking and / or recording cameras. Additionally, wearable device 500 includes one or more inertial, haptic, and / or health sensors, one or more eye-tracking cameras, one or more high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of wearable device 500 may include durable translucent display technology and hardware.

[0189] Wearable device 500 also includes connectivity hardware, which may represent one or more network interfaces supporting multi-mode connectivity, such as 4G communication, 5G communication, Bluetooth, etc. Wearable device 500 also includes one or more ambient light sensors and bone conduction sensors. In some cases, wearable device 500 may also include one or more passive and / or active cameras with fisheye lenses and / or telephoto lenses. Although not explicitly stated... Figure 5 As shown, wearable device 500 may also include one or more light-emitting diode (LED) lights. In some examples, the LED lights may be referred to as "super-bright" LED lights. In some embodiments, wearable device 500 may also include one or more rear-facing cameras. It should be understood that wearable device 500 may exhibit a variety of different form factors.

[0190] Furthermore, tracking and recording cameras and other sensors can facilitate the determination of translation distance. Although in Figure 5 The example is not shown, but the wearable device 500 may include other types of sensors for detecting translational distance.

[0191] Although specific examples of wearable devices have been described, such as the VR device 500 discussed above regarding the example in Figure 2, and in Figure 1A and Figure 1B Other devices illustrated in the examples, but those skilled in the art will understand that... Figure 1A , 1B The description relating to Figure 2 can be applied to other examples of wearable devices. For example, other wearable devices, such as smart glasses, may include sensors to obtain translational head movements. As another example, other wearable devices, such as smartwatches, may include sensors to obtain translational movements. Therefore, the techniques described in this disclosure should not be limited to a particular type of wearable device, but any wearable device can be configured to perform the techniques described in this disclosure.

[0192] Figure 6A and Figure 6B This is a diagram illustrating an example system that can perform various aspects of the techniques described in this disclosure. Figure 6A An example is shown in which the source device 12 also includes a camera 600. The camera 600 can be configured to capture video data and provide the captured raw video data to the content capture device 20. The content capture device 20 can provide the video data to another component of the source device 12 for further processing into portions for viewport partitioning.

[0193] exist Figure 6A In this example, content consumer device 14 also includes wearable device 400. It should be understood that, in various embodiments, wearable device 100 may be included in or externally coupled to content consumer device 14. Wearable device 400 includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewports) and for rendering audio data.

[0194] Figure 6B An example is shown, where Figure 6A The audio renderer 32 shown is replaced by a binaural renderer 42 capable of performing binaural rendering using one or more HRTFs or other functions capable of rendering to the left and right speaker feeds 43. The audio playback system 16C can output the left and right speaker feeds 43 to headphones 44.

[0195] The headphones 44 can be connected via wired connection (such as a standard 3.5mm audio jack, Universal System Bus (USB) connection, optical audio jack, or other forms of wired connection) or wireless connection (such as via Bluetooth). TM The headphones 44 are coupled to the audio playback system 16C via a connection (such as a wireless network connection). The headphones 44 can reconstruct the sound field represented by the audio data 19' based on the left and right speaker feeds 43. The headphones 44 may include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the respective left and right speaker feeds 43.

[0196] Figure 7 This is a block diagram illustrating one or more example components of the source device and content consumer device shown in the example of Figure 1. Figure 7 In one example, device 710 includes a processor 712 (which may be referred to as "one or more processors" or "(multiple) processors"), a graphics processing unit (GPU) 714, system memory 716, a display processor 718, one or more integrated speakers 740, a display 703, a user interface 720, an antenna 721, and a transceiver module 722. In an example where device 712 is a mobile device, the display processor 718 is a mobile display processor (MDP). In some examples, such as where the source device 710 is a mobile device, the processor 712, GPU 714, and display processor 718 may be configured as an integrated circuit (IC).

[0197] For example, an IC can be considered a processing chip within a chip package and can be a system-on-a-chip (SoC). C). In some examples, two of the processor 712, GPU 714, and display processor 718 may be packaged together in the same IC, while the other may be packaged in a different integrated circuit (i.e., a different chip package), or all three may be packaged in different ICs or on the same IC. However, in the example where device 710 is a mobile device, the processor 712, GPU 714, and display processor 718 may all be packaged in different integrated circuits.

[0198] Examples of processor 712, GPU 714, and display processor 718 include, but are not limited to, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Processor 712 may be the central processing unit (CPU) of source device 710. In some examples, GPU 714 may be dedicated hardware including integrated and / or discrete logic circuits that provide GPU 714 with massively parallel processing capabilities suitable for graphics processing. In some instances, GPU 714 may also include general-purpose processing capabilities and may be referred to as a general-purpose GPU (GPGPU) when performing general-purpose processing tasks (i.e., non-graphics-related tasks). Display processor 718 may also be application-specific integrated circuit hardware designed to retrieve image content from system memory 716, composite the image content into image frames, and output the image frames to display 703.

[0199] Processor 712 can execute various types of applications. Examples of these applications include web browsers, email applications, spreadsheets, video games, other applications that generate viewable objects for display, or any of the application types listed in more detail above. System memory 716 can store instructions for executing applications. Executing one of the applications on processor 712 causes processor 712 to generate graphics data for image content to be displayed and audio data 19 to be played (possibly via integrated speaker 740). Processor 712 can transfer the graphics data of the image content to GPU 714 for further processing based on the instructions or commands transferred from processor 712 to GPU 714.

[0200] Processor 712 can communicate with GPU 714 according to a specific application processing interface (API). Examples of such APIs include... of API, Khronos Group or OpenGL and OpenCL TM However, the aspects of this disclosure are not limited to DirectX, OpenGL, or OpenCL APIs, and can be extended to other types of APIs. Furthermore, the techniques described in this disclosure do not require operation based on an API, and the processor 712 and GPU 714 can communicate using any process.

[0201] System memory 716 may be a memory used in device 710. System memory 716 may include one or more computer-readable storage media. Examples of system memory 716 include, but are not limited to, random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other media that can be used to carry or store required program code in the form of instructions and / or data structures and that can be accessed by a computer or processor.

[0202] In some examples, system memory 716 may include instructions that cause processor 712, GPU 714, and / or display processor 718 to perform the functions assigned to processor 712, GPU 714, and / or display processor 718 in this disclosure. Therefore, system memory 716 may be a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors (e.g., processor 712, GPU 714, and / or display processor 718) to perform various functions.

[0203] System memory 716 may include a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that system memory 716 is immovable or that its contents are static. As an example, system memory 716 can be removed from source device 710 and moved to another device. As another example, a memory substantially similar to system memory 716 can be inserted into device 710. In some examples, the non-transitory storage medium may store data that can change over time (e.g., in RAM).

[0204] User interface 720 may represent one or more hardware or virtual (meaning a combination of hardware and software) user interfaces through which a user intersects with device 710. User interface 720 may include physical buttons, switches, toggle keys, lights, or virtual versions thereof. User interface 720 may also include a physical or virtual keyboard, a touch interface—such as a touchscreen, haptic feedback, etc.

[0205] Processor 712 may include one or more hardware units (including so-called "processing cores") configured to perform all or part of the operations discussed above regarding any module, unit, or other functional component of the content creator device and / or content consumer device. Antenna 721 and transceiver module 722 may represent units configured to establish and maintain a connection between content consumer device 12 and content consumer device 14. Antenna 721 and transceiver module 722 may represent one or more receivers and / or one or more transmitters capable of operating according to one or more wireless communication protocols (e.g., fifth-generation (5G) cellular standards), human area network (PAN) protocols (e.g., Bluetooth). TMWireless communication can be performed using (or other open-source, proprietary, or other communication standards). That is, transceiver module 722 can represent a separate transmitter, a separate receiver, both a separate transmitter and a separate receiver, or a combination of transmitter and receiver. Antenna 721 and transceiver 722 can be configured to receive encoded audio data. Similarly, antenna 721 and transceiver 722 can be configured to transmit encoded audio data.

[0206] Figures 8A-8C It is shown Figure 1A and Figure 1B The example shown is a flowchart illustrating the example operations of the stream selection unit when performing various aspects of stream selection technology. First, refer to... Figure 8A For example, the stream selection unit 44 can obtain audio stream 27 from all enabled receivers (this refers to microphones in another way, such as microphone 18), where audio stream 27 may include corresponding audio metadata, such as CLI45A (800). The stream selection unit 44 can perform energy analysis for each audio stream 27 to calculate the corresponding energy map (802).

[0207] The stream selection unit 44 can then iterate over different combinations (804) of the receiver (defined in CM 47) based on the proximity to the audio source 308 (as defined by audio source distances 306A and / or 306B) and the receiver (as defined by proximity distances discussed above). Figure 8A As shown, receivers can be sorted or otherwise associated with different access permissions. Stream selection unit 44 can iterate in the manner described above based on the listener location represented by VLI 45B (which is another way of referring to "virtual location") and the receiver location represented by CLI 45A to identify whether a larger subset or a smaller subset of audio stream 27 is needed (806, 808).

[0208] When a larger subset of audio stream 27 is needed, stream selection unit 44 may add a receiver to audio stream 19', or in other words, add an additional audio stream (810). When a reduced subset of audio stream 27 is needed, stream selection unit 44 may remove a receiver from audio stream 19', or in other words, remove an existing audio stream (812).

[0209] In some examples, the stream selection unit 44 may determine that the current receiver constellation is the optimal set (or, in other words, the existing audio stream 19' will remain the same as the audio stream 19' produced by the selection process described herein) (804). However, when an audio stream is added to or removed from audio stream 19', the stream selection unit 44 may update CM 47 (814) to generate a constellation history (815).

[0210] Furthermore, the stream selection unit 44 can determine whether privacy settings enable or disable the addition of receivers (where privacy settings may refer to digital access permissions, such as passwords, authorization levels or grades, time, etc., restricting access to one or more of the audio streams 27) (816, 818). When privacy settings allow the addition of receivers, the stream selection unit 44 can add the receiver to the updated CM 47 (which refers to adding the audio stream to audio stream 19') (820). When privacy settings disable the addition of receivers, the stream selection unit 44 can remove the receiver from the updated CM 47 (which refers to removing the audio stream from audio stream 19') (822). In this way, the stream selection unit 44 can identify a new set of enabled receivers (824).

[0211] The stream selection unit 44 can iterate in this manner and update various inputs according to any given frequency. For example, the stream selection unit 44 can update privacy settings at the user interface rate (meaning that updates are driven by updates via user interface inputs). As another example, the stream selection unit 44 can update position at the sensor rate (meaning that position changes by the movement of the receiver). The stream selection unit 44 can further update the energy map at the audio frame rate (meaning that the energy map is updated once per frame).

[0212] Next reference Figure 8B For example, the flow selection unit 44 can be configured as described above regarding... Figure 8A The operation is described in a manner different in that the stream selection unit 44 may not be determined based on CM 47 on the energy map. Therefore, the stream selection unit 44 can obtain audio stream 27 from all enabled receivers (this refers to microphones in another way, such as microphone 18), where audio stream 27 may include corresponding audio metadata, such as CLI 45A (840). The stream selection unit 44 can determine whether privacy settings enable or disable the addition of receivers (where privacy settings may refer to digital access permissions, such as passwords, authorization levels or grades, time, etc., restricting access to one or more of the audio streams 27) (842, 844).

[0213] When the privacy setting enables receiver addition, the stream selection unit 44 can add a receiver to the updated CM 47 (which refers to adding an audio stream to audio stream 19') (846). When the privacy setting disables receiver addition, the stream selection unit 44 can remove a receiver from the updated CM 47 (which refers to removing an audio stream from audio stream 19') (848). In this way, the stream selection unit 44 can identify a new set of enabled receivers (850). The stream selection unit 44 can iterate over different combinations of receivers in the CM 47 to determine the constellation history representing audio stream 19' (854).

[0214] The stream selection unit 44 can iterate in this way and update various inputs at any given frequency. For example, the stream selection unit 44 can update privacy settings at the user interface rate (meaning that updates are driven by updates via user interface inputs). As another example, the stream selection unit 44 can update position at the sensor rate (meaning that position changes by the movement of the receiver).

[0215] Next reference Figure 8C For example, the flow selection unit 44 can be configured as described above regarding... Figure 8A The operation is described in a different manner, except that the stream selection unit 44 may not be based on the determination of the enabled receivers of CM 47. Therefore, the stream selection unit 44 can obtain audio stream 27 from all enabled receivers (this refers to microphones in another sense, such as microphone 18), where audio stream 27 may include corresponding audio metadata, such as CLI 45A (860). The stream selection unit 44 can perform energy analysis for each audio stream 27 to calculate the corresponding energy map (862).

[0216] The stream selection unit 44 can then iterate over different combinations of receivers (defined in CM 47) (864) based on proximity to the audio source 308 (as defined by audio source distances 306A and / or 306B) and the receiver (as defined by proximity distances discussed above). Figure 4 As shown in C, receivers can be sorted or otherwise associated with different access permissions. Stream selection unit 44 can iterate in the manner described above based on the listener location represented by VLI 45B (which is also another way of referring to the "virtual location" discussed above) and the receiver location represented by CLI 45A to identify whether a larger subset or a smaller subset of audio stream 27 is needed (866, 868).

[0217] When a larger subset of audio stream 27 is needed, stream selection unit 44 may add a receiver, or in other words, add an additional audio stream to audio stream 19' (870). When a reduced subset of audio stream 27 is needed, stream selection unit 44 may remove a receiver, or in other words, remove an existing audio stream from audio stream 19' (872).

[0218] In some examples, the stream selection unit 44 may determine that the receiver’s current constellation is the optimal set (or, in other words, the existing audio stream 19' will remain the same as the audio stream 19' produced by the selection process described herein) (864). However, when an audio stream is added to or removed from audio stream 19', the stream selection unit 44 may update CM 47 (874) to generate a constellation history (875).

[0219] The stream selection unit 44 can iterate in this manner and update various inputs according to any given frequency. For example, the stream selection unit 44 can update the position at the sensor rate (meaning the position changes by the movement of the receiver). The stream selection unit 44 can further update the energy map at the audio frame rate (meaning the energy map is updated once per frame).

[0220] Figure 9 An example of a privacy-supporting wireless communication system 100 according to aspects of this disclosure is shown. The wireless communication system 100 includes a base station 105, a user interface unit (UE) 115, and a core network 130. In some examples, the wireless communication system 100 may be a Long Term Evolution (LTE) network, an LTE-A Advanced (LTE-A) network, an LTE-A Pro network, or a New Radio (NR) network. In some cases, the wireless communication system 100 may support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low-latency communication, or communication with low-cost and low-complexity devices.

[0221] Base station 105 can wirelessly communicate with UE 115 via one or more base station antennas. Base station 105 described herein may include, or may be referred to by those skilled in the art as, a base station transceiver, radio base station, access point, radio transceiver, NodeB, eNodeB (eNB), next-generation NodeB or giga-NodeB (any of which may be referred to as gNB), home NodeB, home eNodeB, or some other suitable term. Wireless communication system 100 may include different types of base station 105 (e.g., macro or small cell base stations). UE 115 described herein may be able to communicate with various types of base station 105 and network devices including macro eNBs, small cell eNBs, gNBs, relay base stations, etc.

[0222] Each base station 105 may be associated with a specific geographic coverage area 110 in which communication with various UEs 115 is supported. Each base station 105 may provide communication coverage to each geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 may utilize one or more carriers. The communication link 125 shown in the wireless communication system 100 may include uplink transmission from the UE 115 to the base station 105, or downlink transmission from the base station 105 to the UE 115. Downlink transmission may also be referred to as forward link transmission, and uplink transmission may also be referred to as reverse link transmission.

[0223] The geographic coverage area 110 of base station 105 can be divided into sectors that constitute part of the geographic coverage area 110, and each sector can be associated with a cell. For example, each base station 105 can provide communication coverage for macro cells, small cells, hotspots, or other types of cells, or various combinations thereof. In some examples, base station 105 can be mobile, thus providing communication coverage for mobile geographic coverage areas 110. In some examples, different geographic coverage areas 110 associated with different technologies can overlap, and overlapping geographic coverage areas 110 associated with different technologies can be supported by the same base station 105 or different base stations 105. Wireless communication system 100 can include, for example, heterogeneous LTE / LTE-A / LTE-A Pro or NR networks, wherein different types of base stations 105 provide coverage for various geographic coverage areas 110.

[0224] UE 115 may be distributed throughout the wireless communication system 100, and each UE 115 may be stationary or mobile. UE 115 may also be referred to as a mobile device, wireless device, remote device, handheld device, or subscriber device, or some other suitable term, wherein "device" may also be referred to as a unit, station, terminal, or client. UE 115 may also be a personal electronic device, such as a cellular phone, personal digital assistant (PDA), tablet computer, laptop computer, or personal computer. In the examples of this disclosure, UE 115 may be any audio source described in this disclosure, including VR headsets, XR headsets, AR headsets, vehicles, smartphones, microphones, microphone arrays, or any other device, including microphones or capable of transmitting captured and / or synthesized audio streams. In some examples, the synthesized audio stream may be an audio stream stored in memory or previously created or synthesized. In some examples, UE 115 may also refer to a wireless local loop (WLL) station, Internet of Things (IoT) device, Internet of Everything (IoE) device, or MTC device, which may be implemented in various items such as home appliances, vehicles, and meters.

[0225] Some UEs 115, such as MTC or IoT devices, may be low-cost or low-complexity devices and may provide automated communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC can refer to data communication technologies that allow devices to communicate with each other or with base station 105 without human intervention. In some examples, M2M communication or MTC may include communication from devices that exchange and / or use audio metadata indicating privacy restrictions and / or password-based privacy data to switch, mute, and / or zero out various audio streams and / or audio sources, as will be described in more detail below.

[0226] In some cases, UE 115 may also communicate directly with other UE 115 (e.g., using peer-to-peer (P2P) or device-to-device (D2D) protocols). One or more of a group of UE 115s utilizing D2D communication may be within the geographic coverage area 110 of base station 105. Other UE 115s in such a group may be outside the geographic coverage area 110 of base station 105 or may not be able to receive transmissions from base station 105. In some cases, a group of UE 115s communicating via D2D communication may utilize a one-to-many (1:M) system, in which each UE 115 transmits to every other UE 115 in the group. In some cases, base station 105 facilitates the scheduling of resources for D2D communication. In other cases, D2D communication is performed between UE 115s without involving base station 105.

[0227] Base station 105 can communicate with core network 130 and with each other. For example, base station 105 can interface with core network 130 via backhaul link 132 (e.g., via S1, N2, N3 or other interfaces). Base station 105 can communicate with each other directly (e.g., directly between base stations 105) or indirectly (e.g., via core network 130) via backhaul link 134 (e.g., via X2, Xn or other interfaces).

[0228] In some cases, wireless communication system 100 may utilize licensed and unlicensed radio frequency spectrum bands. For example, wireless communication system 100 may employ Licensed Assisted Access (LAA), LTE Unlicensed (LTE-U) radio access technology, or NR technology in unlicensed bands such as the 5 GHz ISM band. When operating in unlicensed radio spectrum bands, wireless devices such as base station 105 and UE 115 may employ a Listen-After-Talk (LBT) procedure to ensure that the channel is clear before transmitting data. In some cases, operation in unlicensed bands may be based on carrier aggregation configurations and component carriers operating in licensed bands (e.g., LAA). Operation in unlicensed spectrum may include downlink transmission, uplink transmission, peer-to-peer transmission, or a combination thereof. Duplexing in unlicensed spectrum may be based on Frequency Division Duplex (FDD), Time Division Duplex (TDD), or a combination of both.

[0229] In this regard, the following terms can be achieved through various aspects of the technology:

[0230] Section 1A. An apparatus configured to play one or more of a plurality of audio streams, the apparatus comprising: a memory configured to store the plurality of audio streams, each of the plurality of audio streams representing a sound field and including one or more substreams; and one or more processors coupled to the memory and configured to: determine, based on the plurality of audio streams, a total number of one or more substreams of all audio streams in the plurality of audio streams; when the total number of one or more substreams is greater than a rendering threshold indicating the total number of substreams supported when a renderer renders the plurality of audio streams to one or more speaker feeds, adapt the plurality of audio streams to reduce the number of one or more substreams and obtain adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold; apply a renderer to the adapted plurality of audio streams to obtain one or more speaker feeds; and output the one or more speaker feeds to one or more speakers.

[0231] Section 2A. In a device pursuant to Section 1A, one or more processors are further configured to avoid removing one or more of the multiple audio streams based on user presets when multiple audio streams are acquired for adaptation.

[0232] Section 3A. A device according to any combination of Sections 1A and 2A, wherein the audio stream includes audio metadata, the audio metadata including start position information identifying the start position of the origin of the audio stream, and wherein one or more processors are configured to adapt multiple audio streams based on the start position information to reduce the total number of one or more sub-streams and obtain multiple adapted audio streams.

[0233] Section 4A. A device according to any combination of Sections 1A-3A, wherein one or more processors are configured to adapt multiple audio streams based on audio data types specified in one or more sub-streams to reduce the total number of one or more sub-streams and obtain multiple adapted audio streams.

[0234] Section 5A. A device pursuant to Section 4A, wherein the type of audio data indicates that the audio data includes stereo reverberation audio data, and wherein one or more processors are configured to perform downgrading on the stereo reverberation audio data to obtain multiple adapted audio streams.

[0235] Section 6A. A device pursuant to Section 4A, wherein the type of audio data indicates that the audio data includes channel-based audio data, and wherein one or more processors are configured to perform downmixing on the channel-based audio data to obtain adapted multiple audio streams.

[0236] Section 7A. A device pursuant to any combination of Sections 1A-6A, wherein one or more processors are configured to adapt multiple audio streams based on privacy settings to remove one or more of the multiple audio streams and obtain the adapted multiple audio streams.

[0237] Section 8A. A device according to any combination of Sections 1A-7A, wherein one or more processors are further configured to apply an overwrite to reduce the number of adapted audio streams such that the total number of substreams is below a rendering threshold and the number of audio streams is reduced.

[0238] Section 9A. A device according to any combination of sections 1A-8A, wherein the adapted plurality of audio streams includes at least one audio stream representing channel-based audio data, wherein the renderer includes a six-degree-of-freedom renderer, wherein one or more processors are further configured to: acquire tracking information representing movement of the device; and, based on the tracking information, modify the six-degree-of-freedom renderer to reflect the movement of the device before applying the six-degree-of-freedom renderer.

[0239] Section 10A. A device according to any combination of sections 1A-8A, wherein the adapted plurality of audio streams include at least one audio stream representing stereo reverberant audio data, wherein the renderer includes a six-degree-of-freedom renderer, wherein one or more processors are further configured to: acquire tracking information representing movement of the device; and, based on the tracking information, modify the six-degree-of-freedom renderer to reflect the movement of the device before applying the six-degree-of-freedom renderer.

[0240] Section 11A. A device according to any combination of Sections 1A-10A, wherein the plurality of audio streams include a first plurality of vehicle-to-any audio streams originating from other vehicles near a device threshold, wherein one or more processors are further configured to: obtain a second plurality of non-vehicle-to-any audio streams representing an additional sound field; render at least one of the second plurality of non-vehicle-to-any audio streams to one or more additional speaker feeds; and output the one or more speaker feeds and the one or more additional speaker feeds to reproduce one or more sound fields and one or more of the additional sound fields.

[0241] Section 12A. A device pursuant to Section 11A, wherein one or more processors are configured to obtain a second or more non-vehicle to any audio stream according to a Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

[0242] Section 13A. Devices pursuant to any combination of Sections 11A and 12A, wherein the first plurality of vehicle-to-any audio streams includes the first plurality of cellular vehicle-to-any audio streams conforming to the Cellular Vehicle-to-Any (C-V2X) protocol.

[0243] Section 14A. Devices pursuant to any combination of sections 1A-13A, wherein the device includes mobile handheld devices.

[0244] Section 15A. Devices according to any combination of sections 1A-13A, wherein the device includes a vehicle audio head unit integrated into the vehicle.

[0245] Section 16A. Devices according to any combination of Sections 1A-15A, wherein at least one of one or more of a plurality of audio streams contains a stereo reverberation coefficient.

[0246] Section 17A. For equipment pursuant to Section 16A, the stereo reverberation factor includes the mixed-order stereo reverberation factor.

[0247] Section 18A. For devices pursuant to Section 16A, the stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

[0248] Section 19A. For devices pursuant to Section 16A, the stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

[0249] Section 20A. A device according to any combination of sections 1A-19A, wherein one or more processors are further configured to: acquire a user audio stream representing the sound field in which the device is located; and output the user audio stream to a second device.

[0250] Section 21A. A method for playing one or more of a plurality of audio streams, the method comprising: storing the plurality of audio streams by one or more processors, each of the plurality of audio streams representing a sound field and including one or more substreams; determining, by one or more processors, based on the plurality of audio streams, a total number of one or more substreams of all audio streams in the plurality of audio streams; adapting the plurality of audio streams to reduce the number of one or more substreams and obtaining adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold by one or more processors; applying a renderer to the adapted plurality of audio streams by one or more processors to obtain one or more speaker feeds; and outputting the one or more speaker feeds to one or more speakers by one or more processors.

[0251] Section 22A. The method according to Section 21A also includes avoiding the removal of one or more of the multiple audio streams based on user presets when obtaining multiple adapted audio streams.

[0252] Section 23A. A method according to any combination of Sections 21A and 22A, wherein the audio stream includes audio metadata, the audio metadata including start position information identifying the start position of the origin of the audio stream, and wherein adapting multiple audio streams includes adapting multiple audio streams based on the start position information to reduce the total number of one or more substreams and obtain multiple adapted audio streams.

[0253] Section 24A. A method according to any combination of sections 21A-23A, wherein adapting multiple audio streams includes adapting multiple audio streams based on audio data types specified in one or more substreams, to reduce the total number of one or more substreams and obtain multiple adapted audio streams.

[0254] Section 25A. The method according to Section 24A, wherein the type of audio data indicates that the audio data includes stereo reverberant audio data, and wherein adapting multiple audio streams includes performing a reduction order on the stereo reverberant audio data to obtain multiple adapted audio streams.

[0255] Section 26A. The method according to Section 24A, wherein the type of audio data indicates that the audio data includes channel-based audio data, and wherein adapting multiple audio streams includes performing downmixing on the channel-based audio data to obtain multiple adapted audio streams.

[0256] Section 27A. A method according to any combination of sections 21A-26A, wherein adapting multiple audio streams includes adapting multiple audio streams based on privacy settings to remove one or more of the multiple audio streams and obtain the adapted multiple audio streams.

[0257] Section 28A. The method according to any combination of sections 21A-27A also includes applying an overwrite to reduce the number of adapted audio streams such that the total number of substreams is below a rendering threshold, and thus obtaining a reduced number of audio streams.

[0258] Section 29A. A method according to any combination of sections 21A-28A, wherein the adapted plurality of audio streams includes at least one audio stream representing channel-based audio data, wherein the renderer includes a six-degree-of-freedom renderer, and wherein the method further includes: acquiring tracking information representing movement of the device; and modifying the six-degree-of-freedom renderer to reflect the movement of the device before applying the tracking information.

[0259] Section 30A. A method according to any combination of sections 21A-28A, wherein the adapted plurality of audio streams includes at least one audio stream representing stereo reverberation audio data, wherein the renderer includes a six-degree-of-freedom renderer, and wherein the method further includes: acquiring tracking information representing the movement of the device; and modifying the six-degree-of-freedom renderer to reflect the movement of the device before applying the tracking information.

[0260] Section 31A. A method according to any combination of Sections 21A-30A, wherein the plurality of audio streams include a first plurality of vehicle-to-any audio streams originating from other vehicles near a device threshold, and wherein the method further comprises: obtaining a second plurality of non-vehicle-to-any audio streams representing an additional sound field; rendering at least one of the second plurality of non-vehicle-to-any audio streams to one or more additional speaker feeds; and outputting one or more speaker feeds and one or more additional speaker feeds to reproduce one or more sound fields and one or more of the additional sound fields.

[0261] Section 32A. The method according to Section 31A, wherein obtaining a second plurality of non-vehicle to any audio stream includes obtaining a second plurality of non-vehicle to any audio stream according to a Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

[0262] Section 33A. The method according to any combination of Sections 31A and 32A, wherein the first plurality of vehicle-to-any audio streams includes the first plurality of cellular vehicle-to-any audio streams conforming to the Cellular Vehicle-to-Any (C-V2X) protocol.

[0263] Section 34A. A method according to any combination of sections 21A-33A, wherein the method is performed by a mobile handheld device.

[0264] Section 35A. A method according to any combination of sections 21A-33A, wherein the method is performed by a vehicle audio head unit integrated into the vehicle.

[0265] Section 36A. A method according to any combination of Sections 21A-35A, wherein at least one of one or more of a plurality of audio streams contains a stereo reverberation coefficient.

[0266] Section 37A. According to the method of Section 36A, the stereo reverberation coefficient includes the mixed-order stereo reverberation coefficient.

[0267] Section 38A. According to the method of Section 36A, the stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

[0268] Section 39A. According to the method of Section 36A, the stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

[0269] Section 40A. The method according to any combination of sections 21A-39A further includes: acquiring a user audio stream representing the sound field in which the device is located; and outputting the user audio stream to a second device.

[0270] Section 41A. An apparatus configured to play one or more of a plurality of audio streams, the apparatus comprising: means for storing the plurality of audio streams, each of the plurality of audio streams representing a sound field and including one or more substreams; means for determining, based on the plurality of audio streams, the total number of one or more substreams of all audio streams in the plurality of audio streams; means for adapting the plurality of audio streams to reduce the number of one or more substreams and obtaining adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold when the total number of one or more substreams is greater than a rendering threshold indicating the total number of substreams supported when a renderer renders the plurality of audio streams to one or more speaker feeds; means for applying a renderer to the adapted plurality of audio streams to obtain one or more speaker feeds; and means for outputting one or more speaker feeds to one or more speakers.

[0271] Section 42A. The device according to Section 41A also includes a component for avoiding the removal of one or more of the multiple audio streams based on a user preset when multiple adapted audio streams are obtained.

[0272] Section 43A. A device pursuant to any combination of Sections 41A and 42A, wherein the audio stream includes audio metadata, the audio metadata including start position information identifying the start position of the origin of the audio stream, and wherein the components for adapting multiple audio streams include components for adapting multiple audio streams based on the start position information to reduce the total number of one or more sub-streams and to obtain multiple adapted audio streams.

[0273] Section 44A. A device pursuant to any combination of sections 41A-43A, wherein the components for adapting multiple audio streams include components for adapting multiple audio streams based on audio data types specified in one or more sub-streams, so as to reduce the total number of one or more sub-streams and obtain multiple adapted audio streams.

[0274] Section 45A. A device according to Section 44A, wherein the type of audio data indicates that the audio data includes stereo reverberant audio data, and wherein the components for adapting multiple audio streams include components for performing downgrading on the stereo reverberant audio data to obtain the adapted multiple audio streams.

[0275] Section 46A. The apparatus according to Section 44A, wherein the type of audio data indicates that the audio data includes channel-based audio data, and wherein the components for adapting multiple audio streams include components for performing down-mixing on the channel-based audio data to obtain the adapted multiple audio streams.

[0276] Section 47A. A device pursuant to any combination of sections 41A-46A, wherein the components for adapting multiple audio streams include components for adapting multiple audio streams based on privacy settings to remove one or more of the multiple audio streams and obtain the adapted multiple audio streams.

[0277] Section 48A. A device according to any combination of sections 41A-47A also includes a component for applying an overwrite to reduce the number of adapted audio streams such that the total number of substreams is below a rendering threshold, and to obtain the reduced number of audio streams.

[0278] Section 49A. A device according to any combination of sections 41A-48A, wherein the adapted plurality of audio streams includes at least one audio stream representing channel-based audio data, wherein the renderer includes a six-degree-of-freedom renderer, and wherein the device further includes: means for acquiring tracking information representing movement of the device; and means for modifying the six-degree-of-freedom renderer to reflect the movement of the device based on the tracking information before applying the six-degree-of-freedom renderer.

[0279] Section 50A. A device according to any combination of sections 41A-48A, wherein the adapted plurality of audio streams include at least one audio stream representing stereo reverberant audio data, wherein the renderer includes a six-degree-of-freedom renderer, and wherein the device further includes: components for acquiring tracking information representing movement of the device; and components for modifying the six-degree-of-freedom renderer to reflect the movement of the device based on the tracking information before applying the six-degree-of-freedom renderer.

[0280] Section 51A. An apparatus according to any combination of sections 41A-50A, wherein the plurality of audio streams include a first plurality of vehicle-to-any audio streams originating from other vehicles near a device threshold, and wherein the apparatus further includes: means for obtaining a second plurality of non-vehicle-to-any audio streams representing an additional sound field; means for rendering at least one of the second plurality of non-vehicle-to-any audio streams to one or more additional speaker feeds; and means for outputting one or more speaker feeds and one or more additional speaker feeds to reproduce one or more sound fields and one or more of the additional sound fields.

[0281] Section 52A. A device pursuant to Section 51A, wherein the components for obtaining a second plurality of non-vehicle-to-any audio streams include components for obtaining a second plurality of non-vehicle-to-any audio streams according to a Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

[0282] Section 53A. Devices pursuant to any combination of Sections 51A and 52A, wherein the first plurality of vehicle-to-any audio streams includes the first plurality of cellular vehicle-to-any audio streams conforming to the Cellular Vehicle-to-Any (C-V2X) protocol.

[0283] Section 54A. Devices pursuant to any combination of sections 41A-53A, wherein the device includes mobile handheld devices.

[0284] Section 55A. Devices pursuant to any combination of sections 41A-53A, wherein the device includes a vehicle audio head unit integrated into the vehicle.

[0285] Section 56A. A device according to any combination of sections 41A-55A, wherein at least one of one or more of a plurality of audio streams contains a stereo reverberation coefficient.

[0286] Section 57A. For equipment pursuant to Section 56A, the stereo reverberation factor includes the mixed-order stereo reverberation factor.

[0287] Section 58A. For devices pursuant to Section 56A, the stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

[0288] Section 59A. For devices pursuant to Section 56A, the stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

[0289] Section 60A. A device according to any combination of sections 41A-59A further includes: a component for acquiring a user audio stream representing the sound field in which the device is located; and a component for outputting the user audio stream to a second device.

[0290] Section 61A. A non-transitory computer-readable storage medium having instructions thereon, which, when executed, cause one or more processors to: store a plurality of audio streams, each of the plurality of audio streams representing a sound field and including one or more substreams; determine, based on the plurality of audio streams, the total number of one or more substreams of all audio streams in the plurality of audio streams; when the total number of one or more substreams is greater than a rendering threshold indicating the total number of substreams supported when a renderer renders the plurality of audio streams to one or more speaker feeds, adapt the plurality of audio streams to reduce the number of one or more substreams and obtain adapted plurality of audio streams including a reduced total number of one or more substreams equal to or less than the rendering threshold; apply a renderer to the adapted plurality of audio streams to obtain one or more speaker feeds; and output the one or more speaker feeds to one or more speakers.

[0291] Section 1B. A device configured to play one or more of a plurality of audio streams, the device comprising: a memory configured to store the plurality of audio streams and corresponding audio metadata, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; and one or more processors coupled to the memory and configured to: determine an arrival direction of each of the plurality of audio streams based on current coordinates of the device relative to the start coordinates corresponding to the plurality of audio streams; render each of the plurality of audio streams to one or more speaker feeds based on each arrival direction, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each arrival direction; and output the one or more speaker feeds to reproduce the one or more sound fields represented by the plurality of audio streams.

[0292] Section 2B. In accordance with Section 1B, the audio metadata also includes privacy restrictions on one or more of a plurality of audio streams, wherein one or more processors are further configured to determine one or more of the plurality of audio streams based on the privacy restrictions.

[0293] Section 3B. In accordance with the device described in Section 2B, one or more processors are further configured to: determine one or more restricted audio streams from audio metadata based on privacy constraints; and not render speaker feeds from the one or more restricted audio streams.

[0294] Section 4B. Devices pursuant to any combination of Sections 2B and 3B, wherein privacy restrictions indicate whether one or more of a plurality of audio streams are restricted or unrestricted.

[0295] Section 5B. A device according to any combination of sections 1B-4B, wherein the device communicates with a vehicle, the vehicle including one or more speakers, the one or more speakers reproducing one or more sound fields represented by one or more of a plurality of audio streams based on speaker feeds.

[0296] Section 6B. The device according to Section 5B, wherein the vehicle includes a first vehicle, and wherein one or more of the plurality of audio streams includes at least one audio stream specifying the speech of an occupant from a second vehicle.

[0297] Section 7B. Equipment according to Section 6B, wherein the first vehicle includes an automatic vehicle that automatically adjusts its operation based on spoken commands.

[0298] Section 8B. Equipment pursuant to any combination of Sections 6B and 7B, wherein the second vehicle includes one of a bicycle, a motorcycle, and a scooter.

[0299] Section 9B. The equipment according to Section 7B, wherein the autonomous vehicle includes at least one speaker configured to output audible commands to a second vehicle.

[0300] Section 10B. A device according to any combination of sections 1B-9B, wherein the plurality of audio streams include a first plurality of vehicle-to-any audio streams originating from other vehicles near a device threshold, wherein one or more processors are further configured to: obtain a second plurality of non-vehicle-to-any audio streams representing an additional sound field; render at least one of the second plurality of non-vehicle-to-any audio streams to one or more additional speaker feeds; and output one or more speaker feeds and one or more additional speaker feeds to reproduce one or more sound fields and one or more of the additional sound fields.

[0301] Section 11B. A device pursuant to Section 10B, wherein one or more processors are configured to obtain a second or more non-vehicle-to-any audio stream according to a Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

[0302] Section 12B. Devices pursuant to any combination of Sections 10B and 11B, wherein the first plurality of vehicle-to-any audio streams includes the first plurality of cellular vehicle-to-any audio streams conforming to the Cellular Vehicle-to-Any (C-V2X) protocol.

[0303] Section 13B. Devices pursuant to any combination of sections 1B-12B, wherein the device includes mobile handheld devices.

[0304] Section 14B. Devices pursuant to any combination of sections 1B-6B, wherein the device includes a vehicle audio head unit integrated into the vehicle.

[0305] Section 15B. Devices according to any combination of Sections 1B-14B, wherein at least one of one or more of a plurality of audio streams contains a stereo reverberation coefficient.

[0306] Section 16B. For equipment pursuant to Section 15B, the stereo reverberation factor includes the mixed-order stereo reverberation factor.

[0307] Section 17B. For devices pursuant to Section 15B, the stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

[0308] Section 18B. For devices pursuant to Section 15B, the stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

[0309] Section 19B. A device according to any combination of sections 1B-18B, wherein one or more processors are further configured to: acquire a user audio stream representing the sound field in which the device is located; and output the user audio stream to a second device.

[0310] Section 20B. A device pursuant to Section 19B, wherein the device includes a first device communicating with a first vehicle, wherein the second device includes a second device communicating with a second vehicle, and wherein the user audio stream includes spoken words from the user of the first device.

[0311] Section 21B. For devices pursuant to Section 20B, the statements therein represent commands specifying the route of action for a user operating the first device.

[0312] Section 22B. A device pursuant to any combination of sections 1B-21B, wherein the capture location of at least one of one or more of a plurality of audio streams indicates that at least one of one or more of the plurality of audio streams will be located in a passenger seat of a vehicle in communication with the device.

[0313] Section 23B. Devices pursuant to any combination of sections 1B-22B also include receivers configured to receive multiple audio streams.

[0314] Section 24B. Devices pursuant to Section 23B, wherein the receiver includes a receiver configured to receive multiple audio streams in accordance with the fifth-generation (5G) cellular standard.

[0315] Section 25B. Devices pursuant to Section 23B, wherein the receiver includes a receiver configured to receive multiple audio streams in accordance with personal area network standards.

[0316] Section 26B. An apparatus according to any combination of sections 1B-25B, wherein the apparatus comprises one or more loudspeakers configured to reproduce one or more of one or more sound fields represented by one or more of a plurality of audio streams based on loudspeaker feeds.

[0317] Section 27B. A method for playing one or more of a plurality of audio streams, the device comprising: storing a plurality of audio streams and corresponding audio metadata in a memory, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; determining, by one or more processors and based on current coordinates of the device relative to the start coordinates corresponding to the one or more of the plurality of audio streams, an arrival direction for each of the one or more of the plurality of audio streams; rendering, by one or more processors and based on each arrival direction, each of the one or more of the plurality of audio streams to one or more speaker feeds, the one or more speaker feeds spatializing the one or more of the plurality of audio streams to make them appear to arrive from each arrival direction; and outputting the one or more speaker feeds by one or more processors to reproduce the one or more sound fields represented by the one or more of the plurality of audio streams.

[0318] Section 28B. The method according to Section 27B, wherein the audio metadata also includes privacy restrictions on one or more of a plurality of audio streams, and wherein the device also includes determining one or more of the plurality of audio streams based on the privacy restrictions.

[0319] Section 29B. The method according to Section 28B also includes: determining one or more restricted audio streams from audio metadata based on privacy restrictions; and not rendering speaker feeds from one or more restricted audio streams.

[0320] Section 30B. A method pursuant to any combination of Sections 28B and 29B, wherein privacy restrictions indicate whether one or more of a plurality of audio streams are restricted or unrestricted.

[0321] Section 31B. A method according to any combination of sections 27B-30B, wherein the method is performed by a device and wherein the device communicates with a vehicle, the vehicle including one or more loudspeakers, the one or more loudspeakers reproducing one or more sound fields represented by one or more of a plurality of audio streams based on loudspeaker feeds.

[0322] Section 32B. The method according to Section 31B, wherein the vehicle includes a first vehicle, and wherein one or more of the plurality of audio streams includes at least one audio stream specifying the speech of an occupant from a second vehicle.

[0323] Article 33B. The method according to Article 32B, wherein the first vehicle includes an automatic vehicle that automatically adjusts its operation according to the spoken words.

[0324] Section 34B. The method according to any combination of Sections 32B and 33B, wherein the second vehicle includes one of a bicycle, a motorcycle, and a scooter.

[0325] Section 35B. The method according to Section 33B, wherein the autonomous vehicle includes at least one speaker configured to output audible commands to a second vehicle.

[0326] Section 36B. A method according to any combination of Sections 27B-35B, wherein the plurality of audio streams includes a first plurality of vehicle-to-any audio streams originating from other vehicles near a device threshold, and wherein the method further comprises: obtaining a second plurality of non-vehicle-to-any audio streams representing an additional sound field; rendering at least one of the second plurality of non-vehicle-to-any audio streams to one or more additional speaker feeds; and outputting one or more speaker feeds and one or more additional speaker feeds to reproduce one or more sound fields and one or more of the additional sound fields.

[0327] Section 37B. The method according to Section 36B, wherein obtaining a second plurality of non-vehicle to any audio stream includes obtaining a second plurality of non-vehicle to any audio stream according to the Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

[0328] Section 38B. The method according to any combination of Sections 36B and 37B, wherein the first plurality of vehicle-to-any audio streams includes the first plurality of cellular vehicle-to-any audio streams conforming to the Cellular Vehicle-to-Any (C-V2X) protocol.

[0329] Section 39B. A method according to any combination of sections 27B-38B, wherein the device includes a mobile handheld device.

[0330] Section 40B. A method according to any combination of sections 27B-32B, wherein the device includes a vehicle audio head unit integrated into the vehicle.

[0331] Section 41B. A method according to any combination of Sections 27B-40B, wherein at least one of one or more of a plurality of audio streams contains a stereo reverberation coefficient.

[0332] Section 42B. The method according to Section 41B, wherein the stereo reverberation coefficient includes the mixed-order stereo reverberation coefficient.

[0333] Section 43B. According to the method of Section 41B, the stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

[0334] Section 44B. According to the method of Section 41B, the stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

[0335] Section 45B. The method according to any combination of sections 27B-44B further includes: acquiring a user audio stream representing the sound field in which the device is located; and outputting the user audio stream to a second device.

[0336] Section 46B. The method according to Section 45B, wherein the device includes a first device communicating with a first vehicle, wherein the second device includes a second device communicating with a second vehicle, and wherein the user audio stream includes spoken words of the user of the first device.

[0337] Section 47B. The method according to Section 46B, wherein the words spoken indicate a command specifying the route of action of the user when operating the first device.

[0338] Section 48B. A method according to any combination of sections 27B-47B, wherein the capture location of at least one of one or more of a plurality of audio streams indicates that at least one of one or more of the plurality of audio streams will be located in a passenger seat of a vehicle communicating with the device.

[0339] Section 49B. The method according to any combination of sections 27B-48B also includes receiving multiple audio streams.

[0340] Section 50B. The method according to Section 49B, wherein receiving multiple audio streams includes receiving multiple audio streams according to fifth-generation (5G) cellular standards.

[0341] Section 51B. The method according to Section 49B, wherein receiving multiple audio streams includes receiving multiple audio streams according to personal area network standards.

[0342] Section 52B. The method according to any combination of sections 27B-51B also includes reproducing one or more of one or more sound fields represented by one or more of a plurality of audio streams based on loudspeaker feed.

[0343] Section 53B. A device configured to play one or more of a plurality of audio streams, the device comprising: means for storing the plurality of audio streams and corresponding audio metadata, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; and means for determining an arrival direction of each of the plurality of audio streams based on current coordinates of the device relative to the start coordinates corresponding to the plurality of audio streams; means for rendering each of the plurality of audio streams to one or more speaker feeds based on each arrival direction, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each arrival direction; and means for outputting the one or more speaker feeds to reproduce one or more sound fields represented by the one or more of the plurality of audio streams.

[0344] Section 54B. A device pursuant to Section 53B, wherein the audio metadata further includes privacy restrictions on one or more of a plurality of audio streams, and wherein the device further includes components for determining one or more of the plurality of audio streams based on the privacy restrictions.

[0345] Section 55B. The device pursuant to Section 54B further includes: components for determining one or more restricted audio streams from audio metadata based on privacy restrictions; and components for not rendering speaker feeds from one or more restricted audio streams.

[0346] Section 56B. Devices pursuant to any combination of Sections 54B and 55B, wherein privacy restrictions indicate whether one or more of a plurality of audio streams are restricted or unrestricted.

[0347] Section 57B. A device according to any combination of sections 53B-56B, wherein the device communicates with a vehicle, the vehicle including one or more loudspeakers, the one or more loudspeakers reproducing one or more sound fields represented by one or more of a plurality of audio streams based on loudspeaker feeds.

[0348] Section 58B. A device pursuant to Section 57B, wherein the vehicle includes a first vehicle, and wherein one or more of a plurality of audio streams include at least one audio stream specifying speech from an occupant of a second vehicle.

[0349] Section 59B. Equipment according to Section 58B, wherein the first vehicle includes an automatic vehicle that automatically adjusts its operation based on spoken commands.

[0350] Section 60B. Equipment pursuant to any combination of Sections 58B and 59B, wherein the second vehicle includes one of a bicycle, a motorcycle, and a scooter.

[0351] Section 61B. Equipment according to Section 59B, wherein the autonomous vehicle includes at least one speaker configured to output audible commands to a second vehicle.

[0352] Section 62B. An apparatus according to any combination of sections 53B-61B, wherein the plurality of audio streams include a first plurality of vehicle-to-any audio streams originating from other vehicles near a device threshold, and wherein the apparatus further includes: means for obtaining a second plurality of non-vehicle-to-any audio streams representing an additional sound field; means for rendering at least one of the second plurality of non-vehicle-to-any audio streams to one or more additional speaker feeds; and means for outputting one or more speaker feeds and one or more additional speaker feeds to reproduce one or more sound fields and one or more of the additional sound fields.

[0353] Section 63B. A device pursuant to Section 62B, wherein the components for obtaining a second plurality of non-vehicle-to-any audio streams include components for obtaining a second plurality of non-vehicle-to-any audio streams according to a Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

[0354] Section 64B. Devices pursuant to any combination of Sections 62B and 63B, wherein the first plurality of vehicle-to-any audio streams includes the first plurality of cellular vehicle-to-any audio streams conforming to the Cellular Vehicle-to-Any (C-V2X) protocol.

[0355] Section 65B. Equipment pursuant to any combination of sections 53B-64B, wherein the equipment includes mobile handheld devices.

[0356] Section 66B. Devices pursuant to any combination of sections 53B-58B, wherein the device includes a vehicle audio head unit integrated into the vehicle.

[0357] Section 67B. A device according to any combination of sections 53B-66B, wherein at least one of one or more of a plurality of audio streams contains a stereo reverberation coefficient.

[0358] Section 68B. For equipment pursuant to Section 67B, the stereo reverberation factor includes the mixed-order stereo reverberation factor.

[0359] Section 69B. For devices pursuant to Section 67B, the stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

[0360] Section 70B. For devices pursuant to Section 67B, the stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

[0361] Section 71B. An apparatus pursuant to any combination of sections 53B-70B further includes: a component for acquiring a user audio stream representing the sound field in which the apparatus is located; and a component for outputting the user audio stream to a second apparatus.

[0362] Section 72B. A device pursuant to Section 71B, wherein the device includes a first device communicating with a first vehicle, wherein the second device includes a second device communicating with a second vehicle, and wherein the user audio stream includes spoken words from a user of the first device.

[0363] Section 73B. The device according to Section 72B, wherein the words stated indicate a command specifying the route of action of the user when operating the first device.

[0364] Section 74B. A device pursuant to any combination of sections 53B-73B, wherein the capture location of at least one of one or more of a plurality of audio streams indicates that at least one of one or more of the plurality of audio streams will be located in a passenger seat of a vehicle in communication with the device.

[0365] Section 75B. Devices pursuant to any combination of sections 53B-74B also include the ability to receive multiple audio streams.

[0366] Section 76B. Equipment pursuant to Section 75B, wherein components for receiving multiple audio streams include components for receiving multiple audio streams according to fifth-generation (5G) cellular standards.

[0367] Section 77B. A device pursuant to Section 75B, wherein components for receiving multiple audio streams include components for receiving multiple audio streams in accordance with personal area network standards.

[0368] Section 78B. An apparatus pursuant to any combination of sections 53B-77B further includes components for reproducing one or more of one or more sound fields represented by one or more of a plurality of audio streams, based on a loudspeaker feed.

[0369] Section 79B. A non-transitory computer-readable storage medium having instructions thereon, which, when executed, cause one or more processors to: store a plurality of audio streams and corresponding audio metadata, each of the plurality of audio streams representing a sound field, and the audio metadata including start coordinates of a corresponding origin for each of the plurality of audio streams; and determine an arrival direction for each of the plurality of audio streams based on the current coordinates of the device relative to the start coordinates corresponding to the plurality of audio streams; render each of the plurality of audio streams to one or more speaker feeds based on each arrival direction, the one or more speaker feeds spatializing the plurality of audio streams to make them appear to arrive from each arrival direction; and output the one or more speaker feeds to reproduce the one or more sound fields represented by the plurality of audio streams.

[0370] It should be recognized that, depending on the examples, certain actions or events of any of the techniques described herein may be performed in a different order, and may be added, combined, or excluded entirely (e.g., not all described actions or events are necessary for the practice of the technique). Furthermore, in some examples, actions or events may be performed in parallel, for example through multithreading, interrupt handling, or multiple processors, rather than sequentially.

[0371] In some examples, the VR device (or streaming device) may use a network interface coupled to the VR / streaming device's memory to transmit exchange messages to an external device, where the exchange messages are associated with multiple available representations of the sound field. In some examples, the VR device may use an antenna coupled to the network interface to receive wireless signals including data packets, audio packets, video packets, or transport protocol data associated with multiple available representations of the sound field. In some examples, one or more microphone arrays may capture the sound field.

[0372] In some examples, multiple available representations of a sound field stored in a storage device may include multiple object-based representations of the sound field, higher-order stereo reverberation representations of the sound field, mixed-order stereo reverberation representations of the sound field, combinations of object-based representations of the sound field and higher-order stereo reverberation representations of the sound field, combinations of object-based representations of the sound field and mixed-order stereo reverberation representations of the sound field, or combinations of mixed-order representations of the sound field and higher-order stereo reverberation representations of the sound field.

[0373] In some examples, one or more of the multiple available representations of the sound field may include at least one high-resolution region and at least one low-resolution region, wherein the selected representation based on the steering angle provides higher spatial accuracy relative to at least one high-resolution region and lower spatial accuracy relative to the low-resolution region.

[0374] In one or more examples, the functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored or transmitted thereon as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium (e.g., a data storage medium), or a communication medium that includes any medium that facilitates the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this way, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.

[0375] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are sent from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but rather refer to non-temporary tangible storage media. Disks and optical discs as used herein include optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0376] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the above-described structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be fully implemented in one or more circuit or logic elements.

[0377] The techniques disclosed herein can be implemented in a variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or IC sets, such as chip sets. Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in a codec hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) along with suitable software and / or firmware.

[0378] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A device configured to play one or more of a plurality of audio streams, the device comprising: The memory is configured to store a plurality of audio streams, each of which represents a sound field and includes one or more sub-streams obtained by an audio decoder; as well as One or more processors are coupled to the memory and configured to: Based on the multiple audio streams, determine the total number of one or more sub-streams of all audio streams in the multiple audio streams; When the total number of the one or more substreams is greater than a rendering threshold indicating the total number of substreams supported by the renderer when rendering the multiple audio streams to one or more speaker feeds, the multiple audio streams are adapted to reduce the number of the one or more substreams and obtain adapted multiple audio streams including the reduced total number of one or more substreams equal to or less than the rendering threshold. The renderer is applied to the adapted multiple audio streams to obtain the one or more speaker feeds; as well as The output of the one or more speakers is fed to the one or more speakers.

2. The device of claim 1, wherein the one or more processors are further configured to avoid removing one or more of the plurality of audio streams based on a user preset when obtaining the adapted plurality of audio streams.

3. The device according to claim 1, in, The audio stream includes audio metadata, which includes start position information identifying the starting position of the audio stream's origin, and... The one or more processors are configured to adapt the plurality of audio streams based on the starting position information to reduce the total number of the one or more sub-streams and obtain the adapted plurality of audio streams.

4. The device of claim 1, wherein the one or more processors are configured to adapt the plurality of audio streams to reduce the total number of the one or more sub-streams and obtain the adapted plurality of audio streams based on the type of audio data specified in the one or more sub-streams.

5. The device according to claim 4, in, The type of the audio data indicates that the audio data includes stereo reverberation audio data, and The one or more processors are configured to perform down-order reduction on the stereo reverb audio data to obtain the adapted multiple audio streams.

6. The device according to claim 4, in, The type of the audio data indicates that the audio data includes channel-based audio data, and The one or more processors are configured to perform downmixing on the channel-based audio data to obtain the adapted multiple audio streams.

7. The device of claim 1, wherein the one or more processors are configured to adapt the plurality of audio streams based on privacy settings to remove one or more of the plurality of audio streams and obtain the adapted plurality of audio streams.

8. The device of claim 1, wherein the one or more processors are further configured to apply an overwrite to reduce the adapted plurality of audio streams such that the total number of substreams is below the rendering threshold and a reduced plurality of audio streams are obtained.

9. The device according to claim 1, in, The adapted multiple audio streams include at least one audio stream representing channel-based audio data. The renderer includes a six-degree-of-freedom renderer, and Wherein, the one or more processors are further configured to: Acquire tracking information representing the movement of the device; and Based on the tracking information, the six-degrees-of-freedom renderer is modified to reflect the movement of the device before it is applied.

10. The device according to claim 1, in, The adapted multiple audio streams include at least one audio stream representing stereo reverberation audio data. The renderer includes a six-degree-of-freedom renderer, and Wherein, the one or more processors are further configured to: Acquire tracking information representing the movement of the device; and Based on the tracking information, the six-degrees-of-freedom renderer is modified to reflect the movement of the device before it is applied.

11. The device according to claim 1, in, The plurality of audio streams includes a first plurality of vehicles originating from other vehicles within the vicinity of the device threshold to any audio stream, and Wherein, the one or more processors are further configured to: Obtain a second, multiple non-vehicle-based audio stream representing an additional sound field; Render at least one of the second plurality of non-vehicle to any audio stream to one or more additional speaker feeds; and Output the one or more loudspeaker feeds and the one or more additional loudspeaker feeds to reproduce one or more of the one or more sound fields and the additional sound fields.

12. The device of claim 11, wherein the one or more processors are configured to obtain the second plurality of non-vehicle to any audio stream according to the Dynamic Adaptive Streaming (DASH) protocol based on Hypertext Transfer Protocol (HTTP).

13. The device of claim 11, wherein the first plurality of vehicle-to-any-audio streams comprises a first plurality of cellular vehicle-to-any-audio streams conforming to a cellular vehicle-to-any-C-V2X protocol.

14. The device of claim 1, wherein the device includes a mobile handheld device.

15. The device of claim 1, wherein the device includes a vehicle audio head unit integrated into a vehicle.

16. The device of claim 1, wherein at least one of one or more of the plurality of audio streams comprises a stereo reverb coefficient.

17. The device of claim 16, wherein the stereo reverberation coefficient includes a mixed-order stereo reverberation coefficient.

18. The device according to claim 16, wherein, The stereo reverberation coefficients include first-order stereo reverberation coefficients associated with spherical basis functions of order 1 or less.

19. The device according to claim 16, wherein, The stereo reverberation coefficients include stereo reverberation coefficients associated with spherical basis functions of order greater than 1.

20. The device of claim 1, wherein the one or more processors are further configured to: Obtain the user audio stream representing the sound field where the device is located; The user's audio stream is output to a second device.

21. A method for playing one or more of a plurality of audio streams, the method comprising: Multiple audio streams are stored by one or more processors, each of the multiple audio streams representing a sound field, and including one or more sub-streams obtained by an audio decoder; The total number of one or more sub-streams of all audio streams in the plurality of audio streams is determined by the one or more processors and based on the plurality of audio streams; The one or more processors adapt the multiple audio streams to reduce the number of the one or more substreams and obtain adapted multiple audio streams including a reduced total number of the one or more substreams equal to or less than the rendering threshold when the total number of the one or more substreams is greater than a rendering threshold. The renderer is applied by the one or more processors to the adapted multiple audio streams to obtain the one or more speaker feeds; as well as The one or more processors feed the one or more speakers to the one or more speakers.

22. The method of claim 21, further comprising avoiding the removal of one or more of the plurality of audio streams based on a user preset when obtaining the adapted plurality of audio streams.

23. The method according to claim 21, in, The audio stream includes audio metadata, which includes start position information identifying the starting position of the audio stream's origin, and... The adaptation of the multiple audio streams includes adapting the multiple audio streams based on the starting position information to reduce the total number of the one or more sub-streams and obtain the adapted multiple audio streams.

24. The method of claim 21, wherein adapting the plurality of audio streams includes adapting the plurality of audio streams based on the type of audio data specified in the one or more sub-streams to reduce the total number of the one or more sub-streams and obtain the adapted plurality of audio streams.

25. The method according to claim 24, in, The type of the audio data indicates that the audio data includes stereo reverberation audio data, and The adaptation of the multiple audio streams includes performing a reduction order on the stereo reverberation audio data to obtain the adapted multiple audio streams.

26. The method according to claim 24, in, The type of the audio data indicates that the audio data includes channel-based audio data, and Adapting the multiple audio streams includes performing downmixing on the channel-based audio data to obtain the adapted multiple audio streams.

27. The method of claim 21, wherein adapting the plurality of audio streams includes adapting the plurality of audio streams based on privacy settings to remove one or more of the plurality of audio streams and obtain the adapted plurality of audio streams.

28. The method of claim 21, further comprising applying an overwrite to reduce the number of adapted audio streams such that the total number of substreams is below the rendering threshold, and obtaining a reduced number of audio streams.

29. A device configured to play one or more of a plurality of audio streams, the device comprising: A component for storing multiple audio streams, each of which represents a sound field and includes one or more sub-streams obtained by an audio decoder; A component for determining the total number of one or more sub-streams of all audio streams in the plurality of audio streams, based on the plurality of audio streams; A component for adapting the multiple audio streams to reduce the number of the one or more substreams and obtaining adapted multiple audio streams including a reduced total number of the one or more substreams equal to or less than the rendering threshold when the total number of the one or more substreams is greater than a rendering threshold. Components for applying the renderer to the adapted multiple audio streams to obtain the feed from the one or more speakers; as well as A component for feeding the output of the one or more speakers to the one or more speakers.

30. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: Store multiple audio streams, each of which represents a sound field and includes one or more sub-streams obtained by an audio decoder; Based on the multiple audio streams, determine the total number of one or more sub-streams of all audio streams in the multiple audio streams; When the total number of the one or more substreams is greater than a rendering threshold indicating the total number of substreams supported by the renderer when rendering the multiple audio streams to one or more speaker feeds, the multiple audio streams are adapted to reduce the number of the one or more substreams and to obtain adapted multiple audio streams including the reduced total number of the one or more substreams equal to or less than the rendering threshold. The renderer is applied to the adapted multiple audio streams to obtain the one or more speaker feeds; as well as The output of the one or more speakers is fed to the one or more speakers.

31. A computer program product comprising computer-readable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 21 to 28.

Citation Information

Patent Citations

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US10405126B2

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US20190007781A1

  • Encoding and rendering of object based audio indicative of game audio content

    CN104520924A

  • Progressive Streaming of Spatial Audio

    US20180315437A1

  • Rendering for computer-mediated reality systems

    US20190116440A1