Apparatus, method and medium for processing an audio stream

By employing authorization level selection technology based on audio streams and surround sound coefficient representation of the sound field in a computer-mediated real-world system, the problem of protecting sensitive information in audio scenes in existing technologies is solved, enabling adaptive processing and rendering of audio data and improving the security and immersion of the audio experience.

CN114041113BActive Publication Date: 2026-01-06QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080047096.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-01
Filing Date
2020-07-02
Publication Date
2026-01-06
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

Existing computer-mediated reality systems struggle to effectively protect sensitive information when providing immersive audio experiences, especially in areas of the audio scene containing sensitive information, leading to an increased risk of information leakage.

Method used

By using audio stream-based authorization level selection technology, a subset of the audio stream can be selectively excluded or adjusted to protect sensitive information from being rendered. Surround sound coefficients are used to represent the sound field, and psychoacoustic coding technology is combined to achieve adaptive processing and rendering of audio data.

Benefits of technology

It effectively protects sensitive information, ensuring that sensitive information is not rendered in audio scenarios, improving the security and immersion of the audio experience, and adapting to user mobility and device changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114041113B_ABST
    Figure CN114041113B_ABST
Patent Text Reader

Abstract

Example devices and methods are disclosed. An example device includes a memory configured to store a plurality of audio streams and an associated authorization level for each of the plurality of audio streams. The device also includes one or more processors implemented in circuitry and communicatively coupled to the memory. The one or more processors are configured to select a subset of the plurality of audio streams based on the associated authorization levels, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefits of U.S. Application No. 16 / 918,386, filed July 1, 2020, and U.S. Provisional Application No. 62 / 870,591, filed July 3, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the processing of media data such as audio data. Background Technology

[0003] Computer-mediated reality systems are being developed to allow computing devices to enhance, add to, remove from, subtract from, or generally modify the user experience of an existing reality. Computer-mediated reality systems (which may also be referred to as “extended reality systems” or “XR systems”) can include, for example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceptual success of computer-mediated reality systems is often related to their ability to provide realistic and immersive experiences in both video and audio, where the video and audio experiences are aligned in a way that the user expects. Although the human visual system is more sensitive than the human auditory system (e.g., in perceptually locating various objects in a scene), ensuring a sufficiently good auditory experience is an increasingly important factor in ensuring realistic and immersive experiences, especially as video experiences improve to allow for better location of video objects, thus enabling users to better identify the source of audio content. Summary of the Invention

[0004] This disclosure generally relates to the auditory aspects of user experience in computer-mediated reality systems, including virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. Various aspects of this technology can provide adaptive audio capture or synthesis and rendering for extending the acoustic space of a reality system. As used herein, an acoustic environment is referred to as an indoor environment or an outdoor environment, or both. An acoustic environment may include one or more subacoustic spaces, which may include various acoustic elements. Examples of outdoor environments may include cars, buildings, walls, forests, etc. An acoustic space can be an example of an acoustic environment and can be an indoor space or an outdoor space. As used herein, audio elements are sounds captured by a microphone (e.g., captured directly from near-field sources or reflections from far-field sources, whether real or synthesized), or previously synthesized sound fields, or monophonic sounds synthesized from text to speech, or reflections of virtual sounds from objects in the acoustic environment.

[0005] In one example, aspects of the technology point to a device comprising: a memory configured to store a plurality of audio streams and an associated license level for each of the plurality of audio streams; and one or more processors implemented in circuitry and communicatively coupled to the memory, and configured to: select a subset of the plurality of audio streams based on the associated license level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0006] In another example, aspects of the technology point to a method comprising: storing a plurality of audio streams and an associated authorization level for each of the plurality of audio streams in a memory; and selecting a subset of the plurality of audio streams by one or more processors based on the associated authorization level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0007] In another example, aspects of the technology refer to a device comprising: components for storing a plurality of audio streams and associated authorization levels for each of the plurality of audio streams; and components for selecting a subset of the plurality of audio streams based on the associated authorization levels, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0008] In another example, aspects of the technology refer to a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: store a plurality of audio streams and an associated authorization level for each of the plurality of audio streams; and select a subset of the plurality of audio streams based on the associated authorization level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0009] Details of one or more examples of this disclosure are set forth in the accompanying drawings and the following description. Other features, objects, and advantages of various aspects of the present technology will be apparent from the specification, drawings, and claims. Attached Figure Description

[0010] Figures 1A to 1C This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure.

[0011] Figure 2 This is a diagram showing an example of a VR device worn by a user.

[0012] Figures 3A to 3E To show in more detail Figures 1A to 1C The example shown is a diagram of the example operation of the stream selection unit.

[0013] Figure 4A and Figure 4B It is shown Figures 1A to 1C The example shown is a flowchart illustrating the operation of a stream selection unit in various aspects of stream selection technology.

[0014] Figure 4C and Figure 4D This is a diagram illustrating various aspects of the technology described in this disclosure with respect to the privacy area.

[0015] Figure 4E and Figure 4F This is a diagram further illustrating the use of privacy zones in various aspects of the technology described in this disclosure.

[0016] Figure 4G and Figure 4H This is a diagram illustrating the exclusion of separate audio streams according to various aspects of the technology described in this disclosure.

[0017] Figure 5 This is a diagram illustrating an example of a wearable device that can operate according to various aspects of the technology described in this disclosure.

[0018] Figure 6A and Figure 6B This is a diagram illustrating other example systems that can perform various aspects of the techniques described in this disclosure.

[0019] Figure 7 It is shown Figures 1A to 1C The example shown is a block diagram of one or more sample components from the source device and content consumer device.

[0020] Figures 8A to 8C It is shown Figures 1A to 1C The example shown is a flowchart illustrating sample operations of the stream selection unit in performing various aspects of the stream selection technique.

[0021] Figure 9 Examples of wireless communication systems supporting privacy zones and license levels according to various aspects of this disclosure are shown. Detailed Implementation

[0022] When an audio scene is rendered using a plurality of audio sources that can be obtained by audio capture devices in the scene or synthesized, certain zones may contain audio sources that may include sensitive information that should be restricted from access. According to the techniques disclosed herein, a subset of multiple audio streams is selected based on an associated authorization level for each of the multiple audio streams. In some examples, one or more audio streams of the multiple audio streams are associated with at least one privacy zone. In some examples, the gain of one or more audio streams in the subset of multiple audio streams can be changed based on the associated authorization level. In some examples, excluded audio streams can be set to empty.

[0023] The techniques disclosed herein can provide the ability to protect sensitive information when rendering audio scenes with multiple audio sources. In some examples, the techniques disclosed herein can provide the ability to protect sensitive information on the rendering side when the capture side cannot restrict access to audio streams containing sensitive information.

[0024] There are several different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound, 7.1 surround sound, 22.2 surround sound, or any other channel-based format that positions audio channels to specific locations around the listener to reconstruct the sound field.

[0025] Object-based audio formats can refer to formats in which audio objects are specified to represent a sound field. These audio objects are typically encoded using Pulse Code Modulation (PCM) and are referred to as PCM audio objects. Such audio objects may include location information (e.g., metadata) identifying the position of the audio object relative to a listener or other reference points in the sound field, allowing the audio object to be rendered to one or more speaker channels for playback when attempting to reconstruct the sound field. The techniques described in this disclosure can be applied to any of the following formats, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.

[0026] Scene-based audio formats can include a set of hierarchical elements that define the sound field in three dimensions. An example of a set of hierarchical elements is a set of spherical harmonic coefficients (SHCs). The following expression demonstrates a description or representation of the sound field using SHCs:

[0027]

[0028] This expression indicates that at any point in the sound field at time t... Pressure p at the point i It can be made by SHC, Unique representation. Here, c is the speed of sound (approximately 343 m / s). It is the reference point (or observation point), j n (·) is an nth-order spherical Bessel function, and These are spherical harmonic basis functions of order n and sub-order m (also called spherical basis functions). It can be recognized that the terms within the square brackets are the frequency domain representation of the signal (e.g., This can be approximated by various time-frequency transforms, such as the Discrete Fourier Transform (DFT), Discrete Cosine Transform, or Wavelet Transform. Other examples of hierarchical sets include wavelet transform coefficient sets and other sets of coefficients for multi-resolution basis functions.

[0029] They can be physically acquired (e.g., recorded) through various microphone array configurations, or alternatively, they can be derived from channel-based or object-based descriptions of the sound field. SHC (also known as ambisonic coefficients) represent scene-based audio, where SHC can be input to an audio encoder to obtain an encoded SHC that facilitates more efficient transmission or storage. For example, a (1+4) approach could be used. 2 A fourth-order representation of 25 (and therefore fourth-order) coefficients.

[0030] As mentioned above, SHC can be derived from microphone recordings using a microphone array. Various examples of how SHC can be physically obtained from a microphone array are described in Poletti, M., “Three-Dimensional Surround Sound Systems Based on Spherical Harmonics” (J. Audio Eng. Soc., Vol. 53, No. 11, November 2005, pp. 1004-1025).

[0031] The following equation illustrates how to derive SHC from an object-based description. The sound field coefficients correspond to individual audio objects. It can be represented as:

[0032]

[0033] Where i is It is an nth-order spherical Hankel function (second kind), and This refers to the object's location. Understanding the object's source energy g(ω) as a function of frequency (e.g., using time-frequency analysis techniques such as performing a Fast Fourier Transform on a Pulse Code Modulation (PCM) stream) allows us to convert each PCM object and its corresponding location into... Furthermore, it can be shown (since the above is a linear and orthogonal decomposition) that each object The coefficients are additive. In this way, multiple PCM objects can be generated by... Each coefficient represents (e.g., the sum of coefficient vectors for individual objects). The coefficients can contain information about the sound field (pressure as a function of three-dimensional (3D) coordinates), and the above represents the distance from individual objects to the observation point. The transformation of the representation of the entire surrounding sound field.

[0034] Computer-mediated reality systems (also known as “extended reality systems” or “XR systems”) are being developed to take advantage of the many potentially beneficial effects offered by surround sound coefficients. For example, surround sound coefficients can represent a sound field in three dimensions in a way that can potentially enable accurate 3D localization of sound sources within the sound field. Therefore, XR devices can render surround sound coefficients to speaker feeds, accurately reproducing the sound field when played through one or more speakers or headphones.

[0035] As another example, surround sound coefficients can be translated or rotated to account for user movement without overly complex mathematical calculations, potentially accommodating the low latency requirements of XR devices. Furthermore, the surround sound coefficients are hierarchical, naturally adapting to scalability through order reduction (which eliminates surround sound coefficients associated with higher orders), and thus potentially enabling dynamic adaptation of the sound field to accommodate the latency and / or battery requirements of XR devices.

[0036] Using surround sound coefficients in XR devices enables the development of many use cases that rely on a more immersive sound field provided by surround sound coefficients, particularly for computer gaming and real-time video streaming applications. In these high-dynamic use cases that rely on low-latency sound field reproduction, XR devices may prefer surround sound coefficients over other representations that are more difficult to manipulate or involve complex rendering. The following discusses… Figures 1A to 1C Provide more information about these use cases.

[0037] Although VR devices are described in this disclosure, various aspects of these technologies can be performed in the context of other devices, such as mobile devices. In this case, a mobile device (such as a so-called smartphone) can present the acoustic space via a screen that can be mounted on the user's head or viewed in the same way as when using a mobile device normally. Therefore, any information on the screen can be part of the mobile device. The mobile device may be able to provide tracking information, allowing both the VR experience (when head-mounted) and the normal experience to view the acoustic space, where the normal experience can still allow the user to view the acoustic space that provides a VR-lite-type experience (e.g., raising the device and rotating or panning it to view different parts of the acoustic space).

[0038] Figures 1A to 1C This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure. For example... Figure 1AAs shown in the example, system 10 includes a source device 12A and a content consumer device 14A. Although described in the context of source device 12A and content consumer device 14A, these techniques can be implemented in any context in which any representation of the sound field is encoded to form a bitstream representing audio data. Moreover, source device 12A can represent any form of computing device capable of generating a sound field representation, and is generally described herein in the context of a VR content creator device. Similarly, content consumer device 14A can represent any form of computing device capable of implementing the rendering techniques and audio playback described in this disclosure, and is generally described herein in the context of a VR client device.

[0039] Source device 12A can be operated by an entertainment company or other entity capable of generating single-channel and / or multi-channel audio content for consumption by an operator of a content consumer device (such as content consumer device 14A). In some VR scenarios, source device 12A combines video content to generate audio content. Source device 12A includes content capture device 20, content editing device 22, and sound field representation generator 24. Content capture device 20 can be configured to interface with microphone 18 or otherwise communicate.

[0040] Microphone 18 can indicate Or other types of 3D audio microphones capable of capturing a sound field and representing it as audio data 19, which may refer to one or more of the scene-based audio data (such as surround sound coefficients), object-based audio data, and channel-based audio data described above. Although described as a 3D audio microphone, microphone 18 may also represent other types of microphones configured to capture audio data 19 (such as omnidirectional microphones, spot microphones, unidirectional microphones, etc.). Audio data 19 may represent an audio stream or include an audio stream.

[0041] In some examples, the content capture device 20 may include an integrated microphone 18 integrated into the housing of the content capture device 20. The content capture device 20 may interface with the microphone 18 wirelessly or via a wired connection. Instead of capturing audio data 19 via, or in combination with, capturing audio data 19 via, the content capture device 20 may process the audio data 19 after it has been input via some type of removable storage, wireless, and / or wired input process. Therefore, various combinations of the capture device 20 and the microphone 18 are possible according to this disclosure.

[0042] Content capture device 20 may also be configured to interface with or otherwise communicate with content editing device 22. In some cases, content capture device 20 may include content editing device 22 (in some cases, it may represent software, or a combination of software and hardware, including software executed by content capture device 20 to configure content capture device 20 to perform a particular form of content editing). Content editing device 22 may represent a unit configured to edit or otherwise alter content 21 (including audio data 19) received from content capture device 20. Content editing device 22 may output the edited content 23 and associated information (e.g., metadata) 25 to sound field representation generator 24.

[0043] The sound field representation generator 24 may include any type of hardware device capable of interfacing with the content editing device 22 (or the content capture device 20). Although not explicitly stated... Figure 1A As shown in the example, the sound field representation generator 24 can use edited content 23, including audio data 19 and information (e.g., metadata) 25, provided by the content editing device 22, to generate one or more bitstreams 27. Focusing on the audio data 19... Figure 1A In the example, the sound field representation generator 24 can generate one or more representations of the same sound field represented by the audio data 19 to obtain a bitstream 27, which includes the representation of the sound field and information (e.g., metadata) 25.

[0044] For example, in order to generate different representations of the sound field using surround sound coefficients (again, an example of audio data 19), the sound field representation generator 24 can use an encoding scheme for the surround sound representation of the sound field, which is called Mixed Order Surround Sound (MOA), as discussed in more detail in U.S. Patent Application Serial No. 15 / 672,058 entitled “MIXED-ORDER AMBISONICS (MOA) AUDIO DATA FOCOMPUTER-MEDIATED REALITY SYSTEMS”, filed August 8, 2017, and published as U.S. Patent Publication No. 20190007781 on January 3, 2019.

[0045] To generate a specific MOA representation of a sound field, the sound field representation generator 24 can generate a subset of the complete set of surround sound coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide accuracy for some regions of the sound field, but lower accuracy for others. In one example, the MOA representation of the sound field may include eight (8) uncompressed surround sound coefficients, while the third-order surround sound representation of the same sound field may include sixteen (16) uncompressed surround sound coefficients. Thus, each MOA representation of the sound field generated as a subset of surround sound coefficients may be less storage-intensive and less bandwidth-intensive (if and when transmitted as part of the bitstream 27 through the illustrated transmission channel) compared to the corresponding third-order surround sound representation of the same sound field generated from surround sound coefficients.

[0046] Although the MOA representation has been described, the techniques disclosed herein can also be performed with respect to the first-order surround sound (FOA) representation, where all surround sound coefficients associated with the first-order and zero-order spherical basis functions are used to represent the sound field. In other words, the sound field representation generator 24 can represent the sound field using all surround sound coefficients of a given order N, instead of using a partial non-zero subset of the surround sound coefficients, resulting in a total of (N+1) 2 The surround sound coefficient.

[0047] In this respect, surround sound audio data (which is another way of referring to surround sound coefficients in MOA representation or full-order representation (such as the first-order representation mentioned above)) can include surround sound coefficients associated with spherical basis functions of order 1 or less (which may be referred to as "first-order surround sound audio data"), surround sound coefficients associated with spherical basis functions having mixed orders and sub-orders (which may be referred to as "MOA representation" discussed above), or surround sound coefficients associated with spherical basis functions of order greater than 1 (which are referred to as "full-order representation" above).

[0048] In some examples, the content capture device 20 or the content editing device 22 may be configured to communicate wirelessly with the sound field representation generator 24. In some examples, the content capture device 20 or the content editing device 22 may communicate with the sound field representation generator 24 via one or both of a wireless or wired connection. Through the connection between the content capture device 20 or the content editing device 22 and the sound field representation generator 24, the content capture device 20 or the content editing device 22 may provide various forms of content, which, for the purposes of discussion, are described herein as part of audio data 19.

[0049] In some examples, content capture device 20 may utilize various aspects of sound field representation generator 24 (in terms of the hardware or software capabilities of sound field representation generator 24). For example, sound field representation generator 24 may include dedicated hardware (or specialized software) configured (or, at execution, to cause one or more processors) to perform psychoacoustic audio coding, such as the Unified Speech and Audio Codec denoted as “USAC,” which is articulated by the Moving Picture Experts Group (MPEG), the MPEG-H 3D Audio Codec Standard, the MPEG-I Immersive Audio Standard, or proprietary standards such as AptX. TM (Including various versions of AptX, such as Enhanced AptX–E-AptX, AptX Live, AptX Stereo, and AptX High Definition–AptX-HD), Advanced Audio Codec (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey's Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Microsoft Media Audio (WMA) or other standards.

[0050] Content capture device 20 may not include dedicated hardware or software for psychoacoustic audio encoders, but may instead provide the audio aspects of content 21 in a non-psychoacoustic audio codec format. Sound field representation generator 24 can assist in capturing content 21 by performing psychoacoustic audio encoding at least partially with respect to the audio aspects of content 21.

[0051] The sound field representation generator 24 can also assist in content capture and transmission by generating one or more bitstreams 27 based at least in part on audio content (e.g., MOA representation and / or first-order surround sound representation) generated from the audio data 19 (where the audio data 19 includes scene-based audio data). The bitstreams 27 can represent compressed versions of the audio data 19 and any other different types of content 21 (such as compressed versions of spherical video data, image data, or text data).

[0052] The sound field representation generator 24 can generate a bitstream 27 for transmission, as an example, via a transmission channel, which can be a wired or wireless channel (such as a Wi-Fi channel, Bluetooth channel, or a channel compliant with 5G cellular standards), a data storage device, etc. Bitstream 27 can represent an encoded version of audio data 19 and can include a main bitstream and another side bitstream, which can be referred to as side channel information or metadata. In some cases, bitstream 27, representing a compressed version of audio data 19 (which can also represent scene-based audio data, object-based audio data, channel-based audio data, or a combination thereof), can conform to a bitstream generated according to the MPEG-H3D audio codec standard and / or the MPEG-I immersive audio standard.

[0053] Content consumer device 14A can be operated by an individual and can represent a VR client device. Although described in relation to a VR client device, content consumer device 14A can represent other types of devices, such as augmented reality (AR) client devices, mixed reality (MR) client devices (or other XR client devices), standard computers, headsets, headphones, mobile devices (including so-called smartphones), or any other device capable of tracking the head movements and / or general translational movements of the individual operating content consumer device 14A. Figure 1A As shown in the example, the content consumer device 14A includes an audio playback system 16A, which can refer to any form of audio playback system capable of rendering audio data for playback as single-channel and / or multi-channel audio content.

[0054] Although Figure 1A The bitstream 27 is shown as being transmitted directly to content consumer device 14A, but source device 12A can output the bitstream 27 to an intermediate device located between source device 12A and content consumer device 14A. The intermediate device can store the bitstream 27 for later delivery to content consumer device 14A, which may request the bitstream 27. The intermediate device can include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing the bitstream 27 for later retrieval by an audio decoder. The intermediate device can exist in a content delivery network capable of streaming the bitstream 27 (and possibly in conjunction with the transmission of corresponding video data bitstreams) to subscribers (such as content consumer device 14A) that requested the bitstream 27.

[0055] Alternatively, source device 12A may store bitstream 27 to a storage medium, such as an optical disc, digital video disc, high-definition video disc, or other storage medium, most of which are computer-readable and therefore may be referred to as computer-readable storage media or non-transitory computer-readable storage media. In this context, a transmission channel may refer to a channel through which the content stored to the medium (e.g., in the form of one or more bitstreams 27) is transmitted (and may include retail stores and other storage-based delivery mechanisms). In no event should the technology disclosed herein be limited to this. Figure 1A Examples.

[0056] As described above, the content consumer device 14A includes an audio playback system 16A. The audio playback system 16A can represent any system capable of playing back single-channel and / or multi-channel audio data. The audio playback system 16A can include multiple different audio renderers 32. Each audio renderer 32 can provide different forms of rendering, wherein different forms of rendering can include one or more of various methods of performing vector-based amplitude shifting (VBAP), and / or one or more of various methods of performing sound field synthesis. As used herein, “A and / or B” means “A or B”, or “both A and B”.

[0057] The audio playback system 16A may also include an audio decoding device 34. Audio decoding device 34 may refer to a device configured to decode bitstream 27 to output audio data 19' (where the apostrophe may indicate that audio data 19' differs from audio data 19 due to lossy compression such as quantization). Similarly, audio data 19' may include scene-based audio data, in some examples, which may form a complete first-order (or higher) order surround sound representation or a subset thereof forming a MOA representation of the same sound field; its decomposition, such as the dominant audio signal, ambient surround sound coefficients, and vector-based signals as described in the MPEG-H3D audio codec standard; or other forms of scene-based audio data. Audio data 19' may include an audio stream or a representation of an audio stream.

[0058] Other forms of scene-based audio data include audio data defined according to the HOA (Higher Order Ambisonics) Transport Format (HTF). More information about HTF can be found in the European Telecommunications Standards Institute (ETSI) technical specification (TS) entitled "Higher Order Ambisonics (HOA) Transport Format" (ETSI TS103 589V1.1.1, dated June 2018 (2018-06)) and U.S. Patent Publication No. 2019 / 0918028 entitled "PRIORITY INFORMATION FOR HIGHER ORDER AMBISONIC AUDIO DATA," filed December 20, 2018. In any case, audio data 19' may resemble a complete set or a subset of audio data 19, but may differ due to lossy operations (e.g., quantization) and / or transmission via a transmission channel.

[0059] As an alternative to or in combination with scene-based audio data, audio data 19' may include channel-based audio data. As an alternative to or in combination with scene-based audio data, audio data 19' may include object-based audio data or channel-based audio. Therefore, audio data 19' may include any combination of scene-based audio data, object-based audio data, and channel-based audio data.

[0060] The audio renderer 32 of the audio playback system 16A can render the audio data 19' to output the speaker feed 35 after the audio decoding device 34 has decoded the bitstream 27 to obtain the audio data 19'. The speaker feed 35 can drive one or more speakers or headphones (for illustrative purposes, it is...) Figure 1A (Not shown in the example). Various audio representations, including scene-based audio data of the sound field (and possibly channel-based and / or object-based audio data), can be normalized in a variety of ways, including N3D, SN3D, FuMa, N2D, or SN2D.

[0061] To select an appropriate renderer, or in some cases generate an appropriate renderer, the audio playback system 16A may obtain speaker information 37 indicating the number of speakers (e.g., loudspeakers or headphone speakers) and / or the spatial geometry of the speakers. In some cases, the audio playback system 16A may use a reference microphone to obtain the speaker information 37 and may drive the speakers in a manner that dynamically determines the speaker information 37 (which may refer to the output of an electrical signal used to cause transducer oscillations). In other cases, or in conjunction with the dynamic determination of the speaker information 37, the audio playback system 16A may prompt the user to interact with the audio playback system 16A and input the speaker information 37.

[0062] The audio playback system 16A can select an audio renderer from the audio renderers 32 based on the speaker information 37. In some cases, when none of the audio renderers 32 falls within a certain threshold similarity metric (in terms of speaker geometry) specified in the speaker information 37, the audio playback system 16A can generate an audio renderer from the audio renderers 32 based on the speaker information 37. In some cases, the audio playback system 16A can generate an audio renderer from the audio renderers 32 based on the speaker information 37 without first attempting to select an existing audio renderer from the audio renderers 32.

[0063] When the speaker feed 35 is output to the headphones, the audio playback system 16A can utilize an audio renderer in renderer 32 to provide binaural rendering, such as a binaural chamber impulse response renderer, using a head-related transfer function (HRTF) or other functions capable of rendering the left and right speaker feeds 35 for headphone speaker playback. The term "speaker" or "transducer" can generally refer to any speaker, including horns, headphone speakers, bone conduction speakers, earbud speakers, wireless headphone speakers, etc. One or more speakers or headphones can then play back the rendered speaker feed 35 to reproduce the sound field.

[0064] Although described as rendering speaker feed 35 by audio data 19', references to rendering speaker feed 35 could refer to other types of rendering, such as rendering directly incorporated into the decoding of audio data from bitstream 27. Examples of alternative rendering can be found in Annex G of the MPEG-H 3D audio standard, where rendering occurs during the formation of the dominant signal and background signal prior to sound field synthesis. Therefore, references to rendering audio data 19' should be understood to refer to either the rendering of the actual audio data 19', or the decomposition or representation of audio data 19' (such as the aforementioned dominant audio signal, ambient surround sound coefficients, and / or vector-based signals—which may also be referred to as V-vectors or multidimensional surround sound space vectors) or both.

[0065] The audio playback system 16A can also adjust the audio renderer 32 based on the tracking information 41. That is, the audio playback system 16A can interface with a tracking device 40 configured to track the head movements and possible translational movements of a user of the VR device. The tracking device 40 can represent one or more sensors (e.g., cameras (including depth cameras), gyroscopes, magnetometers, accelerometers, light-emitting diodes (LEDs), etc.) configured to track the head movements and possible translational movements of a user of the VR device. The audio playback system 16A can adjust the audio renderer 32 based on the tracking information 41 so that the speaker feed 35 reflects the changes in the user's head and possible translational movements to accurately reproduce the sound field in response to such movements.

[0066] Figure 1B This is a block diagram illustrating another exemplary system 50 configured to perform various aspects of the techniques described in this disclosure. System 50 is similar to... Figure 1A System 10 shown, except Figure 1A In addition to replacing the audio renderer 32 shown, which is in the audio playback system 16B of the content consumer device 14B, the binaural renderer 42 is capable of performing binaural rendering using one or more head-related transfer functions (HRTF) or other functions that can render to the left and right speaker feeds 43.

[0067] The audio playback system 16B can output left and right speaker feeds 43 to a headset 48, which can represent another example of a wearable device and can be coupled to additional wearable devices to facilitate sound field reproduction, such as watches, the aforementioned VR headsets, smart glasses, smart clothing, smart rings, smart bracelets, or any other type of smart jewelry (including smart necklaces). The headset 48 can be coupled to the additional wearable device wirelessly or via a wired connection.

[0068] Additionally, the headphones 48 can be connected via wired connections (such as a standard 3.5mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other forms of wired connection) or wirelessly (such as via Bluetooth). TM (Connectivity, wireless network connection, etc.) is coupled to the audio playback system 16B. The headphones 48 can recreate the sound field represented by the audio data 19' based on the left and right speaker feeds 43. The headphones 48 may include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the corresponding left and right speaker feeds 43.

[0069] Figure 1C This is a block diagram illustrating another exemplary system 60. Exemplary system 60 is similar to... Figure 1A The exemplary system 10 is used, but the source device 12B of system 60 does not include a content capture device. Source device 12B includes a compositing device 29. Content developers can use the compositing device 29 to generate a synthesized audio source. The synthesized audio source may have associated positioning information that identifies the position of the audio source relative to a listener or other reference points in the sound field, allowing the audio source to be rendered to one or more speaker channels for playback in an attempt to reconstruct the sound field. In some examples, the compositing device 29 may also synthesize visual or video data.

[0070] For example, content developers can generate synthesized audio streams for video games. Although Figure 1C Examples and Figure 1A The example content of consumer device 14A is shown together, but Figure 1C Example source device 12B can be with Figure 1B It is used together with the content consumer device 14B. In some examples, Figure 1C The source device 12B may also include a content capture device, such that the bitstream 27 may simultaneously contain both captured audio stream(s) and synthesized audio stream(s).

[0071] As described above, content consumer device 14A or 14B (either of which may be referred to as content consumer device 14 below) may represent a VR device in which a human wearable display (which may also be referred to as a “head-mounted display”) is mounted in front of the eyes of the user operating the VR device. Figure 2This is a diagram illustrating an example of a VR device 1100 worn by a user 1102. The VR device 1100 is coupled to or otherwise includes a headset 1104, which can reproduce the sound field represented by audio data 19' via playback from a speaker feed 35. The speaker feed 35 can represent an analog or digital signal that enables the diaphragm within the transducer of the headset 1104 to vibrate at various frequencies, a process commonly referred to as driving the headset 1104.

[0072] Video, audio, and other sensory data can play a significant role in VR experiences. To participate in a VR experience, user 1102 can wear VR device 1100 (which may also be referred to as VR client device 1100) or other wearable electronic devices. VR client devices (such as VR device 1100) may include tracking devices (e.g., tracking device 40) configured to track the head movements of user 1102 and adjust the video data displayed via VR device 1100 to account for head movements, thereby providing an immersive experience where user 1102 can experience an acoustic space displayed in visual three-dimensionality within the video data. The acoustic space can refer to a virtual world (where all worlds are simulated), an augmented world (where a portion of the world is augmented by virtual objects), or a physical world (where real-world images are virtualized for navigation).

[0073] While VR (and other forms of AR and / or MR) can allow users 1102 to visually reside in a virtual world, VR devices 1100 may typically lack the ability to place the user in an audible acoustic space. In other words, a VR system (which may include a computer responsible for rendering video and audio data—not shown for illustrative purposes)... Figure 2 As shown in the example, VR device 1100 may not audibly (and in some cases realistically, in a way that reflects the display scene presented to the user via VR device 1100) support full 3D immersion.

[0074] Although VR devices are described in this disclosure, various aspects of these technologies can be performed in the context of other devices, such as mobile devices. In this case, a mobile device (such as a so-called smartphone) can present the acoustic space via a screen that can be mounted on the head of user 1102 or viewed in the manner of normal mobile device use. Therefore, any information on the screen can be part of the mobile device. The mobile device can be able to provide tracking information 41, thereby allowing both the VR experience (when worn) and the normal experience to view the acoustic space, where the normal experience can still allow the user to view the acoustic space that provides the VR-simplified experience (e.g., raising the device and rotating or panning the device to view different parts of the acoustic space).

[0075] Regardless, returning to the context of VR devices, the audio aspects of VR have been categorized into three distinct immersive categories. The first category provides the lowest level of immersion and is known as three degrees of freedom (3DOF). 3DOF refers to audio rendering that takes into account head movements in three degrees of freedom (yaw, pitch, and roll), allowing users to freely look around in any direction. However, 3DOF cannot account for translational head movements that are not centered on the optical and acoustic center of the sound field.

[0076] The second category, called 3DOF plus (3DOF+), provides three degrees of freedom (yaw, pitch, and roll) as well as limited spatial translational motion that deviates from the optical and acoustic centers within the sound field due to head movements. 3DOF+ can support perceptual effects such as motion parallax, which can enhance immersion.

[0077] The third category, called six degrees of freedom (6DOF), renders audio data in a way that considers three degrees of freedom in terms of head movement (yaw, pitch, and roll) but also the user's translations in space (x, y, and z translations). Spatial translations can be induced by sensors that track the user's position in the physical world or by input controllers.

[0078] 3DOF rendering is the latest technology in VR audio. Therefore, VR audio is less immersive than video, potentially reducing the overall immersiveness of the user experience. However, VR is rapidly evolving and may quickly develop to support both 3DOF+ and 6DOF, which could open up opportunities for additional use cases.

[0079] For example, interactive gaming applications can leverage 6DOF to facilitate fully immersive gaming, where users move freely within the VR world and interact with virtual objects by walking around them. Furthermore, interactive live streaming applications can utilize 6DOF to allow VR client devices to experience live concerts or sporting events as if they were actually at the concert, thus allowing users to move freely within the concert or sporting event.

[0080] There are many challenges associated with these use cases. In the case of fully immersive gaming, latency may need to be kept at low levels to achieve gameplay that does not cause nausea or motion sickness. Moreover, from an audio perspective, latency in audio playback that causes desynchronization with video data can reduce immersion. Furthermore, for certain types of game applications, spatial accuracy can be important for allowing accurate responses, including how users perceive sound, as this allows users to predict actions that are not currently in their field of vision.

[0081] In the context of a live streaming application, a large number of source devices 12A or 12B (any of which may be referred to as source device 12 hereinafter) can stream content 21, where source devices 12 can have a wide variety of different capabilities. For example, one source device 12 may be a smartphone with a digital fixed-lens camera and one or more microphones, while another source device may be a production-grade television setup capable of delivering video at a higher resolution and quality than a smartphone. However, in the context of a live streaming application, all source devices 12 can provide streams of different qualities, from which the VR device can attempt to select a suitable stream to provide the desired experience.

[0082] Furthermore, similar to gaming applications, latency in audio data can cause it to become out of sync with video data, potentially reducing immersion. Spatial accuracy can also be important, allowing users to better understand the background or location of different audio sources. Additionally, privacy can become an issue when users are live-streaming using cameras and microphones, as they may not want their streams to be completely public.

[0083] In the context of streaming applications (real-time or recorded), there may be a large number of audio streams associated with different levels of quality and / or content. Audio streams can represent any type of audio data, including scene-based audio data (e.g., surround sound audio data, including FOA, MOA, and / or HOA audio data), channel-based audio data, and object-based audio data. Selecting only one of the potentially large number of audio streams to recreate the sound field may not provide an experience that ensures a sufficient level of immersion. However, selecting multiple audio streams can be disruptive due to their different spatial positioning, potentially reducing immersion.

[0084] According to the technology described herein, the audio decoding device 34 can adaptively select among audio streams available via bitstream 27 (which is represented by bitstream 27 and therefore bitstream 27 may be referred to as "audio stream 27"). The audio decoding device 34 can also be based on audio location information (ALI) (e.g., Figures 1A to 1C ALI 45A selects among different audio streams of audio stream 27. In some examples, audio location information may be included as metadata accompanying audio stream 27, wherein the audio location information may define the capture coordinates of the microphone capturing the corresponding audio stream 27 in acoustic space, or the virtual capture coordinates of the synthesized audio stream in acoustic space. ALI 45A may represent the location of a corresponding audio stream in acoustic space among the captured or synthesized audio streams 27. Audio decoding device 34 may select a subset of audio streams 27 based on ALI 45A, wherein the subset of audio streams 27 excludes at least one audio stream from audio stream 27. Audio decoding device 34 may output the subset of audio streams 27 as audio data 19'.

[0085] Additionally, the audio decoding device 34 can obtain tracking information 41, which the content consumer device 14 can convert into device location information (DLI) (e.g., Figures 1A to 1C (45B in the original text). DLI 45B can represent the virtual or physical location of the content consumer device 14 in the acoustic space, and can be defined as one or more device coordinates in the acoustic space. The content consumer device 14 can provide DLI 45B to the audio decoding device 34. The audio decoding device 34 can then select audio data 19' from the audio stream 27 based on ALI 45A and DLI 45B. The audio playback system 16A or 16B can then reproduce the corresponding sound field based on the audio data 19'.

[0086] In this respect, the audio decoding device 34 can adaptively select a subset of the audio stream 27 to obtain audio data 19' that can lead to a more immersive experience (compared to selecting a single audio stream or all audio data 19'). Therefore, various aspects of the techniques described in this disclosure can improve the operation of the audio decoding device 34 (and the audio playback system 16A or 16B and the content consumer device 14) itself by potentially enabling the audio decoding device 34 to better spatialize sound sources within the sound field, thereby enhancing immersion.

[0087] In operation, the audio decoding device 34 can interact with one or more source devices 12 to determine the ALI 45A for each audio stream 27. For example... Figure 1A As shown in the example, the audio decoding device 34 may include a stream selection unit 44, which may represent a unit configured to perform various aspects of the audio stream selection technology described in this disclosure.

[0088] The stream selection unit 44 can generate a constellation diagram (CM) 47 based on the ALI 45A. The CM 47 can define the ALI 45A for each audio stream 27. The stream selection unit 44 can also perform energy analysis with respect to each audio stream 27 to determine the energy map of each audio stream 27, storing the energy map along with the ALI 45A in the CM 47. The energy maps can collectively define the energy of a common sound field represented by the audio streams 27.

[0089] The stream selection unit 44 can then determine the distance between the device location represented by the DLI 45B and the capture location or synthesis location represented by the ALI 45A associated with at least one (and possibly each) of the audio streams 27. The stream selection unit 44 can then select audio data 19' from the audio streams 27 based on the distance(s), as described below. Figures 3A to 3E To be discussed in more detail.

[0090] Furthermore, in some examples, the stream selection unit 44 can also select audio data 19' from the audio stream 27 based on the energy map stored in the CM 47, ALI 45A, and DLI 45B (where ALI 45A and DLI 45B are presented together in the form of the aforementioned distances, also referred to as "relative distances"). For example, the stream selection unit 44 can analyze the energy map presented in the CM 47 to determine the audio source location (ASL) 49 of the audio source in the common sound field that emitted the sound, captured or synthesized by a microphone (e.g., microphone 18) and represented by the audio stream 27. The stream selection unit 44 can then determine the audio data 19' from the audio stream 27 based on ALI 45A, DLI 45B, and ASL 49. The following is about Figures 3A to 3E More information is discussed about how the stream selection unit 44 can select streams.

[0091] Figures 3A to 3E To show in more detail Figure 1A The example shown is a diagram illustrating the example operation of the flow selection unit 44. Figure 3A As shown in the example, the streaming selection unit 44 can determine that the DLI 45B indicates that the content consumer device 14 (shown as VR device 1100) is at a virtual location 300A. The streaming selection unit 44 can then determine one or more of the audio elements 302A-302J (collectively referred to as audio elements 302), which can represent not only microphones, but also other audio elements such as... Figure 1A The microphone 18 shown can also represent other types of capture devices, including other XR devices, mobile phones (including so-called smartphones, etc.), or synthetic sound fields, etc.

[0092] As described above, the stream selection unit 44 can obtain the audio stream 27. The stream selection unit 44 can interface with audio components 302A-302J to obtain the audio stream 27. In some examples, the stream selection unit 44 can be configured according to fifth-generation (5G) cellular standards, such as Bluetooth. TM Personal Area Networks (PANs) or other open-source, proprietary, or standardized communication protocols are used to interact with interfaces (such as receivers, transmitters, and / or transceivers) to obtain audio streams.27 Wireless communication of audio streams in… Figures 3A to 3E and Figures 4E to 4H In the example, it is represented as lightning, where the selected audio data 19' is shown as being communicated from one or more selected audio elements 302 to VR device 1100.

[0093] In any case, the stream selection unit 44 can then obtain an energy map in the manner described above, analyze the energy map to determine the audio source location 304, which can represent Figure 1AThe example shown is an example of ASL 49. An energy map can represent audio source location 304 because the energy at audio source location 304 may be higher than the surrounding area. Assuming each energy map can represent this higher energy, the stream selection unit 44 can triangulate the audio source location 304 based on the higher energy in the energy map.

[0094] Next, the streaming selection unit 44 can determine the audio source distance 306A as the distance between the audio source location 304 and the virtual location 300A of the VR device 1100. The streaming selection unit 44 can compare the audio source distance 306A with an audio source distance threshold. In some examples, the streaming selection unit 44 can derive the audio source distance threshold based on the energy of the audio source 308. That is, when the audio source 308 has high energy (or in other words, when the audio source 308 is louder), the streaming selection unit 44 can increase the audio source distance threshold. When the audio source 308 has low energy (or in other words, when the audio source 308 is quieter), the streaming selection unit 44 can decrease the audio source distance threshold. In other examples, the streaming selection unit 44 can obtain a statically defined audio source distance threshold, which can be statically defined or specified by the user 1102.

[0095] In any case, when the audio source distance 306A is greater than the audio source distance threshold (assumed for illustrative purposes in this example), the stream selection unit 44 can select a single audio stream from the audio stream 27 of audio elements 302A-302J (“audio element 302”). The stream selection unit 44 can output the corresponding audio stream from the audio stream 27, which the audio decoding device 34 can decode and output as audio data 19'.

[0096] Suppose user 1102 moves from virtual location 300A to virtual location 300B, the stream selection unit 44 can determine the audio source distance 306B as the distance between the audio source location 304 and the virtual location 300B. In some examples, the stream selection unit 44 can update only after a configurable release time, which may refer to the time after the listener stops moving.

[0097] In any case, the stream selection unit 44 can again compare the audio source distance 306B with the audio source distance threshold. When the audio source distance 306B is less than or equal to the audio source distance threshold (assumed for illustrative purposes in this example), the stream selection unit 44 can select multiple audio streams from the audio stream 27 of audio elements 302A-302J (“audio elements 302”). The stream selection unit 44 can output the corresponding audio stream from the audio stream 27, which the audio decoding device 34 can decode and output as audio data 19'.

[0098] The stream selection unit 44 can also determine one or more proximity distances between the virtual position 300A and one or more (and possibly each) capture positions (or synthesis positions) represented by ALI. The stream selection unit 44 can then compare one or more proximity distances with a threshold proximity distance. When one or more proximity distances are greater than the threshold proximity distance, the stream selection unit 44 can select a smaller number of audio streams 27 to obtain audio data 19' compared to when one or more proximity distances are less than or equal to the threshold proximity distance. However, when one or more proximity distances are less than or equal to the threshold proximity distance, the stream selection unit 44 can select a larger number of audio streams 27 to obtain audio data 19' compared to when the proximity distances are greater than the threshold proximity distance.

[0099] In other words, the stream selection unit 44 can attempt to select those audio streams 27 that most closely align the audio data 19' with and surround the virtual position 300B. A proximity threshold can be defined such a threshold, which the user 1102 of the VR headset 1100 can set, or the stream selection unit 44 can dynamically determine the threshold based on the quality of the audio elements 302F-302J, the gain or loudness of the audio source 308, tracking information 41 (e.g., to determine whether the user 1102 is facing the audio source 308), or any other factor.

[0100] In this respect, when the listener is at position 300B, the stream selection unit 44 can increase the accuracy of audio spatialization. Furthermore, when the listener is at position 300A, the stream selection unit 44 can reduce the bit rate because it uses only the audio stream of audio element 302A instead of multiple audio streams of audio elements 302B-302J to reproduce the sound field.

[0101] Next reference Figure 3B For example, stream selection unit 44 can determine that the audio stream of audio element 302A is corrupted, noisy, or unavailable. Stream selection unit 44 can remove the audio stream from CM 47 and repeatedly iterate through audio stream 27 according to the techniques described in more detail above to select a single audio stream in audio stream 27 (e.g., Figure 3B (The audio stream of audio element 302B in the example), provided that the distance from the audio source to 306A is greater than the audio source distance threshold.

[0102] Next reference Figure 3CFor example, stream selection unit 44 can obtain a new audio stream (the audio stream of audio element 302K) and corresponding new information (e.g., metadata) including ALI 45A. Stream selection unit 44 can add the new audio stream to CM 47 representing audio stream 27. Stream selection unit 44 can iteratively repeat within audio stream 27 according to the techniques described in more detail above to select a single audio stream (e.g., ...) within audio stream 27. Figure 3C (The audio stream of element 302B in the example), provided that the distance of the audio source from 306A is greater than the audio source distance threshold.

[0103] exist Figure 3D In the examples, audio element 302 is replaced by specific example devices 320A-320J (“Device 320”), where Device 320A represents a dedicated microphone 320A, and Devices 320B, 320C, 320D, 320G, 320H, and 320J represent smartphones. Devices 320E, 320F, and 320I may represent VR devices. Each of Devices 320 may include a microphone that captures an audio stream 27 to be selected according to various aspects of the stream selection techniques described in this disclosure.

[0104] In many situations, some audio streams may be inappropriate or offensive to certain individuals. For example, during a live sporting event, someone might use offensive language within the venue. The same might be true in some video games. Sensitive discussions may occur at other live events, such as rallies. By using authorization levels, the stream selection unit 44 can filter or otherwise exclude unwanted or sensitive audio streams from playback for the user of the content consumer device 14. Authorization levels can be associated with individual audio streams or privacy zones (regarding...). Figure 4C (To be discussed in more detail).

[0105] Authorization levels can take several different forms. For example, authorization levels can be similar to the Motion Picture Association of America (MPAA) ratings, or they can be similar to security licenses.

[0106] Another way to implement authorization levels is based on a contact list. This contact list can contain multiple contacts and may also include one or more contact rank or ranking. One or more processors can store the contact list in memory within the content consumer device 14. In this example, an authorization level is satisfied if the content creator or content source is in the contact list on the content consumer device 14 (e.g., people listed in the content list). Otherwise, the authorization level is not satisfied. In another example, authorization levels can be based on rank. For example, authorization can occur when a contact has at least a predetermined rank or ranking.

[0107] In some cases, source device 12 can set authorization levels. For example, in a gathering where a sensitive discussion is taking place, the content creator or source can create and apply authorization levels so that only certain people with appropriate privileges can hear the information. For others without appropriate privileges, stream selection unit 44 can filter out or otherwise exclude one or more audio streams of the discussion.

[0108] In other cases, such as the sports event example, the content consumer device 14 can create authorization levels. Therefore, users can exclude offensive language during audio playback.

[0109] Figure 3E This is a conceptual diagram illustrating an example concert with three or more audio elements. Figure 3E In the example, multiple musicians are depicted on stage 323. Singer 312 is positioned behind audio element 310A. String section 314 is depicted behind audio element 310B. Drummer 316 is depicted behind audio element 310C. Other musicians 318 are depicted behind audio element 310D. Audio elements 310A-310D can include audio streams corresponding to sounds received by a microphone. In some examples, audio elements 310A-310D can represent synthesized audio streams. For example, audio element 310A can represent one or more audio streams primarily associated with singer 312, but this audio stream(s) may also include sounds produced by other band members, such as string section 314, drummer 316, or other musicians 318, while audio element 310B can represent one or more audio streams primarily associated with string section 314, but may also represent sounds produced by other band members. In this way, each of audio elements 310A-310D can represent different audio stream(s).

[0110] Multiple devices are also depicted. These devices represent user devices located at multiple different listening positions. Headphones 321 are located near audio element 310A, but between audio element 310A and audio element 310B. Therefore, according to the technology of this disclosure, the stream selection unit 44 can select at least one audio stream from the audio streams to generate an audio stream that is compatible with the listening position of the user of headphones 321. Figure 3E The user in the same position as the headset 321 experiences a similar audio experience. Similarly, the VR goggles 322 are shown positioned behind the audio element 310C and between the drummer 316 and other musicians 318. The stream selection unit 44 can select at least one audio stream to produce an audio experience for the user of the VR goggles 322 that is positioned similarly to the user in the headset 321. Figure 3E The user in the VR goggles 322 location has a similar audio experience.

[0111] The smart glasses 324 are shown positioned fairly centrally between audio elements 310A, 310C, and 310D. The stream selection unit 44 can select at least one audio stream to generate an audio stream for the user of the smart glasses 324 that is positioned within the audio elements 310A, 310C, and 310D. Figure 3E The user in the location of the smart glasses 324 experiences a similar audio experience. Furthermore, device 326 (which can represent any device capable of implementing the techniques of this disclosure, such as a mobile handheld device, speaker array, headset, VR goggles, smart glasses, etc.) is shown positioned in front of audio element 310B. Stream selection unit 44 can select at least one audio stream to generate an audio experience for the user of device 326 similar to that of the user in the location of the smart glasses 324. Figure 3E Users in the location of device 325 will have a similar audio experience. While specific devices are discussed regarding particular locations, any device depicted can provide a similar audio experience to... Figure 3E The instructions describe the different desired listening locations.

[0112] Figure 4A This illustrates the technology according to this disclosure. Figures 1A to 1C The flowchart illustrates an example of the operation of the stream selection unit. One or more processors of the content consumer device 14 may store multiple audio streams and an associated license level (350) for each audio stream in memory on the content consumer device 14. For example, audio streams may have associated license levels. In some examples, the license level may be directly associated with the audio stream. In some examples, the license level may be associated with a privacy zone associated with the audio stream, and the audio stream may be associated with the license level in this way. In some examples, multiple audio streams are stored in encoded form. In other examples, multiple audio streams are stored in decoded form.

[0113] In some examples, the authorization level of an audio stream is based on location information associated with the audio stream. For example, audio decoding device 34 may store in memory location information associated with the coordinates of the acoustic space in which a corresponding audio stream of a plurality of audio streams is captured or synthesized. The authorization level of each of the plurality of audio streams may be determined based on the location information. In some examples, multiple privacy zones may be defined in the acoustic space. Each privacy zone may have an associated authorization level. In some examples, the authorization level of at least one of the plurality of audio streams may be determined based on the authorization level of the privacy zone containing the location of at least one of the plurality of audio streams being captured or synthesized. In some examples, the authorization level of at least one of the plurality of audio streams is equal to the authorization level of the privacy zone containing the location of one of the audio streams being captured or synthesized.

[0114] One or more processors of content consumer device 14 can select a subset of multiple audio streams based on an associated authorization level (352) in a manner that excludes at least one audio stream from the plurality of audio streams. In some examples, the excluded streams are associated with one or more privacy zones. For example, user 1102 may not have authorization to listen to audio sources in one or more privacy zones, and one or more processors of stream selection unit 44 can exclude those audio streams from the subset of multiple audio streams. In some examples, one or more processors of stream selection unit can exclude audio streams from the subset of multiple audio streams by setting the audio streams to be excluded to null.

[0115] One or more processors of stream selection unit 44 can output a subset of multiple audio streams to one or more speakers or headphones (354). For example, one or more processors of stream selection unit 44 can output a subset of multiple audio streams to headphones 48.

[0116] In some examples, content consumer device 14 may receive authorization levels from source device 12. For example, the authorization level may be included in metadata associated with the audio stream, or it may be otherwise included in bitstream 27. In other examples, one or more processors of the content consumer device may generate the authorization levels. In some examples, one or more processors of stream selection unit 44 may compare the authorization level associated with each audio stream with the authorization level of the device or its user (e.g., user 1102), and select a subset of multiple audio streams based on the comparison result. In some examples, the authorization level includes more than two levels, rather than authorized or unauthorized. In such examples, one or more processors of stream selection unit 44 may select a subset of multiple audio sources by comparing the level of each of the multiple audio streams with the level of the user (e.g., user 1102), and selecting a subset of multiple audio streams based on that comparison. For example, the user's level may indicate that the user is not authorized to listen to audio streams with different levels, and stream selection unit 44 may not select such audio streams. In other examples, the authorization level may be based on multiple contacts, which may be stored in the memory of the content consumer device 14. In such examples, one or more processors of the stream selection unit 44 may select a subset of multiple audio streams by determining whether the source of one or more audio streams is associated with one or more contacts among the multiple contacts; and selecting a subset of the multiple audio streams based on this comparison. In some examples, the multiple contacts include a preference level ranking. In some examples, one or more processors of the stream selection unit 44 may select a subset of multiple audio sources by determining whether the source of one or more audio streams is associated with one or more contacts among the multiple contacts having at least a predetermined preference level ranking; and selecting a subset of the multiple audio streams based on this comparison. In some examples, when the privacy zone does not have an associated authorization level, the content consumer device 14 may suppress decoding of the audio stream associated with the privacy zone.

[0117] In some examples, content consumer device 14 may be configured to receive multiple audio streams and an associated license level for each audio stream. In some examples, content consumer device 14 may be configured to receive license levels of a device user. In some examples, content consumer device 14 may select a subset of multiple audio streams by selecting those audio streams having an associated license level no greater than the license level of the receiving device user, and send the selected subset of multiple audio streams to an audible output device (e.g., headphones 48) for audible output of the selected subset of multiple audio streams.

[0118] In some examples, a subset of the multiple audio streams includes a reproduced audio stream based on encoded information received in a bitstream (e.g., bitstream 27) decoded by one or more processors of the content consumer device 14. In other examples, the audio stream may be unencoded.

[0119] Figure 4A This illustrates the technology according to this disclosure. Figures 1A to 1C The flowchart shows another example of the operation of the stream selection unit (4000) shown in the example. One or more processors may store the audio stream and information associated with the audio stream, including location information and authorization level (400), in the memory of the content consumer device 14. In some examples, the information associated with the audio stream may be metadata. The stream selection unit 44 may obtain the location information (401). As described above, this location information may be associated with capture coordinates in acoustic space. In some examples, the stream selection unit 44 may obtain the location information by reading the location information from the memory, for example when the location information is associated with a specific audio stream, or in other examples, the stream selection unit 44 may obtain the location information by calculating the location information (if necessary), for example when the location information is associated with a privacy zone.

[0120] Authorization levels can be associated with each audio stream or privacy zone (which will be related to...) Figure 4D (For a more thorough discussion). For example, sensitive discussions may occur during live events, or inappropriate language may be used or inappropriate topics may be discussed for certain audience members. By assigning an authorization level to each audio stream or privacy zone, stream selection unit 44 can filter out relevant audio streams or otherwise exclude them so as not to reproduce them. Stream selection unit 44 can determine whether a stream is authorized for use by user 1102 (402). For example, stream selection unit 44 can determine whether an audio stream is authorized based on the authorization level associated with it (e.g., directly or by associating it with a privacy zone that has an associated authorization level). In some examples, the authorization level can be a tier, as discussed below with respect to Tables 1 and 2. In other examples, the authorization level can be based on a contact list. In examples where authentication is performed using a contact list, stream selection unit 44 can filter out or otherwise exclude (one or more) audio streams or privacy zones when the content creator or source is not in the contact list or does not have a sufficiently high favorability rating.

[0121] In one example, audio playback system 16 (for simplicity, it may refer to audio playback system 16A or audio playback system 16B) may allow a user to override an authorization level. Audio playback system 16 may receive a request from user 1102 to override at least one authorization level and determine whether to override that authorization level (404). When an authorization level is overridden, stream selection unit 44 may select or add an audio stream (403), and the audio stream or privacy zone may be included in the audio output. When an authorization level is not overridden, the corresponding audio stream or privacy zone may not be included in the output; for example, stream selection unit 44 will not select an audio stream (405). In some examples, some users may have the ability to override authorization levels, while others may not. For example, parents may have the ability to override authorization levels, while children may not. In some examples, super users may have the ability to override authorization levels, while ordinary users may not. In one example, audio playback system 16 may send a message to source device 12, instructing source device 12 or a base station to stop sending excluded audio stream(s)(409). In this way, bandwidth within the transmission channel can be saved.

[0122] When a user does not have sufficient authorization level for a given audio stream or privacy zone, the stream selection unit 44 can exclude (e.g., not select) that audio stream or privacy zone. In one example, the audio playback system 16 can change the gain based on the authorization level of the audio stream or privacy zone, thereby enhancing or attenuating the audio output (406). In some examples, the audio playback system 16 can empty or zero a given audio stream or privacy zone. The audio decoding device 34 can combine two or more selected audio streams together (407). For example, the combination of selected audio streams can be done by mixing or interpolation. The audio decoding device 34 can then output the selected stream (408).

[0123] Figure 4C and Figure 4D This diagram illustrates various aspects of the technology described in this disclosure regarding the privacy zone. A static audio source 441, such as an open microphone, is shown. The static audio source 441 can be a live audio source or a synthesized audio source. A dynamic audio source 442, for example, is set up by a user in a user-operated mobile handheld device while the audio source is being recorded, is also shown. The dynamic audio source 442 can be a live audio source or a synthesized audio source. One or more of the static audio source 441 and / or the dynamic audio source 442 can capture or synthesize audio information 443. The source device can send the audio information to a controller 444. The controller 444 can process the audio information. Figure 4C In the diagram, controller 444 is shown implemented in processor 449A, which may be located in content consumer device 14. Figure 4DIn this design, controller 444 is shown implemented in processor 450, which may be in source device 12A or 12B, rather than in processor 449B, which may be in content consumer device 14. For example, controller 444 may partition audio information into corresponding zones (e.g., privacy zones), create an audio stream, and label the audio stream with location information about the positions of audio sources 441 and 442, and, for example, with compartmentalization (including zone boundaries) using centroid and radius data. In some examples, the location information may be metadata. Controller 444 may perform these functions online or offline. Controller 444 may then send the location information to priority unit 445 via a separate link 452 and the audio stream to priority unit 445 via link 453, or it may send the location information and audio stream together via a single link.

[0124] In one example, priority unit 445 could be the place where authorization levels are created and assigned. For instance, priority unit 445 could determine which privacy regions' gains can be changed and which privacy regions can be left empty or excluded from rendering.

[0125] Covering unit 446 is shown. This coverage allows users to override the authorization level for a given privacy zone.

[0126] The content consumer device 14 can determine user location and orientation information 447, and use user location and orientation information 447, audio stream, location information, area boundaries and authorization level to create rendering 448.

[0127] Figure 4E and Figure 4FThis is a diagram further illustrating the concept of a privacy zone according to various aspects of this disclosure. User 460 is shown near several groups of audio elements, each group representing an audio stream. In some examples, it may be useful to authorize which audio streams are used to create a grouped rather than individual audio experience for user 460. In some examples, there may be multiple audio elements positioned close to each other. For example, in the example of a gathering, multiple audio elements positioned close to each other may be receiving sensitive information. Therefore, privacy zones can be created and the same authorization level can be assigned to each audio stream associated with a given privacy zone. In some examples, the authorization level may be directly associated with each audio stream in a given privacy zone. In other examples, the authorization level and the audio stream may be associated with the privacy zone. As used in this disclosure, when an authorization level is referred to as being associated with an audio stream, the authorization level may be directly associated with the audio stream or may be associated with the privacy zone associated with the audio stream. For example, when the memory can store multiple audio streams and the authorization level of each audio stream, the memory can store multiple audio streams and: 1) authorization levels of one or more privacy zones, and the association between the multiple audio streams and one or more privacy zones; 2) authorization levels of each audio stream; or 3) any combination thereof.

[0128] For example, source device 12 can assign a user an authorization level that can be a grade. Figure 4C and Figure 4D Prioritization unit 445 can assign gain, attenuation, and null information (e.g., metadata), and in this example, assign a level for each privacy zone. For example, privacy zone 461 can contain audio streams 4611, 4612, and 4613. Privacy zone 462 can contain audio streams 4621, 4622, and 4623. Privacy zone 463 can contain audio streams 4631, 4632, and 4633. As shown in Table 1, controller 444 can mark these audio streams as belonging to their respective privacy zones. Prioritization unit 445 can also associate gain and null information (e.g., metadata) with audio streams. As shown in Table 1, G is gain, and N is null or exclusion. In this example, user 460 has a level of 2 for privacy zones 461 and 463, but a level of 3 for privacy zone 462. As shown in the table, stream selection unit 44 will exclude or empty privacy region 462, and audio elements (or audio sources) within privacy region 462 (e.g., audio streams 4621-4623) will not be available for rendering, such as Figure 4E As shown, unless user 460 wants to override this authorization level, in which case audio rendering will be as follows: Figure 4F As shown in Table 1, although authorization levels are presented as grades, authorization levels can be implemented in other ways, such as based on a contact list.

[0129] district mark Metadata grade 461,463 4611-4613,4631-4633 G-20dB, N=0 2 462 4621-4623 G–N / A, N=1 3

[0130] Table 1

[0131] Figure 4G and Figure 4H This diagram illustrates the exclusion of individual audio streams instead of privacy zones. In this example, the audio streams are not clustered but rather spaced far apart, and controller 444 can label them individually, each with its own authorization level. For example, audio streams 471, 472, 473, and 474 may not contain overlapping information. In some examples, each of audio streams 471, 472, 473, and 474 may have a different authorization level. Referring to Table 2, in this example, controller 444 can label each of audio streams 471, 472, 473, and 474 with a separate authorization level and not assign any of them to a privacy zone. Prioritization unit 445 can assign gain and null information (e.g., metadata) to each audio stream. In one example, content consumer device 14 can assign a level to each user of each audio stream. As can be seen from Table 2, stream selection unit 44 can nullify audio stream 474 or exclude it from rendering for user 470, such as... Figure 4G As shown, unless user 470 overrides this priority, user 470's rendering will be as follows. Figure 4H As shown in the image. In other examples, a contact list instead of a rank can be used as the authorization level, as described above.

[0132] district mark Metadata grade N / A 471 G = 0 dB, N = 0 2 N / A 472 G = 0 dB, N = 0 2 N / A 473 G = 0 dB, N = 0 2 N / A 474 G = N / A, N = 1 1

[0133] Table 2

[0134] Figure 5This is a diagram illustrating an example of a wearable device 500 that can operate according to various aspects of the techniques described in this disclosure. In various examples, wearable device 500 may represent a VR headset (such as VR device 1100 described above), an AR headset, a MR headset, or any other type of extended reality (XR) headset. Augmented Reality “AR” may refer to computer-rendered images or data overlaid on the real world in which the user actually resides. Mixed Reality “MR” may refer to computer-rendered images or data locked to a specific location in the real world, or may refer to a variant of VR in which a combination of computer-rendered 3D elements and filmed real elements is used to create an immersive experience that simulates the user’s physical presence in the environment. Extended Reality “XR” may refer to VR, AR, and MR collectively. More information on the XR terminology can be found in Jason Peterson’s document entitled “Virtual Reality, Augmented Reality, and Mixed Reality Definitions” dated July 7, 2017.

[0135] Wearable device 500 can refer to other types of devices, such as watches (including so-called "smartwatches"), glasses (including so-called "smart glasses"), headphones (including so-called "wireless headphones" and "smart headphones"), smart clothing, smart jewelry, etc. Whether referring to VR devices, watches, glasses, and / or headphones, wearable device 500 can communicate with computing devices that support wearable device 500 via wired or wireless connections.

[0136] In some cases, the computing device supporting the wearable device 500 may be integrated within the wearable device 500, and therefore, the wearable device 500 may be considered the same device as the computing device supporting the wearable device 500. In other cases, the wearable device 500 may communicate with a separate computing device that can support the wearable device 500. In this regard, the term "support" should not be construed as requiring a separate dedicated device, but rather as meaning that one or more processors configured to perform various aspects of the technologies described in this disclosure may be integrated within the wearable device 500 or integrated within a computing device separate from the wearable device 500.

[0137] For example, when wearable device 500 refers to VR device 1100, a separate dedicated computing device (such as a personal computer including one or more processors) can render audio and visual content, while wearable device 500 can determine translational head movements. In this case, the dedicated computing device can render audio content (as a speaker feed) based on the translational head movements according to various aspects of the technology described in this disclosure. As another example, when wearable device 500 refers to smart glasses, wearable device 500 may include one or more processors (via interfaces within one or more sensors of wearable device 500) that determine translational head movements and render speaker feeds based on the determined translational head movements.

[0138] As shown in the figure, wearable device 500 includes one or more directional speakers and one or more tracking and / or recording cameras. Additionally, wearable device 500 includes one or more inertial, haptic, and / or health sensors, one or more eye-tracking cameras, one or more high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of wearable device 500 may include durable translucent display technology and hardware.

[0139] Wearable device 500 also includes connectivity hardware, which may represent one or more network interfaces supporting multi-mode connectivity such as 4G communication, 5G communication, Bluetooth, etc. Wearable device 500 also includes one or more ambient light sensors, one or more cameras and night vision sensors, and one or more bone conduction sensors. In some cases, wearable device 500 may also include one or more passive and / or active cameras with fisheye lenses and / or telephoto lenses. Although in Figure 5 Not shown, but wearable device 500 may also include one or more light-emitting diode (LED) lights. In some examples, the LED lights may be referred to as "super bright" LED lights. In some embodiments, wearable device 500 may also include one or more rear cameras. It should be understood that wearable device 500 may exhibit a variety of different form factors.

[0140] Furthermore, tracking and recording cameras, along with other sensors, can facilitate the determination of translational distance. Although not in Figure 5 The example shown is not applicable, but the wearable device 500 may include other types of sensors for detecting translational distances.

[0141] Although specific examples of wearable devices have been described, such as those mentioned above... Figure 2 The examples discussed include VR devices 1100 and Figures 1A to 1C Other devices illustrated in the examples, but those skilled in the art will understand that, with Figures 1A to 1Cand Figure 2 The related description can be applied to other examples of wearable devices. For example, other wearable devices such as smart glasses may include sensors through which translational head movements are obtained. As another example, other wearable devices such as smartwatches may include sensors through which translational movements are obtained. Therefore, the techniques described in this disclosure should not be limited to a particular type of wearable device, but any wearable device can be configured to perform the techniques described in this disclosure.

[0142] Figure 6A and Figure 6B This is a diagram illustrating an example system capable of performing various aspects of the techniques described in this disclosure. Figure 6A An example is shown in which the source device 12C also includes a camera 600. The camera 600 can be configured to capture video data and provide the captured raw video data to the content capture device 20. The content capture device 20 can provide the video data to another component of the source device 12C for further processing into viewport-divided portions.

[0143] exist Figure 6A In the example, content consumer device 14C also includes VR device 1100. It will be understood that, in various specific embodiments, VR device 1100 may be included in or coupled externally to content consumer device 14C. VR device 1100 includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewports) and for rendering audio data.

[0144] Figure 6B An example is shown, where Figure 6A The audio renderer 32 shown is replaced by a binaural renderer 42 capable of performing binaural rendering using one or more HRTFs or other functions capable of rendering to the left and right speaker feeds 43. The audio playback system 16C of the content consumer device 14D can output the left and right speaker feeds 43 to headphones 48.

[0145] The headphones 48 can be connected via wired connection (such as a standard 3.5mm audio jack, Universal System Bus (USB) connection, optical audio jack, or other forms of wired connection) or wirelessly (such as via Bluetooth). TM(Connectivity, wireless network connectivity, etc.) is coupled to the audio playback system 16B. The headphones 48 can recreate the sound field represented by the audio data 19' based on the left and right speaker feeds 43. The headphones 48 may include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the corresponding left and right speaker feeds 43. It should be noted that the content consumer device 14C and content consumer device 14D can be coupled with… Figure 1C Use it together with the source device 12B.

[0146] Figure 7 It is shown Figures 1A to 1C The example shown is a block diagram of one or more sample components from the source device and content consumer device. Figure 7 In some examples, device 710 includes a processor 712 (which may be referred to as "one or more processors" or "(one or more) processors"), a graphics processing unit (GPU) 714, system memory 716, a display processor 718, one or more integrated speakers 740, a display 703, a user interface 720, an antenna 721, and a transceiver module 722. In examples where device 710 is a mobile device, the display processor 718 is a mobile display processor (MDP). In some examples, such as in the example where device 710 is a mobile device, processor 712, GPU 714, and display processor 718 may be configured as integrated circuits (ICs).

[0147] For example, an IC can be considered a processing chip within a chip package and can be a system-on-a-chip (SoC). In some examples, two of the processor 712, GPU 714, and display processor 718 can be housed together in the same IC, while the third can be housed in a different integrated circuit (e.g., a different chip package), or all three can be housed in different ICs or on the same IC. However, in the example where device 710 is a mobile device, the processor 712, GPU 714, and display processor 718 can all be housed in different integrated circuits.

[0148] Examples of processor 712, GPU 714, and display processor 718 include, but are not limited to, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Processor 712 may be the central processing unit (CPU) of device 710. In some examples, GPU 714 may be specialized hardware that includes integrated and / or discrete logic circuitry to provide GPU 714 with massively parallel processing capabilities suitable for graphics processing. In some cases, GPU 714 may also include general-purpose processing capabilities and may be referred to as a general-purpose GPU (GPGPU) when implementing general-purpose processing tasks (e.g., non-graphics-related tasks). Display processor 718 may also be specialized integrated circuit hardware designed to retrieve image content from system memory 716, combine the image content into image frames, and output the image frames to display 703.

[0149] Processor 712 can execute various types of applications. Examples of applications include web browsers, email applications, spreadsheets, video games, other applications that generate visual objects for display, or any of the application types listed above in more detail. System memory 716 can store instructions for executing applications. Executing an application on processor 712 causes processor 712 to generate graphical data of image content to be displayed and audio data 19 to be played (possibly via integrated speaker 740). Processor 712 can transfer the graphical data of the image content to GPU 714 for further processing based on the instructions or commands transferred from processor 712 to GPU 714.

[0150] Processor 712 can communicate with GPU 714 according to a specific application processing interface (API). Examples of such APIs include... of API, Khronos Group or and OpenCL TM However, aspects of this disclosure are not limited to DirectX, OpenGL, or OpenCL APIs and can be extended to other types of APIs. Furthermore, the techniques described in this disclosure do not require an API to function, and the processor 712 and GPU 714 can communicate using any process.

[0151] System memory 716 may be a memory used in device 710. System memory 716 may include one or more computer-readable storage media. Examples of system memory 716 include, but are not limited to, random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other media that may be used to carry or store desired program code in the form of instructions and / or data structures that can be accessed by a computer or processor.

[0152] In some examples, system memory 716 may include instructions that cause processor 712, GPU 714, and / or display processor 718 to perform the functions attributed to processor 712, GPU 714, and / or display processor 718 in this disclosure. Therefore, system memory 716 may be a computer-readable storage medium on which instructions are stored, which, when executed, cause one or more processors (e.g., processor 712, GPU 714, and / or display processor 718) to perform various functions.

[0153] System memory 716 may include a non-transitory storage medium. The term "non-transitory" means that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that system memory 716 is immovable or that its contents are static. As an example, system memory 716 can be removed from device 710 and moved to another device. As another example, a memory substantially similar to system memory 716 can be inserted into device 710. In some examples, the non-transitory storage medium may store data that can change over time (e.g., in RAM).

[0154] User interface 720 may represent one or more hardware or virtual (meaning a combination of hardware and software) user interfaces through which a user interacts with device 710. User interface 720 may include physical buttons, switches, toggle keys, lights, or virtual versions thereof. User interface 720 may also include physical or virtual keyboards, touch interfaces (such as touchscreens), haptic feedback, etc.

[0155] Processor 712 may include one or more hardware units (including so-called "processing cores") configured to perform all or part of the operations discussed above regarding any module, unit, or other functional component of the content creator device and / or content consumer device. Antenna 721 and transceiver module 722 may represent units configured to establish and maintain a connection between source device 12 and content consumer device 14. Antenna 721 and transceiver module 722 may represent capabilities according to one or more wireless communication protocols (such as fifth-generation (5G) cellular standards, such as Bluetooth). TMA transceiver module 722 is a receiver and / or a transmitter that communicates wirelessly using a Personal Area Network (PAN) protocol, or other open-source, proprietary, or other communication standards. For example, transceiver module 722 can receive and / or transmit wireless signals. Transceiver module 722 can represent a separate transmitter, a separate receiver, both a separate transmitter and a separate receiver, or a combination of transmitter and receiver. Antenna 721 and transceiver module 722 can be configured to receive encoded audio data. Similarly, antenna 721 and transceiver module 722 can be configured to transmit encoded audio data.

[0156] Figures 8A to 8C It is shown Figures 1A to 1C The flowchart shown in the example illustrates the operation of the stream selection unit 44 when performing various aspects of the stream selection technique. First, refer to... Figure 8A For example, the stream selection unit 44 can obtain audio stream 27 from all enabled audio elements (also called receivers), where audio stream 27 may include corresponding information (e.g., metadata), such as ALI 45A (800). The stream selection unit 44 can perform energy analysis with respect to each audio stream 27 to calculate the corresponding energy map (802).

[0157] The stream selection unit 44 can then iterate (804) among different combinations of audio elements (defined in CM 47) based on proximity to audio source 308 (as defined by audio source distances 306A and / or 306B) and proximity to audio elements (as defined by the proximity distances discussed above), and this process can return to 802. Figure 8A As shown, audio elements can be sorted or otherwise associated with different access permissions. Stream selection unit 44 can iterate as described above based on the listener location represented by DLI 45B (which is another way of referring to "virtual location" or "device location") and the audio element location represented by ALI 45A to identify whether a larger subset of audio stream 27 or a smaller subset of audio stream 27 is needed (806, 808).

[0158] When a larger subset of the audio stream 27 is needed, the stream selection unit 44 can add (one or more) audio elements to the audio data 19', or in other words, add additional audio streams (such as when the user is closer). Figure 3A (810) When a subset of the audio stream 27 needs to be reduced, the stream selection unit 44 can remove (one or more) audio elements from the audio data 19', or in other words, remove (one or more) existing audio streams (such as when the user leaves). Figure 3A (812) When the audio source in the example is far away.

[0159] In some examples, the stream selection unit 44 can determine that the current constellation of the audio element is the optimal set (or, in other words, the existing audio data 19' will remain the same because the selection process described herein results in the same audio data 19') (804), and the process can return to 802. However, when an audio stream is added to or removed from audio data 19', the stream selection unit 44 can update CM 47 (814) to generate a constellation history (815) (including positions, energy maps, etc.).

[0160] Additionally, the stream selection unit 44 can determine whether the privacy setting enables or disables the addition of audio elements (wherein the privacy setting may refer to digital access permissions that restrict access to one or more audio streams in audio stream 27, for example, via password, authorization level or grade, time, etc.) (816, 818). When the privacy setting enables the addition of audio elements, the stream selection unit 44 can add (one or more) audio elements to the updated CM 47 (which refers to adding (one or more) audio streams to audio data 19') (820). When the privacy setting disables the addition of audio elements, the stream selection unit 44 can remove (one or more) audio elements from the updated CM 47 (which refers to removing (one or more) audio streams from audio data 19') (822). In this way, the stream selection unit 44 can identify a new set of enabled audio elements (824).

[0161] The stream selection unit 44 can iterate and update various inputs according to any given frequency in this manner. For example, the stream selection unit 44 can update privacy settings at the user interface rate (meaning updates are driven by updates via user interface input). As another example, the stream selection unit 44 can update position at the sensor rate (meaning position changes by movement of audio elements). The stream selection unit 44 can also update the energy map at the audio frame rate (meaning the energy map is updated every frame).

[0162] Next reference Figure 8B For example, besides the flow selection unit 44 being able to determine CM 47 without relying on the energy map, the flow selection unit 44 can be as described above regarding... Figure 8A The operation is described in the manner described. Therefore, the stream selection unit 44 can obtain audio stream 27 from all enabled audio elements, where audio stream 27 may include corresponding information (e.g., metadata), such as ALI45A (840). The stream selection unit 44 can determine whether privacy settings are enabled or disabled for the addition of audio elements (where privacy settings may refer to digital access permissions that restrict access to one or more of the audio streams 27, e.g., via password, authorization level or grade, time, etc.) (842, 844).

[0163] When the privacy setting enables the addition of audio elements, the stream selection unit 44 can add one or more audio elements to the updated CM 47 (which means adding one or more audio streams to audio data 19') (846). When the privacy setting disables the addition of audio elements, the stream selection unit 44 can remove one or more audio elements from the updated CM 47 (which means removing one or more audio streams from audio data 19') (848). In this way, the stream selection unit 44 can identify a new set of enabled audio elements (850). The stream selection unit 44 can iterate over different combinations of audio elements in the CM 47 (852) to determine the constellation history representing the audio data 19' (854).

[0164] The stream selection unit 44 can iterate and update various inputs at any given frequency in this manner. For example, the stream selection unit 44 can update privacy settings at the user interface rate (meaning updates are driven by updates via user interface input). As another example, the stream selection unit 44 can update position at the sensor rate (meaning position is changed by movement of audio elements).

[0165] Next reference Figure 8C For example, besides the stream selection unit 44 being able to determine CM 47 without relying on the audio element with privacy settings enabled, the stream selection unit 44 can, as mentioned above, determine CM 47. Figure 8A The operation is described in the manner described. Therefore, the stream selection unit 44 can obtain audio streams 27 from all enabled audio elements, where each audio stream 27 may include corresponding information (e.g., metadata), such as ALI 45A (860). The stream selection unit 44 can perform energy analysis with respect to each audio stream 27 to calculate the corresponding energy map (862).

[0166] The stream selection unit 44 can then iterate (864) on different combinations of audio elements (defined in CM 47) based on proximity to audio source 308 (as defined by audio source distances 306A and / or 306B) and proximity to audio elements (as defined by the proximity distances discussed above), and this process can return to 862. Figure 8C As shown, audio elements can be sorted or otherwise associated with different access permissions. Stream selection unit 44 can iterate in the manner described above based on the listener location represented by DLI 45B (which also refers to another way of the "virtual location" or "device location" discussed above) and the audio element location represented by ALI 45A to identify whether a larger subset of audio stream 27 or a smaller subset of audio stream 27 (866, 868) is needed.

[0167] When a larger subset of the audio stream 27 is needed, the stream selection unit 44 can add one or more audio elements to the audio data 19', or in other words, add additional one or more audio streams (such as when the user is closer). Figure 3A (e.g., in the example of the audio source) (870). When a subset of the audio stream 27 needs to be reduced, the stream selection unit 44 can remove (one or more) audio elements from the audio data 19', or in other words, remove (one or more) existing audio streams (such as when the user leaves). Figure 3A (872) When the audio source in the example is far away.

[0168] In some examples, the stream selection unit 44 can determine that the current constellation of the audio element is the optimal set (or, in other words, the existing audio data 19' will remain the same because the selection process described herein results in the same audio data 19') (864), and the process can return to 862. However, when an audio stream is added to or removed from audio data 19', the stream selection unit 44 can update CM 47 (874) to generate a constellation history (875).

[0169] The stream selection unit 44 can iterate and update various inputs according to any given frequency in this manner. For example, the stream selection unit 44 can update the position at the sensor rate (meaning changing the position by moving the audio element). The stream selection unit 44 can also update the energy map at the audio frame rate (meaning the energy map is updated every frame).

[0170] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for technical practice). Moreover, in some examples, actions or events may be performed concurrently, for example, through multithreaded processing, interrupt handling, or multiple processors, rather than sequentially.

[0171] In some examples, the VR device (or streaming device) may use a network interface coupled to the VR / streaming device's memory to exchange messages with external devices, where the exchanged messages are associated with multiple available representations of the sound field. In some examples, the VR device may use an antenna coupled to the network interface to receive wireless signals, which include data packets, audio packets, video packets, or transport protocol data associated with multiple available representations of the sound field. In some examples, one or more microphone arrays may capture the sound field.

[0172] In some examples, multiple available representations of a sound field stored in a memory device may include multiple object-based representations of the sound field, higher-order surround sound representations of the sound field, mixed-order surround sound representations of the sound field, a combination of object-based representations of the sound field and higher-order surround sound representations of the sound field, a combination of object-based representations of the sound field and mixed-order surround sound representations of the sound field, or a combination of mixed-order representations of the sound field and higher-order surround sound representations of the sound field.

[0173] In some examples, one or more of the multiple available representations of the sound field may include at least one high-resolution region and at least one low-resolution region, wherein the selected representation based on the steering angle provides greater spatial accuracy with respect to at least one high-resolution region and less spatial accuracy with respect to the low-resolution region.

[0174] Figure 9 An example of a wireless communication system 100 supporting privacy restrictions according to various aspects of this disclosure is shown. The wireless communication system 100 includes a base station 105, a user interface device (UE) 115, and a core network 130. In some examples, the wireless communication system 100 may be a Long Term Evolution (LTE) network, an Advanced LTE (LTE-A) network, an LTE-A Pro network, a fifth-generation (5G) cellular network, or a New Radio (NR) network. In some cases, the wireless communication system 100 may support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low-latency communication, or communication with low-cost and low-complexity devices.

[0175] Base station 105 can wirelessly communicate with UE 115 via one or more base station antennas. Base station 105 described herein may include, or may be referred to by those skilled in the art as, a base station transceiver, radio base station, access point, radio transceiver, NodeB, eNodeB (eNB), next-generation NodeB or giga-NodeB (any of which may be referred to as gNB), home NodeB, home eNodeB, or other suitable terms. Wireless communication system 100 may include different types of base stations 105 (e.g., macro cell base stations or small cell base stations). UE 115 described herein can communicate with various types of base stations 105 and network devices, including macro eNBs, small cell eNBs, gNBs, relay base stations, etc.

[0176] Each base station 105 may be associated with a specific geographic coverage area 110 in which communication with various UEs 115 is supported. Each base station 105 may provide communication coverage to the corresponding geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 may utilize one or more carriers. The communication link 125 shown in the wireless communication system 100 may include uplink transmission from the UE 115 to the base station 105, or downlink transmission from the base station 105 to the UE 115. Downlink transmission may also be referred to as forward link transmission, and uplink transmission may also be referred to as reverse link transmission.

[0177] The geographic coverage area 110 of base station 105 can be divided into sectors that constitute part of the geographic coverage area 110, and each sector can be associated with a cell. For example, each base station 105 can provide communication coverage for macro cells, small cells, hotspots, or other types of cells, or various combinations thereof. In some examples, base station 105 can be mobile and thus provide communication coverage for mobile geographic coverage areas 110. In some examples, different geographic coverage areas 110 associated with different technologies can overlap, and overlapping geographic coverage areas 110 associated with different technologies can be supported by the same base station 105 or different base stations 105. Wireless communication system 100 can include, for example, heterogeneous LTE / LTE-A / LTE-APro, 5G cellular, or NR networks, wherein different types of base stations 105 provide coverage for individual geographic coverage areas 110.

[0178] UE 115 may be distributed throughout the wireless communication system 100, and each UE 115 may be fixed or mobile. UE 115 may also be referred to as a mobile device, wireless device, remote device, handheld device, or subscriber device, or some other suitable term, wherein "device" may also be referred to as a unit, station, terminal, or client. UE 115 may also be a personal electronic device, such as a cellular phone, personal digital assistant (PDA), tablet computer, laptop computer, or personal computer. In the examples of this disclosure, UE 115 may be any audio source described in this disclosure, including VR headsets, XR headsets, AR headsets, vehicles, smartphones, microphones, microphone arrays, or any other device including a microphone or capable of transmitting captured and / or synthesized audio streams. In some examples, the synthesized audio stream may be an audio stream stored in memory or previously created or synthesized. In some examples, UE 115 may also refer to a wireless local loop (WLL) station, Internet of Things (IoT) device, Internet of Everything (IoE) device, or machine-type communication (MTC) device, which may be implemented in various articles of manufacture such as appliances, vehicles, meters, etc.

[0179] Some UEs 115, such as MTC or IoT devices, can be low-cost or low-complexity devices and can provide automated communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC can refer to data communication technologies that allow devices to communicate with each other or with base station 105 without human intervention. In some examples, M2M communication or MTC may include communication from devices that exchange and / or use information indicating privacy zones and / or authorization levels (e.g., metadata), affecting the gain and / or neutralization of various audio streams and / or audio sources, such as regarding... Figures 4A to 4H As described.

[0180] In some cases, UE 115 may also be able to communicate directly with other UE 115s (e.g., using point-to-point (P2P) or device-to-device (D2D) protocols). One or more UEs in a group of UEs 115s utilizing D2D communication may be within the geographic coverage area 110 of base station 105. Other UEs 115s in such a group may be outside the geographic coverage area 110 of base station 105, or otherwise unable to receive transmissions from base station 105. In some cases, a group of UEs 115s communicating via D2D communication may utilize a one-to-many (1:M) system, where each UE 115 sends to every other UE 115 in the group. In some cases, base station 105 facilitates the scheduling of resources for D2D communication. In other cases, D2D communication is performed between UEs 115s without the involvement of base station 105.

[0181] Base station 105 can communicate with core network 130 and with each other. For example, base station 105 can interface with core network 130 via backhaul link 132 (e.g., via S1, N2, N3, or other interfaces). Base stations 105 can communicate with each other directly (e.g., directly between base stations 105) or indirectly (e.g., via core network 130) via backhaul link 134 (e.g., via X2, Xn, or other interfaces).

[0182] In some cases, wireless communication system 100 may utilize licensed and unlicensed radio spectrum bands. For example, wireless communication system 100 may use Licensed Assisted Access (LAA), LTE Unlicensed (LTE-U) radio access technology, or NR technology in unlicensed bands such as the 5 GHz Industrial, Scientific, and Medical (ISM) band. When operating in unlicensed radio spectrum bands, wireless devices such as base station 105 and UE 115 may employ a pre-talk listen-to (LBT) procedure to ensure that the frequency channel is open before transmitting data. In some cases, operation in unlicensed bands may be based on a combination of carrier aggregation configuration and component carriers operating in licensed bands (such as LAA). Operation in unlicensed spectrum may include downlink transmission, uplink transmission, point-to-point transmission, or a combination thereof. Duplexing in unlicensed spectrum may be based on Frequency Division Duplex (FDD), Time Division Duplex (TDD), or a combination of both.

[0183] This disclosure includes the following examples.

[0184] Example 1. An apparatus configured to play one or more audio streams from a plurality of audio streams, the apparatus comprising: a memory configured to store a plurality of audio streams and corresponding audio metadata including an authorization level for each audio stream, and positioning information associated with the coordinates of an acoustic space in which a corresponding audio stream from the plurality of audio streams is captured; and one or more processors coupled to the memory and configured to: select a subset of the plurality of audio streams based on the authorization level and positioning information in the audio metadata, the subset of the plurality of audio streams excluding at least one audio stream from the plurality of audio streams.

[0185] Example 2. The device according to Example 1, wherein one or more processors are further configured to obtain location information.

[0186] Example 3. The device according to Example 2, wherein one or more processors obtain location information by reading location information from memory.

[0187] Example 4. The device according to Example 2, wherein the excluded stream is associated with one or more privacy zones, and one or more processors obtain location information by determining location information.

[0188] Example 5. The device according to any combination of Examples 1 to 4, wherein one or more processors are further configured to output a subset of multiple audio streams to one or more speakers.

[0189] Example 6. A device according to any combination of Examples 1 to 5, wherein one or more processors are further configured to change the gain of one or more audio streams in a subset of multiple audio streams based on an authorization level in the audio metadata.

[0190] Example 7. A device according to any combination of Examples 1 to 6, wherein one or more microprocessors are further configured to determine a privacy zone based on the authorization level in the audio metadata and location information associated with the coordinates of the acoustic space.

[0191] Example 8. The device according to Example 7, wherein one or more microprocessors are further configured to determine the privacy region by obtaining the privacy region.

[0192] Example 9. A device according to any combination of Examples 1 to 6, wherein one or more microprocessors are further configured to generate a privacy zone based on the authorization level in the audio metadata and location information associated with the coordinates of the acoustic space.

[0193] Example 10. The device according to any combination of Examples 1 to 9, wherein one or more microprocessors are further configured to send a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0194] Example 11. The device according to any combination of Examples 1 to 10, wherein one or more processors are further configured to combine at least two audio streams from a subset of multiple audio streams.

[0195] Example 12. The device according to Example 11, wherein one or more processors combine at least two audio streams from a subset of multiple audio streams by at least one of mixing or interpolation.

[0196] Example 13. A device according to any combination of Examples 1 to 12, wherein one or more processors are further configured to override the authorization level in the audio metadata.

[0197] Example 14. The device according to Example 13, wherein one or more processors are configured to output multiple audio streams to one or more speakers based on the authorization level being overridden in the audio metadata.

[0198] Example 15. The device described according to any combination of Examples 1 to 14, wherein the authorization level in the audio metadata is received from the source device.

[0199] Example 16. A device according to any combination of Examples 1 to 15, wherein one or more processors are further configured to generate the authorization level in audio metadata.

[0200] Example 17. The device according to any combination of Examples 1 to 16 further includes a display device.

[0201] Example 18. The device according to Example 17 further includes a microphone, wherein one or more processors are further configured to receive voice commands from the microphone and control the display device based on the voice commands.

[0202] Example 19. The device according to any combination of Examples 1 to 18 further includes one or more speakers.

[0203] Example 20. A device according to any combination of Examples 1 to 19, wherein the device includes extended reality headphones, and wherein the acoustic space includes a scene represented by video data captured by a camera.

[0204] Example 21. The device according to any combination of Examples 1 to 19, wherein the device includes extended reality headphones and wherein the acoustic space includes a virtual world.

[0205] Example 22. The device according to any combination of Examples 1 to 21 further includes a head-mounted device configured to present an acoustic space.

[0206] Example 23. The device according to any combination of Examples 1 to 19, wherein the device includes a mobile handheld device.

[0207] Example 24. The device according to any combination of Examples 1 to 23 further includes a wireless transceiver coupled to one or more processors and configured to receive wireless signals.

[0208] Example 25. The device according to Example 24, wherein the wireless signal is Bluetooth.

[0209] Example 26. The device according to Example 24, wherein the wireless signal conforms to the fifth generation (5G) cellular protocol.

[0210] Example 27. A method for playing one or more audio streams of a plurality of audio streams, comprising: storing, by means of a memory, a plurality of audio streams and corresponding audio metadata including an authorization level for each audio stream, and positioning information associated with coordinates of an acoustic space in which a corresponding audio stream of the plurality of audio streams is captured; and by means of one or more processors and based on the authorization level and positioning information of the audio metadata, the subset of the plurality of audio streams excluding at least one audio stream of the plurality of audio streams.

[0211] Example 28. The method according to Example 27 further includes obtaining location information by one or more processors.

[0212] Example 29. The method according to Example 28, wherein the location information is obtained by reading the location information from the memory.

[0213] Example 30. The method according to Example 28, wherein the location information is obtained by determining the location information, and wherein the excluded stream is associated with one or more privacy zones.

[0214] Example 31. The method according to any combination of Examples 27 to 30 further includes outputting a subset of the multiple audio streams to one or more speakers by one or more processors.

[0215] Example 32. The method according to any combination of Examples 27 to 31 further includes changing the gain of one or more audio streams in a subset of multiple audio streams based on the authorization level in the audio metadata by one or more processors.

[0216] Example 33. The method according to any combination of Examples 27 to 32 further includes determining a privacy zone by one or more processors based on the authorization level in the audio metadata and location information associated with the coordinates of the acoustic space.

[0217] Example 34. The method described in Example 33, wherein the privacy region is determined by obtaining the privacy region.

[0218] Example 35. The method according to any combination of Examples 27 to 34 further includes generating a privacy region by one or more processors based on the authorization level in the audio metadata and location information associated with the coordinates of the acoustic space.

[0219] Example 36. The method according to any combination of Examples 27 to 35 further includes sending a signal from one or more processors to at least one of the source device or base station instructing the transmission of the excluded audio stream to stop.

[0220] Example 37. The method according to any combination of Examples 27 to 36 further includes combining at least two audio streams from a subset of multiple audio streams by one or more processors.

[0221] Example 38. The method according to Example 37, wherein at least two audio streams in a subset of multiple audio streams are combined by at least one of mixing or interpolation.

[0222] Example 39. The method according to any combination of Examples 27 to 38 further includes having one or more processors override the authorization level in the audio metadata.

[0223] Example 40. The method according to Example 39 further includes outputting multiple audio streams to one or more speakers by one or more processors based on the authorization level being overridden in the audio metadata.

[0224] Example 41. The method according to any combination of Examples 27 to 40 further includes receiving an authorization level from the source device.

[0225] Example 42. The method according to any combination of Examples 27 to 41 further includes generating an authorization level by one or more processors.

[0226] Example 43. The method according to any combination of Examples 27 to 42 further includes receiving a voice command from a microphone and controlling a display device based on the voice command.

[0227] Example 44. The method according to any combination of Examples 27 to 43 further includes outputting a subset of the multiple audio streams to one or more speakers.

[0228] Example 45. The method according to any combination of Examples 27 to 44, wherein the method is performed on extended reality headphones, and wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0229] Example 46. The method according to any combination of Examples 27 to 45, wherein the method is performed on extended reality headphones, and wherein the acoustic space includes a virtual world.

[0230] Example 47. The method according to any combination of Examples 27 to 46, wherein the method is performed on a head-mounted device configured to present an acoustic space.

[0231] Example 48. The method according to any combination of Examples 27 to 47, wherein the method is performed on a mobile handheld device.

[0232] Example 49. The method according to any combination of Examples 27 to 48 further includes receiving a wireless signal.

[0233] Example 50. The method described in Example 49, wherein the wireless signal is Bluetooth.

[0234] Example 51. The method described in Example 49, wherein the wireless signal conforms to the fifth-generation (5G) cellular protocol.

[0235] Example 52. An apparatus configured to play one or more audio streams of a plurality of audio streams, the apparatus comprising: means for storing the plurality of audio streams and corresponding audio metadata including an authorization level for each audio stream, and positioning information associated with the coordinates of an acoustic space in which a corresponding audio stream of the plurality of audio streams is captured; and means for selecting a subset of the plurality of audio streams based on the authorization level and positioning information in the audio metadata, the subset of the plurality of audio streams excluding at least one audio stream of the plurality of audio streams.

[0236] Example 53. The device according to Example 52 further includes a component for obtaining positioning information.

[0237] Example 54. The device according to Example 53, wherein the location information is obtained by reading the location information from a memory.

[0238] Example 55. The device according to Example 53, wherein the location information is obtained by determining the location information, and wherein the excluded stream is associated with one or more privacy zones.

[0239] Example 56. The device according to any combination of Examples 52 to 55 further includes a component for outputting a subset of a plurality of audio streams to one or more speakers.

[0240] Example 57. The device according to any combination of Examples 52 to 56 further includes a component for changing the gain of one or more audio streams in a subset of a plurality of audio streams based on an authorization level in the audio metadata.

[0241] Example 58. The device according to any combination of Examples 52 to 57 further includes components for determining a privacy zone based on the authorization level in the audio metadata and location information associated with the coordinates of the acoustic space.

[0242] Example 59. The device according to Example 58, wherein the privacy region is determined by obtaining the privacy region.

[0243] Example 60. The device according to any combination of Examples 52 to 59 further includes components for generating a privacy zone based on the authorization level in the audio metadata and location information associated with the coordinates of the acoustic space.

[0244] Example 61. The device according to any combination of Examples 52 to 60 further includes a component for transmitting a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0245] Example 62. The device according to any combination of Examples 52 to 61 further includes a component for combining at least two audio streams from a subset of the plurality of audio streams.

[0246] Example 63. The device according to Example 62, wherein at least two audio streams from a subset of a plurality of audio streams are combined by at least one of mixing or interpolation.

[0247] Example 64. The device according to any combination of Examples 52 to 63 further includes a component for overriding the authorization level in the audio metadata.

[0248] Example 65. The device according to Example 64 further includes a component for outputting multiple audio streams to one or more speakers based on the authorization level being overridden in the audio metadata.

[0249] Example 66. The device according to any combination of Examples 52 to 65 further includes a component for receiving an authorization level from a source device.

[0250] Example 67. The device according to any combination of Examples 52 to 66 further includes components for generating license levels.

[0251] Example 68. A device according to a combination of Examples 52 to 67, further comprising a component for receiving voice commands from a microphone and a component for controlling a display device based on voice commands.

[0252] Example 69. The device according to any combination of Examples 52 to 68 further includes a component for outputting a subset of multiple audio streams to one or more speakers.

[0253] Example 70. The device according to any combination of Examples 52 to 69, wherein the device includes extended reality headphones, and wherein the acoustic space includes a scene represented by video data captured by a camera.

[0254] Example 71. The device according to any combination of Examples 52 to 70, wherein the device includes extended reality headphones and wherein the acoustic space includes a virtual world.

[0255] Example 72. The device according to any combination of Examples 52 to 71, wherein the device includes a head-mounted device configured to present an acoustic space.

[0256] Example 73. The device according to any combination of Examples 52 to 69, wherein the device includes a mobile handheld device.

[0257] Example 74. The device according to any combination of Examples 52 to 73 further includes a component for receiving wireless signals.

[0258] Example 75. The device according to Example 74, wherein the wireless signal is Bluetooth.

[0259] Example 76. The device according to Example 74, wherein the wireless signal conforms to the fifth generation (5G) cellular protocol.

[0260] Example 77. A non-transitory computer-readable storage medium having instructions thereon, which, when executed, cause one or more processors to: store a plurality of audio streams and corresponding audio metadata including an authorization level for each audio stream and location information associated with the coordinates of an acoustic space in which a corresponding audio stream of the plurality of audio streams is captured; and select a subset of the plurality of audio streams based on the authorization level and location information in the audio metadata, the subset of the plurality of audio streams excluding at least one audio stream of the plurality of audio streams.

[0261] Example 78. A non-transitory computer-readable storage medium according to Example 77, wherein instructions, when executed, cause one or more processors to obtain location information.

[0262] Example 79. A non-transitory computer-readable storage medium according to Example 78, wherein one or more processors obtain location information by reading location information from memory.

[0263] Example 80. A non-transitory computer-readable storage medium according to Example 78, wherein the excluded stream is associated with one or more privacy zones, and one or more processors obtain location information by determining location information.

[0264] Example 81. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 80, wherein instructions, when executed, cause one or more processors to output a subset of a plurality of audio streams to one or more speakers.

[0265] Example 82. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 81, wherein instructions, when executed, cause one or more processors to change the gain of one or more audio streams in a subset of a plurality of audio streams based on an authorization level in audio metadata.

[0266] Example 83. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 82, wherein instructions, when executed, cause one or more processors to determine a privacy zone based on the authorization level in audio metadata and location information associated with coordinates in the acoustic space.

[0267] Example 84. A non-transitory computer-readable storage medium according to Example 83, wherein instructions, when executed, cause one or more processors to determine a privacy region by obtaining the privacy region.

[0268] Example 85. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 84, wherein instructions, when executed, cause one or more processors to generate a privacy region based on the authorization level in audio metadata and location information associated with coordinates in the acoustic space.

[0269] Example 86. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 85, wherein the instructions, when executed, cause one or more processors to send a signal to at least one of a source device or a base station instructing the cessation of transmission of an excluded audio stream.

[0270] Example 87. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 86, wherein the instructions, when executed, cause one or more processors to combine at least two audio streams from a subset of a plurality of audio streams.

[0271] Example 88. A non-transitory computer-readable storage medium according to Example 87, wherein the instructions, when executed, cause one or more processors to combine at least two audio streams from a subset of a plurality of audio streams by at least one of mixing or interpolation.

[0272] Example 89. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 88, wherein instructions, when executed, cause one or more processors to override the authorization level in audio metadata.

[0273] Example 90. A non-transitory computer-readable storage medium according to Example 89, wherein instructions, when executed, cause one or more processors to output multiple audio streams to one or more speakers based on an authorization level overridden in audio metadata.

[0274] Example 91. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 90, wherein the license level in the audio metadata is received from the source device.

[0275] Example 92. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 91, wherein instructions, when executed, cause one or more processors to generate an authorization level in audio metadata.

[0276] Example 93. A non-transitory computer-readable storage medium according to any one of Examples 77 to 92, wherein instructions, when executed, cause one or more processors to control a display device based on voice commands.

[0277] Example 94. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 93, wherein instructions, when executed, cause one or more processors to output a subset of a plurality of audio streams to one or more speakers.

[0278] Example 95. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 95, wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0279] Example 96. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 95, wherein the acoustic space includes a virtual world.

[0280] Example 97. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 96, wherein instructions, when executed, cause one or more processors to present an acoustic space on a head-mounted device.

[0281] Example 98. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 96, wherein instructions, when executed, cause one or more processors to present an acoustic space on a mobile handheld device.

[0282] Example 99. A non-transitory computer-readable storage medium according to any combination of Examples 77 to 98, wherein instructions, when executed, cause one or more processors to receive wireless signals.

[0283] Example 100. A non-transitory computer-readable storage medium according to Example 99, wherein the wireless signal is Bluetooth.

[0284] Example 101. A non-transitory computer-readable storage medium according to Example 99, wherein the wireless signal conforms to a fifth-generation (5G) cellular protocol.

[0285] Example 102. An apparatus configured to play one or more audio streams from a plurality of audio streams originating from a source, the apparatus comprising: a memory configured to store a plurality of contacts, a plurality of audio streams and corresponding audio metadata, and location information associated with the coordinates of an acoustic space corresponding to one of the plurality of audio streams captured therein; and one or more processors coupled to the memory and configured to: determine whether a source is associated with one of the plurality of contacts; and when a source is not associated with one of the plurality of contacts, select a subset of the plurality of audio streams based on the location information, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0286] Example 103. The device according to Example 102, wherein one or more processors are further configured to obtain location information.

[0287] Example 104. The device according to any combination of Examples 102 to 103, wherein one or more processors are further configured to output a subset of multiple audio streams to one or more speakers.

[0288] Example 105. A device according to any combination of Examples 102 to 104, wherein one or more processors are further configured to change the gain of a subset of multiple audio streams based on whether the source is associated with one of a plurality of contacts.

[0289] Example 106. The device according to any combination of Examples 102 to 105, wherein one or more microprocessors are further configured to obtain a privacy region.

[0290] Example 107. A device according to any combination of Examples 102 to 105, wherein one or more microprocessors are further configured to determine a privacy zone based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in the acoustic space.

[0291] Example 108. A device according to any combination of Examples 102 to 107, wherein one or more microprocessors are further configured to generate a privacy zone based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in the acoustic space.

[0292] Example 109. The device according to any combination of Examples 102 to 108, wherein one or more microprocessors are further configured to send a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0293] Example 110. The device according to any combination of Examples 102 to 109, wherein one or more processors are further configured to combine at least two audio streams from a subset of multiple audio streams.

[0294] Example 111. The device according to Example 110, wherein one or more processors combine at least two audio streams from a subset of multiple audio streams by at least one of mixing or interpolation.

[0295] Example 112. The device according to any combination of Examples 102 to 111, wherein one or more microprocessors are further configured to send a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0296] Example 113. The device according to any combination of Examples 102 to 112, wherein multiple contacts include a preference level ranking.

[0297] Example 114. The device according to Example 113, wherein one or more processors are further configured to: determine whether a source device is associated with a contact among a plurality of contacts having at least a predetermined favorability rating; and when the source is not associated with a contact among a plurality of contacts having at least a predetermined favorability rating, select a subset of a plurality of audio streams.

[0298] Example 115. The device according to any combination of Examples 102 to 114 further includes a display device.

[0299] Example 116. The device according to any combination of Examples 102 to 115 further includes a microphone, wherein one or more processors are further configured to receive voice commands from the microphone and control the display device based on the voice commands.

[0300] Example 117. The device according to any combination of Examples 102 to 116 further includes one or more speakers.

[0301] Example 118. The device according to any combination of Examples 102 to 117, wherein the device includes extended reality headphones, and wherein the acoustic space includes a scene represented by video data captured by a camera.

[0302] Example 119. The device according to any combination of Examples 102 to 117, wherein the device includes extended reality headphones and wherein the acoustic space includes a virtual world.

[0303] Example 120. The device according to any combination of Examples 102 to 119 further includes a head-mounted device configured to present an acoustic space.

[0304] Example 121. The device according to any combination of Examples 102 to 117, wherein the device includes a mobile handheld device.

[0305] Example 122. The device according to any combination of Examples 102 to 121 further includes a wireless transceiver coupled to one or more processors and configured to receive wireless signals.

[0306] Example 123. The device according to Example 122, wherein the wireless signal is Bluetooth.

[0307] Example 124. The device according to Example 122, wherein the wireless signal conforms to the fifth generation (5G) cellular protocol.

[0308] Example 125. A method for playing one or more audio streams from a plurality of audio streams originating from a source, the method comprising: storing, by means of a plurality of contacts, a plurality of audio streams and corresponding audio metadata, and location information associated with coordinates of an acoustic space in which a corresponding audio stream of the plurality of audio streams is captured; determining whether the source is associated with one of the plurality of contacts; and when the source is not associated with one of the plurality of contacts, selecting, by means of one or more processors and based on the location information, a subset of the plurality of audio streams, the subset of the plurality of audio streams excluding at least one audio stream of the plurality of audio streams.

[0309] Example 126. The method according to Example 125 further includes obtaining positioning information by one or more processors.

[0310] Example 127. The method according to any combination of Examples 125 to 126 further includes outputting a subset of the plurality of audio streams to one or more speakers by one or more processors.

[0311] Example 128. The method according to any combination of Examples 125 to 127 further includes, by one or more processors, changing the gain of a subset of the multiple audio streams based on whether the source is associated with one of the multiple contacts.

[0312] Example 129. The method according to any combination of Examples 125 to 128 further includes obtaining a privacy region by one or more processors.

[0313] Example 130. The method according to any combination of Examples 125 to 129 further includes determining a privacy zone by one or more processors based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in the acoustic space.

[0314] Example 131. The method according to any combination of Examples 125 to 130 further includes generating a privacy zone by one or more processors based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in the acoustic space.

[0315] Example 132. The method according to any combination of Examples 125 to 131 further includes sending a signal to at least one of the source device or base station instructing to stop transmitting the excluded audio stream.

[0316] Example 133. The method according to any combination of Examples 125 to 132 further includes combining at least two audio streams from a subset of multiple audio streams by one or more processors.

[0317] Example 134. The method according to Example 133, wherein at least two audio streams in a subset of multiple audio streams are combined by at least one of mixing or interpolation.

[0318] Example 135. The method according to any combination of Examples 125 to 134 further includes sending a signal from one or more processors to at least one of the source device or base station instructing the transmission of the excluded audio stream to cease.

[0319] Example 136. The method according to any combination of Examples 125 to 135, wherein the multiple contacts include a preference level sort.

[0320] Example 137. The method according to Example 136 further includes: determining by one or more processors whether a source device is associated with a contact among a plurality of contacts having at least a predetermined favorability rating; and when the source is not associated with a contact among a plurality of contacts having at least a predetermined favorability rating, selecting a subset of the plurality of audio streams.

[0321] Example 138. The method according to any combination of Examples 125 to 137 further includes receiving a voice command from a microphone and controlling a display device based on the voice command by one or more processors.

[0322] Example 139. The method according to any combination of Examples 125 to 138 further includes outputting a subset of the plurality of audio streams to one or more speakers.

[0323] Example 140. The method according to any combination of Examples 125 to 139, wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0324] Example 141. The method according to any combination of Examples 125 to 139, wherein the acoustic space includes a virtual world.

[0325] Example 142. The method according to any combination of Examples 125 to 141 further includes presenting an acoustic space on a head-mounted device.

[0326] Example 143. The method according to any combination of Examples 125 to 139, wherein the method is performed on a mobile handheld device.

[0327] Example 144. The method according to any combination of Examples 125 to 143 further includes receiving a wireless signal.

[0328] Example 145. The method described in Example 144, wherein the wireless signal is Bluetooth.

[0329] Example 146. The method described in Example 144, wherein the wireless signal conforms to the fifth-generation (5G) cellular protocol.

[0330] Example 147. An apparatus configured to play one or more audio streams from a plurality of audio streams originating from a source, the apparatus comprising: components for storing a plurality of contacts, a plurality of audio streams and corresponding audio metadata, and positioning information associated with the coordinates of an acoustic space in which a corresponding audio stream of the plurality of audio streams is captured; components for determining whether the source is associated with one of the contacts; and components for selecting a subset of the plurality of audio streams based on the positioning information when the source is not associated with one of the contacts, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0331] Example 148. The device according to Example 147 further includes a component for obtaining positioning information.

[0332] Example 149. The device according to any combination of Examples 147 to 148 further includes a component for outputting a subset of a plurality of audio streams to one or more speakers.

[0333] Example 150. The device according to any combination of Examples 147 to 149 further includes a component for changing the gain of a subset of multiple audio streams based on whether the source is associated with one of the multiple contacts.

[0334] Example 151. The device according to any combination of Examples 147 to 150 further includes components for obtaining a privacy zone.

[0335] Example 152. The device according to any combination of Examples 147 to 151 further includes components for determining a privacy zone based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in the acoustic space.

[0336] Example 153. The device according to any combination of Examples 147 to 152 further includes components for generating a privacy zone based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in the acoustic space.

[0337] Example 154. The device according to any combination of Examples 147 to 153 further includes a component for transmitting a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0338] Example 155. The device according to any combination of Examples 147 to 154 further includes a component for combining at least two audio streams from a subset of the plurality of audio streams.

[0339] Example 156. The device according to Example 155, wherein at least two audio streams from a subset of a plurality of audio streams are combined by at least one of mixing or interpolation.

[0340] Example 157. The device according to any combination of Examples 147 to 156 further includes a component for transmitting a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0341] Example 158. A device according to any combination of Examples 147 to 157, wherein multiple contacts include a preference level ranking.

[0342] Example 159. The device according to Example 158 further includes: a component for determining whether a source device is associated with a contact among a plurality of contacts having at least a predetermined favorability rating; and a component for selecting a subset of a plurality of audio streams when the source is not associated with a contact among a plurality of contacts having at least a predetermined favorability rating.

[0343] Example 160. The device according to any combination of Examples 147 to 159 further includes components for receiving voice commands from a microphone and components for controlling a display device based on voice commands.

[0344] Example 161. The device according to any combination of Examples 147 to 160 further includes a component for outputting a subset of a plurality of audio streams to one or more speakers.

[0345] Example 162. The device according to any combination of Examples 147 to 161, wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0346] Example 163. The device according to any combination of Examples 147 to 161, wherein the acoustic space includes a virtual world.

[0347] Example 164. The device according to any combination of Examples 147 to 163 further includes components for presenting an acoustic space on the head-mounted device.

[0348] Example 165. The device according to any combination of Examples 147 to 164 further includes components for presenting an acoustic space on a mobile handheld device.

[0349] Example 166. The device according to any combination of Examples 147 to 165 further includes a component for receiving wireless signals.

[0350] Example 167. The device according to Example 166, wherein the wireless signal is Bluetooth.

[0351] Example 168. The device according to Example 166, wherein the wireless signal conforms to the fifth generation (5G) cellular protocol.

[0352] Example 169. A non-transitory computer-readable storage medium having instructions thereon that, when executed, cause one or more processors to: store a plurality of contacts, a plurality of audio streams and corresponding audio metadata, and location information associated with the coordinates of an acoustic space corresponding to one of the plurality of audio streams in which the audio streams are captured; determine whether a source is associated with one of the plurality of contacts; and when the source is not associated with one of the plurality of contacts, select a subset of the plurality of audio streams based on the location information, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0353] Example 170. A non-transitory computer-readable storage medium according to Example 169, wherein instructions, when executed, cause one or more processors to obtain location information.

[0354] Example 171. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 170, wherein instructions, when executed, cause one or more processors to output a subset of a plurality of audio streams to one or more speakers.

[0355] Example 172. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 171, wherein instructions, when executed, cause one or more processors to change the gain of a subset of multiple audio streams based on whether the source is associated with one of a plurality of contacts.

[0356] Example 173. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 172, wherein instructions, when executed, cause one or more processors to acquire a privacy region.

[0357] Example 174. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 173, wherein instructions, when executed, cause one or more processors to determine a privacy zone based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in acoustic space.

[0358] Example 175. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 174, wherein instructions, when executed, cause one or more processors to generate a privacy region based on whether the source is not associated with one of a plurality of contacts and location information associated with coordinates in acoustic space.

[0359] Example 176. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 175, wherein the instructions, when executed, cause one or more processors to send a signal to at least one of a source device or a base station instructing the cessation of transmission of an excluded audio stream.

[0360] Example 177. A non-transitory computer-readable storage medium according to any combination of Examples 172 to 178, wherein instructions, when executed, cause one or more processors to combine at least two audio streams from a subset of a plurality of audio streams.

[0361] Example 178. A non-transitory computer-readable storage medium according to Example 177, wherein the instructions, when executed, cause one or more processors to combine at least two audio streams from a subset of a plurality of audio streams by at least one of mixing or interpolation.

[0362] Example 179. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 178, wherein the instructions, when executed, cause one or more processors to send a signal to at least one of a source device or a base station instructing the cessation of transmission of an excluded audio stream.

[0363] Example 180. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 179, wherein a plurality of contacts include a preference rating order.

[0364] Example 181. A non-transitory computer-readable storage medium according to Example 180, wherein the instructions, when executed, cause one or more processors to: determine whether a source device is associated with a contact among a plurality of contacts having at least a predetermined favorability rating; and when the source is not associated with a contact among a plurality of contacts having at least a predetermined favorability rating, select a subset of a plurality of audio streams.

[0365] Example 182. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 181, wherein instructions, when executed, cause one or more processors to control a display device based on voice commands.

[0366] Example 183. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 182, wherein instructions, when executed, cause one or more processors to output a subset of a plurality of audio streams to one or more speakers.

[0367] Example 184. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 183, wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0368] Example 185. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 183, wherein the acoustic space includes a virtual world.

[0369] Example 186. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 185, wherein instructions, when executed, cause one or more processors to present an acoustic space on a head-mounted device.

[0370] Example 187. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 186, wherein instructions, when executed, cause one or more processors to present an acoustic space on a mobile handheld device.

[0371] Example 188. A non-transitory computer-readable storage medium according to any combination of Examples 169 to 187, wherein instructions, when executed, cause one or more processors to receive wireless signals.

[0372] Example 189. A non-transitory computer-readable storage medium according to Example 188, wherein the wireless signal is Bluetooth.

[0373] Example 190. A non-transitory computer-readable storage medium according to Example 188, wherein the wireless signal conforms to a fifth-generation (5G) cellular protocol.

[0374] Example 191. An apparatus comprising: a memory configured to store a plurality of audio streams and an associated license level for each of the audio streams; and one or more processors implemented in circuitry and communicatively coupled to the memory, and configured to: select a subset of the plurality of audio streams based on the associated license level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0375] Example 192. The device according to Example 191, wherein the memory is further configured to store positioning information associated with the coordinates of the acoustic space corresponding to one of the plurality of audio streams captured or synthesized therein.

[0376] Example 193. The device according to Example 192, wherein the device includes extended reality headphones, and wherein the acoustic space includes a scene represented by video data captured by a camera.

[0377] Example 194. The device according to Example 192, wherein the device includes extended reality headphones, and wherein the acoustic space includes a virtual world.

[0378] Example 195. The device according to Example 192, wherein the device includes extended reality headphones, and wherein the acoustic space includes the physical world.

[0379] Example 196. The device according to Example 192, wherein the selected subset of multiple audio streams is further based on location information.

[0380] Example 197. The device according to any combination of Examples 191 to 196, wherein the excluded flow is associated with one or more privacy zones.

[0381] Example 198. The device according to any combination of Examples 191 to 197, wherein one or more processors are further configured to output a subset of multiple audio streams to one or more speakers or headphones.

[0382] Example 199. A device according to any combination of Examples 191 to 198, wherein one or more processors are further configured to change the gain of one or more audio streams in a subset of a plurality of audio streams based on an associated license level.

[0383] Example 200. A device according to any combination of Examples 191 to 199, wherein one or more processors are further configured to empty the audio stream to be excluded.

[0384] Example 201. The device according to any combination of Examples 191 to 200, wherein one or more processors are further configured to send a signal to at least one of the source device or base station instructing the cessation of transmission of the excluded audio stream.

[0385] Example 202. The device according to any combination of Examples 191 to 201, wherein one or more processors are further configured to combine at least two audio streams from a subset of multiple audio streams by at least one of mixing or interpolation.

[0386] Example 203. The device according to any combination of Examples 191 to 202, wherein one or more processors are further configured to: obtain a request from a user to cover at least one authorization level; and based on the request, add at least one audio stream from the excluded audio streams associated with at least one authorization level to a subset of the plurality of audio streams.

[0387] Example 204. The device according to Example 203, wherein one or more processors are configured to output a subset of the multiple audio streams to one or more speakers or headphones based on adding at least one audio stream from the excluded audio streams to a subset of the multiple audio streams.

[0388] Example 205. The device according to any combination of Examples 191 to 204, wherein an authorization level is received from a source device.

[0389] Example 206. The device according to any combination of Examples 191 to 205, wherein one or more processors are further configured to generate an associated license level.

[0390] Example 207. The device described according to any combination of Examples 191 to 206, wherein the associated authorization level includes a rating.

[0391] Example 208. The device according to Example 207, wherein one or more processors select a subset of multiple audio streams by comparing the rating of each of the multiple audio streams with the rating of a user; and selecting a subset of the multiple audio streams based on the comparison.

[0392] Example 209. A device according to any combination of Examples 191 to 208, wherein the memory is further configured to store a plurality of contacts, and wherein the associated authorization level is based on the plurality of contacts.

[0393] Example 210. The device according to Example 209, wherein one or more processors select a subset of multiple audio streams by: determining whether the source of one or more audio streams is associated with one or more contacts among a plurality of contacts; and selecting a subset of the multiple audio streams based on comparison.

[0394] Example 211. The device according to any combination of Examples 209-210, wherein a plurality of contacts include a favorability rating order, and wherein one or more processors select a subset of a plurality of audio streams by: determining whether the source of one or more audio streams in the plurality of audio streams is associated with one or more contacts in the plurality of contacts having at least a predetermined favorability rating order; and selecting a subset of the plurality of audio streams based on comparison.

[0395] Example 212. A device according to any combination of Examples 191-211, wherein the device is a content consumer device, and the content consumer device suppresses decoding of the audio stream associated with the privacy zone when the privacy zone does not have an associated authorization level.

[0396] Example 213. A device according to any combination of Examples 191-212, wherein the device is a content consumer device and a subset of multiple audio streams includes a reproduced audio stream based on encoded information received in a bitstream decoded by one or more processors.

[0397] Example 214. The device described according to any combination of Examples 191-213, wherein the device is a source device and multiple audio streams are unencoded.

[0398] Example 215. The device according to any combination of Examples 191-214, wherein one or more processors select a subset of the multiple audio streams based on the fact that at least one of the multiple audio streams is not authorized by the user, in order to exclude at least one of the multiple audio streams.

[0399] Example 216. A device according to any combination of Examples 191-215, wherein the associated authorization level is contained in the metadata associated with each audio stream or otherwise in the bitstream.

[0400] Example 217. The device according to any combination of Examples 191 to 216 further includes a display device.

[0401] Example 218. The device according to any combination of Examples 191 to 217 further includes a microphone, wherein one or more processors are further configured to receive voice commands from the microphone and control the display device based on the voice commands.

[0402] Example 219. The device according to any combination of Examples 191 to 218 further includes one or more speakers.

[0403] Example 220. The device according to any combination of Examples 191 to 219, wherein the device includes a mobile handheld device.

[0404] Example 221. The device according to any combination of Examples 191 to 220 further includes a wireless transceiver coupled to one or more processors and configured to receive a wireless signal, wherein the wireless signal is one of Bluetooth, Wi-Fi, or conforms to a fifth-generation (5G) cellular protocol.

[0405] Example 222. A device according to any combination of Examples 191 to 221, wherein the selection of a subset of multiple audio streams is based on a comparison of an associated license level with the license level of the device or the user of the device.

[0406] Example 223. The device according to any combination of Examples 191 to 222, wherein the memory is further configured to store location information associated with the coordinates of the acoustic space of a corresponding audio stream among a plurality of audio streams captured or synthesized therein, and wherein the authorization level of one of the audio streams is determined based on the location information.

[0407] Example 224. The device according to Example 223, wherein a plurality of privacy zones are defined in an acoustic space, each privacy zone having an associated authorization level, wherein the authorization level of one of the plurality of audio streams is determined based on the authorization level of the privacy zone containing the location of the capture or synthesis of one of the plurality of audio streams.

[0408] Example 225. The device according to Example 224, wherein the authorization level of one of the multiple audio streams is equal to the authorization level of a privacy zone containing the location of the capture or synthesis of one of the multiple audio streams.

[0409] Example 226. A device according to any combination of Examples 191 to 225, wherein the device is configured to: receive a plurality of audio streams and an associated license level for each of the plurality of audio streams; and receive a license level for a device user, wherein one or more processors are configured to: select a subset of the plurality of audio streams by selecting those audio streams having an associated license level not greater than the received license level for the device user; and send the selected subset of the plurality of audio streams to an audible output device for audible output of the selected subset of the plurality of audio streams.

[0410] Example 227. A method comprising: storing in memory a plurality of audio streams and an associated license level for each of the audio streams; and selecting, by one or more processors, a subset of the plurality of audio streams based on the associated license level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0411] Example 228. The method according to Example 227 further includes, by storing in a memory, positioning information associated with the coordinates of the acoustic space of a corresponding audio stream in which the plurality of audio streams are captured or synthesized.

[0412] Example 229. The method according to Example 228, wherein the method is performed on extended reality headphones, and wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0413] Example 230. The method according to Example 228, wherein the method is performed on extended reality headphones, and wherein the acoustic space includes a virtual world.

[0414] Example 231. The method according to Example 228, wherein the method is performed on a head-mounted device configured to present an acoustic space.

[0415] Example 232. The method according to any combination of Examples 227 to 231, wherein the selected subset of multiple audio streams is further based on location information.

[0416] Example 233. The method according to any combination of Examples 227 to 232, wherein the excluded flow is associated with one or more privacy zones.

[0417] Example 234. The method according to any combination of Examples 227 to 233 further includes outputting a subset of the multiple audio streams to one or more speakers or headphones by one or more processors.

[0418] Example 235. The method according to any combination of Examples 227 to 234 further includes changing the gain of one or more audio streams in a subset of a plurality of audio streams based on an associated license level by one or more processors.

[0419] Example 236. The method according to any combination of Examples 227 to 235 further includes emptying audio streams to be excluded by one or more processors.

[0420] Example 237. The method according to any combination of Examples 227 to 236 further includes sending a signal from one or more processors to at least one of the source device or base station instructing the transmission of the excluded audio stream to stop.

[0421] Example 238. The method according to any combination of Examples 227 to 237 further includes combining at least two audio streams from a subset of multiple audio streams by one or more processors through at least one of mixing or interpolation.

[0422] Example 239. The method according to any combination of Examples 227 to 238 further includes: obtaining a request from a user to cover at least one authorization level; and based on the request, having one or more processors add at least one audio stream from the excluded audio streams associated with at least one authorization level to a subset of the plurality of audio streams.

[0423] Example 240. The method according to any combination of Examples 227 to 239 further includes having one or more processors add at least one audio stream from the excluded audio streams to a subset of the plurality of audio streams, and outputting the subset of the plurality of audio streams to one or more speakers or headphones.

[0424] Example 241. The method according to any combination of Examples 227 to 240 further includes receiving an authorization level from the source device.

[0425] Example 242. The method according to any combination of Examples 227 to 241 further includes generating an authorization level by one or more processors.

[0426] Example 243. The method according to any combination of Examples 227 to 241, wherein the associated authorization level includes a grade.

[0427] Example 244. The method according to Example 243, wherein selecting a subset of multiple audio streams includes: comparing the rating of each of the multiple audio streams with the user's rating by one or more processors; and selecting a subset of the multiple audio streams based on the comparison by one or more processors.

[0428] Example 245. The method according to any combination of Examples 227 to 244 further includes storing a plurality of contacts in a memory, wherein the associated authorization level is based on the plurality of contacts.

[0429] Example 246. The method according to Example 245, wherein selecting a subset of multiple audio streams includes: determining by one or more processors whether the source of one or more audio streams is associated with one or more contacts among a plurality of contacts; and selecting a subset of multiple audio streams by one or more processors based on comparison.

[0430] Example 247. The method according to any combination of Examples 245 to 246, wherein the plurality of contacts includes a favorability rating order, and wherein selecting a subset of the plurality of audio streams includes: determining by one or more processors whether the source of one or more audio streams among the plurality of audio streams is associated with one or more contacts among the plurality of contacts having at least a predetermined favorability rating order; and selecting a subset of the plurality of audio streams by one or more processors based on comparison.

[0431] Example 248. The method according to any combination of Examples 227 to 247 further includes suppressing decoding of the audio stream associated with the privacy zone when the privacy zone does not have an associated authorization level.

[0432] Example 249. The method according to any combination of Examples 227 to 248, wherein a subset of the plurality of audio streams includes a reproduced audio stream based on encoded information received in a bitstream decoded by one or more processors.

[0433] Example 250. The method according to any combination of Examples 227 to 249, wherein multiple audio streams are unencoded.

[0434] Example 251. The method according to any combination of Examples 227 to 250, wherein a subset of the multiple audio streams is selected to exclude at least one audio stream from the multiple audio streams based on the fact that at least one audio stream among the multiple audio streams is not authorized by the user.

[0435] Example 252. The method according to any combination of Examples 227 to 251 further includes receiving a voice command from a microphone and controlling a display device based on the voice command.

[0436] Example 253. The method according to any combination of Examples 227 to 252, wherein the method is performed on a mobile handheld device.

[0437] Example 254. The method according to any combination of Examples 227 to 253 further includes receiving a wireless signal, wherein the wireless signal is one of Bluetooth or Wi-Fi, or conforms to a fifth-generation (5G) cellular protocol.

[0438] Example 255. The method according to any combination of Examples 227 to 254, wherein the selection of a subset of multiple audio streams is based on a comparison of an associated authorization level with the authorization level of the device or the device user.

[0439] Example 256. The method according to any combination of Examples 227 to 255 further includes location information stored in a memory associated with the coordinates of an acoustic space corresponding to one of the multiple audio streams captured or synthesized therein, wherein the authorization level of one of the multiple audio streams is determined based on the location information.

[0440] Example 257. The method according to Example 256, wherein a plurality of privacy zones are defined in an acoustic space, each privacy zone having an associated authorization level, wherein the authorization level of one of the plurality of audio streams is determined based on the authorization level of the privacy zone containing the location of one of the plurality of audio streams being captured or synthesized.

[0441] Example 258. The method according to Example 257, wherein the authorization level of one of the multiple audio streams is equal to the authorization level of a privacy zone containing the location of the capture or synthesis of one of the multiple audio streams.

[0442] Example 259. The method according to any combination of Examples 227 to 258, further comprising: receiving, by one or more processors, a plurality of audio streams and an associated authorization level for each of the plurality of audio streams; receiving, by one or more processors, an authorization level of a device user; and sending, by one or more processors, a selected subset of the plurality of audio streams to an audible output device for audible output of the selected subset of the plurality of audio streams, wherein selecting the subset of the plurality of audio streams includes selecting those audio streams having an associated authorization level not greater than the received authorization level of the device user.

[0443] Example 260. An apparatus comprising: means for storing a plurality of audio streams and an associated license level for each of the audio streams; and means for selecting a subset of the plurality of audio streams based on the associated license level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0444] Example 261. A non-transitory computer-readable storage medium having instructions thereon that, when executed, cause one or more processors to: store a plurality of audio streams and an associated authorization level for each of the audio streams; and select a subset of the plurality of audio streams based on the associated authorization level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams.

[0445] It should be noted that the methods described herein depict possible implementations, and the operations and steps can be rearranged or otherwise modified, and other implementations are also possible. Furthermore, aspects from two or more methods can be combined.

[0446] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium; or a communication medium, including any medium that facilitates, for example, transferring a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory; or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.

[0447] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave may be included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient, tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0448] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be fully implemented in one or more circuit or logic elements.

[0449] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or IC sets (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in any codec hardware unit, or provided by a collection of interoperable hardware units including one or more processors as described above, combined with suitable software and / or firmware.

[0450] Various examples have been described. These examples, as well as others, are within the scope of the following claims.

Claims

1. A device for processing audio streams, comprising: a memory configured to store a plurality of audio streams and an associated authorization level for each of the plurality of audio streams; and one or more processors implemented in circuitry and communicatively coupled to the memory and configured to: select a subset of the plurality of audio streams based on the associated authorization level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams, wherein the memory is further configured to store a plurality of contacts, and wherein the associated authorization level is based on the plurality of contacts; wherein the one or more processors select the subset of the plurality of audio streams by: determining whether a source of one or more of the plurality of audio streams is associated with one or more of the plurality of contacts; and selecting the subset of the plurality of audio streams based on the determination.

2. The device of claim 1, wherein the memory is further configured to store positioning information associated with coordinates of an acoustic space in which a corresponding one of the plurality of audio streams was captured or synthesized.

3. The device of claim 2, wherein the device comprises an extended reality headset, and wherein the acoustic space comprises a scene represented by video data captured by a camera.

4. The device of claim 2, wherein the device comprises an extended reality headset, and wherein acoustic space comprises a virtual world.

5. The device of claim 2, wherein the device comprises an extended reality headset, and wherein acoustic space comprises a physical world.

6. The device of claim 2, wherein the selected subset of the plurality of audio streams is further based on the positioning information.

7. The device of claim 1, wherein the one or more processors are further configured to output the subset of the plurality of audio streams to one or more speakers or headphones.

8. The device of claim 1, wherein the one or more processors are further configured to change a gain of one or more of the subset of the plurality of audio streams based on the associated authorization level.

9. The device of claim 1, wherein the one or more processors are further configured to null out an excluded audio stream.

10. The device of claim 1, wherein the one or more processors are further configured to send a signal to at least one of a source device or a base station indicating to stop transmitting an excluded audio stream.

11. The device of claim 1, wherein the one or more processors are further configured to combine at least two of the subset of the plurality of audio streams by at least one of mixing or interpolating.

12. The device of claim 1, wherein the one or more processors are further configured to: obtain a request from a user to override at least one authorization level; and based on the request, add at least one of the excluded audio streams associated with the at least one authorization level to the subset of the plurality of audio streams.

13. The device of claim 12, wherein the one or more processors are configured to: output the subset of the plurality of audio streams to one or more speakers or headphones based on adding at least one of the audio streams to be excluded to the subset of the plurality of audio streams.

14. The device of claim 1, wherein the plurality of contacts includes a likability rank ordering, and wherein the one or more processors select the subset of the plurality of audio streams by: determining whether a source of one or more of the plurality of audio streams is associated with one or more of the plurality of contacts having at least a predetermined likability rank ordering; and selecting the subset of the plurality of audio streams based on the comparison.

15. The device of claim 1, wherein the device is a content consumer device and the subset of the plurality of audio streams includes reproduced audio streams based on encoded information received in a bitstream decoded by the one or more processors.

16. The device of claim 1, wherein the device is a source device and the plurality of audio streams are not encoded.

17. The device of claim 1, further comprising a display device.

18. The device of claim 1, further comprising a microphone, wherein the one or more processors are further configured to receive voice commands from the microphone and control a display device based on the voice commands.

19. The device of claim 1, further comprising one or more speakers.

20. The device of claim 1, wherein the device comprises a mobile handset.

21. The device of claim 1, further comprising a wireless transceiver coupled to the one or more processors and configured to receive wireless signals, wherein the wireless signals are one of Bluetooth, or Wi-Fi, or conform to a Fifth Generation (5G) cellular protocol.

22. A method for processing audio streams, comprising: storing, by a memory, a plurality of audio streams and an associated authorization level for each of the plurality of audio streams; and selecting, by one or more processors and based on the associated authorization level, a subset of the plurality of audio streams, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams, further comprising storing, by the memory, a plurality of contacts, and wherein the associated authorization level is based on the plurality of contacts; wherein selecting the subset of the plurality of audio streams comprises: determining, by the one or more processors, whether a source of one or more of the plurality of audio streams is associated with one or more of the plurality of contacts; and selecting, by the one or more processors, the subset of the plurality of audio streams based on the determination.

23. The method of claim 22, further comprising storing, by the memory, localization information associated with coordinates of an acoustic space in which a corresponding one of the plurality of audio streams was captured or synthesized.

24. The method of claim 23, wherein the method is performed on an extended reality headset, and wherein the acoustic space comprises a scene represented by video data captured by a camera.

25. The method of claim 23, wherein the method is performed on an extended reality headset, and wherein the acoustic space comprises a virtual world.

26. The method of claim 23, wherein the method is performed on a head-mounted device configured to present the acoustic space.

27. The method of claim 23, wherein the selected subset of the plurality of audio streams is further based on the positioning information.

28. The method of claim 22, further comprising outputting, by the one or more processors, the subset of the plurality of audio streams to one or more loudspeakers or headphones.

29. The method of claim 22, further comprising changing, by the one or more processors, a gain of one or more audio streams within the subset of the plurality of audio streams based on the associated authorization level.

30. The method of claim 22, further comprising nulling, by the one or more processors, excluded audio streams.

31. The method of claim 22, further comprising sending, by the one or more processors, a signal to at least one of a source device or a base station indicating to stop transmitting excluded audio streams.

32. The method of claim 22, further comprising combining, by the one or more processors, at least two audio streams in the subset of the plurality of audio streams by at least one of mixing or interpolating.

33. The method of claim 22, further comprising: obtaining a request from a user to override at least one authorization level; and based on the request, adding, by the one or more processors, at least one of the excluded audio streams associated with the at least one authorization level to the subset of the plurality of audio streams.

34. The method of claim 22, further comprising outputting, by the one or more processors, the subset of the plurality of audio streams to one or more loudspeakers or headphones based on adding at least one of the excluded audio streams to the subset of the plurality of audio streams.

35. The method of claim 22, wherein the plurality of contacts comprises a likability rank ordering, and wherein selecting the subset of the plurality of audio streams comprises: determining, by the one or more processors, whether a source of one or more audio streams in the plurality of audio streams is associated with one or more of the plurality of contacts having at least a predetermined likability rank ordering; and selecting, by the one or more processors, the subset of the plurality of audio streams based on the comparison.

36. The method of claim 22, wherein the subset of the plurality of audio streams comprises reproduced audio streams based on encoded information received in a bitstream decoded by the one or more processors.

37. The method of claim 22, wherein the plurality of audio streams are not encoded. ​ ​ 38. The method of claim 22, further comprising receiving a voice command from a microphone and controlling a display device based on the voice command.

39. The method of claim 22, wherein the method is performed on a mobile handset.

40. The method of claim 22, further comprising receiving a wireless signal, wherein the wireless signal is one of Bluetooth, or Wi-Fi, or complies with a fifth generation (5G) cellular protocol.

41. A device for processing audio streams, comprising: means for storing a plurality of audio streams and an associated authorization level for each of the plurality of audio streams; and means for selecting a subset of the plurality of audio streams based on the associated authorization level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams, further comprising means for storing a plurality of contacts, and wherein the associated authorization level is based on the plurality of contacts; wherein selecting the subset of the plurality of audio streams comprises: determining whether a source of one or more of the plurality of audio streams is associated with one or more of the plurality of contacts; and selecting the subset of the plurality of audio streams based on the determination.

42. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: storing a plurality of audio streams and an associated authorization level for each of the plurality of audio streams; and select a subset of the plurality of audio streams based on the associated authorization level, the subset of the plurality of audio streams excluding at least one of the plurality of audio streams, wherein the instructions, when executed, further cause one or more processors to store a plurality of contacts, and wherein the associated authorization level is based on the plurality of contacts; wherein the one or more processors select the subset of the plurality of audio streams by: determining whether a source of one or more of the plurality of audio streams is associated with one or more of the plurality of contacts; and selecting the subset of the plurality of audio streams based on the determination.

Citation Information

Patent Citations

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US10405126B2

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US20190007781A1

  • Filtering sounds for conferencing applications

    CN107810646A

  • Method Of Augmenting An Audio Content

    US20170017460A1