Selecting an audio stream based on motion
By selecting a subset of the audio stream that matches the user's location and processing the audio stream using immersive sound coefficients, the problem of inaccurate audio positioning in computer-mediated reality systems is solved, improving the immersiveness of the audio experience and device efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-11
- Publication Date
- 2026-03-03
AI Technical Summary
Existing computer-mediated reality systems suffer from inaccurate auditory localization in terms of audio experience, especially with improvements in video experience. This makes it difficult for users to accurately identify the source of audio content, affecting the immersive experience.
By selecting and reproducing a subset of the audio stream based on user motion, and using panoramic sound coefficients for audio stream processing and rendering, the system selects an audio stream subset that matches the user's position, thereby improving the positioning accuracy and immersion of the audio stream.
It enhances the immersiveness of the audio experience, reduces sound field reproduction and positioning errors, lowers resource utilization, and improves the operational efficiency of audio playback devices.
Smart Images

Figure CN114747231B_ABST
Abstract
Description
[0001] Priority is claimed in accordance with 35 U.S.SC § 119
[0002] This patent application claims priority to non-provisional application No. 16 / 714,150, filed on December 13, 2019, entitled “SELECTING AUDIOSTREAMS BASED ON MOTION”, which has been assigned to the assignee of this application and is expressly incorporated herein by reference. Technical Field
[0003] This disclosure pertains to the processing of audio data. Background Technology
[0004] Computer-mediated reality systems are being developed to allow computing devices to enhance, add to, remove from, reduce, or generally modify the existing reality experienced by users. Computer-mediated reality systems (which may also be referred to as “extended reality systems” or “XR systems”) can include, for example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceptual success of computer-mediated reality systems is often related to their ability to provide truly immersive experiences in both video and audio, where the video and audio experiences are aligned in a way that the user expects. Although the human visual system is more sensitive than the human auditory system (e.g., in perceptual localization of various objects within a scene), ensuring a sufficient auditory experience is an increasingly important factor in ensuring a truly immersive experience, especially as video experiences improve to allow for better localization of video objects that enable users to better identify the source of audio content. Summary of the Invention
[0005] In summary, this disclosure relates to techniques for selecting an audio stream from one or more existing audio streams based on user motion. These techniques can improve the listener experience and also reduce sound field reproduction localization errors because the selected audio stream better reflects the listener's position relative to existing audio streams, thereby improving the operation of the playback device itself (which performs the techniques for reproducing the sound field).
[0006] In one example, the technology relates to a device configured to process one or more audio streams, the device comprising: one or more processors configured to: obtain a current position of the device; obtain a plurality of capture positions, each of the plurality of capture positions identifying a position where a corresponding audio stream of a plurality of audio streams is captured; select a subset of the plurality of audio streams based on the current position and the plurality of capture positions, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and reproduce a sound field based on the subset of the plurality of audio streams; and a memory coupled to the processor and configured to store the subset of the plurality of audio streams.
[0007] In another example, the technology relates to a method for processing one or more audio streams, the method comprising: obtaining a current location of a device; obtaining a plurality of capture locations, each of the plurality of capture locations identifying a location where a corresponding audio stream of a plurality of audio streams is captured; selecting a subset of the plurality of audio streams based on the current location and the plurality of capture locations, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and reproducing a sound field based on the subset of the plurality of audio streams.
[0008] In another example, the technology relates to a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors of a device to: obtain the current location of the device; obtain a plurality of capture locations, each of the plurality of capture locations identifying a location where a corresponding audio stream of a plurality of audio streams is captured; select a subset of the plurality of audio streams based on the current location and the plurality of capture locations, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and reproduce a sound field based on the subset of the plurality of audio streams.
[0009] In another example, the technology relates to a device configured to process one or more audio streams, the device comprising: a unit for obtaining a current position of the device; a unit for obtaining a plurality of capture positions, each of the plurality of capture positions identifying a position where a corresponding audio stream of a plurality of audio streams is captured; a unit for selecting a subset of the plurality of audio streams based on the current position and the plurality of capture positions, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and a unit for reproducing a sound field based on the subset of the plurality of audio streams.
[0010] Details of one or more examples of this disclosure are set forth in the accompanying drawings and the following description. Other features, objects, and advantages of the various aspects of the technology will become apparent from the description, the drawings, and the claims. Attached Figure Description
[0011] Figure 1A and 1B This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure.
[0012] Figure 2A-2G To show in more detail Figure 1A The example shown is a diagram illustrating the example operation of the stream selection unit when performing various aspects of the stream selection techniques described in this disclosure.
[0013] Figure 3A It is shown Figure 1A and 1B A block diagram illustrating further example operations of the interpolation device when performing various aspects of the audio stream interpolation techniques described in this disclosure.
[0014] Figure 3B It is shown Figure 1A and 1B A block diagram illustrating further example operations of the interpolation device when performing various aspects of the audio stream interpolation techniques described in this disclosure.
[0015] Figure 3C It is shown Figure 1A and 1B A block diagram illustrating further example operations of the interpolation device when performing various aspects of the audio stream interpolation techniques described in this disclosure.
[0016] Figure 4A To show in more detail Figure 1A A diagram illustrating how an interpolation device of -2 can perform various aspects of the techniques described in this disclosure.
[0017] Figure 4B To show in more detail Figure 1A A block diagram illustrating how an interpolation device of -2 can perform various aspects of the techniques described in this disclosure.
[0018] Figure 5A and 5B This is a diagram showing an example of a VR device.
[0019] Figure 6A and 6B This is a diagram illustrating an example system that can perform various aspects of the techniques described in this disclosure.
[0020] Figure 7 It is shown Figure 1A ,1B A flowchart illustrating example operations of a -6B system when performing various aspects of the audio interpolation techniques described in this disclosure.
[0021] Figure 8 yes Figure 1A and 1B The example shown is a block diagram of an audio playback device performing various aspects of the techniques described in this disclosure.
[0022] Figure 9 Examples of wireless communication systems supporting audio streaming are shown, based on various aspects of this disclosure. Detailed Implementation
[0023] There are several different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound, 7.1 surround sound, 22.2 surround sound, or any other channel-based format that positions audio channels to specific locations around the listener in order to recreate the sound field.
[0024] Object-based audio formats can refer to a format in which audio objects (typically encoded using Pulse Code Modulation (PCM) and referred to as PCM audio objects) are specified to represent a sound field. Such audio objects may include metadata identifying the position of the audio object relative to a listener or other reference points in the sound field, allowing the audio object to be rendered onto one or more speaker channels for playback in an effort to recreate the sound field. The techniques described in this disclosure are applicable to any of the formats described above, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.
[0025] Scene-based audio formats can include a set of hierarchical elements that define the sound field in three dimensions. An example of such a set is the set of spherical harmonic coefficients (SHCs). The following expression illustrates a description or representation of the sound field using SHCs:
[0026]
[0027] This expression indicates that at any point in the sound field at time t... Pressure p at the point i It can be done Uniquely indicated. Here, c is the speed of sound (~343 m / s). It is the reference point (or observation point), j n (·) is a spherical Bessel function of order n, and These are spherical harmonic basis functions of order n and sub-order m (they can also be called spherical basis functions). It can be recognized that the terms in square brackets are the frequency domain representation of the signal (i.e., This can be approximated by various time-frequency transforms, such as the Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), or Wavelet Transform. Other examples of hierarchical sets include wavelet transform coefficient sets and other sets of coefficients for multi-resolution basis functions.
[0028] These can be physically acquired (e.g., recorded) through various microphone array configurations, or alternatively, they can be derived from channel-based or object-based descriptions of the sound field. SHC (which can also be called immersive sound coefficients) represent scene-based audio, where SHC can be input into an audio encoder to obtain an encoded SHC that can facilitate more efficient transmission or storage. For example, a fourth-order representation involving (1+4)² (25, and therefore fourth-order) coefficients can be used.
[0029] As mentioned above, SHC can be derived from microphone recordings using a microphone array. Various examples of how SHC can be physically obtained from a microphone array are described in the following document: Poletti, M., “Three-Dimensional Surround Sound Systems Based on Spherical Harmonics,” J. Audio Eng. Soc., Vol. 53, No. 11, November 2005, pp. 1004-1025.
[0030] The following equation illustrates how SHC can be derived from an object-based description. The coefficients used represent the sound field corresponding to a single audio object. It can be expressed as:
[0031]
[0032] Where i is It is a spherical Hankel function of order n (of the second kind), and It refers to the location of the object. Knowing the object's source energy g(ω) as a function of frequency (e.g., using time-frequency analysis techniques, such as performing a Fast Fourier Transform on a Pulse Code Modulated (PCM) stream), it is possible to convert each PCM object and its corresponding location into... Furthermore, it can be shown (since the above is a linear and orthogonal decomposition) for each object The coefficients are added together. In this way, multiple PCM objects can be generated by... Coefficients are used to represent this (e.g., as the sum of coefficient vectors for a single object). These coefficients can contain information about the sound field (pressure as a function of 3D coordinates), and the above indicates the position at the viewpoint. The transformation of the representation from a single object to the entire sound field.
[0033] Computer-mediated reality systems (also known as “extended reality systems” or “XR systems”) are being developed to take advantage of the many potential benefits offered by atmos coefficients. For example, atmos coefficients can represent a sound field in three dimensions, potentially enabling accurate three-dimensional (3D) localization of sound sources within the sound field. Therefore, XR devices can render atmos coefficients as speaker feeds, which accurately reproduce the sound field when played back through one or more speakers.
[0034] Using Atmos coefficients in XR enables the development of multiple use cases that rely on the more immersive sound field provided by Atmos coefficients, particularly for computer gaming applications and real-time video streaming applications. In these highly dynamic use cases that depend on low-latency reproduction of the sound field, XR devices may prefer Atmos coefficients (compared to other representations that are more difficult to manipulate or involve complex rendering). See below for more information. Figure 1A and 1B More information about these use cases is provided.
[0035] Although VR devices are described in this disclosure, various aspects of the technology can be implemented in the context of other devices such as mobile devices. In this case, the mobile device (such as a so-called smartphone) can present the display world via a screen that can be mounted on the user 102's head or viewed as is typically done when using a mobile device. Therefore, any information on the screen can be part of the mobile device. The mobile device is capable of providing tracking information 41, thereby allowing both a VR experience (when mounted on the head) and a normal experience of viewing the display world, wherein the normal experience can still allow the user to view the display world as a simplified version of the VR experience (e.g., raising the device and rotating or panning the device to view different parts of the display world).
[0036] Figure 1A and 1B This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure. For example... Figure 1AAs illustrated in the example, system 10 includes a source device 12 and a content consumer device 14. While described in the context of source device 12 and content consumer device 14, these techniques can be implemented in any context in which any hierarchical representation of a sound field is encoded to form a bitstream representing audio data. Furthermore, source device 12 can represent any form of computing device capable of generating hierarchical representations of sound fields, and is generally described herein in the context of a VR content creator device. Similarly, content consumer device 14 can represent any form of computing device capable of implementing the audio stream interpolation techniques and audio playback described in this disclosure, and is generally described herein in the context of a VR client device.
[0037] Source device 12 may be operated by an entertainment company or other entity capable of generating multi-channel audio content for consumption by content consumer devices (such as content consumer device 14). In many VR scenarios, source device 12 combines video content to generate audio content. Source device 12 includes content capture device 300 and content sound field representation generator 302.
[0038] The content capture device 300 can be configured to interface with or otherwise communicate with one or more microphones 5A-5N (“microphone 5”). Microphone 5 can represent... Alternatively, other types of 3D audio microphones may be used, capable of capturing sound fields and representing them as corresponding scene-based audio data 11A-11N (which may also be referred to as immersive sound coefficients 11A-11N or "immersive sound coefficients 11"). In the context of scene-based audio data 11 (which is another way of referring to immersive sound coefficients 11), each microphone 5 may represent a cluster of microphones arranged within a single housing according to a set geometry that facilitates the generation of immersive sound coefficients 11. Therefore, the term microphone can refer to a microphone cluster (which is essentially a geometrically arranged array of sensors) or a single microphone (which may be referred to as a point microphone).
[0039] The Atmos coefficient 11 can represent one example of an audio stream. Therefore, the Atmos coefficient 11 can also be referred to as audio stream 11. Although the description is primarily about the Atmos coefficient 11, the techniques described can be performed with respect to other types of audio streams, including Pulse Code Modulation (PCM) audio streams, channel-based audio streams, object-based audio streams, etc.
[0040] In some examples, the content capture device 300 may include an integrated microphone integrated into the housing of the content capture device 300. The content capture device 300 may interface with the microphone 5 wirelessly or via a wired connection. The content capture device 300 may process the immersive sound coefficients 11 after inputting, generating, or otherwise creating (from stored sound samples, such as those commonly found in gaming applications, etc.) via some type of removable storage device (wirelessly and / or via a wired input process or alternatively or in combination with the aforementioned input processes), rather than capturing or combining audio data captured via the microphone 5. Therefore, various combinations of the content capture device 300 and the microphone 5 are possible.
[0041] The content capture device 300 can also be configured to interface with or otherwise communicate with the sound field representation generator 302. The sound field representation generator 302 can include any type of hardware device capable of interfaced with the content capture device 300. The sound field representation generator 302 can use the immersive sound coefficients 11 provided by the content capture device 300 to generate various representations of the same sound field represented by the immersive sound coefficients 11.
[0042] For example, in order to generate different representations of the sound field using atmos coefficients (which are also an example of audio streams), the sound field representation generator 24 can use an encoding scheme for the atmos representation of the sound field, known as Mixed-Order Atmos (MOA), as discussed in more detail in the following document: U.S. Application Serial No. 15 / 672,058, filed August 8, 2017 and published January 3, 2019 as U.S. Patent Publication No. 20190007781 entitled “MIXED-ORDER AMBISONICS (MOA) AUDIODATA FO COMPUTER-MEDIATED REALITY SYSTEMS”.
[0043] To generate a specific MOA representation of a sound field, the sound field representation generator 24 can generate a subset of the complete set of atmos coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide accuracy for some regions of the sound field, but lower accuracy for others. In one example, the MOA representation of the sound field may include eight (8) uncompressed atmos coefficients, while the third-order atmos representation of the same sound field may include sixteen (16) uncompressed atmos coefficients. Thus, each MOA representation of the sound field generated as a subset of the atmos coefficients may be less storage-intensive and less bandwidth-intensive (if and when transmitted as part of bitstream 27 via the illustrated transmission channel) (compared to the corresponding third-order atmos representation of the same sound field generated from the atmos coefficients).
[0044] Although the MOA representation has been described, the techniques of this disclosure can also be performed with respect to a first-order Atmos (FOA) representation, where all Atmos coefficients associated with the first-order and zero-order spherical basis functions are used to represent the sound field. In other words, the sound field representation generator 302 can represent the sound field using all Atmos coefficients of a given order N (resulting in a total Atmos coefficients equal to (N+1)²), instead of using a partial non-zero subset of the Atmos coefficients.
[0045] In this regard, atmos audio data (which refers to atmos coefficients in MOA representation or full-order representation (such as the first-order representation mentioned above) may include atmos coefficients associated with spherical basis functions of order one or less (which may be referred to as "first-order atmos audio data"), atmos coefficients associated with spherical basis functions of mixed order and sub-order (which may be referred to as "MOA representation" discussed above), or atmos coefficients associated with spherical basis functions of order greater than one (which are referred to as "full-order representation" above).
[0046] In some examples, the content capture device 300 can be configured to communicate wirelessly with the sound field representation generator 302. In some examples, the content capture device 300 can communicate with the sound field representation generator 302 via one or both of a wireless or wired connection. Through the connection between the content capture device 300 and the sound field representation generator 302, the content capture device 300 can provide various forms of content, which, for the purposes of discussion, are described herein as part of the immersive sound coefficient 11.
[0047] In some examples, the content capture device 300 may utilize various aspects of the sound field representation generator 302 (in terms of the hardware or software capabilities of the sound field representation generator 302). For example, the sound field representation generator 302 may include dedicated hardware configured to perform psychoacoustic audio coding (or dedicated software that, when executed, causes one or more processors to perform psychoacoustic audio coding) (such as the Unified Speech and Audio Decoder denoted as "USAC" by the Moving Picture Experts Group (MPEG), the MPEG-H 3D Audio Decoding Standard, the MPEG-I Immersive Audio Standard, or proprietary standards such as AptX). TMThis includes various versions of AptX, such as Enhanced AptX (E-AptX), AptX Live, AptX Stereo, and AptX Highdefinition (AptX HD), Advanced Audio Codec (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus, and Windows Media Audio (WMA).
[0048] The content capture device 300 may not include dedicated hardware or software for a psychoacoustic audio encoder; instead, it may provide the audio aspects of the content 301 in a non-psychoacoustic audio decoding format. The sound field representation generator 302 may assist in capturing the content 301 by performing psychoacoustic audio encoding at least partially with respect to the audio aspects of the content 301.
[0049] The sound field representation generator 302 can also assist in content capture and transmission by generating one or more bitstreams 21 based at least in part on audio content generated from the atlas coefficients 11 (e.g., MOA representation, third-order atlas representation, and / or first-order atlas representation). The bitstreams 21 can represent compressed versions of the atlas coefficients 11 (and / or a subset thereof of the MOA representation used to form the sound field) and any other different types of content 301 (such as compressed versions of spherical video data, image data, or text data).
[0050] The sound field representation generator 302 can generate a bitstream 21 for transmission, for example, across a transmission channel (which may be a wired or wireless channel), a data storage device, etc. Bitstream 21 may represent an encoded version of the atmos coefficients 11 (and / or a subset thereof of the MOA representation used to form the sound field) and may include a main bitstream and another side bitstream (which may be referred to as side-channel information). In some cases, bitstream 21 representing a compressed version of the atmos coefficients 11 may conform to a bitstream generated according to the MPEG-H 3D audio decoding standard.
[0051] Content consumer device 14 can be operated by an individual and can represent a VR client device. Although described in relation to a VR client device, content consumer device 14 can represent other types of devices, such as augmented reality (AR) client devices, mixed reality (MR) client devices (or any other type of head-mounted display or extended reality (XR) device), standard computers, headsets, headphones, or any other device capable of tracking the head movements and / or general translational movements of the individual operating client consumer device 14. Figure 1A As shown in the example, the content consumer device 14 includes an audio playback system 16A, which can refer to any form of audio playback system capable of rendering atmos coefficients (whether in the form of first-order, second-order, and / or third-order atmos representations and / or MOA representations) for playback of multi-channel audio content.
[0052] Content consumer device 14 can retrieve bitstream 21 directly from source device 12. In some examples, content consumer device 12 can interface with networks including fifth-generation (5G) cellular networks to retrieve bitstream 21 or otherwise enable source device 12 to send bitstream 21 to content consumer device 14.
[0053] Despite Figure 1A The bitstream 21 is shown as being sent directly to content consumer device 14, but source device 12 can output the bitstream 21 to an intermediate device located between source device 12 and content consumer device 14. The intermediate device can store the bitstream 21 for later delivery to content consumer device 14, which can request the bitstream. The intermediate device can include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing the bitstream 21 for later retrieval by an audio decoder. The intermediate device can reside in a content delivery network capable of streaming the bitstream 21 (and possibly in conjunction with sending a corresponding video data bitstream) to subscribers (such as content consumer device 14) requesting the bitstream 21.
[0054] Alternatively, source device 12 may store bitstream 21 to a storage medium, such as compressed optical disc, digital video optical disc, high-definition video optical disc, or other storage media, most of which are computer-readable and therefore may be referred to as computer-readable storage media or non-transitory computer-readable storage media. In this context, transmission channels may refer to those channels through which the content stored on the medium is transmitted (and may include retail stores and other store-based delivery mechanisms). In any case, the technology of this disclosure should therefore not be limited in this respect. Figure 1A Examples.
[0055] As mentioned above, the content consumer device 14 includes an audio playback system 16A. The audio playback system 16A can represent any system capable of playing back multichannel audio data. The audio playback system 16A can include multiple different audio renderers 22. Each renderer 22 can provide different forms of audio rendering, which may include one or more of various methods of performing vector-based amplitude shifting (VBAP) and / or one or more of various methods of performing sound field synthesis. As used herein, “A and / or B” means “A or B” or both “A and B”.
[0056] The audio playback system 16A may also include an audio decoding device 24. The audio decoding device 24 may be configured to decode the bitstream 21 to output reconstructed atmos coefficients 11A'-11N' (which may form a complete first, second, and / or third-order atmos representation or a subset thereof of the same sound field MOA representation, such as the main audio signal, ambient atmos coefficients, and vector-based signals as described in the MPEG-H 3D audio decoding standard and / or the MPEG-I immersive audio standard).
[0057] Therefore, the immersive sound coefficients 11A'-11N' (“immersive sound coefficients 11”) may resemble the complete set or a subset of the immersive sound coefficients 11, but may differ due to lossy operations (e.g., quantization) and / or transmission via a transmission channel. The audio playback system 16 may, after decoding the bitstream 21 to obtain the immersive sound coefficients 11', obtain immersive sound audio data 15 from different streams of the immersive sound coefficients 11', and render the immersive sound audio data 15 to output a speaker feed 25. The speaker feed 25 may drive one or more speakers (for illustration purposes, ...). Figure 1A (Not shown in the example). The panoramic sound representation of the sound field can be normalized in a variety of ways (including N3D, SN3D, FuMa, N2D, or SN2D).
[0058] To select an appropriate renderer, or in some cases, to generate an appropriate renderer, the audio playback system 16A may obtain speaker information 13 indicating the number of speakers and / or the spatial geometry of the speakers. In some cases, the audio playback system 16A may use a reference microphone to obtain the speaker information 13 and output a signal to activate (or in other words, drive) the speakers in a manner that dynamically determines the speaker information 13 via the reference microphone. In other cases, or in conjunction with the dynamic determination of the speaker information 13, the audio playback system 16A may prompt the user to dock with the audio playback system 16A and input the speaker information 13.
[0059] The audio playback system 16A may select one of the audio renderers 22 based on the speaker information 13. In some cases, when no audio renderer 22 is within a certain threshold similarity metric (in terms of speaker geometry) to the speaker geometry specified in the speaker information 13, the audio playback system 16A may generate one of the audio renderers 22 based on the speaker information 13. In some cases, the audio playback system 16A may generate one of the audio renderers 22 based on the speaker information 13 without first attempting to select one of the existing audio renderers 22.
[0060] When the speaker feed 25 is output to headphones, the audio playback system 16A can use a renderer in renderer 22 that uses the Head-Related Transfer Function (HRTF) or other functions capable of rendering as left and right speaker feeds 25 to provide binaural rendering for headphone speaker playback. The term "speaker" or "transducer" can generally refer to any speaker, including loudspeakers and headphone speakers. One or more speakers can then play back the rendered speaker feed 25.
[0061] Although described as rendering speaker feed 25 from Atmos audio data 15, the reference to rendering speaker feed 25 can refer to other types of rendering, such as rendering directly incorporated into the Atmos audio data 15 decoded from bitstream 21. Examples of alternative rendering can be found in Appendix G of the MPEG-H 3D audio decoding standard, where rendering occurs during the formation of the main signal and background signal prior to the synthesis of the sound field. Therefore, the reference to rendering Atmos audio data 15 should be understood to refer to the rendering of the actual Atmos audio data 15 or a decomposition or representation of Atmos audio data 15 (such as the main audio signal, ambient Atmos coefficients, and / or vector-based signals (which may also be referred to as V-vectors) mentioned above).
[0062] As described above, content consumer device 14 can refer to a VR device in which a human wearable display is mounted in front of the eyes of a user operating the VR device. Figure 5A and 5B This is a diagram illustrating examples of VR devices 400A and 400B. Figure 5A In the example, VR device 400A is coupled to or otherwise includes headphones 404, which can reproduce the sound field represented by Atmos audio data 15 (another way of referring to Atmos coefficients 15) via playback from speaker feed 25. Speaker feed 25 can represent analog or digital signals that cause the diaphragm within the transducer of headphones 404 to vibrate at various frequencies. Such a process is commonly referred to as driving headphones 404.
[0063] Video, audio, and other sensory data can play a significant role in VR experiences. To participate in a VR experience, user 402 can wear a VR device 400A (which may also be referred to as a VR headset 400A) or other wearable electronic devices. The VR client device (such as the VR headset 400A) can track the user 402's head movements and adapt the video data displayed via the VR headset 400A to take head movements into account, thereby providing an immersive experience where user 402 can experience the virtual world presented by the video data in three visual dimensions.
[0064] While VR (and other forms of AR and / or MR, which are often referred to as computer-mediated reality devices) allows users to visually reside in a virtual world, VR headsets typically lack the ability to audibly place the user in that virtual world. In other words, a VR system (which may include a computer responsible for rendering video and audio data) requires a different setup. Figure 5A (The computer is not shown in the example.) and the VR headset 400A may not be able to support full 3D immersion in an audible way.
[0065] Figure 5B This is a diagram illustrating an example of a wearable device 400B that can operate according to various aspects of the technologies described in this disclosure. In various examples, the wearable device 400B can represent a VR headset (such as the VR headset 400A described above), an AR headset, a MR headset, or any other type of XR headset. Augmented Reality “AR” can refer to computer-rendered images or data overlaid on the real world in which the user actually resides. Mixed Reality “MR” can refer to computer-rendered images or data locked to a specific location in the real world, or can refer to a variant of VR in which a combination of computer-rendered 3D elements and filmed real elements is combined to create an immersive experience that simulates the user’s physical presence in the environment. Extended Reality “XR” can be a general term for VR, AR, and MR. More information on the terminology used for XR can be found in the following document: Jason Peterson, entitled “Virtual Reality, Augmented Reality, and Mixed Reality Definitions”, dated July 7, 2017.
[0066] Wearable device 400B can refer to other types of devices, such as watches (including so-called "smartwatches"), glasses (including so-called "smart glasses"), headphones (including so-called "wireless headphones" and "smart headphones"), smart clothing, smart jewelry, etc. Whether referring to VR devices, watches, glasses, and / or headphones, wearable device 400B can communicate with computing devices that support wearable device 400B via wired or wireless connections.
[0067] In some cases, the computing device supporting the wearable device 400B can be integrated within the wearable device 400B, and therefore, the wearable device 400B can be considered the same device as the computing device supporting the wearable device 400B. In other cases, the wearable device 400B can communicate with a separate computing device that can support the wearable device 400B. In this regard, the term "support" should not be understood as requiring a separate dedicated device, but one or more processors configured to perform various aspects of the technologies described in this disclosure can be integrated within the wearable device 400B, or integrated within a computing device separate from the wearable device 400B.
[0068] For example, when wearable device 400B represents an example of VR device 400B, a separate dedicated computing device (such as a personal computer including one or more processors) can render audio and video content, and wearable device 400B can determine translational head movements according to various aspects of the technology described in this disclosure, wherein based on the translational head movements, the dedicated computing device can render audio content (as a speaker feed). As another example, when wearable device 400B represents smart glasses, wearable device 400B can include one or more processors that both determine translational head movements (by docking within one or more sensors of wearable device 400B) and render speaker feeds based on the determined translational head movements.
[0069] As shown in the figure, wearable device 400B includes one or more directional speakers and one or more tracking and / or recording cameras. Additionally, wearable device 400B includes one or more inertial, haptic, and / or health sensors, one or more eye-tracking cameras, one or more high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of wearable device 400B may include durable translucent display technology and hardware.
[0070] The wearable device 400B also includes connectivity hardware, which may represent one or more network interfaces supporting multi-mode connectivity, such as 4G communication, 5G communication, Bluetooth, etc. The wearable device 400B also includes one or more ambient light sensors and bone conduction transducers. In some cases, the wearable device 400B may also include one or more passive and / or active cameras with fisheye lenses and / or telephoto lenses. Although in Figure 5B Not shown, but wearable device 400B may also include one or more light-emitting diode (LED) lights. In some examples, the LED lights may be referred to as “super bright” LED lights. In some implementations, wearable device 400B may also include one or more rear cameras. It will be understood that wearable device 400B can be represented by a variety of different shape factors.
[0071] Furthermore, tracking and recording cameras, along with other sensors, can facilitate the determination of translation distance. Although in Figure 5B The example is not shown, but the wearable device 400B may include other types of sensors for detecting translational distance.
[0072] Although specific examples of wearable devices (such as those mentioned above) Figure 5B The example discussed is the VR device 400B and in Figure 1A and Figure 1B The description will be based on other devices illustrated in the examples, but those skilled in the art will understand that... Figure 1A-4B The related descriptions can be applied to other examples of wearable devices. For example, other wearable devices (such as smart glasses) may include sensors for obtaining translational head movement. As another example, other wearable devices (such as smartwatches) may include sensors for obtaining translational movement. Therefore, the techniques described in this disclosure should not be limited to a particular type of wearable device, but any wearable device can be configured to perform the techniques described in this disclosure.
[0073] In any case, the audio aspects of VR have been divided into three separate categories of immersion. The first category provides the lowest level of immersion and is known as three degrees of freedom (3DOF). 3DOF refers to audio rendering that takes into account head movement in three degrees of freedom (yaw, pitch, and roll), thus allowing the user to freely look around in any direction. However, 3DOF cannot take into account translational head movements where the head is not centered on the optical and acoustic center of the sound field.
[0074] In addition to limited spatial translation due to the head's distance from the optical and acoustic centers within the sound field, the second category (known as 3DOF plus (3DOF+)) offers three degrees of freedom (yaw, pitch, and roll). 3DOF+ can provide support for perceptual effects such as motion parallax, which can enhance immersion.
[0075] The third category (referred to as six degrees of freedom (6DOF)) renders audio data in a manner that considers three degrees of freedom of head movement (yaw, pitch, and roll), as well as the user's translation in space (x, y, and z translations). Spatial translation can be sensed by sensors tracking the user's position in the physical world or via an input controller.
[0076] 3DOF rendering is the latest technology for audio in VR. Therefore, VR audio is less immersive than video, potentially reducing the overall immersion experienced by the user. It also introduces positioning errors (e.g., when the auditory playback does not match or is not fully relevant to the visual scene).
[0077] According to the techniques described in this disclosure, various methods are described for selecting a subset of existing audio streams 11, thereby allowing 6DOF immersion. As described below, these techniques can improve the listener experience while also reducing sound field reproduction localization errors because the selected subset of audio streams 11 better reflects the listener's position relative to the existing audio streams, thus improving the operation of the playback device itself (which performs the techniques for reproducing the sound field). Furthermore, by selecting only a subset of the available audio streams 11, these techniques can reduce resource utilization (in terms of processor cycles, memory, and bus bandwidth consumption) because not all audio streams 11 need to be rendered in order to reproduce the sound field at sufficient resolution.
[0078] like Figure 1A As shown in the example, the audio playback system 16A may include an interpolation device 30 (“INT DEVICE 30”), which can be configured to process one or more audio streams 11' to obtain an interpolated audio stream 15 (this refers to another way of referring to the immersive audio data 15). Although shown as a separate device, the interpolation device 30 may be integrated or otherwise incorporated into one of the audio decoding devices 24.
[0079] The interpolation device 30 may be implemented by one or more processors, including fixed-function processing circuitry and / or programmable processing circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits.
[0080] Interpolation device 30 can first obtain one or more microphone positions, each of the one or more microphone positions identifying the position of the corresponding one or more microphones capturing one or more audio streams 11'. About Figures 3A-3C The following example illustrates more information about the operation of interpolation device 30.
[0081] However, instead of processing every single audio stream 11', the interpolation device 30 can invoke the stream selection unit 32 (“SSU 32”) (which can select a non-zero subset of audio streams 11'), where the non-zero subset of audio streams 11' can include fewer audio streams than the total number of audio streams provided as audio streams 11'. By reducing the number of audio streams 11' interpolated by the interpolation device 30, the SSU 32 can reduce resource utilization (in terms of processing cycles, memory, and bus bandwidth) while potentially maintaining accurate reproduction of the sound field.
[0082] In operation, SSU 32 can obtain (e.g., via tracking device 306) the current position 17 of content consumer device 14 (which may also be referred to as listener position 17). In some examples, SSU 32 can convert the current position 17 of content consumer device 14 to a different coordinate system, for example, from a real coordinate system to a virtual coordinate system. That is, one or more capture positions of audio stream 11' can be defined relative to the virtual coordinate system so that audio stream 11' can be correctly rendered by audio playback system 16B to reflect the virtual world experienced by the consumer when using content consumer device 14 (e.g., VR device 14).
[0083] SSU 32 can also obtain a capture position indicating the location where the corresponding audio stream 11' within the audio stream 11' is captured. In some examples, the capture position is defined in a virtual coordinate system, where the virtual coordinate system may reflect the position in the virtual world rather than the position in the physical world where the content consumer device 14 is located. Therefore, as described above, the audio playback system 16A can convert the current position 17 from the real-world coordinate system to the virtual coordinate system before selecting a subset of the audio stream 11'.
[0084] In any case, SSU 32 can select a subset of audio stream 11' based on the current position 17 and the capture position of audio stream 11', where, again, the subset of audio stream 11' may have fewer audio streams compared to audio stream 11'. In some cases, SSU 32 can determine the distance between the current position 17 and the capture position of audio stream 11' to obtain a number (or more) distances. SSU 32 can select a subset of audio stream 11' based on distances, such as audio streams 11' with corresponding distances less than a threshold distance.
[0085] Combining the distance-based selection described above, or as an alternative, the SSU 32 can determine the angular position of each capture location relative to the current position (which may include a viewing angle defining zero degrees or a forward angle). When performing both distance-based and angular position-based selection, the SSU 32 can select from the nearest number (which may be user-, application-, or operating system-defined, as several examples) of audio streams 11' that provide a sufficient distribution of audio streams 11' around the listener operating the content consumption device 14 (e.g., regarding...). Figure 2A-2G (The example shown is described in more detail). When none of the audio streams 11' are within the threshold distance and based on the angular position, the SSU 32 may choose to provide a subset of the audio streams 11' distributed sufficiently around the listener of the operating content consumption device 14.
[0086] In some examples, SSU 32 can perform some analysis on the angular position of each capture position relative to the current position. For example, SSU 32 can determine the entropy of the angular position of each capture position relative to the current position. SSU 32 can select a subset of the audio stream 11' to maximize the entropy of the angular positions, where a relatively high entropy indicates that the capture positions are uniformly distributed in the sphere, and a relatively low entropy indicates that the capture positions are not uniformly distributed in the sphere.
[0087] SSU 32 can output a subset of the selected audio stream 11' to interpolation device 30, which can perform the aforementioned interpolation with respect to the subset of audio stream 11'. Considering that the subset of audio stream 11' does not include all audio streams 11', interpolation device 30 may consume fewer resources (such as processing cycles, memory, and bus bandwidth) to perform interpolation, thereby potentially improving the operation of the interpolation device itself.
[0088] Interpolation device 30 can output an interpolated subset of audio stream 11' as immersive audio data 15. Audio playback system 16A can invoke renderer 22 to reproduce the sound field represented by immersive audio data 15 based on immersive audio data 15. That is, renderer 22 can apply one or more rendering algorithms to convert immersive audio data 15 from the immersive (or in other words, spherical harmonic) domain to the spatial domain, thereby generating one or more speaker feeds 25 configured to drive one or more speakers ( Figure 1A (not shown in the example) or other types of transducers (including bone conduction transducers). Regarding Figure 2A-2G The example illustrates more information about the selection of a subset of audio stream 11'.
[0089] Figure 2A-2G To show in more detail Figure 1AThe example shown is a diagram illustrating the example operation of the stream selection unit when performing various aspects of the stream selection techniques described in this disclosure. Figure 2A In the example, user 52 can wear a VR device (such as content consumer device 14) to navigate a virtual world 49, where audio stream 11 is captured via microphones 50A-50F (“microphones 50”) at capture positions 51A-51F (“capture positions 51”).
[0090] As shown with respect to example microphone 50A, microphone 50A can be incorporated into or otherwise included in one or more devices, such as VR headset 60, cellular phone (including so-called smartphone) 62, camera 64, etc. Although shown only with respect to microphone 50A, each microphone 50 can be included in VR device 60, smartphone 62, camera 64, or any other type of device capable of including a microphone for capturing audio stream 11. Microphone 50 can represent the microphone mentioned above. Figure 1A The example discussed is of microphone 5. Although three example devices 60-64 are shown, microphone 50 may be included in a single device in devices 60-64 or in multiple devices in devices 60-64.
[0091] In any case, when user 52 operates content consumer device 14 at starting position 55A, SSU 32 can select a first subset 54A of microphones 50 (including microphones 50A-50D, which have fewer than all microphones 50). SSU 32 can select the first subset 54A of microphones 50 by determining distances 60A-60F from the current position 55A of content consumer device 14 and each of the plurality of capture positions 51 (wherein, for ease of illustration, in...). Figure 2A The example only shows distance 60A, but the individual distance 60B from the current position 55A to the capture position 51B can be determined, the distance 60C from the current position 55A to the capture position 51C can be determined, and so on.
[0092] Next, SSU 32 can select a subset 54A of audio stream 11' based on distances 60A-60F (“distances 60”). As an example, SSU 32 can calculate the total distance as the sum of distances 60, and then calculate the inverse distance for each distance 60 to obtain the inverse distance. Next, SSU 32 can determine the ratio for each distance 60 as the corresponding inverse distance divided by the total distance to obtain multiple corresponding ratios. Throughout this disclosure, this ratio may also be referred to as a weight. Furthermore, regarding... Figure 3A-6B Further discussion on how to calculate the weights is provided.
[0093] SSU 32 can select a subset 54A of audio stream 11' based on ratios. In this example, when one of the ratios exceeds a threshold, SSU 32 can assign the corresponding audio stream 11' from audio stream 11' to subset 54A. In other words, when the distance between content consuming device 14 and capture position 51 is small (since the inverse distance produces a larger number for a smaller distance), SSU 32 can select the audio stream 11' that is closer to user 52 / content consuming device 14. Therefore, for starting position 55A, SSU 32 can select microphones 50A-50D, thus assigning microphones 50A-50D to subset 54A.
[0094] User 52 can move from left to right along movement path 53 (where the slot indicates the direction user 52 is facing). As user 52 moves along movement path 53, SSU 32 can update the subset of microphones to transition from subset 54A of microphones 50 to subset 54B. That is, when user 52 reaches the end position 55B of the end of movement path 53, SSU 32 can recalculate the aforementioned ratio (or in other words, weight) of each microphone 50 to select subset 54B of microphones 50 (i.e., Figure 2A The example shows microphones 50C-50F and the corresponding audio stream 11'.
[0095] Next reference Figure 2B For example, user 52 is operating content consumer device 14 in virtual world 68A, where microphones 70A-70G (“microphone 70”) are located at capture positions 71A-71G (“capture position 71”). Microphone 70 can be further represented as... Figure 1A Microphone 5 is shown in the example.
[0096] In this example, SSU 32 can select a subset of microphones 70 to include microphones 70A, 70B, 70C, and 70E, where the selection is based on both the distance and angular position of microphones 70 relative to the current position 75 of user 52. Although described as both distance and angular position, SSU 32 can perform the selection based on distance, angular position, or a combination of distance and angular position. When performing the selection using both distance and angular position, in some examples, SSU 32 can first select a subset of microphones 70 based on distance and then refine the subset of microphones 70 to obtain the maximum (or at least a threshold) angular diversity (or, in some examples described in more detail below, variance and / or entropy).
[0097] To illustrate, SSU 32 can first form a subset of audio streams 11' whose contributions (or in other words, those with computational weights) are above a threshold, for example, selecting only streams whose contributions are above 10% of the aggregation value. Then, SSU 32 can perform the selection of the ending subset of audio streams 11', such that the ending subset provides a defined or threshold angle extension.
[0098] Therefore, SSU 32 can determine the angular position of each capture position 71 relative to the current position 71 to obtain the angular position. Figure 2B In the example, assume that the notch of user 52 defines a zero-degree angle, and SSU 32 determines the angular position relative to the zero-degree angle defined by the direction that user 52 is looking, or in other words, the direction it is facing. The angular position can also be referred to as the azimuth. In any case, SSU 32 can then select a subset of microphones 70 (which again includes microphones 70A, 70B, 70C, and 70E) based on the angular position to obtain a corresponding subset of audio stream 11'.
[0099] In one example, SSU 32 can determine the variance of different subsets of angular positions to obtain the variance. SSU 32 can assign audio stream 11' to subsets of audio stream 11' based on the variance. SSU 32 can select the subset of audio stream 11' that provides the highest angular (or in other words, azimuth) variance (or at least a variance threshold) in order to provide a complete (in terms of angular variance) reproduction of the 360-degree sound field.
[0100] As an alternative to or in combination with the aforementioned variance-based selection, SSU 32 can determine the entropy of different subsets of angular positions to obtain entropy. SSU 32 can assign corresponding audio streams 11' from audio stream 11' to subsets of audio stream 11' based on entropy. Again, SSU 32 can select subsets of audio stream 11' that provide the highest angular (or in other words, azimuth) entropy (or at least entropy exceeding a certain entropy threshold) to provide a complete (in terms of angular variance) reproduction of the 360-degree sound field.
[0101] like Figure 2C As shown in the example, user 52 is operating content consumer device 14 in virtual world 68B. Virtual world 68B is similar to virtual world 68B except that microphones 70A-70C have been removed. Microphone 70 can again represent Figure 1A Microphone 5 is shown in the example.
[0102] In this example, SSU 32 can select a subset of microphones 70 to include microphones 70C, 70D, 70E, and 70G, where the selection is based on both the distance and angular position of microphones 70 relative to the current position 75 of user 52. Although described as both distance and angular position, as previously stated, SSU 32 can perform the selection based on distance, angular position, or a combination of distance and angular position.
[0103] Therefore, SSU 32 can determine the angular position of each capture position 71 relative to the current position 71 to obtain the angular position. Figure 2B In the example, assume that the notch of user 52 defines a zero-degree angle, and SSU 32 determines the angular position relative to the zero-degree angle defined by the direction that user 52 is looking at, or in other words, the direction it is facing. The angular position can also be referred to as the azimuth. In any case, SSU 32 can then select a subset of microphones 70 (which again includes microphones 70A, 70B, 70C, and 70E) based on the angular position to obtain a corresponding subset of audio stream 11' in a manner similar to that discussed above.
[0104] Although a description has been given regarding the selection of a subset of audio streams 11' comprising four audio streams 11', the technique can be applied to any subset of audio streams 11' having fewer than the total number of audio streams 11', wherein this number can be defined by the user 52, the content creator, dynamically defined based on processor, memory, or other resource utilization, typically dynamically defined based on some other criteria, etc. Therefore, the technique should not be limited to a statically defined subset of audio streams 11' comprising only four audio streams 11'.
[0105] Additionally, user 52 can optionally or otherwise input various biases to support audio streams 11' captured by different microphones 70 in the microphone 70. User 52 can then pre-tune the different microphones 70 in the microphone 70 based on the perceived importance of the microphones 70. For example, one microphone 70 in the microphone 70 may be located near more audio sources, and user 52 can bias the audio stream selection to select the microphone 70 associated with more audio sources. In this respect, user 52 can use biases to cover the distance and / or angular position selection process to varying degrees to insert some user preference into the audio stream selection process.
[0106] Next reference Figure 2D-2E The example shown is as follows: Figure 2DAs shown, user 52 can reside in the first audio partition 80A, identified by microphones 80A, 80B, and 80C (where microphones 80A-80D represent microphone 5 shown in the example of Figure 1). In this example (i.e., when user 52 resides in the first audio partition 80A), SSU 32 can select the audio stream 11' captured by microphones 80A, 80B, and 80C as a subset of audio stream 11'. Therefore, SSU 32 can select a valid area (or in other words, partition) based on user position 85A and the capture position of microphone 80, thereby removing microphone 80D based on ROV (in this example).
[0107] exist Figure 2E In the example, user 52 has moved from the first audio partition 82A to the current position 85B. Interpolation unit 30 can invoke SSU 32 to determine the new ROV (i.e., ...) based on the microphone 80's current position 85B and the capture position. Figure 2E (e.g., the second audio partition 82B in the example). Then, the SSU 32 can determine a subset of the audio stream 11' captured by microphones 80A, 80B, and 80D based on the identification of the second audio partition 82B, thereby removing the audio stream 11' captured by microphone 80C.
[0108] Next reference Figure 2F and 2G For example, additional microphones 80E and 80F are added to the virtual world, creating three audio partitions 82C, 82D, and 82E. User 52 is operating content consumer device 14 at the current location 85C. Interpolation unit 30 can invoke SSU 32 to select audio partition 82D based on the current location 85C and capture location of microphone 80. Based on audio partition 82D, SSU 32 can select a subset of audio streams 11' to include audio streams 11' captured by microphones 80B-80E, thereby removing any audio streams 11' captured by microphones 80A and 80F.
[0109] exist Figure 2G In the example, user 52 is operating content consumer device 14 at current location 85D. Interpolation unit 30 can invoke SSU 32 to select audio partition 82G based on the current location 85D of microphone 80 and the capture location. Based on audio partition 82G, SSU 32 can select a subset of audio streams 11' to include audio streams 11' captured by microphones 80A, 80B, 80D, and 80F, thereby removing any audio streams 11' captured by microphones 80C and 80E.
[0110] The aforementioned audio stream selection technology can have a variety of different applications in a wide range of situations. For example, the technology can be applied to the recording of live events, such as concerts, where listeners (e.g., user 52) can approach different instruments and move around the scene. As another example, the technology can be applied to AR, where there is a mix of live and synthetic (or generated) content.
[0111] Additionally, the technology can facilitate low-cost devices because the audio stream selection technology can reduce latency and complexity (due to fewer available audio streams 11'). Furthermore, user 52 can use the video stream to bias weights or adapt to user preferences to create spatial effects based on various aspects of the technology, while also enabling user 52 to preset weight biases for artistic effects based on user 52's location and potential time.
[0112] Figures 3A-3C It is shown Figure 1A and 1B A block diagram illustrating example operations of the interpolation device 30 when performing various aspects of the audio stream interpolation techniques described in this disclosure. Figure 3A In the example, interpolation device 30 receives from SSU 32 a subset of the panoramic sound audio stream 11' (shown as "panoramic sound stream 11'") captured by microphone 5 (as described above, this can represent a cluster or array of microphones). As described above, the signal output by microphone 5 can undergo a conversion from microphone format to HOA format (shown by the box labeled "MicAmbisonics") to produce panoramic sound audio stream 11'.
[0113] The interpolation device 30 can also receive audio metadata 511A-511N (“audio metadata 511”), which may include microphone locations identifying the locations of corresponding microphones 5A-5N that captured a corresponding audio stream 11'. Microphone 5 can provide the microphone location, the operator of microphone 5 can input the microphone location, a device coupled to the microphone (e.g., content capture device 300) can specify the microphone location, or some combination thereof. The content capture device 300 can specify the audio metadata 511 as part of content 301. In any case, SSU 32 can parse the audio metadata 511 from the bitstream 21 representing content 301.
[0114] SSU 32 can also obtain listener location 17, which identifies the listener's location, such as Figure 5A As shown in the example. Audio metadata can be specified as follows: Figure 3AThe example shows the microphone's position and orientation, or only the microphone position can be specified. Additionally, listener position 17 can include listener position (or in other words, location) and orientation, or only listener position. Brief Recap Figure 1A The audio playback system 16A can interface with the tracking device 306 to obtain the listener's location 17. The tracking device 306 can refer to any device capable of tracking a listener and can include one or more of the following: a Global Positioning System (GPS) device, a camera, a sonar device, an ultrasonic device, an infrared transmitting and receiving device, or any other type of device capable of obtaining the listener's location 17.
[0115] Next, SSU 32 can perform the aforementioned audio stream selection to obtain a subset of audio stream 11'. SSU 32 can then output the subset of audio stream 11' to interpolation device 30.
[0116] Next, interpolation device 30 can perform interpolation on a subset of audio stream 11' based on one or more microphone positions and listener positions 17 to obtain interpolated audio stream 15. Audio stream 11' can initially be stored in the memory of interpolation device 30, and SSU 32 can use pointers or other data constructs to reference a subset of audio stream 11' instead of retrieving a subset of audio stream 11' and sending it to interpolation device 30. To perform interpolation, interpolation device 30 can read a subset of audio stream 11' from memory and determine the weight of each audio stream (denoted as weight(1)...weight(n)) based on one or more microphone positions and listener positions 17 (which can also be stored in memory).
[0117] When identifying a subset of audio stream 11' as described above, the SSU 32 can utilize this weight. In some examples, the SSU 32 can determine the weight and provide it to the interpolation device 30 to perform interpolation.
[0118] In any case, to determine the weights, the interpolation device 30 can calculate each weight as the ratio of the inverse distance of the corresponding audio stream 11' to the listener's position 17 to the total inverse distance of all other audio streams 11', except in edge cases where the listener is in the same location as a microphone 5 represented in the virtual world. That is, the listener may be navigating the virtual world or a real-world location represented on the device's display that has the same location as the location where one of the microphones 5 captures audio stream 11'. When the listener is in the same location as one of the microphones 5, the interpolation unit 30 can calculate the weight of one of the audio streams 11' captured by the microphone 5 where the listener is in the same location as one of the microphones 5, and set the weights of the remaining audio streams 11' to zero.
[0119] Otherwise, interpolation device 30 can calculate each weight as follows:
[0120] Weight(n) = (1 / (distance from microphone n to listener position)) / (1 / (distance from microphone 1 to listener position) + ... + 1 / (distance from microphone n to listener position)). In the above text, the listener position refers to listener position 17, weight(n) refers to the weight of audio stream 11N', and the distance from microphone <number> to listener position refers to the absolute value of the difference between the corresponding microphone position and listener position 17.
[0121] Next, the interpolation device 30 can multiply the weights by a corresponding audio stream 11' in a subset of audio streams 11' to obtain one or more weighted audio streams. The interpolation device 30 can then sum these weighted audio streams together to obtain the interpolated audio stream 15. The above can be mathematically represented by the following equation:
[0122] Weight(1) * audio stream 1 + ... + weight(n) * audio stream n = interpolated audio stream.
[0123] Here, weight (<number>) represents the weight of the corresponding audio stream <number>, and the interpolated panoramic audio data refers to the interpolated audio stream 15. The interpolated audio stream can be stored in the memory of the interpolation device 30, and can also be used for playback by speakers (e.g., VR or AR devices or headsets worn by the listener). The interpolation equation represents... Figure 3A The example shows a weighted average of the immersive audio. It should be noted that in some configurations, interpolation may be performed on the non-immersive audio stream; however, if interpolation is not performed on the immersive audio data, there may be a loss of audio quality or resolution.
[0124] In some examples, interpolation device 30 may determine the aforementioned weights frame by frame. In other examples, interpolation device 30 may determine the aforementioned weights more frequently (e.g., in a certain subframe) or less frequently (e.g., after a certain set number of frames). In these and other examples, interpolation device 30 may calculate weights solely in response to detecting some change in the listener's position and / or orientation or in response to some other feature of the basic immersive audio stream (which may enable and disable various aspects of the interpolation techniques described in this disclosure).
[0125] In some examples, the aforementioned techniques may be enabled only for audio stream 11' with certain characteristics. For instance, when the audio source represented by audio stream 11' is located in a different location than microphone 5, interpolation device 30 may interpolate only audio stream 11'. The following section discusses... Figure 4A and 4B More information on this aspect of the technology is provided.
[0126] Figure 4A To show in more detail Figure 1A , 1B A diagram illustrating how interpolation devices of type 3A can perform various aspects of the techniques described in this disclosure. (See diagram for reference.) Figure 4A As shown, the listener 52 can be within the area 94 defined by the microphones (shown as "microphone array") 5A-5E. In some examples, the microphones 5 (including when the microphones 5 represent a cluster (or in other words, a microphone array)) can be more than 5 feet apart. In any case, when the sound sources 90A-90D (such as...) Figure 4A When the “sound source 90” or “audio source 90” shown is outside the region 94 defined by microphones 5A-5E, given the mathematical constraints imposed by the above equation, the interpolation device 30 (refer to) Figure 3A It can perform interpolation.
[0127] Return to Figure 4A For example, listener 52 can input or otherwise issue one or more navigation commands (possibly by walking or by using a controller or other interface device, including a smartphone, etc.) to navigate within area 94 (along line 96). Tracking devices (such as...) Figure 3A The tracking device 306 shown in the example can receive these navigation commands and generate listener location 17.
[0128] When listener 52 begins navigating from the starting position, interpolation device 30 can generate an interpolated audio stream 15 to reweight the audio stream 11C' captured by microphone 5C, and assign relatively less weight to the audio stream 11B' captured by microphone 5B and the audio stream 11D' captured by microphone 5D, and still assign relatively less weight to (and possibly no weight to) the audio streams 11A' and 11E' captured by the corresponding microphones 5A and 5E (according to the audio stream selection technique discussed above, SSU 32 can exclude audio streams 11A' and 11E' from a subset of audio streams 11').
[0129] As listener 52 navigates along line 96 next to the position of microphone 5B, interpolation device 30 can assign more weight to audio stream 11B', relatively less weight to audio stream 11C', and even less weight to (or possibly no weight to) audio streams 11A', 11D', and 11E'. As listener 52 navigates towards the end of line 96 (where the notch indicates the direction listener 52 is moving) closer to the position of microphone 5E, interpolation device 30 can assign more weight to audio stream 11E', relatively less weight to audio stream 11A', and still relatively less weight to audio streams 11B', 11C', and 11D' (and possibly no weight, as SSU 32 may exclude these audio streams).
[0130] In this respect, the interpolation device 30 can perform interpolation based on changes in the listener position 17, according to navigation commands issued by the listener 32, to assign different weights to the audio streams 11A'-11E' over time. The changing listener position 17 may result in different emphasis within the interpolated audio stream 15, thereby promoting better auditory localization within region 94.
[0131] Although not described in the above example, the technique can also adapt to changes in the microphone's position. In other words, the microphone can be manipulated during the recording process to change its position and orientation. Since the above equation only considers the difference between the microphone position and the listener's position 17, the interpolation device 30 can continue to perform interpolation even if the microphone has been manipulated to change its position and / or orientation.
[0132] Figure 4B To show in more detail Figure 1A , 1B A block diagram illustrating how interpolation devices of type 3A can perform various aspects of the techniques described in this disclosure. Figure 4B The example shown is the same as Figure 4A The example shown is similar, except that microphone 5 is replaced by wearable devices 500A-500E (which may represent examples of wearable devices 400A and / or 400B). Wearable devices 500A-500E may each include a microphone that captures the audio stream described in more detail above.
[0133] Figure 3B It is shown Figure 1A and 1B A block diagram illustrating further example operations of the interpolation device when performing various aspects of the audio stream interpolation techniques described in this disclosure. Figure 3B The interpolation device 30A shown in the example is... Figure 3A Similar to the examples shown, except Figure 3AThe interpolation device 30 shown receives an audio stream 11' that is not captured from the microphone (as well as pre-captured and / or mixed). Figure 3A The interpolation device 30 shown in the example represents an example use during real-time capture (for real-time events such as sporting events, concerts, lectures, etc.), while Figure 3B The interpolation device 30A shown in the example represents an example of use during a pre-recorded or generated event (such as a video game, movie, etc.). The interpolation device 30A may include memory for storing the audio stream, such as... Figure 3B As shown.
[0134] Figure 3C It is shown Figure 1A and 1B A block diagram illustrating further example operations of the interpolation device when performing various aspects of the audio stream interpolation techniques described in this disclosure. Figure 3C The example shown is the same as Figure 3B The example shown is similar, except that wearable devices 500A-500N can capture audio streams 11A-11N (which are compressed and decoded into audio streams 11A'-11N'). Interpolation device 30B may include memory for storing the audio streams, such as... Figure 3B As shown.
[0135] Figure 1B This is a block diagram illustrating another example system 100 configured to perform various aspects of the technologies described in this disclosure. System 100 and Figure 1A The system 10 shown is similar, except that... Figure 1A The audio renderer 22 shown is replaced by a binaural renderer 102 that can perform binaural rendering using one or more HRTFs, or other functions that can render to the left and right speaker feeds 103.
[0136] The audio playback system 16B can output left and right speaker feeds 103 to headphones 104. Headphones 104 can represent another example of a wearable device and can be coupled to additional wearable devices to facilitate sound field reproduction, such as watches, the aforementioned VR headsets, smart glasses, smart clothing, smart rings, smart bracelets, or any other type of smart jewelry (including smart necklaces). Headphones 104 can be coupled to additional wearable devices wirelessly or via a wired connection.
[0137] Furthermore, the headphones 104 can be coupled to the audio playback system 16 via a wired connection (such as a standard 3.5mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other forms of wired connection) or wirelessly (e.g., via Bluetooth™ connection, a wireless network connection, etc.). The headphones 104 can recreate the sound field represented by the immersive sound coefficient 11 based on the left and right speaker feeds 103. The headphones 104 may include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the corresponding left and right speaker feeds 103.
[0138] Despite Figure 5A and 5B The example described pertains to VR devices, but the technology can be implemented by other types of wearable devices, including watches (such as so-called "smartwatches"), glasses (such as so-called "smart glasses"), headphones (including wireless headphones coupled via a wireless connection, or smart headphones coupled via a wired or wireless connection), and any other type of wearable device. Therefore, the technology can be implemented by any type of wearable device through which a user can interact with the device while wearing it.
[0139] Figure 6A and 6B This is a diagram illustrating an example system that can perform various aspects of the techniques described in this disclosure. Figure 6A An example is shown in which the source device 12 also includes a camera 200. The camera 200 can be configured to capture video data and provide the captured raw video data to a content capture device 300. The content capture device 300 can provide the video data to another component of the source device 12 for further processing into viewport-segmented portions.
[0140] exist Figure 6A In the example, content consumer device 14 also includes wearable device 800. It should be understood that in various implementations, wearable device 800 may be included within content consumer device 14 or externally coupled to content consumer device 14. (As mentioned above regarding...) Figure 5A and 5B The wearable device 800 discussed includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewports) and for rendering audio data.
[0141] Figure 6B It shows the relationship with Figure 6A Similar examples are shown, except Figure 6AThe audio renderer 22 shown is replaced by a binaural renderer 102 capable of performing binaural rendering using one or more HRTFs, or by other functionality capable of rendering to the left and right speaker feeds 103. The audio playback system 16 can output the left and right speaker feeds 103 to the headphones 104.
[0142] Headphones 104 can be coupled to audio playback system 16 via a wired connection (such as a standard 3.5mm audio jack, Universal System Bus (USB) connection, optical audio jack, or other forms of wired connection) or a wireless connection (such as via Bluetooth™ connection, wireless network connection, etc.). Headphones 104 can recreate the sound field represented by the immersive sound coefficient 11 based on the left and right speaker feeds 103. Headphones 104 may include a left headphone speaker and a right headphone speaker, which are powered (or in other words, driven) by the corresponding left and right speaker feeds 103.
[0143] Figure 7 It is shown Figure 1A-6B A flowchart illustrating example operations of an audio playback system when performing various aspects of the audio interpolation techniques described in this disclosure. Figure 1A The SSU 32 shown in the example can first obtain one or more capture positions (950), each of the one or more capture positions identifying the position of one or more microphones corresponding to each of the one or more audio streams 11' (in the virtual coordinate system) that are being captured. The SSU 32 can then obtain the current position 17 (952) of the content consumer device 14.
[0144] As described in more detail above, SSU 32 can select a subset of multiple audio streams 11' based on the current position 17 and multiple capture positions (954). Audio playback system 16 can then invoke audio renderer 22 to obtain one or more speaker feeds 25 based on a subset of the multiple audio streams 11' (e.g., immersive audio data 15). Audio playback system 16 can output one or more speaker feeds 25 to drive transducers (e.g., speakers) or otherwise power them. In this way, audio playback system 16 can reproduce the sound field based on a subset of the multiple audio streams 11' (956).
[0145] Figure 8 yes Figure 1A and 1B The example shown is a block diagram of an audio playback device performing various aspects of the techniques described in this disclosure. Audio playback device 16 may represent examples of audio playback device 16A and / or audio playback device 16B. Audio playback system 16 may include audio decoding device 24 in conjunction with a 6DOF audio renderer 22A, which may represent... Figure 1AAn example of audio renderer 22 is shown in the example.
[0146] Audio decoding device 24 may include a low-latency decoder 900A, an audio decoder 900B, and a local audio buffer 902. The low-latency decoder 900A can process the XR audio bitstream 21A to obtain an audio stream 901A, wherein the low-latency decoder 900A can perform relatively low-complexity decoding (compared to the audio decoder 900B) to facilitate low-latency reconstruction of the audio stream 901A. The audio decoder 900B can perform relatively higher-complexity decoding relative to the audio bitstream 21B (compared to the audio decoder 900A) to obtain the audio stream 901B. The audio decoder 900B can perform audio decoding conforming to the MPEG-H 3D audio decoding standard. The local audio buffer 902 may represent a unit configured to buffer local audio content, and the local audio buffer 902 can output the local audio content as an audio stream 903.
[0147] Bitstream 21 (composed of one or more of XR audio bitstream 21A and / or audio bitstream 21B) may also include XR metadata 905A (which may include the microphone location information mentioned above) and 6DOF metadata 905B (which may specify various parameters related to 6DOF audio rendering). The 6DOF audio renderer 22A can obtain audio streams 901A, 901B, and / or 903, as well as XR metadata 905A and 6DOF metadata 905B, and render speaker feeds 25 and / or 103 based on the listener's position and microphone position. Figure 8 In the example, the 6DOF audio renderer 22A includes an interpolation device 30 that can perform various aspects of the audio stream selection and / or interpolation techniques described in more detail above to facilitate 6DOF audio rendering.
[0148] Figure 9 An example of a wireless communication system 100 supporting audio streaming according to various aspects of this disclosure is shown. The wireless communication system 100 includes a base station 105, a user interface unit (UE) 115, and a core network 130. In some examples, the wireless communication system 100 may be a Long Term Evolution (LTE) network, an improved LTE (LTE-A) network, an LTE-A Pro network, or a New Radio (NR) network. In some cases, the wireless communication system 100 may support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low-latency communication, or communication with low-cost and low-complexity devices.
[0149] Base station 105 can wirelessly communicate with UE 115 via one or more base station antennas. Base station 105 described herein may include, or may be referred to by those skilled in the art as, a base transceiver, radio base station, access point, radio transceiver, Node B, evolved Node B (eNB), next-generation Node B, or gigabit Node B (any of which may be referred to as gNB), home Node B, home evolved Node B, or some other suitable term. Wireless communication system 100 may include different types of base stations 105 (e.g., macro cell base stations or small cell base stations). UE 115 described herein is capable of communicating with various types of base stations 105 and network devices (including macro eNBs, small cell eNBs, gNBs, relay base stations, etc.).
[0150] Each base station 105 may be associated with a specific geographic coverage area 110 in which communication with each UE 115 is supported. Each base station 105 may provide communication coverage to the corresponding geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 may utilize one or more carriers. The communication link 125 shown in the wireless communication system 100 may include: an uplink transmission from the UE 115 to the base station 105, or a downlink transmission from the base station 105 to the UE 115. The downlink transmission may also be referred to as a forward link transmission, and the uplink transmission may also be referred to as a reverse link transmission.
[0151] The geographic coverage area 110 for base station 105 can be divided into sectors, which constitute part of the geographic coverage area 110, and each sector can be associated with a cell. For example, each base station 105 can provide communication coverage for macro cells, small cells, hotspots, or other types of cells, or various combinations thereof. In some examples, base station 105 can be mobile, and therefore, communication coverage is provided for mobile geographic coverage areas 110. In some examples, different geographic coverage areas 110 associated with different technologies can overlap, and overlapping geographic coverage areas 110 associated with different technologies can be supported by the same base station 105 or different base stations 105. The wireless communication system 100 can include, for example, heterogeneous LTE / LTE-A / LTE-A Pro, or NR networks, wherein different types of base stations 105 provide coverage for individual geographic coverage areas 110.
[0152] UE 115 may be distributed throughout the wireless communication system 100, and each UE 115 may be stationary or mobile. UE 115 may also be referred to as a mobile device, wireless device, remote device, handheld device, or subscriber device, or some other suitable term, wherein "device" may also be referred to as a unit, station, terminal, or client. UE 115 may also be a personal electronic device, such as a cellular phone, personal digital assistant (PDA), tablet computer, laptop computer, or personal computer. In the examples of this disclosure, UE 115 may be any audio source described in this disclosure, including VR headsets, XR headsets, AR headsets, vehicles, smartphones, microphones, microphone arrays, or any other device including a microphone, or capable of transmitting captured and / or synthesized audio streams. In some examples, the synthesized audio stream may be an audio stream stored in memory or previously created or synthesized. In some examples, UE 115 may also refer to a wireless local loop (WLL) station, Internet of Things (IoT) device, Internet of Everything (IoE) device, or MTC device, which can be implemented in various items such as appliances, vehicles, meters, etc.
[0153] Some UE 115 (e.g., MTC or IoT devices) may be low-cost or low-complexity devices and may provide automated communication between machines (e.g., machine-to-machine (M2M) communication). M2M communication or MTC can refer to data communication technologies that allow devices to communicate with each other or base station 105 without human intervention. In some examples, M2M communication or MTC may include communication from devices exchanging and / or using audio metadata that indicates privacy restrictions and / or password-based privacy data to switch, mask, and / or empty various audio streams and / or audio sources, as will be described in more detail below.
[0154] In some cases, UE 115 can also communicate directly with other UE 115 (e.g., using peer-to-peer (P2P) or device-to-device (D2D) protocols). One or more UE 115s in a group utilizing D2D communication can be within the geographic coverage area 110 of base station 105. Other UE 115s in such a group may be outside the geographic coverage area 110 of base station 105, or otherwise unable to receive transmissions from base station 105. In some cases, multiple groups of UE 115 communicating via D2D communication can utilize a one-to-many (1:M) system, where each UE 115 transmits to every other UE 115 in the group. In some cases, base station 105 facilitates the scheduling of resources for D2D communication. In other cases, D2D communication is performed between UE 115s without involving base station 105.
[0155] Base station 105 can communicate with core network 130 and with each other. For example, base station 105 can interface with core network 130 via backhaul link 132 (e.g., via S1, N2, N3 or other interfaces). Base station 105 can communicate with each other directly (e.g., directly between base stations 105) or indirectly (e.g., via core network 130) on backhaul link 134 (e.g., via X2, Xn or other interfaces).
[0156] In some cases, wireless communication system 100 may utilize both licensed and unlicensed radio frequency spectrum bands. For example, wireless communication system 100 may employ Licensed Assisted Access (LAA), LTE Unlicensed (LTE-U) radio access technology, or NR technology in an unlicensed spectrum band (e.g., a 5 GHz ISM band). When operating in an unlicensed radio frequency spectrum band, wireless devices (e.g., base station 105 and UE 115) may employ a Listen-Before-Speak (LBT) procedure before transmitting data to ensure that the frequency channel is idle. In some cases, operation in the unlicensed spectrum band may be based on carrier aggregation configurations that combine component carriers operating in a licensed spectrum band (e.g., LAA). Operation in the unlicensed spectrum may include downlink transmission, uplink transmission, peer-to-peer transmission, or a combination of these. Duplexing in the unlicensed spectrum may be based on Frequency Division Duplex (FDD), Time Division Duplex (TDD), or a combination of both.
[0157] In this regard, various aspects of the techniques for implementing one or more of the following examples are described:
[0158] Example 1. An apparatus configured to process one or more audio streams, the apparatus comprising: a memory configured to store the one or more audio streams; and a processor coupled to the memory and configured to: obtain one or more microphone positions, each of the one or more microphone positions identifying the position of a corresponding microphone or microphone that captures each of the one or more audio streams; obtain a listener position identifying the position of a listener; perform interpolation with respect to the audio streams based on the one or more microphone positions and the listener position to obtain an interpolated audio stream; obtain one or more speaker feeds based on the interpolated audio streams; and output the one or more speaker feeds.
[0159] Example 2, the device according to Example 1, wherein the one or more processors are configured to: determine a weight for each audio stream in the audio stream based on the one or more microphone locations and the listener location; and obtain the interpolated audio stream based on the weights.
[0160] Example 3: The device according to Example 1, wherein the one or more processors are configured to: determine a weight for each audio stream in the audio stream based on the one or more microphone locations and the listener location; multiply the weight by a corresponding audio stream in the one or more audio streams to obtain one or more weighted audio streams; and obtain the interpolated audio stream based on the one or more weighted audio streams.
[0161] Example 4. The device according to Example 1, wherein the one or more processors are configured to: determine a weight for each audio stream in the audio stream based on the one or more microphone locations and the listener location; multiply the weight by a corresponding audio stream in the one or more audio streams to obtain one or more weighted audio streams; and sum the one or more weighted audio streams together to obtain the interpolated audio stream.
[0162] Example 5: A device according to any combination of Examples 2-4, wherein the one or more processors are configured to: determine the difference between each of the one or more microphone locations and the listener location; and determine the weight of each audio stream in the audio stream based on the difference between each of the one or more microphone locations and the listener location.
[0163] Example 6: A device according to any combination of Examples 2-5, wherein the one or more processors are configured to: determine the weight of each audio frame of the one or more audio streams.
[0164] Example 7: A device according to any combination of Examples 1-6, wherein the audio source represented by the audio stream is located outside the one or more microphones.
[0165] Example 8: A device according to any combination of Examples 1-7, wherein the one or more processors are configured to obtain the listener's location from a computer-mediated reality device.
[0166] Example 9. The device according to Example 8, wherein the computer-mediated reality device includes a head-mounted display device.
[0167] Example 10: A device according to any combination of Examples 1-9, wherein the one or more processors are configured to: obtain audio metadata identifying the location of the one or more microphones from a bitstream including the audio stream.
[0168] Example 11: A device according to any combination of Examples 1-10, wherein at least one of the one or more microphone positions changes to reflect the movement of a corresponding microphone among the one or more microphones.
[0169] Example 12: A device according to any combination of Examples 1-11, wherein the one or more audio streams include a immersive audio stream (including higher-order, mixed-order, first-order, and second-order), and wherein the interpolated audio stream includes an interpolated immersive audio stream (including higher-order, mixed-order, first-order, and second-order).
[0170] Example 13: A device according to any combination of claims 1-11, wherein the one or more audio streams include an Atmos audio stream, and wherein the interpolated audio stream includes an interpolated Atmos audio stream.
[0171] Example 14: A device according to any combination of Examples 1-13, wherein the listener's location changes based on navigation commands issued by the listener.
[0172] Example 15, a device according to any combination of Examples 1-14, wherein the one or more processors are configured to: receive audio metadata specifying the microphone locations, each microphone location identifying the location of a microphone cluster that captures the corresponding one or more audio streams.
[0173] Example 16: The device according to any combination of Example 15, wherein the microphone clusters are each located at a distance greater than 5 feet from each other.
[0174] Example 17: A device according to any combination of Examples 1-14, wherein each microphone is located at a distance greater than 5 feet from each other.
[0175] Example 18. A method for processing one or more audio streams, the method comprising: obtaining one or more microphone locations, each of the one or more microphone locations identifying the location of a corresponding microphone or microphone capturing each of the one or more audio streams; obtaining a listener location identifying the location of a listener; performing interpolation on the audio streams based on the one or more microphone locations and the listener location to obtain an interpolated audio stream; obtaining one or more speaker feeds based on the interpolated audio stream; and outputting the one or more speaker feeds.
[0176] Example 19. The method according to Example 18, wherein performing the interpolation includes: determining a weight for each audio stream in the audio stream based on the one or more microphone locations and the listener location; and obtaining the interpolated audio stream based on the weights.
[0177] Example 20: The method according to Example 18, wherein performing the interpolation includes: determining a weight for each audio stream in the audio stream based on the one or more microphone locations and the listener location; multiplying the weight by a corresponding audio stream in the one or more audio streams to obtain one or more weighted audio streams; and obtaining the interpolated audio stream based on the one or more weighted audio streams.
[0178] Example 21, according to the method of Example 18, wherein performing the interpolation includes: determining a weight for each audio stream in the audio stream based on the one or more microphone locations and the listener location; multiplying the weight by a corresponding audio stream in the one or more audio streams to obtain one or more weighted audio streams; and summing the one or more weighted audio streams together to obtain the interpolated audio stream.
[0179] Example 22, the method according to any combination of Examples 19-21, wherein determining the weights comprises: determining the difference between each of the one or more microphone locations and the listener location; and determining the weight of each audio stream in the audio stream based on the difference between each of the one or more microphone locations and the listener location.
[0180] Example 23, the method according to any combination of Examples 19-22, wherein determining the weights includes: determining the weights of each audio frame of the one or more audio streams.
[0181] Example 24: The method according to any combination of Examples 18-23, wherein the audio source represented by the audio stream is located outside the one or more microphones.
[0182] Example 25, the method according to any combination of Examples 18-24, wherein obtaining the listener location includes: obtaining the listener location from a computer-mediated reality device.
[0183] Example 26, the method according to Example 25, wherein the computer-mediated reality device includes a head-mounted display device.
[0184] Example 27, the method according to any combination of Examples 18-26, wherein obtaining the location of the one or more microphones includes: obtaining audio metadata identifying the location of the one or more microphones from a bitstream including the audio stream.
[0185] Example 28: The method according to any combination of Examples 18-27, wherein at least one of the one or more microphone positions changes to reflect the movement of a corresponding microphone among the one or more microphones.
[0186] Example 29. The method according to any combination of Examples 18-28, wherein the one or more audio streams include a panoramic audio stream (including higher-order, mixed-order, first-order, and second-order), and wherein the interpolated audio stream includes an interpolated panoramic audio stream (including higher-order, mixed-order, first-order, and second-order).
[0187] Example 30: The method according to any combination of claims 18-28, wherein the one or more audio streams include a panoramic audio stream, and wherein the interpolated audio stream includes an interpolated panoramic audio stream.
[0188] Example 31: The method according to any combination of Examples 18-30, wherein the listener's location changes based on navigation commands issued by the listener.
[0189] Example 32, the method according to any combination of Examples 18-31, wherein obtaining the microphone location includes: receiving audio metadata specifying the microphone location, each microphone location identifying the location of a microphone cluster that captures the corresponding one or more audio streams.
[0190] Example 33: The method according to Example 32, wherein each of the microphone clusters is located at a distance greater than 5 feet from each other.
[0191] Example 34: The method according to any combination of Examples 18-31, wherein the microphones are each located at a distance greater than 5 feet from each other.
[0192] Example 35. An apparatus configured to process one or more audio streams, the apparatus comprising: a unit for obtaining one or more microphone locations, each of the one or more microphone locations identifying the location of a corresponding one or more microphones capturing each of the corresponding one or more audio streams; a unit for obtaining a listener location identifying the location of a listener; a unit for performing interpolation with respect to the audio streams based on the one or more microphone locations and the listener location to obtain an interpolated audio stream; a unit for obtaining one or more speaker feeds based on the interpolated audio stream; and a unit for outputting the one or more speaker feeds.
[0193] Example 36, the device according to Example 35, wherein the unit for performing the interpolation includes: a unit for determining a weight of each audio stream in the audio stream based on the one or more microphone positions and the listener position; and a unit for obtaining the interpolated audio stream based on the weights.
[0194] Example 37, the device according to Example 35, wherein the unit for performing the interpolation includes: a unit for determining a weight of each audio stream in the audio stream based on the one or more microphone positions and the listener position; a unit for multiplying the weight by a corresponding audio stream in the one or more audio streams to obtain one or more weighted audio streams; and a unit for obtaining the interpolated audio stream based on the one or more weighted audio streams.
[0195] Example 38, the device according to Example 35, wherein the unit for performing the interpolation includes: a unit for determining a weight of each audio stream in the audio stream based on the one or more microphone positions and the listener position; a unit for multiplying the weight by a corresponding audio stream in the one or more audio streams to obtain one or more weighted audio streams; and a unit for summing the one or more weighted audio streams together to obtain the interpolated audio stream.
[0196] Example 39, a device according to any combination of Examples 36-38, wherein the unit for determining the weight comprises: a unit for determining the difference between each of the one or more microphone locations and the listener location; and a unit for determining the weight of each audio stream in the audio stream based on the difference between each of the one or more microphone locations and the listener location.
[0197] Example 40, a device according to any combination of Examples 36-39, wherein the unit for determining the weights comprises: a unit for determining the weights of each audio frame of the one or more audio streams.
[0198] Example 41: A device according to any combination of Examples 35-40, wherein the audio source represented by the audio stream is located outside the one or more microphones.
[0199] Example 42, the device according to any combination of Examples 35-41, wherein the unit for obtaining the listener's location includes: a unit for obtaining the listener's location from a computer-mediated reality device.
[0200] Example 43, the device according to Example 42, wherein the computer-mediated reality device includes a head-mounted display device.
[0201] Example 44, a device according to any combination of Examples 35-43, wherein the unit for obtaining the location of the one or more microphones comprises: a unit for obtaining audio metadata identifying the location of the one or more microphones from a bitstream including the audio stream.
[0202] Example 45, the device according to any combination of Examples 35-44, wherein at least one of the one or more microphone positions changes to reflect the movement of a corresponding microphone among the one or more microphones.
[0203] Example 46. A device according to any combination of Examples 35-45, wherein the one or more audio streams include a immersive audio stream (including higher-order, mixed-order, first-order, and second-order), and wherein the interpolated audio stream includes an interpolated immersive audio stream (including higher-order, mixed-order, first-order, and second-order).
[0204] Example 47: The device according to any combination of claims 35-44, wherein the one or more audio streams include an Atmos audio stream, and wherein the interpolated audio stream includes an interpolated Atmos audio stream.
[0205] Example 48, a device according to any combination of Examples 35-47, wherein the listener's location changes based on navigation commands issued by the listener.
[0206] Example 49, a device according to any combination of Examples 35-48, wherein the unit for obtaining the microphone location includes: a unit for receiving audio metadata specifying the microphone location, each microphone location identifying the location of a microphone cluster that captures the corresponding one or more audio streams.
[0207] Example 50, the device according to any combination of Example 49, wherein the microphone clusters are each located at a distance greater than 5 feet from each other.
[0208] Example 51: The device according to any combination of Examples 35-48, wherein each microphone is located at a distance greater than 5 feet from each other.
[0209] Example 52. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed, causing one or more processors to: obtain one or more microphone locations, each of the one or more microphone locations identifying the location of a corresponding microphone or microphone capturing each of the one or more corresponding audio streams; obtain a listener location identifying the location of a listener; perform interpolation on the audio streams based on the one or more microphone locations and the listener location to obtain an interpolated audio stream; obtain one or more speaker feeds based on the interpolated audio stream; and output the one or more speaker feeds.
[0210] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for implementing the technique). Furthermore, in some examples, actions or events may be performed concurrently rather than sequentially, for example, through multithreaded processing, interrupt handling, or multiple processors.
[0211] In some examples, the VR device (or streaming device) can use a network interface coupled to the VR / streaming device's memory to exchange messages with external devices, where the exchanged messages are associated with multiple available representations of the sound field. In some examples, the VR device can use an antenna coupled to the network interface to receive wireless signals including data packets, audio packets, video packets, or transport protocol data associated with multiple available representations of the sound field. In some examples, one or more microphone arrays can capture the sound field.
[0212] In some examples, multiple available representations of a sound field stored in a memory device may include multiple object-based representations of the sound field, a high-order immersive representation of the sound field, a mixed-order immersive representation of the sound field, a combination of an object-based representation of the sound field and a high-order immersive representation of the sound field, a combination of an object-based representation of the sound field and a mixed-order immersive representation of the sound field, or a combination of a mixed-order representation of the sound field and a high-order immersive representation of the sound field.
[0213] In some examples, one or more of the multiple available representations of the sound field may include at least one high-resolution region and at least one low-resolution region, and wherein the representation selected based on the steering angle provides higher spatial accuracy with respect to at least one high-resolution region and lower spatial accuracy with respect to the low-resolution region.
[0214] In one or more examples, the described functionality can be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored or transmitted as one or more instructions or code on or through a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products can include computer-readable media.
[0215] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (e.g., infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (e.g., infrared, radio, and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer instead to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically copy data, while optical discs utilize lasers to optically copy data. Combinations of the above items should also be included within the scope of computer-readable media.
[0216] Instructions can be executed by one or more processors, including fixed-function processing circuitry and / or programmable processing circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, the techniques can be implemented entirely within one or more circuit or logic elements.
[0217] The technologies disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed technologies, but they do not necessarily need to be implemented through different hardware units. Specifically, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.
[0218] Various examples have been described. These and other examples are within the scope of the appended claims.
Claims
1. A device configured to process one or more audio streams, the device comprising: one or more processors configured to: obtain a current position of the device, the current position comprising a view angle defining zero degrees or a forward angle; obtain a plurality of capture positions, each capture position of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams is captured; determine an angular position of each capture position of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; determine a variance of different subsets of the plurality of angular positions to obtain one or more variances; select a subset of the plurality of audio streams that provides a highest variance or a variance that exceeds a variance threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and reproduce a sound field based on the subset of the plurality of audio streams; and a memory coupled to the processors and configured to store the subset of the plurality of audio streams. The device comprises one of a head-mounted display, a virtual reality (VR) headset, an augmented reality (AR) headset, and a mixed reality (MR) headset.
3. A device configured to process one or more audio streams, the device comprising:
2. The apparatus of claim 1, wherein, one or more processors configured to: obtain a current position of the device, the current position comprising a view angle defining zero degrees or a forward angle; obtain a plurality of capture positions, each capture position of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams is captured; determine an angular position of each capture position of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; determine an entropy of different subsets of the plurality of angular positions to obtain one or more entropies; select a subset of the plurality of audio streams that provides a highest entropy or an entropy that exceeds an entropy threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and reproduce a sound field based on the subset of the plurality of audio streams; and a memory coupled to the processors and configured to store the subset of the plurality of audio streams.
4. A method of processing one or more audio streams, the method comprising: obtaining a current position of a device, the current position comprising a view angle defining zero degrees or a forward angle; obtaining a plurality of capture positions, each capture position of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams is captured; determining an angular position of each capture position of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; determining a variance of different subsets of the plurality of angular positions to obtain one or more variances; selecting a subset of the plurality of audio streams that provides a highest variance or a variance that exceeds a variance threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and reproducing a sound field based on the subset of the plurality of audio streams. The device comprises one of a head-mounted display, a virtual reality (VR) headset, an augmented reality (AR) headset, and a mixed reality (MR) headset. 5. The method of claim 4, wherein, 6. A method of processing one or more audio streams, the method comprising: obtaining a current position of a device, the current position including a view angle defining zero degrees or a forward angle; obtaining a plurality of capture positions, each of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams was captured; determining an angular position of each of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; determining an entropy of different subsets of the plurality of angular positions to obtain one or more entropies; selecting a subset of the plurality of audio streams that provides a highest entropy or an entropy that exceeds an entropy threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and rendering a sound field based on the subset of the plurality of audio streams.
7. A computer-readable medium having stored thereon instructions that, when executed, cause one or more processors of a device to: obtain a current position of the device, the current position including a view angle defining zero degrees or a forward angle; obtain a plurality of capture positions, each of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams was captured; determine an angular position of each of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; determine a variance of different subsets of the plurality of angular positions to obtain one or more variances; select a subset of the plurality of audio streams that provides a highest variance or a variance that exceeds a variance threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and render a sound field based on the subset of the plurality of audio streams.
8. A computer-readable medium having stored thereon instructions that, when executed, cause one or more processors of a device to: obtain a current position of the device, the current position including a view angle defining zero degrees or a forward angle; obtain a plurality of capture positions, each of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams was captured; determine an angular position of each of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; determine an entropy of different subsets of the plurality of angular positions to obtain one or more entropies; select a subset of the plurality of audio streams that provides a highest entropy or an entropy that exceeds an entropy threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and render a sound field based on the subset of the plurality of audio streams.
9. A device configured to process one or more audio streams, the device comprising: means for obtaining a current position of a device, the current position including a view angle defining zero degrees or a forward angle; means for obtaining a plurality of capture positions, each of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams was captured; means for determining an angular position of each of the plurality of capture positions relative to the view angle to obtain a plurality of angular positions; means for determining a variance of different subsets of the plurality of angular positions to obtain one or more variances; means for selecting a subset of the plurality of audio streams that provides a highest variance or a variance that exceeds a variance threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and means for rendering a sound field based on the subset of the plurality of audio streams.
10. A device configured to process one or more audio streams, the device comprising: means for obtaining a current position of a device, the current position including a perspective that defines a zero or forward angle; means for obtaining a plurality of capture positions, each capture position of the plurality of capture positions identifying a position at which a respective one of a plurality of audio streams was captured; means for determining an angular position of each capture position of the plurality of capture positions relative to the perspective to obtain a plurality of angular positions; means for determining an entropy of different subsets of the plurality of angular positions to obtain one or more entropies; means for selecting a subset of the plurality of audio streams that provides a highest entropy or an entropy that exceeds an entropy threshold, the subset of the plurality of audio streams having fewer audio streams than the plurality of audio streams; and means for rendering a sound field based on the subset of the plurality of audio streams.
Citation Information
Patent Citations
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US10405126B2
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US20190007781A1
Adapting distributed audio recording for end user free viewpoint monitoring
CN110140170A
Dynamic binaural sound capture and reproduction
US20040076301A1
Ambisonic navigation of sound fields from an array of microphones
WO2018064528A1