Parameter setting adjustment for extended reality experiences
By using energy maps to adjust the parameter settings of audio components in an extended reality system, the lack of coordination between audio components is resolved, resulting in a balanced and immersive audio experience, and optimizing the audio capture and rendering process.
Patent Information
- Application Number
- CN202080047177.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-01
- Filing Date
- 2020-07-02
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2040-07-02
AI Technical Summary
In existing extended reality systems, the lack of coordination between audio components results in an inability to provide an immersive audio experience, and users may become disoriented or perceive discrepancies in audio rendering.
The XR device receives the energy map of the audio components, forms a composite energy map, and adjusts the parameter settings of the audio components based on the energy map to match the sound in the audio environment. It also disables or removes unqualified audio components, thus optimizing the audio capture and rendering process.
It achieves a balanced and immersive audio experience, saves processing and storage resources, and improves the coordination of audio components and the quality of audio data.
Smart Images

Figure CN114391263B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Application No. 16 / 918,754, filed July 1, 2020, and U.S. Provisional Application No. 62 / 870,570, filed July 3, 2019, the entire contents of each of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the processing of media data, such as audio data. Background Technology
[0003] Computer-mediated reality systems are being developed to allow computing devices to enhance, add to, remove from, subtract from, or generally modify the existing reality of the user experience. Computer-mediated reality systems (also known as “extended reality systems” or “XR systems”) can include, for example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceptual success of computer-mediated reality systems is often related to their ability to provide realistic and immersive experiences in both video and audio, where the video and audio experiences match the user's expectations. Although the human visual system is more sensitive than the human auditory system (e.g., in perceiving and locating various objects in a scene), ensuring a sufficient auditory experience is an increasingly important factor in ensuring realistic and immersive experiences, especially as video experiences improve to allow for better localization of video objects, enabling users to better identify the source of audio content. Summary of the Invention
[0004] This disclosure generally relates to the auditory aspects of user experience in computer-mediated reality systems, including virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. Various aspects of this technology can provide adaptive audio capture, rendering of extended reality systems, and compensation for differences in parameter settings via one or more parameter adjustments. Various aspects of this technology can provide adaptive audio capture or synthesis and rendering of acoustic spaces for extended reality (XR) systems. As used herein, an acoustic environment is referred to as an indoor environment or an outdoor environment, or both. An acoustic environment may include one or more subacoustic spaces, which may include various acoustic elements. Examples of outdoor environments may include cars, buildings, walls, forests, etc. An acoustic space may be an example of an acoustic environment and may be an indoor space or an outdoor space. As used herein, audio elements are sounds captured by microphones (e.g., captured directly from near-field sources or reflected from far-field sources, whether real or synthesized), or previously synthesized sound fields, or sounds synthesized from text into speech, or reflections of virtual sounds from objects in the acoustic environment.
[0005] In one example, aspects of the technology relate to a device configured to determine parameter adjustments for audio capture, the device including a memory configured to store at least one energy map corresponding to one or more audio streams; and one or more processors coupled to the memory and configured to: access the at least one energy map corresponding to the one or more audio streams; determine parameter adjustments with respect to at least one audio element based at least in part on the at least one energy map, the parameter adjustments being configured to adjust audio capture through the at least one audio element; and output the parameter adjustments.
[0006] In another example, aspects of the technology relate to a method for determining parameter adjustments for audio capture, the method comprising: accessing at least one energy map corresponding to one or more audio streams; determining parameter adjustments with respect to at least one audio element based at least in part on the at least one energy map, the parameter adjustments being configured to adjust audio capture through the at least one audio element; and outputting an indication of the parameter adjustments with respect to the at least one audio element.
[0007] In another example, aspects of the technology relate to a device configured to determine parameter adjustments for audio capture, the device comprising: components for accessing at least one energy map corresponding to one or more audio streams; components for determining parameter adjustments with respect to at least one audio element based at least in part on the at least one energy map, the parameter adjustments being configured to adjust audio capture through the at least one audio element; and components for outputting an indication of the parameter adjustments with respect to the at least one audio element.
[0008] In another example, aspects of the technology relate to a non-transitory computer-readable storage medium on which instructions are stored, which, when executed, cause one or more processors to: access at least one energy map corresponding to one or more audio streams; determine, at least in part based on the at least one energy map, a parameter adjustment with respect to at least one audio element, the parameter adjustment being configured to adjust audio capture through the at least one audio element; and output an indication of the parameter adjustment with respect to the at least one audio element.
[0009] In another example, aspects of the technology relate to a device configured to generate a sound field, the device including a memory configured to store audio data representing the sound field; and one or more processors coupled to the memory and configured to: send an audio stream to one or more source devices; determine instructions for adjusting parameter settings of audio elements; and adjust the parameter settings to adjust the generation of the sound field.
[0010] In another example, aspects of the technology relate to a method for adjusting parameter settings for sound field generation, the method comprising: sending an audio stream to one or more source devices; determining instructions for adjusting parameter settings of audio elements; and adjusting the parameter settings to adjust the generation of the sound field.
[0011] In another example, aspects of the technology relate to a device configured to generate a sound field, the device comprising: components for sending an audio stream to one or more source devices; components for determining instructions for adjusting parameter settings of audio elements; and components for adjusting the parameter settings to adjust the generation of the sound field.
[0012] In another example, aspects of the technology relate to a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: send an audio stream to one or more source devices; determine instructions for adjusting parameter settings of audio elements; and adjust the parameter settings to adjust the generation of the sound field.
[0013] Details of one or more examples of this disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of various aspects of the present invention will be apparent from the detailed description, the accompanying drawings, and the claims. Attached Figure Description
[0014] Figures 1A to 1C This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure.
[0015] Figure 2 This is a diagram showing an example of a VR device worn by a user.
[0016] Figures 3A to 3D To show in more detail Figures 1A to 1C The example shown is a diagram of the example operation of the stream selection unit.
[0017] Figures 4A to 4B It is shown Figures 1A to 1C The example shown is a flowchart illustrating the operation of an audio decoding device when performing various aspects of the adjustment technique.
[0018] Figures 5A to 5D To show in more detail Figures 1A to 1C The example shown is a diagram illustrating the operation of an audio decoding device.
[0019] Figure 6 This is a diagram illustrating an example of a wearable device that can operate according to various aspects of the technology described in this disclosure.
[0020] Figure 7A and7B This is a diagram illustrating other example systems that can perform various aspects of the techniques described in this disclosure.
[0021] Figure 8 It is shown Figures 1A to 1C The example shown is a block diagram of one or more sample components from the source device and content consuming device.
[0022] Figures 9A to 9C It is shown Figures 1A to 1C The example shown is a flowchart illustrating the example operations of the stream selection unit when performing various aspects of the stream selection technique.
[0023] Figure 10 An example of a wireless communication system with supporting parameters adjusted according to various aspects of this disclosure is shown. Detailed Implementation
[0024] The techniques disclosed herein generally relate to the adjustment of certain audio elements configured to facilitate audio rendering in an extended reality (XR) system. Specifically, the disclosed techniques relate to determining ideal parameter settings for audio elements configured to capture or synthesize audio data for an XR system. Multiple audio elements can work together to provide an audio experience for an XR experience. In examples, an XR system can utilize various audio elements, such as audio receivers (e.g., microphones) or audio synthesizers, configured to capture and / or generate (e.g., produce, reproduce, recreate, synthesize, etc.) audio data representing a specific sound field in an audio environment. In examples, an XR system can utilize audio elements configured to synthesize audio data to provide audio in an XR experience. In some examples, a user can utilize a computer program to generate audio for an XR experience. In any case, audio elements configured to capture or generate audio in an XR system can do so based on an application that adjusts the adjustable parameter settings of the audio signal or audio elements. When properly compensated across devices, the audio stream can be provided in a uniform or equalized manner. If there is no proper compensation between audio components, the audio components may fail to provide an immersive XR experience and may ultimately disorient or confuse users who are trying to experience XR spaces (e.g., XR worlds, virtual worlds, AR worlds, etc.).
[0025] The parameter settings of audio components may not initially be coordinated or compatible with other audio components configured to contribute to the audio stream to render an immersive audio experience. In one example, two microphones capturing audio within a common sound field may apply different gain settings. In another example, two microphones from different manufacturers or vendors may apply similar gain settings, but due to manufacturing differences, they may do so in a way that still causes variations in the generated audio data. In yet another example, a source device may provide synthesized audio that will be included in another audio rendering, such as audio captured by a microphone or other audio receiving device. In such examples, synchronized parameter settings may be necessary so that a user experiencing the audio may not perceive differences in the audio rendering from various different audio components. This lack of coordination between audio components can be particularly noticeable when the user manually changes parameter settings (such as when the user adjusts the gain for high-frequency sounds for an audio receiving device or audio synthesizer, or when, as mentioned above, the audio component system includes audio components from different manufacturers or vendors).
[0026] According to the technology disclosed herein, an XR device can receive an energy map of each audio element (e.g., a microphone, synthesized sound source, etc.) in a constellation of audio elements. The energy map corresponds to an audio representation of audio captured or synthesized via the audio elements. The XR device can also form a composite energy map comprising several energy maps corresponding to different audio elements implementing an audio stream in an XR environment. Based on the energy maps, the XR device may adjust parameter settings for one or more audio elements, where the energy map differs from the energy maps of other audio elements in the same audio environment. The XR device can induce parameter adjustments by sending adjustment instructions to the audio elements, such as instructions to adjust the gain of a microphone in the environment as determined according to the energy map to match the sound generated by other audio elements (e.g., microphones, etc.) in the environment. In some examples, the XR device may induce parameter adjustments during decoding of audio data received from a source device or during rendering of audio using an audio renderer.
[0027] Additionally, XR devices can determine their operational status based on various audio components implemented in the environment. In the example, the XR device may receive audio samples from a microphone or other status data indicating the current operational status of the audio components. The operational status may include the signal-to-noise ratio (SNR), which indicates whether the microphone is currently operating to generate audio that meets or does not meet a predefined SNR threshold.
[0028] In illustrative and non-limiting examples, because the first audio element (e.g., a microphone) is in a person's pocket during audio capture, it may fail to generate a high-quality audio signal. Therefore, the XR device can determine that the operating state of the first audio element indicates that its SNR is below a predefined SNR threshold (e.g., it does not meet the SNR threshold). In such examples, the XR device can remove the unqualified first audio element from the constellation set of other audio elements before forming or updating the composite energy map of the constellation. Thus, the XR device can identify an audio element as unqualified, where the audio element is, for example, damaged, noisy (e.g., poor SNR), does not generate sound, etc. In another example, the XR device can disable or remove the audio stream of the first audio element from multiple audio streams before forming or updating the composite energy map of the constellation. In this way, the XR device can form a composite energy map, which is configured to use as a baseline for comparison with additional energy maps of various other audio elements.
[0029] In some examples, an XR device can determine the effective (e.g., qualified) audio elements to be sent to a constellation set of audio elements based on a composite energy map. In one example, the XR device can compare the composite energy map with the energy maps of the audio elements, and based on this comparison, the XR device can determine the parameter adjustments (e.g., gain adjustment, etc.) used to regulate the audio data obtained from the audio elements. In this way, the XR device can effectively reduce variations between individual energy maps based on the energy map determined for the audio stream of other audio elements, such as a composite energy map generated from multiple energy maps.
[0030] According to one or more of the various techniques disclosed herein, an XR device can determine certain parameter adjustments for audio elements. The XR device can be configured to apply these parameter adjustments during audio data capture, during audio data synthesis, or while the XR device is rendering audio data, such as to render aspects of the audio experience to provide an XR experience to the user. In examples, such as in cases where audio elements generate corrupted or noisy audio, parameter adjustments may include adjusting the gain parameter settings of specific audio elements within a constellation set of audio elements, or disabling the audio elements.
[0031] In some examples, parameter tuning may also include disabling or excluding audio elements, such as preventing audio from a particular audio element from being used if the user may set certain privacy restrictions. In such instances, the XR device is configured to exclude the energy map of disabled audio elements when forming a composite energy map. After making specific parameter adjustments, the user can experience a balanced and immersive XR experience when using the XR device. Additionally, the XR device can save processing and memory resources by identifying and excluding certain audio elements from a constellation set of audio elements configured to capture audio data from and / or within a common sound field. This is because the XR device can efficiently utilize those resources to manage and analyze the energy map only for those audio elements capable of providing a balanced and immersive XR experience.
[0032] There are several different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound, 7.1 surround sound, 22.2 surround sound, or any other channel-based format that positions audio channels to specific locations around the listener in order to generate a sound field.
[0033] Object-based audio formats can refer to formats in which specified audio objects (typically encoded using pulse code modulation (PCM) and referred to as PCM audio objects) are used to represent a sound field. Such audio objects may include location information (e.g., metadata) identifying the position of the audio object relative to the listener or other reference points in the sound field, allowing the audio object to be rendered onto one or more speaker channels for playback to generate a sound field. The techniques described in this disclosure can be applied to any of the following formats, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.
[0034] Scene-based audio formats can include a set of hierarchical elements that define a three-dimensional (3D) sound field. An example of a hierarchical element set is a set of spherical harmonic coefficients (SHCs). The following expression demonstrates a description or representation of the sound field using SHCs:
[0035]
[0036] This expression represents any point in the sound field. Pressure p at time t i It can be by Unique representation. Here, c is the speed of sound (approximately 343 m / s). It is the reference point (or observation point), j n (·) is a spherical Bessel function of order n, and These are spherical harmonic basis functions of order n and sub-order m (also referred to as spherical basis functions). It can be recognized that the terms within the square brackets are signals (e.g., The frequency domain representation of a time-frequency basis function (TF-F) can be approximated by various time-frequency transforms such as the Discrete Fourier Transform (DFT), Discrete Cosine Transform (DCT), or Wavelet Transform. Other examples of hierarchical sets include wavelet transform coefficient sets and other coefficient sets of multi-resolution basis functions.
[0037] Physical acquisition (e.g., recording) can be configured through various microphone arrays, or alternatively, they can be derived from channel-based or object-based sound field descriptions. SHC (also known as surround sound coefficients) represents scene-based audio, where SHC can be input to an audio encoder to obtain a post-encoded SHC that can facilitate more efficient transmission or storage. For example, a (1+4) approach could be used. 2 (25, therefore it is a fourth-order representation of (4) coefficients.
[0038] As mentioned above, SHC can be derived from microphone recordings using microphone arrays. Various examples of how SHC can be physically obtained from microphone arrays are described in "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics" by Poletti, M., published in J. Audio Eng. Soc., Vol. 53, No. 11, pp. 1004–1025, November 2005.
[0039] The following equations illustrate how SHC can be derived from an object-based description. The coefficients of the sound field corresponding to each audio object. It can be expressed as:
[0040]
[0041] Where i is It is a (second kind) spherical Hankel function of order n, and It is the location of the object. (For example, using time-frequency analysis techniques, such as performing a Fast Fourier Transform on a stream of pulse code modulated PCM) knowing the object source energy g(ω) as a function of frequency makes it possible to convert each PCM object and its corresponding location into Furthermore, the coefficients of each object can be shown (since the above is a linear and orthogonal decomposition). Yes, it can be added. In this way, the number of PCM objects can be increased by... The coefficients are represented (e.g., as the sum of the coefficient vectors of the individual objects). The coefficients can contain information about the sound field (as a function of 3D coordinates), and the above represents the distance from each object to the viewpoint. The transformation of the entire sound field representation in the vicinity.
[0042] Computer-mediated reality systems (also known as “extended reality systems” or “XR systems”) are being developed to take advantage of the many potential benefits offered by surround sound coefficients. For example, surround sound coefficients can represent a 3D sound field by potentially enabling accurate 3D localization of sound sources within the sound field. Therefore, XR devices can render surround sound coefficients to a speaker feed that, when played through one or more speakers or headphones, can accurately generate the sound field.
[0043] As another example, surround sound coefficients can be translated or rotated to account for user movement without overly complex mathematical calculations, potentially adapting to the low latency requirements of XR devices. Furthermore, surround sound coefficients are hierarchical, thus naturally adapting to scalability through order reduction (which eliminates surround sound coefficients associated with higher orders), potentially enabling dynamic adjustment of the sound field to suit the latency and / or battery requirements of XR devices.
[0044] Especially for computer gaming and real-time video streaming applications, using surround sound coefficients in XR devices enables the development of many use cases that rely on a more immersive sound field provided by the surround sound coefficients. In these highly dynamic use cases that rely on low latency generation (e.g., reproduction) of the sound field, XR devices may prefer surround sound coefficients over other representations that are more difficult to manipulate or involve complex rendering. (See below for reference.) Figures 1A to 1C More information about these use cases is provided.
[0045] While VR devices are described in this disclosure, various aspects of these technologies can be performed in the context of other devices, such as mobile devices, speakers, audio elements (e.g., microphones, synthesized audio sources, etc.) or other XR devices. In illustrative and non-limiting examples, a mobile device (such as a so-called smartphone) can render an acoustic space (e.g., via speakers, one or more headsets, etc.). The mobile device, or at least a portion thereof, can be mounted on a user's head or viewed as is done when using a mobile device normally. That is, any information generated via speakers, headsets, or audio elements, as well as any information on the screen of the mobile device, can be considered part of the mobile device. The mobile device may be able to provide tracking information, thereby allowing both an XR experience (when mounted on the head) and a normal experience to experience the acoustic space, where the normal experience can still allow the user to experience the acoustic space providing a simplified XR experience (e.g., lifting the device and rotating, moving, or panning the device to experience different parts of the acoustic space). Additionally, the technologies of this disclosure can also be used with a displayed world that can correspond to the acoustic space in some instances, where the displayed world can be presented on the screen of an XR device (e.g., a mobile device, VR device, etc.).
[0046] Figures 1A to 1C This is a diagram illustrating a system capable of performing various aspects of the techniques described in this disclosure. For example... Figure 1A As shown in the example, system 10 includes a source device 12A and a content consuming device 14A. Although described in the context of source device 12A and content consuming device 14A, this technique can be implemented in any context in which any representation of a sound field is encoded to form a bitstream representing audio data (e.g., an audio stream). Furthermore, source device 12A can represent any form of computing device capable of generating a sound field representation, and is generally described herein in the context of a VR content creator device. Similarly, content consuming device 14A can represent any form of computing device capable of implementing the audio compensation techniques and audio playback described in this disclosure, and is generally described herein in the context of a VR client device.
[0047] Source device 12A can be operated by an entertainment company or other entity capable of generating multi-channel audio content for consumption by operators of content consumption devices such as content consumption device 14A. In some VR scenarios, source device 12A generates audio content in conjunction with video content. Source device 12A includes content capture device 20, content editing device 22, and sound field representation generator 24. Content capture device 20 can be configured to interface with or otherwise communicate with microphone 18 or other audio components.
[0048] Microphone 18 can represent Or other types of three-dimensional (3D) audio microphones capable of capturing a sound field and representing it as audio data 19, which may refer to one or more of the aforementioned scene-based audio data (such as surround sound coefficients), object-based audio data, and channel-based audio data. Although described as a 3D audio microphone, microphone 18 may also represent other types of microphones configured to capture audio data 19 (such as omnidirectional microphones, point microphones, unidirectional microphones, etc.). Audio data 19 may represent an audio stream or include an audio stream.
[0049] In some examples, the content capture device 20 may include an integrated microphone 18 integrated into the housing of the content capture device 20. The content capture device 20 may interface with the microphone 18 wirelessly or via a wired connection. Instead of capturing or combining captured audio data 19 via the microphone 18, the content capture device 20 may process the audio data 19 after it has been input via some type of removable storage, wirelessly, and / or via a wired input process. In examples, the content capture device 20 may process the audio data 19 after inputting it, and in combination with processing the input audio data 19, the content capture device 20 may capture the audio data 19 via the microphone 18. In some examples, the audio data 19 may include an audio type layer. In examples, the content capture device 20 may output the audio data 19 as including previously stored audio data 19, such as previously recorded audio input, combined with an audio layer captured in combination with real-time or near real-time processing of the previously stored audio data 19. It should be understood that various other combinations of the content capture device 20 and the microphone 18 are possible according to this disclosure.
[0050] Content capture device 20 may also be configured to interface with or otherwise communicate with content editing device 22. In some instances, content capture device 20 may include content editing device 22 (in some instances, this may represent software or a combination of software and hardware, including software executed by content capture device 20 to configure content capture device 20 to perform a particular form of content editing (e.g., signal conditioning)). In some examples, content editing device 22 is a device physically separate from content capture device 20.
[0051] Content editing device 22 can represent a unit configured to edit or otherwise modify content 21 (including audio data 19) received from content capture device 20. Content editing device 22 can output the edited content 23 and associated metadata 25 to sound field representation generator 24. Metadata 25 may include privacy restriction metadata, feasibility metadata, parameter setting information (PSI), audio location information, and other audio metadata. In the example, content editing device 22 can apply parameter adjustments (such as adjustments that can be defined by PSI) to audio data 19 or to content 21 (e.g., gain parameters, frequency response parameters, SNR parameters, etc.) and generate the edited content 23 therefrom.
[0052] In some examples, the content editing device 22 may apply parameter settings (such as gain, frequency response, compression, compression ratio, noise reduction, directional microphone, transpilation / compression, and / or equalization settings) to modify or adjust the capture of incoming audio and / or modify or adjust the outgoing audio stream (e.g., the sound field is synthesized to be rendered as if the audio stream was captured at a specific location in a virtual or non-virtual world or other generated sound field). The parameter settings may be defined by a PSI 46A. The PSI 46A may include information received from the content consuming device 14A via side channels 33 or via bitstream 27. The PSI 46A may define adjustments to the parameter settings, such as gain adjustment, frequency response adjustment, compression adjustment, or equalization settings.
[0053] In another example, content consuming device 14A can send one or more energy maps, such as composite energy maps, to source device 12A. Source device 12A can receive one or more energy maps and determine PSI 46A based on one or more energy maps. Source device 12A can apply adjusted parameter settings to the capture of audio data 19, wherein the adjusted parameter settings are defined by PSI 46A. Source device 12A can then send audio data 19 to content consuming device 14A via bitstream 27, wherein bitstream 27 has been adjusted based on the determined PSI 46A. Thus, content consuming device 14A can receive bitstream 27 (e.g., audio stream) conforming to one or more energy maps from source device 12A without performing additional adjustments on bitstream 27 (e.g., audio signal) to match the audio stream with other source devices 12A (e.g., source devices 12A ... Figure 1C One or more source devices 12B, other source devices 12A, Figure 7A Matching other audio streams (or one or more source devices of 7B, 12C, etc.).
[0054] In some examples, content editing device 22 can generate edited content 23 including audio data 19, where PSI 46A is applied to the audio data 19. Additionally, content editing device 22 can generate metadata 25 that may include PSI 46A. In such examples, source device 12A may send parameter settings applied via PSI 46A to content consuming device 14A before or after adjusting PSI 46A based on PSI 46B. In this way, content consuming device 14A can determine adjustments to parameter settings based on the current parameter settings of source device 12A, since those settings relate to the energy map of the audio stream (e.g., bitstream 27) and the difference between the energy map and the composite energy map that has been formed and / or stored in constellation diagram (CM) 47.
[0055] In the example, content consuming device 14A (e.g., an XR device) can determine PSI 46B based on the energy map of one or more audio streams of an audio element. Content consuming device 14A can determine PSI 46B and use PSI 46B to adjust from source device 12A or from another source device (e.g., Figure 1C The content consuming device 14A receives an audio stream from the source device 12A. It can receive an energy map from the source device 12A and determine the energy map of the source device 12A, or a combination thereof, based on the audio stream received from the source device 12A. In this example, the content consuming device 14A can receive energy maps from several source devices 12A and can determine the energy maps of other source devices 12A. The content consuming device 14A can store the energy map in CM 47 or another storage location of the audio decoding device 34. In some instances, the audio decoding device 34 may include PSI 46B as part of the audio data 19', allowing the audio renderer 32 to apply PSI 46B to the audio data 19' when rendering it.
[0056] Alternatively, content consuming device 14A may output PSI 46B to source device 12A. Source device 12A may store information as PSI 46A, which in some instances may simply involve updating a PSI 46A previously applied by source device 12A. In some instances, source device 12A may be reconfigured or otherwise adjusted based on PSI 46A.
[0057] The sound field representation generator 24 may include any type of hardware device capable of interfacing with the content editing device 22 (or the content capture device 20). Although Figure 1AThe example is not shown, but the sound field representation generator 24 can use edited content 23, which includes audio data 19 and information (e.g., metadata 25) provided by the content editing device 22 to generate one or more bitstreams 27. Focusing on audio data 19... Figure 1A In some examples, the sound field representation generator 24 can generate one or more representations of the same sound field represented by the audio data 19 to obtain a bitstream 27 including the sound field representations. In some examples, the bitstream 27 may also include metadata 25 (e.g., audio metadata).
[0058] For example, in order to generate different representations of the sound field using surround sound coefficients (which is also an example of audio data 19), the sound field representation generator 24 can use a decoding scheme for the surround sound representation of the sound field, called Mixed Order Surround Sound (MOA), as discussed in more detail in U.S. Patent Application No. 15 / 672,058 entitled “MIXED-ORDER AMBISONICS (MOA) AUDIO DATA FOR COMPUTER-MEDIATED REALITY SYSTEMS”, filed August 8, 2017 and published as U.S. Patent Application Publication No. 2019 / 0007781, filed January 3, 2019.
[0059] To generate a specific MOA representation of a sound field, the sound field representation generator 24 can generate a partial subset of the entire set of surround sound coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide accuracy for some regions of the sound field, but lower accuracy for others. In the example, the MOA representation of the sound field may include eight (8) uncompressed surround sound coefficients, while the third-order surround sound representation of the same sound field may include sixteen (16) uncompressed surround sound coefficients. Thus, the storage density and bandwidth density of each MOA representation of the sound field (e.g., generated as a partial subset of surround sound coefficients) may be lower than (e.g., in the case where the MOA representation of the sound field is transmitted as part of bitstream 27 via the shown transmission channel) the corresponding third-order surround sound representation of the same sound field generated from the surround sound coefficients.
[0060] While the MOA representation has been described, the techniques of this disclosure can also be performed on a first-order surround sound (FOA) representation, wherein all surround sound coefficients associated with the first-order and zero-order spherical basis functions are used to represent the sound field. In other words, the sound field representation generator 24 can represent the sound field using all surround sound coefficients of a given order N, rather than using a partial non-zero subset of the surround sound coefficients, resulting in a total surround sound coefficient equal to (N+1). 2
[0061] In this regard, surround sound audio data (which refers to another way of representing surround sound coefficients in MOA representation or full-order representation, such as the first-order representation mentioned above) may include surround sound coefficients associated with spherical basis functions of first order or lower (which may be referred to as "1st-order surround sound audio data"), surround sound coefficients associated with spherical basis functions of mixed order and sub-order (which may be referred to as "MOA representation" discussed above), or surround sound coefficients associated with spherical basis functions of greater than first order (which are referred to as "full-order representation" in this paper).
[0062] In some examples, the content capture device 20 or the content editing device 22 may be configured to communicate wirelessly with the sound field representation generator 24. In some examples, the content capture device 20 or the content editing device 22 may communicate with the sound field representation generator 24 via one or both of a wireless or wired connection. Through the connection between the content capture device 20 or the content editing device 22 and the sound field representation generator 24, the content capture device 20 or the content editing device 22 may provide content in various forms, which, for the purposes of discussion, are described herein as part 19 of audio data.
[0063] In some examples, content capture device 20 may utilize various aspects of sound field representation generator 24 (in terms of the hardware or software capabilities of sound field representation generator 24). For example, sound field representation generator 24 may include dedicated hardware configured (or dedicated software that, when executed, causes one or more processors to execute) to perform psychoacoustic audio coding (such as that by the Moving Picture Experts Group (MPEG), MPEG-H3D audio decoding standard, MPEG-I immersive audio standard, or proprietary standards (such as AptX)). TM (Including various versions of AptX, such as Enhanced AptX–E-AptX, AptX Live, AptX Stereo, and AptX High Definition–AptX-HD), Advanced Audio Decoder (AAC), Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Streaming (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey's Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3), Opus and Windows Media Audio (WMA), or other standards) The unified speech and audio decoder referred to as “USAC”.
[0064] In some examples, content capture device 20 may not include dedicated hardware or software for a psychoacoustic audio encoder, but may instead provide the audio aspects of content 21 in a non-psychoacoustic audio decoding form. Sound field representation generator 24 can assist in capturing content 21 by performing psychoacoustic audio encoding on the audio aspects of content 21, at least partially. In some examples, sound field representation generator 24 may apply PSI 46A to the audio aspects of content 21 to generate a bitstream 27 (e.g., an audio stream) conforming to initial or adjusted parameter settings (such as gain settings or adjusted gain settings) of PSI 46A.
[0065] The sound field representation generator 24 can also assist in content capture and transmission by generating one or more bitstreams 27 based at least in part on audio content (e.g., MOA representation and / or first (or higher) order surround sound representation) generated from audio data 19 (where audio data 19 includes scene-based audio data). The bitstreams 27 can represent compressed versions of audio data 19 and any other different types of content 21 (such as compressed versions of spherical video data, image data, or text data).
[0066] The sound field representation generator 24 can generate a bitstream 27 for transmission, for example, across a transmission channel, which can be a wired or wireless channel, such as Wi-Fi. TM Audio channels Audio channels, or audio channels conforming to fifth-generation (5G) cellular standards, data storage devices, etc. Bitstream 27 may represent an encoded version of audio data 19 and may include a main bitstream and another side bitstream, which may be referred to as sidechannel information (e.g., metadata), as shown via sidechannel 33. In some instances, bitstream 27, representing a compressed version of audio data 19 (which may also represent scene-based audio data, object-based audio data, channel-based audio data, or a combination thereof), may be generated according to the MPEG-H 3D audio decoding standard and / or the MPEG-I immersive audio standard.
[0067] In some examples of this disclosure, source device 12A can be configured to generate multiple audio streams for transmission to content consuming device 14A. Source device 12A can be configured to generate each of the multiple audio streams via a single content capture device 20 and / or a cluster of content capture devices 20 (e.g., multiple content capture devices). In some use cases, it may be desirable to be able to control which audio streams of the multiple audio streams generated by source device 12A are available for playback by content consuming device 14A.
[0068] For example, audio from certain capture devices of content capture device 20 may contain sensitive information and / or audio from certain capture devices of content capture device 20 may not imply exclusive access (e.g., unrestricted access for all users). In some examples, it may be necessary to restrict access to audio from certain capture devices of content capture device 20 based on the type of information captured by content capture device 20 and / or based on the location of the physical area where content capture device 20 is located. Such privacy restrictions may play a role in whether content consuming device 14A can utilize one or more audio streams from certain audio elements to form a composite energy map, where privacy restrictions or other types of restrictions cause content consuming device 14A to exclude such audio elements when forming the composite energy map.
[0069] According to the example techniques of this disclosure, source device 12A may also include a controller 31 configured to generate metadata 25. In the example, metadata 25 may indicate privacy restrictions (e.g., privacy restriction metadata). In some examples, source device 12 and content consuming device 14 may be configured to communicate via side channel 33. In the example, content consuming device 14 may send a PSI to source device 12. In another example, content consuming device 14 may send at least one energy map (e.g., a composite energy map) to source device 12. In such examples, source device 12 may access at least one energy map. Source device 12 may determine PSI 46A based on a comparison of the energy map (e.g., an energy map corresponding to source device 12) with at least one accessed energy map.
[0070] In some examples, metadata 25 may correspond to one or more of a plurality of bitstreams 27 generated by source device 12A. In the examples, privacy restriction metadata may indicate when one or more of the plurality of bitstreams 27 are restricted or unrestricted audio streams.
[0071] In some examples, controller 31 may generate only privacy-restricting metadata to indicate whether bitstream 27 includes a restricted or unrestricted audio stream. In such examples, content consuming device 14 may infer that an audio stream without privacy-restricting metadata (e.g., metadata indicating a restricted audio stream) is unrestricted. Content consuming device 14 may receive privacy-restricting metadata and determine, based on the privacy restrictions, one or more bitstreams 27 (e.g., audio streams) that can be used for decoding and / or playback. Content consuming device 14A may generate a corresponding sound field based on one or more bitstreams 27 determined to be usable for decoding and / or playback.
[0072] exist Figure 1A In one example, controller 31 sends privacy-restricted metadata in side channel 33. In another example, controller 31 may send privacy-restricted metadata in bitstream 27.
[0073] In some examples, controller 31 does not need to be a separate physical unit. Instead, controller 31 can be integrated into content editing device 22 or sound field representation generator 24. In another example, controller 31 can receive data from content consuming device 14A, such as PSI 46B. Controller 31 can then reconfigure content editing device 22, content capture device 20, and / or sound field representation generator 24 based on PSI 46A. After parameter adjustment (e.g., reconfiguration), source device 12 can then generate an audio stream represented by an energy map that has been compensated to match other energy maps and / or composite energy maps (e.g., an energy map formed by multiple energy maps).
[0074] In other examples, controller 31 may be configured to use a password to determine the audio stream available for playback by content consuming device 14. Content consuming device 14 may be configured to issue the password to controller 31 (e.g., via side channel 33). In some examples, content consuming device 14 may be configured to receive one or more of a plurality of bitstreams 27 (e.g., audio streams) based on privacy restrictions associated with the password, and to generate a corresponding sound field based on one or more of the plurality of audio streams.
[0075] In some examples, controller 31 can be configured to generate (or cause other structural units of source device 12 to generate) one or more of multiple bitstreams 27 based on privacy restrictions associated with cryptography. Various cryptographic techniques can be performed or combined with privacy-restricted audio metadata techniques. Additional examples of privacy restrictions (e.g., permission states) are described herein. In these examples, certain privacy restrictions may affect the formation of the composite energy map used to determine PSI 46A or PSI 46B, where PSI parameter settings can be defined and adjusted.
[0076] Content consuming device 14A can be operated by an individual and can represent a VR client device. While a VR client device has been described, content consuming device 14A can represent other types of devices, such as augmented reality (AR) client devices, mixed reality (MR) client devices (or other XR client devices), standard computers, audio speakers, headphones, headsets, mobile devices (including so-called smartphones), or any other device capable of generating (e.g., reproducing) a sound field based on bitstream 27 (e.g., audio stream) and / or tracking head movements and / or general translational movements of an individual operating content consuming device 14A. Figure 1A As shown in the example, the content consumption device 14A includes an audio playback system 16A, which can refer to any form of audio playback system capable of rendering audio data 19' for playback of mono or multi-channel audio content.
[0077] Content consumption device 14A may include a user interface (UI). The UI may include one or more input devices and one or more output devices. Output devices may include, for example, one or more speakers, one or more display devices, one or more haptic devices, etc., configured to output information for user perception. Output devices may be integrated with content consumption device 14A or may be separate devices coupled to content consumption device 14A.
[0078] In some examples, the content consuming device 14A can provide a visual depiction of the energy map. In such examples, a user can manually identify problematic devices in the XR space. In the examples, the content consuming device 14A can provide a visual depiction of the energy map via a UI, indicating that a particular audio element in a constellation set of audio elements (e.g., a set of audio elements configured to capture a common sound field) is not functioning correctly and does not accept parameter adjustments or otherwise continues to generate bitstream 27 that, after parameter adjustments, does not have a corresponding energy map that conforms to the expected energy map (e.g., a composite energy map). In some examples, this type of non-compliance can indicate a calibration failure of the audio element.
[0079] One or more input devices may include any suitable device with which a user can interact to provide input to the content consumption device 14A. For example, one or more input devices may include a microphone, mouse, pointer, game controller, remote control, touchscreen, linear slider potentiometer, rocker switch, button, scroll wheel, knob, etc. In examples where one or more user input devices include a touchscreen, the touchscreen may allow selection of one or more capture device representations based on a single touch input (e.g., touch, swipe, tap, long press, and / or circle an area of the graphical user interface). In some implementations, the touchscreen may allow multi-touch input. In these examples, the touchscreen may allow selection of multiple areas of the graphical user interface based on multiple touch inputs.
[0080] Despite Figure 1AThe bitstream 27 is shown as being sent directly to content consumption device 14A, but source device 12A can output bitstream 27 to an intermediate device located between source device 12A and content consumption device 14A. The intermediate device can store bitstream 27 for later delivery to content consumption device 14A, which can request bitstream 27. The intermediate device can include a file server, web server, desktop computer, laptop computer, tablet computer, mobile phone, smartphone, or any other device capable of storing bitstream 27 for later retrieval or delivery to an audio decoding device (e.g., audio decoding device 34 of content consumption device 14A). The intermediate device can reside in a content delivery network capable of streaming bitstream 27 (and possibly in conjunction with sending a corresponding video data bitstream) to a subscriber requesting bitstream 27, such as by sending bitstream 27 to content consumption device 14A.
[0081] Alternatively, source device 12A may store bitstream 27 to a storage medium, such as an optical disc, digital video disc, high-definition video disc, or other storage medium, most of which are computer-readable and therefore may be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, a transmission channel may refer to the channel through which the content stored to the medium (e.g., in the form of one or more bitstreams 27) is transmitted (and may include retail stores and other store-based delivery mechanisms). In any case, the technology disclosed herein should therefore not be limited in this respect. Figure 1A Examples.
[0082] As noted herein, the content consumption device 14A includes an audio playback system 16A. The audio playback system 16A can represent any system capable of playing back mono and / or multi-channel audio data. The audio playback system 16A can include multiple different audio renderers 32. Each audio renderer 32 can provide different forms of rendering, which can include one or more of various methods of performing vector basis amplitude shift (VBAP) and / or one or more of various methods of performing sound field synthesis. As used herein, “A and / or B” means “A or B”, or “both A and B”.
[0083] The audio playback system 16A may also include an audio decoding device 34. Audio decoding device 34 may refer to a device configured to decode bitstream 27 to output audio data 19' (where the apostrophe may indicate that audio data 19' differs from audio data 19 due to lossy compression (such as quantization)). Audio decoding device 34 may be part of the same physical device as audio renderer 32, or it may be part of a physically separate device and configured to communicate with audio renderer 32 via a wireless or wired connection. Furthermore, audio data 19' may include scene-based audio data, in some examples of which may form a full first-order (or higher-order) surround sound representation or a subset of the full first-order (or higher-order) surround sound representation forming the same sound field, a decomposition of the full first-order (or higher-order) surround sound representation, such as the primary audio signal described in the MPEG-H 3D audio decoding standard or other forms of scene-based audio data, ambient surround sound coefficients, and vector-based signals (which may refer to a multidimensional spherical harmonic vector having multiple elements representing the spatial characteristics of the corresponding primary audio signal). Audio data 19' may include an audio stream or a representation of an audio stream.
[0084] In some examples, audio decoding device 34 can decode bitstream 27 according to PSI 46B. In the examples, audio decoding device 34 can determine a composite energy map according to CM 47, and can determine parameter adjustments for a specific audio element (e.g., source device 12A) and the bitstream 27 received from the audio element (e.g., an audio stream) based on the composite energy map. In an illustrative example, when decoding bitstream 27, audio decoding device 34 can adjust the frequency response of bitstream 27 to generate audio data 19' for subsequent audio rendering.
[0085] Other forms of scene-based audio data include audio data defined according to the HOA (Higher Order Ambisonics (HOA) Transport Format) standard. More information about HTF can be found in the technical specification (TS) entitled “Higher Order Ambisonics (HOA) Transport Format”, published by the European Telecommunications Standards Institute (ETSI) in ETSI TS 103 589 V1.1.1 in June 2018 (2018-06), and in U.S. Patent Application Publication No. 2019 / 0918028 entitled “PRIORITY INFORMATION FORHIGHER ORDER AMBISONIC AUDIO DATA” filed on December 20, 2018. In any case, audio data 19' may resemble the entire set or a subset of audio data 19, but may differ due to lossy operations (e.g., quantization) and / or transmission via the transmission channel.
[0086] Audio data 19' may include channel-based audio data as an alternative to scene-based audio data, or a combination thereof. Audio data 19' may include object-based audio data or channel-based audio data as an alternative to scene-based audio data, or a combination thereof. Thus, audio data 19' may include any combination of scene-based audio data, object-based audio data, and channel-based audio data.
[0087] The audio renderer 32 of the audio playback system 16A can render the audio data 19' to output the speaker feed 35 after the audio decoding device 34 has decoded the bitstream 27 to obtain the audio data 19'. In some examples, the audio data 19' may include PSI 46B. In such examples, the audio renderer 32 can render the audio data 19' according to the PSI 46B. The speaker feed 35 can drive one or more speakers or headphones (for illustration purposes, ...). Figure 1A (Not shown in the example). Various audio representations (including scene-based audio data of the sound field (and possibly channel-based and / or object-based audio data) can be normalized in a variety of ways (including N3D, SN3D, FuMa, N2D, or SN2D). In the example, the audio renderer 32 can normalize the sound field based on PSI 46B. In this way, the audio renderer 32 can further provide a sound field with uniform parameter settings so that the user does not perceive gain differences when listening to audio streams of different audio elements.
[0088] To select an appropriate renderer, or in some instances, to generate an appropriate renderer, the audio playback system 16A may obtain speaker information 37 indicating the number of speakers (e.g., loudspeakers or headphone speakers) and / or the spatial geometry of the speakers. In some instances, the audio playback system 16A may use a reference microphone to obtain the speaker information 37 and may drive the speakers in a manner that dynamically determines the speaker information 37 (which may refer to the output of an electrical signal to induce transducer vibration). In other instances, or in conjunction with the dynamic determination of the speaker information 37, the audio playback system 16A may prompt a user to interact with the audio playback system 16A and input the speaker information 37.
[0089] Audio playback system 16A may select one of the audio renderers 32 based on speaker information 37. In some instances, audio playback system 16A may generate one of the audio renderers 32 based on speaker information 37 when no audio renderer 32 is within a certain threshold similarity metric (in terms of speaker geometry) of the speaker geometry specified in speaker information 37. In some instances, audio playback system 16A may generate one of the audio renderers 32 based on speaker information 37 without first attempting to select an existing audio renderer from the audio renderers 32. In some examples, speaker information 37 (such as speaker volume or number) may lead to further adjustments to audio data 19', where audio data 19' includes PSI 46B. In examples, audio renderer 32 may apply a version of PSI 46B to audio data 19' to suit a specific speaker configuration and / or speaker settings.
[0090] When the speaker feed 35 is output to the headphones, the audio playback system 16A can utilize one of the audio renderers 32 to provide binaural rendering using a Head-Related Transfer Function (HRTF) or other functions capable of rendering to the left and right speaker feeds 35 for playback by the headphone speakers, such as a binaural room impulse response renderer. The term "speaker" or "transducer" can generally refer to any speaker, including loudspeakers, headphone speakers, bone conduction speakers, earbud speakers, wireless headphone speakers, etc. One or more speakers or headphones can then play back the rendered speaker feed 35 to generate a sound field. In the example, one or more speakers can be placed near a user and can generate a sound field configured to immerse the user in the sound field or depict a sound field at a location near the user where the user can perceive the sound field emanating from various different locations as defined by the audio data 19'.
[0091] Although described as rendering speaker feed 35 from audio data 19', the reference to rendering speaker feed 35 can refer to other types of rendering, such as rendering directly incorporated into the decoding of audio data from bitstream 27. Examples of alternative rendering can be found in Annex G of the MPEG-H 3D audio decoding standard, where rendering occurs during the formation of the main signal and the background signal formation prior to sound field composite. Therefore, the reference to rendering audio data 19' should be understood to refer to rendering or decomposition or representation of the actual audio data 19' (such as the aforementioned main audio signal, ambient surround sound coefficients, and / or vector-based signals – also referred to as V-vectors or multidimensional surround sound space vectors).
[0092] The audio playback system 16A can also adjust the audio renderer 32 based on the tracking information 41. That is, the audio playback system 16A can interface with a tracking device 40 configured to track the head movements and possible translational movements of the user of the VR device. The tracking device 40 can represent one or more sensors (e.g., cameras—including depth cameras, gyroscopes, magnetometers, accelerometers, light-emitting diodes—LEDs, etc.) configured to track the head movements and possible translational movements of the user of the VR device. The audio playback system 16A can adjust the audio renderer 32 based on the tracking information 41 such that the speaker feed 35 reflects changes in the user's head movements and possible translational movements to generate a sound field in response to such movements.
[0093] Figure 1B This is a block diagram illustrating another example system 50 configured to perform various aspects of the techniques described in this disclosure. Besides Figure 1A The audio renderer 32 shown is replaced by a binaural renderer 42 capable of performing binaural rendering using one or more HRTFs or other functions capable of rendering to the left and right speaker feeds 43 (in the audio playback system 16B of the content consumption device 14B), and the system 50 is similar to Figure 1A System 10 shown.
[0094] The audio playback system 16B can output the left and right speaker feeds 43 to the headset 48. The headset 48 represents another example of a wearable device that can be coupled to additional wearable devices, such as XR devices (e.g., VR headsets), smart glasses, smart clothing, smart jewelry (e.g., watches, rings, bracelets, necklaces, etc.), to facilitate the generation (e.g., reproduction) of a sound field. The headset 48 can be coupled to the additional wearable device wirelessly or via a wired connection.
[0095] Additionally, the headphones 48 can be connected via wired connections (such as a standard 3.5mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other forms of wired connection) or wirelessly (such as via...). The headset 48 is coupled to the audio playback system 16B via a connection (such as a wireless network connection). The headset 48 can generate a sound field represented by audio data 19' based on the left and right speaker feeds 43. The headset 48 may include a left headset speaker and a right headset speaker, which are powered (or in other words, driven) by the corresponding left and right speaker feeds 43.
[0096] Figure 1C This is a block diagram illustrating another example system 60. Example system 60 is similar to... Figure 1AExample system 10, but the source device 12B of system 60 does not include a content capture device. Source device 12B includes a compositing device 29. Content developers can use compositing device 29 to generate composite audio sources. The composite audio source can have associated location information that identifies the position of the audio source relative to the listener or other reference points in the sound field, allowing the audio source to be rendered to one or more speaker channels for playback in an effort to generate the sound field. In some examples, compositing device 29 can also composite visual or video data.
[0097] For example, content developers can generate synthesized audio streams for video games. Although Figure 1C Examples and Figure 1A The example content consumption device 14A is shown together, but Figure 1C Example source device 12B can be with Figure 1B It is used together with the content consumption device 14B. In some examples, Figure 1C The source device 12B may also include a content capture device, such that the bitstream 27 may contain both a captured audio stream and a synthesized audio stream.
[0098] As described above, content consumption device 14A or 14B (either of which may be referred to as content consumption device 14 below) can refer to a VR device in which a human wearable display (which may also be referred to as a "head-mounted display") is mounted in front of the user operating the VR device. Figure 2 This is a diagram illustrating an example of a VR device 204. In this illustrative example, the VR device 204 is depicted as a headset worn by a user 202. While described as such, the technology disclosed herein is not limited thereto, and those skilled in the art will understand that VR devices can take many different forms. In the example, the VR device 204 may include one or more speakers (e.g., a headset worn by the user 202, an external speaker assembly, one or more mountable speakers, etc.).
[0099] In some examples, the VR device 204 is coupled to or otherwise includes a headset 206, which can generate a sound field represented by audio data 19' through playback of speaker feed 35. Speaker feed 35 can represent analog or digital signals that enable the diaphragm within the transducer of the headset 206 to vibrate at various frequencies, a process commonly referred to as driving the headset 206.
[0100] Video, audio, and other sensor data can play a significant role in XR experiences. For example, to participate in a VR experience, user 202 can wear VR device 204 (also referred to as a VR client device) or other wearable electronic devices. The VR client device (such as VR device 204) can include a tracking device (e.g., tracking device 40) configured to track the user 202's head movements and adjust the video data displayed via VR device 204 to account for head movements, providing an immersive experience where user 202 can experience an acoustic space, a displayed world, or both. The displayed world can refer to a virtual world (where all worlds are simulated), an augmented world (where parts of the world are augmented by virtual objects), or a physical world (where real-world images are virtualized for navigation).
[0101] While VR (and other forms of AR and / or MR) allows users 202 to visually reside in a virtual world, VR devices 204 may typically lack the ability to audibly place the user in an acoustic space. In other words, a VR system (which may include a computer responsible for rendering video and audio data - not shown for illustration) Figure 2 As shown in the example, VR device 204 may not audibly support full 3D immersion (and in some instances may not realistically support it in a way that reflects the scene presented to the user via VR device 204).
[0102] While VR devices have been described in this disclosure, various aspects of these technologies can be performed in the context of other devices, such as mobile devices, speakers, audio elements (e.g., microphones, synthesized audio sources, etc.) or other XR devices. In this example, a mobile device can render an acoustic space (e.g., via speakers, one or more headsets, etc.). The mobile device, or at least a portion thereof, can be mounted on the head of user 202 or viewed as is done when using a mobile device normally. Thus, any information on the screen can be part of the mobile device, as well as any information generated via speakers, headsets, or audio elements. The mobile device may be able to provide tracking information 41, thereby allowing both a VR experience (when mounted on the head) and a normal experience to experience the acoustic space, where the normal experience can still allow the user to experience the acoustic space that provides a simplified VR experience (e.g., raising the device and rotating or panning it to view different parts of the displayed world).
[0103] The audio aspects of XR have been categorized into three distinct immersive categories. The first category offers the lowest level of immersion and is known as three degrees of freedom (3DOF). 3DOF refers to audio rendering that takes head movement into account in 3DOF (yaw, pitch, and roll), thus allowing users to freely look around in any direction. However, 3DOF cannot account for translational head movements that are not centered on the optical and acoustic center of the sound field.
[0104] The second category (called 3DOF plus (3DOF+)) provides 3DOF (yaw, pitch, and roll) in addition to the limited spatial translation caused by head movement away from the optical and acoustic centers within the sound field. 3DOF+ can support perceptual effects such as motion parallax, which can enhance immersion.
[0105] The third category (called six degrees of freedom (6DOF)) renders audio data in a way that takes into account 3DOF in terms of head movements (yaw, pitch, and roll) as well as the user's translations in space (x, y, and z translations). Spatial translations can be induced by sensors that track the user's position in the physical world or by input controllers.
[0106] 3DOF rendering is the latest technology in VR audio. Therefore, VR audio immersion is not as strong as video immersion, potentially reducing the overall immersive experience. However, VR is rapidly evolving and may quickly develop to support both 3DOF+ and 6DOF, which could open up opportunities for other use cases.
[0107] For example, interactive gaming applications can leverage 6DOF to facilitate fully immersive gaming, where users move freely within the VR world and interact with virtual objects by walking towards them. Furthermore, interactive live streaming applications can utilize 6DOF to allow VR client devices to experience live streams of concerts or sporting events as if they were actually at the concert, allowing users to move freely within the concert or sporting event.
[0108] There are several challenges associated with these use cases. In fully immersive games, latency may need to be kept low to ensure gameplay that doesn't cause nausea or motion sickness. Furthermore, from an audio perspective, audio playback latency that causes desynchronization with video data can reduce immersion. Additionally, for certain types of game applications, spatial accuracy can be important for allowing accurate responses, including responses to how the user perceives sound, as this allows the user to predict actions that are not currently in their field of vision.
[0109] In the context of a live streaming application, a large number of source devices 12A or 12B (either of which may be referred to as source device 12 hereinafter) can stream content 21, where source devices 12 can have a wide range of different capabilities. For example, one source device 12 may be a smartphone with a digital fixed-lens camera and one or more microphones, while another source device may be a production-grade television setup capable of delivering video at a higher resolution and quality than a smartphone. However, in the context of a live streaming application, all source devices 12 can provide streams of different qualities, from which the VR device can attempt to select a suitable stream to deliver the desired experience.
[0110] Furthermore, similar to gaming applications, audio data latency, leading to desynchronization with video data, can reduce immersion. Spatial accuracy can also be important, allowing users to better understand the context or location of different audio sources. Additionally, privacy can become an issue when users are live-streaming using cameras and audio components (e.g., microphones), as users may not want their streams to be completely public.
[0111] In the context of modulating the audio stream, audio components (e.g., XR devices, audio receiving devices, audio synthesis devices, etc.) can be fitted with various parameter settings, such as gain, frequency response, and / or other adjustment settings, to modify the audio capture and generate an immersive XR experience. In some instances, the parameter settings of the audio receiving device can be fully compensated to allow the audio decoding device 34 to generate a sufficient sound field for the XR environment based on the audio stream.
[0112] In some examples, the parameter settings of an audio element may initially be incompatible or discriminatory with other audio elements, source devices, or accessory devices (such as wearable devices, mobile devices, etc.) that may also be equipped with audio receiving devices (e.g., microphones) that apply various parameter settings. For example, the parameter settings of one audio element may not initially correspond to the parameter settings of another audio element, resulting in poor audio capture and distorted sound field performance.
[0113] This lack of coordination can be particularly noticeable when users manually change parameter settings (e.g., adjusting gain for high-frequency sounds for an audio receiving device or XR device, or when an audio element system includes audio elements from different manufacturers or suppliers). In examples involving microphones as at least one audio element, it's possible that not all microphones across the microphone constellation streaming audio to the XR device have the same gain, frequency response, or other parameters, or may not be compensated for, thus not coordinating with other microphones or other audio elements in the same constellation. In another example, not all microphones across the microphone constellation streaming audio to the XR device can be compensated to provide a noise-free audio stream, or audio elements may not be compensated for to coordinate with audio elements in other audio element constellations. That is, if audio elements are not properly compensated for, for example, through parameter adjustments (e.g., equalization, calibration, etc.), they may fail to provide an immersive XR experience and could ultimately disorient or confuse users utilizing XR devices.
[0114] According to the technology described in this disclosure, audio decoding device 34 can access an energy map corresponding to an audio stream (represented by bitstream 27 and therefore referred to as "audio stream 27") available via bitstream 27. Audio decoding device 34 can utilize the energy map to determine parameter adjustments, such as gain or frequency response adjustments, for at least one audio element (e.g., an audio receiving device, such as a microphone or other receiver, or an audio generating device, such as a virtual speaker or other virtual device configured to synthesize an audio sound field in a virtual environment). In this way, audio decoding device 34 can automatically adjust the audio element when a difference in parameter settings is detected. Audio decoding device 34 can apply the parameter adjustments as adjusted parameter settings to modify audio capture and compensate for differences in applied parameter settings for the audio element. Parameter adjustments can be determined based on an analysis of an energy map relative to a reference energy map (e.g., a baseline energy map, a composite energy map, etc.). In some examples, parameter adjustments can be configured to shift a target energy map according to the parameter adjustments, such that the target energy map overlaps with the reference energy map as much as possible.
[0115] Audio elements may be able to improve both audio spatialization accuracy and 6DOF rendering quality through parameter tuning. In operation, the audio decoding device 34 can interface with one or more source devices 12 to determine the parameter tuning for each audio element. For example... Figures 1A to 1CAs shown in the examples, audio playback system 16A or 16B may include audio decoding device 34, while source device 12A or 12B may include content editing device 22 or synthesis device 29. These systems and devices may individually or collectively represent one or more audio elements configured to perform various aspects of the audio compensation techniques described in this disclosure.
[0116] In some instances, audio elements (e.g., microphones, audio field synthesizers, or other XR devices) may include source devices 12A or 12B, in which case content editing device 22 or synthesis device 29 may determine and apply parameter adjustments based on one or more energy maps corresponding to the bitstream 27. In such instances, content editing device 22 may be configured to perform some or all aspects of adjustment techniques for compensating source devices 12A or 12B (which may generally be referred to herein as "source device 12"). Similarly, in instances where content consuming devices 14A or 14B (which may generally be referred to herein as content consuming device 14) include one or more audio elements (e.g., microphones, etc.), audio decoding device 34 may perform some or all aspects of adjustment techniques for compensating content consuming device 14. In some instances, source device 12 and content consuming device 14 may be integrated into audio elements (e.g., standalone microphones). For example, content consuming device 14 may include an XR device, while the source device may include a microphone 18 that interfaces with the XR device as an audio element. In a non-limiting example, an XR device may include a microphone attachment that captures the user's voice as the user navigates or otherwise experiences the XR space.
[0117] Furthermore, in some examples, the audio decoding device 34 can determine operational status information (e.g., diagnostic data) of one or more audio elements and / or may receive operational status information from one or more audio elements. Similarly, audio elements may send operational status information (e.g., self-diagnostic data) to the audio decoding device 34. The operational status information can provide information about the quality of the received audio signal, the permission status of the audio element (e.g., permission to access the audio element's energy map, permission to access the audio stream of the audio element, etc.), and / or other feasibility characteristics. For example, the operational status information may include SNR information or gain information. Operational status information may indicate that the audio element is currently inactive, such as by being detected in a user's pocket without receiving a clear audio stream or based on accelerometer, light detection, or other sensor data.
[0118] Operational status (e.g., diagnostic data, feasibility data, etc.) can also indicate the permission status of audio components (e.g., privacy settings). For example, operational status information can indicate that a particular audio component is restricting the transmission of a certain amount or type of audio data. Permission status can indicate whether one or more audio streams are restricted or unrestricted. In some examples, privacy settings can refer to digital access permissions that restrict access to one or more of the bitstreams 27, for example, through passwords, authorization levels or grades, time, etc.
[0119] Operational status information can also indicate that using a specific audio element would be infeasible when determining parameter adjustments. For example, a specific audio element may not be configured to allow manipulation of parameter settings used to adjust audio capture. In other words, the audio element may not have configurable settings and may only be able to apply a single parameter setting that can be programmed by the manufacturer. In another example, the feasibility status can indicate the microphone's position relative to the audio decoding device 34. When the microphone is too far from the audio decoding device 34, the audio decoding device 34 can determine that using the audio stream from that specific microphone to determine parameter adjustments for another audio element (e.g., an audio element corresponding to the audio decoding device 34) would be infeasible.
[0120] In some examples, operational status information may include tracking information (e.g., to determine whether the user is facing the audio source 308). In such examples, the audio decoding device 34 may use the tracking information to determine the feasibility status.
[0121] The audio decoding device 34 can exclude at least one bitstream 27 (e.g., at least one audio stream) and / or energy map from the energy map set, at least in part, based on operational state information, such that the excluded audio stream does not contribute to the parameter adjustment determination. For example, an audio stream from a defective microphone can be excluded from the set of bitstreams 27 used to determine the parameter adjustment.
[0122] Audio decoding device 34 can perform energy analysis on each of the bitstreams 27 to determine an energy map for each of the bitstreams 27, and store the energy maps in CM 47. The energy maps can collectively define the energy of a common sound field represented by the bitstreams 27. In some instances, audio decoding device 34 can receive one or more energy maps from source device 12. Alternatively, audio decoding device 34 can generate a single energy map based on multiple bitstreams 27. In other instances, audio decoding device 34 can aggregate multiple energy maps corresponding to multiple bitstreams 27 (e.g., multiple audio streams) to determine a composite energy map. In some instances, a single energy map may include a composite energy map corresponding to multiple energy maps.
[0123] In some examples, audio decoding device 34 may receive one or more and / or other individual energy maps from a composite energy map from source device 12. In some instances, a single energy map (such as a composite energy map) may include multiple energy map components. For example, energy map components may be based on energy analysis of one or more audio streams. In another example, an energy map component may be another energy map. For example, multiple energy maps may form a single composite energy map (e.g., multiple energy maps may be fused together or synthesized into a single composite energy map). In this way, an energy map may include multiple energy map components corresponding to one or more audio streams, wherein the energy map components may be audio streams or energy maps corresponding to audio streams.
[0124] Audio decoding device 34 can store one or more of the composite energy map, individual energy maps, or bitstreams 27 into a memory for subsequent access and analysis. For example, audio playback systems 16A or 16B can be configured to store the energy map and audio stream into CM 47. In some examples, source device 12 can be configured to store the energy map and audio stream into its own memory device. Source device 12 can be configured to send the energy map and / or audio stream to content consumption device 14.
[0125] Audio decoding device 34 can analyze one or more energy maps to determine parameter adjustments for audio elements (e.g., one of the microphones 18, etc.). For example, audio decoding device 34 can analyze gain or frequency response defined by one or more energy maps to determine whether parameter adjustments are needed to equalize the microphone and / or sound generation device. In some instances, audio decoding device 34 can perform energy map comparisons to determine whether parameter settings need adjustment (e.g., increase or decrease of gain). In one example, audio decoding device 34 can compare energy maps derived from multiple audio streams with another energy map to determine appropriate adjustments to parameter settings for audio elements such as microphones to adequately compensate for insufficient audio capture from the audio elements. In some examples, audio decoding device 34 can compare a composite energy map with another energy map, where the other energy map was also used to generate the composite energy map. In some examples, audio decoding device 34 can utilize long-term energy maps to handle transient situations, such as shouts or loud engines.
[0126] Audio decoding device 34 can analyze differences in the energy graph and determine parameter adjustments to compensate for these differences. In some examples, audio decoding device 34 can identify differences, such as discontinuities in the audio stream, while analyzing the energy graph. For example, audio decoding device 34 can detect gaps in the frequency response of the audio stream from a comparison of the energy graphs. Audio decoding device 34 can determine adjustments to the parameter settings of the audio receiving device to remove or compensate for the differences based on the analysis of the energy graph. In some instances, a remote server can analyze the energy graph and determine adjustments, in which case the remote server sends the adjustments to one or more audio elements. Alternatively, the remote server can generate one or more energy graphs based on bitstream 27, including composite energy graphs.
[0127] Audio decoding device 34 can output instructions indicating adjustments to one or more parameters, including details of the parameter adjustments such as how much gain is applied, whether the gain is specified to a certain frequency range, compression settings, frequency response settings, and / or other settings configured to adjust the parameter settings of audio components in the presence of other audio components (e.g., physically or virtually present). For example, when compensating audio components, the settings can be configured to, for example, coordinate, optimize, equalize, calibrate, normalize, modify, or otherwise increase compatibility between audio components by adjusting audio capture or audio generation. In one example, according to some techniques disclosed herein, one or more processors can optimize parameter settings for an XR experience. In such examples, an XR device can then be configured to provide a balanced, immersive experience after parametric adjustments (e.g., gain adjustment, frequency response adjustment, enabling, disabling, etc.) to the audio elements in the constellation of audio elements. Additionally, the XR device will then require less processing for damaged audio elements (e.g., when disabled), and in some examples, may also require less processing when forming a composite energy map, such as when using only energy maps with high or relatively high SNR, thus those energy maps can be used to form a composite energy map.
[0128] Parameter adjustments can be implemented to apply them to captured audio. For example, implementation may include changing the microphone's current parameter setting to another parameter setting based on the parameter adjustment. Implementation of parameter adjustment may include replacing the current parameter setting with another parameter setting based on the parameter adjustment. In a non-limiting example, parameter adjustment may be one of adjustments to the gain of one or more microphones. One or more processors of the microphones may determine or receive a parameter adjustment instruction and implement the gain adjustment, thereby achieving parameter adjustment.
[0129] Parameter adjustments can be applied to audio captured by a microphone (such as microphone 18) during capture or before generation (e.g., at the sound source). For example, audio decoding device 34 can utilize one or more parameter settings (adjusted or unadjusted) to regulate audio received via one or more microphones. Parameter adjustments can be determined only for a specific microphone having an audio stream energy map that does not correspond to one or more other energy maps corresponding to audio captured by one or more other microphones.
[0130] In some instances, audio decoding device 34 can send parameter adjustments to source device 12, allowing source device 12 to utilize these adjustments when generating audio (e.g., streaming content from source device 12 to content consumption device 14). In any case, a specific parameter adjustment will correspond to a microphone that differs in the energy map compared to one or more other energy maps, such that the parameter adjustment is configured to optimize the specific microphone and compensate for or remove the differences. When differences are removed by implementing the adjusted parameter settings, the microphone will be able to generate a sound field that more closely resembles a mirror image of the true sound field.
[0131] In some examples, parameter adjustments will correspond to the microphone capturing the audio stream based on the degree to which the energy map corresponding to the audio captured from the microphone deviates from a normal or reference energy map. In some instances, the audio decoding device 34 may also receive identification details about the microphones in the system, such as model number, manufacturer, etc., which can help the audio decoding device 34 determine parameter adjustments. For example, the audio decoding device 34 may apply a confidence score where the identification details indicate that the microphones are from different OEMs (Original Equipment Manufacturers). In such instances, the initial parameter settings are more likely to be incompatible with each other, and one or two initial parameter settings should be adjusted (e.g., through normalization, equalization, or calibration) to be mirrored or closer to a mirror image of the true sound field.
[0132] Based at least in part on parameter settings (adjusted or unadjusted), audio decoding device 34 can output bitstream 27 as audio data 19'. Additionally, audio decoding device 34 can select which bitstreams 27 to adjust according to parameter settings based on quality characteristics. Audio decoding device 34 can apply parameter settings to adjust one or more of the bitstreams 27 to generate audio data 19'. In such instances, audio decoding device 34 can generate audio data 19' that includes parameter setting information or already has parameter setting information applied to audio data 19'. Audio playback system 16A or 16B can then use audio data 19' to generate a sound field. In another example, when rendering audio data 19' to speaker feed 35 or speaker feed 43, audio renderer 32 or binaural renderer 42 can apply parameter settings to audio data 19'.
[0133] In some examples, the audio decoding device 34 can generate an energy curve overlay based on an energy map. In some examples, the energy curve overlay can be based on a composite energy map. The energy curve overlay can also include an overlay of a specific energy map as it is compared with a composite energy map or another energy map corresponding to a different audio stream. The audio decoding device 34 can provide energy curves for display to a user. For example, the audio decoding device 34 can output an energy curve overlay or output the overlay as part of a UI. Thus, the audio decoding device 34 can be configured to generate UI data, which includes energy map data, audio data, and / or parameter setting data. The audio decoding device 34 can work in conjunction with a UI generation device to cause a UI to be displayed on an XR device or another audio element (e.g., the microphone of a mobile phone).
[0134] In some examples, the audio decoding device 34 can be configured to test the success of adjustments to one or more stray microphones to synchronize the microphones with other audio rendering devices (e.g., content consuming device 14) and / or receiving devices (e.g., source device 12). The audio decoding device 34 can monitor newly configured microphones to determine whether the microphones are receiving audio based on parameter adjustments. The audio decoding device 34 can output a signal indicating successful adjustment. However, in other instances, the audio decoding device 34 can output a warning signal indicating unsuccessful adjustment. The warning signal can be based on another comparison of the energy map after parameter adjustments were performed. The signal can take the form of a notification displayed on the UI. For example, the audio decoding device 34 can generate feedback for a live streamer regardless of whether the streamer's audio is corrupted or whether the adjustment of audio components (e.g., microphones, etc.) was successful. Similarly, feedback can be given to the user regarding successful XR device calibration. For example, see the following reference... Figures 3A to 3D , Figures 4A to 4B and Figures 5A to 5D This section discusses more information about how audio decoding device 34 can adjust audio capture.
[0135] Figures 3A to 3D To show in more detail Figures 1A to 1C The example shown is a diagram illustrating the example operation of the flow selection unit 44. Figure 3A As shown in the example, the flow selection unit 44 can determine device location information (DLI) (e.g., Figures 1A to 1C 45B indicates that the content consumption device 14 (shown as VR device 204) is at virtual location 300A. The stream selection unit 44 can then determine audio location information (ALI) 45A for one or more of the audio elements 302A to 302J (collectively referred to as audio elements 302), which may represent not only the microphone, but also other elements such as... Figure 1AOr the microphone 18 shown in 1B, and also represents other types of capture devices, including other XR devices, mobile phones—including so-called smartphones—etc., or devices that can define a synthetic sound field, such as Figure 1C The audio data 19 generated by the synthesis device 29 based on the PSI 46A.
[0136] The stream selection unit 44 can then obtain an energy map in the manner described above, analyze the energy map to determine the audio source location 304, which can represent... Figures 1A to 1C The example shown is an example of the ASL49. An energy map can represent the audio source location 304. In this example, the stream selection unit 44 can represent the audio source location 304 based on at least one energy map (e.g., a composite energy map), where the energy at the audio source location 304 may be higher than the surrounding area. That is, the stream selection unit 44 can determine a higher energy location based on at least one energy map and can determine the audio source location 304 as corresponding to a higher energy location (e.g., a virtual or physical location). In some examples, the stream selection unit 44 can represent the audio source location 304 based on multiple energy maps. Given that each of the energy maps can represent that higher energy corresponding to the audio source location 304, the stream selection unit 44 can triangulate the audio source location 304 based on the higher energy in the energy maps.
[0137] Next, the streaming selection unit 44 can determine the audio source distance 306A as the distance between the audio source location 304 and the virtual location 300A of the VR device 204. The streaming selection unit 44 can compare the audio source distance 306A with an audio source distance threshold. In some examples, the streaming selection unit 44 can derive the audio source distance threshold based on the energy of the audio source 308. That is, when the audio source 308 has high energy (or in other words, when the audio source 308 is noisy), the streaming selection unit 44 can increase the audio source distance threshold. When the audio source 308 has low to high energy (or in other words, when the audio source 308 is quiet), the streaming selection unit 44 can decrease the audio source distance threshold. In other examples, the streaming selection unit 44 can obtain a statically defined audio source distance threshold, which can be statically defined or specified by the user (e.g., user 202).
[0138] In any case, stream selection unit 44 may select a single audio stream of bitstream 27 of audio elements 302A to 302J (“audio element 302”) when the audio source distance 306A is greater than an audio source distance threshold (assumed for illustrative purposes in this example). For example, stream selection unit 44 may select the audio element with the shortest distance to virtual location 300 (e.g., Figure 3AThe example audio element 302A) outputs bitstream 27. Stream selection unit 44 can output the corresponding bitstream in bitstream 27, which audio decoding device 34 can decode and output as audio data 19'.
[0139] Suppose a user (e.g., user 202) moves from virtual location 300A to virtual location 300B, the stream selection unit 44 can determine the audio source distance 306B as the distance between the audio source location 304 and the virtual location 300B. In some examples, the stream selection unit 44 can update only after some configurable release time, which can refer to the time after the listener stops moving.
[0140] In any case, the stream selection unit 44 can again compare the audio source distance 306B with the audio source distance threshold. When the audio source distance 306B is less than or equal to the audio source distance threshold (assumed for illustrative purposes in this example), the stream selection unit 44 can select multiple audio streams of the bitstream 27 of the audio element 302. The stream selection unit 44 can output the corresponding bitstream from the bitstream 27, which the audio decoding device 34 can decode and output as audio data 19'.
[0141] Stream selection unit 44 can also determine one or more proximity distances between virtual position 300A and one or more (and possibly each) of capture positions (or synthesis positions) represented by ALI 45A to obtain one or more proximity distances. Stream selection unit 44 can then compare one or more proximity distances with a threshold proximity distance. Stream selection unit 44 can select fewer bitstreams 27 when one or more proximity distances are greater than the threshold proximity distance compared to when audio data 19' is obtained when one or more proximity distances are greater than the threshold proximity distance. However, stream selection unit 44 can select a larger number of bitstreams 27 when one or more proximity distances are less than the threshold proximity distance compared to when audio data 19' is obtained when one or more proximity distances are greater than the threshold proximity distance.
[0142] In other words, the stream selection unit 44 can attempt to select those bitstreams in the bitstream 27 such that the audio data 19' most closely matches and surrounds the virtual position 300B. A proximity threshold can be defined, which the user 202 can set, or the stream selection unit 44 can re-determine the threshold based on the quality of the audio elements 302F to 302J, the gain or loudness of the audio source 308, tracking information 41 (e.g., to determine whether the user 202 is facing the audio source 308), or any other factor.
[0143] In this respect, when the listener is at position 300B, the stream selection unit 44 can increase the accuracy of audio spatialization. Furthermore, when the listener is at position 300A, the stream selection unit 44 can reduce the bit rate because it uses only the audio stream of audio element 302A instead of multiple audio streams from audio elements 302B to 302J to generate the sound field.
[0144] Next reference Figure 3B For example, stream selection unit 44 can determine that the audio stream of audio element 302A is corrupted, noisy, or unavailable. Given that the audio source distance 306A is greater than an audio source distance threshold, stream selection unit 44 can remove the audio stream from CM 47 according to the techniques described in more detail above and iterate through bitstream 27 to select a single bitstream from bitstream 27 (e.g., in...). Figure 3B In the example, the audio stream of audio element 302B.
[0145] Next reference Figure 3C For example, stream selection unit 44 can obtain a new audio stream (audio stream of audio element 302K) including ALI 45A and corresponding new information (e.g., metadata). Stream selection unit 44 can add the new audio stream to CM 47 representing bitstream 27. Given that the audio source distance 306A is greater than the audio source distance threshold, stream selection unit 44 can then iterate through bitstream 27 according to the technique described in more detail above to select a single bitstream in bitstream 27 (e.g., in...). Figure 3C In the example, the audio stream of audio element 302B.
[0146] exist Figure 3D In the examples, audio element 302 is replaced by specific example devices 320A to 320J (“Device 320”), where Device 320A represents a dedicated microphone 320A, and Devices 320B, 320C, 320D, 320G, 320H, and 320J represent smartphones. Devices 320E, 320F, and 320I may represent XR devices (e.g., VR devices). Each of Devices 320 may include audio element 302, which captures or synthesizes bitstream 27 (e.g., audio stream) selected or excluded according to various aspects of the stream selection and parameter adjustment techniques described in this disclosure.
[0147] In some examples, device 320 may also include one or more audio speakers. Although not in Figure 3D The example is shown, but it should be understood that Figure 3DIt may also include audio elements 302 corresponding to the generated audio source, such as audio generated via a computer program. The parameter settings of the audio elements (such as microphones or synthesized audio sources) configured to generate audio data 19 can be adjusted based on energy map analysis and comparison of each corresponding audio element, so that user 202 experiences the audio in a manner that strictly matches the intended audio experience (e.g., equalized audio). In the example, user 202 may speak into microphone 320A, and another user may then be able to hear user 202 speaking through headphones or other speaker devices. When microphone 320A is not compensated for other sounds (e.g., generated audio streams) in the other user's XR environment (e.g., at a virtual concert), the other user may experience or perceive user 202's spoken words as having inappropriate volume or gain, or as a noisy or distorted signal, which could then lead to an unpleasant experience for the other user with the XR system.
[0148] In an illustrative example, audio decoding device 34 may compensate microphone 320A of user 202 based on energy maps of one or more sounds in an XR environment (e.g., energy maps of each microphone that captures audio in a concert environment), such that when generating a sound field for another user, audio decoding device 34 can generate an audio stream of user 202’s speech during joint viewing of a virtual concert, where the other user can hear user 202 speaking after various gain adjustments to microphone 320A.
[0149] In some examples, audio decoding device 34 can compensate microphone 320A by sending parameter adjustments (e.g., via side channel 33), at least one audio stream for generating an energy map, or at least one energy map (e.g., a composite energy map) to source device 12. Source device 12 can implement parameter adjustments to compensate for one or more differences between the energy map corresponding to the audio element (e.g., microphone 18, synthesizer 29, etc.) and at least one other energy map (e.g., a composite energy map). In another example, the sound field representation generator 24 of source device 12 or the audio decoding device 34 of content consumption device 14 can adjust the audio data 19 generated from the audio element based on energy map analysis to ultimately generate audio data 19' representing an energy map that matches one or more other energy maps (including a composite energy map). In this illustrative example, another user may be physically separated from user 202, but in the virtual world, may be sitting next to user 202 and enjoying the same concert.
[0150] Figure 4A It is shown Figures 1A to 1C The example shown is a flowchart illustrating the operation of the audio decoding device 34 when performing various aspects of parameter adjustment techniques. Figure 4AIn this example, audio decoding device 34 can obtain bitstream 27 (e.g., an audio stream) from all enabled audio elements in a specifically defined set of audio elements (such as a constellation set defined by the proximity of the audio elements to content consuming device 14 or to the spatial proximity of the sound field of the audio). In this example, audio decoding device 34 can obtain bitstream 27 from each audio receiving device (e.g., this refers to another way of referring to microphones such as microphone 18). Bitstream 27 may include corresponding information (e.g., metadata). Audio decoding device 34 can perform energy analysis on each of the bitstream 27 to calculate the corresponding energy map and store the energy map in a storage location, such as CM47.
[0151] In some examples, the audio decoding device 34 may access at least one energy map (402). At least one energy map may include a composite energy map formed by each of the respective energy maps stored in a memory location. In some instances, accessing at least one energy map may include the audio decoding device 34 receiving at least one energy map from another device, receiving more than one energy map from another device, generating one or more energy maps based on an audio stream, generating a composite energy map, obtaining an audio stream and generating an energy map therefrom, or any combination thereof.
[0152] The audio decoding device 34 can then use the accessed energy map to determine parameter adjustments (404) for audio elements such as microphones or synthesizers 29. In some examples, the audio decoding device 34 can compare a first energy map corresponding to the audio of the first audio element with a comparison energy map. As discussed above, the comparison energy map may be based on one or more energy maps and may include the first energy map or may be a composite energy map based on one or more energy maps, which may or may not include the first energy map.
[0153] In some examples, the audio decoding device 34 can determine a difference score for a comparison of energy maps. For example, the difference score can represent the degree to which a particular audio element deviates from a baseline energy map (e.g., a composite energy map). In some examples, the audio decoding device 34 can adjust the difference score based on the difference in the comparison. For example, when there is a discontinuity regarding a first energy map or a discontinuity between the first energy map and one or more other energy maps, the audio decoding device 34 can increase the difference score, indicating a discontinuity regarding an audio stream or multiple audio streams. The audio decoding device 34 can compare the difference score to a difference threshold to determine whether parameter adjustment is needed. In other examples, the audio decoding device 34 can use the difference score to determine parameter adjustment regardless of whether the score exceeds a difference threshold. For example, the audio decoding device 34 can use a lookup table or a compensation formula to determine parameter adjustment based on the difference score.
[0154] In some examples, audio decoding device 34 can determine parameter adjustments by determining the difference between the energy map of an audio element and a composite energy map. In such examples, the difference may indicate the frequency-dependent equalizer (EQ) gain that has been or will be applied to the audio signal of a target audio element (e.g., an audio element targeted for parameter adjustment). In the examples, the difference is the difference between the expected energy at the location of the audio element given by the composite energy map and the measured energy of the audio element signal. In illustrative examples, audio decoding device 34 may determine the expected energy at a specific location of the audio element based at least in part on the composite energy map, and then may perform energy analysis on the signal generated by the audio element to determine the difference (e.g., energy difference). Audio decoding device 34 may then send parameter adjustments to the audio element, which may include the difference as part of the parameter adjustments. When implementing parameter adjustments, the audio element may apply the difference (in decibels (dB)) directly as a gain factor to the audio signal generated by the audio element.
[0155] In some instances, the audio decoding device 34 may determine the operating state (408) of one or more audio elements before determining parameter adjustments. The operating state may include the signal-to-noise ratio (SNR) of the audio elements. In another example, the audio decoding device 34 may utilize self-diagnostic data received from one or more microphones 18 to determine the array of microphones 18 that can be used to establish baseline readings. In one example, the audio decoding device 34 may utilize this data before or after accessing one or more energy maps. In the former case, the audio decoding device 34 may selectively access only those energy maps that satisfy the criteria defined by the operating state information and as previously discussed. In the latter case, the audio decoding device 34 may modify the energy map or audio stream based on the operating state (410). For example, the audio decoding device 34 may remove or exclude certain audio streams or energy maps from consideration when forming a composite energy map. In any case, the audio decoding device 34 may remove noise signals or remove those devices that generate noisy signals from consideration.
[0156] The audio decoding device 34 can update various inputs at any given frequency. For example, the audio decoding device 34 can update all or some energy maps at the audio frame rate (meaning the energy maps are updated once per frame). In some instances, the audio decoding device 34 can periodically update energy maps and composite energy maps. In some examples, the audio decoding device 34 can update energy maps in response to triggers (such as detecting a new audio element or detecting that a previously unavailable, unresponsive, noisy, or otherwise damaged audio element has now become available for consideration). For example, a user can update privacy settings, which allows the audio decoding device 34 to use the new audio stream and corresponding energy map to determine parameter adjustments, or requests the audio decoding device 34 to now exclude an audio stream or energy map from consideration. In another example, the audio decoding device 34 can update permission / privacy settings at the UI rate (meaning the update is driven by an update via UI input). As yet another example, the audio decoding device 34 can update position at the sensor rate (meaning the position changes with the movement of audio elements).
[0157] In some examples, the audio decoding device 34 can output parameter adjustments (406) corresponding to the microphone. For example, the audio decoding device 34 can send parameter adjustments to the microphone corresponding to the audio stream capture that needs adjustment. In another example, the audio decoding device 34 can directly apply parameter adjustments to the microphone corresponding to the audio decoding device 34. In some examples, the audio decoding device 34 can output the parameter adjustments to a storage location and store the parameter adjustments for later access. In some instances, the audio decoding device 34 can decode and output audio data 19' corresponding to one of the bitstreams 27 based on parameter settings adjusted or maintained due to strong energy map readings (e.g., near energy map comparison).
[0158] In another example, audio decoding device 34 can adjust the frequency-dependent gain of each audio element (e.g., each receiver). In this example, audio decoding device 34 can determine the adjustment of the frequency-dependent gain based on a comparison of a composite energy map with an individual energy map obtained for each audio element.
[0159] like Figure 4B As shown, the audio decoding device 34 can repeat this process in a loop configuration. For example, the audio decoding device 34 can then access and / or determine the energy map (420) of each audio element 302 (e.g., audio capture receiver, audio synthesizer). In some examples, the audio decoding device 34 can then determine the operating state (422) of the audio element 302. In the example, the audio decoding device 34 can receive operating state information (e.g., self-diagnostics) and check if any audio capture receiver is malfunctioning (e.g., noisy or silent).
[0160] In some examples, audio decoding device 34 can determine whether to remove any audio element from consideration as a valid audio element (424). In the examples, audio decoding device 34 can remove those unqualified audio elements from consideration. That is, audio decoding device 34 can use or consider only the energy map of valid audio elements (e.g., receiver, synthesizer). Thus, audio decoding device 34 can use the energy map obtained via valid audio elements to determine a composite energy map (426). Valid audio elements may include audio elements that are undamaged, noisy, silent, or otherwise unusable for forming a composite energy map for baseline comparison.
[0161] In some examples, audio decoding device 34 can determine a composite energy map. While described with reference to audio decoding device 34, the techniques of this disclosure are not limited thereto, and it should be understood that other devices or processing systems of content consuming device 14, source device 12, or remote device (e.g., remote server 504) can perform one or more of the various techniques of this disclosure. In an illustrative example involving source device 12 performing one or more of the various techniques of this disclosure, controller 31 can receive multiple energy maps via side channel 33 from content consuming device 14A or from other source device 12.
[0162] In another example, a specific controller 31 for source device 12 may receive audio streams (e.g., bitstream 27) from other source devices 12 or from content consuming device 14. The specific controller 31 may then determine a composite energy map based on multiple energy maps corresponding to audio element 302 or from multiple audio streams corresponding to it. In another example, controller 31 may pass multiple energy maps and / or multiple audio streams (e.g., as metadata 25 sent from controller 31 or from sound field representation generator 24 to content editing device 22) to content editing device 22. Additionally, content editing device 22 or content capture device 20 may receive multiple energy maps and / or multiple audio streams from controller 31 or from sound field representation generator 24, and may then determine a composite energy map based on these multiple energy maps or based on multiple audio streams determined for each of the multiple source devices 12. As illustrated, controller 31 may be integrated with one or more of sound field representation generator 24 and / or content editing device 22.
[0163] However, to avoid confusion and as described herein, various techniques of this disclosure are described with reference to audio decoding device 34, wherein audio decoding device 34 can determine parameter adjustments for a particular source device 12 (e.g., a valid source device 12) by analyzing one or more corresponding energy maps based on a composite energy map, and pass the parameter adjustments to various source devices 12, wherein source devices 12 can receive parameter adjustments (e.g., as PSI 46A) and implement the parameter adjustments to compensate for differences (e.g., discrepancies) between the generation of audio data 19 via source device 12 or to provide compensated edited content 23. In some instances, parameter adjustments may include instructions for disabling a particular source device 12, wherein source device 12 is generating a corrupted, noisy, or otherwise confusing bitstream 27 and / or sending it to content consumption device 14 or a remote server (e.g., remote server 504).
[0164] In some examples, to generate a composite energy map, audio decoding device 34 may calculate a roll-off (e.g., frequency) based on multiple energy maps and combine the energy maps to form a composite energy map based on the roll-off frequency. In examples, audio decoding device 34 may interpolate between the roll-off values of multiple energy maps to determine a single composite energy map, such as an energy map composed of at least two energy maps from multiple energy maps (e.g., two energy maps with the highest SNR, a comparison energy map, and one or more reference energy maps from a pre-compensated reference audio element, etc.). The composite energy map may be different from and separate from the energy map used to form the composite energy map. In another example, the composite energy map may include individual energy maps specific to a particular source device 12. In any case, the composite energy map is formed as a baseline for determining when the energy map of another device does not match or is not compensated for by other devices in the constellation set of audio elements 302.
[0165] In an illustrative example, audio decoding device 34 can calculate theoretical and / or location-dependent roll-offs from multiple energy maps to form a composite energy map. The roll-off calculation can be based on a linear or logarithmic scale (e.g., decibels, etc.), depending on tuning preference information (e.g., based on PSI 46 including tuning preference information). In some examples, audio decoding device 34 can determine the composite energy map from the roll-off information in a variety of different ways, including determining a reference energy map, interpolating between energy maps, such as interpolating between frequency data of the corresponding energy maps, etc.
[0166] In some examples, audio decoding device 34 may determine a first audio element (e.g., audio element 302A) that has been signaled as having been pre-compensated (e.g., pre-calibrated, pre-equalized). In illustrative and non-limiting examples, the first audio element may include microphone 18 and / or content capture device 20 (e.g., Figure 1A (Or 1B). In any case, the audio decoding device 34 can use the energy map corresponding to the first audio element as a reference energy map. The audio decoding device 34 can then calculate the composite energy map by calculating the roll-off of other audio elements (e.g., receivers, synthesized sound fields, etc.) relative to the reference energy map. In an example where no audio element signals for pre-compensation, the audio decoding device 34 can calculate the centroid positions of multiple audio elements in the audio element set. The audio decoding device 34 can then determine one or more audio elements closest to the centroid position as reference elements to provide a reference energy map for performing energy map comparisons (e.g., roll-off frequency comparisons).
[0167] In another example, audio decoding device 34 (or a remote server in some examples) may receive multiple audio streams (e.g., multiple bitstreams 27) from multiple source devices 12. Alternatively, audio decoding device 34 may receive multiple energy maps and / or SNR information regarding one or more energy maps. Audio decoding device 34 may determine multiple energy maps and / or determine SNR information for multiple energy maps (e.g., energy maps of different audio elements 12) based on the multiple audio streams. Based on the SNR information associated with the multiple energy maps and / or the SNR information of one or more energy maps, audio decoding device 34 may identify one or more audio elements 302 from a set of audio elements (e.g., a constellation) that have the highest SNR energy map relative to the energy maps of other audio elements 302 in the constellation.
[0168] In one example, the audio decoding device 34 can identify the top "N" audio elements 302 with a higher SNR energy map relative to other audio elements 302 in the constellation. In another example, the audio decoding device 34 can identify the top "N" audio elements 302 in the constellation that exceed an SNR threshold. In some examples, the audio decoding device 34 can receive SNR information (e.g., as audio metadata) indicating the operational state of one or more audio elements 302, and then determine a composite energy map from the SNR information based on the audio element 302 with the highest SNR energy map relative to any other audio element 302 in the set of audio elements 302.
[0169] In the illustrative example, audio decoding device 34 may be a decoding device of device 204 (e.g., headphones, audio speakers, XR devices, etc.) that receives audio streams and / or energy maps from audio elements 302A to 302K, as exemplified by reference. Figures 3A to 3CThe decoding devices described herein. In another example, the techniques of this disclosure can be performed by an audio element 302 that implements the functionality of source device 12, such as in cases where audio element 302 includes content capture device 20 (e.g., microphone 18) and / or synthesis device 29. Audio element 302 can determine a composite energy map and determine one or more parameter adjustments that audio element 302 can send to other audio elements 302 in its constellation.
[0170] In another illustrative example, refer to Figure 5B To clarify, any one or more of devices 504A, 204A, 204B, 504B, etc., can determine a composite energy map based on the first "N" audio elements 302 or based on a reference energy map in order to subsequently determine the parameter adjustments for audio element 302A. Figure 5B In the illustrative example, this parameter adjustment includes disabling audio element 302A from generating audio for the XR experience for user 202 of device 204A or 204B, or excluding the energy map of audio element 302A from the generation of a composite energy map. This may be because audio element 302A has, for example, poor quality characteristics, operational or feasibility status (e.g., in a pocket, under password protection), energy map differences, etc., causing audio element 302A to be disabled and / or signaled as unqualified until audio element 302A is reconfigured to subsequently improve quality, operational status, energy map differences, or other factors that cause audio decoding device 34 to provide parameter adjustments for audio element 302A that was initially disabled or marked as unqualified.
[0171] In any case, once generated, audio decoding device 34, or another device in another example, can access the composite energy map that, once generated, determines the parameter adjustments of one or more of the audio elements as described herein. As additional audio elements enter the constellation set of audio element 302 and / or as a particular audio element 302 becomes eligible or becomes ineligible, audio decoding device 34 can update the composite energy map over time (e.g., audio element 302A can be...). Figure 5B In the example, upon re-entering the constellation, the audio decoding device 34 of devices 504 and 204, or another audio element, is prompted to update or modify one or more composite energy maps. (For example, the audio decoding device 34 of device 204A) can then use the composite energy map to determine differences in the energy maps of the audio elements in the constellation set and / or can send the composite energy map to other devices (e.g., server 504, device 204B, other audio elements 302, etc.) for further energy map and / or parameter adjustment processing.
[0172] When determining the composite energy map, the audio decoding device 34 is further configured to utilize the top "N" audio elements 302 (e.g., the audio element 302 with the highest SNR energy map) and may utilize the corresponding energy maps corresponding to audio elements 302 in the set of audio elements 302, whose energy map satisfies an SNR threshold or whose SNR value is originally higher than the SNR of the energy maps corresponding to other audio elements 302 in the set of audio elements 302. The audio decoding device 34 can then interpolate between the N energy maps to form the composite energy map. In an example, the audio decoding device 34 may access multiple energy maps corresponding to multiple audio streams, where the energy maps have SNR values that satisfy the SNR threshold (e.g., have higher quality characteristics relative to other energy maps), and then the audio decoding device 34 may interpolate between the multiple energy maps to form the composite energy map. In some examples, the audio decoding device 34 may determine an average value among the multiple energy maps and form the composite energy map based on that average value and / or based on the interpolation.
[0173] In the illustrative example, "N" can be based on the total number of audio elements 302 detected within a threshold distance from each other in a geographic area. In any example, the total number of audio elements 302 may include fifteen microphones and two synthesized sound fields residing on a stage (e.g., stage 523) in a concert hall. The audio decoding device 34 may be configured to select a specific number less than the total as "N" (e.g., the first five) or to select a score of the total as "N" (e.g., the first half or the first third of a total of seventeen audio elements 302 in this illustrative example).
[0174] In some examples, the audio decoding device 34 can compare the energy map of each valid audio element with a composite energy map (428). In the example, the audio decoding device 34 can then analyze the energy maps to examine frequency-related differences between them.
[0175] In some instances, the audio decoding device 34 may not perform an operational status check on the audio element 302. In such examples, the audio decoding device 34 may compare the energy maps of the audio elements to determine if any frequency-dependent differences exist between the energy maps and the composite energy map. The audio decoding device 34 may adjust the frequency-dependent gain of each receiver accordingly. However, the audio decoding device 34 may determine, based on the analysis of the frequency-dependent differences, that the audio element does not require parameter adjustment. In such examples, the audio decoding device 34 may return to analyze the energy maps of all valid audio elements.
[0176] In some examples, the audio decoding device 34 may receive or determine operational status information before accessing the energy map. In such cases, the audio decoding device 34 may remove some energy maps or some audio streams from consideration based on the operational status information. In other instances, the audio decoding device 34 may analyze the energy map and operational status information in parallel to determine which energy maps should be analyzed and which should not. The audio decoding device 34 may then automatically equalize or normalize the real-time XR capture device for 6DOF listening and / or rendering.
[0177] In another example, audio decoding device 34 can access a composite energy map of multiple audio elements generated from an energy map set, each energy map corresponding to one of the multiple audio elements; determine the difference between the configuration signature of at least one of the multiple audio elements and the composite energy map; and generate instructions for adjusting parameter settings based at least in part on the difference.
[0178] In some examples, audio decoding device 34 can compile an energy atlas comprising energy maps of at least two audio elements from a plurality of audio elements. Audio decoding device 34 can then generate a composite energy map of the plurality of audio elements, generated from the energy atlas, with each energy map corresponding to one audio element from the plurality of audio elements. Content consuming device 14 can then send the composite energy map to source device 12.
[0179] In another example, controller 31 can generate an energy map and send it to content consuming device 14. Controller 31 can then receive instructions from content consuming device 14A to adjust parameter settings. Content consuming device 14 can determine the instructions based on a comparison of the energy map and a composite energy map. In response to receiving the instructions, source device 12 can adjust the parameter settings.
[0180] While the discussion focuses on audio decoding device 34, any number of different audio devices (including one or more processors of source device 12, one or more processors of content consumption device, or any other audio-related device) can perform the various techniques disclosed herein. The audio device should be configured at least to access and / or analyze one or more energy maps.
[0181] Figures 5A to 5D To show in more detail Figures 1A to 1C The example diagram illustrates the operation of the audio decoding device 34. Figure 5AAs shown in the example, audio decoding device 34 can determine the presence of multiple audio elements 302. Audio decoding device 34 may correspond to one of VR device 204A or VR device 204B (“VR device 204”), audio source 308, one of audio elements 302, or one of remote server 504. Audio decoding device 34 can determine one or more audio elements 302 (which may represent not only microphones, but also...) Figure 1A and 1B The microphone 18 shown also represents parameter settings for other types of audio receiving devices, including other XR devices, mobile phones (including so-called smartphones, etc., or the generated sound field).
[0182] As described above, the audio decoding device 34 can decode audio elements 302 (which may represent not only microphones, but also other audio components such as microphones, ... Figure 1A The microphone 18 shown, and also other types of capture devices, including other XR devices, mobile phones (including so-called smartphones, etc.), or the generated sound field, obtain bitstream 27. Audio decoding device 34 can interface with audio element 302 to obtain bitstream 27. In some examples, stream selection unit 44 can be configured according to 5G cellular standards, personal area networks (PANs) (such as... Interacting with other open-source, proprietary, or standardized communication protocols and interfaces (such as receivers, transmitters, and / or transceivers) to obtain a bitstream27. Wireless transmission of the audio stream and / or transmission of other audio data (such as energy maps or parameter adjustments) in... Figures 5A to 5C In the example, it is represented as lightning, where selected audio data 19' is shown as being transmitted from one or more audio elements 302 to and from VR device 204, to and from remote server 504, and to and from audio source 308.
[0183] In some instances, audio source 308 may include a user, a streaming media source (e.g., a smart TV), or other audio generation sources, such as an environment with sound. In some instances, audio element 302G may be located in an environment remote from the user, where the user can experience that environment in an XR space. For example, a user might remotely watch a movie in a remote cinema, or multiple people might remotely experience an XR space together, where one or more microphones can be placed in the cinema to capture one or more audio streams.
[0184] The audio decoding device 34 can access at least one energy map in the manner described above by retrieving an energy map from a memory or generating at least one energy map. For example, the audio decoding device 34 can perform energy analysis on one of the bitstreams 27 to determine, for example, at least one energy map corresponding to the corresponding audio stream or, in some instances, multiple audio streams when a composite energy map is determined.
[0185] Next reference Figure 5B For example, audio decoding device 34 can receive operational status information from audio element 302A. Using the operational status information, audio decoding device 34 can determine that the audio stream captured by audio element 302A is corrupted, noisy, or unusable. In some examples, audio decoding device 34 can compare the SNR of the audio element with a threshold SNR to determine whether the audio element is corrupted, noisy, or unusable. Similarly, the audio element can be subjected to a minimum gain check for silence (e.g., the device is on but in the user's pocket or wallet). Audio decoding device 34 can remove the audio stream and / or the corresponding energy map from CM 47.
[0186] In some instances, the audio decoding device 34 may remove the audio stream and / or energy map before generating the composite energy map. In other instances, the audio decoding device 34 may regenerate the composite energy map if the audio stream and / or energy map are not considered unavailable. In any case, the audio decoding device 34 may update the energy map periodically relative to the audio frame rate, update the energy map individually, or update it as a composite energy map. The composite energy map may be the average of multiple energy maps, such that the composite energy map is most likely to provide the most accurate depiction of the common sound field.
[0187] exist Figure 5C In the examples, audio element 302 is replaced by specific devices 320A to 320E (“Device 320”), where Device 320A represents a dedicated microphone 320A, and Devices 320B and 320C represent mobile devices 320 (e.g., smartphones or mobile handheld devices). Devices 320E and 320F may represent VR devices 320. Each of Devices 320 may include audio element 302, which captures a bitstream 27 modulated according to various aspects of the parameter adjustment techniques described in this disclosure. In some examples, audio element 302 may be enabled to receive audio. In another example, audio element 302 may be generated by synthesis device 29. Furthermore, Device 320 may include wearable devices, mobile handheld terminals, XR devices, and audio receivers.
[0188] Device 320 may be coupled to one or more speakers. Speakers can be configured to generate sound fields, such as by reproducing, recreating, generating, playing, storing, or otherwise representing the sound field. According to various aspects of this disclosure, device 320 may be configured to provide a 3DOF, 3DOF+, or 6DOF user experience. In some instances, device 320 may include a receiver configured to receive audio streams according to 5G cellular standards and / or personal area network standards. Additionally, device 320 may be configured to receive data via a wireless link, such as via a 5G air interface or Bluetooth interface. In other examples, device 320 may be configured to receive data via a wired link. In some instances, one of devices 320 may include a remote server configured to perform tuning techniques. Additionally, device 320 may include source device 12 or content consuming device 14. For example, device 320 may generate audio via one or more speakers, and therefore may include source device 12.
[0189] In some examples, mobile devices such as smartphones can receive audio streams from multiple source devices 12 and determine parameter adjustments for the source devices 12 and the content consuming device 14. The smartphone can generate an energy map and render an overlay of energy curves to provide a visual representation of the energy map and sound field. The smartphone can access the energy map from another audio device, an external server, or by generating the energy map using bitstream 27.
[0190] In some examples, the audio decoding device 34 may use peer-to-peer communication to share parameter adjustments (e.g., calibration adjustments) from one or more devices that have already been tuned and to share those parameter adjustments with new devices introduced into the area. In some examples, the audio decoding device 34 may utilize an electronic communication network (ECNS). The audio decoding device 34 may be configured to prevent feedback from noise or other microphones to the loudspeakers.
[0191] In some examples, room-dependent parameter adjustments may be present. For example, audio decoding device 34 may be configured to reduce room-mode resonance. In some examples, a particular room resonance may occur at one or more specific nodes. In some examples, a particular room resonance may occur at one or more nodes where the sound wavelength is a multiple of the room size. When an audio element is located at a destructive interference node, audio decoding device 34 may then apply parameter adjustments to boost the frequencies that have been affected by one or more specific nodes. In such examples, the affected audio element may implement (e.g., apply) bandpass equalization gain to boost the frequencies affected by the node.
[0192] In another example, when the audio element is located at a node where constructive interference exists, the audio decoding device 34 can determine a parameter adjustment to reduce the gain at room-dependent frequencies. In this way, the audio element can implement parameter adjustments (e.g., to reduce gain, boost one or more affected frequencies, etc.) to increase the quality characteristics of the signal generated by the audio element. In the example, reducing the gain of the audio element at the room-dependent frequencies at a node where constructive interference exists can allow the audio element to generate a sound field that doesn't sound "boomy," or at least doesn't sound as "boomy" as before the parameter adjustment.
[0193] Figure 5D This is a conceptual diagram illustrating an example concert with three or more audio elements. Figure 5D In the example, several musicians are depicted on stage 523. Singer 512 is located behind audio element 510A. String section 514 is depicted behind audio element 510B. Drummer 516 is depicted behind audio element 510C. Other musicians 518 are depicted behind audio element 510D. Audio elements 510A to 510D can capture audio streams corresponding to the sounds received by the microphone. In some examples, audio elements 510A to 510D can represent generated audio streams (e.g., synthesized audio streams).
[0194] Audio element 510A can represent one or more captured audio streams primarily associated with singer 512. Alternatively, example audio streams may also include sounds generated by other band members, such as string section 514, drummer 516, or other musicians 518. Alternatively, audio element 510B can represent one or more audio streams primarily associated with string section 514, and may also represent sounds generated by other band members. In this way, each of audio elements 510A through 510D can represent a different audio stream.
[0195] Numerous devices are also depicted. These devices represent user devices located at multiple different target listening positions. Headphones 521 are located near audio element 510A, but between audio element 510A and audio element 510B. Therefore, according to the technology of this disclosure, the stream selection unit 44 can select at least one audio stream to generate an audio experience for the user of headphones 521, similar to the user being located at... Figure 5D The headset 521 is positioned as shown in the image. Similarly, the VR goggles 522 are shown positioned behind the audio element 510C and between the drummer 516 and other musicians 518. The stream selection unit 44 can select at least one audio stream to generate an audio experience for the user of the VR goggles 522, similar to the user being positioned in... Figure 5D The location of VR goggles 522 in the image.
[0196] The smart glasses 524 are shown centered between audio elements 510A, 510C, and 510D. The stream selection unit 44 can select at least one audio stream to generate an audio experience for the user of the smart glasses 524, similar to a user positioned in a different location. Figure 5D The location of the smart glasses 524 is shown. Additionally, device 526 (which can represent any device capable of implementing the technologies of this disclosure, such as a mobile handheld terminal, speaker array, headset, VR goggles, smart glasses, etc.) is shown positioned in front of audio element 510B. Stream selection unit 44 can select at least one audio stream to generate an audio experience for a user of device 526 (e.g., user 202), similar to the user being positioned in front of audio element 510B. Figure 5D The location of device 526 in the diagram. Although a specific device is discussed for a specific location, any device depicted can provide [something related to] ... Figure 5D The text describes different desired listening locations.
[0197] In such examples, content consuming device 14 and / or source device 12 may coordinate with each other to form a composite energy map of each audio element 510 and audio element 521 (e.g., headphone microphone) to determine whether to disable any audio element 510 or 521. In some examples, such as in the case of a noisy audio element, content consuming device 14 and / or source device 12 may further determine whether to remove any audio element 510 or 521 before generating the composite energy map. In an example, audio element 510A may include a microphone of a listener member, wherein the microphone is in the listener member's pocket and may therefore have a poor SNR reading or other quality metrics below a predetermined SNR threshold. Thus, content consuming device 14 and / or source device 12 may receive an energy map of audio element 510A, but may exclude that energy map when generating a composite energy map for the remaining and active audio elements relative to stage 523. In some examples, musician 518 may not be physically present on stage 523, but audio element 510D may include an audio stream generated at the location of the musician shown in Figure 5. However, the content consumption device 14 and / or the source device 12 can determine a composite energy map that includes an energy map corresponding to the synthesized audio element 510D.
[0198] Figure 6This is a diagram illustrating an example of a wearable device 602 that can operate according to various aspects of the techniques described in this disclosure. In various examples, wearable device 602 may represent an XR device (e.g., a VR device 204 described herein, an AR headset, an MR headset, or any other type of XR headset). Augmented Reality “AR” may refer to computer-rendered images or data superimposed on the real world in which the user actually resides. Mixed Reality “MR” may refer to computer-rendered images or data locked to a specific location in the real world, or may refer to a variant of VR in which a combination of computer-rendered 3D elements and filmed real elements creates an immersive experience that simulates the user’s physical presence in the environment. Extended Reality “XR” may refer to VR, AR, and MR collectively. More information on the XR terminology can be found in the document entitled “Virtual Reality, Augmented Reality, and Mixed Reality Definitions” published by Jason Peterson on July 7, 2017.
[0199] Wearable device 602 can refer to other types of devices, such as watches (including so-called "smartwatches"), glasses (including so-called "smart glasses"), headphones (including so-called "wireless headphones" and "smart headphones"), smart clothing, smart jewelry, etc. Regardless of whether it refers to a VR device, watch, glasses, and / or headphones, wearable device 602 can communicate with a computing device that enables wearable device 602 via a wired or wireless connection.
[0200] In some instances, the computing device supporting the wearable device 602 may be integrated within the wearable device 602; therefore, the wearable device 602 can be considered the same device as the computing device supporting the wearable device 602. In other instances, the wearable device 602 may communicate with a separate computing device that can support the wearable device 602. In this regard, the term "support" should not be construed as requiring a separate dedicated device, but rather as meaning that one or more processors configured to perform various aspects of the technologies described in this disclosure may be integrated within the wearable device 602 or integrated within a computing device separate from the wearable device 602.
[0201] For example, when wearable device 602 refers to a VR device, a separate dedicated computing device (such as a personal computer including one or more processors) can render audio and visual content, while wearable device 602 can determine translational head movements, and the dedicated computing device can render audio content (as a speaker feed) based on these translational head movements, according to various aspects of the technology described in this disclosure. As another example, when wearable device 602 refers to smart glasses, wearable device 602 may include one or more processors that determine translational head movements (via docking within one or more sensors of wearable device 602) and render speaker feeds based on the determined translational head movements.
[0202] As shown in the figure, wearable device 602 includes one or more directional speakers and one or more tracking and / or recording cameras. Additionally, wearable device 602 includes one or more inertial, haptic, and / or health sensors, one or more eye-tracking cameras, one or more high-sensitivity audio elements (e.g., one or more microphones), and optical / projection hardware. The optical / projection hardware of wearable device 602 may include durable translucent display technology and hardware.
[0203] Wearable device 602 also includes connectivity hardware, which may represent one or more network interfaces supporting multi-mode connectivity, such as 4G communication, 5G communication, Wi-Fi TM Wearable device 602 also includes one or more ambient light sensors, one or more cameras and night vision sensors, and one or more bone conduction sensors. In some instances, wearable device 602 may also include one or more passive and / or active cameras with fisheye lenses and / or telephoto lenses. Although not described in Figure 6 As shown, wearable device 602 may also include one or more light-emitting diode (LED) lights. In some examples, the LED lights may be referred to as “ultra-bright” LED lights. In some embodiments, wearable device 602 may also include one or more rear cameras. It should be understood that wearable device 602 may exhibit a variety of different shape factors.
[0204] Furthermore, tracking and recording cameras, along with other sensors, can facilitate the determination of translation distance. Although not in Figure 6 The example shown is not applicable, but wearable device 602 may include other types of sensors for detecting translational distance.
[0205] Although specific examples of wearable devices (such as those discussed in this article) Figure 2 The example discussed in this article is VR device 204 and in this paper. Figures 1A to 1C Other devices described in the examples are illustrated, but those skilled in the art will understand that, with Figures 1A to 1C The description related to section 2 can be applied to other examples of wearable devices. For example, other wearable devices (such as smart glasses) may include sensors that detect translational head movement. As another example, other wearable devices (such as smartwatches) may include sensors that detect translational movement. Therefore, the techniques described in this disclosure should not be limited to a particular type of wearable device, but any wearable device can be configured to perform the techniques described in this disclosure.
[0206] Figure 7A and 7B This is a diagram illustrating an example system that can perform various aspects of the techniques described in this disclosure. Figure 7A An example is shown in which the source device 12C also includes a camera 702. The camera 702 can be configured to capture video data and provide the captured raw video data to the content capture device 20. The content capture device 20 can provide the video data to another component of the source device 12C for further processing into viewport-divided portions.
[0207] exist Figure 7A In the example, content consuming device 14C also includes VR device 204. It will be understood that in various implementations, VR device 204 may be included in or externally coupled to content consuming device 14C. VR device 204 includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewports) and for rendering audio data.
[0208] Figure 7B It shows that Figure 7A The audio renderer 32 shown is replaced with an example of a binaural renderer 42 capable of performing binaural rendering using one or more HRTFs or other functions capable of rendering to the left and right speaker feeds 43. The audio playback system 16C of the content consumption device 14D can output the left and right speaker feeds 43 to headphones 48.
[0209] The headphones 48 can be connected via a wired connection (such as a standard 3.5mm audio jack, a universal system bus (USB) connection, an optical audio jack, or other forms of wired connection) or wirelessly (such as via... The headset 48 can generate a sound field represented by audio data 19' based on the left and right speaker feeds 43. The headset 48 may include a left and right headphone speaker, which are powered (or driven) by the corresponding left and right speaker feeds 43. It should be noted that the content consumption device 14C and / or content consumption device 14D can be coupled to the audio playback system 16C. Figures 1A to 1C Use it together with source device 12.
[0210] Figure 8 It is shown Figures 1A to 1C The example shown is a block diagram of one or more sample components from the source device and content consuming device. Figure 8 In some examples, device 710 includes a processor 712 (which may be referred to as "one or more processors" or "processor"), a graphics processing unit (GPU) 714, system memory 716, a display processor 718, one or more integrated speakers 740, a display 703, a UI 720, an antenna 721, and a transceiver module 722. In examples where device 710 is a mobile device, the display processor 718 is a mobile display processor (MDP). In some examples, such as those where device 710 is a mobile device, processor 712, GPU 714, and display processor 718 may be configured as integrated circuits (ICs).
[0211] For example, an IC can be considered a processing chip within a chip package and can be a system-on-a-chip (SoC). In some examples, two of the processor 712, GPU 714, and display processor 718 may be housed together in the same IC, while the third may be housed in a different integrated circuit (e.g., in a different chip package), or all three may be housed in different ICs or on the same IC. However, in the example where device 710 is a mobile device, the processor 712, GPU 714, and display processor 718 may all be housed in different integrated circuits.
[0212] Examples of processor 712, GPU 714, and display processor 718 include, but are not limited to, one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Processor 712 may be the central processing unit (CPU) of device 710. In some examples, GPU 714 may be dedicated hardware that includes integrated and / or discrete logic circuitry to provide GPU 714 with massively parallel processing capabilities suitable for graphics processing. In some instances, GPU 714 may also include general-purpose processing capabilities and may be referred to as a general-purpose GPU (GPGPU) when performing general-purpose processing tasks (e.g., non-graphics-related tasks). Display processor 718 may also be application-specific integrated circuit hardware designed to retrieve image content from system memory 716, combine the image content into image frames, and output the image frames to display 703.
[0213] Processor 712 can execute various types of applications. Examples of applications include web browsers, email applications, spreadsheets, video games, other applications that generate visual objects for display, or any application types listed in more detail herein. System memory 716 can store instructions for executing applications. Execution of one of the applications on processor 712 causes processor 712 to generate graphical data of image content to be displayed and audio data 19 to be played (possibly via integrated speaker 740). Processor 712 can send the graphical data of the image content to GPU 714 for further processing based on instructions or commands sent by processor 712 to GPU 714.
[0214] Processor 712 can communicate with GPU 714 according to a specific application processing interface (API). Examples of such APIs include... of API, Khronos Group or and OpenCL TM However, aspects of this disclosure are not limited to DirectX, OpenGL, or OpenCL APIs and can be extended to other types of APIs. Furthermore, the techniques described in this disclosure do not require an API to function, and the processor 712 and GPU 714 can communicate using any process.
[0215] System memory 716 may be the memory of device 710. System memory 716 may include one or more computer-readable storage media. Examples of system memory 716 include, but are not limited to, random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other media that can be used to carry or store desired data in the form of instructions and / or data structures and that can be accessed by a computer or processor.
[0216] In some examples, system memory 716 may include instructions that cause processor 712, GPU 714, and / or display processor 718 to perform the functions attributed to processor 712, GPU 714, and / or display processor 718 in this disclosure. Therefore, system memory 716 may be a computer-readable storage medium storing instructions thereon that, when executed, cause one or more processors (e.g., processor 712, GPU 714, and / or display processor 718) to perform various functions.
[0217] System memory 716 may include a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that system memory 716 is immovable or that its contents are static. As an example, system memory 716 may be removed from device 710 and moved to another device. As another example, a memory substantially similar to system memory 716 may be inserted into device 710. In some examples, the non-transitory storage medium may store data that changes over time (e.g., stored in RAM).
[0218] UI 720 can represent one or more hardware or virtual (meaning a combination of hardware and software) UIs through which a user can interact with device 710. UI 720 may include physical buttons, switches, toggle switches, lights, or virtual versions thereof. UI 720 may also include physical or virtual keyboards, touch interfaces—such as touchscreens, haptic feedback, etc.
[0219] Processor 712 may include one or more hardware units (including so-called "processing cores") configured to perform all or part of the operations discussed above regarding any one or more of the modules, units, or other functional components of source device 12 (e.g., a content creator device) and / or content consumer device 14. For example, processor 712 may implement the operations discussed above regarding... Figures 3A to 3D , Figures 4A to 4B , Figures 5A to 5D , Figure 6 , Figures 7A to 7B and Figures 9A to 9C as well as Figure 10 The functionality described in the section regarding parameter adjustment and / or energy diagrams. Antenna 721 and transceiver module 722 can represent units configured to establish and maintain a connection between source device 12 and content consumption device 14. Antenna 721 and transceiver module 722 can represent the ability to operate according to one or more wireless communication protocols (such as 5G cellular standards, PAN protocols, etc.). One or more receivers and / or one or more transmitters (or other open-source, proprietary, or other communication standards) for wireless communication. Therefore, transceiver module 722 can be configured to receive and / or transmit wireless signals. In some examples, transceiver module 722 can represent a separate transmitter, a separate receiver, a separate transmitter and a separate receiver, or a combination of both. Antenna 721 and transceiver module 722 can be configured to receive encoded audio data. Similarly, antenna 721 and transceiver module 722 can be configured to transmit encoded audio data.
[0220] Figures 9A to 9C It is shown Figures 1A to 1CThe example shown is a flowchart illustrating the operation of the stream selection unit 44 in performing various aspects of stream selection and audio element compensation techniques. First, refer to... Figure 9A For example, stream selection unit 44 may obtain bitstream 27 from all enabled audio elements (e.g., receivers, such as microphone 18, audio synthesizers, such as synthesis device 29, etc.), wherein bitstream 27 may include corresponding information (e.g., metadata), such as ALI 45A (800). Stream selection unit 44 may perform energy analysis on each of the bitstreams 27 to calculate the corresponding energy map (802). In an illustrative example, stream selection unit 44 may determine a composite energy map based on a combination of at least two energy maps (e.g., energy maps determined for multiple audio elements in a constellation of audio elements). In an illustrative example, audio decoding device 34 may interpolate between multiple energy maps to form a composite energy map. In another example, audio decoding device 34 may compare the roll-off of an energy map with the roll-off of another reference energy map of a pre-compensated audio element to form a composite energy map.
[0221] The stream selection unit 44 can iterate through different combinations of audio elements (defined in CM 47) (804) based on proximity to audio source 308 (as defined by audio source distances 306A and / or 306B) and proximity to audio elements (as defined by proximity distances discussed herein). Figure 9A As shown, audio elements can be sorted or otherwise associated with different access permissions. Stream selection unit 44 can iterate in the manner described above based on the listener position represented by DLI 45B (which is another way of referring to "virtual position" or "device position") and the audio element position represented by ALI 45A to identify whether a larger subset or a smaller subset of bitstream 27 is needed (806, 808).
[0222] When a larger subset of bitstream 27 is needed, stream selection unit 44 can add audio elements to audio data 19', or in other words, add additional audio streams (such as when the user is closer). Figure 3A (810) When a reduced subset of bitstream 27 is needed, stream selection unit 44 can remove audio elements from audio data 19', or in other words, remove one or more existing audio streams (such as when the user is further away). Figure 3A (812) in the example of the audio source.
[0223] In some examples, the stream selection unit 44 may determine that the current constellation of the audio element is the optimal set (or, in other words, the existing audio data 19' will remain the same as the selection process described herein, resulting in the same audio data 19') (804), and the process may return to 802. However, when an audio stream is added to or removed from the audio data 19', the stream selection unit 44 may update CM 47 (814) to generate a constellation history (815) (including positions, energy maps, etc.).
[0224] Additionally, the stream selection unit 44 can determine whether privacy settings enable or disable the addition of audio elements (where privacy settings may refer to digital access permissions that restrict access to one or more of the bitstreams 27, such as through a password, authorization level or grade, or time) (816, 818). When privacy settings enable the addition of audio elements, the stream selection unit 44 can add the audio elements to the updated CM 47 (which means adding an audio stream to audio data 19') (820). When privacy settings disable the addition of audio elements, the stream selection unit 44 can remove audio elements from the updated CM 47 (which means removing one or more audio streams from audio data 19') (822). In this way, the stream selection unit 44 can identify the enabling of a new set of audio elements (824).
[0225] The stream selection unit 44 can iterate and update various inputs at any given frequency in this way. For example, the stream selection unit 44 can update privacy settings at the UI rate (meaning updates are driven by updates via UI input). The stream selection unit 44 can update position at the sensor rate (meaning position changes as audio elements move). The stream selection unit 44 can further update the energy map at the audio frame rate (meaning the energy map is updated once per frame).
[0226] Next reference Figure 9B For example, the flow selection unit 44 can be used as described above. Figure 9A The operation differs from the described method in that the stream selection unit 44 may not determine CM 47 based on the energy map. Therefore, the stream selection unit 44 can obtain bitstream 27 from all enabled audio elements, where bitstream 27 may include corresponding information (e.g., metadata), such as ALI 45A (840). The stream selection unit 44 can determine whether privacy settings are enabled or disabled for the addition of audio elements (where privacy settings may refer to digital access permissions, such as passwords, authorization levels or grades, time, etc., restricting access to one or more of the bitstreams 27) (842, 844).
[0227] When the privacy settings enable the addition of receivers, the stream selection unit 44 can add audio elements to the updated CM 47 (which means adding an audio stream to audio data 19') (846). When the privacy settings disable the addition of receivers, the stream selection unit 44 can remove audio elements from the updated CM 47 (which means removing one or more audio streams from audio data 19') (848). In this way, the stream selection unit 44 can identify the enabling of a new set of audio elements (850). The stream selection unit 44 can iterate (852) through different combinations of audio elements in CM 47 to determine the constellation history (854) representing audio data 19'.
[0228] The stream selection unit 44 can iterate and update various inputs at any given frequency in this way. For example, the stream selection unit 44 can update privacy settings at the UI rate (meaning updates are driven by updates via UI input). The stream selection unit 44 can update position at the sensor rate (meaning position changes as audio elements move). The stream selection unit 44 can further update the energy map at the audio frame rate (meaning the energy map is updated once per frame).
[0229] Next reference Figure 9C For example, the flow selection unit 44 can be configured as described above. Figure 9A The operation differs from the described method in that the stream selection unit 44 can determine CM 47 without relying on the privacy settings of the enabled audio elements. Therefore, the stream selection unit 44 can obtain bitstream 27 (e.g., audio stream) from all enabled audio elements, where bitstream 27 may include corresponding metadata (e.g., audio metadata, PSI, ALI 45A, etc.) (860). The stream selection unit 44 can perform energy analysis on each of the bitstream 27 to calculate the corresponding energy map (862).
[0230] The stream selection unit 44 can then iterate through different combinations of audio elements (defined in CM 47) (864) based on proximity to audio source 308 (as defined by audio source distances 306A and / or 306B) and proximity to audio elements (as defined by proximity distances discussed above). Figure 9C As shown, audio elements can be sorted or otherwise associated with different access permissions. Stream selection unit 44 can iterate in the manner described above based on the listener position represented by DLI 45B (which again refers to another way of the “virtual position” or “device position” discussed above) and the audio element position represented by ALI 45A to identify whether a larger subset or a reduced subset of bitstream 27 is needed (866, 868).
[0231] When a larger subset of bitstream 27 is needed, stream selection unit 44 can add audio elements to audio data 19', or in other words, add additional audio streams (such as when the user is closer). Figure 3A (as in the example audio source) (870). When a reduced subset of bitstream 27 is needed, stream selection unit 44 can remove audio elements from audio data 19', or in other words, remove one or more existing audio streams (such as when the user is further away). Figure 3A (872) in the example of the audio source.
[0232] In some examples, the stream selection unit 44 may determine that the current constellation of the audio element is the optimal set (or, in other words, the existing audio data 19' will remain the same as the selection process described herein, resulting in the same audio data 19') (864), and the process may return to 862. However, when an audio stream is added to or removed from the audio data 19', the stream selection unit 44 may update CM 47 (874), generating a constellation history (875).
[0233] The stream selection unit 44 can iterate and update various inputs according to any given frequency in this way. For example, the stream selection unit 44 can update the position at the sensor rate (meaning the position changes as the audio element moves). The stream selection unit 44 can further update the energy map at the audio frame rate (meaning the energy map is updated once per frame).
[0234] It should be recognized that, depending on the example, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or excluded entirely (e.g., not all described actions or events are necessary for technical practice). Furthermore, in some examples, actions or events may be performed concurrently, for example, through multithreaded processing, interrupt handling, or multiple processors, rather than sequentially.
[0235] In some examples, the VR device (or streaming device) may use a network interface coupled to the VR / streaming device's memory to transmit exchange messages to an external device, where the exchange messages are associated with multiple available representations of the sound field. In some examples, the VR device may use an antenna coupled to the network interface to receive wireless signals, including data packets, audio packets, video protocol data, or transport protocol data associated with multiple available representations of the sound field. In some examples, one or more microphone arrays may capture the sound field.
[0236] In some examples, multiple available representations of a sound field stored in a memory device may include multiple object-based representations of the sound field, higher-order surround sound representations of the sound field, mixed-order surround sound representations of the sound field, a combination of object-based representations of the sound field and higher-order surround sound representations of the sound field, a combination of object-based representations of the sound field and mixed-order surround sound representations of the sound field, or a combination of mixed-order representations of the sound field and higher-order surround sound representations of the sound field.
[0237] In some examples, one or more of the multiple available representations of the sound field may include at least one high-resolution region and at least one low-resolution region, and wherein the selected representation based on the steering angle provides greater spatial accuracy with respect to at least one high-resolution region and lower spatial accuracy with respect to the low-resolution region.
[0238] Figure 10 An example of a wireless communication system 1002 adapted to support parameters according to various aspects of this disclosure is shown. The wireless communication system 1002 includes a base station 105, a user equipment (UE) 115, and a core network 130. In some examples, the wireless communication system 1002 may be a Long Term Evolution (LTE) network, an Advanced LTE (LTE-A) network, an LTE-APro network, a 5G cellular network, or a New Radio (NR) network. In some cases, the wireless communication system 1002 may support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low latency communication, or communication with low-cost and low-complexity devices.
[0239] Base station 105 can wirelessly communicate with UE 115 via one or more base station antennas. Base station 105 described herein may include, or be referred to by those skilled in the art as, a base station transceiver, radio base station, access point, radio transceiver, NodeB, eNodeB (eNB), next-generation NodeB or giga-NodeB (any of which may be referred to as gNB), home NodeB, home eNodeB, or other suitable terms. Wireless communication system 1002 may include different types of base stations 105 (e.g., macro cell base stations or small cell base stations). UE 115 described herein may be able to communicate with various types of base stations 105 and network devices, including macro eNBs, small cell eNBs, gNBs, and relay base stations.
[0240] Each base station 105 may be associated with a specific geographic coverage area 110 in which communication with various UEs 115 is supported. Each base station 105 may provide communication coverage to the corresponding geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 may utilize one or more carriers. The communication link 125 shown in the wireless communication system 1002 may include uplink transmission from the UE 115 to the base station 105, or downlink transmission from the base station 105 to the UE 115. Downlink transmission may also be referred to as forward link transmission, and uplink transmission may also be referred to as reverse link transmission.
[0241] The geographic coverage area 110 of base station 105 can be divided into sectors that constitute part of the geographic coverage area 110, and each sector can be associated with a cell. For example, each base station 105 can provide communication coverage for macro cells, small cells, hotspots, or other types of cells, or various combinations thereof. In some examples, base station 105 can be mobile and thus provide communication coverage for mobile geographic coverage areas 110. In some examples, different geographic coverage areas 110 associated with different technologies can overlap, and the same base station 105 or different base stations 105 can support overlapping geographic coverage areas 110 associated with different technologies. The wireless communication system 1002 can include, for example, heterogeneous LTE / LTE-A / LTE-A Pro, 5G cellular, or NR networks, wherein different types of base stations 105 provide coverage for various geographic coverage areas 110.
[0242] UE 115 may be distributed throughout the wireless communication system 1002, and each UE 115 may be fixed or mobile. UE 115 may also be referred to as a mobile device, wireless device, remote device, handheld device, or subscriber device, or some other suitable term, wherein "device" may also be referred to as a unit, station, terminal, or client. UE 115 may also be a personal electronic device, such as a cellular phone, personal digital assistant (PDA), tablet computer, laptop computer, or personal computer. In the examples of this disclosure, UE 115 may be any audio source described in this disclosure, including VR headsets, XR headsets, AR headsets, vehicles, smartphones, microphones, microphone arrays, or any other device including microphones or capable of transmitting captured and / or synthesized audio streams. In some examples, the synthesized audio stream may be an audio stream stored in memory or previously generated (e.g., created, synthesized, etc.). In some examples, UE 115 may also refer to a wireless local loop (WLL) station, Internet of Things (IoT) device, Internet of Everything (IoE) device, or machine-type communication (MTC) device, which can be implemented in various products such as appliances, vehicles, and instruments.
[0243] Some UEs 115, such as MTC or IoT devices, can be low-cost or low-complexity devices and can provide automated communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC can refer to data communication technologies that allow devices to communicate with each other or with base station 105 without human intervention. In some examples, M2M communication or MTC may include communication from switching and / or using parameter settings and adjustments, such as gain or frequency response adjustments, which indicate parameter adjustments and / or energy curve overlay data adjustments of one or more microphones and / or audio sources (e.g., audio elements) capturing various audio streams.
[0244] In some cases, UE 115 may also be able to communicate directly with other UE 115s (e.g., using peer-to-peer (P2P) or device-to-device (D2D) protocols). One or more of a group of UEs 115s utilizing D2D communication may be within the geographic coverage area 110 of base station 105. Other UEs 115s in this group may be outside the geographic coverage area 110 of base station 105 or unable to receive transmissions from base station 105. In some cases, multiple groups of UEs 115s communicating via D2D communication may utilize a one-to-many (1:M) system, where each UE 115 transmits to every other UE 115 in the group. In some cases, base station 105 facilitates the scheduling of resources for D2D communication. In other cases, D2D communication between UEs 115s is performed without the involvement of base station 105.
[0245] Base station 105 can communicate with core network 130 and with each other. For example, base station 105 can interface with core network 130 via backhaul link 132 (e.g., via S1, N2, N3 or other interfaces). Base stations 105 can communicate with each other directly (e.g., directly between base stations 105) or indirectly (e.g., via core network 130) via backhaul link 134 (e.g., via X2, Xn or other interfaces).
[0246] In some cases, wireless communication system 1002 may utilize both licensed and unlicensed radio spectrum bands. For example, wireless communication system 1002 may employ Licensed Assisted Access (LAA), unlicensed LTE (LTE-U) Radio Access Technology (RAT), or NR technology in unlicensed bands such as the 5 GHz Industrial, Scientific and Medical (ISM) band. When operating in unlicensed radio frequency spectrum bands, wireless devices such as base station 105 and UE 115 may employ a Listen-Before-Talk (LBT) procedure to ensure that the channel is cleared before transmitting data. In some cases, operation in unlicensed bands may be based on a combination of carrier aggregation configuration and component carriers operating in licensed bands (e.g., LAA). Operation in unlicensed spectrum may include downlink transmission, uplink transmission, peer-to-peer transmission, or a combination thereof. Duplexing in unlicensed spectrum may be based on Frequency Division Duplex (FDD), Time Division Duplex (TDD), or a combination of both.
[0247] This disclosure includes the following examples:
[0248] Example 1A: An audio device configured to determine parameter adjustments for audio capture, the audio device comprising: a memory configured to store at least one energy map corresponding to one or more audio streams; and one or more processors coupled to the memory and configured to: access the at least one energy map corresponding to the one or more audio streams; determine parameter adjustments with respect to at least one microphone based at least in part on the at least one energy map, the parameter adjustments being configured to adjust the audio capture through the at least one microphone; and output an indication of the parameter adjustments with respect to the at least one microphone.
[0249] Example 2A: The audio device of claim 1A, wherein the one or more processors are configured to perform energy analysis on the one or more audio streams to determine the at least one energy map.
[0250] Example 3A: An audio device according to any combination of Examples 1A and 2A, wherein the one or more processors are configured to: compare the at least one energy map with one or more other energy maps corresponding to audio captured by the at least one microphone; and determine the parameter adjustment based at least in part on the comparison between the at least one energy map and the one or more other energy maps.
[0251] Example 4A: An audio device according to any combination of Examples 1A to 3A, wherein the one or more processors are configured to receive at least one of the following from one or more source devices: the at least one energy map and the one or more other energy maps.
[0252] Example 5A: An audio device according to any combination of Examples 1A to 4A, wherein the at least one energy map comprises a plurality of energy map components.
[0253] Example 6A: The audio device according to Example 5A, wherein the energy map component corresponds to the one or more audio streams.
[0254] Example 7A: An audio device according to any combination of Examples 1A to 6A, wherein one or more processors are configured to analyze at least one of the following in determining the parameter adjustments: gain and frequency response.
[0255] Example 8A: An audio device according to any combination of Examples 1A to 7A, wherein the one or more processors are configured to: determine the parameter adjustment to modify the capture of the one or more audio streams.
[0256] Example 9A: An audio device according to any combination of Examples 1A to 8A, wherein the parameter adjustment includes adjustment of the gain of the at least one microphone.
[0257] Example 10A: An audio device according to Example 9A, wherein the gain is frequency-dependent.
[0258] Example 11A: An audio device according to any combination of Examples 1A to 10A, wherein one or more processors are configured to receive audio by adjusting one or more parameter settings of the at least one microphone according to the parameters.
[0259] Example 12A: An audio device according to any combination of Examples 1A to 11A, wherein one or more processors are configured to send the parameter adjustment to a first source device corresponding to the at least one microphone.
[0260] Example 13A: An audio device according to any combination of Examples 1A to 12A, wherein determining the parameter adjustment includes determining a difference score with respect to the one or more audio streams.
[0261] Example 14A: The audio device according to Example 13A, wherein the difference score increases when there is a discontinuity with respect to at least one of the one or more audio streams.
[0262] Example 15A: An audio device according to Example 14A, wherein the discontinuity includes gaps in the frequency response of the at least one audio stream.
[0263] Example 16A: An audio device according to any combination of Examples 13A to 15A, wherein one or more processors are configured to: compare the difference score with a difference threshold; and determine the parameter adjustment based at least in part on the comparison of the difference score with the difference threshold.
[0264] Example 17A: An audio device according to any combination of Examples 1A to 16A, wherein determining the parameter adjustment includes determining a change in the gain of the one or more audio streams.
[0265] Example 18A: An audio device according to any combination of Examples 1A to 17A, wherein one or more processors are configured to render an energy graph overlay based at least in part on the at least one energy graph.
[0266] Example 19A: An audio device according to Example 18A, wherein one or more processors are configured to output an energy graph overlay for display to a user.
[0267] Example 20A: An audio device according to any combination of Examples 1A to 19A, wherein the one or more processors are configured to: access diagnostic data of at least one of the one or more audio streams; determine quality characteristics of the one or more audio streams based at least in part on the diagnostic data; modify at least one of the following: the at least one energy map and the one or more audio streams based at least in part on the quality characteristics; and determine parameter adjustments based at least in part on the modifications.
[0268] Example 21A: An audio device according to any combination of Examples 1A to 20A, wherein the one or more processors are configured to: determine a license state corresponding to at least one of the one or more audio streams; modify at least one of the following based at least in part on the license state: the at least one energy map and the one or more audio streams; and determine the parameter adjustment based at least in part on the modification.
[0269] Example 22A: An audio device according to Example 21A, wherein the license status indicates whether the one or more audio streams are restricted or unrestricted.
[0270] Example 23A: An audio device according to any combination of Examples 1A to 22A, wherein the one or more processors are configured to: determine a feasibility state of the one or more microphones, the feasibility state indicating a feasibility score of the one or more microphones; modify at least one of the following based at least in part on the feasibility state: the at least one energy map and the one or more audio streams; and determine the parameter adjustment based at least in part on the modification.
[0271] Example 24A: An audio device according to any combination of Examples 20A to 23A, wherein the modification includes adjusting the number of energy map components used to determine the at least one energy map.
[0272] Example 25A: An audio device according to any combination of Examples 20A to 24A, wherein the modification includes removing at least one audio stream from the one or more audio streams.
[0273] Example 26A: An audio device according to any combination of Examples 20A to 25A, wherein one or more processors are configured to receive the diagnostic data as self-diagnostic data.
[0274] Example 27A: An audio device according to any combination of Examples 20A to 26A, wherein the diagnostic data includes at least one of the following: signal-to-noise ratio information and gain.
[0275] Example 28A: An audio device according to any combination of Examples 20A to 27A, wherein determining the quality characteristic includes marking at least one of the one or more audio streams as a non-compliant audio stream.
[0276] Example 29A: An audio device according to any combination of Examples 1A to 28A, wherein one or more processors are configured to receive an adjustment state.
[0277] Example 30A: An audio device according to Example 29A, wherein the adjustment state indicates a successful adjustment of the at least one microphone receiving audio according to the parameters.
[0278] Example 31A: An audio device according to any combination of Examples 29A and 30A, wherein the adjustment state indicates that the at least one microphone is receiving audio.
[0279] Example 32A: An audio device according to any combination of Examples 1A to 31A, wherein one or more processors are configured to periodically update the at least one energy map with respect to the audio frame rate.
[0280] Example 33A: An audio device according to any combination of Examples 1A to 32A, wherein the audio device includes a wearable device.
[0281] Example 34A: An audio device according to any combination of Examples 1A to 33A, wherein the audio device includes a mobile device.
[0282] Example 35A: An audio device according to any combination of Example 35A, wherein the mobile device includes a mobile handheld terminal.
[0283] Example 36A: An audio device according to any combination of Examples 1A to 35A, wherein the audio device includes the at least one microphone.
[0284] Example 37A: An audio device according to any combination of Examples 1A to 36A, wherein the audio device includes headphones coupled to one or more speakers.
[0285] Example 38A: An audio device according to any combination of Examples 1A to 37A, wherein the audio device includes one or more speakers.
[0286] Example 39A: An audio device according to any combination of Examples 1A to 38A, wherein the audio device includes an extended reality (XR) headset coupled to one or more speakers.
[0287] Example 40A: The audio device according to Example 39A, wherein the XR headset includes one or more of augmented reality headsets, virtual reality headsets, or mixed reality headsets.
[0288] Example 41A: An audio device according to any combination of Examples 1A to 41A, wherein the audio device includes one or more speakers configured to generate a sound field.
[0289] Example 42A: An audio device according to any combination of Examples 1A to 41A, wherein the at least one microphone is configured to provide a six-degrees-of-freedom user experience.
[0290] Example 43A: An audio device according to any combination of Examples 1A to 42A, wherein the audio device includes an audio receiver that is enabled to receive audio.
[0291] Example 44A: An audio device according to any combination of Examples 1A to 43A, wherein the audio device includes a receiver configured to receive the one or more audio streams.
[0292] Example 45A: An audio device according to Example 44A, wherein the receiver includes a receiver configured to receive the one or more audio streams according to a 5G cellular standard.
[0293] Example 46A: An audio device according to Example 44A, wherein the receiver includes a receiver configured to receive the one or more audio streams according to a personal area network standard.
[0294] Example 47A: An audio device according to any combination of Examples 1A to 46A, wherein the one or more processors are configured to receive at least one of the following via a wireless link: the one or more audio streams and the at least one energy map.
[0295] Example 48A: The audio device according to Example 47A, wherein the wireless link is via a 5G air interface.
[0296] Example 49A: The audio device according to Example 47A, wherein the wireless link is via a Bluetooth interface.
[0297] Example 50A: An audio device according to any combination of Examples 1A to 49A, wherein the audio device includes a remote server configured to determine the at least one energy map.
[0298] Example 51A: A method for determining parameter adjustments for audio capture, the method comprising: accessing at least one energy map corresponding to one or more audio streams; determining parameter adjustments with respect to at least one microphone based at least in part on the at least one energy map, the parameter adjustments being configured to adjust the audio capture through the at least one microphone; and outputting an indication of the parameter adjustments with respect to the at least one microphone.
[0299] Example 52A: According to the method of Example 51A, the method further includes: performing an energy analysis on the one or more audio streams to determine the at least one energy map.
[0300] Example 53A: The method according to any combination of Examples 51A and 52A, the method further comprising: comparing the at least one energy map with one or more other energy maps, the one or more other energy maps corresponding to audio captured by the at least one microphone; and determining the parameter adjustment based at least in part on the comparison between the at least one energy map and the one or more energy maps.
[0301] Example 54A: The method according to any combination of Examples 51A to 53A, the method further comprising: receiving from one or more source devices at least one of the following: the at least one energy map and the other energy map.
[0302] Example 55A: The method according to any combination of Examples 51A to 54A, wherein the at least one energy map comprises a plurality of energy map components.
[0303] Example 56A: The method according to Example 55A, wherein the energy map component corresponds to the one or more audio streams.
[0304] Example 57A: The method according to any combination of Examples 51A to 56A, the method further comprising: analyzing at least one of the following in determining the parameter adjustment: gain and frequency response.
[0305] Example 58A: The method according to any combination of Examples 51A to 57A, the method further comprising: determining the parameter adjustment to modify the capture of the one or more audio streams.
[0306] Example 59A: The method according to any combination of Examples 51A to 58A, wherein the parameter adjustment includes adjusting the gain of the at least one microphone.
[0307] Example 60A: The method described in Example 59A, wherein the gain is frequency-dependent.
[0308] Example 61A: The method according to any combination of Examples 51A to 60A, the method comprising: receiving audio by adjusting one or more parameter settings of the at least one microphone according to the parameters.
[0309] Example 62A: The method according to any combination of Examples 51A to 61A, the method further comprising: sending the parameter adjustment to a first source device corresponding to the at least one microphone.
[0310] Example 63A: The method according to any combination of Examples 51A to 62A, wherein determining the parameter adjustment includes: determining a difference score with respect to the one or more audio streams.
[0311] Example 64A: According to the method of Example 63A, the method further includes: increasing the difference score when there is a discontinuity with respect to at least one of the one or more audio streams.
[0312] Example 65A: According to the method of Example 64A, the discontinuity includes gaps in the frequency response of the at least one audio stream.
[0313] Example 66A: The method according to any combination of Examples 63A to 65A, the method further comprising: comparing the difference score with a difference threshold; and determining the parameter adjustment based at least in part on the comparison of the difference score with the difference threshold.
[0314] Example 67A: The method according to any combination of Examples 51A to 66A, wherein determining the parameter adjustment includes: determining a change in the gain of the one or more audio streams.
[0315] Example 68A: The method according to any combination of Examples 51A to 67A, the method comprising: rendering an energy curve overlay based at least in part on the at least one energy map.
[0316] Example 69A: According to the method described in Example 68A, the method further includes: outputting the energy curve overlay to display to the user.
[0317] Example 70A: The method according to any combination of Examples 51A to 69A, the method further comprising: accessing diagnostic data of at least one of the one or more audio streams; determining quality characteristics of the one or more audio streams based at least in part on the diagnostic data; modifying at least one of the following: the one or more energy maps and the plurality of audio streams based at least in part on the quality characteristics; and determining the parameter adjustment based at least in part on the modification.
[0318] Example 71A: The method according to any combination of Examples 51A to 70A, the method further comprising: determining a license state corresponding to at least one of the one or more audio streams; modifying at least one of the following: the one or more energy maps and the plurality of audio streams, at least in part based on the license state; and determining the parameter adjustment, at least in part based on the modification.
[0319] Example 72A: The method according to Example 71A, wherein the permission status indicates whether the one or more audio streams are restricted or unrestricted.
[0320] Example 73A: The method according to any combination of Examples 51A to 72A, the method further comprising: determining a feasibility state of the one or more microphones, the feasibility state indicating a feasibility score of the one or more microphones; modifying at least one of the following: the at least one energy map and the one or more audio streams, at least in part based on the feasibility state; and determining the parameter adjustment, at least in part based on the modification.
[0321] Example 74A: The method according to any combination of Examples 70A to 73A, wherein the modification includes: adjusting the number of energy map components used to determine the at least one energy map.
[0322] Example 75A: The method according to any combination of Examples 70A to 74A, wherein the modification includes removing at least one audio stream from the one or more audio streams.
[0323] Example 76A: The method according to any combination of Examples 70A to 75A, the method further comprising: receiving the diagnostic data as self-diagnostic data.
[0324] Example 77A: The method according to any combination of Examples 70A to 76A, wherein the diagnostic data includes at least one of the following: signal-to-noise ratio information and gain information.
[0325] Example 78A: The method according to any combination of Examples 70A to 77A, wherein determining the quality characteristic includes: marking at least one of the one or more audio streams as an audio stream.
[0326] Example 79A: The method according to any combination of Examples 51A to 78A, the method further comprising: determining an adjustment state.
[0327] Example 80A: The method according to Example 79A, wherein the adjustment state indicates that the at least one microphone is receiving audio.
[0328] Example 81A: The method according to Example 80A, wherein the adjustment state is adjusted according to the parameters to indicate that at least one microphone is receiving audio.
[0329] Example 82A: The method according to any combination of Examples 51A to 81A, the method further comprising: periodically updating the at least one energy map with respect to the audio frame rate.
[0330] Example 83A: The method according to any combination of Examples 51A to 82A, wherein at least one audio device performs the method, the at least one audio device including a wearable device.
[0331] Example 84A: The method according to any combination of Examples 51A to 83A, wherein at least one audio device performs the method, the at least one audio device including a mobile device.
[0332] Example 85A: The method according to Example 84A, wherein the mobile device includes a mobile handheld terminal.
[0333] Example 86A: The method according to any combination of Examples 51A to 85A, wherein at least one audio device performs the method, the at least one audio device including the at least one microphone.
[0334] Example 87A: The method according to any combination of Examples 51A to 86A, wherein at least one audio device performs the method, the at least one audio device including headphones configured to be coupled to the one or more speakers.
[0335] Example 88A: The method according to any combination of Examples 51A to 87A, wherein at least one audio device performs the method, the at least one audio device comprising one or more speakers.
[0336] Example 89A: The method according to any combination of Examples 51A to 88A, wherein at least one audio device performs the method, the at least one audio device including an extended reality (XR) headset configured to be coupled to the one or more speakers.
[0337] Example 90A: The method according to Example 89A, wherein the XR headset includes one or more of augmented reality headsets, virtual reality headsets, or mixed reality headsets.
[0338] Example 91A: The method according to any combination of Examples 51A to 90A, wherein at least one audio device performs the method, the at least one audio device including one or more speakers configured to generate a sound field.
[0339] Example 92A: The method according to any combination of Examples 51A to 91A, wherein at least one audio device performs the method, wherein the at least one audio device is configured to provide a six-degrees-of-freedom user experience.
[0340] Example 93A: The method according to any combination of Examples 51A to 92A, wherein at least one audio device performs the method, the at least one audio device including at least one audio receiver, the audio receiver being enabled to receive audio from one or more source devices.
[0341] Example 94A: The method according to any combination of Examples 51A to 93A, wherein at least one audio device performs the method, the at least one audio device including at least one audio receiver configured to receive the one or more audio streams.
[0342] Example 95A: The method according to Example 94A, wherein the receiver includes a receiver configured to receive the one or more audio streams according to a 5G cellular standard.
[0343] Example 96A: The method according to Example 94A, wherein the receiver includes a receiver configured to receive the one or more audio streams according to a personal area network standard.
[0344] Example 97A: The method according to any combination of Examples 51A to 96A, the method further comprising: receiving at least one of the following via a wireless link: the one or more audio streams and the at least one energy map.
[0345] Example 98A: The method described in Example 97A, wherein the wireless link is via a 5G air interface.
[0346] Example 99A: The method described in Example 97A, wherein the wireless link is via a Bluetooth interface.
[0347] Example 100A: The method according to any combination of Examples 51A to 99A, wherein the audio device includes a remote server configured to determine the at least one energy map.
[0348] Example 101A: An audio device configured to adjust audio capture, the audio device comprising: means for accessing at least one energy map, the at least one energy map corresponding to one or more audio streams; means for determining parameter adjustments with respect to at least one microphone based at least in part on the at least one energy map, the parameter adjustments being configured to adjust the audio capture through the at least one microphone; and means for outputting an indication of the parameter adjustments with respect to the at least one microphone.
[0349] Example 102A: The audio device according to Example 101A further includes: a component for performing energy analysis on the one or more audio streams to determine the at least one energy map.
[0350] Example 103A: An audio device according to any combination of Examples 101A and 102A, the audio device further comprising: means for comparing the at least one energy map with one or more other energy maps corresponding to audio captured by the at least one microphone; and means for determining the parameter adjustment based at least in part on the comparison between the at least one energy map and the one or more other energy maps.
[0351] Example 104A: The audio device according to any combination of Examples 101A to 103A, the audio device further comprising: a component for receiving at least one of the following from one or more source devices: the at least one energy map and the one or more energy maps.
[0352] Example 105A: An audio device according to any combination of Examples 101A to 104A, wherein the at least one energy map comprises a plurality of energy map components.
[0353] Example 106A: An audio device according to Example 105A, wherein the energy map component corresponds to the one or more audio streams.
[0354] Example 107A: The audio device according to any combination of Examples 101A to 106A, wherein the component for determining the parameter adjustment further includes a component for analyzing at least one of the following: gain and frequency response.
[0355] Example 108A: An audio device according to any combination of Examples 101A to 107A, wherein the parameter adjustment is configured to modify the capture of one or more audio streams.
[0356] Example 109A: An audio device according to any combination of Examples 101A to 108A, wherein the parameter adjustment includes adjustment of the gain of the at least one microphone.
[0357] Example 110A: An audio device according to Example 109A, wherein the gain is frequency-dependent.
[0358] Example 111A: The audio device according to any combination of Examples 101A to 110A further includes: a component for adjusting one or more parameter settings using the at least one microphone according to the parameters.
[0359] Example 112A: The audio device according to any combination of Examples 101A to 111A further includes: a component for sending the parameter adjustment to a first source device corresponding to the at least one microphone.
[0360] Example 113A: The audio device according to any combination of Examples 101A to 112A, wherein the component for determining the parameter adjustment further includes: a component for determining a difference score with respect to one or more audio streams.
[0361] Example 114A: The audio device according to Example 113A, wherein the difference score increases when there is a discontinuity between at least one of the one or more audio streams.
[0362] Example 115A: An audio device according to Example 114A, wherein the discontinuity includes gaps in the frequency response of the at least one audio stream.
[0363] Example 116A: An audio device according to any combination of Examples 113A to 115A, the audio device further comprising: a component for comparing the difference score with a difference threshold; and a component for determining the parameter adjustment based at least in part on the comparison of the difference score with the difference threshold.
[0364] Example 117A: An audio device according to any combination of Examples 101A to 116A, wherein the component for determining the parameter adjustment further includes: a component for determining a change in the gain of the one or more audio streams, wherein the difference is at least partially based on a change in the gain of the at least one audio stream.
[0365] Example 118A: The audio device according to any combination of Examples 101A to 117A further includes: a component for rendering an energy curve overlay based at least in part on the at least one energy map.
[0366] Example 119A: The audio device according to Example 118A further includes a component for outputting the energy curve overlay for display to a user.
[0367] Example 120A: An audio device according to any combination of Examples 101A to 119A, the audio device further comprising: means for accessing diagnostic data of at least one of the one or more audio streams; means for determining quality characteristics of the one or more audio streams at least in part based on the diagnostic data; means for modifying at least one of the following at least in part based on the quality characteristics: accessing the at least one energy map and the one or more audio streams; and means for determining the parameter adjustment at least in part based on the modification.
[0368] Example 121A: An audio device according to any combination of Examples 101A to 120A, the audio device further comprising: components for determining a license state corresponding to at least one of the one or more audio streams; components for modifying at least one of the following based at least in part on the license state: the at least one energy map and the one or more audio streams; and components for determining the parameter adjustment based at least in part on the modification.
[0369] Example 122A: An audio device according to Example 121A, wherein the license status indicates whether the one or more audio streams are restricted or unrestricted.
[0370] Example 123A: An audio device according to any combination of Examples 101A to 122A, the audio device further comprising: components for determining a feasibility state of the one or more microphones, the feasibility state indicating a feasibility score of the one or more microphones; components for modifying at least one of the following based at least in part on the feasibility state: the at least one energy map and the one or more audio streams; and components for determining the parameter adjustment based at least in part on the modification.
[0371] Example 124A: The audio device according to any combination of Examples 120A to 123A, wherein the modification component further includes: adjusting the number of energy map components used to determine the at least one energy map.
[0372] Example 125A: An audio device according to any combination of Examples 120A to 124A, wherein the modification component further includes a component for removing at least one audio stream from the one or more audio streams.
[0373] Example 126A: The audio device according to any combination of Examples 120A to 125A, the audio device further includes: a component for receiving the diagnostic data as self-diagnostic data.
[0374] Example 127A: An audio device according to any combination of Examples 120A to 126A, wherein the diagnostic data includes at least one of the following: signal-to-noise ratio information and gain level information.
[0375] Example 128A: The audio device according to any combination of Examples 120A to 127A, the audio device further comprising: a component for marking at least one of the one or more audio streams as a defective audio stream, at least in part based on the quality characteristic.
[0376] Example 129A: The audio device according to any combination of Examples 101A to 128A, the audio device further includes: a component for determining an adjustment state.
[0377] Example 130A: An audio device according to Example 129A, wherein the component for determining the adjustment state includes a component for indicating successful adjustment of the at least one microphone receiving audio according to the parameters.
[0378] Example 131A: An audio device according to any combination of Examples 129A and 130A, wherein the component for determining the adjustment state includes a component for indicating that the at least one microphone is receiving audio.
[0379] Example 132A: An audio device according to any combination of Examples 101A to 131A, the audio device further comprising: a component for periodically updating the at least one energy map with respect to the audio frame rate.
[0380] Example 133A: An audio device according to any combination of Examples 101A to 132A, wherein the audio device includes a wearable device.
[0381] Example 134A: An audio device according to any combination of Examples 101A to 133A, wherein the audio device includes a mobile device.
[0382] Example 135A: An audio device according to any combination of Examples 101A to 134A, wherein the mobile device includes a mobile handheld terminal.
[0383] Example 136A: An audio device according to any combination of Examples 101A to 135A, wherein the audio device includes at least one of the microphones.
[0384] Example 137A: An audio device according to any combination of Examples 101A to 136A, wherein the audio device includes headphones coupled to one or more speakers.
[0385] Example 138A: An audio device according to any combination of Examples 101A to 137A, wherein the audio device includes one or more speakers.
[0386] Example 139A: An audio device according to any combination of Examples 101A to 138A, wherein the audio device includes an extended reality (XR) headset coupled to one or more speakers.
[0387] Example 140A: The audio device according to Example 139A, wherein the XR headset includes one or more of augmented reality headsets, virtual reality headsets, or mixed reality headsets.
[0388] Example 141A: An audio device according to any combination of Examples 101A to 140A, wherein the audio device includes components for generating a sound field.
[0389] Example 142A: The audio device according to any combination of Examples 101A to 141A further includes components for providing a six-degrees-of-freedom user experience.
[0390] Example 143A: An audio device according to any combination of Examples 101A to 142A, wherein the audio device includes at least one audio receiver, the audio receiver including components for receiving audio from one or more source devices.
[0391] Example 144A: An audio device according to any combination of Examples 101A to 143A, wherein the audio device includes a receiver, the receiver including components for receiving the one or more audio streams.
[0392] Example 145A: An audio device according to Example 144A, wherein the receiver includes a receiver comprising components for receiving the one or more audio streams according to a 5G cellular standard.
[0393] Example 146A: An audio device according to Example 144A, wherein the receiver includes a receiver comprising components for receiving the one or more audio streams according to a personal area network standard.
[0394] Example 147A: An audio device according to any combination of Examples 101A to 146A, further comprising: a component for receiving at least one of the following via a wireless link: the one or more audio streams and the at least one energy map.
[0395] Example 148A: The audio device according to Example 147A, wherein the component for receiving via a wireless link includes a 5G air interface.
[0396] Example 149A: The audio device according to Example 147A, wherein the component for receiving via a wireless link includes a Bluetooth interface.
[0397] Example 150A: An audio device according to any combination of Examples 111A to 149A, wherein the audio device includes a remote server, the remote server including components for determining at least one energy map.
[0398] Example 151A: A non-transitory computer-readable storage medium having instructions thereon that, when executed, cause one or more processors of an audio device to: access at least one energy map corresponding to one or more audio streams; determine, at least in part, a parameter adjustment with respect to at least one microphone based on the at least one energy map, the parameter adjustment being configured to adjust the audio capture through the at least one microphone; and output an indication of the parameter adjustment with respect to the at least one microphone.
[0399] Example 1B: An audio device configured to generate a sound field, the audio device comprising: a memory configured to store audio data representing the sound field; and one or more processors coupled to the memory and configured to: send an audio stream to one or more source devices; determine instructions for adjusting parameter settings of the audio device; and adjust the parameter settings to adjust the generation of the sound field.
[0400] Example 2B: An audio device according to Example 1B, wherein the one or more processors are configured to send diagnostic data to the one or more source devices.
[0401] Example 3B: An audio device according to any combination of Examples 1B and 2B, wherein the one or more processors are configured to: perform energy analysis on the audio stream to determine at least one energy map; and send the at least one energy map to the one or more source devices.
[0402] Example 4B: An audio device according to any combination of Examples 1B to 3B, wherein the parameter settings are configured to modify the capture of the audio stream.
[0403] Example 5B: An audio device according to any combination of Examples 1B to 4B, wherein the parameter settings include adjustment of the frequency-dependent gain of the audio device.
[0404] Example 6B: An audio device according to any combination of Examples 1B to 5B, wherein one or more processors are configured to render an energy graph overlay based at least in part on the at least one energy graph.
[0405] Example 7B: The device according to Example 6B, wherein the energy graph overlay is rendered at least in part based on a composite energy graph, the composite energy graph being at least in part based on the at least one energy graph.
[0406] Example 8B: An audio device according to any combination of Examples 6B and 7B, wherein one or more processors are configured to output an energy graph overlay for display to a user.
[0407] Example 9B: An audio device according to any combination of Examples 1B to 8B, wherein the one or more processors are configured to: determine the quality characteristics of the audio stream; and send the quality characteristics to the one or more source devices.
[0408] Example 10B: An audio device according to any combination of Examples 1B to 9B, wherein the one or more processors are configured to send at least one of the following to the one or more source devices: a license status and a feasibility status.
[0409] Example 11B: An audio device according to any combination of Examples 1B to 10B, wherein the instructions are received from the one or more source devices.
[0410] Example 12B: An audio device according to any combination of Examples 1B to 11B, wherein the one or more processors are configured to receive a composite energy map from the one or more source devices.
[0411] Example 13B: An audio device according to any combination of Examples 1B to 12B, wherein the instructions are at least partially based on a composite energy map.
[0412] Example 14B: An audio device according to any combination of Examples 1B to 13B, wherein the one or more processors are configured to receive diagnostic data from the one or more source devices, wherein the diagnostic data includes at least one of the following: signal-to-noise ratio information and sound level information.
[0413] Example 15B: An audio device according to any combination of Examples 1B to 14B, wherein one or more processors are configured to: send an adjustment state.
[0414] Example 16B: An audio device according to any combination of Example 15B, wherein the adjustment state indicates that the audio device is receiving audio.
[0415] Example 17B: An audio device according to any combination of Examples 1B to 16B, wherein the audio device includes a wearable device.
[0416] Example 18B: An audio device according to any combination of Examples 1B to 17B, wherein the audio device includes a mobile device.
[0417] Example 19B: An audio device according to any combination of Example 18B, wherein the mobile device includes a mobile handheld terminal.
[0418] Example 20B: An audio device according to any combination of Examples 1B to 19B, wherein the audio device includes at least one microphone.
[0419] Example 21B: An audio device according to any combination of Examples 1B to 20B, wherein the audio device includes headphones coupled to one or more speakers.
[0420] Example 22B: An audio device according to any combination of Examples 1B to 21B, wherein the audio device includes one or more speakers.
[0421] Example 23B: An audio device according to any combination of Examples 1B to 22B, wherein the audio device includes an extended reality (XR) headset coupled to one or more speakers.
[0422] Example 24B: The audio device according to Example 23B, wherein the XR headset includes one or more of augmented reality headsets, virtual reality headsets, or mixed reality headsets.
[0423] Example 25B: An audio device according to any combination of Examples 1B to 24B, wherein the audio device includes one or more speakers configured to generate a sound field.
[0424] Example 26B: An audio device according to any combination of Examples 1B to 25B, wherein the one or more source devices include a plurality of audio receivers.
[0425] Example 27B: An audio device according to any combination of Examples 1B to 26B, wherein the audio device includes a transmitter configured to transmit data according to a 5G cellular standard.
[0426] Example 28B: An audio device according to any combination of Examples 1B to 27B, wherein the audio device includes a transmitter configured to transmit data according to a personal area network standard.
[0427] Example 29B: An audio device according to any combination of Examples 1B to 28B, wherein one or more processors are configured to transmit at least one of the following via a wireless link: the audio stream and the at least one energy map.
[0428] Example 30B: An audio device according to Example 29B, wherein the wireless link is via a 5G air interface.
[0429] Example 31B: An audio device according to Example 29B, wherein the wireless link is via a Bluetooth interface.
[0430] Example 32B: An audio device according to any combination of Examples 1B to 31B, wherein the audio device includes a remote server configured to determine the at least one energy map.
[0431] Example 33B: A method for configuring an audio device configured to adjust audio capture, the method comprising: sending an audio stream to one or more source devices; determining instructions for adjusting parameter settings of the audio device; and adjusting the parameter settings to adjust the sound field.
[0432] Example 34B: According to the method of Example 33B, the method further includes: sending diagnostic data to the one or more source devices.
[0433] Example 35B: The method according to any combination of Examples 33B and 34B, the method further comprising: performing energy analysis on the audio stream to determine at least one energy map; and sending the energy map to the one or more source devices.
[0434] Example 36B: The method according to any combination of Examples 33B to 35B, wherein the parameter settings are configured to modify the capture of the audio stream.
[0435] Example 37B: The method according to any combination of Examples 33B to 36B, wherein the parameter settings include adjusting the frequency-dependent gain of the audio device.
[0436] Example 38B: The method according to any combination of Examples 33B to 37B, the method further comprising: rendering an energy curve overlay based at least in part on the at least one energy map.
[0437] Example 39B: The method according to Example 38B, wherein the energy graph overlay is rendered based at least in part on a composite energy graph, the composite energy graph being at least in part based on the one or more energy graphs.
[0438] Example 40B: The method according to any combination of Examples 38B and 39B, the method further comprising: outputting the energy curve overlay for display to the user.
[0439] Example 41B: The method according to any combination of Examples 33B to 40B, the method further comprising: determining the quality characteristics of the audio stream; and sending the quality characteristics to the one or more source devices.
[0440] Example 42B: The method according to any combination of Examples 33B to 41B, the method further comprising: sending at least one of the following to the one or more source devices: a license status and a feasibility status.
[0441] Example 43B: The method according to any combination of Examples 33B to 42B, the method further comprising: receiving the instruction from the one or more source devices.
[0442] Example 44B: The method according to any combination of Examples 33B to 43B, the method further comprising: receiving a composite energy map from the one or more source devices.
[0443] Example 45B: The method according to any combination of Examples 33B to 44B, wherein the instructions are at least partially based on a composite energy map.
[0444] Example 46B: The method according to any combination of Examples 33B to 45B, the method further comprising: receiving diagnostic data from the one or more source devices, wherein the diagnostic data includes at least one of: signal-to-noise ratio information and sound level information.
[0445] Example 47B: The method according to any combination of Examples 33B to 46B, the method further comprising: sending an adjustment status.
[0446] Example 48B: The method described in Example 47B, wherein the adjustment state indicates that the audio device is receiving audio.
[0447] Example 49B: The method according to any combination of Examples 33B to 48B, wherein at least one audio device performs the method, the at least one audio device including a wearable device.
[0448] Example 50B: The method according to any combination of Examples 33B to 49B, wherein at least one audio device performs the method, the at least one audio device including a mobile device.
[0449] Example 51B: The method according to any combination of Examples 33B to 50B, wherein at least one audio device performs the method, the at least one audio device including headphones configured to be coupled to the one or more speakers.
[0450] Example 52B: The method according to any combination of Examples 33B to 51B, wherein at least one audio device performs the method, the at least one audio device comprising one or more speakers.
[0451] Example 53B: The method according to any combination of Examples 33B to 52B, wherein at least one audio device performs the method, the at least one audio device including an extended reality (XR) headset configured to be coupled to the one or more speakers.
[0452] Example 54B: According to the method of Example 53B, the XR headset includes one or more of augmented reality headsets, virtual reality headsets, or mixed reality headsets.
[0453] Example 55B: The method according to any combination of Examples 33B to 54B, wherein at least one audio device performs the method, the at least one audio device including one or more speakers configured to generate a sound field.
[0454] Example 56B: The method according to any combination of Examples 33B to 55B, wherein at least one audio device performs the method, the at least one audio device comprising a plurality of audio receivers.
[0455] Example 57B: The method according to any combination of Examples 33B to 56B, wherein at least one audio device performs the method, the at least one audio device including the at least one microphone.
[0456] Example 58B: The method according to any combination of Examples 33B to 57B, wherein at least one audio device performs the method, the at least one audio device including headphones configured to be coupled to the one or more speakers.
[0457] Example 59B: The method according to any combination of Examples 33B to 58B, wherein at least one audio device performs the method, the at least one audio device comprising one or more speakers.
[0458] Example 60B: The method according to any combination of Examples 33B to 59B, wherein at least one audio device performs the method, the at least one audio device comprising a plurality of audio receivers.
[0459] Example 61B: The method according to any combination of Examples 33B to 60B, wherein at least one audio device performs the method, the at least one audio device including a transmitter configured to transmit data according to a 5G cellular standard.
[0460] Example 62B: The method according to any combination of Examples 33B to 61B, wherein at least one audio device performs the method, the at least one audio device including a transmitter configured to transmit data according to a personal area network standard.
[0461] Example 63B: The method according to any combination of Examples 33B to 62B, wherein at least one audio device performs the method, the at least one audio device being configured to transmit at least one of the following via a wireless link: the audio stream and the at least one energy map.
[0462] Example 64B: The method described in Example 63B, wherein the wireless link is via a 5G air interface.
[0463] Example 65B: The method described in Example 63B, wherein the wireless link is via a Bluetooth interface.
[0464] Example 66B: The method according to any combination of Examples 33B to 65B, wherein at least one audio device performs the method, the at least one audio device including a remote server configured to determine the at least one energy map.
[0465] Example 67B: An audio device configured to generate a sound field, the audio device comprising: components for sending an audio stream to one or more source devices; components for determining instructions for adjusting parameter settings of the audio device; and components for adjusting the parameter settings to adjust the generation of the sound field.
[0466] Example 68B: The audio device according to Example 67B further includes: a component for sending diagnostic data to the one or more source devices.
[0467] Example 69B: The audio device according to any combination of Examples 67B and 68B, the audio device further comprising: components for performing energy analysis on the audio stream to determine an energy map; and components for transmitting the energy map to the one or more source devices.
[0468] Example 70B: An audio device according to any combination of Examples 67B to 69B, wherein the parameter settings are configured to modify the capture of the audio stream.
[0469] Example 71B: An audio device according to any combination of Examples 67B to 70B, wherein the parameter settings include adjustment of the frequency-dependent gain of the audio device.
[0470] Example 72B: The audio device according to any combination of Examples 67B to 71B, the audio device further comprising: a component for rendering an energy curve overlay based at least in part on one or more energy maps.
[0471] Example 73B: The audio device according to Example 72B further includes: a component for rendering the energy graph overlay based at least in part on a composite energy graph, the composite energy graph being at least in part based on the one or more energy graphs.
[0472] Example 74B: The audio device according to any combination of Examples 72B and 73B, the audio device further includes: a component for outputting the energy curve overlay for display to a user.
[0473] Example 75B: The audio device according to any combination of Examples 67B to 74B, the audio device further comprising: components for determining quality characteristics of the audio stream; and components for transmitting the quality characteristics to the one or more source devices.
[0474] Example 76B: The audio device according to any combination of Examples 67B to 75B further includes: a component for sending at least one of the following to the one or more source devices: a license status and a feasibility status.
[0475] Example 77B: The audio device according to any combination of Examples 67B to 76B, the audio device further comprising: a component for receiving the instructions from the one or more source devices.
[0476] Example 78B: The audio device according to any combination of Examples 67B to 58B, the audio device further comprising: a component for receiving a composite energy map from the one or more source devices.
[0477] Example 79B: An audio device according to any combination of Examples 67B to 78B, wherein the instructions are at least partially based on a composite energy map.
[0478] Example 80B: The audio device according to any combination of Examples 67B to 79B, the audio device further comprising: a component for receiving diagnostic data from the one or more source devices, wherein the diagnostic data includes at least one of: signal-to-noise ratio information and sound level information.
[0479] Example 81B: The audio device according to any combination of Examples 67B to 80B, the audio device further includes: a component for transmitting an adjustment state.
[0480] Example 82B: An audio device according to any combination of Example 81B, wherein the adjustment state indicates that the audio device is receiving audio.
[0481] Example 83B: An audio device according to any combination of Examples 67B to 82B, wherein the audio device includes a wearable device.
[0482] Example 84B: An audio device according to any combination of Examples 67B to 83B, wherein the audio device includes a mobile device.
[0483] Example 85B: An audio device according to any combination of Examples 67B to 84B, wherein the audio device includes headphones coupled to one or more speakers.
[0484] Example 86B: An audio device according to any combination of Examples 67B to 85B, wherein the audio device includes one or more speakers.
[0485] Example 87B: An audio device according to any combination of Examples 67B to 86B, wherein the audio device includes an extended reality (XR) headset coupled to one or more speakers.
[0486] Example 88B: The audio device according to Example 87B, wherein the XR headset includes one or more of augmented reality headsets, virtual reality headsets, or mixed reality headsets.
[0487] Example 89B: An audio device according to any combination of Examples 67B to 88B, wherein the audio device includes components for generating a sound field.
[0488] Example 90B: An audio device according to any combination of Examples 67B to 89B, wherein the audio device includes a plurality of audio receivers.
[0489] Example 91B: An audio device according to any combination of Examples 67B to 90B, wherein the audio device includes a plurality of audio receivers.
[0490] Example 92B: An audio device according to any combination of Examples 67B to 90B, wherein the audio device includes at least one microphone.
[0491] Example 93B: An audio device according to any combination of Examples 67B to 92B, wherein the audio device includes headphones coupled to one or more speakers.
[0492] Example 94B: An audio device according to any combination of Examples 67B to 93B, wherein the audio device includes one or more speakers.
[0493] Example 95B: An audio device according to any combination of Examples 67B to 94B, wherein the audio device includes a plurality of audio receivers.
[0494] Example 96B: An audio device according to any combination of Examples 67B to 95B, the audio device including components for transmitting data according to a 5G cellular standard.
[0495] Example 97B: An audio device according to any combination of Examples 67B to 96B, wherein the audio device includes components for transmitting data according to a personal area network standard.
[0496] Example 98B: An audio device according to any combination of Examples 67B to 97B, wherein the at least one audio device includes: a component for transmitting at least one of the following via a wireless link: the audio stream and the at least one energy map.
[0497] Example 99B: An audio device according to Example 98B, wherein the wireless link is via a 5G air interface.
[0498] Example 100B: An audio device according to Example 98B, wherein the wireless link is via a Bluetooth interface.
[0499] Example 101B: An audio device according to any combination of Examples 67B to 100B, wherein the audio device includes a remote server having components for determining at least one energy map.
[0500] Example 102B: A non-transitory computer-readable storage medium having instructions thereon that, when executed, cause one or more processors of an audio device to: send an audio stream to one or more source devices; determine instructions for adjusting parameter settings of the audio device; and adjust the parameter settings to adjust the generation of a sound field.
[0501] It should be noted that the methods described herein depict possible embodiments, and the operations and steps can be rearranged or otherwise modified, and other embodiments are possible. Furthermore, aspects from two or more methods can be combined.
[0502] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including any medium that facilitates, for example, the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium (such as a signal or carrier wave). A data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0503] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage media or other magnetic storage devices, flash memory, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, optical fiber cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the definition of medium includes coaxial cable, optical fiber cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient, tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically generate data magnetically, and optical discs generate data optically by means of lasers. The above combinations should also be included within the scope of computer-readable media.
[0504] Instructions can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuit systems. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding or incorporated into combined codecs. Moreover, the techniques can be fully implemented in one or more circuit or logic elements.
[0505] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. More precisely, as described above, various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware.
[0506] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A device configured to determine parameter adjustments for audio capture, the device comprising: A memory configured to store at least one energy map corresponding to one or more audio streams; as well as One or more processors, said one or more processors being coupled to the memory and configured to: Access the at least one energy map corresponding to the one or more audio streams; The parameter adjustments for at least one audio element are determined, at least in part, based on the at least one energy map, and the parameter adjustments are configured to adjust the audio capture through the at least one audio element. Specifically, the at least one energy map is compared with one or more other energy maps, wherein the at least one energy map and the one or more other energy maps correspond to different audio elements; as well as The parameter adjustment is determined at least in part based on a comparison between the at least one energy map and the one or more other energy maps; and The output parameters are adjusted.
2. The device according to claim 1, wherein, The one or more processors are configured to: Perform energy analysis on the one or more audio streams to determine the at least one energy map.
3. The device according to claim 1, wherein, The one or more processors are configured to: Audio is received by adjusting one or more parameter settings of the at least one audio element according to the parameters.
4. The device according to claim 1, wherein, The one or more processors are configured to: The parameter adjustment is sent to a first source device corresponding to the at least one audio element.
5. The device according to claim 1, wherein, The one or more processors are configured to: Determine the quality characteristics of the one or more audio streams; The following at least one of the following is modified, at least in part, based on the quality characteristics: the at least one energy map and the one or more audio streams; as well as The parameter adjustment is determined at least in part based on the modification.
6. The device according to claim 1, wherein, The one or more processors are configured to: Determine the license status corresponding to at least one of the one or more audio streams; The following at least one of the following may be modified, at least in part, based on the licensed state: the at least one energy map and the one or more audio streams; as well as The parameter adjustment is determined at least in part based on the modification.
7. The device according to claim 1, wherein, The one or more processors are configured to: Determine the feasibility status of the one or more audio elements, wherein the feasibility status indicates the feasibility score of the one or more audio elements; The following at least one of the following may be modified, at least in part, based on the feasibility status: the at least one energy map and the one or more audio streams; as well as The parameter adjustment is determined at least in part based on the modification.
8. The device according to claim 1, wherein, The device includes one or more speakers.
9. The device according to claim 1, wherein, The device includes an augmented reality (XR) headset.
10. The device according to claim 1, wherein, The device includes the at least one audio element, wherein the at least one audio element is configured to receive audio.
11. The device according to claim 1, wherein, The at least one audio element includes at least one microphone, which is configured to receive the one or more audio streams.
12. The device according to claim 1, wherein, The one or more processors are configured to: Receive at least one of the following via a wireless link: one or more audio streams and at least one energy map.
13. The device according to claim 1, wherein, The device includes a remote server configured to determine the at least one energy map.
14. A method for determining parameter adjustments for audio capture, the method comprising: Access at least one energy graph, the at least one energy graph corresponding to one or more audio streams; The parameter adjustments for at least one audio element are determined, at least in part, based on the at least one energy map, and the parameter adjustments are configured to adjust the audio capture through the at least one audio element. Specifically, the at least one energy map is compared with one or more other energy maps. The at least one energy map and the one or more other energy maps correspond to different audio elements; and The parameter adjustment is determined at least in part based on a comparison between the at least one energy map and the one or more other energy maps; and The output indicates an indication of the parameter adjustment for the at least one audio element.
15. The method according to claim 14, further comprising: Perform energy analysis on the one or more audio streams to determine the at least one energy map.
16. The method of claim 14, further comprising: Receive at least one of the following from one or more source devices: the at least one energy map and the one or more other energy maps.
17. The method of claim 14, further comprising: In determining the parameter adjustments, at least one of the following is analyzed: the gain and frequency response of the at least one audio element.
18. The method according to claim 14, wherein, The at least one audio element includes a microphone, and the parameter adjustment includes adjusting the gain of the microphone.
19. The method of claim 14, wherein the method comprises: Audio is received by adjusting one or more parameter settings of the at least one audio element according to the parameters.
20. The method of claim 14, further comprising: The parameter adjustment is sent to a first source device corresponding to the at least one audio element.
21. The method according to claim 14, wherein, Determining the parameter adjustment includes: Determine the difference scores for the one or more audio streams.
22. The method according to claim 21, further comprising: The difference score is compared with the difference threshold; as well as The parameter adjustment is determined at least in part based on the comparison between the difference score and the difference threshold.
23. The method according to claim 14, further comprising: Determine the quality characteristics of the one or more audio streams; The following at least one of the following is modified, at least in part, based on the quality characteristics: the at least one energy map and the one or more audio streams; as well as The parameter adjustment is determined at least in part based on the modification.
24. The method according to claim 14, further comprising: Determine the license status corresponding to at least one of the one or more audio streams; The following at least one of the following may be modified, at least in part, based on the licensed state: the at least one energy map and the one or more audio streams; as well as The parameter adjustment is determined at least in part based on the modification.
25. The method according to claim 14, further comprising: Determine the feasibility status of the at least one audio element, wherein the feasibility status indicates the feasibility score of the at least one audio element; The following at least one of the following may be modified, at least in part, based on the feasibility status: the at least one energy map and the one or more audio streams; as well as The parameter adjustment is determined at least in part based on the modification.
26. The method according to claim 14, further comprising: Receive at least one of the following via a wireless link: one or more audio streams and at least one energy map.
27. A device configured to adjust audio capture, the device comprising components for performing the method according to any one of claims 14-26.
28. A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any one of claims 14-26.
29. A computer program product comprising computer instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 14-26.
Citation Information
Patent Citations
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US10405126B2
Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems
US20190007781A1
Audio processing apparatus and method thereof to provide hearing protection
US20100014682A1