Sound Field Adaptation for Virtual Reality Audio

By using the rotation information of the motion sensor to rotate the spatial component and reconstructing the three-dimensional acoustic signal from the rotating spatial component and audio source, the problems of computing resource consumption and power density in the prior art are solved, and a more efficient audio playback system is realized.

CN114731483BActive Publication Date: 2025-06-27QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080078575.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-18
Filing Date
2020-11-19
Publication Date
2025-06-27
Estimated Expiration
2040-11-19

AI Technical Summary

Technical Problem

Existing psychoacoustic decoders have difficulty rotating spatial components and audio objects separately in the ambient stereo domain, resulting in problems of computing resource consumption and power-intensive.

Method used

By using the rotation information obtained by the motion sensor, the spatial component is rotated to form the rotational spatial component, and the three-dimensional acoustic signal is reconstructed from the rotational spatial component and the audio source, using the spatial characteristics represented by the spherical harmonic function domain.

Benefits of technology

It reduces the demand for computing resources, reduces the cost of information encoding, improves encoding quality, and realizes a more efficient audio playback system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114731483B_ABST
    Figure CN114731483B_ABST
Patent Text Reader

Abstract

An example apparatus includes a memory configured to store at least one spatial component and at least one audio source within a plurality of audio streams. The apparatus further includes one or more processors coupled to the memory. The one or more processors are configured to receive rotation information from a motion sensor. The one or more processors are configured to rotate at least one spatial component based on the rotation information to form at least one rotated spatial component. The one or more processors are further configured to reconstruct an ambient stereo signal from the at least one rotated spatial component and the at least one audio source, wherein the at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Application No. 16 / 951,662, filed on Nov. 18, 2020, which claims the benefit of U.S. Provisional Application No. 62 / 939,477, filed on Nov. 22, 2019, the entire contents of each of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to the processing of media data, such as audio data. Background Art

[0003] Computer-mediated reality systems are being developed to allow a computing device to add to or augment, remove or subtract from, or generally modify an existing reality experienced by a user. Computer-mediated reality systems (which may also be referred to as “augmented reality systems” or “XR systems”) can include, by way of example, virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceived success of a computer-mediated reality system generally relates to the ability of such a computer-mediated reality system to provide a realistically immersive experience in terms of both video and audio experiences, where the video and audio experiences are aligned in the manner expected by the user. Although the human visual system is more sensitive than the human auditory system (e.g., with respect to the perceived localization of various objects within a scene), ensuring an adequate auditory experience is an increasingly important factor in ensuring a realistically immersive experience, particularly as video experiences improve to allow better localization of video objects, which enables a user to better identify the source of audio content. Summary of the Invention

[0004] The present disclosure generally relates to the auditory aspects of the user experience of computer-mediated reality systems, including virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. Aspects of the technology can provide for adaptive audio capture and for rendering of an acoustic space for an extended reality system.

[0005] In one example, aspects of the technology relate to an apparatus configured to play one or more audio streams out of a plurality of audio streams, the apparatus including: a memory configured to store at least one spatial component and at least one audio source within the plurality of audio streams; and one or more processors coupled to the memory and configured to: receive rotation information from a motion sensor; rotate the at least one spatial component based on the rotation information to form at least one rotated spatial component; and reconstruct a three-dimensional sound signal from the at least one rotated spatial component and the at least one audio source, wherein the at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation.

[0006] In another example, aspects of the technology relate to a method of playing one or more of a plurality of audio streams, the method including storing, by a memory, at least one spatial component and at least one audio source within the plurality of audio streams; receiving, by one or more processors, rotation information from a motion sensor; rotating, by one or more processors, at least one spatial component based on the rotation information to form at least one rotated spatial component; and reconstructing, by one or more processors, a three-dimensional sound signal from at least one rotated spatial component and at least one audio source, wherein the at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation.

[0007] In another example, aspects of the technology relate to an apparatus configured to play one or more of a plurality of audio streams, the apparatus including: means for storing at least one spatial component and at least one audio source within the plurality of audio streams; means for receiving rotation information from a motion sensor; means for rotating at least one spatial component to form at least one rotated spatial component; and means for reconstructing a three-dimensional sound signal from at least one rotated spatial component and at least one audio source, wherein the at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation.

[0008] In another example, aspects of the technology are directed to a non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: store at least one spatial component and at least one audio source within a plurality of audio streams; receive rotation information from a motion sensor; rotate at least one spatial component based on the rotation information to form at least one rotated spatial component; and reconstruct a three-dimensional sound signal from at least one rotated spatial component and at least one audio source, wherein the at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation.

[0009] Details of one or more examples of the present disclosure are set forth in the following drawings and description. Other features, objects, and advantages of aspects of the technology will be apparent from the description and drawings and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figures 1A - 1C is a diagram of a system that can execute aspects of the technology described in the present disclosure.

[0011] Figure 2 is a diagram of an example of a VR device worn by a user.

[0012] Figure 3 illustrates an example of a wireless communication system 100 that supports an apparatus and method in accordance with aspects of the present disclosure.

[0013] Figure 4 is a block diagram illustrating an example audio playback system in accordance with the techniques described in the present disclosure.

[0014] Figure 5 is a block diagram further illustrating aspects of the techniques of the present disclosure in an example audio playback system.

[0015] Figure 6 is a block diagram further illustrating aspects of the techniques of the present disclosure in an example audio playback system.

[0016] Figure 7 is a block diagram further illustrating aspects of the techniques of the present disclosure in an example audio playback system.

[0017] Figure 8 is a conceptual diagram illustrating an example concert with three or more audio receivers.

[0018] Figure 9 is a flowchart illustrating an example of using rotational information in accordance with the techniques of the present disclosure.

[0019] Figure 10 is a diagram illustrating an example of a wearable device that may operate in accordance with aspects of the techniques described in the present disclosure.

[0020] Figure 11A and Figure 11B are diagrams illustrating other example systems that may perform aspects of the techniques described in the present disclosure.

[0021] Figure 12 is a diagram illustrating Figures 1A - 1C block diagrams of example components of one or more of a source device and a content consumer device shown in an example. DETAILED DESCRIPTION

[0022] Current psychoacoustic decoders may not be able to separately rotate spatial components and audio objects in an ambient stereo domain. Thus, current psychoacoustic decoders may have to perform a domain conversion to a pulse code modulation (PCM) domain and other processing to rotate such components. These operations may be computationally expensive and power intensive.

[0023] According to the techniques of the present disclosure, a psychoacoustic decoder may rotate at least one spatial component based on rotation information from a motion sensor to form at least one rotated spatial component. The psychoacoustic decoder may also construct an ambisonic signal from at least one rotated spatial component and at least one audio source. The at least one spatial component is represented in a spherical harmonic domain to describe spatial characteristics associated with the at least one audio source. In this way, in a VR platform, a previous spatial vector before motion rotation may be used for a multi-channel environment. According to the techniques of the present disclosure, an audio playback system may receive rotation information from a rotation sensor and use the rotation information to create a rotated spatial vector, such as a V-vector, in a spatial vector domain. This may reduce the computational resources required, may reduce the information that otherwise must be encoded in a bitstream, and may improve the encoding quality.

[0024] In some examples, the audio playback system may jointly decode stereo without the encoder sending inter-channel phase information. The joint stereo operation may utilize spatial placement information obtained from a rotation sensor.

[0025] Encoding efficiency may be improved by leveraging rotation information. First, in phase difference quantization, compression efficiency may be improved by using rotation sensor data. This may be achieved by adding phase information to the rotation sensor data. For example, the interaural phase difference (IPD) in the pulse code modulation / modified discrete cosine transform (PCM / MDCT) domain may be input into a residual coupling / decoupling rotator together with the rotation sensor data, and the residual coupling / decoupling rotator may characterize the residual coupling for stereo vector quantization. Second, using rotation information may improve encoding quality because phase quantization bits may be dynamically reallocated to improve encoding quality by relying on the rotation sensor data for residual coupling. According to the techniques of the present disclosure, if rotation information is available at the decoder, residual coupling may be performed without the encoder sending the phase difference.

[0026] There are multiple different ways to represent a sound field. Example formats include channel-based audio formats, object-based audio formats, and scene-based audio formats. Channel-based audio formats refer to 5.1 surround sound formats, 7.1 surround sound formats, 22.2 surround sound formats, or any other channel-based format that locates audio channels at specific positions around a listener to recreate a sound field.

[0027] An object-based audio format can refer to a format in which audio objects, which are typically encoded using pulse code modulation (PCM) and are referred to as PCM audio objects, are specified to represent a sound field. Such audio objects can include information such as metadata that identifies the position of the audio object relative to a listener or other reference points in the sound field, such that when attempting to recreate the sound field, the audio objects can be rendered to one or more speaker channels for playback. The techniques described in this disclosure can be applied to any of the above formats, including scene-based audio formats, channel-based audio formats, object-based audio formats, or any combination thereof.

[0028] A scene-based audio format can include a hierarchical set of elements that define a sound field in three-dimensional space. An example of a hierarchical set of elements is a set of spherical harmonic coefficients (SHC). The following expression shows a description or representation of a sound field using SHC.

[0029]

[0030] This expression shows the pressure p at any point i in the sound field at time t can be uniquely represented by SHC, where. Here, c is the speed of sound (about 343 m / s), is a reference point (or observation point), j n (·) is the spherical Bessel function of order n, and is the spherical harmonic basis function of order n and sub-order m (which can also be referred to as a spherical basis function). It can be recognized that the terms in the brackets are the frequency-domain representation of the signal (i.e., which can be approximated by various time-frequency transforms, such as the discrete Fourier transform (DFT), discrete cosine transform (DCT), or wavelet transform. Other examples of hierarchical sets include sets of wavelet transform coefficients and other sets of coefficients of multi-resolution basis functions.

[0031] SHC can be physically acquired (e.g., recorded) through various microphone array configurations, or alternatively, they can be derived from a channel-based or object-based description of the sound field. SHC (which can also be referred to as ambisonic coefficients) represents scene-based audio, where SHC can be input into an audio encoder to obtain encoded SHC that can facilitate more efficient transmission or storage. For example, a fourth-order representation involving (1 + 4) 2 (25, and thus fourth-order) coefficients can be used.

[0032] As described above, SHC can be derived from microphone recordings using a microphone array. Various examples of how SHC can be physically obtained from a microphone array are described in Poletti, M., "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics", J. Audio Eng. Soc., Vol. 53, No. 11, November 2005, pp. 1004-1025.

[0033] The following equation can illustrate how SHC can be derived from an object-based description. Coefficients for sound fields corresponding to individual audio objects can be expressed as:

[0034]

[0035] where i is is the (second kind) spherical Hankel function of order n, and is the position of the object. The object source energy g(ω) known as a function of frequency (e.g., using time-frequency analysis techniques such as performing a fast Fourier transform on a pulse code modulation - PCM stream) can enable the conversion of each PCM object and the corresponding position into SHC Additionally, it can be shown (since the above is a linear and orthogonal decomposition) that the coefficients for each object are additive. In this way, multiple PCM objects can be represented by coefficients (e.g., as the sum of coefficient vectors for individual objects). The coefficients can include information about the sound field (pressure as a function of 3D coordinates), and the above represents a transformation from individual objects to a representation of the overall sound field near the observation point nearby.

[0036] Computer-mediated reality systems are being developed (which can also be referred to as "extended reality systems" or "XR systems") to take advantage of the many possible benefits provided by ambisonic coefficients. For example, ambisonic coefficients can represent the sound field in three dimensions in a way that potentially enables precise three-dimensional (3D) localization of audio sources within the sound field. Thus, an XR device can feed rendered ambisonic coefficients to speakers, which, when played via one or more speakers, accurately reproduce the sound field.

[0037] As another example, the ambisonic coefficients can be transformed (e.g., rotated) to account for user movement without overly complex mathematics, thereby potentially meeting the low-latency requirements of XR. Additionally, the ambisonic coefficients are hierarchical, thereby naturally accommodating scalability by downmixing (which can eliminate the ambisonic coefficients associated with higher orders), thereby potentially enabling dynamic adaptation of the sound field to meet the latency and / or battery requirements of the XR device.

[0038] The use of ambisonic coefficients for XR can enable the development of multiple use cases that rely on a more immersive sound field provided by the ambisonic coefficients, particularly for computer game applications and live video streaming applications. In these highly dynamic use cases that rely on low-latency reproduction of the sound field, the XR device may prefer ambisonic coefficients over other representations that are more difficult to manipulate or involve complex rendering. More information regarding these use cases is provided below with respect to Figures 1A - 1C Provided.

[0039] Although described in the context of VR devices in this disclosure, various aspects of the technology can be implemented in the context of other devices, such as mobile devices. In this case, a mobile device (such as a so-called smartphone) can render the world shown on a screen, which can be mounted to the head of user 102 or viewed as in normal use of the mobile device. Thus, any information on the screen is part of the mobile device. The mobile device is capable of providing tracking information 41, thereby allowing both a VR experience (when head-mounted) and a normal experience to view the world shown, where the normal experience can still allow the user to view the world shown, demonstrating a VR-lite-type experience (e.g., lifting the device and rotating or translating the device to view different parts of the world shown). Additionally, although the world shown is mentioned in various examples of this disclosure, the techniques of this disclosure can also be used for acoustic spaces that do not correspond to or in which there is no world shown.

[0040] Figures 1A - 1C is a diagram of a system that can perform various aspects of the techniques described in this disclosure. As Figure 1A shown in the example of, system 10 includes a source device 12 and a content consumer device 14. Although described in the context of source device 12 and content consumer device 14, the technology can be implemented in any context in which any representation of an encoded sound field is formed into a bitstream representation of audio data. Additionally, source device 12 can represent any form of computing device capable of generating a representation of a sound field, and is generally described herein in the context of being a VR content creator device. Similarly, content consumer device 14 can represent any form of computing device capable of implementing the rendering techniques and audio playback described in this disclosure, and is generally described herein in the context of being a VR client device.

[0041] The source device 12 can be operated by an entertainment company or other entity that can generate multi-channel audio content for consumption by an operator of a content consumer device, such as the content consumer device 14. In some VR scenarios, the source device 12 generates audio content in combination with video content. The source device 12 includes a content capture device 20, a content editing device 22, and a sound field representation generator 24. The content capture device 20 can be configured to interface or otherwise communicate with the microphone 18.

[0042] The microphone 18 can represent a 3D audio microphone or other type that can capture and represent a sound field as audio data 19. The audio data 19 can refer to one or more of the above-mentioned scene-based audio data (such as ambisonic coefficients), object-based audio data, and channel-based audio data. Although described as a 3D audio microphone, the microphone 18 can also represent other types of microphones configured to capture the audio data 19 (such as omnidirectional microphones, point microphones, unidirectional microphones, etc.).

[0043] In some examples, the content capture device 20 can include an integrated microphone 18 integrated into the housing of the content capture device 20. The content capture device 20 can interface with the microphone 18 wirelessly or via a wired connection. Instead of or in combination with capturing audio data via the microphone 18, after inputting the audio data 19 via wireless and / or wired input processing by certain types of removable storage devices, the content capture device 20 can process the audio data 19. Thus, according to the present disclosure, different combinations of the content capture device 20 and the microphone 18 are possible.

[0044] The content capture device 20 can also be configured to interface or otherwise communicate with the content editing device 22. In some cases, the content capture device 20 can include the content editing device 22 (in some cases, this can represent software or a combination of software and hardware, including software executed by the content capture device 20 to configure the content capture device 20 to perform a specific form of content editing). The content editing device 22 can represent a unit configured to edit or otherwise change the content 21 received from the content capture device 20, including the audio data 19. The content editing device 22 can output the edited content 23 and associated audio information 25 (such as metadata) to the sound field representation generator 24.

[0045] The sound field representation generator 24 can include any type of hardware device capable of interfacing with the content editing device 22 (or the content capture device 20). Although in Figure 1AThis is not shown in the example, but the sound field representation generator 24 may use the edited content 23 including the audio data 19 and the audio information 25 provided by the content editing device 22 to generate one or more bitstreams 27. In the example focusing on the audio data 19 Figure 1A In the example, the sound field representation generator 24 may generate one or more representations of the same sound field represented by the audio data 19 to obtain the bitstream 27 including the representation of the edited content 23 and the audio information 25.

[0046] For example, to generate different representations of the sound field using the ambisonic coefficients (which is again an example of the audio data 19), the sound field representation generator 24 may use an encoding scheme for the ambisonic representation of the sound field, called mixed-order ambisonics (MOA), as detailed in U.S. Application No. 15 / 672,058, filed on August 8, 2017, and titled "MIXED-ORDER AMBISONICS (MOA) AUDIO DATA FOR COMPUTER-MEDIATED REALITY SYSTEMS", and published as U.S. Patent Publication No. 20190007781 on January 3, 2019.

[0047] To generate a specific MOA representation of the sound field, the sound field representation generator 24 may generate a partial subset of the complete set of ambisonic coefficients. For example, each MOA representation generated by the sound field representation generator 24 may provide a certain level of accuracy with respect to some regions of the sound field, but less accuracy in other regions. In one example, an MOA representation of the sound field may include eight (8) uncompressed ambisonic coefficients, while a third-order ambisonic representation of the same sound field may include sixteen (16) uncompressed ambisonic coefficients. Thus, each MOA representation of the sound field generated as a partial subset of the ambisonic coefficients may be less storage-intensive and less bandwidth-intensive than the corresponding third-order ambisonic representation of the same sound field generated from the ambisonic coefficients (if and when transmitted over the illustrated transmission channel as part of the bitstream 27).

[0048] Although described with respect to the MOA representation, the techniques of the present disclosure may also be performed with respect to the first-order ambisonics (FOA) representation, where all ambisonic coefficients associated with the first-order spherical basis functions and the zero-order spherical basis functions are used to represent the sound field. In other words, instead of using a non-zero subset of the partial ambisonic coefficients to represent the sound field, the sound field representation generator 24 may use all ambisonic coefficients of a given order N to represent the sound field, resulting in a total of (N + 1) 2 ambisonic coefficients.

[0049] In this regard, ambisonic audio data (which is another way of referring to ambisonic coefficients in terms of MOA representation or full order representation, such as the first order representation mentioned above) can include ambisonic coefficients associated with spherical basis functions of first order or less order (which can be referred to as "first order ambisonic audio data"), ambisonic coefficients associated with spherical basis functions having a mixed order and sub - order (which can be referred to as the "MOA representation" discussed above), or ambisonic coefficients associated with spherical basis functions having an order greater than one (which was referred to above as "full order representation").

[0050] In some examples, the sound field representation generator 24 can represent an audio encoder configured to compress or otherwise reduce the number of bits used to represent the content 21 in the bitstream 27. Although not shown, in some examples, the sound field representation generator can include a psychoacoustic audio coding device that complies with any of the various standards discussed herein.

[0051] In this example, the sound field representation generator 24 can apply SVD to the ambisonic coefficients to determine a decomposed version of the ambisonic coefficients. The decomposed version of the ambisonic coefficients can include one or more primary audio signals and one or more corresponding spatial components that describe the spatial characteristics of the associated primary audio signals, such as direction, shape, and width. Thus, the sound field representation generator 24 can apply the decomposition to the ambisonic coefficients to decouple the energy (represented by the primary audio signals) from the spatial characteristics (represented by the spatial components).

[0052] The sound field representation generator 24 can analyze the decomposed version of the ambisonic coefficients to identify various parameters, which can facilitate the re - ordering of the decomposed version of the ambisonic coefficients. The sound field representation generator 24 can re - order the decomposed version of the ambisonic coefficients based on the identified parameters, where it is assumed that the transform can re - order the ambisonic coefficients across frames of the ambisonic coefficients (where a frame typically includes M samples of the decomposed version of the ambisonic coefficients, and in some examples, M is), and this re - ordering can improve the coding efficiency.

[0053] After reordering the decomposed versions of the ambisonic coefficients, the sound field representation generator 24 may select one or more of the decomposed versions of the ambisonic coefficients as the representation of the foreground (or, in other words, different, dominant, or significant) components of the sound field. The sound field representation generator 24 may specify a decomposed version of the ambisonic coefficients that represents the foreground component (which may also be referred to as the "primary sound signal", "primary audio signal", or "primary sound component") and the associated direction information (which may also be referred to as the "spatial component", or in some cases, the so-called "V-vector" that identifies the spatial characteristics of the corresponding audio object). The spatial component may represent a vector having a plurality of different elements (which may be referred to as "coefficients" in terms of a vector), and thus may be referred to as a "multidimensional vector".

[0054] The sound field representation generator 24 may then perform a sound field analysis on the ambisonic coefficients in order to at least partially identify the ambisonic coefficients that represent one or more background (or, in other words, ambient) components of the sound field. The background components may also be referred to as "background audio signals" or "ambient audio signals". Assuming that in some examples, the background audio signal may only include a subset of any given sample of the ambisonic coefficients (e.g., those corresponding to zero-order and first-order spherical basis functions without those corresponding to second-order or higher-order spherical basis functions), the sound field representation generator 24 may perform energy compensation on the background audio signal. When performing downsampling, in other words, the sound field representation generator 24 may enhance the remaining background ambisonic coefficients (e.g., add energy to it / subtract energy from it) to compensate for the change in the total energy caused by performing downsampling.

[0055] The sound field representation generator 24 may then perform a form of interpolation on the foreground direction information (which is another way of referring to the spatial component), and then perform downsampling on the interpolated foreground direction information to generate downsampled foreground direction information. The sound field representation generator 24 may further perform quantization on the downsampled foreground direction information in some examples, outputting encoded foreground direction information. In some cases, this quantization may include scalar / entropy quantization that may be in the form of vector quantization. The sound field representation generator 24 may then output the intermediate-formatted audio data as the background audio signal, the foreground audio signal, and the quantized foreground direction information to the psychoacoustic audio coding device in some examples.

[0056] In any case, in some examples, the background audio signal and the foreground audio signal can include a transmission channel. That is, the sound field representation generator 24 can output each frame of the ambisonic coefficients including the respective background audio signals (e.g., M samples of one of the ambient stereo coefficients corresponding to the zero-order or first-order spherical basis functions) and each frame of the foreground audio signal (e.g., M samples of the audio object decomposed from the ambisonic coefficients) of the transmission channel. The sound field representation generator 24 can further output side information (which can also be referred to as "sideband information"), which includes the quantized spatial components corresponding to each foreground audio signal.

[0057] Collectively, the transmission channel and the side information can be represented as ambisonic transmission format (ATF) audio data (which is another way of referring to intermediate-formatted audio data) in the examples of Figure 1A In other words, the AFT audio data can include the transmission channel and the side information (which can also be referred to as "metadata"). As an example, the ATF audio data can conform to the higher order ambisonics (HOA) transport format (HTF). More information about the HTF can be found in the technical specification (TS) of the European Telecommunications Standards Institute (ETSI) titled "Higher Order Ambisonics (HOA) Transport Format", ETSI TS 103 589 V1.1.1 dated June 2018 (2018-06). Thus, the ATF audio data can be referred to as HTF audio data.

[0058] In an example where the sound field representation generator 24 does not include a psychoacoustic audio coding device, the sound field representation generator 24 can then send or otherwise output the ATF audio data to a psychoacoustic audio coding device (not shown). The psychoacoustic audio coding device can perform psychoacoustic audio coding on the ATF audio data to generate a bitstream 27. The psychoacoustic audio coding device can operate according to a standardized, open-source, or proprietary audio coding process. For example, the psychoacoustic audio coding device can perform psychoacoustic audio coding according to AptX TM 、various other versions of AptX (e.g., enhanced AptX – E-AptX, AptX live, AptX stereo, and AptX high definition – AptX-HD), or advanced audio coding (AAC) and its derivatives. The source device 12 can then send the bitstream 27 to the content consumer device 14 via the transmission channel.

[0059] In some examples, the psychoacoustic audio encoding device may represent one or more instances of a psychoacoustic audio encoder, each for encoding a transmission channel of the ATF audio data. In some cases, the psychoacoustic audio encoding device may represent one or more instances of an AptX encoding unit (as described above). The psychoacoustic audio encoder unit may, in some cases, invoke an instance of the AptX encoding unit for each transmission channel of the ATF audio data.

[0060] In some examples, the content capture device 20 or the content editing device 22 may be configured to communicate wirelessly with the sound field representation generator 24. In some examples, the content capture device 20 or the content editing device 22 may communicate with the sound field representation generator 24 via one or both of a wireless connection and a wired connection. Via the connection between the content capture device 20 and the sound field representation generator 24, the content capture device 20 may provide content in various forms, described herein as part of the audio data 19 for discussion.

[0061] In some examples, the content capture device 20 may utilize aspects of the sound field representation generator 24 (in terms of the hardware or software capabilities of the sound field representation generator 24). For example, the sound field representation generator 24 may include dedicated hardware configured to perform psychoacoustic audio encoding (or dedicated software that causes one or more processors to perform psychoacoustic audio encoding when executed).

[0062] In some examples, the content capture device 20 may not include dedicated hardware or dedicated software for a psychoacoustic audio encoder and instead may provide the audio aspect of the content 21 in a non-psychoacoustic audio encoding form. The sound field representation generator 24 may assist in the capture of the content 21 by performing psychoacoustic audio encoding at least in part with respect to the audio aspect of the content 21.

[0063] The sound field representation generator 24 may also assist in content capture and transmission by generating one or more bitstreams 27 at least in part based on audio content (e.g., MOA representation and / or third-order ambisonic representation) generated from the audio data 19 (in cases where the audio data 19 includes scene-based audio data). The bitstream 27 may represent a compressed version of the audio data 19 and any other different types of content 21 (such as a compressed version of spherical video data, image data, or text data).

[0064] As an example, the sound field representation generator 24 can generate a bitstream 27 for transmission across a transmission channel, a data storage device, etc., and the transmission channel can be a wired or wireless channel. The bitstream 27 can represent an encoded version of the audio data 19 and can include a primary bitstream and a side bitstream, which can be referred to as side channel information or metadata. In some cases, the bitstream 27 representing a compressed version of the audio data 19 (which again can represent scene-based audio data, object-based audio data, channel-based audio data, or a combination thereof) can conform to a bitstream generated according to the MPEG-H 3D audio coding standard and / or the MPEG-I immersive audio standard.

[0065] The content consumer device 14 can be operated by an individual and can represent a VR client device. Although described with respect to a VR client device, the content consumer device 14 can represent other types of devices, such as an augmented reality (AR) client device, a mixed reality (MR) client device (or other XR client device), a standard computer, a head-mounted device, headphones, a mobile device (including so-called smart phones), or any other device capable of tracking the head movement and / or the general translational movement of the individual operating the content consumer device 14. As Figure 1A shown in the example, the content consumer device 14 includes an audio playback system 16A, which can refer to any form of audio playback system capable of rendering audio data for playback as multi-channel audio content.

[0066] Although Figure 1A shown as being sent directly to the content consumer device 14, the source device 12 can output the bitstream 27 to an intermediate device located between the source device 12 and the content consumer device 14. The intermediate device can store the bitstream 27 for subsequent transmission to the content consumer device 14 that can request the bitstream 27. The intermediate device can include a file server, a web server, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a smart phone, or any other device capable of storing the bitstream 27 for subsequent retrieval by an audio decoder. The intermediate device can be located in a content delivery network that can stream the bitstream 27 (and possibly in combination with a corresponding video data bitstream for transmission) to a user requesting the bitstream 27, such as the content consumer device 14.

[0067] Alternatively, the source device 12 can store the bitstream 27 to a storage medium, such as a compact disc, digital video disc, high definition video disc, or other storage medium, most of which can be read by a computer and thus can be referred to as a computer-readable storage medium or a non-transitory computer-readable storage medium. In this context, the transmission channel can refer to the channel through which the content stored to the medium (e.g., in the form of one or more bitstreams 27) is sent (and can include retail stores and other storage-based delivery mechanisms). Thus, in any case, the techniques of the present disclosure should not be limited in this regard to Figure 1A examples of

[0068] As described above, the content consumer device 14 includes an audio playback system 16A. The audio playback system 16A can represent any system capable of playing back multi-channel audio data. The audio playback system 16A can include a plurality of different renderers 32. Each renderer 32 can provide for a different form of rendering, where different forms of rendering can include one or more of various ways of performing vector-based amplitude panning (VBAP) and / or one or more of various ways of performing sound field synthesis. As used herein, "A and / or B" means "A or B", or both "A and B".

[0069] The audio playback system 16A can further include an audio decoding device 34. The audio decoding device 34 can represent a device configured to decode the bitstream 27 to output audio data 19' (where the prime symbol can indicate that the audio data 19' is different from the audio data 19 due to lossy compression, such as quantization). Again, the audio data 19' can include scene-based audio data, which in some examples, can form a full-one (or higher)-order ambisonic representation or a subset of a MOA representation of the same sound field, such as a primary audio signal, its decomposition of surround ambisonic coefficients, and vector-based signals described in the MPEG-H 3D audio coding standard, or other forms of scene-based audio data.

[0070] Other forms of scene-based audio data include audio data defined according to the Higher Order Ambisonics (HOA) Transport Format (HTF). More information about the HTF can be found in the Technical Specification (TS) of the European Telecommunications Standards Institute (ETSI) titled "higherOrder Ambisonics (HOA) Transport Format", ETSI TS 103 589 V1.1.1, dated June 2018 (2018-06), and in U.S. Patent Publication No. 2019 / 0918028 titled "PRIORITY INFORMATION FOR HIGHER ORDER AMBISONIC AUDIO DATA" filed on December 20, 2018. In any case, the audio data 19' can be similar to the entire set or a partial subset of the audio data 19', but may be different due to lossy operations (e.g., quantization) and / or transmission via a transmission channel.

[0071] As an alternative to or in combination with scene-based audio data, the audio data 19' can include channel-based audio data. As an alternative to or in combination with scene-based audio data, the audio data 19' can include object-based audio data. Thus, the audio data 19' can include any combination of scene-based audio data, object-based audio data, and channel-based audio data.

[0072] The audio renderer 32 of the audio playback system 16A can render the audio data 19' to output a speaker feed 35 after the audio decoding device 34 has decoded the bitstream 27 to obtain the audio data 19'. The speaker feed 35 can drive one or more speakers (not shown in the example for illustrative purposes in Figure 1A The various audio representations of the sound field, including scene-based audio data (and possibly channel-based audio data and / or object-based audio data), can be normalized in many ways, including N3D, SN3D, FuMa, N2D, or SN2D.

[0073] To select an appropriate renderer or, in some cases, generate an appropriate renderer, the audio playback system 16A can obtain speaker information 37 indicative of the number of speakers (e.g., loudspeakers or headphone speakers) and / or the spatial geometry of the speakers. In some cases, the audio playback system 16A can use a reference microphone to obtain the speaker information 37 and can drive the speakers (which can refer to the output of an electrical signal to cause a transducer to vibrate) in a manner that dynamically determines the speaker information 37. In other instances, or in combination with the dynamic determination of the speaker information 37, the audio playback system 16A can prompt the user to interface with the audio playback system 16A and input the speaker information 37.

[0074] The audio playback system 16A can select one of the audio renderers 32 based on the speaker information 37. In some cases, when no audio renderer 32 is within certain threshold similarity metrics (in terms of speaker geometry) of the speaker geometry specified in the speaker information 37, the audio playback system 16A can generate one of the audio renderers 32 based on the speaker information 37. In some cases, the audio playback system 16A can generate one of the audio renderers 32 based on the speaker information 37 without first attempting to select an existing one of the audio renderers 32.

[0075] When outputting the speaker feed 35 to headphones, the audio playback system 16A can utilize one of the renderers 32 that uses a head-related transfer function (HRTF) or other functionality capable of rendering the left and right speaker feeds 35 to provide binaural rendering for headphone speaker playback, such as a binaural room impulse response renderer. The term "speaker" or "transducer" can generally refer to any speaker, including loudspeakers, headphone speakers, bone conduction speakers, earbud speakers, wireless headphone speakers, etc. One or more speakers can then play back the rendered speaker feed 35 to reproduce the sound field.

[0076] Although described as rendering the speaker feed 35 from the audio data 19', the rendering of the reference speaker feed 35 can refer to other types of rendering, such as rendering that is directly included in the decoding of the audio data 19 from the bitstream 27. Examples of alternative renderings can be found in Appendix G of the MPEG-H 3D Audio standard, where rendering occurs during the main signal formatting and background signal formation prior to the synthesis of the sound field. Thus, the rendering of the reference audio data 19' should be understood to involve the rendering of the actual audio data 19' or both its decomposition or representation (such as the main audio signal, ambisonic coefficients, and / or vector-based signals - which can also be referred to as V-vectors or multi-dimensional ambisonic spatial vectors) mentioned above.

[0077] The audio playback system 16A can also adapt the audio renderer 32 based on the tracking information 41. That is, the audio playback system 16A can interface with a tracking device 40 configured to track the head movement and possible translational movement of a user of the VR device. The tracking device 40 can represent one or more sensors configured to track the head movement and possible translational movement of a user of the VR device (e.g., cameras - including depth cameras, gyroscopes, magnetometers, accelerometers, light-emitting diodes - LEDs, etc.). The audio playback system 16A can adapt the audio renderer 32 based on the tracking information 41 such that the speaker feeds 35 reflect changes in the user's head and possible translational movement to correctly reproduce the sound field in response to such movement.

[0078] Figure 1C is a block diagram illustrating another example system 60. The example system 60 is similar to Figure 1A example system 10 of, however, the source device 12B of system 60 does not include a content capture device. The source device 12B includes a synthesizing device 29. The synthesizing device 29 can be used by a content developer to generate a synthesized audio source. The synthesized audio source can have location information associated therewith, which can identify the location of the audio source relative to a listener or other reference point in the sound field such that the audio source can be rendered to one or more speaker channels for playback in an effort to recreate the sound field. In some examples, the synthesizing device 29 can also synthesize visual or video data.

[0079] For example, a content developer can generate a synthesized audio stream for a video game. Although the content consumer device 14A of Figure 1A example shows Figure 1C example, however Figure 1C example of the source device 12B can be used with Figure 1B content consumer device 14B of. In some examples, Figure 1C source device 12B of can also include a content capture device such that the bitstream 27 can include both a captured audio stream and a synthesized audio stream.

[0080] As described above, the content consumer device 14A or 14B (for simplicity purposes, both may be referred to hereinafter as the content consumer device 14) can represent a VR device in which a human-wearable display (which may also be referred to as a "head-mounted display") is mounted in front of the eyes of a user operating the VR device. Figure 2FIG. is an example diagram of a VR device 400 worn by a user 402. The VR device 400 is coupled to, or otherwise includes, headphones 404, which can reproduce a sound field represented by audio data 19' through playback of a speaker feed 35. The speaker feed 35 can represent an analog or digital signal capable of causing a diaphragm within a transducer of the headphones 404 to vibrate at various frequencies, where this process is commonly referred to as driving the headphones 404.

[0081] Video, audio, and other sensor data can play important roles in a VR experience. To participate in a VR experience, the user 402 can wear a VR device 400 (which can also be referred to as a VR headset 400) or other wearable electronic device. A VR client device, such as the VR headset 400, can include a tracking device (e.g., tracking device 40), which is configured to track the head movement of the user 402 and adapt the video data presented via the VR headset 400 to account for the head movement, providing an immersive experience in which the user 402 can experience the displayed world shown in the video data in a visual three-dimensional space. The displayed world can refer to a virtual world (where the entire world is simulated), an augmented world (where parts of the world are enhanced by virtual objects), or a physical world (where real-world images are virtually navigated).

[0082] Although VR (and other forms of AR and / or MR) can allow the user 402 to be visually located within a virtual world, typically the VR headset 400 may lack the ability to auditorily place the user within the displayed world. In other words, a VR system (which can include a computer responsible for rendering video data and audio data - not shown in the example for illustrative purposes, and the VR headset 400) may not support fully three-dimensional auditory immersion (and in some cases may not actually reflect the presented scene to the user via the VR headset 400 in a way that matches the displayed scene). Figure 2 Although described in the context of a VR device in this disclosure, various aspects of the technology can be implemented in the context of other devices, such as a mobile device. In this case, a mobile device (such as a so-called smartphone) can present the displayed world via a screen, which can be mounted to the head of the user 402 or viewed as in normal use of the mobile device. Thus, any information on the screen is part of the mobile device. The mobile device is capable of providing tracking information 41, thereby allowing both a VR experience (when head-mounted) and a normal experience of viewing the displayed world, where the normal experience can still allow the user to view the displayed world, demonstrating a VR-lite-type experience (e.g., lifting the device and rotating or translating the device to view different parts of the displayed world).

[0083]

[0084] ​In any case, returning to the VR device context, the audio aspect of VR has been classified into three separate levels of immersion. The first level provides the lowest level of immersion and is called three degrees of freedom (3DOF). 3DOF refers to audio rendering that interprets head movement in three degrees of freedom (yaw, pitch, and roll), thereby allowing the user to freely look around in any direction. However, 3DOF cannot interpret head movement that involves translation where the head is not centered on the optical and acoustic center of the sound field.

[0085] The second level is called 3DOF Plus (3DOF+), which provides three degrees of freedom (yaw, pitch, and roll) in addition to limited spatial translation movement due to head movement away from the optical and acoustic centers within the sound field. 3DOF+ can provide support for perceptual effects such as motion parallax, which can enhance the sense of immersion.

[0086] The third level is called six degrees of freedom (6DOF), which renders audio data in a way that interprets three degrees of freedom of head movement (yaw, pitch, and roll) and also interprets the user's translation in space (x, y, and z translations). The spatial translations can be derived by sensors that track the user's position in the physical world or by means of an input controller.

[0087] 3DOF rendering is the current state of the art in the audio aspect of VR. Thus, the audio aspect of VR is less immersive than the video aspect, potentially reducing the overall immersion of the user experience. However, VR is changing rapidly and can quickly evolve to support both 3DOF+ and 6DOF, which may expose opportunities for additional use cases.

[0088] For example, interactive gaming applications can utilize 6DOF to facilitate fully immersive games where the user moves around within the VR world and can interact with virtual objects by walking up to them. Additionally, interactive live stream applications can utilize 6DOF to allow VR client devices to experience live streams of concerts or sports events as if they were in attendance, allowing the user to move around within the concert or sports event.

[0089] There are multiple difficulties associated with these use cases. In the case of fully immersive games, latency may need to be kept low so that the game play does not cause dizziness or motion sickness. Additionally, from an audio perspective, latency in audio playback that causes loss of synchronization with the video data may reduce immersion. Moreover, for certain types of gaming applications, spatial accuracy may be important to allow for precise responses, including regarding how sound is perceived by the user as it allows the user to anticipate actions that are not currently in view.

[0090] In the context of live streaming applications, a large number of source devices 12A or 12B (for simplicity purposes, both are hereinafter referred to as source device 12) can stream content 21, where the source devices 12 can have very different performances. For example, one source device may be a smart phone with a digital fixed lens camera and one or more microphones, while another source device may be a production-level television device capable of obtaining video with a much higher resolution and quality than a smart phone. However, in the context of live streaming applications, all source devices can provide streams of varying quality, and a VR device can attempt to select an appropriate one from the streams of varying quality to provide a desired experience.

[0091] Figure 3 FIG. illustrates an example of a wireless communication system 100 that supports devices and methods according to aspects of the present disclosure. The wireless communication system 100 includes a base station 105, a UE 115, and a core network 130. In certain examples, the wireless communication system 100 can be a Long Term Evolution (LTE) network, an Advanced LTE (LTE-A) network, an LTE-A Pro network, a Fifth Generation (5G) cellular network, or a New Radio (NR) network. In some cases, the wireless communication system 100 can support enhanced broadband communication, ultra-reliable (e.g., mission-critical) communication, low-latency communication, or communication with low-cost and low-complexity devices.

[0092] The base station 105 can communicate wirelessly with the UE 115 via one or more base station antennas. The base station 105 described herein can include or can be what a person skilled in the art refers to as a base transceiver station, radio base station, access point, radio transceiver, Node B, eNodeB (eNB), next-generation Node B, or giga Node B (both can be referred to as gNB), home Node B, home eNodeB, or some other appropriate term. The wireless communication system 100 can include different types of base stations 105 (e.g., macro or small cell base stations). The UE 115 described herein is capable of communicating with various types of base stations 105 and network devices including macro eNBs, small cell eNBs, gNBs, relay base stations, etc.

[0093] Each base station 105 can be associated with a specific geographic coverage area 110 in which communication with various UEs 115 is supported. Each base station 105 can provide communication coverage for the respective geographic coverage area 110 via a communication link 125, and the communication link 125 between the base station 105 and the UE 115 can utilize one or more carriers. The communication link 125 shown in the wireless communication system 100 can include an uplink transmission from the UE 115 to the base station 105, or a downlink transmission from the base station 105 to the UE 115. The downlink transmission can also be referred to as a forward link transmission, while the uplink transmission can also be referred to as a reverse link transmission.

[0094] The geographical coverage area 110 for the base station 105 can be divided into sectors that form part of the geographical coverage area 110, and each sector can be associated with a cell. For example, each base station 105 can provide communication coverage for macro cells, small cells, hotspots, or other types of cells or various combinations thereof. In some examples, the base station 105 can be movable and thus provide communication coverage for a mobile geographical coverage area 110. In some examples, different geographical coverage areas 110 associated with different technologies can overlap, and the overlapping geographical coverage areas 110 associated with different technologies can be supported by the same base station 105 or by different base stations 105. The wireless communication system 100 can include, for example, different types of LTE / LTE-A / LTE-A Pro, 5G cellular, or NR networks in which different types of base stations 105 provide coverage for various geographical coverage areas 110.

[0095] UEs 115 can be dispersed throughout the wireless communication system 100, and each UE 115 can be stationary or mobile. The UE 115 can also be referred to as a mobile device, wireless device, remote device, handheld device, or user device or some other suitable term, where a "device" can also be referred to as a unit, station, terminal, or client. The UE 115 can also be a personal electronic device, such as a cellular phone, personal digital assistant (PDA), tablet computer, laptop computer, or personal computer. In examples of the present disclosure, the UE 115 can be any audio source described in the present disclosure, including VR headsets, XR headsets, AR headsets, vehicles, smartphones, microphones, arrays of microphones, or any other device including a microphone, or capable of transmitting a captured and / or synthesized audio stream. In some examples, the synthesized audio stream can be an audio stream stored in a memory or previously created or synthesized. In some examples, the UE 115 can also be referred to as a wireless local loop (WLL) station, Internet of Things (IoT) device, Internet of Everything (IoE) device, or machine type communication (MTC) device, etc., which can be implemented in various articles such as instruments, vehicles, meters, etc.

[0096] Certain UEs 115, such as MTC or IoT devices, can be low-cost or low-complexity devices and can provide for automated communication between machines (e.g., via machine-to-machine (M2M) communication). M2M communication or MTC can refer to data communication technologies that allow devices to communicate with each other or with the base station 105 without human intervention. In some examples, M2M communication or MTC can include communication from devices that exchange and / or use audio information, such as metadata, for various audio streams and / or audio sources indicating privacy restrictions and / or password-based privacy data for switching, masking, and / or emptying, as will be described in more detail below.

[0097] In some cases, UE 115 can also communicate directly with other UEs 115 (e.g., using a peer-to-peer (P2P) or device-to-device (D2D) protocol). One or more of a group of UEs 115 utilizing D2D communication can be within the geographical coverage area 110 of base station 105. Other UEs 115 in such a group can be outside the geographical coverage area 110 of base station 105 or otherwise unable to receive transmissions from base station 105. In some cases, a group of UEs 115 communicating via D2D communication can utilize a one-to-many (1:M) system in which each UE 115 in the group sends to each other UE 115 in the group. In some cases, base station 105 facilitates the scheduling of resources for D2D communication. In other cases, D2D communication occurs between UEs 115 without involving base station 105.

[0098] Base stations 105 can communicate with core network 130 and with each other. For example, base stations 105 can interface with core network 130 via a backhaul link 132 (e.g., via S1, N2, N3, or other interfaces). Base stations 105 can communicate with each other directly (e.g., directly between base stations 105) or indirectly (e.g., via core network 130) via a backhaul link 134 (e.g., via X2, Xn, or other interfaces).

[0099] In some cases, wireless communication system 100 can utilize licensed and unlicensed radio frequency bands. For example, in an unlicensed band such as the 5 GHz ISM band, wireless communication system 100 can employ licensed-assisted access (LAA), LTE-unlicensed (LTE-U) radio access technology, 5G cellular technology, or NR technology. When operating in an unlicensed radio frequency spectrum band, wireless devices such as base station 105 and UE 115 can employ a listen-before-talk (LBT) procedure to ensure that the frequency channel is clear before transmitting data. In some cases, operation in the unlicensed band can be based on a carrier aggregation configuration (e.g., LAA) combined with operation in a licensed band. Operation in the unlicensed spectrum can include downlink transmissions, uplink transmissions, peer-to-peer transmissions, or a combination of these. Duplexing in the unlicensed spectrum can be based on frequency-division duplexing (FDD), time-division duplexing (TDD), or a combination of both.

[0100] When, for example Figure 2When the user 402 of the head-mounted device of the VR head-mounted device 400 in [[]] moves their head in the direction of the sound, they may expect to experience the movement of the sound. For example, if the user 402 hears a car leaving from their left, then when the user 402 turns to their left, they may expect to hear the car as if in front of them after having turned to face the sound. To move the sound field, the content consumer device 14 may translate the sound field in the PCM domain. However, translating the sound field in the PCM domain may consume computing resources (such as processing cycles, memory bandwidth, memory, and / or storage space, etc.), because translation in the PCM domain may be computationally complex.

[0101] According to various aspects of the techniques described in the present disclosure, for example, the content consumer device 14, which may be a VR head-mounted device 400, may translate the sound field in the spatial vector domain. By translating the sound field in the spatial vector domain rather than in the PCM domain, computing resources can be saved.

[0102] In operation, the content consumer device 14 may receive rotation information from a motion sensor. The motion sensor may be located, for example, within the head-mounted display. The rotation information may include the roll, pitch, and / or yaw of the user 402's head. The audio playback system 16 of the content consumer device 14 may multiply the rotation information by a spatial vector, such as a V-vector. In this way, the content consumer device 14 can achieve translation of the sound field without the high-cost processing of translating the sound field in the PCM domain.

[0103] After the audio playback system 16 of the content consumer device 14 rotates or performs some form of translation with respect to the spatial vector, the content consumer device 14 may perform ambisonic decoding of the sound field based on the rotated spatial vector and audio data (which may include a U-vector decomposed from the ambient stereo audio data 19). More information regarding various aspects of the translation technique is discussed below with respect to Figure 4 discussion.

[0104] Figure 4 is a block diagram that more specifically illustrates example audio playback systems, such as Figures 1A - 1C the audio playback system 16A or the audio playback system 16B of Figure 4 As shown in the example of

[0105] The spatial vector rotator 205 may represent a unit configured to receive rotation information about the movement of the head of user 402, such as roll, pitch, and / or yaw information, and to generate a rotated spatial vector signal using the rotation information. For example, the spatial vector rotator 205 may rotate the spatial vector signal in the spatial vector domain so that the audio playback system 16 can avoid the high cost translation of the sound field in the PCM domain (in terms of processing cycles, memory space, and / or bandwidth including memory bandwidth).

[0106] The HOA reconstructor 230 may represent Figures 1A - 1C an example of all or part of the audio decoding device 34 shown in the example of. In some examples, the HOA reconstructor 230 may operate as all or part of a high-order ambisonics (HOA) transmission format (HTF) decoder according to the HTF audio standard discussed elsewhere in this disclosure.

[0107] As further shown in the example of Figure 4 the audio playback system 16 may interface with a rotation sensor 200, which may be included within a head-mounted device such as Figure 2 the VR head-mounted device 400 and / or Figures 1A - 1C within the tracking device 40 of. When mounted on the user's head, the rotation sensor 200 may monitor the rotational movement of the user's head. For example, the rotation sensor 200 may measure the pitch, roll, and yaw (theta, phi, and psi) of the head when the user 402 moves their head. The measurement of the rotational movement of the head (rotation information) may be sent to the spatial vector rotator 205. The spatial vector rotator 205 may be part of the audio playback system 16, which may be represented as 16A or 16B in the content consumer device 14 as shown in Figures 1A - 1C respectively.

[0108] The spatial vector rotator 205 may receive the rotation information of the user's head. The spatial vector rotator 205 may also receive from Figures 1A - 1CThe source device 12 receives the spatial vector 220 in a bitstream, such as bitstream 27. The spatial vector rotator 205 can use the rotation information to rotate the spatial vector 220. For example, the spatial vector rotator 205 can rotate the spatial vector 220 by multiplying the spatial vector by the rotation information via a series of left shifts, via a lookup table, via matrix multiplication, row-by-row multiplication, or by accessing an array and multiplying by individual numbers. In this way, the spatial vector rotator 205 can move the sound field to where the user 402 expects it to be. Information on how to create a rotation compensation matrix can be found in Matthias Kronlachner and Franz Zotter's Enhanced Spatial Transform for Ambisonic Recording, and when implemented, the rotation compensation matrix can be used by the spatial vector rotator 205 to rotate the spatial vector 220 via matrix multiplication. Although the audio playback system 16 is described here as moving the sound field to where the user 402 would expect it to be, this is not necessary. For example, the content creator may wish to have more control over the rendering to create specific audio effects or reduce the movement of the sound field due to the micro-movements of the user 402. In these cases, rendering metadata can be added to the bitstream 27 to limit or modify the ability of the spatial vector rotator to rotate the sound field.

[0109] The spatial vector rotator 205 can then provide the rotated spatial vector to the HOA reconstructor 230. The HOA reconstructor 230 can receive a representation of the audio source 225, such as a U-vector, from the bitstream 27 or from other parts of the audio decoding device 34 of the audio playback system 16, from Figures 1A - 1C the source device 12 and reconstruct the rotated HOA signal. The HOA reconstructor 230 can then output the reconstructed HOA signal to be rendered. Although Figure 4 it has been described with respect to HOA signals, it can also be applied to MOA signals and FOA signals.

[0110] Figure 5 is a block diagram of an example audio playback system further illustrating various aspects of the techniques of the present disclosure. Figure 5 can represent Figure 4 a more detailed view, where, for example, a representation of an audio source, such as a U-vector, is reconstructed in the audio decoding device 34 of the audio playback system 16. The audio source or the audio source as used herein can refer to a representation of an audio source, such as a U-vector, or a representation of multiple audio sources, such as multiple U-vectors. As in Figure 4 the audio playback system 16 receives rotation information from the rotation sensor 200. The spatial vector rotator 205 can receive the rotation information and the spatial vector received in the bitstream 27 and, for example, as described above with respect to Figure 4The rotation of the spatial vector is formed in the manner described above. The HOA reconstructor 230 may receive the rotated spatial vector from the spatial vector rotator 205.

[0111] The multi-channel vector dequantizer 232 may receive the quantized reference residual vector signal (REF VQ) and a plurality of quantized side information signals (REF / 2 (not shown)-REF / M) relative to the reference residual vector. In this example, the audio playback system 15 is shown as processing M reference side information signals. M may be any integer. The multi-channel vector dequantizer 232 may dequantize the reference residual vector (REF VQ) and the side information (REF / 2–REF / / M), and provide the dequantized reference residual vector (REF VD) to each of a plurality of residual decouplers 233A-233M. The multi-channel vector dequantizer 232 may also provide the dequantized side information for its respective channels 2-M to each of the plurality of residual decouplers 233B (not shown for simplicity)-233M. For example, the multi-channel vector dequantizer 232 may provide the dequantized side information (REF / MD SIDE) for channel M to the residual decoupler 233M. Each of the residual decouplers 233A-233M may also receive the energy dequantization of the reference residual vector or for its respective channels 2-M. The residual decouplers 233A-233M decouple the residual from the reference residual vector and begin to reconstruct the reference audio source, such as the reference U-vector, and the audio sources for channels 2-M. The even / odd subband synthesizers (E / O SUB) 236A-236M receive the outputs of the residual decouplers 233A-233M, and may separate the even coefficients from the odd coefficients, thereby avoiding phase distortion in the reconstructed audio sources. The gain / shape synthesizers (GAIN / SHAPE SYNTH) 238A-238M may receive the outputs of the even / odd subband synthesizers, and change the gain and / or shape of the signals received by the gain / shape synthesizers 238A-238M, thereby reconstructing one or more reference audio sources for channels 2-M. The HOA reconstructor 230 may receive one or more reference audio sources for channels 2-M, and reconstruct a high-order ambisonic signal based on the received rotated spatial vector and the received audio sources.

[0112] Figure 6 is a block diagram of an example audio playback system further illustrating various aspects of the techniques of the present disclosure. Figure 6 The example of Figure 5 is similar to the example of Figure 4 and Figure 5In this case, the audio playback system 16 receives rotation information from the rotation sensor 200. The spatial vector rotator 205 can, for example, receive rotation information and spatial vectors from the bitstream 27 and form rotated spatial vectors in a manner such as that described above with respect to Figure 4 in the manner described. The HOA reconstructor 230 can receive the rotated spatial vectors from the spatial vector rotator 205.

[0113] The Residual Coupling / Decoupling Rotator (RESID C / D ROTATOR) 240 receives a plurality of side information signals relative to a reference for each of channels 2 - M. The Residual Coupling / Decoupling Rotator 240 can also receive rotation information from the rotation sensor 200 and the rotated spatial vectors from the spatial vector rotator 205. The Residual Coupling / Decoupling Rotator can create a projection matrix for each of the side information of channels 2 - M relative to a reference residual vector and provide the projection matrix for each channel to the associated Projection - Based Residual Decoupler (PROJ - BASED RESID DECOUPLER) 234A - 234M. The projection matrix can be an energy - preserving rotation matrix, which can be used to decouple the reconstructed channels from the reference residual vector. The projection matrix can be created using the Karhunen - Loève Transform (KLT) or Principal Component Analysis (PCA) or other methods.

[0114] The Reference Vector Dequantizer (REF VECTOR DEQUANT) 242 can receive the quantized reference residual vector and dequantize the quantized reference residual vector. The Reference Vector Dequantizer 242 can provide the dequantized reference residual vector to the plurality of Projection - Based Residual Decouplers 234A - 234M. The Reference Vector Dequantizer 242 can also provide the dequantized reference residual vector to the Gain / Shape Synthesizer (GAIN / SHAPE SYNTH) 238R. The Projection - Based Residual Decouplers 234A - 234M decouple the rotated side information from the reference residual vector and output the residual coupling components for channels 2 - M. The Even / Odd Sub - band Synthesizers (E / O SUB) 236A - 236M receive the residual coupling components output by the Projection - Based Residual Decouplers 234A - 234M and separate the even coefficients from the odd coefficients. The Gain / Shape Synthesizers 238A - 238M receive the output of the Even / Odd Sub - band Synthesizers and the dequantized energy signals for channels 2 - M respectively. The Gain / Shape Synthesizers 238A - 238M synthesize the residual coupling components with the dequantized energy components to create the rotated audio sources for channels 2 - M.

[0115] In addition to the quantized reference residual vector, the gain / shape synthesizer (GAIN / SHAPE SYNTH) 238R may also receive the de-quantized energy of the reference residual signal. The gain / shape synthesizer 238R may synthesize the reference residual vector and the de-quantized energy of the reference residual signal to reconstruct and output the reconstructed reference audio source. The gain / shape synthesizers 238A-238M may output the rotated reconstructed audio sources for channels 2-M. The HOA reconstructor 230 may receive the reconstructed reference residual audio sources and the rotated reconstructed audio sources for channels 2-M, and reconstruct the high-order ambisonic signal based on the reconstructed reference audio sources, the rotated reconstructed audio sources, and the rotated spatial vectors for channels 2-M.

[0116] Figure 7 is a block diagram of an example audio playback system further illustrating aspects of the techniques of the present disclosure. Figure 7 may be Figure 6 a more detailed example of the example of, including the energy de-quantization component and the residual component. As in Figures 4 - 6 the audio playback system 16 may receive rotation information from the rotation sensor 200. The HTF decoder 248 may decode the information in the bitstream 27 to obtain the spatial vector. The HTF decoder 248 may provide the spatial vector to the spatial vector rotator (SPAT VECTORROTATOR) 205. The spatial vector rotator 205 may also receive rotation information from the rotation sensor 200. The spatial vector rotator 205 may form the rotated spatial vector in a manner such as described above with respect to Figure 4 The HOA reconstructor 230 may receive the rotated spatial vector from the spatial vector rotator 205.

[0117] The residual coupling / de-coupling rotator (RESID C / D ROT) 240 may also receive rotation information from the rotation sensor 200. The residual side temporal decoder (RESID SIDE TEMPORAL DECODER) 246 may receive side information for channels 2-M with respect to the reference residual vector from the bitstream 27. The residual side temporal decoder 246 may determine, for example via stereo coupling analysis, the temporal phase information for each of channels 2-M, and send the temporal phase information for each of channels 2-M to the residual coupling / de-coupling rotator 240. The residual coupling / de-coupling rotator 240 may create a projection matrix for each of channels 2-M based on the rotation information from the rotation sensor 200 and the temporal phase information from the residual side temporal decoder 246. Thus, Figure 7 the projection matrix in the example of

[0118] The multi-channel energy decoder 244 can receive a multi-channel energy bitstream from the bitstream 27. The multi-channel energy decoder 244 can decode the multi-channel energy bitstream and provide an energy reference signal to the gain / shape synthesizer (GAIN / SHAPE SYNTH) 238R. The multi-channel energy decoder 244 can also provide energy signals for respective channels 2-M to each of the projection-based residual decouplers (PROJ-BASED RESIDDECOUPLER) 234A-M and each of the gain / shape synthesizers (GAIN / SHAPE SYNTH) 238A-238M. The projection-based residual decouplers 234A-234M, even / odd subband separators (E / O SUB) 236A-236M, gain / shape synthesizers 238A-238M and 238R, and the HOA reconstructor 230 can work similarly to Figure 6 the projection-based residual decouplers 234A-234M, even / odd subband separators 236A-236M, gain / shape synthesizers 238A-238M and 238R, and the HOA reconstructor 230 in the example of

[0119] Figure 8 is a conceptual diagram illustrating an example concert with three or more audio receivers. In Figure 8 the example of, a plurality of musicians are shown on the stage 323. The singer 312 is located behind the microphone 310A. The string section 314 is shown behind the microphone 310B. The drummer 316 is shown behind the microphone 310C. The other musicians 318 are shown behind the microphone 310D. The microphones 310A-301D can capture audio streams corresponding to the sounds received by the microphones. In some examples, the microphones 310A-310D can represent synthesized audio streams. For example, the microphone 310A can capture an audio stream mainly associated with the singer 312, but the audio stream can also include sounds produced by other band members, such as the string section 314, the drummer 316, or the other musicians 318, while the microphone 310B can capture an audio stream mainly associated with the string section 314, but includes sounds produced by other band members. In this way, each of the microphones 310A-310D can capture different audio streams.

[0120] A plurality of devices are also shown. These devices represent user devices located at a plurality of different desired listening positions. The headphones 320 are located near the microphone 310A, but between the microphone 310A and the microphone 310B. Thus, according to the techniques of the present disclosure, the content consumer device can select at least one audio stream to produce a similar effect as if the user were located at the headphones 320 in Figure 8The local audio experience for the user of the headphones 320. Similarly, the VR goggles 322 are shown to be located behind the microphone 310C and between the drummer 316 and the other musicians 318. The content consumer device can select at least one audio stream to produce an audio experience similar to that of a user located where the VR goggles 322 are Figure 8 The local audio experience for the user of the VR goggles 322.

[0121] The smart glasses 324 are shown to be located rather centrally between the microphones 310A, 310C, and 310D. The content consumer device can select at least one audio stream to produce an audio experience similar to that of a user located where the smart glasses 324 are Figure 8 The local audio experience for the user of the smart glasses 324. Additionally, the device 326 (which can represent any device capable of implementing the technology of the present disclosure, such as a mobile phone, a speaker array, headphones, VR goggles, smart glasses, etc.) is shown to be located in front of the microphone 310B. The content consumer device can select at least one audio stream to produce an audio experience similar to that of a user located where the device 326 is Figure 8 The local audio experience for the user of the device 326. Although specific devices are discussed in relation to specific locations, the use of any of the devices shown can provide an indication different from the Figure 8 Desired listening position shown in. Figure 8 Any device can be used to implement the technology of the present disclosure.

[0122] Figure 9 Is a flowchart illustrating an example of using rotation information according to the technology of the present disclosure. The audio playback system 16 can store at least one spatial component and at least one audio source (250). For example, the audio playback system can receive multiple audio streams in a bitstream 27. The multiple audio streams can include at least one spatial component and at least one audio component. The audio playback system 16 can store at least one spatial component and at least one audio source in a memory.

[0123] The audio playback system 16 can receive rotation information (252) from a motion sensor such as a rotation sensor 200. For example, the rotation sensor 200 can measure the pitch, roll, and yaw (theta, phi, and psi) of the head when the user 402 moves their head. The measurement of the rotational movement of the head (rotation information) can be received by the audio playback system 16. The audio playback system 15 can rotate at least one spatial component based on the rotation information (254). For example, the spatial vector rotator 205 can rotate at least one spatial component by multiplying at least one spatial component by the rotation information via a series of left shifts, via a lookup table, via matrix multiplication, row-by-row multiplication, or by accessing an array and multiplying by individual numbers.

[0124] The audio playback system 15 can reconstruct an ambisonic signal (256) from at least one spatial component and at least one audio source that rotates. For example, the HOA reconstructor 230 can receive a representation of the audio source 225, such as a U-vector, from the bitstream 27 or from other parts of the audio decoding device 34, from the Figures 1A - 1C source device 12 and reconstruct the rotating HOA signal. In some examples, the at least one spatial component includes a V-vector and the at least one audio source includes a U-vector. In some examples, the audio playback system 15 can apply a projection matrix to the reference residual vector and the dequantized energy signal to reconstruct the U-vector. In some examples, the projection matrix includes temporal and spatial rotation data. For example, Figure 7 the residual coupling / decoupling rotator 240 can create a projection matrix for each of the channels 2-M based on the rotation information from the rotation sensor 200 and the temporal phase information from the residual side temporal decoder 246. In some examples, the audio playback system 15 can output a representation of at least one audio source, such as based on the ambisonic signal, to one or more speakers (258). In some examples, the audio playback system can combine at least two representations of at least one audio source by at least one combination of mixing or interpolation before outputting the ambisonic signal. In some examples, the content consumer device 14 can receive a voice command from a microphone and control the display device based on the voice command. In some examples, the content consumer device 14 can receive a wireless signal, such as a wireless bitstream similar to the bitstream 27.

[0125] Figure 10 is a diagram illustrating an example of a wearable device 500 that can operate in accordance with various aspects of the techniques described in this disclosure. In various examples, the wearable device 500 can represent a VR headset (such as the VR headset 400 described above), an AR headset, an MR headset, or any other type of extended reality ((XR) headset. Augmented reality "AR" can refer to computer-rendered images or data that overlay the real world in which the user is actually located. Mixed reality "MR" can refer to computer-rendered images or data in which the world is locked to a specific location in the real world, or can refer to a variant of VR in which a combination of partially computer-rendered 3D elements and partially captured real elements creates an immersive experience that simulates the user's physical presence in the environment. Extended reality "XR" can represent an umbrella term for VR, AR, and MR. More information about the terminology used for XR can be found in the document "Virtual Reality, Augmented Reality, and Mixed Reality Definitions" by Jason Peterson on July 7, 2017.

[0126] The wearable device 500 may represent other types of devices, such as watches (including so-called "smartwatches"), glasses (including so-called "smart glasses"), headphones (including so-called "wireless headphones" and "smart headphones"), smart clothing, smart jewelry, and so on. Regardless of the representation of the VR device, watch, glasses, and / or headphones, the wearable device 500 may communicate with a computing device that supports the wearable device 500 via a wired connection or a wireless connection.

[0127] In some cases, the computing device that supports the wearable device 500 may be integrated within the wearable device 500. Thus, the wearable device 500 may be considered the same device as the computing device that supports the wearable device 500. In other instances, the wearable device 500 may communicate with a separate computing device that may support the wearable device 500. In this regard, the term "support" should not be construed as requiring a separate dedicated device, but rather should be construed as one or more processors configured to perform various aspects of the techniques described in the present disclosure may be integrated within the wearable device 500 or within a computing device separate from the wearable device 500.

[0128] For example, when the wearable device 500 represents the VR device 1100, a separate dedicated computing device (such as a personal computer including one or more processors) may render audio and visual content, and the wearable device 500 may determine translational head movements, and the dedicated computing device may render audio content (such as speaker feeds) according to various aspects of the techniques described in the present disclosure based on the translational head movements. As another example, when the wearable device 500 represents smart glasses, the wearable device 500 may include one or more processors that determine translational head movements (by interfacing within one or more sensors of the wearable device 500) and render speaker feeds based on the determined translational head movements.

[0129] As shown in the figure, the wearable device 500 includes one or more directional speakers and one or more tracking and / or recording cameras. Additionally, the wearable device 500 includes one or more inertial, tactile, and / or health sensors, one or more eye-tracking cameras, one or more high-sensitivity audio microphones, and optical / projection hardware. The optical / projection hardware of the wearable device 500 may include durable semi-transparent display technology and hardware.

[0130] The wearable device 500 also includes connectivity hardware, which may represent one or more network interfaces that support multi-mode connectivity, such as 4G communication, 5G communication, Bluetooth, Wi-Fi, etc. The wearable device 500 also includes one or more ambient light sensors and bone conduction sensors. In some cases, the wearable device 500 may also include one or more passive and / or active cameras with a fish-eye lens and / or a telephoto lens. Although Figure 10 is not shown, the wearable device 500 may also include one or more light-emitting diode (LED) lights. In certain examples, the LED lights may be referred to as "ultra-bright" LED lights. In certain implementations, the wearable device 500 may also include one or more rear cameras. It will be appreciated that the wearable device 500 may exhibit various different form factors.

[0131] In addition, tracking and recording cameras and other sensors can facilitate the determination of the translation distance. Although not shown in the Figure 10 example, the wearable device 500 may include other types of sensors for detecting the translation distance.

[0132] Although described with respect to specific examples of wearable devices, such as the VR device 1100 discussed above with respect to the Figure 10 example and other devices mentioned in the Figures 1A - 1C example, those skilled in the art will appreciate that the descriptions related to Figures 1A - 1C and Figure 2 can be applied to other examples of wearable devices. For example, other wearable devices such as smart glasses may include sensors through which translational head movements are obtained. As another example, other wearable devices such as smart watches may include sensors through which translational movements are obtained. Thus, the techniques described in this disclosure should not be limited to a specific type of wearable device, but any wearable device can be configured to perform the techniques described in this disclosure.

[0133] Figure 11A and Figure 11B are diagrams of example systems that can perform various aspects of the techniques described in this disclosure. Figure 11A Diagram of an example in which the source device 12 further includes a camera 600. The camera 600 may be configured to capture video data and provide the captured raw video data to the content capture device 20. The content capture device 20 may provide the video data to another component of the source device 12 for further processing into viewpoint-divided portions.

[0134] In Figure 11AIn the example, the content consumer device 14 further includes a wearable device 300. It will be understood that in various implementations, the wearable device 300 may be included within the content consumer device 14 or externally coupled to the content consumer device 14. The wearable device 300 includes display hardware and speaker hardware for outputting video data (e.g., associated with various viewpoints) and for rendering audio data.

[0135] Figure 11B Illustrating where Figure 11A Shown is an example where the audio renderer 32 is replaced with a binaural renderer 42 that is capable of performing binaural rendering using one or more HRTFs or other functions that can be performed on the left and right speaker feeds 43. The audio playback system 16C can output the left and right speaker feeds 43 to the headphones 44.

[0136] The headphones 44 can be coupled to the audio playback system 16C via a wired connection (such as a standard 3.5 mm audio jack, a Universal System Bus (USB) connection, an optical audio jack, or other forms of wired connections) or wirelessly (such as via Bluetooth TM connection, a wireless network connection, etc.). The headphones 44 can recreate the sound field represented by the audio data 19' based on the left and right speaker feeds 43. The headphones 44 can include a left headphone speaker and a right headphone speaker powered (or, in other words, driven) by the respective left and right speaker feeds 43.

[0137] Figure 12 Is an illustration Figures 1A - 1C Of an example block diagram of one or more example components of a source device and a content consumer device shown in the example. In Figure 12 The example, the device 710 includes a processor 712 (which may be referred to as "one or more processors" or "processor"), a Graphics Processing Unit (GPU) 714, a system memory 716, a display processor 718, one or more integrated speakers 740, a display 703, a user interface 720, an antenna 721, and a transceiver module 722. In an example where the device 710 is a mobile device, the display processor 718 is a Mobile Display Processor (MDP). In certain examples, such as where the device 710 is a mobile device, the processor 712, GPU 714, and display processor 718 may be formed as an integrated circuit (IC).

[0138] For example, an IC can be considered a processing chip within a chip package and can be a system-on-chip (SoC). In some examples, two of the processor 712, GPU 714, and display processor 718 can be packaged together in the same IC, and the other in a different integrated circuit (i.e., a different chip package), or all three can be in different ICs or on the same IC. However, in examples where the device 710 is a mobile device, it is possible that the processor 712, GPU 714, and display processor 718 are all in different integrated circuits.

[0139] Examples of the processor 712, GPU 714, and display processor 718 include, but are not limited to, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. The processor 712 can be the central processing unit (CPU) of the device 710. In some examples, the GPU 714 can be a dedicated hardware including integrated and / or discrete logic circuitry that provides a large parallel processing capability suitable for graphics processing to the GPU 714. In some cases, the GPU 714 can also include general processing capabilities and can be referred to as a general-purpose GPU (GPGPU) when implementing general processing tasks (i.e., non-graphics-related tasks). The display processor 718 can also be application-specific integrated circuit hardware designed to retrieve image content from the system memory 716, compose the image content into image frames, and output the image frames to the display 703.

[0140] The processor 712 can execute various types of applications. Examples of applications include a web browser, an email application, a spreadsheet, a video game, other applications that generate viewable objects for display, or any of the application types listed in more detail above. The system memory 716 can store instructions for the execution of the applications. The execution of one of the applications on the processor 712 causes the processor 712 to generate graphic data for the image content to be displayed and audio data 19 to be played (possibly via the integrated speaker 740). The processor 712 can send the graphic data of the image content to the GPU 714 for further processing based on instructions or commands sent by the processor 712 to the GPU 714.

[0141] The processor 712 can communicate with the GPU 714 according to a specific application programming interface (API). Examples of such APIs include API, of the Khronos Group or OpenGL and OpenCL TM; however, aspects of the present disclosure are not limited to the DirectX, OpenGL, or OpenCL APIs and can be extended to other types of APIs. Additionally, the techniques described in the present disclosure do not need to operate according to an API, and the processor 712 and GPU 714 can utilize any processing for communication.

[0142] The system memory 716 can be the memory for the device 710. The system memory 716 can include one or more computer-readable storage media. Examples of the system memory 716 include, but are not limited to, random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other media that can be used to carry or store desired program code in the form of instructions and / or data structures and can be accessed by a computer or a processor.

[0143] In some examples, the system memory 716 can include instructions that cause the processor 712, GPU 714, and / or display processor 718 to perform the functions attributed to the processor 712, GPU 714, and / or display processor 718 in the present disclosure. Thus, the system memory 716 can be a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors (e.g., the processor 712, GPU 714, and / or display processor 718) to perform various functions.

[0144] The system memory 716 can include non-transitory storage media. The term "non-transitory" indicates that the storage media does not specifically embody a carrier wave or a propagated signal. However, the term "non-transitory" should not be construed to mean that the system memory 716 is immovable or that its contents are static. As an example, the system memory 716 can be removed from the device 710 and moved to another device. As another example, a memory substantially similar to the system memory 716 can be inserted into the device 710. In some examples, the non-transitory storage media can store data that may change over time (e.g., in RAM).

[0145] The user interface 720 can represent one or more hardware or virtual (meaning a combination of hardware and software) user interfaces through which a user can interface with the device 710. The user interface 720 can include physical buttons, switches, triggers, lights, or their virtual versions. The user interface 720 can also include a physical or virtual keyboard, a touch interface - such as a touchscreen, haptic feedback, etc.

[0146] The processor 712 may include one or more hardware units (including so-called "processing cores") configured to perform all or some of the operations discussed above with respect to any of the modules, units, or other functional components of the content creator device and / or the content consumer device. The antenna 721 and the transceiver module 722 may represent units configured to establish and maintain a connection between the source device 12 and the content consumer device 14. The antenna 721 and the transceiver module 722 may represent one or more receivers and / or one or more transmitters capable of wireless communication according to one or more wireless communication protocols, such as the fifth generation (5G) cellular standard, Wi-Fi, personal area network (PAN) protocols such as Bluetooth TM or other open source, proprietary, or other communication standards. For example, the transceiver module 722 may receive and / or transmit wireless signals. The transceiver module 722 may represent a separate transmitter, a separate receiver, both a separate transmitter and a separate receiver, or a combined transmitter and receiver. The antenna 721 and the transceiver module 722 may be configured to receive encoded audio data. Similarly, the antenna 721 and the transceiver module 722 may be configured to transmit encoded audio data.

[0147] It should be recognized that depending on the example, certain actions or events of any of the techniques described herein may be performed in a different sequence, may be added, combined, or omitted altogether (e.g., actions or events not all described may not be required for the practice of the technique). Additionally, in certain examples, instead of being performed sequentially, the actions or events may be performed simultaneously, e.g., via multi-threading, interrupt processing, or multiple processors.

[0148] In certain examples, a VR device (or a streaming device) may communicate and exchange messages with an external device using a network interface coupled to the memory of the VR / streaming device, where the exchanged messages are associated with multiple available representations of the sound field. In certain examples, the VR device may receive wireless signals associated with multiple available representations of the sound field, including data packets, audio packets, video protocols, or transport protocol data, using an antenna coupled to the network interface. In certain examples, one or more microphone arrays may capture the sound field.

[0149] In certain examples, the multiple available representations of the sound field stored in the memory device may include multiple object-based representations of the sound field, multiple high-order ambisonic representations of the sound field, multiple hybrid-order ambisonic representations of the sound field, a combination of an object-based representation of the sound field and a high-order ambisonic representation of the sound field, a combination of an object-based representation of the sound field and a hybrid-order ambisonic representation of the sound field, or a combination of a hybrid-order representation of the sound field and a high-order ambisonic representation of the sound field.

[0150] In some examples, one or more sound field representations of the plurality of available representations of the sound field may include at least one high-resolution region and at least one low-resolution region, and wherein the rendering selected based on the steering angle provides greater spatial accuracy with respect to at least one high-resolution region and less spatial accuracy with respect to the low-resolution region.

[0151] The present disclosure includes the following examples.

[0152] Clause 1. An apparatus configured to play one or more audio streams out of a plurality of audio streams, the apparatus comprising: a memory configured to store at least one spatial component and at least one audio source within the plurality of audio streams; and one or more processors coupled to the memory and configured to: receive rotation information from a motion sensor; rotate at least one spatial component based on the rotation information to form at least one rotated spatial component; and construct an ambisonic signal from the at least one rotated spatial component and the at least one audio source.

[0153] Wherein the at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation.

[0154] Clause 1.5. The apparatus of clause 1, wherein the at least one spatial component includes a V-vector and the at least one audio source includes a U-vector.

[0155] Clause 1.6. The apparatus of clause 1.5, wherein the one or more processors are further configured to reconstruct the U-vector.

[0156] Clause 1.7. The apparatus of clause 1.6, wherein the one or more processors are further configured to reconstruct the U-vector by applying a projection matrix to a reference residual vector and a dequantized energy signal.

[0157] Clause 1.8. The apparatus of clause 1.7, wherein the projection matrix includes time and spatial rotation data.

[0158] Clause 2. The apparatus of clause 1, wherein the one or more processors are further configured to output the at least one audio source to one or more speakers.

[0159] Clause 3. The apparatus of any combination of clauses 1-2, wherein the one or more processors are further configured to combine at least two of the at least one audio source.

[0160] Clause 4. The apparatus of clause 3, wherein the one or more processors combine at least two of the at least one audio source by at least one of mixing or interpolation.

[0161] Clause 5. The apparatus of any combination of clauses 1-4, further comprising a display device.

[0162] Clause 6. The apparatus of Clause 5 further includes a microphone, wherein one or more processors are further configured to receive voice commands from the microphone and control the display device based on the voice commands.

[0163] Clause 7. The apparatus of any combination of Clauses 1 - 6 further includes one or more speakers.

[0164] Clause 8. The apparatus of any combination of Clauses 1 - 7, wherein the apparatus includes a mobile phone.

[0165] Clause 9. The apparatus of any combination of Clauses 1 - 7, wherein the apparatus includes an extended reality headset, and wherein the acoustic space includes a scene represented by video data captured by a camera.

[0166] Clause 10. The apparatus of any combination of Clauses 1 - 7, wherein the apparatus includes an extended reality headset, and wherein the acoustic space includes a virtual world.

[0167] Clause 11. The apparatus of any combination of Clauses 1 - 10 further includes a head-mounted device configured to present the acoustic space.

[0168] Clause 12. The apparatus of any combination of Clauses 1 - 11 further includes a wireless transceiver coupled to one or more processors and configured to receive wireless signals.

[0169] Clause 13. The apparatus of Clause 12, wherein the wireless signal complies with a personal area network standard.

[0170] Clause 13.5. The apparatus of Clause 13, wherein the personal area network standard includes the AptX standard.

[0171] Clause 14. The apparatus of Clause 12, wherein the wireless signal complies with a fifth-generation (5G) cellular protocol.

[0172] Clause 15. A method of playing one or more audio streams out of a plurality of audio streams, including: storing, by a memory, at least one spatial component and at least one audio source within the plurality of audio streams; receiving, by one or more processors, rotation information from a motion sensor; rotating, by one or more processors, at least one spatial component based on the rotation information to form at least one rotated spatial component; and constructing, by one or more processors, an ambient stereo signal from at least one rotated spatial component and at least one audio source, wherein at least one spatial component describes spatial characteristics associated with at least one audio source in a spherical harmonic function domain.

[0173] Clause 15.5. The method of Clause 15, wherein at least one spatial component includes a V-vector and at least one audio source includes a U-vector.

[0174] Clause 15.6. The method of Clause 15.5 further includes reconstructing the U-vector.

[0175] Clause 15.7. The method of Clause 15.6, wherein reconstructing the U-vector includes applying a projection matrix to a reference residual vector and a dequantized energy signal.

[0176] Clause 15.8. The apparatus of Clause 15.7, wherein the projection matrix includes temporal and spatial rotation data.

[0177] Clause 16. The method of Clause 15 further includes outputting at least one audio source to one or more speakers by one or more processors.

[0178] Clause 17. The method of any combination of Clauses 15-16 further includes combining at least two of at least one audio source by one or more processors.

[0179] Clause 18. The method of Clause 17, wherein combining at least two of at least one audio source is by at least one of mixing or interpolation.

[0180] Clause 19. The method of any combination of Clauses 15-18 further includes receiving a voice command from a microphone and controlling a display device based on the voice command.

[0181] Clause 20. The method of any combination of Clauses 15-19, wherein the method is executed on a mobile phone.

[0182] Clause 21. The method of any combination of Clauses 15-19, wherein the method is executed on an extended reality headset, and wherein the acoustic space includes a scene represented by video data captured by a camera.

[0183] Clause 22. The method of any combination of Clauses 15-19, wherein the method is executed on an extended reality headset, and wherein the acoustic space includes a virtual world.

[0184] Clause 23. The method of any combination of Clauses 15-22, wherein the method is executed on a head-mounted device configured to present an acoustic space.

[0185] Clause 24. The method of any combination of Clauses 15-23 further includes receiving a wireless signal.

[0186] Clause 25. The method of Clause 24, wherein the wireless signal complies with a personal area network standard.

[0187] Clause 25.5. The method of Clause 25, wherein the personal area network standard includes the AptX standard.

[0188] Clause 26. The method of Clause 24, wherein the wireless signal complies with the fifth-generation (5G) cellular protocol.

[0189] Clause 27. An apparatus configured to play one or more audio streams out of a plurality of audio streams, the apparatus comprising: means for storing at least one spatial component and at least one audio source within the plurality of audio streams; means for receiving rotation information from a motion sensor; means for rotating at least one spatial component to form at least one rotated spatial component; and means for constructing an ambient stereo signal from at least one rotated spatial component and at least one audio source, wherein at least one spatial component describes spatial characteristics associated with at least one audio source in a spherical harmonic function domain.

[0190] Clause 27.5. The apparatus of Clause 27, wherein at least one spatial component comprises a V-vector and at least one audio source comprises a U-vector.

[0191] Clause 27.6. The apparatus of Clause 27.5, further comprising means for reconstructing the U-vector.

[0192] Clause 27.7. The apparatus of Clause 27.6, wherein the means for reconstructing the U-vector applies a projection matrix to a reference residual vector and a dequantized energy signal.

[0193] Clause 27.8. The apparatus of Clause 27.7, wherein the projection matrix comprises time and spatial rotation data.

[0194] Clause 28. The apparatus of Clause 27, further comprising means for outputting at least one audio source to one or more speakers.

[0195] Clause 29. The apparatus of any combination of Clauses 27-28, further comprising means for combining at least two of at least one audio source.

[0196] Clause 30. The apparatus of Clause 29, wherein combining at least two of at least one audio source is by at least one of mixing or interpolation.

[0197] Clause 31. The apparatus of the combination of Clauses 27-30, further comprising means for receiving a voice command from a microphone and means for controlling a display device based on the voice command.

[0198] Clause 32. The apparatus of any combination of Clauses 27-31, wherein the apparatus comprises an extended reality head-mounted device, and wherein the acoustic space comprises a scene represented by video data captured by a camera.

[0199] Clause 33. The apparatus of any combination of Clauses 27-32, wherein the apparatus comprises a mobile phone.

[0200] Clause 34. Apparatus of any combination of Clauses 27 - 32, wherein the apparatus includes an extended reality headset, and wherein the acoustic space includes a virtual world.

[0201] Clause 35. Apparatus of any combination of Clauses 27 - 34, wherein the apparatus includes a head-mounted device configured to present an acoustic space.

[0202] Clause 36. Apparatus of any combination of Clauses 27 - 35, further comprising means for receiving a wireless signal.

[0203] Clause 37. The apparatus of Clause 36, wherein the wireless signal complies with a personal area network standard.

[0204] Clause 37.5. The apparatus of Clause 37, wherein the personal area network standard includes the AptX standard.

[0205] Clause 38. The apparatus of Clause 36, wherein the wireless signal complies with a fifth-generation (5G) cellular protocol.

[0206] Clause 39. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: store at least one spatial component and at least one audio source within a plurality of audio streams;

[0207] receive rotation information from a motion sensor; rotate at least one spatial component based on the rotation information to form at least one rotated spatial component; and construct an ambient stereo signal from at least one rotated spatial component and at least one audio source, wherein at least one spatial component describes spatial characteristics associated with at least one audio source in a spherical harmonic function domain.

[0208] Clause 39.5. The non-transitory computer-readable storage medium of Clause 39, wherein at least one spatial component includes a V-vector and at least one audio source includes a U-vector.

[0209] Clause 39.6. The non-transitory computer-readable storage medium of Clause 39.5, further having instructions stored thereon that, when executed, cause one or more processors to reconstruct the U-vector.

[0210] Clause 39.7. The non-transitory computer-readable storage medium of Clause 39.6, further having instructions stored thereon that, when executed, cause one or more processors to reconstruct the U-vector including by applying a projection matrix to a reference residual vector and a dequantized energy signal.

[0211] Clause 39.8. The non-transitory computer-readable storage medium of Clause 39.7, wherein the projection matrix includes time and spatial rotation data.

[0212] Clause 40. The non-transitory computer-readable storage medium of Clause 39, wherein the instructions, when executed, cause one or more processors to output at least one audio source to one or more speakers.

[0213] Clause 41. The non-transitory computer-readable storage medium of any combination of Clauses 39-40, wherein the instructions, when executed, cause one or more processors to combine at least two of at least one audio source.

[0214] Clause 42. The non-transitory computer-readable storage medium of Clause 41, wherein the instructions, when executed, cause one or more processors to combine at least two of at least one audio source by at least one of mixing or interpolation.

[0215] Clause 43. The non-transitory computer-readable storage medium of any of Clauses 39-42, wherein the instructions, when executed, cause one or more processors to control a display device based on a voice command.

[0216] Clause 44. The non-transitory computer-readable storage medium of any combination of Clauses 39-43, wherein the instructions, when executed, cause one or more processors to present an acoustic space on a mobile phone.

[0217] Clause 45. The non-transitory computer-readable storage medium of any combination of Clauses 39-44, wherein the acoustic space includes a scene represented by video data captured by a camera.

[0218] Clause 46. The non-transitory computer-readable storage medium of any combination of Clauses 39-44, wherein the acoustic space includes a virtual world.

[0219] Clause 47. The non-transitory computer-readable storage medium of any combination of Clauses 39-46, wherein the instructions, when executed, cause one or more processors to present an acoustic space on a head-mounted device.

[0220] Clause 48. The non-transitory computer-readable storage medium of any combination of Clauses 39-47, wherein the instructions, when executed, cause one or more processors to receive a wireless signal.

[0221] Clause 49. The non-transitory computer-readable storage medium of Clause 48, wherein the wireless signal complies with a personal area network standard.

[0222] Clause 49.5. The non-transitory computer-readable storage medium of Clause 49, wherein the personal area network standard includes the AptX standard.

[0223] Clause 50. The non-transitory computer-readable storage medium of Clause 48, wherein the wireless signal complies with a fifth-generation (5G) cellular protocol.

[0224] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, such as according to a communication protocol. In this manner, a computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for use in implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0225] By way of example, and not limitation, such a computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead refer to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0226] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Further, the techniques can be fully implemented in one or more circuits or logic elements.

[0227] The techniques of the present disclosure can be implemented in a variety of apparatus or devices, including a wireless handset, an integrated circuit (IC), or a group of ICs (e.g., a chip set). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of an apparatus configured to perform the disclosed techniques, but need not be implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware device or provided by a collection of interoperating hardware units, including one or more processors as described above in conjunction with appropriate software and / or firmware.

[0228] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A device configured to play one or more audio streams out of a plurality of audio streams, the audio streams including at least one decomposed version of an ambisonics coefficient, the at least one decomposed version of the ambisonics coefficient including at least one spatial component and at least one audio source, wherein, The at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation, and the apparatus comprises: a memory configured to store the at least one spatial component and the at least one audio source within the plurality of audio streams; and one or more processors coupled to the memory and configured to: receive rotation information from a motion sensor; rotate the at least one spatial component based on the rotation information to form at least one rotated spatial component; and reconstruct an ambient stereo signal from the at least one rotated spatial component and the at least one audio source.

2. The device according to claim 1, wherein, The at least one spatial component includes a V-vector identifying spatial characteristics of a corresponding audio object, and the at least one audio source includes a U-vector representing the audio source.

3. The device according to claim 2, wherein The one or more processors are further configured to reconstruct the U-vector by applying a projection matrix to a reference residual vector and a dequantized energy signal.

4. The apparatus according to claim 3, wherein The projection matrix includes time and spatial rotation data.

5. The apparatus according to claim 1, wherein, The one or more processors are further configured to output a representation of the at least one audio source to one or more speakers.

6. The device according to claim 1, wherein The one or more processors are further configured to combine at least two representations of the at least one audio source by at least one of mixing or interpolation.

7. The apparatus according to claim 1, further comprising a display device.

8. The apparatus according to claim 7, further comprising a microphone, wherein, The one or more processors are further configured to receive a voice command from the microphone and control the display device based on the voice command.

9. The apparatus according to claim 1, further comprising one or more speakers.

10. The device according to claim 1, wherein, The apparatus includes a mobile phone.

11. The apparatus according to claim 1, Among them, The apparatus includes an extended reality head-mounted device, and wherein the acoustic space includes a scene represented by video data captured by a camera.

12. The apparatus according to claim 1, Among them, The apparatus includes an extended reality head-mounted device, and wherein the acoustic space includes a virtual world.

13. The apparatus according to claim 1, further comprising a head-mounted device configured to present the acoustic space.

14. The apparatus according to claim 1, further comprising a wireless transceiver coupled to the one or more processors and configured to receive wireless signals, the wireless signals including one or more signals compliant with a fifth-generation cellular standard, a Bluetooth standard, or a Wi-Fi standard.

15. A method of playing one or more audio streams out of a plurality of audio streams, the audio streams including at least one decomposed version of an ambisonics coefficient, the at least one decomposed version of the ambisonics coefficient including at least one spatial component and at least one audio source, wherein, The at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation, and the method comprises: storing, by a memory, the at least one spatial component and the at least one audio source within the plurality of audio streams; receiving, by one or more processors, rotation information from a motion sensor; rotating, by one or more processors, the at least one spatial component based on the rotation information to form at least one rotated spatial component; and reconstructing, by the one or more processors, an ambient stereo signal from the at least one rotated spatial component and the at least one audio source.

16. The method according to claim 15, wherein, The at least one spatial component includes a V-vector identifying spatial characteristics of a corresponding audio object, and the at least one audio source includes a U-vector representing the audio source.

17. The method of claim 16, further comprising reconstructing the U-vector by applying a projection matrix to a reference residual vector and a dequantized energy signal.

18. The method according to claim 17, wherein, The projection matrix includes temporal and spatial rotation data.

19. The method of claim 15, further comprising outputting, by the one or more processors, a representation of the at least one audio source to one or more speakers.

20. The method of claim 15, further comprising combining, by the one or more processors, at least two representations of the at least one audio source by at least one of mixing or interpolation.

21. The method of claim 15, further comprising receiving a voice command from a microphone and controlling a display device based on the voice command.

22. The method according to claim 15, wherein, The method is executed on a mobile phone.

23. The method according to claim 15, wherein, The method is executed on an extended reality head-mounted device, and wherein the acoustic space includes a scene represented by video data captured by a camera.

24. The method according to claim 15, wherein, The method is executed on an extended reality head-mounted device, and wherein the acoustic space includes a virtual world.

25. The method according to claim 15, wherein The method is executed on a head-mounted device configured to present an acoustic space.

26. The method of claim 15, further comprising receiving a wireless signal, the wireless signal including one or more signals compliant with a fifth generation cellular standard, a Bluetooth standard, or a Wi-Fi standard.

27. A device configured to play one or more audio streams out of a plurality of audio streams, the audio streams including at least one decomposed version of an ambisonics coefficient, the at least one decomposed version of the ambisonics coefficient including at least one spatial component and at least one audio source, wherein, The at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation, the apparatus comprising: means for storing at least one spatial component and at least one audio source within a plurality of audio streams; means for receiving rotation information from a motion sensor; means for rotating the at least one spatial component to form at least one rotated spatial component; and means for reconstructing an ambient stereo signal from the at least one rotated spatial component and the at least one audio source.

28. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: Store at least one spatial component and at least one audio source within a plurality of audio streams, the audio streams including at least one decomposed version of an ambisonics coefficient, the at least one decomposed version of the ambisonics coefficient including the at least one spatial component and the at least one audio source, wherein, The at least one spatial component describes spatial characteristics associated with the at least one audio source in a spherical harmonic domain representation; receive rotation information from a motion sensor; rotate the at least one spatial component based on the rotation information to form at least one rotated spatial component; and reconstruct an ambient stereo signal from the at least one rotated spatial component and the at least one audio source.

29. The non-transitory computer-readable storage medium according to claim 28, wherein, The at least one spatial component includes a V-vector identifying spatial characteristics of a corresponding audio object and the at least one audio source includes a U-vector representing the audio source.

30. The non-transitory computer-readable storage medium of claim 29, further having instructions stored thereon that, when executed, cause the one or more processors to reconstruct the U-vector, including by applying a projection matrix to a reference residual vector and a dequantized energy signal.

Citation Information

Patent Citations

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US10405126B2

  • Mixed-order ambisonics (MOA) audio data for computer-mediated reality systems

    US20190007781A1

  • Adaptive ambisonic binaural rendering

    US20160241980A1

  • Fast and memory efficient encoding of sound objects using spherical harmonic symmetries

    US20190069110A1