Pose correction metadata for interactive headtracking
Patent Information
- Authority / Receiving Office
- IL · IL
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2024-12-16
- Publication Date
- 2026-07-01
AI Technical Summary
Existing audio processing technologies face challenges in efficiently rendering immersive audio experiences across various devices, particularly in split-rendering topologies, where accurate head-tracking and binaural signal reconstruction are crucial to maintain audio quality and prevent comb filtering artifacts.
The proposed solution involves encoding and decoding immersive audio content using a method that includes obtaining input audio signals with immersive content, rendering sets of binaural signals based on head poses, computing reconstruction metadata to enable the reconstruction of binaural signals, and encoding these signals and metadata in a bitstream for transmission between devices.
This approach enables efficient reconstruction of immersive audio experiences across devices, effectively mitigates comb filtering artifacts, and maintains the interaural level difference (ILD) critical for listener perception, thereby enhancing the overall audio rendering quality.
Smart Images

Figure 00000059_0000 
Figure 00000060_0000 
Figure 00000061_0000
Abstract
Description
Docket No.: D23168WO01 POSE CORRECTION METADATA FOR INTERACTIVE HEADTRACKING CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No.63 / 613,663, filed December 21, 2023, and U.S. Provisional Application No.63 / 733,993, filed December 13, 2024, each of which is hereby incorporated by reference. TECHNICAL FIELD
[0002] This application relates generally to audio processing, and more specifically to audio rendering in split-rendering topologies. BACKGROUND
[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.
[0004] Voice and video encoder / decoder (codec) standards have recently focused on developing codecs for immersive audio services, such as for immersive voice and audio services (IVAS) and the 3GPP TS 26.253 standard. The immersive audio services are expected to support a range of audio service capabilities, including but not limited to mono to stereo upmixing and fully immersive audio encoding, decoding, and rendering.
[0005] A wide range of devices, including endpoints and network nodes, are expected to support immersive audio services.Some example devices include, but are not limited to, mobile devices such as smart phones and electronic tablets, personal computers, video and audio conferencing, home theater devices, television and video gaming devices, extended reality (XR) devices, and other suitable devices. These devices, end points, and network nodes may have various acoustic interfaces for sound capture, transmission, relay, and rendering.
[0006] Extended reality (XR) applications, such as augmented reality (AR), mixed reality (MR), and virtual reality (VR), may be implemented for use on any variety of physical devices, where such XR applications support immersive audio. More specifically, XR devices may leverage immersive audio services to enhance the user experience with a fully immersive interactive experience. In such XR applications, the audio rendering may be adjusted responsive to the motion of the user. For example, a user’s head position and / or head movement may be tracked, and the audio rendering may be adjusted in response to the user’s tracked motion. Thus, anDocket No.: D23168WO01 immersive audio experience may process head movements using models with three degrees of freedom (3DoF) or six degrees of freedom (6DoF).
[0007] It is with respect to these and other considerations that the disclosure made herein is presented. BRIEF SUMMARY OF THE DISCLOSURE
[0008] Techniques for coding, signaling, and decoding audio signals are described herein. In some examples, these techniques may be applied to split-rendering of audio between an upstream device and a downstream device, such as for an IVAS split-rendering topology.
[0009] Briefly stated, systems, methods and devices for encoding and decoding immersive audio content are disclosed. Some examples provide methods for encoding an input audio signal. An example method includes obtaining an input audio signal with immersive audio content, obtaining a first set of head poses, and rendering a first set of binaural signals. The example method may also include obtaining a second set of head poses and rendering a second set of binaural signals. The example method may include computing, in a first frequency range, a first reconstruction metadata to enable the reconstruction of both magnitude and phase of the second set of binaural signals and computing, in a second frequency range, a second reconstruction metadata to enable the reconstruction of the magnitude of the second set of binaural signals. The method may also include encoding the first set of binaural signals, the first reconstruction metadata, and the second reconstruction metadata in a bitstream.
[0010] In some described embodiments, a method to encode an input audio signal with immersive audio content in a bitstream from a first device to a second device is described, the method for the first device comprising: obtaining the input audio signal with immersive audio content; obtaining a first set of head poses; rendering a first set of binaural signals with the immersive audio content based on the first set of head poses; obtaining a second set of head- poses that are different from the first set of head poses; rendering a second set of binaural signals with the immersive audio content based on the second set of head poses; computing, in a first frequency range of a banded domain, a first reconstruction metadata (M) to enable the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals; computing, in a second frequency range of the banded domain, a second reconstruction metadata (D) to enable the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; encoding the first set of binaural signals,Docket No.: D23168WO01 the first reconstruction metadata (M) and the second reconstruction metadata (D) in thebitstream; and outputting the bitstream.
[0011] In some embodiments, a method to decode immersive audio content from a bitstream from a first device to a second device is described, the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain: a first set of binaural signals of the immersive audio content in a banded domain, a first reconstruction metadata (M) corresponding to a first frequency range of a banded domain, wherein the first reconstruction metadata (M) enables the reconstruction of both magnitude and phase of a second set of binaural signals from the first set of binaural signals, and a second reconstruction metadata (D) corresponding to a second frequency range of the banded domain, wherein the second reconstruction metadata (D) enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals, wherein the second set of binaural signals correspond to a second set of head poses such that the second set of head poses are different than the first set of head poses; obtaining a first set of head poses associated with the first set of binaural signals; obtaining the second set of head poses associated with the second set of binaural signals; detecting a current head pose (P) associated with a user of the second device; reconstructing, in the first frequency range, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata (M), and a relationship between the second set of head poses, the first set of head poses, and the current head pose (P); reconstructing, in the second frequency range, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata (D) and a relationship between the second set of head poses, the first set of head poses, and the current head pose (P); and generating a binaural output by combining the first binaural audio inthe first frequency range and the second binaural audio in the second frequency range.
[0012] In some additional embodiments, a method to decode immersive audio content from a bitstream from a first device to a second device is described, the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain: a first set of binaural signals of the immersive audio content in a banded domain; and a first reconstruction metadata (M), wherein the first reconstruction metadata (M) enables the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals; obtaining a first set of head poses associated with the first set of binaural signals; obtaining a second set of head poses associated with a second set of binaural signals, wherein the second set of head poses are different than the first set of head poses; detecting a current head pose (P) associated with a user of the second device; reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata (M),Docket No.: D23168WO01 and a relationship between the second set of head poses, the first set of head poses, and the current head pose (P); generating, in a second frequency range of the banded domain, a second reconstruction metadata (D) from the first reconstruction metadata (M) based on the first set of binaural signals, wherein the second reconstruction metadata (D) enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; reconstructing, in the second frequency range of the banded domain, a second output binaural audio based on the first set of binaural signals, the second reconstruction metadata (D) and a relationship between the second set of head poses, the first set of head poses and the current head pose (P); and generating a binaural output by combining the first binaural audio in the firstfrequency range and the second binaural audio in the second frequency range.
[0013] In still other embodiments, a method to decode immersive audio content from a bitstream from a first device to a second device is described, the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain a first set of binaural signals of the immersive audio content in a banded domain; obtaining a first set of head poses associated with the first set of binaural signals; detecting a current head pose (P) associated with a user of the second device; reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, and a relationship between the first set of head poses and the current head pose (P); selecting a reference pose from the first set of head poses and a corresponding reference binaural signal from the first set of binaural signals; generating, in a second frequency range of the banded domain, a first reconstruction metadata (D) from the first set of binaural signals, wherein the first reconstruction metadata (D) enables the reconstruction of the magnitude of the first set of binaural signals from the reference binaural signal; reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the first reconstruction metadata (D) and a relationship between the first set of head poses and the current head pose (P); and generating a binaural output by combining the first binaural audio in the first frequency range and the second binauralaudio in the second frequency range.
[0014] In some embodiments, the first frequency range of the banded domain corresponds to frequency values below a frequency threshold, and the second frequency range of the banded domain corresponds to frequency values above the frequency threshold. The frequency threshold may be a constant value in a range of approximately 1000 Hz to 3000 Hz. In some embodiments, the banded domain has a bandwidth of approximately 24 kHz sampled at 48 kHz, which is divided into 60 bands with a resolution of approximately 400 Hz per band.Docket No.: D23168WO01
[0015] In some further embodiments, the banded domain is one of a complex low delay filter bank (CLDFB) domain, a modified discrete Fourier transform (MDFT) domain, and a modified discrete cosine transform (MDCT) domain. Example methods may include converting the first set of binaural signals and the second set of binaural signals from a time domain to the banded domain by processing the binaural signals in a current frame with a CLDFB filterbank that includes a plurality of subbands, by applying a MDFT to the binaural signals in the time domain, or by applying a MDCT to the binaural signals in the time domain.
[0016] In some embodiments, head poses, including sets of head poses, are transmitted from the second device (e.g., the downstream device) to the first device (e.g., the upstream device).
[0017] Various aspects of the present disclosure provide for processing of audio signals, and effect improvements in at least the technical fields of audio processing, audio encoding, audio decoding, virtual reality, and the like.
[0018] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.
[0019] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims. DESCRIPTION OF THE DRAWINGS
[0020] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:
[0021] FIG.1 illustrates a block diagram of an IVAS coder / decoder framework for encoding and decoding IVAS bitstreams in which various aspects of the present disclosure can be practiced.
[0022] FIG.2 illustrates an example comb filtering artifact for a sine tone signal added to its delayed version.Docket No.: D23168WO01
[0023] FIG.3 illustrates another example comb filtering artifact for a sine tone signal added to its delayed version.
[0024] FIG.4 illustrates yet another example comb filtering artifact for a sine tone signal added to its delayed version.
[0025] FIG.5 illustrates a block diagram of an example system for split rendering with interactive head tracking, in accordance with some aspects of the present disclosure.
[0026] FIG.6 illustrates a block diagram of various example methods for encoding an input audio signal with immersive audio content in a bitstream from a first device to a second device according to some aspects of the present disclosure.
[0027] FIG.7 illustrates a block diagram of various example methods for decoding an immersive audio content from a bitstream from a first device to a second device according to some aspects of the present disclosure.
[0028] FIG.8 illustrates a block diagram of various example methods for decoding an immersive audio content from a bitstream from a first device to a second device according to some aspects of the present disclosure.
[0029] FIG.9 illustrates a block diagram of various example methods for decoding an immersive audio content from a bitstream from a first device to a second device according to some aspects of the present disclosure.
[0030] FIG.10A illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.
[0031] FIG.10B illustrates a schematic block diagram of an example CPU implemented in the device architecture of FIG.10A that may be used to implement various aspects of the present disclosure. DETAILED DESCRIPTION
[0032] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments with reference to the accompanying drawings. The illustrative embodiments in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes made, without departing from the spirit or scope of the present disclosure. In light of the present disclosure, it will be apparent to one of ordinary skill in the art that the various described features andDocket No.: D23168WO01 implementations may be practiced without many of these specific details. In some instances, well-known methods, procedures, components, and circuits, have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described hereafter that can each be used independently of one another or with any combination of other features. Thus, the features may be arranged, substituted, combined, separated, or designed into other configurations, which is contemplated in light of the present disclosure.
[0033] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0034] Throughout this disclosure, various terms are used to describe upstream and downstream devices that may cooperatively provide a split-rendering solution for immersive audio experiences. In a typical implementation for split rendering, the solution may include an upstream device that may be considered a “heavyweight processing device”, while the downstream device may be considered a “lightweight processing device.” The terms “heavyweight” and “lightweight” may take on a number of meanings based on the context presented. In some examples, a “lightweight processing device" or “heavyweight processing device” may refer to the physical weight of the device. In other examples, a “lightweight processing device" and “heavyweight processing device” may refer to the form factor of being portable versus less portable in terms of size. In still other examples, a “lightweight processing device" and “heavyweight processing device” may refer to different battery or power requirements, while in other examples the distinctions are based on computational speed and / or complexity.Docket No.: D23168WO01
[0035] Various Acronyms that may appear throughout this disclosure and in the associated claims and / or drawings are listed below. Other commonly used acronyms and terms of art may be excluded from this list in the interest of brevity. Thus, a short list of acronyms is provided below as an easy reference for the reader. IVAS – Immersive Voice and Audio Services ISAR – Immersive Audio for Split Rendering Scenarios CLDFB – Complex Low Delay Filter Bank HOA – Higher Order Ambisonics FOA – First Order Ambisonics SPAR – Spatial Reconstructor DirAC – Directional Audio Coding MD – Metadata BS – Bitstream EVS – Enhanced Voice Services MDFT – Modified Discrete Fourier Transform MDCT – Modified Discrete Cosine Transform ITD – Interaural Time Difference ILD – Interaural Level Difference IPD – Interaural Phase Difference
[0036] An immersive voice and audio services (IVAS) system is expected to support a range of audio service capabilities, including but not limited to mono to stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including but not limited to mobile and smart phones, electronic tablets, personal computers, conference phones, extended reality (XR) devices, home theatre devices, and other suitable devices. Additionally, IVAS is expected to support a broad and varied set of communication topologies, including but not limited to a mix of wired communications and wireless communications. The wireless communications may include a mix of cellular, Wi-Fi, and Bluetooth technologies.
[0037] FIG.1 illustrates a block diagram of an IVAS coder / decoder (“codec”) framework 100 for encoding and decoding IVAS bitstreams, according to one or more embodiments. The example IVAS codec framework 100 includes an IVAS encoder 101 and an IVAS decoder 104. The IVAS encoder 101 may be considered a first device that is located upstream from a second device (e.g., a mobile device) such as the IVAS decoder 104. Thus, the first device may also beDocket No.: D23168WO01 referred to as an upstream device or an upstream encoder device, while the second device may be referred to as a downstream device or a downstream decoder device.
[0038] The IVAS encoder 101 includes a spatial encoder 102 and a core audio encoder 103. The input of the spatial encoder 102 corresponds to a first path 110. The spatial encoder 102 receives N channels of input spatial audio (e.g., input audio content, First Order Ambisonics [FOA], Higher Order Ambisonics [HOA]) via the first path 110. The spatial encoder 102 processes and encodes the received input audio. In some implementations, the spatial encoder 102 implements Spatial Reconstructor (SPAR) and Directional Audio Coding (DirAC) for analyzing / downmixing the received input audio into N_dmx downmix audio channels, as described in further detail below. The N_dmx downmix audio channels may be mono channels or spatial channels. The outputs of the spatial encoder 102 correspond to a second path 111 and a third path 112. The spatial encoder 102 is coupled to the core audio encoder 103 via the second path 111. The spatial encoder 102 is coupled to the IVAS decoder 104 via the third path 112. The output of the spatial encoder 102 includes a spatial metadata (MD) bitstream (BS) and N_dmx channels of downmix. The N_dmx channels of downmix are provided by the spatial encoder 102 to the core audio encoder 103 via the second path 111. The spatial MD BS is provided by the spatial encoder 102 to the IVAS decoder 104 via the third path 112. The spatial MD is quantized and entropy coded. In some implementations, quantization can include fine, moderate, coarse, and extra coarse quantization strategies and entropy coding can include Huffman or Arithmetic coding. The framework permits not more than three levels of quantization at a given operating mode; however, with decreasing bitrates, the three levels become increasingly coarser overall, to meet bitrate requirements. Some example bitrates in IVAS may be in a range from about 13.2kbps to about 512kbps for immersive modes, and a variable bitrate from about 5.9kbps to about 128kbps for EVS modes, with each frame representing about 20ms of speech or audio data.
[0039] The input of the core audio encoder 103 corresponds to the second path 111. The output of the core audio encoder 103 corresponds to a fourth path 113. The core audio encoder 103 (e.g., based on mono Enhanced Voice Services (EVS) encoding unit or core coders described in the TS 26.253 IVAS specification) encodes N_dmx channels (N_dmx = 1-16 channels) of the spatial downmix into an audio bitstream, which is combined (via the fourth path 113) with the spatial MD bitstream into an IVAS encoded bitstream transmitted to IVAS decoder 104 via the third path 112.Docket No.: D23168WO01
[0040] The IVAS decoder 104 includes a core audio decoder 105 (e.g., an EVS decoder or core decoders described in the TS 26.253 IVAS specification) and a spatial decoder / renderer 106 (e.g., SPAR / DirAC). The input of the core audio decoder 105 corresponds to a fifth path 114. The core audio decoder 105 receives the audio bitstream via the fifth path 114. The core audio decoder 105 is configured to decode the audio bitstream extracted from the IVAS bitstream to recover the N_dmx audio channels. The output of the core audio decoder 105 corresponds to a sixth path 115. The core audio decoder 105 is coupled to the spatial decoder / renderer 106 via the sixth path 115. The core audio decoder 105 is configured to provide the N_dmx audio channels (e.g., the decoded spatial downmix) to the spatial decoder / renderer 106 via the sixth path 115.
[0041] The inputs of the spatial decoder / renderer 106 correspond to the third path 112 and the sixth path 115. The spatial decoder / renderer 106 receives the spatial MD bitstream from the IVAS encoder 101 via the third path 112 and receives the decoded spatial downmix from the core audio decoder 105 via the sixth path 115. The spatial decoder / renderer 106 decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD and synthesizes (e.g., renders) output audio channels using the spatial MD and a spatial upmix for playback on various audio systems with different speaker configurations and capabilities. The output of the spatial decoder / renderer 106 corresponds to a seventh path 116. The spatial decoder / renderer 106 is configured to provide the output audio (e.g., decoded audio) via the seventh path 116.
[0042] Interactive headtracking is a technique where a user device can be confgured to monitor (or track) the movement and / or position of the user’s head. In a split rendering system, such as may be applicable for an IVAS or Immersive Audio for Split Rendering Scenarios (ISAR) codec, the headtracking information may be encoded into metadata that is communicated between an upstream device and a downstream device. By processing the user movement and / or position in the upstream device, audio pre-rendering may be accomplished by the upstream device and transmitted to the downstream device. Accordingly, immersive audio can be experienced by the user on the downstream device without requiring significant processing resources on the downstream device (e.g., power resources, heat dissipation devices, expensive processing devices, etc.).
[0043] Examples described herein relate to binaural signals having left and right channels. The left and right channels of a binaural signal may have Interaural Time Difference (ITD) and interaural phase difference (IPD) depending on the source location and direction of arrival.Docket No.: D23168WO01 Moreover, the left and right channels of a first binaural signal s1 may have a different time delay and phase as compared to the left and right channel of a second binaural signal s2. Interpolation or weighted addition of such signals may lead to comb filtering in certain frequencies, and more specifically in high frequencies. For example, if a first binaural signal s1 and a second binaural signal s2 have the same source audio, and signal s2 is delayed by 0.1 ms with respect to signal s1, then an averaged signal s3 based on the combination of signals s1 and s2 corresponds to s3 = (s1+s2) / 2, where the averaged signal s3 for this example has a notch at 5000 Hz. Consequently, if pose correction metadata (described below with respect to FIG.5) to reconstruct BIN signal s1’ from the first binaural signal s1 is M1 and the pose correction metadata to reconstruct BIN signal s2’ from the first binaural signal s1 is M2such that M1and M2can have a phase difference at any frequency, then interpolation between M1 and M2 can also result in comb filtering artifacts leading to unwanted attenuation at corresponding comb frequencies in the final head tracked output.
[0044] FIG.2 illustrates an example comb filtering artifact for a sine tone signal added to its delayed version. The x-axis represents sample numbers. The y-axis represents a normalized amplitude of the signals. Specifically, a first sine tone signal s1 is illustrated with a frequency of 5000 Hz with a corresponding period of 0.2ms. A second sine tone signal s2 is illustrated with a frequency of 5000 Hz and has a delay of approximately 0.05 ms or about 10 sample points in FIG.2 relative to the first sine tone signal s1. This time delay of 0.05ms correponds to a phase difference between signals s1 and s2 of about 90 degrees. An average signal s3 is a result of the combination of first sine tone signal s1 and the second sine tone signal s2, which is given as s3 = (s1+s2) / 2. The average signal s3 in FIG.2 is shown to include an attenuated amplitude of about 70% of the original amplitude, which is a result of the 0.05 ms time delay and corresponding 90 degree phase difference between signals s1 and s2.
[0045] FIG.3 illustrates another example comb filtering artifact for a sine tone signal added to its delayed version. The x-axis represents sample numbers. The y-axis represents an amplitude of the signals. Specifically, a first sine tone signal s1 is illustrated with a frequency of 5000 Hz with a corresponding period of 0.2ms. A second sine tone signal s2 is illustrated with a frequency of 5000 Hz and has a delay of approximately 0.075 ms or about 15 sample points in FIG.3 relative to the first sine tone signal s1. This time delay of 0.075ms correponds to a phase difference between signals s1 and s2 of about 135 degrees. An average signal s3 is a result of combination of the first sine tone signal s1 and second sine tone signal s2, which is given as s3 = (s1+s2) / 2. The average signal s3 in FIG.3 is shown to include an attenuated amplitude of aboutDocket No.: D23168WO01 40% of the original amplitude, which is a result of the 0.075ms time delay and corresponding 135 degree phase difference between signals s1 and s2.
[0046] FIG.4 illustrates yet another example comb filtering artifact for a sine tone signal added to its delayed version. The x-axis represents sample numbers. The y-axis represents an amplitude of the signals. Specifically, a first sine tone signal s1 is illustrated with a frequency of 5000 Hz with a corresponding period of 0.2ms. A second sine tone signal s2 has a frequency of 5000 Hz and has a delay of approximately 0.1 ms or about 20 sample points in FIG.3 compared to the first sine tone signal s1. This time delay of 0.01ms correponds to a phase difference between signals s1 and s2 of about 180 degrees. An average signal s3 is a result of the combination of first sine tone signal s1 and second sine tone signal s2, which is given as s3 = (s1+s2) / 2. The average signal s3 in FIG.4 is shown to include an attenuated amplitude of about 0% of the original amplitude, which is a result of the 0.1ms time delay and corresponding 180 degree phase difference between signals s1 and s2.
[0047] As shown in FIGS.2-4, due to the phase differences between signals s1 and s2, the amplitude of signal s3 is lower than the amplitude of signals s1 or s2. For example, the smaller phase difference of about 90 degrees observed in FIG.2 results in an attenuation to about 70% of the maximum signal amplitude for signal s3; while the larger the phase difference of about 135 degrees observed in FIG.3 results in an attenuation to about 40% of the maximum signal amplitude for signal s3.A worst case scenario is observed in FIG.4, where a phase difference of 180 degrees (i.e., completely out of phase) results in a complete attenuation of the amplitude of signal s3 to zero, since the sum of signals s1 and s2 cancel one another when they are completely out of phase.
[0048] Additionally, the interaural phase difference (IPD) in a high frequency range is largely unimportant with respect to a listener’s perception. Accordingly, even if the phase difference in a high frequency range (e.g., above about 2 kHz) in the left and right channels of a binaural signal S are different than the phase difference in the high frequency range in the left and right channels of a binaural signal S’, both S and S’ may provide the same listening perception. However, the Interaural Level Difference (ILD) in a high frequency range is critical with respect to listener’s perception such as perceived direction to a sound source, and should be preserved when possible.
[0049] Examples, aspects, and instances described herein achieve improved generation of pose correction metadata such that interpolation (e.g., linear interpolation) may be performed on the pose correction metadata without resulting in comb filtering artifacts. Additionally, theDocket No.: D23168WO01 frequency range of binaural signals may be considered when generating pose correction metadata such that the ILD is maintained when critical to listener perception. Additional benefits other than those explicitly summarized may also be realized by the following example implementations.
[0050] FIG.5 illustrates a block diagram of an example system 500 for split rendering with interactive head tracking, arranged in accordance with some embodiments. As illustrated, the example system 500 may include a first device 510 (e.g., an upstream device or pre-renderer) and a second device 520 (e.g., a downstream device, lightweight processing device, or post- renderer).
[0051] The first device 510 includes a decoder 511, a downmixer 512, a head pose decoder 513, a binaural renderer 514, a metadata generator 515, a first encoder 516, a second encoder 517, and a multiplexer 518. The first device 510 may be, for example, a server. The first device 510 may correspond to, for example, the IVAS decoder 104 of FIG.1. The second device 520 includes a demultiplexer 521, a first decoder 522, a second decoder 523, a head-tracker 524, a pose information encoder 525, and a binaural reconstruction block 526. The second device 520 may be a user-held device, a lightweight device such as earbuds or AR glasses, a VR headset, or the like.
[0052] In the first device 510, the input of the decoder 511 corresponds to a first path 551. The input of decoder 511 receives a bitstream b1from the first path 551. The decoder 511 is configured to decode an immersive audio content A from the bitstream b1. The output of the decoder 511 corresponds to a second path 552. The output of decoder 511 is coupled to a first input of downmixer 512 and a first input of binaural renderer 514 via the second path 552. The decoder 511 is configured to provide the audio content A to the downmixer 512 and the binaural renderer 514 via the second path 552.
[0053] The inputs of the downmixer 512 correspond to the second path 552 and a third path 553. A first input of downmixer 512 is coupled to the output of decoder 511 via the second path 552. A second input of the downmixer 512 is coupled to the output of the head pose decoder 513 via the third path 553. The downmixer 512 receives the audio content A from the decoder 511 via the second path 552. The downmixer 512 is configured to generate a downmix representation Dmx of the audio content A. The downmixer 512 may be a binaural renderer that renders a reference binaural signal BIN’ corresponding to the head pose obtained from the head pose decoder 513 using the immersive audio content A. The output of the downmixer 512 corresponds to a fourth path 554. The output of downmixer 512 is coupled to the input of theDocket No.: D23168WO01 first encoder 516 and a first input of the metadata generator 515 via the fourth path 554. The downmixer 512 is configured to provide the downmix representation Dmx to the first encoder 516 via the fourth path 554.
[0054] The input of the head pose decoder 513 corresponds to a fifth path 555. The input of the head pose decoder 513 is coupled to pose information encoder 525 via the fifth path 555. The head pose decoder 513 receives a pose bitstream bp, which includes head pose information, from the pose information encoder 525 via the fifth path 555. The head pose decoder 513 is configured to decode the pose bitstream bp and generate a first head pose P’ (e.g., a reference head pose). The output of the head pose decoder 513 corresponds to the third path 553. The output of the head pose decoder 513 is coupled to the second input of the downmixer 512 and the second input of the binaural renderer 514 via the third path 553. The head pose decoder 513 is configured to provide the first head pose P’ to the downmixer 512 and the binaural renderer 514 via the third path 553.
[0055] The inputs of the binaural renderer 514 correspond to the second path 552 and the third path 553. The first input of the binaural renderer 514 is coupled to the output of the decoder 511 via the second path 552. The binaural renderer 514 receives the audio content A from the decoder 511 via the second path 552. The second input of the binaural renderer 514 is coupled to the head pose decoder 513 via the third path 553. The binaural renderer 514 receives the first head pose P’ from the head pose decoder 513 via the third path 553. The binaural renderer 514 is configured to render one or several binaural representations BINn corresponding to a second set of head poses P2 using the immersive audio content A, where in the second set of head poses P2are generated by adding offset (e.g., fixed offsets known to both device 510 and 520 or dynamic offsets that are coded in metadata bitstream at the output of the second encoder 517) to the first head pose P’. The output of the binaural renderer 514 corresponds to a sixth path 556. The output of the binaural renderer 514 is coupled to the second input of the metadata generator 515 via the sixth path 556. The binaural renderer 514 is configured to provide the binaural representations BINn to the metadata generator 515 via the sixth path 556.
[0056] The inputs of the metadata generator 515 correspond to the fourth path 554 and the sixth path 556. The first input of the metadata generator 515 is coupled to the output of the downmixer 512 via the fourth path 554. The metadata generator 515 receives the downmix Dmx from the downmixer 512 via the fourth path 554. The second input of the metadata generator 515 is coupled to the binaural renderer 514 via the sixth path 556. The metadata generator 515 receives the binaural representations BINn from the binaural renderer 514 via the sixth path 556. TheDocket No.: D23168WO01 metadata generator 515 is configured to generate reconstruction metadata M indicating (e.g., instructing how to) reconstruction of the binaural representations BINnfrom the downmix Dmx. In some instances, the metadata M includes the first head pose P’, phase information, amplitude information, or combinations thereof. The output of the metadata generator 515 corresponds to a seventh path 557. The output of the metadata generator 515 is coupled to the input of the second encoder 517 via the seventh path 557. The metadata generator 515 is configured to provide the metadata M to the second encoder 517 via the seventh path 557. In some instances, Dmx is the reference binaural signal BIN’ corresponding to first head pose P’ and metadata M includes a first reconstruction metadata ^^^at frequencies ^^, as shown below in equation (5), such that0 ≤ ^^ < ^^, where ^^ is a constant ranging from 1000^^ ≤ ^^ ≤ 3000^^, and such that ^^^reconstructs both phase and magnitude of the binaural representations BINn from the reference binaural signal BIN’. Metadata M further includes a second reconstruction metadata ^^^atfrequencies ^^, as shown below in equation (7), such that ^^ ≤ ^^, where ^^ is a constant rangingfrom 1000^^ ≤ ^^ ≤ 3000^^, and such that ^^^ reconstructs only magnitude of the binauralrepresentations BINn from the reference binaural signal BIN’. The computation of the second reconstruction metadata ^^^is performed so that only magnitude parameters are interpolated at high frequencies in the binaural reconstruction block 526 in the second device 520, as shown below in equation (9), which avoids comb filtering artifact in the reconstructed output, as shown below in equation (10), as there is no interpolation of phase.
[0057] The input of the first encoder 516 corresponds to the fourth path 554. The input of the first encoder 516 is coupled to the output of the downmixer 512 via the fourth path 554. The first encoder 516 receives the downmix Dmx from the downmixer 512 via the fourth path 554. The first encoder 516 is configured to encode the downmix Dmx as a first encoded bitstream b11. The output of the first encoder 516 corresponds to an eighth path 558. The output of the first encoder 516 is coupled to the first input of the multiplexer 518 via the eighth path 558. The first encoder 516 is configured to provide the first encoded bitstream b11 to the multiplexer 518 via the eighth path 558. In some instances, Dmx is the reference binaural signal BIN’ corresponding to first head pose P’ and the first encoded bitstream b11 corresponds to the encoded reference binaural signal BIN’.
[0058] The input of the second encoder 517 corresponds to the seventh path 557. The input of the second encoder 517 is coupled to the output of the metadata generator 515 via the seventh path 557. The second encoder 517 receives the metadata M (which may include the first head pose P’) from the metadata generator 515 via the seventh path 557. The second encoder 517 isDocket No.: D23168WO01 configured to encode the metadata M as a second encoded bitstream b12. The output of the second encoder 517 corresponds to a ninth path 559. The output of the second encoder 517 is coupled to the second input of the multiplexer 518 via the ninth path 559. The second encoder 517 is configured to provide the second encoded bitstream b12to the multiplexer 518 via the ninth path 559.
[0059] The inputs of the multiplexer 518 correspond to the eighth path 558 and the ninth path 559. The first input of the multiplexer 518 is coupled to the first encoder 516 via the eighth path 558. The multiplexer 518 receives the first encoded bitstream b11 from the first encoder 516 via the eighth path 558. The second input of the multiplexer 518 is coupled to the second encoder 517 via the ninth path 559. The multiplexer 518 receives the second encoded bitstream b12 from the second encoder 517 via the ninth path 559. The multiplexer 518 is configured to combine (e.g., multiplex) the encoded bits of the first encoded bitstream b11and the second encoded bitstream b12 into a combined bitstream b2. The output of the multiplexer 518 corresponds to a tenth path 560. The output of the multiplexer 518 is coupled to an input of the second device 520 (and, more specifically, the demultiplexer 521) via the tenth path 560. The multiplexer 518 is configured to provide the combined bitstream b2to the second device 520 via the tenth path 560. The tenth path 560 may be associated with, for example, an interface that provides transmission of the combined bitstream b2to another device external to the first device 510 (e.g., the second device 520).
[0060] In the second device 520, the input of the demultiplexer 521 corresponds to the tenth path 560. The input of the demultiplexer 521 is coupled to the output of the first device 510 (and, more specifically, the multiplexer 518) via the tenth path 560. The demultiplexer 521 receives the combined bitstream b2 from the first device 510 via the tenth path 560. The demultiplexer 521 is configured to separate (e.g., demultiplex) the combined bitstream b2into a first separated bitstream b21 and a second separated bitstream b22. The outputs of the demultiplexer 521 correspond to an eleventh path 561 and a twelfth path 562. The firat output of the demultiplexer 521 is coupled to the first decoder 522 via the eleventh path 561. The demultiplexer 521 is configured to provide the first separated bitstream b21to the first decoder 522 via the eleventh path 561. The second output of the demultiplexer 521 is coupled to the second decoder 523 via the twelfth path 562. The demultiplexer 521 is configured to provide the second separated bitstream b22to the second decoder 523 via the twelfth path 562.
[0061] The input of the first decoder 522 corresponds to the eleventh path 561. The input of the first decoder 522 is coupled to the first output of the demultiplexer 521 via the eleventh path 561.Docket No.: D23168WO01 The first decoder 522 receives the first separated bitstream b21 from the demultiplexer 521 via the eleventh path 561. The first decoder 522 is configured to decode the first separated bitstream b21 to generate (e.g., obtain, construct) a decoded downmix signal Dmx’. The output of the first decoder 522 corresponds to a thirteenth path 563. The output of the first decoder 522 is coupled to the first input of the binaural reconstruction block 526 via the thirteenth path 563. The first decoder 522 is configured to provide the decoded downmix signal Dmx’ to the binaural reconstruction block 526 via the thirteenth path 563. In some instances, Dmx’ is the decoded reference binaural signal BIN’ corresponding to first head pose P’.
[0062] The input of the second decoder 523 corresponds to the twelfth path 562. The input of the second decoder 523 is coupled to the demultiplexer 521 via the twelfth path 562. The second decoder 523 is configured to receive the second separated bitstream b22 from the demultiplexer 521 via the twelfth path 562. The second decoder 523 is configured to decode the second separated bitstream b22 to generate (e.g., obtain, construct) a decoded metadata M’. The decoded metadata M’ may include the first head pose P’ and may further include the dynamic offsets to compute the second set of head poses P2. The second set of head poses P2 may also be computed by adding static offsets to the first head pose P’. The output of the second decoder 523 corresponds to a fourteenth path 564. In some instances, decoded metadata M’ is the quantized representation of metadata M generated by the metadata generator 515 of the first device 510 and includes a first reconstruction metadata ^^^at frequencies ^^, as shown below inequation (5), such that 0 ≤ ^^ < ^^, where ^^ is a constant ranging from 1000^^ ≤ ^^ ≤3000^^, such that ^^^both phase and magnitude of the binaural representations BINn from the reference binaural signal BIN’. M’ further includes a second reconstructionmetadata ^^^ at frequencies ^^, as shown below in equation (7), such that ^^ ≤ ^^, where ^^ is aconstant ranging from 1000^^ ≤ ^^ ≤ 3000^^, such that ^^^ reconstructs only magnitude ofthe binaural representations BINn from the reference binaural signal BIN’. The output of the second decoder 523 is coupled to the second input of the binaural reconstruction block 526 via the fourteenth path 564. The second decoder 523 is configured to provide the decoded metadata M’ to the binaural reconstruction block 526 via the fourteenth path 564.
[0063] The head-tracker 524 is configured to sense a user head position and responsively generate pose information. The pose information may include a current (e.g., actual) user head pose P. The output of the head-tracker 524 corresponds to a fifteenth path 565. The output of the head-tracker 524 is coupled to the input of the pose information encoder 525 and the third input of the binaural reconstruction block 526 via the fifteenth path 565. The head-tracker 524 isDocket No.: D23168WO01 configured to provide the current user head pose P to the pose information encoder 525 and the binaural reconstruction block 526 via the fifteenth path 565.
[0064] The input of the pose information encoder 525 corresponds to the fifteenth path 565. The input of the pose information encoder 525 is coupled to the output of the head-tracker 524 via the fifteenth path 565. The pose information encoder 525 receives the current user head pose P from the head-tracker 524 via the fifteenth path 565. The pose information encoder 525 is configured to encode the current user head pose P in the pose bitstream bp. The output of the pose information encoder 525 corresponds to the fifth path 555. The output of the pose information encoder 525 is coupled to the first device 510 (and, more specifically, the decoder 513) via the fifth path 555. The pose information encoder 525 is configured to provide the pose bitstream bp to the decoder 513 via the fifth path 555.
[0065] The inputs of the binaural reconstruction block 526 correspond to the thirteenth path 563, the fourteenth path 564, and the fifteenth path 565. The first input of the binaural reconstruction block 526 is coupled to the output of the first decoder 522 via the thirteenth path 563. The binaural reconstruction block 526 receives the decoded downmix Dmx’ from the first decoder 522 via the thirteenth path 563. The second input of the binaural reconstruction block 526 is coupled to the output of the second decoder 523 via the fourteenth path 564. The binaural reconstruction block 526 receives the decoded metadata M’ from the second decoder 523 via the fourteenth path 564. The third input of the binaural reconstruction block 526 is coupled to the output of the head-tracker 524 via the fifteenth path 565. The binaural reconstruction block 526 receives the current user head pose P from the head-tracker 524 via the fifteenth path 565. The binaural reconstruction block 526 is configured to interpolate (e.g., calculate, estimate, generate) the decoded metadata M’ based on the difference between the first head pose, and / or the second set of head poses and the current user head pose P to generate metadata Mp and determine (e.g., calculate, estimate, generate) a binaural output BINout based on the decoded downmix Dmx’, the interpolated metadata Mp. In some instances, Dmx’ is the decoded reference binaural signal BIN’ corresponding to first head pose P’ and M’ includes a first reconstruction metadata ^^^atfrequencies ^^, as shown below in equation (5), such that 0 ≤ ^^ < ^^, where ^^ is a constantranging from 1000^^ ≤ ^^ ≤ 3000^^, and such that ^^^ reconstructs both phase andmagnitude of the binaural representations BINn from the reference binaural signal BIN’. M’ further includes a second reconstruction metadata ^^^at frequencies ^^, as shown below inequation (7), such that ^^ ≤ ^^, where ^^ is a constant ranging from 1000^^ ≤ ^^ ≤ 3000^^,such that ^^^reconstructs only magnitude of the binaural representations BINn from theDocket No.: D23168WO01 reference binaural signal BIN’, wherein the computation of second reconstruction metadata ^^^is done so that only magnitude parameters are interpolated at high frequencies by the binaural reconstruction block 526, as shown below in equation (9), which avoids comb filtering artifact in the reconstructed output, as shown below in equation (10), as there is no interpolation of phase. The output of the binaural reconstruction block 526 corresponds to a sixteenth path 566. The binaural reconstruction block 526 is configured to provide the binaural output BINoutvia the sixteenth path 566. The sixteenth path 566 may correspond to the seventh path 116 of FIG.1.
[0066] In general, a pose P (e.g., the current user head pose P, the first head pose P’) may be data representing the orientation and / or head position of a user of a device (e.g., a user of the second device 520). The second device 520 receives and encodes a pose P and sends the pose P as coded data in the pose bitstream bp to the first device 510 through a data channel (e.g., a back channel, the fifth path 555). The first device 510 receives and decodes the pose P from the pose bitstream bp, thereby obtaining first head pose P’, which may be considered a delayed and quantized version of pose P due to a transmission delay between the second device 520 and the first device 510.
[0067] In some implementations, the pose bitstream bp may include information in addition to the current user head pose P, such as one or more parameters V associated with head motion. For example, parameters V may include rotation data including angular velocity, acceleration or deceleration of the user’s head rotation, and the like. Thus, the second device 520 may receive and encode both the current user head pose P and parameters V, and encode both the current user head pose P and the parameters V in the pose bitstream bp transmitted to the first device 510. The first device 510 (and, specifically, the decoder 513) may thus responsively receive and decode both the first head pose P’ and parameters V’ from the bitstream bp.
[0068] Signals described herein may be obtained from either a banded domain or transformed from a time domain into a banded domain. For example, the metadata M for binaural channels corresponding to poses P described above with respect to FIG.5 may be transformed from a time domain representation into a filterbank domain representation that includes a number of bands. In one example, the filterbank domain may have a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band. The banded domain may correspond to a complex low delay filterbank (CLDFB) domain. In some examples, the audio input and / or pose information from the time domain can be transformed into a banded domain using a modified discrete Fourier transform (MDFT) process by a subband processor that creates a complex modulated multi-channel filter bank. In some other examples,Docket No.: D23168WO01 the audio input is transformed into a banded domain including time domain-based filters or using a modified discrete cosine transform (MDCT) process, without requiring complex values. In other examples, the audio input signal may already be obtained in the banded domain, and no subband processor would be required.
[0069] In some examples, interactive headtracking may include generating multiple binaural representations corresponding to various head poses P at a main or pre-renderer device (e.g., the first device 510) and computing metadata M which can be used along with a reference binaural signal to reconstruct binaural output corresponding to any given pose at a secondary or post renderer device (e.g., the second device 520). At the first device 510, such approach requires generating, with the binaural renderer 514, one or more binaural signals BINn with one or more head poses P, herein referred to as a first set of binaural signals BIN1 and a first set of head poses P1, and a second set of binaural signals BIN2with a second set of head poses P2, generating, with the metadata generator 515, a first pose correction metadata M1 such that the second set of binaural signals BIN2can be reconstructed from the first binaural signals BIN1and the first pose correction metadata M1, and transmitting, with the muiltiplexer 518 over the tenth path 560 and in the combined bitstream b2, the first set of binaural signals BIN1and the first pose correction metadata M1 to the second device 520.
[0070] At the second device 520, such approach requires obtaining, with the head-tracker 524, the current head pose P of the user, also referred to as the third head pose P3, generating, with the binaural reconstruction block 526, a second pose correction metadata M2 by interpolating or extrapolating the first pose correction metadata M1 based on the difference in the first set of head poses P1, the second set of head poses P2and the third head pose P3, and computing, with the binaural reconstruction block 526, the head tracked binaural output BINout at the second device 520 by applying the second pose correction metadata M2to the first set of binaural signals BIN1.
[0071] In an example implementation, the first device 510 obtains, using the first path 551, an immersive audio content A and renders, using the binaural renderer 514, a reference binaural signal BINp’ corresponding to a reference pose P’. The reference pose P’ may be either an assumed pose or a pose transmitted from the second device 520 to the first device 510 and decoded by the decoder 513. In addition to generating the reference binaural signal BINp’, the first device 510 may also generate, using the metadata generator 515, pose correction metadata Mp such that a post renderer, such as the binaural reconstruction block 526, can perform pose correction from P’ to the actual, current pose P and generate binaural signal BINpfrom theDocket No.: D23168WO01 reference binaural signal BINp’ using the pose correction metadata Mp. The binaural signal BINp may have all the spatial cues as per current pose P.
[0072] The deviation between reference pose P’ and current pose P may be along one or more of the yaw, pitch, and roll axis. The ITD cues may change with deviation along the yaw axis and the roll axis, but not with deviation along the pitch axis. Hence, comb filtering may occur only while correcting the pose along the yaw axis or roll axis. To generate the pose correction metadata along the yaw or roll axis, the pre-renderer (e.g., the binaural renderer 514) may first generate a second set of binaural signals, BINp1 to BINpN, corresponding to a set of second poses P2(which may include, for example, P21, P22, … P2N). The pre-renderer (e.g., the metadata generator 515) then computes pose correction metadata Mp such that the binaural signals BINp1 to BINpN can be reconstructed from BINp’ using the pose correction metadata Mp.
[0073] In one example, the first set of head poses P1 is the reference pose P’ that is received by the head pose decoder 513. In such an example, the second set of head poses P2(which may include, for example, P21, P22, … P2N) may be obtained (e.g., calculated) by the binaural renderer 514 by applying offsets to the reference pose P’ such that all the head poses in the second set of head poses P2 vary from the reference pose P’ along one axis of rotation. For example, if the reference pose P’ has {yaw, pitch, roll} values of {0, 0, 0}, then the second set of head poses P2may include {15, 0 ,0}, {-15, 0 ,0}, {0, 15, 0}, {0, -15, 0}, {0, 0, 15}, and {0, 0, -15}. These offsets may be static such that both the first device 510 and the second device 520 store these offsets in memory. In another instance, the offsets may be dynamic and are transmitted by the second device 520 to the first device 510 in the pose bitstream bp.
[0074] In another example, the first set of head poses P1 includes the reference pose P’ and several additional head poses. For example, if the reference pose P’ is {0, 0, 0}, then the first set of head poses P1 may be {0, 0, 0}, {90, 0, 0}, and {-90, 0, 0}. The second set of head poses P2 may be obtained (e.g., calculated) by the head pose decoder 513 by applying offsets (static or dynamic offsets) to each head pose within the first set of head poses P1. In this example, given that three poses are provided in the first set of head poses P1 which generate three corresponding binaural signals, pose correction may be performed over a larger range.
[0075] In another example, the first set of head poses P1includes the reference pose P’, that is received by the head pose decoder 513, and a first set of head poses (which may include, for example, P11, P12, … P1N). In such an example, the first set of head poses P1may be obtained (e.g., calculated) by the head pose decoder 513 by applying offsets to the reference pose P’ such that all the head poses in P11, P12, … P1Nvary from the reference pose P’ along one axis ofDocket No.: D23168WO01 rotation. For example, if the reference pose P’ has {yaw, pitch, roll} values of {0, 0, 0}, then the set of head poses may include {15, 0, 0}, {-15, 0, 0}, {0, 15, 0}, {0, -15, 0}, {0, 0, 15} and {0, 0, -15}. These offsets may be static such that both the first device 510 and the second device 520 store these offsets in memory. In another instance, the offsets may be dynamic and are transmitted by the second device 520 to the first device 510 in the pose bitstream bp.
[0076] The pose correction metadata Mp may be generated by the metadata generator 515 based on frequencies of the binaural signals. As one example of generating pose correction metadataMp at a pre-renderer (e.g., the first device 510), at frequencies ^^, such that 0 ≤ ^^ < ^^, where ^^is a constant ranging from 1000^^ ≤ ^^ ≤ 3000^^, the pose correction metadata Mp mayinclude prediction parameters. To generate the prediction parameters, first, all binaural signals may be converted (e.g., “F{…}) by the metadata generator 515 from the time domain to a frequency domain or a filterbank domain (e.g., a CLDFB filterbank), as provided by Equation (1): ^^^^^^ = ^^^^^^^^, ^^^^^^ = ^{^^^^^} Equation (1)
[0077] Then, the prediction parameters are computed by the metadata generator 515 according to Equation (2): ^^^^ = ^ &',^!, ^^ "^ ,^!,^! + ^$% Equation (2)where:^^^^are the prediction parameters, ^,^!,^! is the covariance matrix of reference binaural signal ^^^^^^ , ^,^!, ^^ ,is the covariance matrix of reference binaural signal ^^^^^^ and binaural signal ^^^^^(, is the covariance matrix of binaural signal ^^^^^(, ^^^^^^ is the reference binaural signal corresponding to reference pose P’, and ^^^^^(are the binaural signals corresponding to the second set of poses P2i.
[0078] With the provided prediction matrix, a post prediction matrix may be computed by the metadata generator 515 according to Equation (3): ^= ^^ ^ ! ! ^^∗),^^, ^^ ^^ ,^ ,^ ^^ Equation (3)Docket No.: D23168WO01 Where: ^^^∗^ is the conjugate transpose of matrix ^^^^, and ^),^^, ^^is the post prediction covariance matrix corresponding to binaural signal ^^^^^( .
[0079] To further energy match the post-prediction matrix, an additional gain matrix Gpimay be by the metadata generator 515 according to Equation (4): -^, 0+^( = , ^(0 -.,^( / Equation (4)8).(1) may then be scaled by the metadata: ^^^= +^^∗ ^^^^Equation (5)
[0081] ^^^ is the complex pose correction metadata at frequencies ^^, where 0 ≤ ^^ < ^^, suchthat ^^^^^can be reconstructed by the binaural reconstruction block 526 according to Equation (6): ^^^^( = ^ ^^^ ^^ ^ ^^^ Equation (6)
[0082] At frequencies ^^, such that ^^ ≤ ^^, where ^^ can be constant ranging from 1000^^ ≤^^ ≤ 3000^^, the pose correction metadata may not include a prediction matrix. Rather, thepose correction metadata may include an ILD matching gain matrix Dpi, computed by the metadata generator 515 according to Equation (7): ;0^^ = , ^,^( / Equation (7)^at frequencies ^^, where ^^ ≤ ^^, such^<=^^^ can be reconstructed by the binaural reconstruction block 526 according to Equation (8): ^^^^^( = ^^^^^^^^^ Equation (8)Docket No.: D23168WO01
[0084] Accordingly, at relatively lower frequencies (e.g., a frequency ^^and a threshold ^^where ^^ ≤ ^^), the pose correction metadata may include complex pose correction metadata ^^^describing magnitude and phase information for reconstructing ^^^^^. At relatively higherfrequencies (e.g., a frequency ^^ and a threshold ^^ where ^^ ≤ ^^), the pose correction metadatamay include an ILD matching gain matrix Dpi. The ILD matching gain matrix Dpimay include magnitude information for reconstructing ^^^^^.
[0085] The first device 510 transmits ^^^ ^^^ or ^^^^^ , ^ ^^> , ^^>^, ^ ^^? , ^^?^, … .. , { ^^B , ^^C}and D′ to the second device 520 (e.g., along the tenth path 560 as part of the combined bitstream b2). The ^^^^^or ^^^^^^ are coded in bitstream b11by the first encoder 516 and metadata^ ^^>, ^^>^, ^ ^^?, ^^?^, … .. , { ^^B , ^^C} and D′ are coded in bitstream b12 by block 517, bothb11 and b12 are multiplexed into bitstream b2 by the multiplexer 518. In some implementations, both the first device 510 and the second device 520 have knowledge of reference pose P’ andhence reference pose P’ may or may not be transmitted in the combined bitstream b2. In some other implementations, both the first device 510 and the second device 520 have prior knowledge of poses D:'3F D:Gwhereas in some other implementations, poses D:'3F D:Gare transmitted in the bitstream b2. At the second device 520, the current user head pose P is obtained from the head-tracker 524, and the demultiplexer 521 receives bitstream b2 and generates b21 and b22. The bitstream b21 is decoded by the first decoder 522, which generates a decoded ^^^^^or ^^^^^^ . The bitstream b22is decoded by the second decoder 523, which generates the metadata M’ which includes a first reconstruction metadata ^^^at frequencies ^^, as shown in equation (5), such that0 ≤ ^^ < ^^, where ^^ is a constant ranging from 1000^^ ≤ ^^ ≤ 3000^^, and such that ^^^reconstructs both phase and magnitude of the binaural representations BINn from the reference binaural signal BIN’. M’ further includes a second reconstruction metadata ^^^at frequencies ^^,as shown in equation (7), such that ^^ ≤ ^^, where ^^ is a constant ranging from 1000^^ ≤ ^^ ≤3000^^, such that ^^^reconstructs only magnitude of the binaural representations BINn from the reference binaural signal BIN’. The computation of second reconstruction metadata ^^^is performed so that only the magnitude parameters are interpolated at high frequencies in the binaural reconstruction block 526, as shown in equation (9), which avoids comb filtering artifact in the reconstructed output, as shown in equation (10), as there is no interpolation of phase. Themetadata {M^ , ^^} are generated by the binaural reconstruction block 526 by interpolation ofthe metadata ^^^and ^^^according to Equation (9): M^LG ^^ = ∑^LM J^^^K , D^ = ∑ LG^LM J^^^K Equation (9)Docket No.: D23168WO01where: 0 ≤ < ≤ ^, J^ are the interpolation parameters which can be obtained with linearinterpolation based on the difference between poses D:'3F D:G, D′ and P. M^Pand D^Pare identity matrices.
[0086] The head tracked output binaural signal is generated by the binaural reconstruction block 526 according to Equation (10): ^^ ^^^^^ R ^^^^ 0 ≤ f < ^^^ = Q ^ Equation (10)f
[0087] signal BINp for pose P is then generated by the binaural reconstruction block 526 as an inverse transform from the filterbank domain to the time domain (e.g., “F-1{…}”) according to Equation (11): ^^^TT} Equation (11)
[0088] In another example implementation, the second device 520 computes the ILD pose correction metadata. For example, the first device 510 computes, with the metadata generator 515, pose correction metadata Mp which includes prediction parameters at all frequencies. The computation of pose correction metadata Mpi may be as provided by Equation (5). The pose correction metadata Mpi may be generated in a frequency domain or a filterbank domain (e.g., a CLDFB filterbank). For example, the first device 510 transmits ^^^^^or ^^^^^^ ,^ ^^>^, ^ ^^?^, … .. , { ^^B} and D′ to the second device 520. The second device 520 converts,using the binaural reconstruction block 526, the complex prediction matrix to an ILD pose correction matrix at high frequencies. For example, first, a post prediction matrix is computed by the binaural reconstruction block 526 at poses P21to P2Naccording to Equation (12): ^),^^, ^^ = ^^^ ^ ,^!,^! ^∗^^ Equation (12)where:
[0089] Then, at frequencies ^^, such that ^^ ≤ ^^, where ^^ is a constant with a value in a rangefrom 1000^^ ≤ ^^ ≤ 3000^^, the binaural reconstruction block 526 computes ILD matchinggain matrix ^^Kaccording to Equation (13): ;0^ = ^,^KDocket No.: D23168WO01 where: 5; 69,7 , 7 (','85 (:,:8^,^K = 0123( ^ ^56,7!,7! (','88 , ;.,^K = 0123(69,7^, 7^56,7!,7! (:,:8). sign^al ^^^^^^),^^, ^^is the post prediction matrix computed in Equation (12).
[0090] The binaural reconstruction block 526 obtains actual user head pose P from the head-tracker 524 and generates M^ , ^^ by interpolation of the pose correction metadata received fromthe first device 510 and the ILD pose correction matrix, according to Equation (14) and Equation (15): M^ = ∑^LG^LM J^^^K Equation (14)K Equation (15)where:J^are the interpolation parameters which can be obtained with linear interpolation based on the difference between poses D:'3F D:G, D′ and P, and M^Pand D^Pare identity matrices.
[0091] The head tracked output binaural signal is then generated by the binaural reconstruction block 526 according to Equation (16): ^^^^ ^ 0 ≤ f < ^^^^^^ = R ^^ ^^ Equation (16)
[0092] signal, BINp, corresponding to pose P, is generated by the binaural reconstruction block 526 according to Equation (17): ^^^T = ^&'{^^^^T} Equation (17)
[0093] In yet another example implementation, the first device 510 computes binaural representations corresponding to a first set of head poses including a reference head pose P’ and first set of head poses D''3F D'G. These N+1 binaural representations BINp’ and BINpi, where 1 ≤ i ≤ N and BINpi corresponds to pose D'(, are coded and transmitted to the second device 520. At the second device 520, the actual user head pose P is obtained from the head-tracker 524, and the N+1 binaural representations are converted by the binaural reconstruction block 526 into frequency domain (or filterbank domain) according to Equation (18):Docket No.: D23168WO01 ^^^^^^ = ^^^^^^^^, ^^^^^^ = ^{^^^^^} Equation (18)where: ^^^^domain representation of ^^^^^.
[0094] At frequencies ^^, such that 0 ≤ ^^ < ^^, where ^^ is a constant ranging from 1000^^ ≤^^ ≤ 3000^^, the head tracked binaural output, corresponding to pose D, is obtained by thebinaural reconstruction block 526 by interpolating ^^^^^ signals according to Equation (19): ^^^^^ = ∑^LG^LM J^^^^^^^ 0 ≤ f < ^^ Equation (19)where:parameters which can be obtained withlinear interpolation or spline interpolation based on the difference between poses D''3F D'G, D′ and D.
[0095] At frequencies ^^, such that f^ ≤ ^^, the ILD matching gain matrix ^^W is computed bythe binaural reconstruction block 526 according to Equation (20): ;^,^ 0^^W = U WV Equation (20)56,7where: ; = 0123( X, 7X (','8, ; = 01256,7X, 7X (:,:8^,^W .,^W 3(^,^P,to pose DM, DMis one of the first set of head poses that is closest to pose D, ^,^X, ^Xis the covariance matrix of binaural signal corresponding to pose DYDYis one of the first set of head poses excluding DM, and1 ≤ Z ≤ ^.
[0096] In some implementations, pose P0 is the same as pose P’. In some other implementations, the ILD matching gain matrix ^^Wis computed at the first device 510 and is transmitted to the second device 520.
[0097] The ILD matching gain matrix Dpcorresponding to the current user head pose P is computed by the binaural reconstruction block 526 by interpolation of the ILD pose correction matrices ^^W, according to Equation (21): ^= ∑YLG^ YLM JY^^W Equation (21)Docket No.: D23168WO01 where: 0 ≤ j ≤ N, JYare the interpolation parameters which can be obtained with linear interpolation based on the difference between poses D'3F DG, DMand P, and D^Mis an identity matrix.
[0098] At frequencies ^^, such that ^^ ≤ ^^, the head tracked output, corresponding to pose D, isobtained by the binaural reconstruction block 526 by interpolating ^^^^^ signals according to Equation (22): ^^^^^ = ^R^^^^^P f^ ≤ f Equation (22)domain binaural signal, ^^^^, corresponding to pose D, is generated by the binaural reconstruction block 526 according to Equation (23): ^^^ &' ^T = ^ {^^^T} Equation (23)
[0100] FIG.6 illustrates a block diagram of various example methods 600 for encoding an input audio signal with immersive audio content in a bitstream from a first device to a second device, which may be performed by the first device 510 of FIG.5. The methods 600 may be performed by one or more processors, which may be configured to perform methods 600 via machine- executable instructions. The methods 600 may be broken into various blocks or partitions, such as blocks 602, 604, 606, 608, 610, 612, 614, 616, and 618. The various process blocks illustrated in FIG.6 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 602.
[0101] At block 602, “Obtaining An Input Audio Signal With Immersive Audio Content,” an example method 600 may include obtaining an input audio signal with immersive audio content. For example, with reference to FIG.5. The decoder 511 receives a bitstream b1via the first path 551. The decoder 511 decodes the bitstream b1 to obtain immersive audio content A. As illustrated in FIG.5, the immersive audio content A may be provided by decoder 511 to downmixer 512 and binaural renderer 514 for further processing. Processing may proceed from block 602 to block 604.Docket No.: D23168WO01
[0102] At block 604, “Obtaining A First Set Of Head Poses,” an example method 600 may include obtaining a first set of head poses. For example, with reference to FIG.5, the head pose decoder 513 decodes pose bitstream bp to obtain a first set of head poses P1. The first set of head poses P1may include the reference pose P’. The first set of head poses may be provided by head pose decoder 513 to dowmixer 512 and binaural renderer 514 for further processing. Processing may proceed from block 604 to block 606.
[0103] At block 606, “Rendering A First Set Of Binaural Signals With The Immersive Audio Content Based On The First Set Of Head Poses,” an example method 600 may include rendering a first set of binaural signals with the immersive audio content based on the first set of head poses. For example, with reference to FIG.5, the downmixer 512 may be another binaural renderer configured to generate a first set of binaural signals BIN1 using the immersive audio content A and the first set of head poses P1. The first set of binaural signals BIN1may be provided by the downmixer 512 to the metadata generator 515 for further processing. Processing may proceed from block 606 to block 608.
[0104] At block 608, “Obtaining A Second Set Of Head Poses That Are Different From The First Set Of Head Poses,” an example method 600 may include obtaining a second set of head poses that are different from the first set of head poses. For example, with reference to FIG.5, the head pose decoder 513 decodes pose bitstream bp to obtain a second set of head poses P2. The second set of head poses may be provided by head pose decoder 513 to dowmixer 512 and binaural renderer 514 for further processing. Processing may proceed from block 608 to block 610.
[0105] At block 610, “Rendering A Second Set Of Binaural Signals With The Immersive Audio Content Based On The Second Set Of Head Poses,” an example method 600 may include rendering a second set of binaural signals with the immersive audio content based on the second set of head poses. For example, with reference to FIG.5, the binaural renderer 514 generates a second set of binaural signals BIN2 using the immersive audio content A and the second set of head poses P2. The second set of binaural signals BIN2 may be provided by binaural renderer 514 to metadata generator 515 for further processing. Processing may proceed from block 610 to block 612.
[0106] At block 612, “Computing, In A First Frequency Range, A First Reconstruction Metadata To Enable The Reconstruction Of Both Magnitude And Phase Of The Second Set Of Binaural Signals From The First Set Of Binaural Signals,” an example method 600 may include computing, in a first frequency range of a banded domain, a first reconstruction metadata toDocket No.: D23168WO01 enable the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals. For example, with reference to FIG.5, the metadata generator 515 generates prediction matrix ^^^as previously described with respect to Equations (1)-(8). The prediction matrix ^^^may be provided by the metadata generator 515 to the second encoder 517 for further processing. Processing may proceed from block 612 to block 614.
[0107] At block 614, “Computing, In A Second Frequency Range, A Second Reconstruction Metadata To Enable The Reconstruction Of The Magnitude Of The Second Set Of Binaural Signals From The First Set Of Binaural Signals,” an example method 600 may include computing, in a second frequency range of the banded domain, a second reconstruction metadata to enable the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals. For example, with reference to FIG.5, the metadata generator 515 generates ILD matching gain matrix Dpi, as previously described with respect to Equations (1)- (8). The ILD matching gain matrix Dpi may be provided by the metadata generator 515 to the second encoder 517 for further processing. Processing may proceed from block 614 to block 616.
[0108] At block 616, “Encoding The First Set Of Binaural Signals, The First Reconstruction Metadata, And The Second Reconstruction Metadata In A Bitstream,” an example method 600 may include encoding the first set of binaural signals, the first reconstruction metadata, and the second reconstruction metadata in a bitstream. For example, with reference to FIG.5, the multiplexer 518 combines the first set of binaural signals BIN1, the first reconstruction metadata ^^^, and the second reconstruction metadata Dpi into a combined bitstream b2. Processing may proceed from block 616 to block 618.
[0109] At block 618, “Outputting The Bitstream,” an example method 600 may include outputting the bitstream. For example, with reference to FIG.5, the first device 510 transmits the combined bitstream b2 from the output of the multiplexer 518 to the second device 520.
[0110] FIG.7 illustrates a block diagram of various example methods 700 for decoding an immersive audio content from a bitstream from a first device to a second device, which may be performed by the second device 520 of FIG.5. The methods 700 may be performed by one or more processors, which may be configured to perform methods 700 via machine-executable instructions. The methods 700 may be broken into various blocks or partitions, such as blocks 702, 704, 706, 708, 710, 712, 714, and 716. The various process blocks illustrated in FIG.7 provide examples of various methods disclosed herein, and it is understood that some blocksDocket No.: D23168WO01 may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 702.
[0111] At block 702, “Obtaining A Bitstream,” an example method 700 may include obtaining a bitstream. For example, with respect to FIG.5, the second device 520 receives the combined bitstream b2 from the first device 510. The bitstream b2 may be obtained by the demultiplexer 521 in device 520, which may in turn separate (e.g., demultiplex) the obtained bitstream b2into a first separated bitstream b21 and a second separated bitstream b22. Processing may proceed from block 702 to block 704.
[0112] At block 704, “Decoding The Bitstream To Obtain A First Set Of Binaural Signals, A First Reconstruction Metadata, And A Second Reconstructed Metadata,” an example method 700 may include decoding the bitstream to obtain a first set of binaural signals of the immersive audio content in a banded domain, a first reconstruction metadata corresponding to a first frequency range of the banded domain, and a second reconstruction metadata corresponding to a second frequency range of the banded domain. The first reconstruction metadata enables the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals. The second reconstruction metadata enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals. For example, with reference to FIG.5, the binaural reconstruction block 526 obtains first set of binaural signals ^^^', the first reconstruction metadata ^^^, and the second reconstruction metadata Dpi. The step of decoding the bitstream may further include decoding the first and seond separated bitstreams b21and b22with the corresponding first and second decoders 522 and 523 of FIG.5, which resposively may generate the decoded downmix signal Dmx’ and the decoded metadata M’, and which may be provided to the binaural reconstruction block 526. In some instances, the downmix signal Dmx’ is the first set of binaural signals ^^^', wherein the first set of binaural signal ^^^'correspond to the reference binaural signal ^^^^^and / or ^^^^^^ , and metadata M’ includes the first reconstruction metadata ^^^and the second reconstruction metadata Dpi. Processing may proceed from block 704 to706.
[0113] At block 706, “Obtaining A First Set Of Head Poses Associated With The First Set Of Binaural Signals,” an example method 700 may include obtaining a first set of head poses associated with the first set of binaural signals. For example, with reference to FIG.5, the second decoder 523 receives the second separated bitstream b22 and decodes the first set of head poses P1, which may be included within the decoded metadata M’. The binaural reconstructionDocket No.: D23168WO01 block 526 obtains the first set of head poses D'associated with first set of binaural signals BIN1from the second decoder 523. In another example, the first set of head poses P1 are known by the second device 520 (for example, stored in a memory) and are acquired by the binaural reconstruction block 526 from the memory. Processing may proceed from block 706 to block 708.
[0114] At block 708, “Obtaining A Second Set Of Head Poses Associated With The Second Set Of Binaural Signals,” an example method 700 may include obtaining a second set of head poses associated with the second set of binaural signals. For example, with reference to FIG.5, the second decoder 523 receives the second separated bitstream b22and decodes the second set of head poses P2, which may be included within the decoded metadata M’. The binaural reconstruction block 526 obtains the second set of head poses D:associated with second set of binaural signals BIN2from the second decoder 523. In another example, the second set of head poses P2 are known by the second device 520 (for example, stored in a memory) and are acquired by the binaural reconstruction block 526 from the memory. In some instances, the second set of head poses D:are obtained by the binaural reconstruction block 526 applying known offsets to the first set of head poses D'. Processing may proceed from block 708 to block 710.
[0115] At block 710, “Detecting A Current Head Pose Associated With A User Of The Second Device,” an example method 700 may include detecting a current head pose associated with a user of the second device. For example, with reference to FIG.5, the binaural reconstruction block 526 obtains the current head pose P from the head-tracker 524 via the fifteenth path 565. Processing may proceed from block 710 to block 712.
[0116] At block 712, “Reconstructing, In A First Frequency Range, A First Binaural Audio Based On The First Set Of Binaural Signals, The First Reconstruction Metadata, And A Relationship Between The Second Set Of Head Poses, The First Set Of Head Poses And The Current Head Pose,” an example method 700 may include reconstructing, in the first frequency range, a first binaural audio based on the first set of binaural signals, the reconstruction metadata, and a relationship between the second set of head poses, the first set of head poses and the current head pose. For example, with reference to FIG.5, the binaural reconstruction block 526 calculates the interpolated metadata M^by interpolation of the first reconstruction metadata ^^^obtained from the second decoder 523, as previously described with respect to Equation (9). The binaural reconstruction block 526 then generates the head tracked output binaural signal^^^^^ for frequency within a range of 0 ≤ f < ^^, as previously described with respect toDocket No.: D23168WO01 Equation (10) using the interpolated metadata ^^and the first set of binaural signals ^^^^^and / or ^^^^^^ obtained from the first decoder 522. Processing may proceed from block 712 to block 714.
[0117] At block 714, “Reconstructing, In A Second Frequency Range, A Second Binaural Audio Based On The First Set Of Binaural Signals, The Second Reconstruction Metadata And A Relationship Between The Second Set Of Head Poses, The First Set Of Head Poses And The Current Head Pose,” an example method 700 may include reconstructing, in the second frequency range, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata and a relationship between the second set of head poses, the first set of head poses and the current head pose. For example, with reference to FIG.5, the binaural reconstruction block 526 calculates the interpolated metadata D^by interpolation of the second reconstruction metadata ^^^obtained from the second decoder 523, as previously described with respect to Equation (9). The binaural reconstruction block 526 then generates the head trackedoutput binaural signal ^^^^^ for frequency within a range of f[ ≤ f, as previously described withrespect to Equation (10) using the interpolated metadata Dp and the first set of binaural signals ^^^^^and / or ^^^^^^ obtained from the first decoder 522. Processing may proceed from block 714 to block 716.
[0118] At block 716, “Generating A Binaural Output By Combining The First Binaural Audio In The First Frequency Range And The Second Binaural Audio In The Second Frequency Range,” an example method 700 may include generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range. For example, with respect to FIG.5, the binaural reconstruction block 526 generates the final binaural output BINout. In one example, the binaural reconstruction block 526 generates the final binaural output BINoutby performing an inverse transform from the filterbank domain to the time domain on the head tracked output binaural signal ^^^^^, as previously described with respect to Equation (11).
[0119] FIG.8 illustrates a block diagram of various example methods 800 for decoding an immersive audio content from a bitstream from a first device to a second device, which may be performed by the second device 520 of FIG.5. The methods 800 may be performed by one or more processors, which may be configured to perform methods 800 via machine-executable instructions. The methods 800 may be broken into various blocks or partitions, such as blocks 802, 804, 806, 808, 810, 812, 814, 816, and 818. The various process blocks illustrated in FIG.8Docket No.: D23168WO01 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 802.
[0120] At block 802, “Obtaining A Bitstream,” an example method 800 may include obtaining a bitstream. For example, with respect to FIG.5, the second device 520 receives the combined bitstream b2from the first device 510. The bitstream b2may be obtained by the demultiplexer 521 in device 520, which may in turn separate (e.g., demultiplex) the obtained bitstream b2 into a first separated bitstream b21and a second separated bitstream b22. Processing may proceed from block 802 to block 804.
[0121] At block 804, “Decoding The Bitstream To Obtain A First Set Of Binaural Signals And A First Reconstructed Metadata,” an example method 800 may include decoding the bitstream to obtain a first set of binaural signals of the immersive audio content in a banded domain and a first reconstruction metadata. The first reconstruction metadata enables the reconstruction of both magnitude and phase of a second set of binaural signals from the first set of binaural signals. For example, with reference to FIG.5, the binaural reconstruction block 526 obtains first binaural signals ^^^'and the first reconstruction metadata ^^^. The step of decoding the bitstream may further include decoding the first and seond separated bistreams b21 and b22 with the corresponding first and second decoders 522 and 523 of FIG.5, which resposively may generate the decoded downmix signal Dmx’ and the decoded metadata M’, and which may be provided to the binaural reconstruction block 526. In some instances, the downmix signal Dmx’ is the first set of binaural signals ^^^', wherein the first set of binaural signal ^^^'corresponds to reference binaural signal ^^^^^and / or ^^^^^^ , and metadata M’ includes the first reconstruction metadata ^^^. Processing may proceed from block 804 to block 806.
[0122] At block 806, “Obtaining A First Set Of Head Poses Associated With The First Set Of Binaural Signals,” an example method 800 may include obtaining a first set of head poses associated with the first set of binaural signals. For example, with reference to FIG.5, the second decoder 523 receives the second separated bitstream b22 and decodes the first set of head poses P1, which may be included within the decoded metadata M’. The binaural reconstruction block 526 obtains the first set of head poses D'associated with first set of binaural signals BIN1from the second decoder 523. In another example, the first set of head poses P1 are known by the second device 520 (for example, stored in a memory) and are acquired by the binauralDocket No.: D23168WO01 reconstruction block 526 from the memory. Processing may proceed from block 806 to block 808.
[0123] At block 808, “Obtaining A Second Set Of Head Poses Associated With The Second Set Of Binaural Signals,” an example method 800 may include obtaining a second set of head poses associated with the second set of binaural signals. For example, with reference to FIG.5, the second decoder 523 receives the second separated bitstream b22 and decodes the second set of head poses P2, which may be included within the decoded metadata M’. The binaural reconstruction block 526 obtains the second set of head poses D:associated with second set of binaural signals BIN2from the second decoder 523. In another example, the second set of head poses P2 are known by the second device 520 (for example, stored in a memory) and are acquired by the binaural reconstruction block 526 from the memory. In some instances, the second set of head poses D:are obtained by applying known offsets to the first set of head poses D'. Processing may proceed from block 808 to block 810.
[0124] At block 810, “Detecting A Current Head Pose Associated With A User Of The Second Device,” an example method 800 may include detecting a current head pose associated with a user of the second device. For example, with reference to FIG.5, the binaural reconstruction block 526 receives the current head pose P from the head-tracker 524 via the fifteenth path 565. Processing may proceed from block 810 to block 812.
[0125] At block 812, “Reconstructing, In A First Frequency Range, A First Binaural Audio Based On The First Set Of Binaural Signals, The First Reconstruction Metadata, And A Relationship Between The Second Set Of Head Poses, The First Set Of Head Poses, And The Current Head Pose,” an example method 800 may include reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata, and a relationship between the second set of head poses, the first set of head poses, and the current head pose. For example, with reference to FIG.5, the binaural reconstruction block 526 calculates the interpolated metadata M^by interpolation of the first reconstruction metadata ^^^obtained from the second decoder 523, as previously described with respect to Equation (14). The binaural reconstruction block 526 then generates the headtracked output binaural signal ^^^^^ for frequency within a range of 0 ≤ f < ^^, as previouslydescribed with respect to Equation (16) using the interpolated metadata ^^and the first set of binaural signals ^^^^^and / or ^^^^^^ obtained from the first decoder 522. Processing may proceed from block 812 to block 814.Docket No.: D23168WO01
[0126] At block 814, “Generating, In A Second Frequency Range, A Second Reconstruction Metadata From The First Reconstruction Metadata Based On The First Set Of Binaural Signals,” an example method 800 may include generating a second reconstruction metadata from the first reconstruction metadata based on the first set of binaural signals. For example, with reference to FIG.5, the binaural reconstruction block 526 computes a post prediction matrix using the second set of poses P21to P2N, as previously described with respect to Equation (12). Then, the binaural reconstruction block 526 computes a second reconstruction metadata as ILD matching gain matrix ^^K, as previously described with respect to Equation (13). Processing may proceed from block 814 to block 816.
[0127] At block 816, “Reconstructing, In The Second Frequency Range, A Second Binaural Audio Based On The First Set Of Binaural Signals, The Second Reconstruction Metadata, And A Relationship Between The Second Set Of Head Poses, The First Set Of Head Poses, And The Current Head Pose,” an example method 800 may include reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata, and a relationship between the second set of head poses, the first set of head poses, and the current head pose. For example, with reference to FIG. 5, the binaural reconstruction block 526 interpolates the second reconstruction metadata ^^Kto generate the interpolated metadata D^as previously described with respect to Equation (15) andgenerates a head tracked output binaural signal ^^^^^ for frequencies within the range f^ ≤ fusing the interpolated metadata ^^, as previously described with respect to Equation (16). Processing may proceed from block 816 to block 818.
[0128] At block 818, “Generating A Binaural Output By Combining The First Binaural Audio In The First Frequency Range And The Second Binaural Audio In The Second Frequency Range,” an example method 800 may include generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range. For example, with respect to FIG.5, the binaural reconstruction block 526 generates the final binaural output BINout. In one example, the binaural reconstruction block 526 generates the final binaural output BINout by performing an inverse transform from the filterbank domain to the time domain on the head tracked output binaural signal ^^^^^, as previously described with respect to Equation (17).
[0129] FIG.9 illustrates a block diagram of various example methods 900 for decoding an immersive audio content from a bitstream from a first device to a second device, which may beDocket No.: D23168WO01 performed by the second device 520 of FIG.5. The methods 900 may be performed by one or more processors, which may be configured to perform methods 900 via machine-executable instructions. The methods 900 may be broken into various blocks or partitions, such as blocks 902, 904, 906, 908, 910, 912, 914, 916, and 918. The various process blocks illustrated in FIG.9 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 902.
[0130] At block 902, “Obtaining A Bitstream,” an example method 900 may include obtaining a bitstream. For example, with respect to FIG.5, the second device 520 receives the combined bitstream b2 from the first device 510. The bitstream b2 may be obtained by the demultiplexer 521 in device 520, which may in turn separate (e.g., the obtained bitstream b2into afirst separated bitstream b21 and a second separated bitstream b22. Processing may proceed from block 902 to block 904.
[0131] At block 904, “Decoding The Bitstream To Obtain A First Set Of Binaural Signals Of Immersive Audio Content In A Banded Domain,” an example method 900 may include decoding the bitstream to obtain a first set of binaural signals of the immersive audio content in a banded domain. For example, with reference to FIG.5, the binaural reconstruction block 526 obtains first set of binaural signals ^^^^' . The step of decoding the bitstream may further include decoding the first and second separated bistreams b21 and b22 with the corresponding first and second decoders 522 and 523 of FIG.5, which responsively may generate the decoded downmix signal Dmx’ and the decoded metadata M’, and which may be provided to the binaural reconstruction block 526. In some instances, the downmix signal Dmx’ is the first set of binaural signals ^^^'and metadata M’ includes the first set of poses P1. Processing may proceed from block 904 to block 906.
[0132] At block 906, “Obtaining A First Set Of Head Poses Associated With The First Set Of Binaural Signals,” an example method 900 may include obtaining a first set of head poses associated with the first set of binaural signals. For example, with reference to FIG.5, the second decoder 523 receives the second separated bitstream b22 and decodes the first set of head poses P1, which may be included within the decoded metadata M’. The binaural reconstruction block 526 obtains the first set of head poses D'associated with first set of binaural signals BIN1 from the second decoder 523. In anotherthe first set of head poses P1are known by the second device 520 (for example, stored in a memory) and are acquired by the binauralDocket No.: D23168WO01 reconstruction block 526 from the memory. Processing may proceed from block 906 to block 908.
[0133] At block 908, “Detecting A Current Head Pose Associated With A User Of The Second Device,” an example method 900 may include detecting a current head pose associated with a user of the second device. For example, with respect to FIG.5, the binaural reconstruction block 526 receives the current head pose P from the head-tracker 524 via the fifteenth path 565. Processing may proceed from block 908 to block 910.
[0134] At block 910, “Reconstructing, In A First Frequency Range, A First Binaural Audio Based On The First Set Of Binaural Signals And A Relationship Between A First Set Of Head Poses And The Current Head Pose,” an example method 900 may include reconstructing a first binaural audio based on the first set of binaural signals and a relationship between a first set of head poses and the current head pose. For example, with reference to FIG.5, the binaural reconstruction block 526 converts N+1 binaural representations BINp’and BINpiinto the filterbank domain ^^^^^^ as previously described with respect to Equation (18). The N+1 binaural signals are then interpolated to generate the head tracked output binaural signal^^^^^ for frequency within a range of 0 ≤ f < ^^, as previously described with respect toEquation (19). Processing may proceed from block 910 to block 912.
[0135] At block 912, “Selecting A Reference Pose From The First Set Of Head Poses And A Corresponding Reference Binaural Signal From The First Set Of Binaural Signals,” an example method 900 may include selecting a reference pose from the first set of head poses. For example, with reference to FIG.5, the binaural reconstruction block 526 selects a reference pose P’ from first set of head poses P1 that may be associated with the first binaural audio BIN1’. Processing may proceed from block 912 to block 914.
[0136] At block 914, “Generating, In A Second Frequency Range, A First Reconstruction Metadata From The First Set Of Binaural Signals,” an example method may include generating, in a second frequency range of the banded domain, a first reconstruction metadata from the first set of binaural signals. The first reconstruction metadata may enable the reconstruction of the magnitude of the first set of binaural signals from the reference binaural signal. For example, with reference to FIG.5, the binaural reconstruction block 526 computes first reconstructionmetadata (e.g., ILD matching gain matrix) ^^K for frequencies f^ ≤ ^^, as previously describedwith respect to Equations (20). Processing may proceed from block 914 to block 916.Docket No.: D23168WO01
[0137] At block 916, “Reconstructing, In The Second Frequency Range, A Second Binaural Audio Based On The First Set Of Binaural Signals, The First Reconstruction Metadata And a Relationship Between The First Set Of Head Poses And The Current Head Pose,” an example method 900 may include reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the first reconstruction metadata, and a relationship between the first set of head poses and the current head pose. For example, with reference to FIG.5, the binaural reconstruction block 526 interpolates the first reconstruction metadata ^^Kto generate an interpolated metadata ^^with respect to Equation (21) and then generates a head tracked output binaural signal ^^^^^for frequencies within therange ^^ ≤ ^^ using the interpolated metadata ^^, as previously described with respect toEquation (22). In some instances, the binaural reconstruction block 526 converts the head tracked output binaural signal ^^^^^to the time domain, as previously described with respect to Equation (23). Processing may proceed from block 916 to block 918.
[0138] At block 918, “Generating A Binaural Output By Combining The First Binaural Audio In The First Frequency Range And The Second Binaural Audio In The Second Frequency Range,” and example method 900 may include generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range. For example, with respect to FIG.5, the binaural reconstruction block 526 combines the head tracked output binaural siganls ^^^^^over both the first frequency range and the second frequency range to generate the final binaural output BINout.
[0139] FIG.10A illustrates a schematic block diagram of an example device architecture 1000 (e.g., an apparatus 1000) that may be used to implement various aspects of the present disclosure. Architecture 1000 includes but is not limited to servers and client devices, systems, and methods as described in reference to FIGS.1-9. As shown, the architecture 1000 includes central processing unit (CPU) 1001 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 1002 or a program loaded from, for example, storage unit 1008 to random access memory (RAM) 1003. The CPU 1001 may be, for example, an electronic processor 1001. In RAM 1003, the data required when CPU 1001 performs the various processes is also stored, as required. CPU 1001, ROM 1002, and RAM 1003 are connected to one another via bus 1004. Input / output interface 1005 is also connected to bus 1004.Docket No.: D23168WO01
[0140] The following components are connected to I / O interface 1005: input unit 1006, that may include a keyboard, a mouse, or the like; output unit 1007 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 1008 including a hard disk, or another suitable storage device; and communication unit 1009 including a network interface card such as a network card (e.g., wired or wireless).
[0141] In some implementations, input unit 1006 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0142] In some implementations, output unit 1007 include systems with various number of speakers. Output unit 1007 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0143] In some embodiments, communication unit 1009 is configured to communicate with other devices (e.g., via a network). Drive 1010 is also connected to I / O interface 1005, as required. Removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 1010, so that a computer program read therefrom is installed into storage unit 1008, as required. A person skilled in the art would understand that although apparatus 1000 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0144] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 1009, and / or installed from the removable medium 1011, as shown in FIG. 10A.
[0145] FIG.10B illustrates a schematic block diagram of an example CPU 1001 implemented in the device architecture 1000 of FIG.10A that may be used to implement various aspects of the present disclosure. The CPU 1001 includes an electronic processor 1020 and a memory 1021. The electronic processor 1020 is electrically and / or communicatively connected to the memoryDocket No.: D23168WO01 1021 for bidirectional communication. The memory 1021 stores an immersive audio coding software 1022 and an immersive audio decoding software 1023. In some examples, memory 1021 may be located internal to the electronic processor 1020, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 1021 may be located external to the electronic processor 1020, such as in a ROM 1002, a RAM 1003, flash memory or a removable medium 1011, or another non-transitory computer readable medium that is contemplated for device architecture 1000. In some instances, the electronic processor 1020 may implement the immersive audio coding software 1022 stored in the memory 1021 to perform, among other things, any of the methods 600 of FIG.6. In some instances, the electronic processor 1020 may implement the immersive audio decoding software 1023 stored in the memory 1021 to perform, among other things, any of the methods 700 of FIG. 7, the methods 800 of FIG.8, and / or the methods 900 of FIG.9.
[0146] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 1001 in combination with other components of FIG.10A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0147] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0148] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instructionDocket No.: D23168WO01 execution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0150] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.
[0151] EEE1. A method to encode an input audio signal with immersive audio content in a bitstream from a first device to a second device, the method for the first device comprising: obtaining the input audio signal with immersive audio content; obtaining a first set of head poses; rendering a first set of binaural signals with the immersive audio content based on the first set of head poses; obtaining a second set of head-poses that are different from the first set of head poses; rendering a second set of binaural signals with the immersive audio content based on the second set of head poses; computing, in a first frequency range of a banded domain, a firstDocket No.: D23168WO01 reconstruction metadata to enable the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals; computing, in a second frequency range of the banded domain, a second reconstruction metadata to enable the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; encoding the first set of binaural signals, the first reconstruction metadata and the second reconstruction metadata in the bitstream; and outputting the bitstream.
[0152] EEE2. The method of EEE1, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
[0153] EEE3. The method of EEE2, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
[0154] EEE4. The method of any one of EEE1 to EEE3, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.
[0155] EEE5. The method of any one of EEE1 to EEE4, further comprising converting the first set of binaural signals and the second set of binaural signals from a time domain into the banded domain.
[0156] EEE6. The method of EEE5, wherein converting the first set of binaural signals and the second set of binaural signals from the time domain into the banded domain comprises: processing the first set of binaural signals and the second set of binaural signals in the time domain to a current frame with a complex low delay filter bank (CLDFB) that includes a plurality of subbands.
[0157] EEE7. The method of EEE5, wherein converting the first set of binaural signals and the second set of binaural signals in the time domain into the banded domain comprises either applying a modified discrete Fourier transform (MDFT) to the first set of binaural signals and the second set of binaural signals in the time domain, or applying a modified discrete cosine transform (MDCT) to the first set of binaural signals and the second set of binaural signals in the time domain.
[0158] EEE8. The method of any one of EEE1 to EEE7, wherein the first set of head poses includes a reference pose.Docket No.: D23168WO01
[0159] EEE9. The method of EEE8, wherein the reference pose corresponds to a current pose obtained from a back channel bitstream (555) from the second device (520) to the first device (510).
[0160] EEE10. The method of EEE8, wherein the first set of head poses further includes a third set of head poses that are obtained by applying offsets to the reference pose.
[0161] EEE11. The method of any one of EEE1 to EEE10, wherein the second set of head poses are obtained by applying offsets to the first set of head poses.
[0162] EEE12. The method of EEE11, wherein the second set of head poses are obtained by applying offsets to the first set of head poses along one rotational axis.
[0163] EEE13. The method of EEE12, wherein the value of the offsets is fixed.
[0164] EEE14. The method of EEE12, wherein the value of the offsets is dynamic.
[0165] EEE15. The method any one of EEE1 to EEE14, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.
[0166] EEE16. A method to decode immersive audio content from a bitstream from a first device to a second device, the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain: a first set of binaural signals of the immersive audio content in a banded domain; a first reconstruction metadata corresponding to a first frequency range of a banded domain, wherein the first reconstruction metadata enables the reconstruction of both magnitude and phase of a second set of binaural signals from the first set of binaural signals, wherein the second set of binaural signals correspond to a second set of head poses such that the second set of head poses are different than the first set of head poses; and a second reconstruction metadata corresponding to a second frequency range of the banded domain, wherein the second reconstruction metadata enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; obtaining a first set of head poses associated with the first set of binaural signals; obtaining the second set of head poses associated with the second set of binaural signals; detecting a current head pose associated with a user of the second device; reconstructing, in the first frequency range, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata, and a relationship between the second set of head poses, the first set of head poses, and the current head pose; reconstructing, in the second frequency range, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata and a relationship between the second set ofDocket No.: D23168WO01 head poses, the first set of head poses, and the current head pose; and generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range.
[0167] EEE17. The method of EEE16, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
[0168] EEE18. The method of EEE17, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
[0169] EEE19. The method of any one of EEE16 to EEE18, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.
[0170] EEE20. The method of any one of EEE16 to EEE19, wherein obtaining the bitstream comprises receiving and storing the bitstream.
[0171] EEE21. The method of any one of EEE16 to EEE20, wherein the banded domain comprises one of: a complex low delay filter bank (CLDFB) domain; a modified discrete Fourier transform (MDFT) domain; and a modified discrete cosine transform (MDCT) domain.
[0172] EEE22. The method of any one of EEE16 to EEE21, wherein: the first set of head poses corresponds to a set of reference poses for the first set of binaural signals, the second set of head poses corresponds to a set of reference poses for the second set of binaural signals; and correction metadata for the current pose is linearly interpolated or extrapolated based on a difference in first set of head poses, the second set of head poses and the current head pose.
[0173] EEE23. The method of any one of EEE16 to EEE22, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.
[0174] EEE24. The method of any one of EEE16 to EEE23, wherein the first set of head poses includes a reference head pose.
[0175] EEE25. The method of EEE24, wherein the reference pose corresponds to a delayed pose sent via a back channel bitstream from the second device to the first device.
[0176] EEE26. The method of any one of EEE16 to EEE25, wherein the second set of head poses are obtained by applying offsets to the first set of head poses.Docket No.: D23168WO01
[0177] EEE27. The method of EEE26, wherein the second set of head poses are obtained by applying offsets to the first set of head poses along one rotational axis.
[0178] EEE28. The method of any one of EEE26 to EEE27, wherein the value of the offsets is fixed.
[0179] EEE29. The method of any one of EEE26 to EEE27, wherein the value of the offsets is dynamic and is obtained from the bitstream.
[0180] EEE30. A method to decode immersive audio content from a bitstream from a first device to a second device, the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain: a first set of binaural signals of the immersive audio content in a banded domain; and a first reconstruction metadata, wherein the first reconstruction metadata enables the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals; obtaining a first set of head poses associated with the first set of binaural signals; obtaining a second set of head poses associated with a second set of binaural signals, wherein the second set of head poses are different than the first set of head poses; detecting a current head pose associated with a user of the second device; reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata, and a relationship between the second set of head poses, the first set of head poses, and the current head pose; generating, in a second frequency range of the banded domain, a second reconstruction metadata from the first reconstruction metadata based on the first set of binaural signals, wherein the second reconstruction metadata enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata and a relationship between the second set of head poses, the first set of head poses and the current head pose; and generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range.
[0181] EEE31. The method of EEE30, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.Docket No.: D23168WO01
[0182] EEE32. The method of EEE31, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
[0183] EEE33. The method of any one of EEE30 to EEE32, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.
[0184] EEE34. The method of any one of EEE30 to EEE33, wherein obtaining the bitstream comprises receiving and storing the bitstream.
[0185] EEE35. The method of any one of EEE30 to EEE34, wherein the banded domain comprises one of: a complex low delay filter bank (CLDFB) domain; a modified discrete Fourier transform (MDFT) domain; and a modified discrete cosine transform (MDCT) domain.
[0186] EEE36. The method of any one of EEE30 to EEE35, wherein: the first set of head poses corresponds to a set of reference poses for the first set of binaural signals, the second set of head poses corresponds to a set of refence poses for the second set of binaural signals; and correction metadata for the current pose is linearly interpolated or extrapolated based on a difference in first set of head poses, the second set of head poses and the current head pose.
[0187] EEE37. The method of any one of EEE30 to EEE36, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.
[0188] EEE38. A method to decode immersive audio content from a bitstream from a first device to a second device , the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain a first set of binaural signals of the immersive audio content in a banded domain; obtaining a first set of head poses associated with the first set of binaural signals; detecting a current head pose associated with a user of the second device; reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, and a relationship between the first set of head poses and the current head pose; selecting a reference pose from the first set of head poses and a corresponding reference binaural signal from the first set of binaural signals; generating, in a second frequency range of the banded domain, a first reconstruction metadata from the first set of binaural signals, wherein the first reconstruction metadata enables the reconstruction of a magnitude of the first set of binaural signals from the reference binaural signal; reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the first reconstruction metadata and a relationship between the first set of head poses and theDocket No.: D23168WO01 current head pose; and generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range.
[0189] EEE39. The method of EEE38, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
[0190] EEE40. The method of EEE39, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
[0191] EEE41. The method of any one of EEE38 to EEE40, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.
[0192] EEE42. The method of any one of EEE38 to EEE41, wherein obtaining the bitstream comprises receiving and storing the bitstream.
[0193] EEE43. The method of any one of EEE38 to EEE42, wherein the banded domain comprises one of: a complex low delay filter bank (CLDFB) domain; a modified discrete Fourier transform (MDFT) domain; and a modified discrete cosine transform (MDCT) domain.
[0194] EEE44. The method of any one of EEE38 to EEE43, wherein the bitstream is an IVAS encoded bitstream.
[0195] EEE45. An apparatus comprising: an electronic processor configured to perform operations including the method of any one of EEE1 to EEE44.
[0196] EEE46. A non-transitory computer-readable storage medium recording a program of instructions that is executable by a device to perform the method of any one of EEE1 to EEE44.
[0197] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.Docket No.: D23168WO01
[0198] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0199] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0200] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
Claims
Docket No.: D23168WO01 CLAIMS What is claimed is:
1. A method (600) to encode an input audio signal with immersive audio content in a bitstream from a first device (510) to a second device (520), the method (600) for the first device comprising: obtaining the input audio signal with immersive audio content; obtaining a first set of head poses; rendering a first set of binaural signals with the immersive audio content based on the first set of head poses; obtaining a second set of head poses that are different from the first set of head poses; rendering a second set of binaural signals with the immersive audio content based on the second set of head poses; computing, in a first frequency range of a banded domain, a first reconstruction metadata (M) to enable the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals; computing, in a second frequency range of the banded domain, a second reconstruction metadata (D) to enable the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; encoding the first set of binaural signals, the first reconstruction metadata (M) and the second reconstruction metadata (D) in the bitstream; and outputting the bitstream.
2. The method (600) of claim 1, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
3. The method (600) of claim 2, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
4. The method (600) of claim 1, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.Docket No.: D23168WO01 5. The method (600) of claim 1, further comprising converting the first set of binaural signals and the second set of binaural signals from a time domain into the banded domain.
6. The method (600) of claim 5, wherein converting the first set of binaural signals and the second set of binaural signals from the time domain into the banded domain comprises: processing the first set of binaural signals and the second set of binaural signals in the time domain to a current frame with a complex low delay filter bank (CLDFB) that includes a plurality of subbands.
7. The method (600) of claim 5, wherein converting the first set of binaural signals and the second set of binaural signals in the time domain into the banded domain comprises either applying a modified discrete Fourier transform (MDFT) to the first set of binaural signals and the second set of binaural signals in the time domain, or applying a modified discrete cosine transform (MDCT) to the first set of binaural signals and the second set of binaural signals in the time domain.
8. The method (600) of claim 1, wherein the first set of head poses includes a reference pose.
9. The method (600) of claim 8, wherein the reference pose corresponds to a current pose obtained from a back channel bitstream (555) from the second device (520) to the first device (510).
10. The method (600) of claim 8, wherein the first set of head poses further includes a third set of head poses that are obtained by applying offsets to the reference pose.
11. The method (600) of claim 1, wherein the second set of head poses are obtained by applying offsets to the first set of head poses.
12. The method (600) of claim 11, wherein the second set of head poses are obtained by applying offsets to the first set of head poses along one rotational axis.
13. The method (600) of claim 12, wherein the value of the offsets is fixed.
14. The method (600) of claim 12, wherein the value of the offsets is dynamic.Docket No.: D23168WO01 15. The method (600) of claim 1, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.
16. A method (700) to decode immersive audio content from a bitstream from a first device (510) to a second device (520), the method (700) for the second device (520) comprising: obtaining the bitstream; decoding the bitstream to obtain: a first set of binaural signals of the immersive audio content in a banded domain; a first reconstruction metadata (M) corresponding to a first frequency range of a banded domain, wherein the first reconstruction metadata (M) enables the reconstruction of both magnitude and phase of a second set of binaural signals from the first set of binaural signals, wherein the second set of binaural signals correspond to a second set of head poses such that the second set of head poses are different than the first set of head poses; and a second reconstruction metadata (D) corresponding to a second frequency range of the banded domain, wherein the second reconstruction metadata (D) enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; obtaining a first set of head poses associated with the first set of binaural signals; obtaining the second set of head poses associated with the second set of binaural signals; detecting a current head pose (P) associated with a user of the second device; reconstructing, in the first frequency range, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata (M), and a relationship between the second set of head poses, the first set of head poses, and the current head pose (P); reconstructing, in the second frequency range, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata (D) and a relationship between the second set of head poses, the first set of head poses, and the current head pose (P); and generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range.
17. The method (700) of claim 16, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the secondDocket No.: D23168WO01 frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
18. The method (700) of claim 17, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
19. The method (700) of claim 16, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.
20. The method (700) of claim 16, wherein obtaining the bitstream comprises receiving and storing the bitstream.
21. The method (700) of claim 16, wherein the banded domain comprises one of: a complex low delay filter bank (CLDFB) domain; a modified discrete Fourier transform (MDFT) domain; and a modified discrete cosine transform (MDCT) domain.
22. The method (700) of claim 16, wherein: the first set of head poses corresponds to a set of reference poses for the first set of binaural signals, the second set of head poses corresponds to a set of reference poses for the second set of binaural signals; and correction metadata for the current pose is linearly interpolated or extrapolated based on a difference in first set of head poses, the second set of head poses and the current head pose.
23. The method (700) of claim 16, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.
24. The method (700) of claim 16, wherein the first set of head poses includes a reference head pose.
25. The method (700) of claim 24, wherein the reference pose corresponds to a delayed pose sent via a back channel bitstream (555) from the second device (520) to the first device (510).Docket No.: D23168WO01 26. The method (700) of claim 16, wherein the second set of head poses are obtained by applying offsets to the first set of head poses.
27. The method (700) of claim 26, wherein the second set of head poses are obtained by applying offsets to the first set of head poses along one rotational axis.
28. The method (700) of claim 26, wherein the value of the offsets is fixed.
29. The method (700) of claim 26, wherein the value of the offsets is dynamic and is obtained from the bitstream.
30. A method (800) to decode immersive audio content from a bitstream from a first device (510) to a second device (520), the method (800) for the second device (520) comprising: obtaining the bitstream; decoding the bitstream to obtain: a first set of binaural signals of the immersive audio content in a banded domain; and a first reconstruction metadata (M), wherein the first reconstruction metadata (M) enables the reconstruction of both magnitude and phase of the second set of binaural signals from the first set of binaural signals; obtaining a first set of head poses associated with the first set of binaural signals; obtaining a second set of head poses associated with a second set of binaural signals, wherein the second set of head poses are different than the first set of head poses; detecting a current head pose (P) associated with a user of the second device; reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, the first reconstruction metadata (M), and a relationship between the second set of head poses, the first set of head poses, and the current head pose (P); generating, in a second frequency range of the banded domain, a second reconstruction metadata (D) from the first reconstruction metadata (M) based on the first set of binaural signals, wherein the second reconstruction metadata (D) enables the reconstruction of the magnitude of the second set of binaural signals from the first set of binaural signals; reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the second reconstruction metadata (D) and a relationship between the second set of head poses, the first set of head poses and the current head pose (P); andDocket No.: D23168WO01 generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range.
31. The method (800) of claim 30, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
32. The method (800) of claim 31, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
33. The method (800) of claim 30, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.
34. The method (800) of claim 30, wherein obtaining the bitstream comprises receiving and storing the bitstream.
35. The method (800) of claim 30, wherein the banded domain comprises one of: a complex low delay filter bank (CLDFB) domain; a modified discrete Fourier transform (MDFT) domain; and a modified discrete cosine transform (MDCT) domain.
36. The method (800) of claim 30, wherein: the first set of head poses corresponds to a set of reference poses for the first set of binaural signals, the second set of head poses corresponds to a set of refence poses for the second set of binaural signals; and correction metadata for the current pose is linearly interpolated or extrapolated based on a difference in first set of head poses, the second set of head poses and the current head pose.
37. The method (800) of claim 30, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.Docket No.: D23168WO01 38. A method (900) to decode immersive audio content from a bitstream from a first device (510) to a second device (520), the method for the second device comprising: obtaining the bitstream; decoding the bitstream to obtain a first set of binaural signals of the immersive audio content in a banded domain; obtaining a first set of head poses associated with the first set of binaural signals; detecting a current head pose (P) associated with a user of the second device; reconstructing, in a first frequency range of the banded domain, a first binaural audio based on the first set of binaural signals, and a relationship between the first set of head poses and the current head pose (P); selecting a reference pose from the first set of head poses and a corresponding reference binaural signal from the first set of binaural signals; generating, in a second frequency range of the banded domain, a first reconstruction metadata (D) from the first set of binaural signals, wherein the first reconstruction metadata (D) enables the reconstruction of a magnitude of the first set of binaural signals from the reference binaural signal; reconstructing, in the second frequency range of the banded domain, a second binaural audio based on the first set of binaural signals, the first reconstruction metadata (D) and a relationship between the first set of head poses and the current head pose (P); and generating a binaural output by combining the first binaural audio in the first frequency range and the second binaural audio in the second frequency range.
39. The method (900) of claim 38, wherein the first frequency range of the banded domain corresponds to frequency values below an upper frequency threshold, and wherein the second frequency range of the banded domain corresponds to frequency values above the upper frequency threshold.
40. The method (900) of claim 39, wherein the upper frequency threshold is a constant value that is in a range of about 1000Hz to about 3000Hz.
41. The method (900) of claim 38, wherein the banded domain has a bandwidth of about 24 kHz, sampled at 48 kHz, which is divided into 60 bands with a resolution of about 400 Hz per band.Docket No.: D23168WO01 42. The method (900) of claim 38, wherein obtaining the bitstream comprises receiving and storing the bitstream.
43. The method (900) of claim 38, wherein the banded domain comprises one of: a complex low delay filter bank (CLDFB) domain; a modified discrete Fourier transform (MDFT) domain; and a modified discrete cosine transform (MDCT) domain.
44. The method (900) of claim 38, wherein the bitstream is an IVAS encoded bitstream or an ISAR encoded bitstream.
45. An apparatus (1000) comprising: an electronic processor (1020) configured to perform operations including the method of any one of claims 1-44.
46. A non-transitory computer-readable storage medium (1021) recording a program of instructions (1022) that is executable by a device (1000) to perform the method of any one of claims 1-44.