Method and apparatus for encoding or decoding scene-based immersive audio content - Patents.com

JP2024543189A5Pending Publication Date: 2025-12-12DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024532160
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2022-11-30
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing coding methods for Ambisonics audio signals, such as SPAR and DirAC, face inefficiencies in terms of bitrate requirements, complexity, and quality at high bitrates, with SPAR lacking in reconstructing higher order signals and DirAC saturating at high bitrates, leading to suboptimal performance.

Method used

A combined coding scheme that integrates Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) systems, where SPAR downmix channels are waveform encoded and parametrically encoded, with DirAC analysis performed on intermediate Ambisonics signals to enhance spatial resolution and quality, particularly at lower bitrates.

Benefits of technology

The combined scheme efficiently encodes and decodes Ambisonics signals, providing high-quality output with reduced bitrate demands and complexity, enabling flexible rendering options like binaural or multi-speaker audio, while maintaining compatibility with standalone SPAR and DirAC codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This specification describes a method (500) for encoding an Ambisonics input audio signal (101). The method (500) includes the step of providing (501) the input audio signal (101) to a SPAR encoder (110, 130) and a DirAC analyzer and parameter encoder (120). The method (500) further includes the step of generating (502) an encoder bitstream (106) based on the output (102, 105) of the SPAR encoder (110, 130) and the output (104) of the DirAC analyzer and parameter encoder (120).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 284,198, filed November 30, 2021, and U.S. Provisional Application No. 63 / 410,587, filed September 27, 2022.

[0002] [Technical field] The present specification relates to a method and corresponding apparatus for processing audio, in particular for coding immersive audio content. [Background technology]

[0003] A sound or sound field in a listening environment of a listener located at a listening position can be described using an Ambisonics audio signal, specifically a first order Ambisonics signal (FOA) or a higher order Ambisonics signal (HOA). An Ambisonics signal can be considered as a multi-channel audio signal where each channel corresponds to a particular directional pattern of the sound field at the listener's listening position. An Ambisonics signal can be described using a three-dimensional (3D) Cartesian coordinate system, where the origin of the coordinate system corresponds to the listening position, the x-axis points forward, the y-axis points left, and the z-axis points up.

[0004] This specification addresses the technical problem of enabling a particularly efficient and flexible coding of Ambisonics audio signals. The technical problem is solved by each of the independent claims. Preferred examples are set out in the dependent claims. Summary of the Invention

[0005] According to one aspect, a method for encoding an Ambisonics input audio signal is described. The method includes providing the input audio signal to a spatial reconstruction (SPAR) encoder and a directional audio coding (DirAC) analyzer and parameter encoder. The method further includes generating an encoder bitstream based on an output of the SPAR encoder and an output of the DirAC analyzer and parameter encoder.

[0006] According to another aspect, a method for decoding an encoder bitstream representing an Ambisonics input audio signal is described. The method includes generating an intermediate Ambisonics signal using a spatial reconstruction (SPAR) decoder based on the encoder bitstream. Further, the method includes processing the intermediate Ambisonics signal using a directional audio coding (DirAC) synthesizer to provide an output audio signal for rendering.

[0007] It should be noted that each of the methods described herein may be implemented, in whole or in part, in software and / or computer readable code on one or more processors.

[0008] According to a further aspect, a software program is described, the software program adapted for execution on a processor and, when executed on the processor, to perform the steps of the methods described herein.

[0009] According to another aspect, a storage medium is described that includes a software program adapted for execution on a processor and, when executed on the processor, to perform the steps of the methods described herein.

[0010] According to a further aspect, a computer program product is described. The computer program may include executable instructions for performing the steps of the methods described herein when executed on a computer.

[0011] According to another aspect, a system is described that includes one or more processors, the system including a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform one or more operations of the methods described herein.

[0012] According to a further aspect, a non-transitory computer-readable medium is described that stores instructions that, when executed by one or more processors, cause the one or more processors to perform one or more operations of the methods described herein.

[0013] According to another aspect, a coding device for coding an Ambisonics input audio signal is described. The coding device is configured to provide the input audio signal to a spatial reconstruction (SPAR) encoder and a directional audio coding (DirAC) analyzer and parameter encoder. The coding device is further configured to generate an encoder bitstream based on an output of the SPAR encoder and an output of the DirAC analyzer and parameter encoder.

[0014] According to a further aspect, a decoding device is described for decoding an encoder bitstream representing an Ambisonics input audio signal. The decoding device is configured to generate an intermediate Ambisonics signal based on the encoder bitstream using a spatial reconstruction (SPAR) decoder. Furthermore, the decoding device is configured to process the intermediate Ambisonics signal using a directional audio coding (DirAC) synthesizer to provide an output audio signal for rendering.

[0015] It should be noted that the methods and systems, including the preferred embodiments, outlined in this patent application can be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application can be combined in any manner. In particular, the features of the claims can be combined with each other in any manner. [Brief description of the drawings]

[0016] The invention will now be described, by way of example only, with reference to the accompanying drawings, in which: [Figure 1] 1 illustrates an exemplary audio encoder. [Diagram 2] 1 illustrates an exemplary audio decoder. [Figure 3a] 1 illustrates an exemplary audio encoder. [Figure 3b] 1 illustrates an exemplary audio decoder. [Figure 4] 1 illustrates an exemplary audio encoder. [Figure 5a] 1 shows a flowchart of an example of a method for encoding an Ambisonics audio signal. [Figure 5b] 1 shows a flowchart of an example of a method for decoding a bitstream representing an Ambisonics audio signal. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] As mentioned above, this specification relates to efficient and flexible coding of Ambisonics audio signals. An example of a coding scheme for Ambisonics audio signals is the so-called SPAR (Spatial Reconstruction) scheme, which is described, for example, in McGrath et al., "Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec," ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 730-734, doi: 10.1109 / ICASSP.2019.8683712, the entire contents of which are incorporated herein by reference. A further coding scheme is the so-called directional audio coding (DirAC) scheme, which is described, for example, in Ahonen, Jukka, et al. "Directional analysis of sound field with linear microphone array and applications in sound reproduction." Audio Engineering Society Convention 124. Audio Engineering Society, 2008, and / or in V. Pulkki et al, "Directional audio coding - perception-based reproduction of spatial sound", International Workshop on the principles and applications of spatial hearing, Nov. 11-13, 2009, Zao, Miyagi, Japan, the contents of which are incorporated herein by reference in their entirety.

[0018] In SPAR, an Ambisonics (FOA or HOA) audio signal may be spatially processed during downmix such that one or more downmix channels are waveform coded and some channels are parametrically coded based on metadata determined by the SPAR encoder. A SPAR decoder performs the inverse operation in that it upmixes one or more received (and decoded) downmix channels with the help of SPAR metadata and reconstructs the original Ambisonics channel. SPAR typically operates on multiple different time / frequency (T / F) tiles.

[0019] Directional Audio Coding (DirAC) is a parametric coding method based on the direction of arrival (DoA) and the diffusivity per T / F tile (i.e., for multiple different T / F tiles). DirAC is generally independent of the input audio format, but can be used with Ambisonics audio. This means that the DirAC parameter analysis can be based on an Ambisonics (FOA or HOA) input audio signal, and the DirAC decoder can reconstruct the Ambisonics signal. A property of DirAC is that it can be adapted to directly generate a binaurally rendered output signal based on a number of received signals (transport channels) and DirAC metadata. The DirAC metadata generation can be partially or fully present at the decoding end, for example operating on the transmitted and received FOA or HOA transport channels (as outlined herein). Furthermore, DirAC can be used to restore Ambisonics audio with a higher order than the original input to the coding system, and thus can be used to improve, for example, the spatial resolution of the output signal compared to the spatial resolution of the input audio signal.

[0020] SPAR is an efficient coding method, which allows Ambisonics signals to be stored and / or transmitted at a relatively low bit rate. For high quality requirements, SPAR can be used to efficiently represent HOA signals (e.g., of order L=3 or higher) while using only a relatively small number of downmix channel signals (e.g., 4 or less). However, SPAR does not provide a solution for recovering and / or generating Ambisonics output signals with increased Ambisonics orders from lower order Ambisonics input audio signals. For example, if the input audio signal is FOA (L=1), it is usually not possible to recover and / or generate HOA2 (L=2) or HOA3 (L=3) signals. In that case, SPAR can usually only reconstruct a relatively high quality FOA input audio signal at a given bit rate.

[0021] DirAC is an efficient coding method with strengths that vary depending on the requirements of the coding system. For example, when the requirement is to restore an input Ambisonics signal of a given order L (FOA or HOA) with the highest possible fidelity after decoding, it has been observed that the coding efficiency of DirAC is generally inferior to SPAR. It has also been observed that the quality of DirAC audio reconstruction saturates at relatively high bit rates, and DirAC does not provide a native solution for obtaining transparent audio quality at relatively high bit rates. To address this issue, coding systems that rely on DirAC may transmit all input channels (e.g., 4 for FOA) as transport channels (at relatively high bit rates) and disable parametric reconstruction using DirAC. This impacts the efficiency of DirAC (compared to SPAR) and leads to a relatively high complexity (in terms of numerical and memory resources) due to the need to code a relatively large number of transport channels using waveform coding (compared to the number of downmix channels coded in a SPAR coding system).

[0022] This document describes a coding scheme that combines the advantages of SPAR and DirAC coding systems in an optimal way. SPAR and DirAC may be combined such that a combined decoder reconstructs a first set of Ambisonics upmix signals based on the received and decoded SPAR downmix signals (using a SPAR decoder) and the SPAR metadata. The reconstructed SPAR upmix signals (herein referred to as intermediate Ambisonics signals) may then be fed to a DirAC decoder for operating on the set of SPAR upmix signals using the DirAC metadata (e.g. to generate an output signal with an increased Ambisonics order).

[0023] Figure 1 shows an exemplary encoding device (also called an encoder or coding unit) 100, and Figures 2 and 4 show an exemplary decoding device (also called a decoder or decoding unit) 200. SPAR and DirAC operate in a parallel structure in the encoding device 100 and in a serial structure in the decoding device 200. By way of example (and not by way of limitation), it can be assumed that the input audio signal 101 is a FOA signal and that the codec (coding / decoding) system 100, 200 operates on frames of 20 ms length.

[0024] In the encoding device 100, frames of an input audio signal 101 can be fed to a SPAR encoder 110, 130 (which may include a downmix unit 110 and a core audio encoder 130) and to an optional DirAC analyzer and parameter encoder 120. Each of these units 110, 120, 130 generates a respective corresponding partial bitstream 102, 104, 105. The SPAR encoders 110, 130, in particular the downmix unit 110 of the SPAR encoders 110, 130, generate a SPAR metadata bitstream (or SPAR metadata) 102 and a set of (one or more) SPAR downmix channel signals 103. The one or more SPAR downmix channel signals 103 are fed to the core audio encoder 130, which is configured to represent these signals 103 using the core audio bitstream 105.

[0025] The core audio encoder 130 of the SPAR encoder 110, 130 is configured to perform waveform encoding of one or more downmix channel signals 103, thereby providing a core audio bitstream 105. Each of the downmix channel signals 103 can be encoded using a mono waveform encoder (e.g., 3GPP EVS encoding), thereby enabling efficient encoding. Further examples for encoding one or more downmix channel signals 103 are MPEG AAC, MPEG HE-AAC and other MPEG Audio codecs, 3GPP codecs, and Dolby Digital / Dolby Digital Plus (AC-3, eAC-3). It is worth noting that both SPAR and DirAC are spatial audio coding frameworks that can work with a variety of different core audio codecs that represent downmix or transport channels, respectively. SPAR and DirAC represent spatial audio information by respective SPAR or DirAC metadata.

[0026] An optional DirAC analyzer and metadata encoder 120 generates an optional DirAC metadata bitstream (or DirAC metadata) 104. In contrast to conventional DirAC encoders for FOA, the encoding device 100 does not include a DirAC transport channel generator or a downmixer (as this information is provided by the SPAR encoders 110, 130). The partial bitstreams 102, 104, 105 can be multiplexed (in a multiplexing unit 140) into a common encoder bitstream 106 and transmitted to the decoding device 200.

[0027] In the decoding device 200, the received (encoder) bitstream 106 can be demultiplexed (in a demultiplexing unit 240) into partial bitstreams 102, 104, 105, in particular the SPAR metadata bitstream 102, the core audio bitstream 105 and the (optional) DirAC metadata bitstream 104. The core audio bitstream 105 is fed to a core audio decoder 230 which reconstructs one or more SPAR downmix channel signals 205. These one or more reconstructed downmix channel signals 205 are fed together with the SPAR metadata bitstream 102 to a SPAR upmix unit 210. The SPAR upmix unit 210 upmixes the one or more reconstructed downmix channel signals 205 to provide a reconstruction 201 of at least a subset of the channels of the original Ambisonics signal 101 (sometimes referred to as an intermediate Ambisonics signal 101). This intermediate Ambisonics signal 201 is typically only an approximation of the original Ambisonics input audio signal 101 of the encoding device 100. The Ambisonics order L of the intermediate Ambisonics signal 201 is approximately the same as, or no greater than, the Ambisonics order of the original input audio signal 101 .

[0028] This intermediate Ambisonics signal 201 can be provided to a DirAC analysis and metadata generation unit 250 of the decoding device 200. This optional DirAC analysis and metadata generation unit 250 can perform DirAC analysis and metadata generation based on the SPAR reconstructed intermediate Ambisonics signal 201. The optional (auxiliary) DirAC metadata 204 (referred to as auxiliary DirAC metadata 204) from the DirAC analysis and metadata generation unit 250, the optional DirAC metadata bitstream 104 received from the encoding device 100 and the SPAR reconstructed intermediate Ambisonics signal 201 can be provided to a DirAC synthesis unit 220. This DirAC synthesis unit 220 can decode the received metadata bitstream 104. Then, DirAC signal synthesis can be performed on the SPAR reconstructed intermediate Ambisonics signal 201 using the available DirAC metadata 104, 204. The DirAC synthesis unit 220 may be configured to synthesize a higher order output Ambisonics signal 211 (compared to the input audio signal 101), to synthesize (render) a binaural output signal 211, or to synthesize (render) a multi-speaker output signal 211.

[0029] As shown in Figures 1, 2 and 4, DirAC analysis can be optionally performed in the encoding device 100 (in the DirAC analyzer and metadata encoder 120) and / or in the decoding device 200 (in the DirAC analysis and metadata generation unit 250). DirAC analysis and metadata encoding can be performed in the encoding device 100 if the transport channel signals (i.e. one or more downmix channel signals 103 and / or intermediate Ambisonics signal 201) are not suitable for performing DirAC analysis after decoding in the decoding device 200. This is the case when the decoded transport channel signals are only single (mono) or stereo audio signals and not Ambisonics signals. In such a situation, it is not possible to perform a Direction of Arrival (DOA) analysis for all spherical or at least cylindrical directions (which is typically performed in the DirAC analysis units 120, 250). This is also the case for certain frequency bands of the transmitted Ambisonics signal, if parametric coding methods (e.g. bandwidth extension, spectral band replication (SBR), etc.) are used in the core audio encoder 130 that make DOA analysis impossible or unreliable for certain frequency bands. The advantage of combining DirAC and SPAR as outlined herein is that in any case the decoded intermediate Ambisonics signal 201 is available for DirAC analysis at the decoder side (in the DirAC analysis and metadata generation unit 250).

[0030] An aspect of SPAR and DirAC coding is that both methods operate on frequency bands (subbands) and frames or subframes, i.e. T / F tiles. Implementations of these methods can use operations in the time domain on subbands, the QMF domain, or the frequency domain, e.g. on (modified) DFT frequency bins or groups of such bins. Thus, all aspects described herein are applicable to any T / F tile. Furthermore, the terms subband, frequency band / bin, or QMF band / bin are interchangeable in the context of this specification. Similarly, the terms subband domain, QMF domain, or frequency domain are synonymous in the context of this specification.

[0031] When combining SPAR and DirAC coding, it may turn out that certain T / F tiles or subbands benefit more from performing DirAC analysis based on the SPAR decoded intermediate Ambisonics signal 201 (in the DirAC analysis and metadata generation unit 250), while in other cases it may be beneficial to perform such analysis in the encoding device 100 (in the DirAC analyzer and metadata encoder 120) and transmit the corresponding metadata bitstream 104 to the decoding device 200. Usually, the DirAC parameter analysis is more reliable in the encoding device 100, since it can be based on the original input audio signal 101. However, in this case, the corresponding metadata bitstream 104 needs to be coded and transmitted. Assuming a certain total bitrate budget, the partial bitrate of the DirAC metadata bitstream 104 comes at the expense of the bitrate available for the SPAR metadata bitstream 102 and the core audio bitstream 105. Therefore, for at least one or more selected T / F tiles or subbands, it may be more beneficial to the performance of the overall coding system to perform DirAC analysis on the corresponding SPAR decoded T / F tile or subband signal of the intermediate Ambisonics signal 201 (in the decoding device 200).

[0032] Whether to select DirAC parameter analysis for a given subband or T / F tile at the encoder side or at the decoder side (to achieve an optimal coding system) depends on one or more characteristics of one or more core audio coded SPAR downmix channel signals 103 and then on the suitability of the reconstructed intermediate Ambisonics signal 201 after upmixing in the upmix unit 210 of the SPAR decoder 210, 230 for performing DirAC parameter analysis (in the DirAC analysis and metadata generation unit 250). It has been observed that subbands and time frames whose coding is waveform-preserving are generally more suitable for DirAC parameter analysis in the decoding device 200 (in the DirAC analysis and metadata generation unit 250) than subbands and time frames whose coding is not waveform-preserving. This is typically the case for low frequency bands and / or for time / frequency signal parts that are tonal rather than noise-like etc. Thus, the codec system 100, 200 may be configured to perform DirAC parameter analysis at the encoder side (in the DirAC analysis and metadata encoder 120) for high frequency bands and / or noise-like time / frequency signal parts, whereas the codec system 100, 200 may be configured to perform DirAC parameter analysis at the decoder side (in the DirAC analysis and metadata generation unit 250) for low frequency bands and / or tonal time / frequency signal parts.

[0033] Thus, the combined SPAR and DirAC coding / decoding system 100, 200 may include adaptation means for adaptively switching between DirAC metadata transmission from the encoding device 100 and DirAC analysis performed in the decoding device 200 for selective T / F tiles, subbands and / or frames. The adaptation may depend on one or more detected characteristics of the input audio signal 101, such as, for example, tones or noise.

[0034] A combination of SPAR and DirAC encoding / decoding systems 100, 200 may include a decoding device 200 operating with a modified number of SPAR upmix channels fed to a subsequent DirAC unit 220, 250. SPAR systems typically upmix to an Ambisonics signal, which for a given Ambisonics order L is (L+1) 2 This means generating upmix channels. Especially for lower bitrate operation (e.g., at <64 kbps), the SPAR decoding and upmix operation (in the upmix unit 210) may result in a lower signal quality, at least for certain T / F tiles or frequency bands. This may affect the subsequent DirAC operation and thus the quality of the audio output signal 211 of the encoding / decoding system 100, 200.

[0035] This problem can be addressed by modifying the SPAR such that the number of upmix channel signals (following the upmix in the upmix unit 210) is reduced (at least for certain T / F tiles or frequency bands). As an example, for the FOA input audio signal 101, the SPAR can be modified to generate only a single upmix channel or two upmix channels, corresponding to the decoded B-format FOA component signal W, or W and Y, respectively, at least for certain T / F tiles or frequency bands. This SPAR modification can be achieved by setting each upmix coefficient (in the SPAR metadata) for the discarded channels (Y, Z, X, respectively Z, X) to 0 and / or, in the two-channel example, by not performing prediction from W to Y, so that the transmitted prediction residual signal Y' is identical to Y.

[0036] The DirAC units 220 and / or 250 of the decoding device 200 can be modified to operate with a correspondingly reduced number of input signals, at least for selected T / F tiles or frequency bands. For the DirAC synthesis unit 220, this means that the number of prototype signals used is reduced accordingly. In the one-channel example, this means that the DirAC synthesis is based on a single (mono) prototype signal, while in the two-channel example, with the W and Y input signals, the DirAC synthesizer can convert these signals into a left / right stereo representation from which the prototype signals can be obtained. DirAC analysis (in the DirAC analysis and metadata generation unit 250) is typically not possible for these T / F tiles or frequency bands. Therefore, for these T / F tiles or frequency bands, the DirAC metadata 104 should be calculated in the encoding device 100 and transmitted in the encoder bitstream 106.

[0037] Therefore, the decoding device 200 detects that the output signal 201 of the upmix unit 210 (i.e., the intermediate Ambisonics signal 201) is (L+1) 2 Alternatively, the partial upmix can be performed on a subset of the T / F tiles or subbands. Alternatively, the partial upmix can be performed on the complete set of T / F tiles or subbands. The option to perform a partial upmix can be used to improve the perceived audio quality of the output signal 211 at a relatively low bitrate.

[0038] As an example, the decoding device 200 may be configured to perform a partial upmix in the upmix unit 210, such that the output signal 201 of the upmix unit 210 is a stereo signal. Alternatively or additionally, the decoding device 200 may be configured to put the DirAC synthesis unit 220 into a pass-through operation mode (in which the DirAC synthesis unit 220 passes the output signal 201 of the upmix unit 210 without modifying and / or performing operations on the output signal 201), which allows an efficient generation of a stereo output signal 211 (e.g. in a multi-speaker case with two speakers).

[0039] The combined SPAR and DirAC encoding / decoding system 100, 200 can be configured to efficiently process head tracking input data and adjust (rotate) the output audio signal 211 in response to such data. In one example, a relatively low-order (e.g. FOA) SPAR reconstructed intermediate Ambisonics signal 201 can be rotated (according to the head tracking data) before being fed to the DirAC unit 220, 250. This is particularly numerically efficient if the DirAC analysis and metadata generation is based only on the SPAR reconstructed intermediate Ambisonics signal 201 available at the decoding device 200. A less numerically efficient alternative is to rotate the higher-order Ambisonics signal 211 after DirAC synthesis (in the DirAC synthesis unit 220). Even if the DirAC metadata 104 is (partially) received from the encoding device 100, this metadata 104 (including the azimuth and elevation angles of the detected dominant direction) can be subjected to an additional adjustment of the reception angle based on the rotation angle obtained from the head tracking device.

[0040] 3a and 3b show examples of a combined SPAR and DirAC encoding device 100 and a combined SPAR and DirAC and a combined SPAR and DirAC decoding device 200. The encoding device 100 and / or the decoding device 200 can be configured to switch between the SPAR and DirAC encoders depending on the bit rate. The encoding device 100 and the decoding device 200 shown in Figs. 3a and 3b cannot provide the synergistic effect described herein.

[0041] The encoding device 100 shown in Fig. 3a comprises a selection unit 300 configured to select a SPAR encoder branch or (alternatively) a DirAC encoder branch depending on a (target) bitrate 301 of the encoder bitstream 106. As an example, if the bitrate 301 is less than or equal to a predefined bitrate threshold, the SPAR encoder branch may be selected. On the other hand, if the bitrate 301 is greater than the bitrate threshold, the DirAC encoder branch may be selected. As a result, the encoder bitstream 106 comprises either the bitstream 102, 105 from the SPAR encoder branch or the bitstream 325, 104 from the DirAC encoder branch.

[0042] The DirAC encoder branch may include a downmix unit 321 configured to downmix multiple input channel signals of the Ambisonics input audio signal 101 into one or more transport channel signals 324. The one or more transport channel signals 324 may be encoded using any (single-channel, dual-channel or multi-channel) waveform encoder 322, thereby providing a core audio bitstream 325.

[0043] FIG. 3b shows a corresponding decoding device 200 including a SPAR decoder branch and a separate DirAC decoder branch, both of which are configured to generate output signals that can be selected (depending on the bit rate 301 of the encoder bitstream 106) to provide an output signal 211 of the decoding device 200.

[0044] The SPAR decoder branch may include an optional rendering unit 320 configured to generate an alternative output signal 311 (different from the intermediate Ambisonics signal 201), such as a stereo or binaural signal. A selection unit 371 may be provided to select between the intermediate Ambisonics signal 201 and the alternative output signal 311.

[0045] The DirAC decoder branch typically includes a metadata decoding unit 340 configured to generate DirAC metadata 304 from the DirAC metadata bitstream 104. Furthermore, the DirAC decoder branch may include a core decoder unit 342 configured to generate one or more reconstructed transport channel signals 344 (corresponding to the one or more transport channel signals 324) based on the core audio bitstream 325. The one or more reconstructed transport channel signals 344 and the DirAC metadata 304 may be used in a DirAC synthesis unit 360 to generate an output signal (e.g., an Ambisonics signal).

[0046] The DirAC decoder branch may further include a DirAC analyzer and metadata generator 350 (similar or equivalent to unit 350) configured to analyze one or more reconstructed transport channel signals 344 to generate auxiliary DirAC metadata 354 that may be used in a DirAC synthesis unit 360 to generate an output signal (for rendering). The output signal of the DirAC synthesis unit 360 may be selected (using a selection unit 372, 300) as the overall output signal 211 of the decoding device 200.

[0047] Furthermore, the DirAC decoder branch may include an i rendering unit 361 (or a harmonic internal rendering unit) configured to generate an alternative output signal (as an alternative to the output signal of the DirAC synthesis unit 360). The alternative output signal may be a binaural signal or a stereo signal (as an alternative to the Ambisonics signal). The rendering unit 361 may be configured to generate the alternative output signal based on the DirAC metadata 304, the auxiliary DirAC metadata 354, and / or the reconstructed transport channel signal 344. The rendering unit 361 may be included within the DirAC synthesis unit 220 of the decoding device 200 of FIG. 2.

[0048] It should be noted that one or more of the components of the encoding device 100 of Figure 3a may be used in the encoding device 100 of Figure 1. Similarly, one or more of the components of the decoding device 200 of Figure 3b may be used in the decoding device 200 of Figures 2 and / or 4.

[0049] In the encoding device 100 of Fig. 1, the SPAR waveform encoder may utilize any core audio coding tool, in particular for all bit rates. SPAR can be performed in combination with DirAC for all bit rates. DirAC decoding (in the DirAC synthesis unit 220) can rely on the intermediate Ambisonics signal 201 reconstructed by SPAR for all bit rates.

[0050] At relatively low bit rates, where the DirAC codec typically uses one or two transport channels, the combined SPAR / DirAC codec can accommodate the following behavior: operating in a specific frequency band with one or two transport channels that require DirAC operations at both the encoder and the decoder to reconstruct the FOA signal; and / or It operates in a specific other frequency band with four SPAR reconstruction signals of the FOA signal.

[0051] Thus, FOA pass-through can be provided.

[0052] At a certain (relatively low) bit rate, a combined SPAR / DirAC codec can be adapted to operate at least in a certain frequency band with SPAR reconstruction of the lower order Ambisonics signal and rely on DirAC to reconstruct the original Ambisonics order, thus providing a HOA pass-through.

[0053] DirAC can be used as a primary tool to enhance the spatial resolution of audio signals based on low-order Ambisonics signals reconstructed by SPAR. In particular, The FOA and / or HOA input audio signals 101 can be converted to HOAm, binaural and / or LS (loudspeaker) signals, where the output Ambisonics order m is greater than the input Ambisonics order n.

[0054] You can provide the option of internal and / or external renderers. For example: An internal renderer that performs comparably to the (supposed) reference renderer for subjective evaluation using a reference test, An internal renderer that doesn't introduce any additional delay, and / or External renderers that provide advanced features that cannot be tested by reference tests against the reference renderer. The renderer can improve pass-through performance.

[0055] The combined SPAR / DirAC codec described herein can be configured to be backward compatible with the standalone SPAR and DirAC codecs. In particular, the original SPAR behavior can be preserved when the decoder-side DirAC synthesis module is in a pass-through mode of operation. Furthermore, the original DirAC behavior can be preserved (e.g., by setting the SPAR prediction coefficients to 0) when the SPAR module is in a pass-through mode of operation.

[0056] By providing FOA pass-through, strictly quality vs. bitrate behavior can be improved. By providing HOA pass-through, the codec can achieve the performance of a pure SPAR encoder for HOA signals. The use of DirAC allows efficient generation of HOA content (e.g., HOA4 resolution). A combined SPAR / DirAC system operates particularly efficiently at low bitrates, since it can rely on an active downmix channel W* that may be generated within the SPAR encoding module.

[0057] As mentioned above, SPAR and / or DirAC processing is usually performed in different subbands and / or T / F files. For this purpose, one or more different types of filter banks (FB) can be used. As an example, a first type of filter bank called FB_A can be used. FB_A can be a QMF (quadrature mirror filter) filter bank, in particular a Complex Low Delay Filter bank (CLDFB). FB_A can include 60 channels that can be grouped into a set of subbands. A second type of filter bank can be called FB_B. FB_B can be a Nyquist filter bank that includes the application of a modified DFT (Discrete Fourier Transform) that can group different bins of the modified DFT into a set of subbands. The filter bank can be applied to the time domain signal with a certain overlap (e.g., 1 ms overlap) to avoid blocking effects. FB_A (analysis and synthesis) may exhibit a delay of 2.5 to 5 ms and / or FB_B (analysis and synthesis) may exhibit a delay of 2 ms.

[0058] In a first example, the downmix unit 110 of the SPAR encoder 110, 130 may use FB_B analysis to generate the SPAR metadata bitstream 102 and FB_B synthesis to generate one or more downmix channel signals 103. Furthermore, the DirAC analyzer and metadata encoder 120 may use FB_A analysis. On the decoder side, the SPAR upmix unit 210 may use FB_B analysis and FB_B synthesis to generate the intermediate (Ambisonics) signal 201. Furthermore, the DirAC unit 220, 250 may use FB_A analysis of the intermediate (Ambisonics) signal 201 and FB_A synthesis (after DirAC processing) to generate the output signal 211.

[0059] The FB_B analysis can be performed at the input of the downmix unit 110 of the SPAR encoder 110, 130, and the FB_B synthesis can be performed at the output of the downmix unit 110 of the SPAR encoder 110, 130 (output providing one or more downmix channel signals 103). Furthermore, the FB_A analysis can be performed at the input of the DirAC analyzer and metadata encoder 120. Furthermore, the FB_B analysis can be performed at the input of the SPAR upmix unit 210 (input of one or more reconstructed downmix channel signals 205), and the FB_B synthesis can be performed at the output of the SPAR upmix unit 210. Furthermore, the FB_A analysis can be performed on the intermediate (Ambisonics) signal 201 (before entering the DirAC analysis and metadata generation unit 250 and / or the DirAC synthesis unit 220), and the FB_A synthesis process can be performed at the output of the DirAC synthesis unit 220.

[0060] In a further example, the Ambisonics input audio signal 101 can be analyzed using FB_B (preferably for both SPAR and DirAC processing before entering the downmix unit 110 of the SPAR encoder 110, 130 and / or the DirAC analyzer and metadata encoder 120). The FB_B synthesis can be used to generate one or more downmix channel signals 103 (at the output of the downmix unit 110 of the SPAR encoder 110, 130). The decoding device 200 can use the filter bank configuration of the first example.

[0061] In a preferred example, FB_B (or alternatively FB_A) analysis can be used to analyze the Ambisonics input audio signal 101 (preferably for SPAR and DirAC processing before entering the downmix unit 110 of the SPAR encoder 110, 130 and / or the DirAC analyzer and metadata encoder 120). FB_B (or alternatively FB_A) synthesis can be used to generate one or more downmix channel signals 103 (and can be performed at the output of the downmix unit 110 of the SPAR encoder 110, 130). On the decoder side, FB_A (or alternatively FB_B) analysis can be used to analyze one or more reconstructed downmix channel signals 205 (at the input of the SPAR upmix unit 210). The intermediate (Ambisonics) signal 201 can be provided to the DirAC processing unit 250, 220 in the filter bank domain, thereby removing the need for a separate filter bank operation. This can reduce the processing load and delay of the decoding device 200. The FB_A (or alternatively FB_B) synthesis may be used at the output of the DirAC synthesis unit 220 to generate the output signal 211 .

[0062] 5a shows a flow chart of an example of a method 500 for encoding an Ambisonics input audio signal 101. The Ambisonics input audio signal 101 comprises a number of different input channel signals, where different channels can be associated with different panning and / or spherical basis functions and / or different directivity patterns. As an example, an L-th order 3D Ambisonics signal may be (L+1) 2 A First Order Ambisonics (FOA) signal is an Ambisonics signal of order L=1, and a Higher Order Ambisonics (HOA) signal is an Ambisonics signal of order L>1.

[0063] The method 500 comprises a step 501 of providing an input audio signal 101 to a spatial reconstruction (SPAR) encoder 110 , 130 and a directional audio coding (DirAC) analyzer and parameter encoder 120 (in parallel).

[0064] The SPAR encoder 110, 130 may be configured to downmix multiple input channel signals of an Ambisonics input audio signal 101 in the subband and / or QMF domain into one or more downmix channel signals 103. Typically, the number of downmix channel signals 103 is less than the number of input channel signals. The one or more downmix channel signals 103 may be encoded by a (waveform) audio encoder 130 to provide an audio bitstream 105.

[0065] Furthermore, the SPAR encoder 110, 130 may be configured to generate a SPAR metadata bitstream 102 associated with a representation of the Ambisonics input audio signal 101 in the subband and / or QMF domain. The SPAR metadata bitstream 102 may be adapted to upmix (in a corresponding decoding device 200) one or more downmix channel signals 103 into a plurality of reconstructed channel signals of a reconstructed intermediate Ambisonics signal 201, which typically correspond (in a one-to-one relationship) to the plurality of input channel signals of the Ambisonics input audio signal 101.

[0066] To determine the SPAR metadata bitstream 102, one or more downmix channel signals 103 may be transformed into and / or processed in the subband domain. Also, the multiple input channel signals of the input sound audio signal 101 may be transformed into the subband domain (including subbands of multiple different frequency bands). The SPAR metadata bitstream 102 may then be determined subband-by-subband (e.g., frequency band-by-frequency band and / or time / frequency tile-by-time) such that approximations of the subband signals of the multiple input channel signals of the input audio signal 101 are obtained, in particular by upmixing the subband signals of the one or more downmix channel signals 103 using the SPAR metadata bitstream 102. SPAR metadata for different subbands (i.e., for different frequency bands and / or different time / frequency tiles) may be combined to form the SPAR metadata bitstream 102.

[0067] The DirAC analyzer and parameter encoder 120 may be configured to perform a direction of arrival analysis (DoA) on the Ambisonics input audio signal 101 in the subband and / or QMF domain to determine a DirAC metadata bitstream 104 indicative of the direction of arrival of one or more dominant components of the Ambisonics input audio signal 101. The DirAC metadata bitstream 104 may indicate the spatial direction of one or more dominant components of the Ambisonics input audio signal 101. The DirAC metadata 104, in particular the spatial direction of one or more dominant components, may be generated for multiple different frequency bands and / or multiple different time / frequency tiles.

[0068] The method 500 further includes a step 502 of generating an encoder bitstream 106 based on the output 102, 105 of the SPAR encoder 110, 130 and the output 104 of the DirAC analyzer and parameter encoder 120. The DirAC analyzer may be configured to perform a direction of arrival (DoA) analysis and / or a diffuseness analysis. In other words, the DirAC analysis may include a DoA analysis and / or a diffuseness analysis. As mentioned above, the output 102, 105 of the SPAR encoder 110, 130 may include a SPAR metadata bitstream 102 and an audio bitstream 105 indicative of a set of SPAR downmix channel signals 103. The output 104 of the DirAC analyzer and parameter encoder 120 may include a DirAC metadata bitstream 104. The step 502 of generating an encoder bitstream 106 may include a step of multiplexing the SPAR metadata bitstream 102, the audio bitstream 105, and the DirAC metadata bitstream 104 into a common encoder bitstream 106. The representation of the encoder bitstream 106 may be transmitted (particularly to the decoder device 200) and / or stored.

[0069] Thus, a method 500 for jointly using SPAR and DirAC coding is described in order to provide a particularly efficient Ambisonics audio encoder with improved perceptual quality. In the context of the method 500, the data provided by the DirAC coding scheme may be limited to DirAC metadata, whereas one or more transport channels of the DirAC coding scheme may be replaced by data provided by the SPAR coding scheme (in particular one or more downmix channel signals and / or SPAR metadata).

[0070] The method 500 may include generating subband data in multiple frequency bands and / or multiple time / frequency tiles, the subband data representing the input audio signal 101. For this purpose, a QMF and / or a subband filter bank may be used.

[0071] Further, the method 500 may include selecting a subset of the frequency bands and / or time / frequency tiles, which may correspond to frequency ranges above a predefined threshold frequency. This may be used to allow selectively operating based on SPAR metadata for one (lower) frequency range and DirAC metadata 104 for another (higher) frequency range.

[0072] Alternatively or additionally, characteristic information regarding characteristics of the input audio signal 101 can be determined (e.g. by analyzing the input audio signal 101), in particular noise-like or timbre-related characteristics of the input audio signal 101. A subset of frequency bands and / or time / frequency tiles can then be selected based on the characteristic information. In particular, threshold frequencies for frequency ranges of the selected frequency bands and / or time / frequency tiles can be determined based on the characteristic information.

[0073] The output 104 of the DirAC analyzer and parameter encoder 120, in particular the DirAC metadata bitstream 104, can then be determined for a selected subset of frequency bands and / or time / frequency tiles, in particular only for the selected subset of frequency bands and / or time / frequency tiles.

[0074] In other words, DirAC metadata may be determined in the encoding device 100 only for a reduced subset of the entire number of frequency bands and / or time / frequency tiles, in particular for frequency bands and / or time / frequency tiles that do not have tonal characteristics and / or have noise-like characteristics and / or for upper frequency bands and / or time / frequency tiles (above a certain threshold frequency), thereby providing a particularly efficient and high-quality Ambisonics coding scheme.

[0075] As mentioned above, SPAR and / or DirAC processing is typically performed in the subband and / or filter bank domain. The method 500 may include generating subband data in multiple frequency bands and / or multiple time / frequency tiles, the subband data representing the input audio signal 101. The subband data may be generated using an analysis filter bank. The subband data is then provided to the SPAR encoders 110, 130 to generate the SPAR metadata bitstream 102 and to the DirAC analyzer and parameter encoder 120 to generate the DirAC metadata 104.

[0076] Thus, a single analysis filterbank can be used to transform the input audio signal 101 into the filterbank domain. The input 101 signal can be represented in the filterbank domain by coefficients and / or samples for different subbands (i.e. by subband data). This subband data can be used as the basis for SPAR and DirAC processing, thereby providing a particularly efficient coding device 100.

[0077] The method 500 may include using a synthesis filter bank to generate one or more downmix channel signals 103 in the SPAR encoder 110, 130. The analysis filter bank and the synthesis filter bank may form a (possibly perfect reconstruction) analysis / synthesis filter bank. The one or more downmix channel signals 103 may be time-domain signals that are encoded in the core audio encoder 130.

[0078] Therefore, a single analysis / synthesis filter bank (e.g., a Nyquist or QMF filter bank) can be used within the encoding device 100 to perform SPAR and DirAC processing, thereby reducing the computational complexity of the encoding device 100 (without affecting perceptual quality).

[0079] Fig. 5b shows a flow chart of an example of a (computer-implemented) method 510 for decoding an encoder bitstream 106 representing an Ambisonics input audio signal 101. The method 510 comprises a step 511 of generating an intermediate Ambisonics signal 201 based on the encoder bitstream 106 using a spatial reconstruction (SPAR) decoder 210, 230. The intermediate Ambisonics signal 201 may have the same order L as the input audio signal 101. The intermediate Ambisonics signal 201 may be a time domain signal. Alternatively, the intermediate Ambisonics signal 201 may be represented in a filter bank or subband domain.

[0080] The SPAR metadata bitstream 102 and the audio bitstream 105 may be extracted from the encoder bitstream 106. The intermediate Ambisonics signal 201 may then be generated from the SPAR metadata bitstream 102 and the audio bitstream 105 using the SPAR decoders 210, 230. In particular, a set of reconstructed downmix channel signals 205 may be generated from the audio bitstream 105 using the (waveform) audio decoder 230. Furthermore, the set of reconstructed downmix channel signals 205 may be generated from the intermediate Ambisonics signal 201 (including multiple (in particular (L+1)) downmix channel signals) based on the SPAR metadata bitstream 102 using the upmix unit 210. 2 The intermediate channel signals of the intermediate Ambisonics signal 201 are typically reconstructions and / or approximations of the input channel signals of the Ambisonics input audio signal 101 or a subset thereof.

[0081] Further, the method 500 includes processing the intermediate Ambisonics signal using a directional audio coding (DirAC) synthesizer 220 (also referred to as a DirAC synthesis unit) to provide an output audio signal 211 for rendering. The output signal 211 may include at least one of an Ambisonics output signal, a binaural output signal, a stereo or a multi-speaker output signal. In particular, a DirAC metadata bitstream 104 may be extracted from the encoder bitstream 106. The intermediate Ambisonics signal 201 may be processed using the DirAC synthesizer 220 in dependence on the DirAC metadata bitstream 104 to provide the output audio signal 211.

[0082] As mentioned above, the intermediate Ambisonics signal 201 can be represented in the time domain. In this case, the DirAC processing can include the application of an analysis filter bank to transform the intermediate Ambisonics signal 201 into a filter bank domain. In a preferred example, the intermediate Ambisonics signal 201 (provided by the SPAR processing) is already represented in the filter bank domain. By doing this, the application of a synthesis filter bank (in the SPAR processing) and the subsequent application of an analysis filter bank (in the DirAC processing) can be eliminated, thereby improving the computational efficiency and the perceptual quality of the decoding device 200.

[0083] Thus, a decoding method 510 is described that utilizes SPAR decoding followed by a DirAC synthesis operation (and possibly a DirAC analysis operation). SPAR decoding can be used to provide one or more transport channels (in particular the intermediate Ambisonics signal 201) in an efficient and high-quality manner. A DirAC synthesizer can be used to provide one or more different types of output signal 211 for rendering the audio signal in a flexible manner. In this context, DoA data of one or more main components of the input audio signal 101 (contained in the DirAC metadata) can be used to generate the output signal 211.

[0084] The DirAC metadata may be provided (at least in part) in the encoder bitstream 106. Alternatively, or in addition, the DirAC metadata may be generated (at least in part) in the decoding device 200.

[0085] Thus, the method 510 may include processing the intermediate Ambisonics signal 201 in a DirAC analyzer 250 (i.e. in the DirAC analysis and metadata generation unit 250) to generate auxiliary DirAC metadata 204. In this context, a DoA analysis may be performed to determine the auxiliary DirAC metadata 204 indicative of the DoA of one or more major components of the intermediate Ambisonics signal 201.

[0086] The intermediate Ambisonics signal 201 can be processed using a DirAC synthesizer 220 in dependence on the auxiliary DirAC metadata 204 to provide an output audio signal 211. By using the DirAC metadata determined in the decoding device 200, the efficiency of the Ambisonics codec can be further improved.

[0087] As mentioned above, (SPAR and / or DirAC) metadata is typically generated for several different frequency bands and / or time / frequency tiles. The codec can be configured to generate DirAC metadata for some of the different frequency bands and / or time / frequency tiles in the encoding device 100 and for other parts of the different frequency bands and / or time / frequency tiles in the decoding device 200 (particularly in a complementary and / or mutually exclusive manner). This can further improve the efficiency and quality of the Ambisonics codec.

[0088] The method 510 may include generating subband data in a number of frequency bands and / or a number of time / frequency tiles (e.g., using a subband transform and / or a QMF filterbank), where the subband data represents the intermediate Ambisonics signal 201 (in the filterbank or subband domain). Additionally, the method 510 may include selecting a subset of the number of frequency bands and / or the number of time / frequency tiles.

[0089] A subset of frequency bands and / or time / frequency tiles may be selected that correspond to a frequency range of frequencies above a predefined threshold frequency.

[0090] Alternatively or additionally, characteristic information regarding characteristics of the input audio signal 101 and / or the intermediate Ambisonics signal 201, in particular noise-like or timbre-related characteristics of the input audio signal 101 and / or the intermediate Ambisonics signal 201, can be determined, for example by analysing the intermediate Ambisonics signal 201. A subset of frequency bands and / or time / frequency tiles can then be determined based on the characteristic information. In particular, a threshold frequency for selecting the subset can be determined based on the characteristic information.

[0091] The method 510 may further include determining auxiliary DirAC metadata 204 for a selected subset of frequency bands and / or time / frequency tiles based on the subband data, in particular for only the selected subset of frequency bands and / or time / frequency tiles.

[0092] Thus, the auxiliary DirAC metadata 204 can be generated directly in the decoding device 200 for a reduced subset of frequency bands and / or time / frequency tiles (without the need to transmit DirAC metadata for these frequency bands and / or time / frequency tiles), which may be the case for low frequency bands, allowing further improvement of the efficiency of the Ambisonics codec.

[0093] The method 510 may include determining orientation data relating to the (spatial) orientation of the listener's head (within the listening environment), in particular using a head tracker. A rotation operation may be performed on the intermediate Ambisonics signal 201 depending on the orientation data to generate a rotated Ambisonics signal. Thus, the intermediate Ambisonics signal may be rotated to take into account the orientation of the listener's head in a resource-efficient manner. Furthermore, auxiliary DirAC metadata may be generated based on the rotated Ambisonics signal (instead of the unrotated intermediate Ambisonics signal).

[0094] The rotated intermediate Ambisonics signal 201 may then be processed using a DirAC synthesizer 220 to provide a (rotated) output audio signal 211 for rendering to the listener. By doing this, head rotation can be taken into account efficiently and accurately.

[0095] As mentioned above, the DirAC metadata bitstream (i.e., DirAC metadata) 104 can be extracted from the encoder bitstream 106. The method 510 can include performing a rotation operation on the DirAC metadata bitstream (i.e., DirAC metadata) 104 in response to the orientation data to generate a rotated DirAC metadata bitstream (i.e., rotated DirAC metadata). The intermediate Ambisonics signal 201 or an Ambisonics signal derived therefrom (in particular the rotated Ambisonics signal) can then be processed in response to the rotated DirAC metadata bitstream (i.e., on the DirAC metadata) using a DirAC synthesizer 220 to provide an output audio signal 211 for rendering to a listener. By doing this, head rotation can be taken into account efficiently and accurately.

[0096] The method 510 may include generating an Ambisonics output signal 211 from the intermediate Ambisonics signal 201 using a DirAC synthesizer 220. For this purpose, the DirAC metadata bitstream 104 (from the encoder bitstream 106) and / or the auxiliary DirAC metadata 204 (generated in the decoding device 200) may be used. The Ambisonics output signal 211 may have an Ambisonics order L that is larger than the Ambisonics order of the input audio signal 101 and / or the intermediate Ambisonics signal 201. By doing this, the quality and flexibility of Ambisonics audio rendering may be efficiently improved.

[0097] As mentioned above, the method 510 may include extracting the audio bitstream 105 from the encoder bitstream 106 and generating a set of reconstructed downmix channel signals 205 from the audio bitstream 105 using the (core) audio decoder 230. In other words, the set of reconstructed downmix channel signals 205 may be derived from the encoder bitstream 106.

[0098] The method 510 may further include applying an analysis filterbank to the set of reconstructed downmix channel signals 205 to transform the set of reconstructed downmix channel signals 205 (from the time domain) into a filterbank domain. The analysis filterbank may be configured to transform the one or more different reconstructed downmix channel signals 205 into different frequency channels or frequency bins that may be grouped into a set of subbands. The one or more different reconstructed downmix channel signals 205 may be represented in the filterbank domain as samples and / or coefficients for different subbands.

[0099] Furthermore, the method 510 may comprise a step 511 of generating an intermediate Ambisonics signal 201 represented in the filterbank domain based on the set of reconstructed downmix channel signals 205 in the filterbank domain. For this purpose, an upmix operation (using the SPAR metadata bitstream 102) may be performed. The intermediate Ambisonics signal 201 may be represented in the filterbank domain as samples and / or coefficients for different subbands.

[0100] The method 510 may further include a step 512 of processing the intermediate Ambisonics signal 201 represented in the filterbank domain using a DirAC synthesizer 220. Thus, the DirAC synthesizer 220 (and possibly the DirAC analyzer 250) can directly operate on the intermediate Ambisonics signal 201 represented in the filterbank domain (without having to perform a separate filterbank operation), thereby allowing the DirAC metadata 104, 204 (already represented in the filterbank domain) to be directly applied to the intermediate Ambisonics signal 201 represented in the filterbank domain.

[0101] Thus, the decoding device 200 can use a single analysis filterbank to transform one or more reconstructed downmix signals 205 into the filterbank domain. SPAR upmixing and / or DirAC processing can then be provided directly within the same filterbank domain. This can provide a particularly efficient decoding device 200. Furthermore, the audio quality of the decoding device 200 can be improved.

[0102] The method 510 may further include processing the intermediate Ambisonics signal 201 represented in the filterbank domain using a DirAC synthesizer 220 to generate an output signal 211 represented in the filterbank domain 512. As mentioned above, the DirAC synthesis may be performed directly in the filterbank domain of an analysis filterbank that is applied to one or more reconstructed downmix signals 205, thereby generating the output signal 211 in this filterbank domain. The output signal 211 may be represented in the filterbank domain as samples and / or coefficients for different subbands of the filterbank domain.

[0103] Furthermore, the method 510 may include applying a synthesis filter bank to the output signal 211 represented in the filter bank domain to generate the output signal 211 in the time domain. The analysis filter bank and the synthesis filter bank typically form a joint analysis / synthesis filter bank, in particular a perfect reconstruction analysis / synthesis filter bank. As an example, the analysis filter bank and the synthesis filter bank may be a Nyquist filter bank or a QMF (quadrature mirror filter) filter bank.

[0104] The encoder bitstream 106 may be generated using a first type of filterbank, in particular a Nyquist filterbank. The analysis filterbank (used in the decoding device 200) may be a second type of filterbank different from the first type, in particular a QMF filterbank. The frequency band boundaries of the first type of filterbank are preferably adjusted and / or matched to the corresponding frequency band boundaries of the second type of filterbank.

[0105] Therefore, different types of analysis / synthesis filterbanks may be used in the encoding device 100 and the decoding device 200. This allows to further improve the perceptual quality of the overall codec while keeping the latency of the codec as low as possible.

[0106] The intermediate Ambisonics signal 201 (in the time domain or in the filter bank domain) may contain fewer channels than the original Ambisonics input audio signal 101. In other words, the SPAR decoders 210, 230 may be used to perform (only) a partial upmix operation to generate the intermediate Ambisonics signal 201 that contains fewer channels than the Ambisonics input audio signal 101.

[0107] The partial upmix operation can be performed in a filter bank domain with multiple subbands and / or multiple time / frequency tiles. The intermediate Ambisonics signal 201 can include fewer channels than the Ambisonics input audio signal 101 for all of the multiple subbands and / or all of the multiple time / frequency tiles. Alternatively, the intermediate Ambisonics signal 201 can include fewer channels than the Ambisonics input audio signal 101 for only a subset of the multiple subbands and / or multiple time / frequency tiles.

[0108] Thus, the decoding device 200 may be configured to have the SPAR decoder 210, 230 generate only a subset of the channels of the original Ambisonics input audio signal 101, for example when the bitrate of the encoder bitstream 106 is equal to or less than a predefined bitrate threshold (e.g. 64 kbs). This subset of channels can then be used in the DirAC synthesis 220 to generate the output signal 211. This allows to improve the audio quality (at a relatively low bitrate) while reducing the numerical complexity and memory requirements of the decoder operation.

[0109] The decoding device 200 may be configured to put the DirAC synthesizer 220 into a pass-through operation mode and / or to bypass the DirAC synthesizer 220. This may be done so that the intermediate Ambisonics signal 201 corresponds to the output audio signal 211 for rendering (where the intermediate Ambisonics signal 201 may for example correspond to a stereo signal resulting from a partial upmix operation), thereby allowing to provide a stereo output in an efficient manner.

[0110] It should be noted that the terms "metadata" and "metadata bitstream" are used interchangeably within this specification such that references to "metadata" also refer to "metadata bitstream" and / or references to "metadata" also refer to "metadata".

[0111] Aspects of the system described herein may be implemented in any suitable computer-based audio processing network environment that processes digital or digitized audio files. Portions of the adaptive audio system may include one or more networks including any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0112] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described in terms of their operation as hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media, using any number of combinations of register transfers, logical components, and / or other characteristics. The computer-readable media on which such formatted data and / or instructions are embodied include various forms of physical (non-transitory) non-volatile storage media, such as, but not limited to, optical, magnetic, or semiconductor storage media.

[0113] While one or more implementations have been described by way of example and in terms of specific embodiments, it is to be understood that the one or more implementations are not limited to the disclosed embodiments. On the contrary, the implementations are intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Thus, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.

[0114] Various aspects and implementations of the present invention can be seen in the following enumerated example embodiments (EEE), which are not claimed.

[0115] (EEE1) A method (500) for encoding an Ambisonics input audio signal, the method (500) comprising: providing (501) the input audio signal (101) to a SPAR encoder (110, 130) and a DirAC analyzer and parameter encoder (120); generating (502) an encoder bitstream (106) based on the output (102, 105) of the SPAR encoder (110, 130) and based on the output (104) of the DirAC analyzer and parameter encoder (120); The method includes:

[0116] (EEE2) the output (102, 105) of the SPAR encoder (110, 130) comprises a SPAR metadata bitstream (102) and an audio bitstream (105) representative of a set of SPAR downmix channel signals (103); and / or The method (500) according to EEE1, wherein the output (104) of the DirAC analyzer and parameter encoder (120) comprises a DirAC metadata bitstream (104).

[0117] (EEE3) The method (500) according to EEE2, wherein the step (502) of generating the encoder bitstream (106) comprises multiplexing the SPAR metadata bitstream (102), the audio bitstream (105) and the DirAC metadata bitstream (104) into a common encoder bitstream (106).

[0118] (EEE4) A method (500) according to any of EEE1 to 3, further comprising a step of transmitting a representation of the encoder bitstream (106), in particular to a decoding device (200), and / or storing the representation of the encoder bitstream (106).

[0119] (EEE5) The method (500) generating subband data in a number of frequency bands and / or a number of time / frequency tiles representing the input audio signal (101); selecting a subset of said plurality of frequency bands and / or said plurality of time / frequency tiles; determining an output (104) of the DirAC analyzer and parameter encoder (120), in particular a DirAC metadata bitstream (104), for the selected subset of frequency bands and / or time / frequency tiles, in particular only for the selected subset of frequency bands and / or time / frequency tiles, based on the subband data; The method (500) according to any one of EEE1 to EEE4,

[0120] (EEE6) The method (500) determining characteristic information relating to characteristics of the input audio signal (101), in particular noise-like or timbral characteristics of the input audio signal (101); selecting said subset of frequency bands and / or time / frequency tiles based on said characteristic information; The method according to EEE5 (500),

[0121] (EEE7) The method (500) according to any of EEE5 to 6, wherein said subset of frequency bands and / or time / frequency tiles corresponds to a frequency range of frequencies above a predefined threshold frequency.

[0122] (EEE8) The method (500) generating subband data in a number of frequency bands and / or a number of time / frequency tiles representing the input audio signal (101) using an analysis filterbank; providing the subband data to the SPAR encoder (110, 130) for generating SPAR metadata (102) and to the DirAC analyzer and parameter encoder (120) for generating DirAC metadata (104); The method (500) according to any one of EEE1 to 7,

[0123] (EEE9) The method (500) comprises: generating one or more downmix channel signals (103) in said SPAR encoder (110, 130) using a synthesis filter bank; The method according to claim 8, further comprising:

[0124] (EEE10) A method (510) for decoding an encoder bitstream (106) representative of an Ambisonics input audio signal (101), the method (510) comprising: generating (511) an intermediate Ambisonics signal (201) based on the encoder bitstream (106) using a SPAR decoder (210, 230); processing (512) the intermediate Ambisonics signal (201) using a DirAC synthesizer (220) to provide an output audio signal (211) for rendering; A method (510) comprising:

[0125] (EEE11) The method (510) comprises: extracting a SPAR metadata bitstream (102) and an audio bitstream (105) from the encoder bitstream (106); generating the intermediate Ambisonics signal (201) from the SPAR metadata bitstream (102) and the audio bitstream (105) using the SPAR decoder (210, 230); The method according to claim 10, further comprising:

[0126] (EEE12) The method (510) comprises: generating a set of reconstructed downmix channel signals (205) from said audio bitstream (105) using an audio decoder (230); upmixing said set of reconstructed downmix channel signals (205) to said intermediate Ambisonics signal (201) based on said SPAR metadata bitstream (102) using an upmix unit (210); The method according to claim 8, further comprising:

[0127] (EEE13) The method (510) comprises: extracting a DirAC metadata bitstream (104) from the encoder bitstream (106); processing (512) the intermediate Ambisonics signal (201) using the DirAC synthesizer (220) in dependence on the DirAC metadata bitstream (104) to provide the output audio signal (211); The method according to any one of EEE10 to 12, comprising the steps of:

[0128] (EEE14) The method (510) comprises: processing said intermediate Ambisonics signal (201) in said DirAC analyzer (250) to generate auxiliary DirAC metadata (204); processing (512) the intermediate Ambisonics signal (201) using the DirAC synthesizer (220) in dependence on the auxiliary DirAC metadata (204) to provide the output audio signal (211); The method according to any one of EEE10 to 13, comprising the steps of:

[0129] (EEE15) The method (510) generating sub-band data in a plurality of frequency bands and / or a plurality of time / frequency tiles representing the intermediate Ambisonics signal (201); selecting a subset of said plurality of frequency bands and / or said plurality of time / frequency tiles; determining the auxiliary DirAC metadata (204) for the selected subset of frequency bands and / or time / frequency tiles, in particular only for the selected subset of frequency bands and / or time / frequency tiles, based on the subband data; The method according to claim 8, further comprising the steps of:

[0130] (EEE16) The method (510) comprises: determining characteristic information relating to characteristics of the input audio signal (101) and / or of the intermediate Ambisonics signal (201), in particular noise-like or timbral characteristics of the input audio signal (101) and / or of the intermediate Ambisonics signal (201); selecting said subset of frequency bands and / or time / frequency tiles based on said characteristic information; The method according to claim 8, further comprising:

[0131] (EEE17) The method (510) according to any of EEE15 to 16, wherein said subset of frequency bands and / or time / frequency tiles corresponds to a frequency range of frequencies below a predefined threshold frequency.

[0132] (EEE18) The method (510) comprises: generating an Ambisonics output signal (211) from the intermediate Ambisonics signal (201) using a DirAC synthesizer (220) having an Ambisonics order greater than an Ambisonics order of the input audio signal (101) and / or of the intermediate Ambisonics signal (201); The method according to any one of EEE10 to 17, comprising the steps of:

[0133] (EEE19) The method (510) according to any one of EEE10 to 18, wherein the output signal (211) comprises at least one of an Ambisonics output signal, a binaural output signal, a stereo or multi-speaker output signal.

[0134] (EEE20) The method (510) determining directional data relating to the orientation of the listener's head, in particular using a head tracker; performing a rotation operation on the intermediate Ambisonics signal (201) in response to the orientation data to generate a rotated Ambisonics signal; processing the rotated Ambisonics signal using the DirAC synthesizer (220) to provide the output audio signal (211) for rendering to the listener; The method according to any one of EEE10 to 19, comprising the steps of:

[0135] (EEE21) The method (510) comprises: determining directional data relating to the orientation of the listener's head, in particular using a head tracker; extracting DirAC metadata (104) from the encoder bitstream (106); performing a rotation operation on the DirAC metadata (104) in response to the orientation data to generate rotated DirAC metadata; processing the intermediate Ambisonics signal (201) or an Ambisonics signal derived from the intermediate Ambisonics signal according to the rotated DirAC metadata using the DirAC synthesizer (220) to provide the output audio signal (211) for rendering to the listener; The method according to any one of EEE10 to 20, comprising the steps of:

[0136] (EEE22) The intermediate Ambisonics signal (201) contains fewer channels than the Ambisonics input audio signal (101), and / or the SPAR decoder (210, 230) is used to perform a partial upmixing operation to generate an intermediate Ambisonics signal (201) that includes fewer channels than the Ambisonics input audio signal (101); A method according to any one of EEE10 to 21 (510).

[0137] (EEE23) The partial upmixing operation is performed in a filter bank domain having multiple subbands and / or multiple time / frequency tiles; the intermediate Ambisonics signal (201) comprises fewer channels than the Ambisonics input audio signal (101) for all of the subbands and / or all of the time / frequency tiles, or the intermediate Ambisonics signal (201) includes fewer channels than the Ambisonics input audio signal (101) only for a subset of the subbands and / or the time / frequency tiles; The method according to EEE22 (510).

[0138] (EEE24) The method (510) comprises: extracting an audio bitstream (105) from the encoder bitstream (106); generating a set of reconstructed downmix channel signals (205) from said audio bitstream (105) using an audio decoder (230); applying an analysis filterbank to the set of reconstructed downmix channel signals (205) to transform the set of reconstructed downmix channel signals (205) into a filterbank domain; generating (511) an intermediate Ambisonics signal (201) represented in the filter bank domain based on the set of reconstructed downmix channel signals (205) in the filter bank domain; processing (512) the intermediate Ambisonics signal (201) represented in the filter bank domain using the DirAC synthesizer (220); The method according to any one of EEE10 to 22, comprising the steps of:

[0139] (EEE25) The method (510) comprises: processing (512) the intermediate Ambisonics signal (201) represented in the filter bank domain using the DirAC synthesizer (220) to generate an output signal (211) represented in the filter bank domain; applying a synthesis filterbank to the output signal (211) represented in the filterbank domain to generate an output signal (211) in the time domain; The method of claim 8, further comprising:

[0140] (EEE26) the analysis filter bank and the synthesis filter bank form a joint analysis / synthesis filter bank, in particular a perfect reconstruction analysis / synthesis filter bank; and / or the analysis filter bank and the synthesis filter bank are Nyquist filter banks or QMF filter banks; The method described in EEE25 (510).

[0141] (EEE27) the encoder bitstream (106) has been generated using a filter bank of a first type, in particular a Nyquist filter bank; the analysis filter bank is a filter bank of a second type different from the first type, in particular a QMF filter bank, A method according to any one of EEE24 to 26 (510).

[0142] (EEE28) The method (510) according to EEE27, wherein frequency band boundaries of the first type filter bank are adjusted to corresponding frequency band boundaries of the second type filter bank.

[0143] (EEE29) A system comprising: one or more processors; a non-transitory computer-readable storage medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the operations described in EEE1-28; and A system including:

[0144] (EEE30) A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the operations described in EEE1-28.

[0145] (EEE31) A coding device (100) for coding an Ambisonics input audio signal (101), the coding device (100) comprising: providing said input audio signal (101) to a SPAR encoder (110, 130) and a DirAC analyzer and parameter encoder (120); generating an encoder bitstream (106) based on the output (102, 105) of said SPAR encoder (110, 130) and based on the output (104) of said DirAC analyzer and parameter encoder (120); The encoding device (100) is configured as follows.

[0146] (EEE32) The Ambisonics input audio signal (101) comprises a plurality of input channel signals, and the SPAR encoder (110, 130) downmixing said plurality of input channel signals in the subband and / or QMF domain into one or more downmix channel signals (130); generating a SPAR metadata bitstream (102) in the subband and / or QMF domain adapted to upmix the one or more downmix channel signals (103) into a plurality of reconstructed channel signals of a reconstructed Ambisonics signal (201); The encoding device (100) according to EEE31,

[0147] (EEE33) An encoding device (100) according to any of EEE31 to 32, wherein the DirAC analyzer and parameter encoder (130) is configured to perform a direction of arrival analysis on an Ambisonics input audio signal (101) in the subband and / or QMF domain to determine a DirAC metadata bitstream (104) indicating the direction of arrival of one or more dominant components of the Ambisonics input audio signal (101).

[0148] (EEE34) A decoding device (200) for decoding an encoder bitstream (106) representing an Ambisonics input audio signal (101), the decoding device (200) comprising: generating an intermediate Ambisonics signal (201) using a SPAR decoder (210, 230) based on said encoder bitstream (106); processing the intermediate Ambisonics signal (210) using a DirAC synthesizer (220) to provide an output audio signal (211) for rendering; A decoding device (200) configured as follows.

[0149] (EEE35) The decoding device (200) placing the DirAC combiner (220) in a pass-through mode of operation; and / or Bypassing the DirAC synthesizer (220); The decoding device (200) according to EEE34, configured in particular for the intermediate Ambisonics signal (201) to correspond to the output audio signal (211) for rendering.

Claims

1. 1. A method for encoding an Ambisonics input audio signal, the method comprising: providing the input audio signal to a SPAR encoder and a DirAC analyzer and parameter encoder; generating an encoder bitstream based on the output of the SPAR encoder and based on the outputs of the DirAC analyzer and parameter encoder; A method comprising:

2. the output of the SPAR encoder comprises a SPAR metadata bitstream and an audio bitstream representing a set of SPAR downmix channel signals; and / or The method of claim 1 , wherein the output of the DirAC analyzer and parameter encoder comprises a DirAC metadata bitstream.

3. The method of claim 2 , wherein generating the encoder bitstream comprises multiplexing the SPAR metadata bitstream, the audio bitstream, and the DirAC metadata bitstream into a common encoder bitstream.

4. The method comprises: generating subband data in a plurality of frequency bands and / or a plurality of time / frequency tiles representing the input audio signal; selecting a subset of said plurality of frequency bands and / or said plurality of time / frequency tiles; - determining an output of the DirAC analyzer and parameter encoder, in particular a DirAC metadata bitstream, for the selected subset of frequency bands and / or time / frequency tiles, in particular only for the selected subset of frequency bands and / or time / frequency tiles, based on the subband data; Further comprising: The method of claim 1 , wherein the subset of frequency bands and / or time / frequency tiles corresponds to a frequency range of frequencies below a predetermined threshold frequency.

5. The method comprises: determining characteristic information relating to characteristics of the input audio signal, in particular noise-like or tonal characteristics of the input audio signal; selecting the subset of frequency bands and / or time / frequency tiles based on the characteristic information; The method of claim 4 further comprising:

6. The method comprises: generating subband data in a plurality of frequency bands and / or a plurality of time / frequency tiles representing the input audio signal using an analysis filterbank; providing the subband data to the SPAR encoder to generate SPAR metadata and to the DirAC analyzer and parameter encoder to generate DirAC metadata; generating one or more downmix channel signals in the SPAR encoder using a synthesis filter bank; The method of any one of claims 1 to 5, further comprising:

7. 1. A method for decoding an encoder bitstream representing an Ambisonics input audio signal, the method comprising: generating an intermediate Ambisonics signal using a SPAR decoder based on the encoder bitstream; processing the intermediate Ambisonics signal using a DirAC synthesizer to provide an output audio signal for rendering; A method comprising:

8. The method comprises: extracting a SPAR metadata bitstream and an audio bitstream from the encoder bitstream; generating the intermediate Ambisonics signal from the SPAR metadata bitstream and the audio bitstream using the SPAR decoder; The method of claim 7 further comprising:

9. The method comprises: generating a set of reconstructed downmix channel signals from the audio bitstream using an audio decoder; upmixing, using an upmix unit, the set of reconstructed downmix channel signals into the intermediate Ambisonics signal based on the SPAR metadata bitstream; The method of claim 8 further comprising:

10. The method comprises: extracting a DirAC metadata bitstream from the encoder bitstream; processing the intermediate Ambisonics signal using the DirAC synthesizer in dependence on the DirAC metadata bitstream to provide the output audio signal; The method of claim 7 further comprising:

11. The method comprises: processing the intermediate Ambisonics signal in a DirAC analyzer to generate auxiliary DirAC metadata; processing the intermediate Ambisonics signal using the DirAC synthesizer in dependence on the auxiliary DirAC metadata to provide the output audio signal; The method of claim 7 further comprising:

12. The method comprises: generating subband data in a plurality of frequency bands and / or a plurality of time / frequency tiles representing the intermediate Ambisonics signal; selecting a subset of said plurality of frequency bands and / or said plurality of time / frequency tiles; - determining the auxiliary DirAC metadata for the selected subset of frequency bands and / or time / frequency tiles, in particular only for the selected subset of frequency bands and / or time / frequency tiles, based on the subband data; The method of claim 11 further comprising:

13. The method comprises: determining characteristic information relating to characteristics of the input audio signal and / or the intermediate Ambisonics signal, in particular noise-like or tonal characteristics of the input audio signal and / or the intermediate Ambisonics signal; selecting the subset of frequency bands and / or time / frequency tiles based on the characteristic information; The method of claim 12 further comprising:

14. The method comprises: determining directional data relating to the orientation of the listener's head, in particular using a head tracker; performing a rotation operation on the intermediate Ambisonics signal in response to the orientation data to generate a rotated Ambisonics signal; processing the rotated Ambisonics signal using the DirAC synthesizer to provide the output audio signal for rendering to the listener; The method of claim 7 further comprising:

15. The method comprises: determining directional data relating to the orientation of the listener's head, in particular using a head tracker; extracting DirAC metadata from the encoder bitstream; performing a rotation operation on the DirAC metadata in response to the orientation data to generate rotated DirAC metadata; processing the intermediate Ambisonics signal or an Ambisonics signal derived from the intermediate Ambisonics signal in accordance with the rotated DirAC metadata using the DirAC synthesizer to provide the output audio signal for rendering to the listener; The method of any one of claims 7 to 14, further comprising: