Method and device for encoding and / or decoding immersive audio signal

The method for encoding immersive audio signals through downmixing and energy compaction addresses the challenge of transmitting high-quality immersive audio efficiently, enabling flexible rendering and reduced bandwidth usage.

JP2025170395APending Publication Date: 2025-11-18DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025145017
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-07-02
Filing Date
2025-09-02
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies face challenges in transmitting and storing immersive audio signals with high perceptual quality in a bandwidth-efficient manner.

Method used

A method for encoding immersive audio signals involves determining downmix channel signals, performing energy compaction, and generating joint encoding metadata to allow efficient encoding and decoding of these signals, using techniques such as energy compaction, joint encoding, and upmixing processes.

Benefits of technology

Enables efficient encoding and decoding of immersive audio signals with high perceptual quality, allowing flexible rendering on various speaker arrangements and reducing bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170395000001_ABST
    Figure 2025170395000001_ABST
Patent Text Reader

Abstract

To provide a method and device for decoding immersive audio signals.SOLUTION: A decoding unit is a device for decoding immersive audio signals, which comprises: a decode module that derives multiple reconstructed channel signals from encoded audio data, derives congruence coding spatial audio resolution reconstruction (SPAR) metadata from encoded metadata, and derives object metadata; and a reconstruction module that derives reconstructed multichannel signals from the congruence coding SPAR metadata and the multiple reconstructed channel signals.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 693,246, filed July 2, 2018, the contents of which are incorporated herein by reference.

[0002] Technical Field This document relates to immersive audio signals that may include sound field representation signals, in particular Ambisonics signals. In particular, this document relates to providing an encoder and a corresponding decoder that allows immersive audio signals to be transmitted and / or stored in a bitrate-efficient manner and / or with high perceptual quality. [Background technology]

[0003] The sound or sound field in a listening environment of a listener positioned at a listening position can be described using Ambisonics signals. Ambisonics signals can be viewed as multi-channel audio signals, where each channel corresponds to a specific directional pattern of the sound field at the listener's listening position. Ambisonics signals may be described using a three-dimensional (3D) Cartesian coordinate system, where the origin of the coordinate system corresponds to the listening position, the x-axis points forward, the y-axis points left, and the z-axis points upward.

[0004] By increasing the number of audio signals or channels and the corresponding number of directivity patterns (and corresponding panning functions), the accuracy of the sound field description can be increased. As an example, a first-order Ambisonics signal contains four channels or waveforms: a W channel that describes the omnidirectional component of the sound field, an X channel that describes the sound field with a dipole directivity pattern corresponding to the x-axis, a Y channel that describes the sound field with a dipole directivity pattern corresponding to the y-axis, and a Z channel that describes the sound field with a dipole directivity pattern corresponding to the z-axis. A second-order Ambisonics signal has nine channels, including the four channels of the first-order Ambisonics signal (also called B format) and five additional channels for different directivity patterns. Generally, an L-order Ambisonics signal is a multiple of the L-order Ambisonics signal. 2 channels and [(L+1) 2 -L 2 ] additional channels (L+1) 2 channels (when using the 3D Ambisonics format). An L-th order Ambisonics signal, for L>1, is sometimes called a Higher-Order Ambisonics (HOA) signal.

[0005] The HOA signal may be used to describe a 3D sound field independent of the arrangement of speakers used to render the HOA signal. Examples of speaker arrangements include headphones, or one or more arrangements of loudspeakers, or a virtual reality rendering environment. Thus, it may be beneficial to provide the HOA signal to an audio renderer to allow the audio rendering to flexibly adapt to different arrangements of speakers. Summary of the Invention [Problem to be solved by the invention]

[0006] A soundfield representation (SR) signal, such as an Ambisonics signal, may be complemented with an audio object and / or a multi-channel (bed) signal to provide an immersive audio (IA) signal. This paper addresses the technical problem of transmitting and / or storing IA signals with high perceptual quality in a bandwidth-efficient manner. Such technical problem is solved by the independent claims. Preferred examples are set out in the dependent claims. [Means for solving the problem]

[0007] According to one aspect, a method for encoding a multi-channel input signal is described. The multi-channel input signal may be part of an immersive audio (IA) signal. The multi-channel input signal may include a sound field representation (SR) signal, particularly a first-order or higher-order Ambisonics signal. The method includes determining a plurality of downmix channel signals from the multi-channel input signal. The method further includes performing energy compaction of the plurality of downmix channel signals to provide a plurality of compacted channel signals. The method further includes determining joint encoding metadata (particularly spatial audio resolution reconstruction (SPAR) metadata) based on the plurality of compacted channel signals and based on the multi-channel input signal, the joint encoding metadata allowing the plurality of compacted channel signals to be upmixed to an approximation of the multi-channel input signal. The method further includes encoding the plurality of compacted channel signals and the joint encoding metadata.

[0008] According to a further aspect, a method is described for determining a reconstructed multi-channel signal from encoded audio data indicative of a plurality of reconstructed channel signals and from encoded metadata indicative of joint encoding metadata, the method including decoding the encoded audio data to provide the plurality of reconstructed channel signals and decoding the encoded metadata to provide the joint encoding metadata, and further including determining the reconstructed multi-channel signal from the plurality of reconstructed channel signals using the joint encoding metadata.

[0009] According to a further aspect, a software program is described, which may be adapted for execution on a processor and, when executed on the processor, to perform the method steps outlined herein.

[0010] According to another aspect, a storage medium is described that may include a software program adapted for execution on a processor and, when executed on the processor, to perform the method steps outlined herein.

[0011] According to a further aspect, a computer program product is described, which may include executable instructions for performing the method steps outlined herein when executed on a computer.

[0012] According to another aspect, an encoding unit or encoding device for encoding a multi-channel input signal and / or an immersive audio (IA) signal is described. The encoding unit is configured to determine a plurality of downmix channel signals from the multi-channel input signal. The encoding unit is further configured to perform energy compaction of the plurality of downmix channel signals to provide a plurality of compacted channel signals. The encoding unit further includes determining joint encoding metadata based on the plurality of compacted channel signals and based on the multi-channel input signal, the joint encoding metadata allowing for upmixing the plurality of compacted channel signals to an approximation of the multi-channel input signal. The encoding unit is further configured to encode the plurality of compacted channel signals and the joint encoding metadata.

[0013] According to another aspect, a decoding unit or decoding apparatus is described for determining a reconstructed multi-channel signal from encoded audio data indicative of a plurality of reconstructed channel signals and from encoded metadata indicative of joint encoding metadata. The decoding unit includes decoding the encoded audio data to provide the plurality of reconstructed channel signals and decoding the encoded metadata to provide the joint encoding metadata. The decoding unit further includes determining the reconstructed multi-channel signal from the plurality of reconstructed channel signals using the joint encoding metadata.

[0014] It should be noted that the methods, devices, and systems outlined in this patent application, including preferred embodiments thereof, may be used independently or in combination with other methods, devices, and systems disclosed herein. Furthermore, all aspects of the methods, devices, and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief explanation of the drawings]

[0015] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1] 1 shows an example of a coding system. [Figure 2] 1 illustrates an exemplary encoding unit for encoding an immersive audio signal. [Figure 3] 1 shows another exemplary decoding unit for decoding an immersive audio signal; [Figure 4] 1 illustrates exemplary encoding and decoding units for encoding and decoding immersive audio signals. [Figure 5] 1 illustrates exemplary encoding and decoding units with mode switching. [Figure 6] 1 illustrates an exemplary reconstruction module. [Figure 7] 1 shows a flowchart of an exemplary method for encoding an immersive audio signal. [Figure 8] 1 shows a flowchart of an exemplary method for decoding an immersive audio signal. DETAILED DESCRIPTION OF THE INVENTION

[0016] As outlined above, this paper relates to the efficient encoding of immersive audio (IA) signals, such as first-order ambisonics (FOA) or HOA signals, multi-channel and / or object-based audio signals, where FOA or HOA signals in particular are more generally referred to herein as soundfield representation (SR) signals.

[0017] As outlined in the introduction, SR signals may contain a relatively large number of channels or waveforms, with different channels relating to different panning functions and / or different directional patterns. As an example, a 3D FOA or HOA signal of order L may have (L+1) 2 The SR signal can be represented in a variety of different formats.

[0018] The sound field can be considered to consist of one or more sound events emanating from any direction around the listening position, and consequently the position of said one or more sound events may be defined on the surface of a sphere (with the listening position or reference position at the center of the sphere).

[0019] Sound field formats like FOA or Higher Order Ambisonics (HOA) are defined in such a way that they allow the sound field to be rendered with any speaker arrangement (i.e., any rendering system). However, rendering systems (such as Dolby Atmos systems) are typically constrained in the sense that the possible speaker heights are fixed to a defined number of planes (e.g., the (horizontal) plane at ear height, the ceiling or upper plane, and / or the floor or lower plane). Thus, the concept of an ideal spherical sound field can be modified to a sound field consisting of sound objects located in different rings (similar to the stacked rings that make up a honeycomb) at various heights on the surface of a sphere.

[0020] As shown in FIG. 1, the audio coding system 100 includes an encoding unit 110 and a decoding unit 120. The encoding unit 110 may be configured to generate a bitstream 101 for transmission to the decoding unit 120 based on an input signal 111, which may include an immersive audio signal (e.g., used for virtual reality (VR) applications). The immersive audio signal 111 may include an SR signal, a multi-channel (bed) signal, and / or multiple objects (each object including an object signal and object metadata). The decoding unit 120 may be configured to provide an output signal 121 based on the bitstream 101, which may include a reconstructed immersive audio signal.

[0021] 2 shows examples of encoding units 110, 200. The encoding unit 200 may be configured to encode an input signal 111, which may be an immersive audio (IA) signal 111. The IA signal 111 may include a multi-channel input signal 201. The multi-channel input signal 201 may include an SR signal and one or more object signals. Furthermore, object metadata 202 for the plurality of object signals may be provided as part of the IA signal 111. The IA input signal 111 may be provided by a content ingestion engine, which may be configured to derive the object and / or SR signals from (composite) VR content.

[0022] The encoding unit 200 comprises a downmix module 210 configured to downmix a multi-channel input signal 201 into a plurality of downmix channel signals 203. The plurality of downmix channel signals 203 may correspond to SR signals, in particular First Order Ambisonics (FOA) signals. The downmix may be performed in the subband domain or in the QMF domain (e.g., using 10 or more subbands).

[0023] The encoding unit 200 further comprises a joint encoding module 230 (in particular a SPAR module) configured to determine joint encoding metadata 205 (in particular SPAR (Spatial Audio Resolution Reconstruction) metadata) configured to reconstruct the multi-channel input signal 201 from the plurality of downmix channel signals 203. The joint encoding module 230 may be configured to determine the joint encoding metadata 205 in the subband domain.

[0024] To determine the joint coding metadata 205, the multiple downmix channel signals 203 may be transformed into and / or processed in the subband domain. Furthermore, the multi-channel input signal 201 may be transformed into the subband domain. The joint coding metadata 205 may then be determined for each subband, and in particular, an approximation of the subband signal of the multi-channel input signal 201 may be obtained by upmixing the subband signals 203 of the multiple downmix channel signals 203 using the joint coding metadata 205. The joint coding metadata 205 for the various subbands may be inserted into the bitstream 101 for transmission to the corresponding decoding unit 120.

[0025] Additionally, the encoding unit 200 may include an encoding module 240 configured to perform waveform encoding of the multiple downmix channel signals 203, thereby providing encoded audio data 206. Each of the downmix channel signals 203 may be encoded using a mono waveform encoder (e.g., 3GPP® EVS encoding), thereby enabling efficient encoding. Further examples of encoding the multiple downmix channel signals 203 include MPEG AAC, MPEG HE-AAC, and other MPEG audio codecs, 3GPP® codecs, Dolby Digital / Dolby Digital Plus (AC-3, eAC-3), Opus, LC-3, and other similar codecs. As a further example, encoding tools included in the AC-4 codec may be configured to perform the operations of the encoding unit 200.

[0026] Furthermore, the encoding module 240 may be configured to perform entropy encoding of the jointly encoded metadata (i.e., SPAR metadata) 205 and the object metadata 202, thereby providing encoded metadata 207. The encoded audio data 206 and the encoded metadata 207 may be inserted into the bitstream 101.

[0027] 3 illustrates an example of a decoding unit 120, 350. The decoding unit 120, 350 may include a receiver that receives a bitstream 101, which may include encoded audio data 206 and encoded metadata 207. The decoding unit 120, 350 may include a processor and / or demultiplexer that demultiplexes the encoded audio data 206 and encoded metadata 207 from the bitstream 101. The decoding unit 350 has a decoding module 360 ​​configured to derive multiple reconstructed channel signals 314 from the encoded audio data 206. The decoding module 360 ​​may be further configured to derive joint encoded metadata 205 and object metadata 202 from the encoded metadata 207.

[0028] Furthermore, the decoding unit 350 includes a reconstruction module 370 configured to derive a reconstructed multi-channel signal 311 from the joint coding metadata 205 and from the multiple reconstructed channel signals 314. The joint coding metadata 205 may convey time- and / or frequency-varying elements of an upmix matrix that allows for the reconstruction of the multi-channel signal 311 from the multiple reconstructed channel signals 314. The upmix process may be performed in the QMF (Quadrature Mirror Filter) subband domain. Alternatively, other time / frequency transforms, particularly those based on the FFT (Fast Fourier Transform), may be used to perform the upmix process. In general, transforms that allow frequency-selective analysis and (upmix) processing may be applied. The upmix process may also include a decorrelator that allows for improved reconstruction of the covariance of the reconstructed multi-channel signal 311, and the decorrelator may be controlled by the additional joint coding metadata 205.

[0029] The reconstructed multi-channel signal 311 may include a reconstructed sound reinforcement signal and one or more reconstructed object signals. The reconstructed multi-channel signal 311 and object metadata may form the reconstructed IA signal 121. The reconstructed IA signal 121 may be used for speaker rendering 330, headphone rendering 331, and / or sound reinforcement rendering 332, for example.

[0030] 4 illustrates an encoding unit 200 and a decoding unit 350. The encoding unit 200 has the components described in the context of FIG. 2. Furthermore, the encoding unit 200 includes an energy compaction module 420 configured to concentrate the energy of the multiple downmix channel signals 203 into one or more downmix channel signals 203. The energy compaction module 420 may transform the downmix channel signals 203 to provide multiple compacted channel signals 404. The transformation may be performed such that one or more of the compacted channel signals 404 have less energy than the corresponding one or more downmix channel signals 203.

[0031] For example, the multiple downmix channel signals 203 may include a W channel signal, an X channel signal, a Y channel signal, and a Z channel signal. The multiple compacted channel signals 404 may include a W channel signal, an X' channel signal, a Y' channel signal, and a Z' channel signal. The X' channel signal, the Y' channel signal, and the Z' channel signal may be determined such that the X' channel signal has less energy than the X channel signal, the Y' channel signal has less energy than the Y channel signal, and / or the Z' channel signal has less energy than the Z channel signal.

[0032] The energy compaction module 420 may be configured to perform energy compaction using a prediction operation. In particular, a first subset of the plurality of downmix channel signals 203 (e.g., the X-channel signal, the Y-channel signal, and the Z-channel signal) may be predicted from a second subset of the plurality of downmix channel signals 203 (e.g., the W-channel signal). The energy compaction may include subtracting a scaled version of one of the downmix channel signals 203 (e.g., the W-channel signal) from the other downmix channel signals 203 (e.g., the X-channel signal, the Y-channel signal, and / or the Z-channel signal). The scaling factor may be determined such that the energy of the other downmix channel signals 203 is reduced, in particular minimized.

[0033] By performing energy compaction, the efficiency for encoding the multiple compacted channel signals 404 may be increased compared to encoding the multiple downmix channel signals 203. The encoding unit 200 is configured to implicitly insert metadata for performing the inverse operation of the energy compaction operation into the joint coding metadata 205. As a result, efficient encoding of the IA input signal 111 is achieved.

[0034] As outlined above, the decoding unit includes a reconstruction module 370. FIG. 6 illustrates an exemplary reconstruction module 370. The reconstruction module 370 receives as input a plurality of reconstructed channel signals 314 (which may, for example, form a first-order Ambisonics signal). A first mixer 611 may be configured to upmix the plurality of reconstructed channel signals 314 (e.g., the four channel signals) into a larger number of signals (e.g., eleven signals representing a second Ambisonics signal and two object signals). The first mixer 611 relies on the jointly encoded metadata 205.

[0035] The reconstruction module 370 may have decorrelators 601, 602 configured to generate two signals from the W channel signal, which are processed in a second mixer 612 to produce an increased number of signals (e.g., 11 signals). The second mixer 612 depends on the jointly encoded metadata 205. The outputs of the first mixer 611 and the second mixer 612 are summed to provide the reconstructed multi-channel signal 311.

[0036] As mentioned above, the joint coding or SPAR metadata 205 may consist of data representing the coefficients of the upmix matrices used by the first mixer 611 and the second mixer 612. The mixers 611, 612 may operate in the subband domain (particularly the QMF domain). In this case, the joint coding or SPAR metadata 205 includes data representing the coefficients of the upmix matrices used by the first mixer 611 and the second mixer 612 for multiple different subbands (e.g., 10 or more subbands).

[0037] 5 shows an encoding unit 200 with two branches, one for encoding a multi-channel input signal 201 and one for encoding object metadata 202 (which forms the IA input signal 111). The upper branch corresponds to the encoding scheme described in the context of FIG. 4. In the lower branch, the joint encoding unit 230 is modified to determine metadata 205 that allows the multiple downmix channel signals 203 to be reconstructed from the multiple compacted channel signals 404. The metadata 205 thus indicates the predictor (in particular the scaling factor or factors) used to generate the multiple compacted channel signals 404 from the multiple downmix channel signals 203. In a variant, the metadata 205 may be provided directly from the energy compaction module 220 (without the need to use the joint encoding module 230).

[0038] The encoding unit 200 of Figure 5 includes a mode switching module 500 configured to switch between a first mode (corresponding to the upper branch) and a second mode (corresponding to the lower branch). The first mode may be used to provide high perceptual quality at an increased bit rate, and the second mode may be used to provide reduced perceptual quality at a reduced bit rate. The mode switching module 500 may be configured to switch between the first mode and the second mode depending on the conditions of the transmission network.

[0039] 5 further illustrates a corresponding decoding unit 350 configured to perform decoding according to a first mode (upper branch) and a second mode (lower branch). The mode switching module 550 may be configured to determine the mode used by the encoding unit 200 (e.g., frame by frame). If the first mode is used, a reconstructed multi-channel signal 311 and object metadata 202 may be determined (as outlined in the context of FIG. 4). On the other hand, if the second mode is used, a plurality of reconstructed downmix channel signals 513 (corresponding to the plurality of downmix channel signals 203) may be determined by the decoding unit 350.

[0040] Thus, an encoding unit 200 is described having a downmix module 210 configured to process the object and HOA input signals 111 to generate a reduced-channel output signal 203, e.g., a first-order Ambisonics signal. A SPAR encoding module 230 generates metadata (i.e., SPAR metadata) 205 that describes how the original inputs 111, 201 (e.g., object signals and HOA) are regenerated from the FOA signal 203. A set of EVS encoders 240 receives the four-channel FOA signal 203 and generates encoded audio data 206 that is inserted into the bitstream 101. The audio data is then decoded by a set of EVS decoders 360 to generate a four-channel FOA signal 314. The SPAR metadata 205 may be provided to the decoder 360 as (entropy-) coded metadata 207 in the bitstream 101. A reconstruction module 370 then regenerates the output 121 consisting of the audio object and HOA signals.

[0041] The low-resolution signal 203 produced by the downmix module 210 may be modified by a WXYZ energy compaction transform (in module 420), which produces an output signal 404 with less inter-channel correlation compared to the output of the downmix module 210. The purpose of the energy compaction filter 420 is to reduce the energy in the XYZ channels so that the W channel can be encoded at a higher bitrate and the low-energy X'Y'Z' channels can be encoded at a lower bitrate. In this way, coding artifacts are more effectively masked, thereby improving audio quality.

[0042] Additionally or alternatively to performing prediction, energy compaction can use the Karhunen-Loeve transform (KLT), principal component analysis (PCA) transform, and / or singular value decomposition (SVD) transform. In particular, an energy compaction filter 420 including a whitening filter, a KLT, a PCA transform, and / or an SVD transform may be used. The whitening filter may be implemented using the prediction scheme described above. In particular, the energy compaction filter 420 may include a combination of a whitening filter and a KLT, PCA, and / or SVD transform, the latter arranged in series with the whitening filter. The KLT, PCA, and / or SVD transform may be applied to the X, Y, and Z channels, in particular to the prediction residual.

[0043] 7 shows a flowchart of an exemplary method 700 for encoding a multi-channel input signal 201. In particular, the method 700 is directed to encoding an IA signal that includes the multi-channel input signal 201. The multi-channel input signal 201 may include a sound field representation (SR) signal. In particular, the multi-channel input signal 201 may include a combination of an SR signal (e.g., an HOA signal, in particular a second-order Ambisonics signal) and one or more (in particular two) object signals of one or more audio objects 303.

[0044] The method 700 includes determining 701 a plurality of downmix channel signals 203 from a multi-channel input signal 201. The plurality of downmix channel signals 203 may include a reduced number of channels compared to the multi-channel input signal 201. As described above, the multi-channel input signal 201 may include an SR signal, in particular an L-th order Ambisonics signal, where L≧1, and one or more object signals of one or more audio objects 303. The plurality of downmix channel signals 203 may be determined by downmixing the multi-channel input signal 201 to an SR signal, in particular a K-th order Ambisonics signal, where L≧K. Thus, the plurality of downmix channel signals 203 may be an SR signal, in particular a K-th order Ambisonics signal.

[0045] In particular, determining 701 the multiple downmix channel signals 203 may involve mixing one or more object signals of one or more audio objects 303 (of the multi-channel input signal 201) with an SR signal (or a downmixed version of the SR signal) of the multi-channel input signal 201. The mixing (in particular panning) may be performed depending on the object metadata 202 of one or more audio objects 303, which indicates the spatial position of the audio objects 303. Downmixing an SR signal may involve mixing an SR signal from an L-th order SR signal [(L+1) 2 -L 2 ] additional channels to provide an (L-1)th order SR signal.

[0046] In a preferred example, the multiple downmix channel signals 203 form a first order Ambisonics signal, in particular in B-format or A-format. The SR signal of the multi-channel input signal 201 may also be a second order (or higher) Ambisonics signal.

[0047] Furthermore, the method 700 comprises performing 702 energy compaction of the plurality of downmix channel signals 203 to provide the plurality of compacted channel signals 404. The number of channels of the plurality of downmix channel signals 203 and the plurality of compacted channel signals 404 may be the same. In particular, the plurality of compacted channel signals 404 may form or be in a first-order Ambisonics signal format, in particular B-format or A-format.

[0048] The energy compaction may be performed such that inter-channel correlation between the different channel signals 203 is reduced. In particular, the plurality of compacted channel signals 404 may exhibit less inter-channel correlation than the plurality of downmix channel signals 203. Alternatively or additionally, the energy compaction may be performed such that the energy of a compacted channel signal is less than or equal to the energy of the corresponding downmix channel signal. This condition may be fulfilled for each channel.

[0049] Performing energy compaction 702 may include predicting a first downmix channel signal 203 (e.g., an X, Y, or Z channel) from a second downmix channel signal (e.g., a W channel) to provide a first predicted channel signal, which may be subtracted from the first downmix channel signal 203 (or vice versa) to provide a first compacted channel signal 404.

[0050] Predicting the first downmix channel signal 203 from the second downmix channel signal 203 may include determining a scaling factor for scaling the second downmix channel signal 203. The scaling factor may be determined such that the energy of the first compacted channel signal 404 is reduced compared to the energy of the first downmix channel signal 203 and / or such that the energy of the first compacted channel signal 404 is minimized. The first predicted channel signal may then correspond to the second downmix channel signal 203 scaled according to the scaling factor. Different scaling factors may be determined for different channels.

[0051] In particular (for first-order Ambisonics signals), performing energy compaction 702 may include predicting the X channel signal, the Y channel signal, and the Z channel signal from the W channel signal of the multiple downmix channel signals 203 to provide a predicted X channel signal, a predicted Y channel signal, and a predicted Z channel signal, respectively. The predicted X channel signal may be subtracted from the X channel signal (or vice versa) to determine the X' channel signal of the multiple compacted channel signals 404. The predicted Y channel signal may be subtracted from the Y channel signal (or vice versa) to determine the Y' channel signal of the multiple compacted channel signals 404. The predicted Z channel signal may be subtracted from the Z channel signal (or vice versa) to determine the Z' channel signal of the multiple compacted channel signals 404. Furthermore, the W channel signal of the multiple downmix channel signals 203 may be used as the W channel signal of the multiple compacted channel signals 404.

[0052] As a result of this, the energy of all channels (except one, namely the W channel) may be reduced, thereby allowing efficient encoding of the multiple compacted channel signals 404.

[0053] The method 700 may further include determining 703 joint encoding metadata (also referred to herein as SPAR metadata) 205 based on the plurality of compacted channel signals 404 and based on the multi-channel input signal 201. The joint encoding metadata 205 may be determined such that the joint encoding metadata 205 allows for upmixing the plurality of compacted channel signals 404 into an approximation of the multi-channel input signal 201. By utilizing the plurality of compacted channel signals 404 to determine the joint encoding metadata, a process for inverting energy compaction is automatically included in the joint encoding metadata 205 (no additional metadata specific to inverting the energy compaction operation needs to be provided).

[0054] The joint coding metadata 205 may include upmix data, in particular one or more upmix matrices, allowing the multiple compacted channel signals 404 to be upmixed into an approximation of the multi-channel input signal 201, which approximation comprises the same number of channels as the multi-channel input signal 201. Furthermore, the joint coding metadata 205 may include decorrelation data allowing the reconstruction of the covariance of the multi-channel input signal 201.

[0055] The joint coding metadata 205 may be determined for multiple different subbands (e.g., for 10 or more subbands, particularly in the QMF domain) of the multi-channel input signal 201. By providing the joint coding metadata 205 for the different subbands (i.e., within different frequency bands), an accurate upmix operation can be performed.

[0056] Further, the method 700 includes encoding 704 the plurality of compacted channel signals 404 and jointly encoded metadata 205 (also known as SPAR metadata). The encoding 704 of the plurality of compacted channel signals 404 may include performing waveform encoding (particularly EVS encoding) of each of the plurality of compacted channel signals 404, particularly using a mono encoder for each compacted channel signal 404. Alternatively or additionally, the jointly encoded metadata 205 may be encoded using an entropy encoder. As mentioned above, the multi-channel input signal 201 may include one or more object signals of one or more audio objects 303. In such a case, the method 700 may include encoding object metadata 202 for the one or more audio objects 303, particularly using an entropy encoder.

[0057] The method 700 allows a multi-channel input signal 201, which may represent an SR signal and / or one or more audio object signals, to be encoded in a bitrate-efficient manner, while enabling a decoder to reconstruct the multi-channel input signal 201 with high perceptual quality.

[0058] Determining the jointly encoded metadata 205 based on the plurality of compacted channel signals 404 and based on the multi-channel input signal 201 may correspond to a first mode for encoding the multi-channel input signal 201.

[0059] Alternatively or additionally to using prediction, performing energy compaction 702 may include applying a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform to at least some of the plurality of downmix channel signals 203. By doing so, the coding efficiency of the plurality of compacted channel signals 404 may be further improved.

[0060] In particular, a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform may be applied to the compacted channel signal 404, which corresponds to a prediction residual derived based on the second downmix channel signal 203 (in particular, based on the W channel signal). In other words, a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform may be applied to the prediction residual.

[0061] As mentioned above, in the context of prediction, the X' channel signal, the Y' channel signal, and the Z' channel signal may be derived based on the W channel signal of the plurality of downmix channel signals 203 forming the Ambisonics signal. In particular, the X' channel signal may correspond to the X channel signal minus a prediction of the X channel signal based on the W channel signal. Similarly, the Y' channel signal may correspond to the Y channel signal minus a prediction of the Y channel signal based on the W channel signal. Similarly, the Z' channel signal may correspond to the Z channel signal minus a prediction of the Z channel signal based on the W channel signal. The plurality of compacted channel signals 404 may be determined based on or correspond to the W channel signal, the X' channel signal, the Y' channel signal, and the Z' channel signal.

[0062] To further increase the coding efficiency of the multiple compacted channel signals 404, a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform may be applied to the X′ channel signal, the Y′ channel signal, and the Z′ channel signal to provide an X″ channel signal, a Y″ channel signal, and a Z″ channel signal. The multiple compacted channel signals 404 may then be determined based on the W channel signal, the X″ channel signal, the Y″ channel signal, and the Z″ channel signal.

[0063] In the second mode, the joint coding metadata 205 may be determined based on the plurality of compacted channel signals 404 and based on the plurality of downmix channel signals 203. The joint coding metadata 205 may be determined such that the joint coding metadata 205 allows for reconstructing the plurality of downmix channel signals 203 from the plurality of compacted channel signals 404. In particular, the joint coding metadata 205 may be determined such that the joint coding metadata 205 (only) reverses or inverts the energy compaction operation (without performing an upmix operation). The second mode may be used to reduce the bit rate (with reduced perceptual quality).

[0064] As described above, the multi-channel input signal 201 may include an SR signal and one or more object signals. The first and second modes may allow reconstruction of the SR signal (based on the multiple compacted channel signals 404). Thus, the listener's overall listening experience may be maintained (even when using the second mode).

[0065] The multi-channel input signal 201 may include a sequence of frames. The processing described herein may be performed on a frame-by-frame basis for each frame of the sequence of frames. In particular, the method 700 may include determining for each frame of the sequence of frames whether to use a first mode or a second mode. This allows the encoding to quickly adapt to changing conditions in the transmission network.

[0066] The method 700 may include generating a bitstream 101 based on encoded audio data 206 derived by encoding 704 the plurality of compacted channel signals 404 and based on encoded metadata 207 derived by encoding 704 the jointly encoded metadata 205. Additionally, the method 700 may include inserting an indication into the bitstream 101 indicating whether the second mode or the first mode was used. The indication may be inserted on a frame-by-frame basis. As a result, a corresponding decoding unit 350 can adapt its decoding in a reliable manner.

[0067] 8 shows a flowchart of an example method 800 for determining a reconstructed multi-channel signal 311 from encoded audio data 206 indicative of a plurality of reconstructed channel signals 314 and from encoded metadata 207 indicative of jointly encoded metadata 205. The method 800 may include extracting the encoded audio data 206 and the encoded metadata 207 from the bitstream 101.

[0068] Further, the method 800 may include decoding 801 the encoded audio data 206 to provide the plurality of reconstructed channel signals 314 and decoding the encoded metadata 207 to provide the jointly encoded metadata 205. In one preferred example, the plurality of reconstructed channel signals 203 form a first-order Ambisonics signal, in particular in B-format or A-format.

[0069] The decoding 801 of the encoded audio data 206 may include waveform decoding of each of the multiple reconstructed channel signals 314, particularly using a mono decoder (e.g., an EVS decoder) for each reconstructed channel signal 314. The encoded metadata 207 may be decoded using an entropy decoder.

[0070] Further, the method 800 may include determining 802 a reconstructed multi-channel signal 311 from the plurality of reconstructed channel signals 314 using the jointly encoded metadata 205. The reconstructed multi-channel signal 311 may include a reconstructed sound field representation (SR) signal. In particular, the reconstructed multi-channel signal 311 corresponds to an approximation or reconstruction of the multi-channel input signal 201. The reconstructed multi-channel signal 311 and the object metadata 202 may together form a reconstructed immersive audio (IA) signal 121.

[0071] Additionally, the method 800 may include rendering the reconstructed multi-channel signal 311 (typically in conjunction with the object metadata 202). The rendering may be performed using headphone rendering, speaker rendering, and / or sound field rendering. As a result, flexible rendering of spatial audio content is enabled (particularly for VR applications).

[0072] As mentioned above, the joint coding metadata 205 may include upmix data, in particular one or more upmix matrices, that enable upmixing of the multiple reconstructed channel signals 404 into a reconstructed multi-channel signal 311. Furthermore, the joint coding metadata 205 may include decorrelation data that enables generation of a reconstructed multi-channel signal 311 having a predetermined covariance. The joint coding metadata 205 may include different metadata for different subbands of the reconstructed multi-channel signal 311. As a result of this, an accurate reconstruction of the multi-channel input signal 201 may be achieved.

[0073] In the corresponding encoder 200, energy compaction may be applied to the multiple downmix channel signals 304. The energy compaction may be performed using prediction and / or using the Karhunen-Loeve transform, the principal component analysis transform, and / or the singular value decomposition transform. The joint coding metadata 205 may be such that, in addition to the upmix, it implicitly performs the inverse of the energy compaction operation. In particular, the joint coding metadata 205 may be such that, in addition to the upmix, it implicitly performs the inverse of the prediction operation and / or the inverse of the Karhunen-Loeve transform, the principal component analysis transform, and / or the singular value decomposition transform.

[0074] In other words, the joint coding metadata 205 may be configured to enable upmixing of the plurality of reconstructed channel signals 404 into the reconstructed multi-channel signal 311 and (implicitly) perform an inverse energy compaction operation on the plurality of reconstructed channel signals 314. In particular, the joint coding metadata 205 may be configured to (implicitly) perform an inverse prediction operation (the inverse of the prediction operation performed by the encoder 200) on at least some of the plurality of reconstructed channel signals 314. Alternatively or additionally, the joint coding metadata 205 may be configured to perform an inverse of a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform (the inverse of the transform performed by the encoder 200) on at least some of the plurality of reconstructed channel signals 314. As a result, a particularly efficient coding scheme may be provided.

[0075] The reconstructed multi-channel signal 311 may include one or more reconstructed object signals (in addition to an SR signal, e.g., an FOA or HOA signal) of one or more audio objects 303. The method 800 may include decoding, particularly using an entropy decoder, object metadata 202 for one or more audio objects 303 from the encoded metadata 207. As a result, the one or more objects 303 may be accurately rendered.

[0076] As described above, the plurality of reconstructed channel signals 314 may form an SR signal, in particular a K-th order Ambisonics signal, where K≧1 (in particular K=1). On the other hand, the reconstructed multi-channel signal 311 may include an SR signal, in particular an L-th order Ambisonics signal, where L≧K (in particular L=K or L=K+1), and one or more (e.g., n=2) reconstructed object signals of one or more audio objects 303. The reconstructed multi-channel signal 311 may be determined by upmixing the plurality of reconstructed channel signals 314 using jointly encoded metadata 205, thereby imparting substantial spatial acoustic events to the reconstructed multi-channel signal 311.

[0077] As mentioned above, the use of upmixing may correspond to a first mode (for high perceptual quality), in which the joint object metadata 205 includes upmix data to enable an upmixing operation, and in a second mode, the reconstructed multi-channel signal 311 may include the same number of channels as the plurality of reconstructed channel signals 314 (and thus no upmixing operation is required).

[0078] In the second mode, the joint coding metadata 205 may include prediction data (e.g., one or more scaling factors) configured to reallocate energy among the different reconstructed channel signals 314. Further, in the second mode, determining 802 the reconstructed multi-channel signal 311 may include reallocating energy among the different reconstructed channel signals 314 using the prediction data. In particular, the inverse operation of the above-mentioned energy compaction operation may be performed using the joint coding metadata 205. As a result, the multiple downmix channel signals 203 may be reconstructed in an efficient and accurate manner.

[0079] As outlined above, the energy compaction operations performed during encoding may include applying a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform to at least some of the multiple downmix channel signals 203. The joint coding metadata 205 may include transform data that enables the decoder 350 to perform the inverse of the Karhunen-Loeve transform, the principal component analysis transform, and / or the singular value decomposition transform. In other words, the transform data indicates the inverse of the Karhunen-Loeve transform, the principal component analysis transform, and / or the singular value decomposition transform to be applied to at least some of the multiple reconstructed channel signals 314 to determine the reconstructed multi-channel signal 311. As a result, the multiple downmix channel signals 203 may be reconstructed in an efficient and accurate manner.

[0080] As described above, the reconstructed multi-channel input signal 311 may include a sequence of frames. The method 800 may include determining, for each frame of the sequence of frames, whether the second mode is used. To this end, an indication of whether the second mode is used may be extracted from the bitstream 101.

[0081] Various exemplary embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. In general, it will be understood that the present disclosure encompasses apparatus suitable for performing the methods described above, for example, an apparatus (spatial renderer) having a memory and a processor coupled to the memory, the processor configured to execute instructions and perform methods according to embodiments of the present disclosure.

[0082] Although various aspects of the exemplary embodiments of the invention are shown and described as block diagrams, flowcharts, or using some other pictorial representations, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controllers, or other computing devices, or some combination thereof.

[0083] Additionally, the various blocks illustrated in the flowcharts may be viewed as method steps and / or as operations resulting from computer program code operations and / or as multiple coupled logic circuit elements configured to perform the associated functions. For example, embodiments of the present invention include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.

[0084] In the context of the present disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0085] Computer program code for carrying out the methods of the present invention may be written in any combination of one or more programming languages. The computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, and when executed by the processor of the computer or other programmable data processing apparatus, the program code causes the computer to perform the functions / acts specified in the flowcharts and / or block diagrams. The program code may run on the computer, partly on the computer, as a stand-alone software package, partly on the computer, partly on a remote computer, or entirely on a remote computer or server.

[0086] Furthermore, while acts are depicted in a particular order, this should not be understood as requiring such acts to be performed in the particular order or sequential order depicted, or that all of the acts depicted be performed to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.

[0087] It should be noted that the specification and drawings merely illustrate the principles of the proposed method and apparatus. Thus, it will be understood that those skilled in the art will be able to devise various arrangements, not explicitly described or shown herein, which embody the principles of the invention and are within its spirit and scope. Furthermore, all examples described herein are expressly intended primarily for educational purposes only to aid the reader in understanding the principles of the proposed method and apparatus and concepts contributed by the inventors to further the art, and are to be construed as being limited to the examples and conditions specifically described as such. Furthermore, all statements herein describing principles, aspects, and embodiments of the invention, as well as specific examples thereof, are intended to encompass equivalents thereof.

[0088] Several aspects will be described. [Aspect 1] A method (700) for encoding a multi-channel input signal (201), the method (700) comprising: determining (701) a plurality of downmix channel signals (203) from the multi-channel input signal (201); performing energy compaction (702) of the plurality of downmix channel signals (203) to provide a plurality of compacted channel signals (404); determining (703) joint encoding metadata (205) based on the plurality of compacted channel signals (404) and on the multi-channel input signal (201), the joint encoding metadata (205) being such that it allows upmixing the plurality of compacted channel signals (404) to an approximation of the multi-channel input signal (201); encoding (704) the plurality of compacted channel signals (404) and the jointly encoded metadata (205), method. [Aspect 2] 2. The method of embodiment 1, wherein the energy compaction is performed such that the energy of the compacted channel signal (404) is lower than the energy of the corresponding downmix channel signal (203). Aspect 3 Energy compactification can be performed: predicting the first downmix channel signal (203) from the second downmix channel signal (203) to provide a first predicted channel signal; subtracting the first predicted channel signal from the first downmix channel signal to provide a first compacted channel signal; 3. The method of embodiment 1 or 2. Aspect 4 predicting the first downmix channel signal (203) from the second downmix channel signal (203) comprises determining a scaling factor for scaling the second downmix channel signal (203); the first predicted channel signal corresponds to the second downmix channel signal (203) scaled according to the scaling factor, The method of embodiment 3. Aspect 5 The scaling factor is the energy of the first compacted channel signal (404) is reduced compared to the energy of the first downmix channel signal (203); and / or The energy of the first compacted channel signal (404) is minimized; The method of embodiment 4, wherein the Aspect 6 To perform energy compactification, determining a number of compacted channel signals (404) based on prediction from said second downmix channel signal (203); applying a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform to the number of compacted channel signals (404); 6. The method of any one of embodiments 3 to 5. Aspect 7 the downmix channel signals (203) are first-order Ambisonics signals, in particular in B-format or A-format; and / or The plurality of compacted channel signals (404) are represented in the format of a first-order Ambisonics signal, in particular in the B-format or the A-format. 7. The method of any one of embodiments 1 to 6. Aspect 8 To perform energy compactification, predicting an X channel signal, a Y channel signal, and a Z channel signal from the W channel signal of the plurality of downmix channel signals (203) to provide a predicted X channel signal, a predicted Y channel signal, and a predicted Z channel signal; subtracting the predicted X channel signal from the X channel signal to determine an X′ channel signal; subtracting the predicted Y channel signal from the Y channel signal to determine a Y′ channel signal; subtracting the predicted Z channel signal from the Z channel signal to determine a Z′ channel signal; determining the plurality of compacted channel signals (404) based on the W channel signal, the X' channel signal, the Y' channel signal, and the Z' channel signal; The method according to embodiment 7. Aspect 9 To perform energy compactification, applying a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform to the X' channel signal, the Y' channel signal, and the Z' channel signal to provide an X" channel signal, a Y" channel signal, and a Z" channel signal; determining the plurality of compacted channel signals (404) based on the W channel signal, the X″ channel signal, the Y″ channel signal, and the Z″ channel signal; The method according to embodiment 8. Aspect 10 10. The method of any one of aspects 1 to 9, wherein performing energy compaction includes applying a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform to at least some of the downmix channel signals (203). Aspect 11 The jointly encoded metadata (205) upmix data, in particular an upmix matrix, that allows the upmixing of the plurality of compacted channel signals (404) to an approximation of the multi-channel input signal (201) that contains the same number of channels as the multi-channel input signal (201); and / or decorrelated data allowing the reconstruction of the covariance of said multi-channel input signal (201) 11. The method of any one of embodiments 1 to 10, comprising: Aspect 12 12. The method of any one of aspects 1-11, wherein the joint encoding metadata (205) is determined for a plurality of different subbands of the multi-channel input signal (201). Aspect 13 13. The method of any one of aspects 1 to 12, wherein encoding (704) the plurality of compacted channel signals (404) includes performing waveform coding of each of the plurality of compacted channel signals (404) using, in particular, a mono encoder for each compacted channel signal (404). Aspect 14 14. The method of any one of aspects 1-13, wherein the jointly encoded metadata (205) is encoded using an entropy encoder. Aspect 15 the multi-channel input signal (201) comprises one or more object signals of one or more audio objects (303); the method (700) comprising encoding object metadata (202) for the one or more audio objects (303), in particular using an entropy encoder; 15. The method of any one of embodiments 1 to 14. Aspect 16 the multi-channel input signal (201) comprises a sound field representation signal, called SR, in particular an L-th order Ambisonics signal, where L≧1, and one or more object signals of one or more audio objects (303); the plurality of downmix channel signals (203) are determined by downmixing the multi-channel input signal (201) to an SR signal, in particular to a K-th order Ambisonics signal, where L≧K, 16. The method of any one of embodiments 1 to 15. Aspect 17 determining (701) the plurality of downmix channel signals (203) comprises mixing the one or more object signals of one or more audio objects (303) into the SR signal of the multi-channel input signal (201) depending on object metadata (202) of the one or more audio objects (303); The object metadata (202) of an audio object (303) indicates the spatial location of the audio object (303); 17. The method of embodiment 16. Aspect 18 the method (700) comprising determining that the multi-channel input signal (201) should be encoded using a second mode; in a second mode, the joint encoding metadata (205) is determined based on the plurality of compacted channel signals (404) and on the plurality of downmix channel signals (203), the joint encoding metadata (205) being such that it allows reconstructing the plurality of downmix channel signals (203) from the plurality of compacted channel signals (404); 18. The method of any one of embodiments 1 to 17. Aspect 19 Determining the jointly encoded metadata (205) based on the plurality of compacted channel signals (404) and based on the multi-channel input signal (201) corresponds to a first mode; the multi-channel input signal (201) comprises a sequence of frames; The method (700) includes determining, for each frame of the sequence of frames, whether to use a first mode or a second mode; 19. The method of embodiment 18. Aspect 20 generating a bitstream (101) based on encoded audio data (206) derived by encoding (704) the plurality of compacted channel signals (404) and based on encoded metadata (207) derived by encoding (704) the joint encoded metadata (205); inserting into the bitstream (101) an indication of whether the second mode was used, 20. The method of any one of embodiments 17 to 19. Aspect 21 1. A method (800) for determining a reconstructed multi-channel signal (311) from encoded audio data (206) indicative of a plurality of reconstructed channel signals (314) and encoded metadata (207) indicative of joint encoding metadata (205), the method (800) comprising: decoding (801) the encoded audio data (206) to provide the plurality of reconstructed channel signals (314) and decoding the encoded metadata (207) to provide the joint encoded metadata (205); determining (802) the reconstructed multi-channel signal (311) from the plurality of reconstructed channel signals (314) using the jointly encoded metadata (205); method. Aspect 22 22. The method of embodiment 21, wherein the plurality of reconstructed channel signals (314) are first-order Ambisonics signals, particularly in B-format or A-format. Aspect 23 The jointly encoded metadata (205) upmix data, in particular an upmix matrix, enabling the upmixing of the plurality of reconstructed channel signals (404) into the reconstructed multi-channel signal (311); and / or decorrelated data that allows generating a reconstructed multi-channel signal (311) with a predetermined covariance. 23. The method of embodiment 21 or 22, comprising: Aspect 24 24. The method of any one of aspects 21 to 23, wherein the jointly encoded metadata (205) comprises different metadata for different subbands of the reconstructed multi-channel signal (311). Aspect 25 25. The method of any one of aspects 21 to 24, wherein decoding (801) the encoded audio data (206) includes performing waveform decoding of each of the plurality of reconstructed channel signals (314), in particular using a mono decoder for each reconstructed channel signal (314). Aspect 26 26. The method of any one of aspects 21 to 25, wherein the encoded metadata (207) is decoded using an entropy decoder. Aspect 27 the reconstructed multi-channel signal (311) comprises one or more reconstructed object signals of one or more audio objects (303); The method (800) comprises decoding object metadata (202) for the one or more audio objects (303) from the encoded metadata (207), in particular using an entropy decoder, 27. The method of any one of embodiments 21 to 26. Aspect 28 The plurality of reconstructed channel signals (314) form a sound field representation signal called SR, in particular a K-th order Ambisonics signal, where K≧1; the reconstructed multi-channel signal (311) is determined by upmixing the plurality of reconstructed channel signals (314) using the jointly encoded metadata (205); the reconstructed multi-channel signal (311) comprises the reconstructed SR signal, in particular an L-th order Ambisonics signal, where L≧K, and one or more reconstructed object signals of one or more audio objects (303), 28. A method according to any one of embodiments 21 to 27. Aspect 29 the jointly encoded metadata (205) is configured to perform an inverse energy compaction operation on the plurality of reconstructed channel signals (314); and / or the joint coding metadata (205) is configured to perform an inverse prediction operation on at least some of the plurality of reconstructed channel signals (314); and / or the jointly encoded metadata (205) is configured to perform an inverse of a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform on at least a portion of the plurality of reconstructed channel signals (314); 29. The method of any one of embodiments 21 to 28. Aspect 30 the method (800) including determining that the reconstructed multi-channel signal (311) should be determined using a second mode; In a second mode, the joint coding metadata (205) comprises prediction data and / or transformation data configured to reallocate energy between different reconstructed channel signals (314): In a second mode, determining (802) the reconstructed multi-channel signal (311) includes reallocating energy between different reconstructed channel signals (314) using the prediction data and / or the transformation data; In a second mode, the reconstructed multi-channel signal (311) contains the same number of channels as the plurality of reconstructed channel signals (314); 30. The method of any one of embodiments 21 to 29. Aspect 31 31. The method of claim 30, wherein the transformation data indicates an inverse of a Karhunen-Loeve transform, a principal component analysis transform, and / or a singular value decomposition transform to be applied to at least a portion of the plurality of reconstructed channel signals (314) to determine the reconstructed multi-channel signal (311). Aspect 32 the reconstructed multi-channel input signal (311) comprises a sequence of frames; The method (800) includes determining for each frame of the sequence of frames whether a second mode should be used; 32. The method of embodiment 30 or 31. Aspect 33 extracting the encoded audio data (206) and the encoded metadata (207) from the bitstream (101); extracting from said bitstream (101) an indication indicating whether the second mode should be used, 33. The method of any one of embodiments 30 to 32. Aspect 34 34. The method of any one of aspects 30 to 33, wherein the method (800) comprises rendering the reconstructed multi-channel signal (311). Aspect 35 An encoding unit (200) for encoding a multi-channel input signal (201), the encoding unit (200) comprising: determining a plurality of downmix channel signals (203) from the multi-channel input signal (201); performing energy compaction of the plurality of downmix channel signals (203) to provide a plurality of compacted channel signals (404); determining joint encoding metadata (205) based on the plurality of compacted channel signals (404) and based on the multi-channel input signal (201), the joint encoding metadata (205) being such that it allows upmixing the plurality of compacted channel signals (404) to an approximation of the multi-channel input signal (201); encoding the plurality of compacted channel signals (404) and the jointly encoded metadata (205), Encoding unit. Aspect 36 a decoding unit (350) for determining a reconstructed multi-channel signal (311) from encoded audio data (206) indicative of a plurality of reconstructed channel signals (314) and encoded metadata (207) indicative of the joint encoding metadata (205), said decoding unit (350) comprising: decoding the encoded audio data (206) to provide the plurality of reconstructed channel signals (314); decoding the encoded metadata (207) to provide the joint encoded metadata (205); configured to determine the reconstructed multi-channel signal (311) from the plurality of reconstructed channel signals (314) using the jointly encoded metadata (205); Decoding unit.

Claims

[Claim 1] 1. A method for determining a reconstructed multi-channel signal from encoded audio data indicative of a plurality of reconstructed channel signals and encoded metadata indicative of jointly encoded metadata, the method comprising: decoding the encoded audio data to provide the plurality of reconstructed channel signals forming a K-th order Ambisonics signal, where K≧1, and decoding the encoded metadata to provide the joint encoded metadata; upmixing the plurality of reconstructed channel signals into an increased number of first upmix channel signals based on the jointly encoded metadata, where the first upmix channel signals include an L-th order Ambisonics signal, where L>K; applying decorrelation to selected channels of the K-th order Ambisonics signal to generate first and second decorrelated signals; upmixing the first and second decorrelated signals into an increased number of second upmix channel signals based on the jointly encoded metadata; determining the reconstructed multi-channel signal by adding corresponding channels of the L-th order Ambisonics signal and the second upmix channel signal. method.