Coding and data structure of layered construction for compressed higher-order ambisonics sound or sound field representation

The method enhances HOA layer-based coding by distributing transport signals across hierarchical layers and providing layer-specific payloads, ensuring high-quality decoding and sound representation even at low bitrates.

JP2025186229APending Publication Date: 2025-12-23DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025134595
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-07-13
Filing Date
2025-08-13
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Current HOA layer-based coding methods fail to adequately enhance lower layers, especially at low bitrates, leading to suboptimal sound quality and incomplete decoding of lower layers due to missing V-vector elements in higher layers.

Method used

A method for layered encoding of compressed HOA representations that allocates transport signals to multiple hierarchical layers, generates HOA extension payloads for each layer, and signals these payloads in the bitstream to enhance the reconstructed sound representation, ensuring proper decoding even at low bitrates.

Benefits of technology

Enables high-quality decoding of HOA representations across all layers, including lower layers, by providing necessary enhancement side information for each layer, thus maintaining sound quality and reducing bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186229000001_ABST
    Figure 2025186229000001_ABST
Patent Text Reader

Abstract

To provide a method for decoding a compressed higher-order ambisonics (HOA) representation of sound or sound fields.SOLUTION: A method includes the steps of: receiving a bitstream containing a compressed HOA representation; determining the highest available layer among multiple layers for decoding; extracting HOA extended payload assigned to the highest available layer; decoding the compressed HOA representation corresponding to the highest available layer, based on layer information and a transport signal assigned to the highest available layer and arbitrary layers lower than the highest available layer; and parametrically enhancing the decoded HOA representation using the side information included in the HOA extension payload assigned to the highest available layer.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority from European Patent Application No. 15306653.5, filed October 15, 2015, the contents of which are incorporated herein by reference in their entirety.

[0002] Technical Field This document relates to methods and apparatus for layered audio coding. In particular, this document relates to methods and apparatus for layered audio coding of frames of compressed Higher-Order Ambisonics (HOA) sound (or sound field) representations. This document further relates to data structures (e.g., bitstreams) for representing frames of compressed HOA sound (or sound field) representations. [Background technology]

[0003] In the current definition of HOA layer-based coding, side information is generated for the HOA decoding tools Spatial Signal Prediction, Sub-band Directional Signal Synthesis, and Parametric Ambience Replication (PAR) decoders to enhance a particular HOA representation. That is, in the current definition of layer-based HOA coding, the data provided only adequately enhances the HOA representation in the top layer (e.g., the highest enhancement layer). For lower layers, including the base layer, these tools do not adequately enhance the partially reconstructed HOA representation.

[0004] The tools for subband directional signal synthesis and parametric ambience replication decoders are especially designed for low data rates where only a few transport signals are available. However, in HOA layered coding, proper enhancement of the (partially) reconstructed HOA representation is not possible, especially for low bitrate layers such as the base layer. This is obviously undesirable from the viewpoint of sound quality at low bitrates.

[0005] Additionally, it has been found that the usual way of handling encoded V vector elements for vector-based signals does not provide proper decoding when CodedVVecLength equal to 1 is signaled in HOADecoderConfig() (i.e., when vector coding mode is active). In this vector coding mode, no V vector elements are transmitted for HOA coefficient indices included in the set ContAddHoaCoeff. This set includes all HOA coefficient indices AmbCoeffIdx[i] with AmbCoeffTransitionState equal to 0. Normally, there is no need to also add a weighted V vector signal, since the original HOA coefficient sequence for these indices is explicitly signaled.

[0006] However, in layered coding mode, the set of consecutive HOA coefficient indices depends on the transport channels that are part of the currently active layer. Additional HOA coefficient indices sent in higher layers may be missing in lower layers. Therefore, the assumption that vector signals should not contribute to the HOA coefficient sequence is incorrect for HOA coefficient indices that belong to the HOA coefficient sequence included in higher layers.

[0007] As a result, the V-vector in layered HOA coding may not be suitable for decoding any layer below the top layer.

[0008] Thus, there is a need for a coding scheme and bitstream adapted to the coding of layered compressed HOA representations of sounds or sound fields. Summary of the Invention [Problem to be solved by the invention]

[0009] This paper addresses the above problems. In particular, a method and encoder / decoder for layered coding of frames of compressed HOA sound or sound field representations and a data structure for representing frames of compressed HOA sound or sound field representations are described. [Means for solving the problem]

[0010] According to one aspect, a method for layered encoding of frames of a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field is described. The compressed HOA representation conforms to the Draft MPEG-H 3D Audio standard and any other future adopted or draft standard. The compressed HOA representation may include a plurality of transport signals. The transport signal may relate to a mono signal, for example, representing either a dominant sound signal or a coefficient sequence of the HOA representation. The method may include allocating the plurality of transport signals to a plurality of hierarchical layers. For example, the transport signal may be distributed among a plurality of layers. The plurality of layers may include a base layer and one or more hierarchical enhancement layers. The plurality of hierarchical layers may be ordered from a base layer through a first enhancement layer, a second enhancement layer, etc., to an overall highest enhancement layer (overall top layer). The method may further include generating, for each layer, a respective HOA extension payload including side information (e.g., enhancement side information). The side information is for parametrically enhancing a reconstructed HOA representation obtained from a transport signal assigned to the respective layer and any layers lower than the respective layer. The method may further include assigning the generated HOA extension payload to each layer. The method may further include signaling the generated HOA extension payload in an output bitstream. The HOA extension payload may be signaled in an HOAEnhFrame() payload. Thus, the side information may be moved from HOAFrame() to HOAEnhFrame().

[0011] Implemented as described above, the proposed method applies layer-based coding to the compressed HOA representation (frames) to enable high-quality decoding thereof even at low bit rates. In particular, the proposed method ensures that each layer includes a suitable HOA extension payload (e.g., enhancement side information) for enhancing the (partially) reconstructed sound representation obtained from the transport signal at any layer up to the current layer. Here, layers up to the current layer are understood to include, for example, the base layer, the first enhancement layer, the second enhancement layer, etc. up to the current layer. Here, layers up to the current layer are understood to include, for example, the base layer, the first enhancement layer, the second enhancement layer, etc. up to the current layer. For example, a decoder is enabled to enhance the (partially) reconstructed sound representation obtained from the base layer by referring to the HOA extension payload assigned to the base layer. In a conventional approach, only the reconstructed HOA representation of the highest enhancement layer can be enhanced by the HOA extension payload. In this way, the decoder is enabled to improve or enhance the (partially) reconstructed sound representation, even if it differs from the complete (e.g., full) sound representation, regardless of the actual highest available layer (e.g., a layer below the lowest layer that has not yet been effectively received; all layers below the highest available layer and the highest available layer itself have been effectively received). In particular, it is sufficient for the decoder to decode the HOA extension payload for only a single layer (i.e., for the highest available layer) to improve or enhance the (partially) reconstructed sound representation that can be obtained based on all transport signals contained in layers up to the actual highest available layer, regardless of the actual highest available layer. Decoding of the HOA extension payloads of higher or lower layers is not required. On the other hand, the proposed method allows to fully take advantage of the reduction in required bandwidth that can be achieved when applying layer-based coding.

[0012] In some embodiments, the method may further include transmitting data payloads for the plurality of layers with respective levels of error protection. The data payloads may include respective HOA extension payloads. The base layer may have the highest error protection, and the one or more enhancement layers may have successively decreasing error protection. This may ensure that at least some lower layers are transmitted reliably while reducing the overall required bandwidth by not applying excessive error protection to the higher layers.

[0013] In embodiments, the HOA extension payload may include bitstream elements for an HOA spatial signal prediction decoding tool. Additionally or alternatively, the HOA extension payload may include bitstream elements for an HOA subband directional signal synthesis decoding tool. Additionally or alternatively, the HOA extension payload may include bitstream elements for an HOA parametric ambient sound replication decoding tool.

[0014] In embodiments, the HOA extension payload may have a usacExtElementType of ID_EXT_ELE_HOA_ENH_LAYER.

[0015] In some embodiments, the method may further include generating an HOA configuration extension payload including bitstream elements for configuring an HOA spatial signal prediction decoding tool, an HOA subband directional signal synthesis decoding tool, and / or an HOA parametric ambient sound replication decoding tool. The HOA configuration extension payload may be included in HOADecoderEnhConfig(). The method may further include signaling the HOA configuration extension payload in the output bitstream.

[0016] In embodiments, the method may further include generating an HOA decoder configuration payload that includes information indicating an allocation of HOA extension payloads to the plurality of layers, and signaling the HOA decoder configuration payload in an output bitstream.

[0017] In some embodiments, the method may further include determining whether a vector coding mode is active. If the vector coding mode is active, the method may further include determining, for each layer, a set of consecutive HOA coefficient indices based on a transport signal assigned to the respective layer. The HOA coefficient indices in the set of consecutive HOA coefficient indices may be HOA coefficient indices included in the set ContAddHOACoeff. The method may further include generating, for each transport signal, a V-vector based on the determined set of consecutive HOA coefficient indices for the layer to which the respective transport signal is assigned. Here, the generated V-vector includes elements for any transport signals assigned to layers higher than the layer to which the respective transport signal is assigned. The method may further include signaling the generated V-vector in an output bitstream.

[0018] According to another aspect, a method for layered encoding of a frame of a compressed Higher-Order Ambisonics (HOA) representation of a sound or sound field is described. The compressed HOA representation may include a plurality of transport signals. The transport signal may relate to a mono signal, e.g., representing either a dominant sound signal or a coefficient sequence of the HOA representation. The method may include assigning the plurality of transport signals to a plurality of hierarchical layers. For example, the transport signals may be distributed among a plurality of layers. The plurality of layers may include a base layer and one or more hierarchical enhancement layers. The method may further include determining whether a vector coding mode is active. If the vector coding mode is active, the method may further include determining, for each layer, a set of consecutive HOA coefficient indices based on the transport signal assigned to the respective layer. The HOA coefficient indices in the set of consecutive HOA coefficient indices may be HOA coefficient indices included in a set ContAddHOACoeff. The method may further include generating, for each transport signal, a V vector based on the determined set of consecutive HOA coefficient indices for the layer to which the respective transport signal is assigned, where the generated V vector includes elements for any transport signals assigned to layers higher than the layer to which the respective transport signal is assigned. The method may further include signaling the generated V vector in an output bitstream.

[0019] Configured in this way, the proposed method ensures that in vector coding mode, suitable V vectors are available for all transport signals belonging to layers up to the highest available layer. Specifically, the proposed method excludes cases where elements of the V vector corresponding to transport signals in higher layers are not explicitly signaled. Thus, the information contained in layers up to the highest available layer is sufficient to decode any transport signals belonging to layers up to the highest available layer. This ensures proper decompression of the respective reconstructed HOA representations for lower layers (low bit-rate layers) even if the higher layers are not effectively received by the decoder. On the other hand, the proposed method allows to fully take advantage of the reduction in required bandwidth that can be achieved when applying layer-based coding.

[0020] According to another aspect, a method for decoding a frame of a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field is described. The compressed HOA representation may be encoded with multiple hierarchical layers. The multiple hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The method may include receiving a bitstream related to a frame of the compressed HOA representation. The method may further include extracting payloads for the multiple layers. Each payload may include a transport signal assigned to the respective layer. The method may further include determining a highest usable layer for decoding from the multiple layers. The method may further include extracting an HOA extension payload assigned to the highest usable layer. The HOA extension payload may include side information for parametrically enhancing a (partially) reconstructed HOA representation corresponding to the highest usable layer. The (partially) reconstructed HOA representation corresponding to the highest usable layer may be obtained based on transport signals assigned to the highest usable layer and any layers lower than the highest usable layer. The method may further include generating a (partially) reconstructed HOA representation corresponding to the highest available layer based on transport signals assigned to the highest available layer and any layers lower than the highest available layer. The method may further include enhancing (e.g., parametrically enhancing) the (partially) reconstructed HOA representation using side information included in the HOA extension payload assigned to the highest available layer. As a result, an enhanced reconstructed HOA representation may be obtained.

[0021] Configured in this way, the proposed method ensures that the final (e.g., enhanced) reconstructed HOA representation is of optimal quality, making maximum use of the available (e.g., effectively received) information.

[0022] In embodiments, the HOA extension payload may include bitstream elements for an HOA spatial signal prediction decoding tool. Additionally or alternatively, the HOA extension payload may include bitstream elements for an HOA subband directional signal synthesis decoding tool. Additionally or alternatively, the HOA extension payload may include bitstream elements for an HOA parametric ambient sound replication decoding tool.

[0023] In embodiments, the HOA extension payload may have a usacExtElementType of ID_EXT_ELE_HOA_ENH_LAYER.

[0024] In some embodiments, the method may further include parsing the bitstream to extract an HOA configuration extension payload, which may include bitstream elements for configuring an HOA spatial signal prediction decoding tool, an HOA subband directional signal synthesis decoding tool, and / or an HOA parametric ambient sound replication decoding tool.

[0025] In some embodiments, the method may further include extracting HOA extension payloads assigned to the plurality of layers, respectively. Each HOA extension payload may include side information for parametrically enhancing the (partially) reconstructed HOA representation corresponding to the assigned layer. The (partially) reconstructed HOA representation corresponding to each assigned layer may be obtained based on transport signals assigned to that layer and any layers below it. The assignment of HOA extension payloads to each layer may be known from configuration information included in the bitstream.

[0026] In some embodiments, determining the highest available layer may involve determining a set of invalid layer indices that indicate layers that have not yet been validly received. It may further involve determining the highest available layer as the layer that is one layer below the layer indicated by the smallest (lowest) index in the set of invalid layer indices. The base layer may have the lowest layer index (e.g., layer index 1), and hierarchical enhancement layers may have successively higher layer indices. The proposed method thereby ensures that the highest available layer is chosen such that all information required to decode a (partially) reconstructed HOA representation from the highest available layer and any layers below the highest available layer is present.

[0027] In some embodiments, determining the highest usable layer may involve determining a set of invalid layer indices indicating layers that have not yet been validly received. It may further involve determining the highest usable layer of a previous frame preceding the current frame. It may further involve determining the highest usable layer as the lower of the highest usable layer of the previous frame and a layer that is further below the layer indicated by the smallest index in the set of invalid layer indices. Thus, even if the current frame was differentially encoded with respect to the previous frame, the highest usable layer for the current frame is chosen such that all information required to decode a (partially) reconstructed HOA representation from the highest usable layer and any layers below the highest usable layer is available.

[0028] In embodiments, the method may further comprise deciding not to perform parametric enhancement of the (partially) reconstructed representation using side information contained in the HOA extension payload assigned to the highest available layer if the highest available layer of the current frame is lower than the highest available layer of said previous frame and if the current frame is differentially coded with respect to said previous frame, so that the reconstructed HOA representation can be decoded without error if the current frame (including the side information contained in the HOA extension payload assigned to the highest available layer) was differentially coded with respect to said previous frame.

[0029] In some embodiments, the set of invalid layer indices may be determined by evaluating the validity flag of the corresponding HOA extension payload. A layer index for a given layer may be added to the set of invalid layer indices if the validity flag for the HOA extension payload assigned to the respective layer is not set. This allows the set of invalid layer indices to be determined in an efficient manner.

[0030] According to another aspect, a data structure (e.g., a bitstream) is described that represents a frame of a compressed Higher-Order Ambisonics (HOA) representation of a sound or sound field. The compressed HOA representation may include a plurality of transport signals. The data structure may include a plurality of HOA frame payloads corresponding to respective layers of a plurality of hierarchical layers. The HOA frame payloads may include respective transport signals. The plurality of transport signals may be assigned (e.g., distributed) to the plurality of layers. The plurality of layers may include a base layer and one or more hierarchical enhancement layers. The data structure may further include, for each layer, a respective HOA extension payload containing side information for parametrically enhancing a (partially) reconstructed HOA representation obtained from transport signals assigned to the respective layer and any layers below the respective layer.

[0031] In some embodiments, the HOA frame payload and HOA extension payload for the multiple layers may be provided with respective levels of error protection, with the base layer having the highest error protection and the one or more enhancement layers having successively decreasing error protection.

[0032] In embodiments, the HOA extension payload may include bitstream elements for an HOA spatial signal prediction decoding tool. Additionally or alternatively, the HOA extension payload may include bitstream elements for an HOA subband directional signal synthesis decoding tool. Additionally or alternatively, the HOA extension payload may include bitstream elements for an HOA parametric ambient sound replication decoding tool.

[0033] In embodiments, the HOA extension payload may have a usacExtElementType of ID_EXT_ELE_HOA_ENH_LAYER.

[0034] In some embodiments, the data structure may further include an HOA configuration extension payload including bitstream elements for configuring an HOA spatial signal prediction decoding tool, an HOA subband directional signal synthesis decoding tool, and / or an HOA parametric ambient sound replication decoding tool.

[0035] In embodiments, the data structure may further include an HOA decoder configuration settings payload that includes information indicating the allocation of HOA extension payloads to the multiple layers.

[0036] In embodiments, methods and apparatus relate to decoding compressed Higher Order Ambisonics (HOA) representations of sounds or sound fields. The apparatus may be configured to perform the following steps, or the method may include the following steps: receiving a bitstream including the compressed HOA representation corresponding to a plurality of hierarchical layers including a base layer and one or more hierarchical enhancement layers, the plurality of layers being assigned components of the basic compressed sound representation of the sound or sound field, the components being assigned to each layer in a respective component group; determining a highest usable layer among the plurality of layers for decoding; extracting an HOA extension payload assigned to the highest usable layer, the HOA extension payload including side information for parametrically enhancing a reconstructed HOA representation corresponding to the highest usable layer, the reconstructed HOA representation corresponding to the highest usable layer being obtainable based on a transport signal assigned to the highest usable layer and any layers lower than the highest usable layer; decoding the compressed HOA representation corresponding to the highest usable layer based on layer information, the highest usable layer, and any layers lower than the highest usable layer; and parametrically enhancing the decoded HOA representation using the side information included in the HOA extension payload assigned to the highest usable layer.

[0037] The HOA extension payload may include a bitstream element for HOA spatial signal prediction decoding tools. The layer information may indicate a number of active directional signals in the current frame of an enhancement layer.

[0038] The layer information may indicate the total number of additional ambient sound HOA coefficients for the enhancement layer. The layer information may include an HOA coefficient index for each additional ambient sound HOA coefficient for the enhancement layer. The layer information may include enhancement information including at least one of spatial signal prediction, subband directional signal synthesis, and parametric ambient sound reproduction decoder. The compressed HOA representation is adapted for a layer-configured coding mode for HOA-based content when CodedVVecLength equal to 1 is signaled in HOADecoderConfig(). Furthermore, v vector elements may not be transmitted for indices equal to the indices of additional HOA coefficients included in the set ContAddHoaCoeff. The set ContAddHoaCoeff may be defined separately for each layer of the multiple hierarchical layers. The layer information includes NumLayers elements, each indicating the number of transport signals included in all layers up to the i-th layer. The layer information may include indicators of all actually used layers for the k-th frame. The layer information may indicate that all of the coefficients for the dominant vector are specified. The layer information may indicate that coefficients of the dominant vector corresponding to more than MinNumOfCoeffsForAmbHOA are specified. The layer information may indicate that not all elements defined in MinNumOfCoeffsForAmbHOA and ContAddHoaCoeff[lay] are transmitted, where lay is the index of the layer containing the vector-based signal corresponding to the vector.

[0039] According to another aspect, an encoder for encoding a layered arrangement of frames of a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field representation is described. The compressed HOA representation may include a plurality of transport signals. The encoder may include a processor configured to perform some or all of the method steps of the methods according to the first and second aspects above.

[0040] According to another aspect, a decoder for decoding frames of a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field representation is described. The compressed HOA representation may be encoded in multiple hierarchical layers, including a base layer and one or more hierarchical enhancement layers. The decoder may include a processor configured to perform some or all of the method steps of the method according to the third aspect above.

[0041] According to another aspect, a software program is described that is adapted for execution on a processor and may be adapted to perform some or all of the method steps outlined herein when executed on a computing device.

[0042] According to yet another aspect, a storage medium is described that may include a software program adapted for execution on a processor and adapted to perform some or all of the method steps outlined herein when executed on a computing device.

[0043] As one skilled in the art will appreciate, statements made with respect to any of the above aspects or embodiments thereof can be understood to apply to other aspects or embodiments as well, and repeating these statements for every single aspect or embodiment has been omitted for the sake of brevity.

[0044] It should be noted that the methods and apparatus, including the preferred embodiments outlined herein, may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and apparatus outlined herein may be combined in any manner. In particular, features of the claims may be combined in any manner with other features.

[0045] It should further be noted that method steps and apparatus features may be interchanged in many ways. In particular, those skilled in the art will appreciate that details of a disclosed method can be implemented as an apparatus adapted to perform some or all of the method steps, and vice versa. [Brief explanation of the drawings]

[0046] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1] FIG. 1 is a block diagram illustrating a schematic allocation of payload to a base layer and M−1 enhancement layers at the encoder side. [Figure 2] FIG. 1 is a block diagram illustrating a schematic example of a receiver and decompression stage. [Figure 3] 1 is a flowchart illustrating an example method for layered encoding of a frame of a compressed HOA representation, according to an embodiment of the present disclosure. [Figure 4] 10 is a flowchart illustrating another example method for layered encoding of a frame of a compressed HOA representation, according to an embodiment of the present disclosure. [Figure 5] 10 is a flowchart illustrating an example method for decoding a frame of a compressed HOA representation according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a block diagram that schematically illustrates an example of a hardware implementation of an encoder according to an embodiment of the present disclosure. [Figure 7] FIG. 2 is a block diagram that schematically illustrates an example of a hardware implementation of a decoder according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0047] We first describe a compressed sound (or sound field) representation to which the methods and encoders / decoders according to the present disclosure may be applied.

[0048] For streaming of compressed sound (or sound field) representations through transmission channels with time-varying conditions, layered coding is a means of adapting the quality of the received sound representation to the transmission conditions and, in particular, avoiding unwanted signal dropouts.

[0049] For layered coding, the compressed sound (or sound field) representation is usually subdivided into a high-priority base layer of relatively small size and additional enhancement layers of decreasing priority and arbitrary size. Each enhancement layer is typically assumed to contain incremental information to complement the information of all lower layers in order to improve the quality of the compressed sound (or sound field) representation. The idea is then to control the amount of error protection for the transmission of each individual layer based on their priority. In particular, the base layer is provided with a high error protection, which is reasonable and acceptable due to its small size.

[0050] In what follows, it is assumed that a complete compressed sound (or sound field) representation generally consists of three components:

[0051] 1. A basic compressed sound (or sound field) representation that is itself made up of several complementary components. This makes up by far the largest proportion of the complete compressed sound (or sound field) representation.

[0052] 2. Basic side information required to decode the basic compressed sound representation, which is assumed to be much smaller in size than the basic compressed sound (or sound field) representation. It is further assumed to consist mostly of the following components, both of which specify the decompression of only one specific component of the basic compressed sound representation: a) The first component is a set of individual complementary components of the basic compressed sound (or sound field) representation, Independently of other complementary components Contains side information describing b) A second (optional) component represents the individual complementary components of the basic compressed sound (or sound field) representation, Dependent on other complementary components It contains side information that describes the dependency. In particular, a dependency has the following attributes: The dependent side information for each individual complementary component of the basic compressed sound (or sound field) representation achieves maximum range when the basic compressed sound (or sound field) representation does not include certain other complementary components. If additional complementary components of some kind are added to the basic compressed sound (or sound field) representation, the dependent side information for each complementary component under consideration becomes a subset of the original, thereby reducing its size.

[0053] 3. Optional enhancement side information to improve the basic compressed sound (or sound field) representation, the size of which is also assumed to be much smaller than the size of the basic compressed sound (or sound field) representation.

[0054] One prominent example of such type of fully compressed sound (or sound field) representation is given by the compressed HOA sound field representation defined by the preliminary version of the MPEG-H 3D Audio standard.

[0055] 1. The basic compressed sound field representation can be identified using several quantized mono signals representing either the so-called dominant sound signal or the coefficient sequences of the so-called ambient sound HOA sound field components.

[0056] 2. The basic side information describes, among other things, for each of these mono signals how it contributes spatially to the sound field. This information can be further separated into two different components: (a) The presence of other mono signals Independent , side information related to a particular individual mono signal. Such a side signal may, for example, specify a mono signal representing a directional signal (i.e., a general plane wave) with a certain direction of incidence. Alternatively, the mono signal may be specified as a coefficient sequence of the original HOA representation with a certain index. (b) In the presence of other mono signals Depends, side information related to a particular individual mono signal. Such a side signal occurs, for example, when the mono signal is specified to be a so-called vector-based signal, i.e., when it is directionally distributed in the sound field and the directional distribution is specified by a vector. In certain modes (i.e., CodedVVecLength=1), certain components of this vector are implicitly set to 0 and are not part of the compressed vector representation. These components are those with indices equal to the coefficient sequences of the original HOA representation that are part of the underlying compressed sound field representation. This means that the total number of individual components of a vector to be coded depends on the underlying compressed sound field representation, in particular which coefficient sequences of the original HOA representation they contain.

[0057] If the coefficient sequence of the original HOA representation is not included in the basic compressed sound field representation, the dependent basic side information for each vector-based signal consists of all vector components and has its maximum size. If the coefficient sequence of the original HOA representation with certain indices is added to the basic compressed sound field representation, the vector components with those indices are removed from the side information for each vector-based signal, thereby reducing the size of the dependent basic side information for the vector-based signal.

[0058] 3. Enhancement side information consists of the following components: Parameters related to so-called (broadband) spatial prediction, which is used to (linearly) predict missing parts of the sound field from directional signals. Parameters related to the so-called subband directional signal synthesis and parametric ambience replication. Subband directional signal synthesis and parametric ambience replication are compression tools that allow the frequency-dependent parametric prediction of additional mono signals to be spatially distributed in order to complement the so far spatially incomplete or missing compressed HOA representations. The prediction is based on the coefficient sequence of the basic compressed sound field representation. An important aspect is that said complementary contributions to the sound field are represented in the compressed HOA representation not by additional quantized signals but by additional side information of a comparably much smaller size. Thus, the two coding tools mentioned above are particularly suitable for the compression of HOA representations at low data rates.

[0059] A second example of a compressed representation of a mono signal with the structure described above may consist of the following components: 1. Some coded spectral information about distinct frequency bands up to some upper frequency limit. This can be considered a basic compressed representation. 2. Some basic side information specifying the above coded spectral information (eg, by the number and width of the coded frequency bands). 3. Some enhanced side information consisting of so-called Spectral Band Replication (SBR) parameters. These parameters describe how to parametrically reconstruct from the basic compressed representation the spectral information for higher frequency bands that are not considered in the basic compressed representation.

[0060] Next, a method for layered coding of the complete compressed sound (or sound field) representation with the above structure is described.

[0061] The compression is assumed to be frame-based, in the sense that it gives compressed representations (e.g., in the form of data packets, or equivalently, frame payloads) of a series of time intervals, e.g., of equal size. These data packets are assumed to contain, in addition to the actual compressed representation data, a validity flag and a value indicating their size. Throughout the following description, we will mostly focus on dealing with single frames; therefore, the frame index is omitted.

[0062] Each frame payload of the considered complete compressed sound (or sound field) representation 1100 is assumed to contain J data packets, each for one component 1110-1, ..., 1110-J of the elementary compressed sound (or sound field) representation, denoted as BSRCj, j = 1, ..., J. Furthermore, each frame payload is I It is assumed that the packet contains independent basic side information 1120 denoted by BSI. I are the specific components of the basic compressed sound representation BSRC independently of other components. j Optionally, each frame payload also specifies the BSI D It is assumed that the packet contains a dependent basic side information denoted BSI. D The specific components of the basic compressed sound representation BSRC depend on other components. j Specify two data packets BSI I and BSI D The information contained within can optionally be grouped into a single data packet BSI.

[0063] Finally, each frame payload contains an enhancement side information payload, denoted ESI, with a description of how to improve the reconstructed sound (or sound field) from the full underlying compressed representation.

[0064] The described scheme for layered coding addresses the required steps to enable both the compressor, which includes packing data packets for transmission, and the receiver and decompressor, each of which is described in detail below.

[0065] Compression and packing for transmission is now discussed. For layered coding (with a total of M layers, i.e., one base layer and M-1 enhancement layers), each component of the complete compressed sound (or sound field) representation 1100 is treated as follows:

[0066] The basic compressed sound (or sound field) representation is subdivided into parts that are assigned to individual layers. Without loss of generality, the grouping is made up of M+1 numbers J m , m=0,…,M, where J0=1, J M =J+1, and BSRC j is J m-1 ≦j <J m is assigned to the m-th layer.

[0067] Due to its small size, it is reasonable to allocate the complete basic side information to the base layer to avoid its unnecessary fragmentation. I is left unchanged for the assignment, while the dependent basic side information needs to be treated specially for layered coding to allow correct decoding at the receiver and to reduce the size of the transmitted dependent side information. D,m , m=1,...,M, where the mth part is the component BSRC of the basic compressed sound representation assigned to the mth layer. j , J m-1 ≦j <J m If the dependent side information is not present, the BSI D,m is assumed to be empty. Side information BSI D,mis the sum of all components in all layers up to the mth layer. j , 1 ≤ j <J m Depends on.

[0068] In the case of layered coding, it is important to realize that the enhancement side information needs to be calculated extra for each layer, since it is intended to enhance the pre-decompressed sound (or sound field). However, it depends on the available layers for decompression. Therefore, compression is performed using ESI. m , m=1,...,M, where the enhancement side information ESI m is calculated to enhance the sound (or sound field) representation obtained from all the data contained in the base layer and in the enhancement layers with indices lower than m.

[0069] In summary, the compression stage must provide a frame data packet, denoted FRAME, with the following composition:

[0070]

number

[0071] The allocation of individual payloads to base and enhancement layers, as already mentioned, is achieved by a so-called transport layer packer, which is shown schematically in FIG.

[0072] Receiving and decompressing will now be described. The corresponding receiver and decompression stages are shown in FIG.

[0073] First, the individual layer packets 1200, 1300-1, ..., 1300-(M-1) are multiplexed to form the received frame packet of the complete compressed sound (or sound field) representation.

[0074]

number

[0075] In the decompressor 2100, the received frame packets are first demultiplexed. For this purpose, information about the size of each payload may be utilized to avoid unnecessary parsing through the data of the individual payloads.

[0076] The next stage is to determine the highest layer number N that is actually used for decompressing the basic sound representation. B is selected. The best one actually used for decompressing the basic sound representation is improvement Layer N B -1. Since each layer contains exactly one enhancement side information payload, it is possible to know from each enhancement side information payload whether the containing layer is valid or not. Thus, the selection is made by selecting all enhancement side information payloads ESI m , m=1,...,M. In addition, the index N of the enhanced side information payload used for decompression E is determined. This is always N B or equal to 0. That is, the improvement is always achieved according to the basic sound expression, or not achieved at all. A more detailed description of the selection is given further below.

[0077] Sequentially, the basic compressed sound representation components BSRC1, ..., BSRC J The payload of the Basic Side Information Payload (i.e., BSI I and BSID,m , m=1,…,M) and value N B The basic representation decompression processing unit 2200 then passes the lowest N B layers (i.e., the base layer and N B - one enhancement layer) to reconstruct the basic sound (or sound field) representation using only the basic compressed sound representation components contained in the individual layers. It is assumed that the required information about which components of the basic compressed sound (or sound field) representation are contained in the individual layers is known to the decompressor 2100 from data packets with configuration information. It is assumed that the configuration information is sent and received before the frame data packets. Each individual subordinate Basic Side Information payload BSI D,m , m=1,…,N B The actual decoding of can be split into two parts:

[0078] 1. Each payload BSI D,m , m=1,…,N B This is the first J contained in the first m layers assumed in the encoding stage. m - 1 basic compressed sound representation component BSRC1, ..., BSRC (Jm)-1 By leveraging the dependency on

[0079] 2. Each payload BSI D,m , m=1,…,N B This is because the first N components have more components than were assumed for preliminary decoding. B > the first J in layer m NB - 1 basic compressed sound representation component BSRC1, ..., BSRC (JNB)-1Therefore, correction can be achieved by discarding outdated information. This is possible due to an initially assumed property of the dependent fundamental side information: if certain complementary components are added to the basic compressed sound (or sound field) representation, then the dependent fundamental side information for each individual complementary component will be a subset of the original.

[0080] Finally, the reconstructed basic sound (or sound field) representation is obtained by adding all the enhancement side information payloads ESI1, ..., ESI M , Basic Side Information Payload BSI I and BSI D,m , m=1,…,M and value N E The enhanced representation decompression processing unit 2300 receives the enhanced side information payload ESI NE Use only N and discard all other enhancement side information payloads to compute the final enhanced sound (or sound field) representation. E If the value of is equal to 0, all enhancement side information payloads are discarded and the final reconstructed enhanced sound (or sound field) representation is equal to the reconstructed basic sound (or sound field) representation.

[0081] Next, we discuss layer selection: if all frame data packets can be decompressed independently of each other, then the highest layer number N that will actually be used for decompressing the basic sound representation is chosen. B and the index N of the enhanced side information payload used for decompression E are both set to the highest numbered valid enhancement side information payload, L, which itself can be determined by evaluating the validity flag within the enhancement side information payload. By leveraging knowledge of the size of each enhancement side information payload, complex parsing through the actual data of the payload to determine validity can be avoided.

[0082] When differential decompression is used, where there is inter-frame dependency, decisions from previous frames must also be taken into account. In differential decompression, independent frame data packets are transmitted at regular time intervals to allow decompression to start from those points. For independent frame data packets, the value N B and N E The determination of is now frame independent and is performed as above.

[0083] To explain the frame-dependent decision in detail, first for the kth frame, Let L(k) be the highest number of valid enhancement side information payloads, N is the highest layer number selected and used for decompressing the basic sound representation. B (k) The number of enhanced side information payloads used for decompression is N. E This is expressed as (k).

[0084] Using this notation, the highest layer number N used for decompression of the basic sound representation B (k) is calculated according to the following formula:

[0085]

number

[0086] Number N of enhanced side information payloads used for decompression E (k) is determined according to the following formula:

[0087]

number

[0088] Alternatively, in decompression, N E If all of the enhanced side information payloads numbered up to (k) are decompressed in parallel, selection rule (4) becomes N E (k)=N B (k) (5) may be replaced by

[0089] Finally, note that for differential decompression, the number of the highest used layer can only be increased in an independent frame data packet, while a decrease is possible in any frame.

[0090] Next, embodiments of the present disclosure relating to layered coding of frames of compressed audio representations and data structures (e.g., bitstreams) representing encoded frames of compressed audio representations are described for the case of compressed HOA representations, and in particular proposed modifications to the layered coding scheme of compressed HOA representations are described.

[0091] As a modification of the layered coding mode for HOA-based content, a new usacExtElementType is defined to better adapt the configuration settings and frame payloads of the HOA decoding tools spatial signal prediction, subband directional signal synthesis, and parametric ambience replication (PAR) decoders to the corresponding HOA enhancement layers. When the layered coding mode for HOA-based content is activated, which is signaled by SingleLayer==0, it is proposed to move the corresponding bitstream elements of these tools into one additional HOA extension payload of this new type for each layer (including the base layer and one or more enhancement layers).

[0092] The need for extensions arises because the side information for these tools is created to enhance a specific HOA representation. In the current definition of layered HOA encoding, the data provided only adequately enhances the HOA representation at the top layer. For lower layers, these tools do not adequately enhance the partially reconstructed HOA representation.

[0093] Therefore, it would be better to provide side information for these tools for each layer to adapt these tools to the reconstructed HOA representation of the corresponding layer.

[0094] Furthermore, the tools for subband directional signal synthesis and parametric ambience replication decoder are specifically designed for low data rates where only a few transport signals are available. Therefore, the proposed extension provides the ability to optimally adapt the side information of these tools to the number of transport signals in a layer. Thus, the sound quality of the reconstructed HOA representation for low bit-rate layers, e.g., the base layer, can be significantly enhanced compared to existing layered approaches.

[0095] Furthermore, if CodedVVecLength equal to 1 is signaled in HOADecoderConfig(), the bitstream syntax for the encoded V vector elements for vector-based signals needs to be adapted for HOA layer configuration coding. In this vector coding mode, no V vector elements are transmitted for HOA coefficient indices included in the set ContAddHoaCoeff. This set includes all HOA coefficient indices AmbCoeffIdx[i] with AmbCoeffTransitionState equal to 0. Since the original HOA coefficient sequence for these indices is explicitly sent, there is no need to add a weighted V vector signal either. Therefore, the V vector elements in the normal approach are set to 0 for these indices.

[0096] However, in layered coding mode, the set of consecutive HOA coefficient indices depends on the transport channel that is part of the currently active layer. That is, additional HOA coefficient indices sent in higher layers are missing in lower layers. And the assumption that the vector signal should not contribute to the HOA coefficient sequence is incorrect for HOA coefficient indices that belong to the HOA coefficient sequence included in higher layers. Therefore, it is proposed to (explicitly) signal the V vector elements for these missing coefficient indices.

[0097] As a result, it is proposed to define a set of ContAddHoaCoeff for each layer and use the set of layers to which the V-vector signal is added (the layer to which its transport signal belongs) for the selection of the active V-vector elements. Nevertheless, it is proposed that the V-vector data stay in HOAFrame() and are not transferred to HOAEnhFrame().

[0098] Next, integration into the MPEG-H bitstream syntax is described. A corresponding encoding method (e.g., a method for encoding a layered structure of frames of a compressed HOA representation of a sound or sound field) according to an embodiment of the present disclosure is described with reference to Figure 3. Proposed modifications to the MPEG-H 3D bitstream are described later in the appendix.

[0099] In layered coding mode, the flag SingleLayer in HOADecoderConfig() is inactive (SingleLayer==0), and the number of layers and the corresponding number of HOA transport signals assigned to those layers are defined. In general, a compressed HOA representation may contain multiple transport signals.

[0100] Thus, in S3010 of FIG. 3, multiple transport signals are assigned to multiple hierarchical layers. In other words, the transport signals are distributed to multiple layers. Each layer may be said to include a respective transport signal assigned to that layer. Each layer may have two or more transport signals assigned to it. The multiple layers may include a base layer and one or more hierarchical enhancement layers. The layers may be ordered from the base layer, through the enhancement layers, to the overall highest enhancement layer (overall top layer).

[0101] It is proposed to add additional HOA configuration extension payloads and HOA frame extension payloads with the newly defined usacExtElementType ID_EXT_ELE_HOA_ENH_LAYER to the MPEG-H bitstream to carry one payload for spatial signal prediction, subband directional signal synthesis, and PAR decoder data for each HOA enhancement layer (including the base layer). These additional payloads immediately follow the payloads of type ID_EXT_ELE_HOA in mpegh3daExtElementConfig() and correspondingly in mpegh3daFrame().

[0102] Therefore, when SingleLayer==0, it is proposed to move the configuration elements for spatial signal prediction, sub-band directional signal synthesis and PAR decoder from HOADecoderConfig() to the newly defined HOADecoderEnhConfig(), and correspondingly move HOAPredictionInfo(), HOADirectionalPredictionInfo() and HOAParInfo() from HOAFrame() to the newly defined HOAEnhFrame().

[0103] Thus, at S3020, a respective HOA extension payload is generated for each layer. The generated HOA extension payload may include side information for parametrically enhancing the reconstructed HOA representation obtained from the transport signal assigned to (e.g., included in) the respective layer. As indicated above, the HOA extension payload may include bitstream elements for one or more of the HOA spatial signal prediction decoding tools, the HOA subband directional signal synthesis decoding tools, and the HOA parametric ambience replication decoding tools. Furthermore, the HOA extension payload may have a usacExtElementType of ID_EXT_ELE_HOA_ENH_LAYER.

[0104] In S3030, the generated HOA extension payload is assigned to each layer.

[0105] Additionally (not shown in FIG. 3), an HOA configuration extension payload may be generated that includes bitstream elements for configuring the HOA spatial signal prediction decoding tool, the HOA subband directional signal synthesis decoding tool, and / or the HOA parametric ambient sound replication decoding tool.

[0106] Additionally (not shown in FIG. 3), an HOA decoder configuration payload may be generated that includes information indicating the allocation of HOA extension payloads to the multiple layers.

[0107] Next, we consider the transmission of layered bitstreams (e.g., MPEG-H bitstreams). Because all extension payloads in an MPEG-H bitstream are byte-aligned and their size is explicitly signaled, assuming an elementLengthPresent flag equal to 1, a depacker can parse the MPEG-H bitstream, extract payloads for layers higher than 1, and transmit those payloads separately through various transmission channels. The base layer comprises (e.g., consists of) an MPEG-H bitstream excluding higher layers. Missing extension payloads are signaled as empty or inactive. For payloads of type ID_USAC_SCE, ID_USAC_CPE, and ID_USAC_LFE, an empty payload is signaled by an elementLength of 0, where elementLengthPresent must be set to 1. An empty payload of type ID_USAC_EXT can be signaled by setting the usacExtElementPresent flag to 0 (false).

[0108] Thus, at S3040, the generated HOA extension payload is signaled (e.g., transmitted or output) in the output bitstream. Generally, the multiple layers and their assigned payloads are signaled (e.g., transmitted or output) in the output bitstream. Additionally, an HOA decoder configuration payload and / or an HOA configuration extension payload may be signaled (e.g., transmitted or output) in the output bitstream.

[0109] The HOA base layer (layer index equal to 1) is assumed to be transmitted with the highest error protection and have a relatively small bit rate. The error protection for subsequent layers (one or more HOA enhancement layers) is progressively reduced with the increasing bit rate of the enhancement layers. Due to poor transmission conditions and lower error protection, the transmission of higher layers may fail, and in the worst case, only the base layer is transmitted correctly. A combined error protection for all payloads of a layer is assumed to be applied. Thus, if the transmission of a layer fails, all payloads of the corresponding layer are missing.

[0110] In other words, data payloads for multiple layers may be transmitted with respective levels of error protection, where the base layer has the highest error protection and the one or more enhancement layers have successively decreasing error protection.

[0111] It will be understood that the steps described above may be performed in any order, and the exemplary order shown in FIG. 3 is not limiting, unless a step requires another step as a prerequisite.

[0112] As indicated above, if CodedVVecLength equal to 1 is signaled in HOADecoderConfig(), the bitstream syntax for encoded V vector elements for vector-based signals needs to be adapted for HOA layered coding. A corresponding encoding method (e.g., layered encoding method for a frame of a compressed HOA representation of a sound or sound field) according to an embodiment of the present disclosure will be described with reference to FIG.

[0113] In S4010 of Figure 4, multiple transport signals are allocated to multiple hierarchical layers. This step may be performed in the same manner as S3010 above.

[0114] In S4020, it is determined whether vector coding mode is active, which may involve determining whether CodedVVecLength==1.

[0115] As shown above, in the conventional approach in vector coding mode, no V vector elements are transmitted for HOA coefficient indices included in the set ContAddHoaCoeff. This set includes all HOA coefficient indices AmbCoeffIdx[i] with AmbCoeffTransitionState equal to 0. Because the original HOA coefficient sequence for these indices is explicitly sent, there is no need to add a weighted V vector signal either. Therefore, the V vector elements in the conventional approach are set to 0 for these indices.

[0116] However, in layered coding mode, the set of consecutive HOA coefficient indices depends on the transport channels that are part of the currently active layer, i.e., additional HOA coefficient indices sent in higher layers are missing in lower layers, and the assumption that vector signals should not contribute to the HOA coefficient sequence is incorrect for HOA coefficient indices that belong to the HOA coefficient sequence included in higher layers.

[0117] Thus, if vector coding mode is active, then in S4030, for each layer, a set of consecutive HOA coefficient indices (eg, ContAddHoaCoeff) is determined (eg, defined) based on the transport signal assigned to the respective layer.

[0118] If vector coding mode is active, then in S4040, for each transport signal, a V-vector is generated based on the determined set of consecutive HOA coefficient indices for the layer to which the transport signal is assigned. Each generated V-vector may include elements for any transport signals assigned to a layer higher than the layer to which the transport signal is assigned. This step may involve using the determined set of consecutive HOA coefficient indices for the layer to which the V-vector signal is applied (the layer to which the V-vector signal's transport signal belongs) for the selection of active V-vector elements. Nevertheless, it is proposed that the V-vector data remain in HOAFrame() and are not transferred to HOAEnhFrame().

[0119] Then, at S4050, the generated V vector (V vector signal) is signaled in the output bitstream, which may involve (explicitly) signaling V vector elements for the missing coefficient indices mentioned above.

[0120] Steps S4020 to S4050 of Figure 4 may also be used in the context of the encoding method shown in Figure 3, for example after S3010, in which case S3040 and S4050 may be combined into a single signalling step.

[0121] It will be understood that the steps described above may be performed in any order, and the exemplary order shown in FIG. 4 is not limiting, unless a step requires another step as a prerequisite.

[0122] At the receiver side, the MPEG-H bitstream packer can re-insert correctly received payloads into the base layer MPEG-H bitstream and pass it to the MPEG-H 3D audio decoder.

[0123] Next, we describe HOA decode initialization (configuration). HOA configuration payloads of type ID_EXT_ELE_HOA and ID_EXT_ELE_HOA_ENH_LAYER, along with their corresponding sizes in bytes, are input to the HOA decoder for its initialization. The HOA encoding tool is configured according to the bitstream elements defined in HOAConfig() parsed from the payload of type ID_EXT_ELE_HOA. Furthermore, the payload contains the layered coding mode used, the number of layers, and the corresponding number of transport signals per layer. Then, if layered coding is activated (SingleLayer==0), HOAEnhConfig() is parsed from the payload of type ID_EXT_ELE_HOA_ENH_LAYER to configure the corresponding spatial signal prediction, subband directional signal synthesis, and parametric ambience replication decoders for each layer.

[0124] The element LayerIdx from HOAEnhConfig(), together with the order of the HOA enhancement layer configuration payloads in mpegh3daExtElementConfig(), indicates the order of the HOA enhancement layers. To unambiguously assign frame payloads to the corresponding layers, the order of HOA enhancement layer frame payloads of type ID_EXT_ELE_HOA_ENH_LAYER in mpegh3daFrame() is the same as the order of the configuration payloads in mpegh3daExtElementConfig().

[0125] If SingleLyaer==1 (single layer coding), payloads of type ID_EXT_ELE_HOA_ENH_LAYER are ignored and the spatial signal prediction, subband directional signal synthesis and parametric ambience replication decoders use the corresponding data from HOADecoderConfig() for their configuration.

[0126] Next, HOA frame decoding in layered mode will be described. A corresponding method of decoding (e.g., a method of decoding a frame of a compressed HOA representation of a sound or sound field) according to an embodiment of the present disclosure will be described with reference to Figure 5. It will be understood that the compressed HOA representation (e.g., the output of the method of Figure 3 or Figure 4 above) may be encoded in multiple hierarchical layers, including a base layer and one or more hierarchical enhancement layers.

[0127] At S5010 of FIG. 5, a bitstream relating to a frame of a compressed HOA representation is received.

[0128] The 3D audio core decoder decodes correctly transmitted HOA transport signals and generates transport signals where all samples are equal to 0 for the corresponding invalid payloads. The decoded transport signals are input to the HOA decoder along with the usacExtElementPresent flag, data, and size of HOA payloads of type ID_EXT_ELE_HOA and ID_EXT_ELE_HOA_ENH_LAYER. Extension payloads from type ID_USAC_EXT with the usacExtElementPresent flag set to false must be signaled to the HOA decoder as missing payloads to ensure allocation of the payload to the corresponding layer.

[0129] At S5020, payloads of multiple layers are extracted, each of which may include a transport signal assigned to a respective layer.

[0130] At this stage, the HOA decoder may parse a HOAFrame() from the payload of type ID_EXT_ELE_HOA.

[0131] Then, valid payloads of type ID_EXT_ELE_HOA_ENH_LAYER and invalid payloads of type ID_EXT_ELE_HOA_ENH_LAYER are determined by evaluating their corresponding usacExtElementPresent flags, where invalid payloads are indicated by a usacExtElementPresent flag equal to false, and the assignment of HOA enhancement payloads to enhancement layer indices is known from the HOA decoder configuration settings.

[0132] In S5030, the highest available layer for decoding among the plurality of layers is determined.

[0133] Because layers are interdependent on the transport signal, the HOA decoder can only decode a layer if all layers with lower indices are received correctly. The highest available layer may be selected at this stage such that all layers up to the highest available layer have been received correctly. This stage is described in more detail below.

[0134] At S5040, the HOA extension payload assigned to the highest available layer is extracted. As indicated above, the HOA extension payload may include side information for parametrically enhancing the reconstructed HOA representation corresponding to the highest available layer. Here, the reconstructed HOA representation corresponding to the highest available layer may be obtained based on transport signals assigned to the highest available layer and any layers lower than the highest available layer.

[0135] Further, HOA extension payloads assigned to the remaining layers of the plurality of layers may be extracted. Each HOA extension payload may include side information for parametrically enhancing the reconstructed HOA representation corresponding to the assigned layer. The reconstructed HOA representation corresponding to the assigned layer may be obtainable from the transport signal assigned to that layer and any layers below it.

[0136] Additionally (not shown in FIG. 5 ), the decoding method may include extracting an HOA configuration extension payload, which may be done by parsing the bitstream. The HOA configuration extension payload may include bitstream elements for configuring an HOA spatial signal prediction decoding tool, an HOA subband directional signal synthesis decoding tool, and / or an HOA parametric ambient sound replication decoding tool.

[0137] In S5050, a (partially) reconstructed HOA representation corresponding to the highest available layer is generated based on the transport signals assigned to the highest available layer and any layers lower than the highest available layer.

[0138] Number of transport signals actually used I ADD,LAY (k) is the index of the highest available layer (M LAY (k)), a first preliminary HOA representation is decoded from the HOAFrame() and from the corresponding transport signal at that layer and any lower layers.

[0139] Then, in S5060, the reconstructed HOA representation is enhanced (parametrically enhanced) according to the side information contained in the HOA extension payload assigned to the highest available layer.

[0140] That is, the HOA representation obtained in S5050 is then enhanced by spatial signal prediction, sub-band directional signal synthesis and a parametric ambient sound replica decoder. LAY (k), i.e., enhanced using HOAEnhFrame() data parsed from the HOA enhancement layer extension payload of the highest available layer type ID_EXT_ELE_HOA_ENH_LAYER.

[0141] The information used in steps S5020 to S5060 may be known as layer information.

[0142] It will be understood that the steps described above may be performed in any order, and the exemplary order shown in FIG. 5 is not limiting, unless a step requires another step as a prerequisite.

[0143] Details of determining (eg, selecting) the best available layer in S5030 are now described.

[0144] As shown above, the layers are interdependent in the transport signal, so the HOA decoder can only decode a layer if all layers with lower indices are received correctly.

[0145] For the selection of the best decodable layer, the HOA decoder can create a set of invalid layer indices, where the smallest index from this set minus 1 is the best decodable enhancement layer index M LAY The set of invalid layer indices may be determined by evaluating the validity flags of the corresponding HOA extension payloads.

[0146] In other words, determining the highest available tier may involve determining a set of invalid tier indices that indicate tiers that were not validly received. It may further involve determining the highest available tier as the tier that is one tier below the tier indicated by the smallest index in the set of invalid tier indices, thereby ensuring that all tiers below the highest available tier have been validly received.

[0147] For frame differential encoding, the highest available layer index of the previous (e.g., immediately preceding) frame needs to be taken into account. First, we describe the situation where the highest available layer index of the previous (e.g., immediately preceding) frame is retained.

[0148] The index of the highest available layer (e.g., highest decodable layer) for the current frame is the layer index M of the previous frame. LAY If it is equal to (k-1), the layer index M of the current frame LAY (k) is M LAY It is set to (k-1).

[0149] Next, the number of transport signals actually used, I ADD,LAY (k) is M LAY (k), a first preliminary HOA representation is decoded from the HOAFrame() and from the corresponding transport signal of that layer and any lower layers, as shown above. This HOA representation is then used to decode the currently active layer M, as shown above. LAY Enhanced by spatial signal prediction, sub-band directional signal synthesis and parametric ambient sound replication decoder using HOAEnhFrame() data parsed from the HOA enhancement layer extension payload of type ID_EXT_ELE_HOA_ENH_LAYER in (k).

[0150] Next, we consider the situation where the highest available layer index for the current frame is switched to a lower index than the highest available layer index of the previous (e.g., immediately preceding) frame, i.e., the highest decodable layer index for the current frame is lower than the layer index M of the previous frame. LAY If it is smaller than (k-1), the HOA decoder LAY Set (k) to the index of the highest decodable layer for the current frame. Payload decoding for spatial signal prediction, subband directional signal synthesis and parametric ambient sound replication decoders for the new layer can start only in the next HOA frame with hoaIndependencyFlag equal to 1. Until such an HOAFrame() is received, the index M LAY The HOA representation of layer (k) is reconstructed without performing spatial signal prediction, subband directional signal synthesis, and parametric ambient sound replica decoder. That is, the number of transport signals actually used, I ADD,LAY (k) is M LAY Based on (k), a first preliminary HOA representation is decoded from HOAFrame() from the corresponding transport signal of that layer and any lower layers. Then, if an HOAFrame() with hoaIndependencyFlag equal to 1 is received, the payloads for spatial signal prediction, subband directional signal synthesis, and parametric ambience replication decoders are parsed and decoded to enhance the preliminary HOA representation, thereby providing the full quality of the currently active layer for this frame.

[0151] Thus, the proposed method may include deciding not to perform parametric enhancement of the reconstructed HOA representation using side information contained in the HOA extension payload assigned to the highest available layer if the highest available layer of the current frame is lower than the highest available layer of the previous frame (in the case where the current frame is differentially coded with respect to the previous frame) (not shown in Figure 5).

[0152] In general, determining the highest usable layer for the current frame may involve determining a set of invalid layer indices indicating layers that have not been validly received for the current frame. It may further include determining the highest usable layer of a previous frame preceding the current frame. It may further include determining the highest usable layer as the lower of the highest usable layer of the previous frame and a layer that is further below the layer indicated by the smallest index in the set of invalid layer indices (if the current frame is differentially coded with respect to the previous frame).

[0153] An alternative solution may always parse all enabled enhancement layer payloads (e.g., HOA extension payloads) in parallel, even if they are currently inactive. This allows for direct switching to layers with lower indices with full quality. Spatial signal prediction, subband directional signal synthesis, and parametric ambience replication (PAR) decoders can be applied directly in the switched frame.

[0154] Next, we discuss the situation where a frame switches to an index higher than the highest available layer index of the previous (e.g., immediately preceding) frame. This switching to a layer with a higher index is only applicable if mpegh3daFrame() has usacIndependencyFlag equal to 1 (e.g., the frame is an independent frame) because all corresponding payload or decoding state of the previous frame is missing. Thus, the HOA decoder will continue to update the HOA layer index M until an mpegh3daFrame() with usacIndependencyFlag equal to 1 (e.g., an independent frame) is received that contains valid data for a higher decodable layer. LAY (k) to M LAY (k-1). Then, M LAY (k) is set to the highest decodable layer index for the current frame, and thus the number of transport signals actually used, IADD,LAY The preliminary HOA representation for that layer is decoded from the HOAFrame() and the corresponding transport signal, and the currently active layer M LAY Enhanced by spatial signal prediction, sub-band directional signal synthesis and parametric ambience replication decoder using HOAEnhFrame() parsed from the HOA enhancement layer extension payload of type ID_EXT_ELE_HOA_ENH_LAYER in (k).

[0155] It will be understood that the proposed method for encoding a layered structure of a compressed audio representation can be implemented by an encoder for encoding a layered structure of a compressed audio representation. Such an encoder may comprise respective units adapted to perform the respective steps described above. An example of such an encoder 6000 is shown schematically in FIG. 6. For example, such an encoder 6000 may comprise a transport signal allocation unit 6010 adapted to perform S3010 described above, an HOA enhancement layer payload generation unit 6020 adapted to perform S3020 described above, an HOA enhancement payload allocation unit 6030 adapted to perform S3030 described above, and a signaling or output unit 6040 adapted to perform S3040 described above. Furthermore, it will be understood that each unit of such an encoder may be embodied by a processor 6100 of a computing device adapted to perform the processing performed by each of said units, i.e. adapted to perform some or all of the above-mentioned steps of the proposed encoding method schematically shown in FIG. 3. Additionally or alternatively, the processor 6100 may be adapted to perform each of the steps of the encoding method schematically illustrated in Figure 4. To this end, the processor 6100 may be adapted to implement each unit of an encoder. The encoder or computing device may further comprise a memory 6200 accessible by the processor 6100.

[0156] It will further be understood that the proposed method for decoding a compressed sound representation encoded in multiple hierarchical layers can be implemented by a decoder for decoding a compressed sound representation encoded in multiple hierarchical layers. Such a decoder may have respective units adapted to perform the respective steps described above. An example of such a decoder 7000 is shown schematically in Figure 7. For example, such a decoder 7000 may have a receiving unit 7010 adapted to perform S5010 described above, a payload extraction unit 7020 adapted to perform S5020 described above, a highest available layer determination unit 7030 adapted to perform S5030 described above, an HOA extended payload extraction unit 7040 adapted to perform S5040 described above, a reconstructed HOA representation generation unit 7050 adapted to perform S5050 described above, and an enhancement unit 7060 adapted to perform S5060 described above. It will further be understood that each unit of such a decoder may be embodied by a processor 7100 of a computing device adapted to perform the processing performed by each of said units, i.e. adapted to perform some or all of the above-mentioned steps of the proposed decoding method. The decoder or computing device may further comprise a memory 7200 accessible by the processor 7100.

[0157] Next, we describe a data structure (e.g., bitstream) for receiving (e.g., representing) a compressed HOA representation in layered coding mode. Such a data structure may result from using the proposed encoding method and may be decoded (e.g., decompressed) by the proposed decoding method.

[0158] The data structure may include multiple HOA frame payloads corresponding to each of multiple hierarchical layers. The multiple transport signals may be assigned to (e.g., belong to) the multiple layers. The data structure may include respective HOA extension payloads containing side information for parametrically enhancing a reconstructed HOA representation obtained from the transport signals assigned to each layer and any layers below the respective layer. As indicated above, the HOA frame payloads and HOA extension payloads for the multiple layers may be provided with respective levels of error protection. Furthermore, the HOA extension payloads may include the bitstream elements described above and may have a usacExtElementType of ID_EXT_ELE_HOA_ENH_LAYER. The data structure may further include an HOA configuration extension payload and / or an HOA decoder configuration payload containing the bitstream elements described above.

[0159] It should be noted that the present description and drawings merely illustrate the principles of the proposed method and apparatus. It is therefore understood that those skilled in the art will be able to devise various configurations that embody the principles of the present invention and are within its spirit and scope, even if not explicitly described or shown herein. Furthermore, all examples described herein are expressly intended solely for educational purposes to aid the reader in understanding the principles of the proposed method and apparatus and the concepts contributed by the inventors to the advancement of the art, and are to be construed without limitation to such specifically described examples and conditions. Furthermore, all statements herein describing principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass equivalents thereof.

[0160] The methods and apparatus described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented, for example, as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or application-specific integrated circuits. Signals emerging from the described methods and apparatus may be stored on media such as random access memory or optical storage media, or transmitted over a network such as a radio, satellite, wireless, or wired network, e.g., the Internet.

[0161] Annex: Proposed MPEG-H 3D Bitstream Modifications Changes are marked with a grey highlight.

[0162] [Table 1-1]

[0163] [Table 1-2] Note (Table 1): For unknown extElementTypes, the default entry for usacExtElementType is used, allowing legacy decoders to accommodate future extensions.

[0164] [Table 2] Note (Table 2): Application-specific usacExtElementType values ​​are required in the space reserved for use outside the ISO range. These are skipped by decoders since minimal structure is required by decoders to skip these extensions.

[0165] [Table 3]

[0166] [Table 4-1]

[0167] [Table 4-2] Note (Table 4): MinAmbHoaOrder=30...37 are reserved. HOAFrameLengthIndicator=3 is reserved.

[0168] [Table 4-3]

[0169] [Table 4-4]

[0170] [Table 4-5]

[0171] [Table 5-1]

[0172] [Table 5-2]

[0173] [Table 5-3] Note (Table 5): If usacIndependencyFlag (see mpegh3daFrame()) is set to 1, the encoder sets hoaIndependencyFlag to 1. Note: If SingleLayer==1, set NumLayers=1. NumOfDirSigsPerLayer[lay] This element determines the number of active directional signals in the current HOAFrame() that are actually used in the HOA enhancement layer lay. AddHoaCoeffPerLayer[lay] This array contains the HOA coefficient index for each additional ambient HOA coefficient actually used in the HOA enhancement layer lay. NumOfAddHoaChansPerLayer[lay] This element signals the total number of additional ambient HOA coefficients actually used in the HOA enhancement layer lay.

[0174] Add this table.

[0175] [Table 5-4] Note: lay is the index of the currently active HOA improvement layer.

[0176] Update this table.

[0177] [Table 6-1]

[0178] [Table 6-2] Note (Table 6): See ○ for calculation of VVecLength.

[0179] [Table 7-1]

[0180] [Table 7-2] Note (Table 7):lay is the index of the currently active HOA improvement layer.

[0181] [Table 7-3]

[0182] [Table 7-4]

[0183] [Table 7-5] Note (Table AMD1.2):lay is the index of the currently active HOA improvement layer.

[0184] [Table 8] codedLayerCh This element indicates the number of transport signals included for the first (i.e., base) layer. That number is given by codedLayerCh+MinNumOfCoeffsForAmbHOA. For higher (i.e., enhancement) layers, this element indicates the number of additional signals included in the enhancement layer compared to the next lower layer. It is given by codedLayerCh+1. HOALyaerChBits This element indicates the number of bits to read codedLayerCh. NumLayers This element indicates the total number of layers in the bitstream (after reading HOADecoderConfig()). NumHOAChannelsLayer This element is an array consisting of NumLayers elements, and the i-th element indicates the number of transport signals included in all layers up to the i-th layer.

[0185] 12.4.1.x Frame and User-Dependent Parameters M LAY(k) The number of all layers actually used for the kth frame (see below) at the decoder side. In the case of layered coding (indicated by SingleLayer==0), this number must be less than or equal to the total number of layers present in the bitstream. That is, M LAY ≦NumLayers. In the case of single layer coding (indicated by SingleLayers==1), M LAY is set to 1.

[0186] M LAY Depending on the choice of (k), there may be additional (i.e., always implicit) Os actually used for spatial HOA decoding. MIN Number of transport channels (additional to the number of channels) I ADD,LAY (k) is calculated as follows:

[0187] [Table 9] VVecLength and VVecCoeffId codedVVecLength indicates: 0) Full vector length (NumOfHoaCoeffs elements). Indicates that all coefficients (NumOfHoaCoeffs) for the dominance vector are specified. 1) All elements defined in ContAddHoaCoeff[lay] of the vector elements 1 to MinNumOfCoeffsForAmbHOA and the currently active layer with index lay=0...NumLayers-1 are not transmitted. For the single layer mode SingleLayer==1, the variable NumLayers must be set equal to 1. This indicates that only coefficients of the dominant vector corresponding to numbers greater than MinNumOfCoeffsForAmbHOA are specified. Additionally, those NumOfContAddAmbHoaChan[lay] coefficients identified in ContAddAmbHoaChan[lay] are subtracted. The list ContAddAmbHoaChan[lay] specifies additional channels corresponding to orders exceeding the order ContAddAmbHoaChan[lay]. 2) Vector element 1 through MinNumOfCoeffsForAmbHOA is not transmitted, meaning that the coefficients of the dominant vector corresponding to numbers greater than MinNumOfCoeffsForAmbHOA are specified.

[0188] If codedVVecLength==1, then both the VVecLength[i] array and the VVecCoeffId[i][m] 2D array are valid for the V vector with index i. Otherwise, the VVecLength element and the VVecCoeffId[m] array are valid for all VVectors in the HOA frame. For the assignment algorithm below, a helper function is defined as follows:

[0189] [Table 10] The first switch statement with three cases (case 0-2) thus provides a way to determine the dominant vector length using the number of coefficients (VVecLength) and index (VVecCoeffId).

[0190] 12.4.1.X Conversion to VVec Elements The type of dequantization of the V vector is signaled by the word NbitsQ. An NbitsQ value of 4 indicates vector quantization. When NbitsQ is equal to 5, uniform 8-bit scalar dequantization is performed. In contrast, an NbitsQ value of 6 or greater indicates the application of Huffman decoding of the scalar-quantized V vector. The prediction mode is represented as PFlag, while CbFlag represents the Huffman table information bits.

[0191] [Table 11-1]

[0192] [Table 11-2]

Claims

[Claim 1] 1. A method for decoding a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field, the method comprising: receiving a bitstream including the compressed HOA representation, the bitstream including a base layer and a plurality of hierarchical layers including one or more hierarchical enhancement layers; determining a highest available layer of the plurality of hierarchical layers for decoding; determining that a parameter CodedVVecLength=0, and based on this determination, determining that all of the coefficients (NumOfHoaCoeffs) for the dominant vector are assigned; extracting an HOA extension payload assigned to the highest available layer, the HOA extension payload including side information for parametrically enhancing a reconstructed HOA representation corresponding to the highest available layer, the reconstructed HOA representation corresponding to the highest available layer being based on transport signals assigned to the highest available layer and any layers lower than the highest available layer; decoding the compressed HOA representation corresponding to the highest available layer based on layer information, the layer information indicating an active enhancement layer, the active enhancement layer being usable to determine a number of active directional signals in a current frame of the active enhancement layer; and parametrically enhancing the decoded HOA representation using side information contained in the HOA extension payload assigned to the highest available layer. method.