Layered encoding for compressed sound or sound field representation

The method for layered encoding of compressed sound representations addresses the challenges of adapting to varying transmission conditions by structuring components into hierarchical layers, ensuring efficient bandwidth use and robust decoding with incomplete data, thus enhancing the quality and efficiency of sound representation.

JP2025128202AActive Publication Date: 2025-09-02DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025089515
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-07-13
Filing Date
2025-05-29
Publication Date
2025-09-02
Estimated Expiration
2036-10-07

AI Technical Summary

Technical Problem

Existing layered coding schemes for compressed sound or sound field representations, such as Higher-Order Ambisonics (HOA), struggle to adapt to time-varying transmission conditions and prevent signal dropouts while efficiently managing error protection and bandwidth reduction.

Method used

A method for layered encoding of compressed sound representations, subdividing components into hierarchical layers with base and enhancement layers, where base side information is included in the base layer, and enhancement side information is structured to allow decoding from any available layer, ensuring efficient bandwidth use and robust decoding even with incomplete data reception.

Benefits of technology

Ensures optimal quality of reconstructed sound representations by maximizing available information use and minimizing bandwidth requirements, allowing decoding with incomplete data without redundant elements, thus enhancing the robustness and efficiency of layered coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128202000001_ABST
    Figure 2025128202000001_ABST
Patent Text Reader

Abstract

To provide a method for layered encoding for compressed sound or sound field representation and an encoder / decoder.SOLUTION: A method for layered encoding of compressed sound representation of sound or sound field subdivides a plurality of components into a plurality of groups and assigns each group into an individual hierarchical layer in a plurality of hierarchical layers. The number of groups corresponds to the number of layers and the plurality of layers includes basic layers and hierarchical enhancement layers. The method adds basic side information to the basic layers, discriminates a plurality of portions of enhancement side information from the enhancement side information and assigns each of the plurality of portions of the enhancement side information to the respective layer of the plurality of layers. Each portion of the enhancement side information includes a parameter which improves reconstituted sound representation, which is obtained from data included in each layer and any layer lower than such layer.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to European Patent Application No. 15306590.9, filed October 15, 2015, and U.S. Patent Application No. 62 / 361,809, the contents of which are incorporated herein by reference in their entireties.

[0002] Technical Field This document relates to methods and apparatus for layered audio coding. In particular, this document relates to methods and apparatus for layered audio coding of compressed sound (or sound field) representations, such as Higher-Order Ambisonics (HOA) sound (or sound field) representations. [Background technology]

[0003] For streaming of sound (or sound field) representations through transmission channels with time-varying conditions, layered coding is a means of adapting the quality of the received sound representation to the transmission conditions and in particular of avoiding unwanted signal dropouts.

[0004] For layered coding, the sound (or sound field) representation is typically subdivided into a high-priority base layer of relatively small size and additional enhancement layers of decreasing priority and arbitrary size. Each enhancement layer is typically assumed to contain incremental information to complement the information of all lower layers in order to improve the quality of the sound (or sound field) representation. The amount of error protection for the transmission of individual layers is controlled based on their priority. In particular, the base layer is provided with high error protection, which is reasonable and acceptable due to its small size.

[0005] However, there is a need for layered coding schemes for (extended versions of) special types of compressed representations of sounds or sound fields, such as compressed HOA sound or sound field representations. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] ISO / IEC JTC1 / SC29 / WG11 23008-3:2015(E), Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, February 2015 [Non-patent document 2] ISO / IEC JTC1 / SC29 / WG11 23008-3:2015 / PDAM3, Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, AMENDMENT 3: MPEG-H 3D Audio Phase 2, July 2015 Summary of the Invention [Problem to be solved by the invention]

[0007] This paper addresses the above problems. In particular, a method and encoder / decoder for layered coding of compressed sound or sound field representations is described. [Means for solving the problem]

[0008] According to one aspect, a method for layered encoding of a compressed sound representation of a sound or sound field is described. The compressed sound representation may include a base compressed sound representation including a plurality of components. The plurality of components may be complementary components. The compressed sound representation may further include base side information for decoding the base compressed sound representation into a base reconstructed sound representation of the sound or sound field. The compressed sound representation may further include enhancement side information including parameters for improving (e.g., enhancing) the base reconstructed sound representation. The method may include subdividing (e.g., grouping) the plurality of components into a plurality of component groups. The method may further include assigning (e.g., adding) each group of the plurality of groups to a respective one of a plurality of hierarchical layers. The assignment may indicate a correspondence between the respective groups and the layers. The components assigned to a respective layer may be said to be included in that layer. The number of groups may correspond to (e.g., be equal to) the number of layers. The plurality of layers may include a base layer and one or more hierarchical enhancement layers. The plurality of hierarchical layers may be ordered from a base layer, through a first enhancement layer, a second enhancement layer, etc., to an overall highest enhancement layer (overall top layer). The method may further include adding base side information to the base layer (e.g., including or allocating the base side information to the base layer, e.g., for transmission or storage purposes). The method may further include determining a plurality of portions of enhancement side information from the enhancement side information. The method may further include allocating (e.g., adding) each of the plurality of portions of enhancement side information to a respective layer of the plurality of layers. Each portion of enhancement side information may include parameters for improving a reconstructed (e.g., decompressed) sound representation obtained from data contained in (e.g., assigned or added to) that respective layer and any layers below that respective layer.The encoding of the layer configuration may be performed for transmission over a transmission channel or for storage on a suitable storage medium, such as a CD, DVD or Blu-ray Disc™.

[0009] Configured as described above, the proposed method allows layer-structured coding to be efficiently applied to compressed audio representations including multiple components and first and enhancement side information (e.g., independent base and enhancement side information) of the above-described nature. In particular, the proposed method allows each layer to include suitable side information for reconstructing a reconstructed audio representation from components contained in any layers up to the layer in question. Here, layers up to the layer in question are understood to include, for example, the base layer, the first enhancement layer, the second enhancement layer, etc. up to the layer in question. In this way, a decoder is enabled to improve or enhance the reconstructed audio representation even if it differs from the complete (e.g., full) audio representation, regardless of the actual highest available layer (e.g., layers below the lowest layer that have not yet been validly received; all layers below the highest available layer and the highest available layer itself have been validly received). In particular, it is sufficient for a decoder to decode the enhancement side information payload for only a single layer (i.e., for the highest available layer) to improve or enhance a reconstructed audio representation that can be obtained based on all components contained in layers up to the actual highest available layer, regardless of the actual highest available layer. That is, for each time interval (e.g., frame), only a single payload of enhancement side information may need to be decoded. On the other hand, the proposed method allows to take full advantage of the reduction in required bandwidth that can be achieved when applying layered coding.

[0010] In embodiments, the components of the elementary compressed sound representation may correspond to a mono signal (e.g., a transport signal or a mono transport signal), which may represent either a dominant sound signal or a coefficient sequence of an HOA representation, and which may be quantized.

[0011] In embodiments, the basic side information may include information that specifies the decoding (e.g., decompression) of one or more of said components individually, independently of the other components. For example, the basic side information may represent side information related to an individual mono signal, independent of other mono signals. Hence, the basic side information may sometimes be referred to as independent basic side information.

[0012] In embodiments, the enhancement side information may represent enhancement side information, which may include prediction parameters for the basic compressed sound representation for improving (e.g. enhancing) the basic compressed sound representation and the basic reconstructed sound representation obtained from the basic side information.

[0013] In some embodiments, the method may further include generating a transport stream for transmission of the data of the multiple layers (e.g., data assigned to, added to, or otherwise included in each layer). The base layer may have the highest transmission priority, and the hierarchical enhancement layers may have decreasing transmission priorities. That is, the transmission priority may decrease from the base layer to the first enhancement layer, from the first enhancement layer to the second enhancement layer, and so on. The amount of error protection for transmission of the data of the multiple layers may be controlled according to the transmission priority of each. This may ensure that at least some lower layers are transmitted reliably while reducing the overall required bandwidth by not applying excessive error protection to the higher layers.

[0014] In embodiments, the method may further include generating, for each layer of the plurality of layers, a transport layer packet including data for the respective layer. For example, for each time interval (e.g., frame), a respective transport layer packet may be generated for each layer of the plurality of layers.

[0015] In embodiments, the compressed sound representation may further include additional basic side information for decoding the basic compressed sound representation into a basic reconstructed sound representation. The additional basic side information may include information specifying the decoding of one or more of the plurality of components depending on the other components. The method may further include decomposing the additional basic side information into multiple portions of additional basic side information. The method may further include adding the portions of the additional basic side information to a base layer (e.g., including or allocating the portions of the additional basic side information to a base layer for transmission or storage). Each portion of the additional basic side information may correspond to a respective layer and may include information specifying the decoding of one or more components assigned to the respective layer depending on (only) the respective other components assigned to the respective layer and any layers below the respective layer. That is, each portion of the additional basic side information specifies the component in the respective layer to which that portion of the additional basic side information corresponds, without reference to any other components assigned to layers above the respective layer.

[0016] So configured, the proposed method avoids fragmentation of the additional basic side information by adding all parts to the base layer. In other words, all parts of the additional basic side information are included in the base layer. The decomposition of the additional basic side information ensures that for each layer, a part of the additional basic side information is available that does not require knowledge of the content of higher layers. Thus, it is sufficient for a decoder to decode the additional basic side information contained in layers up to the highest available layer, regardless of the actual highest available layer.

[0017] In embodiments, the additional basic side information may include information specifying the decoding (e.g., decompression) of one or more of the components depending on other components. For example, the additional basic side information may represent side information related to an individual mono signal depending on other mono signals. Thus, the additional basic side information may be referred to as dependent basic side information.

[0018] In embodiments, the compressed sound representation may be processed for a series of time intervals, for example time intervals of equal size. The series of time intervals may be frames. Thus, the method may operate on a frame basis, i.e., the compressed sound representation may be encoded frame by frame. A compressed sound representation may be available for each successive time interval (e.g., for each time frame), i.e., the compression operation by which said compressed sound representation was obtained may operate on a frame basis.

[0019] In some embodiments, the method may further include generating configuration information for each layer that indicates the components of the underlying compressed sound representation that are assigned to that layer. In this way, a decoder can easily access the information needed for decoding without unnecessary parsing through the received data payload.

[0020] According to another aspect, a method for layered encoding of a compressed sound representation of a sound or sound field is described. The compressed sound representation may include a basic compressed sound representation including a plurality of components. The plurality of components may be complementary components. The compressed sound representation may further include basic side information (e.g., independent basic side information) and third information (e.g., dependent basic side information) for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The basic side information may include information specifying decoding of one or more components of the plurality of components individually and independently of other components. The additional basic side information may include information specifying decoding of one or more components of the plurality of components dependent on each other component. The method may include subdividing (e.g., grouping) the plurality of components into a plurality of component groups. The method may further include assigning (e.g., adding) each group of the plurality of groups to a respective one of a plurality of hierarchical layers. The assignment may indicate a correspondence between the respective groups and the layer. The components assigned to a respective layer may be said to be included in that layer. The number of groups may correspond to (e.g., be equal to) the number of layers. The multiple layers may include a base layer and one or more hierarchical enhancement layers. The method may further include adding the basic side information to the base layer (e.g., including the basic side information in the base layer or allocating the basic side information to the base layer, e.g., for transmission or storage purposes). The method may further include decomposing the additional basic side information into multiple portions of the additional basic side information and adding those portions of the additional basic side information to the base layer (e.g., including those portions of the additional basic side information in the base layer or allocating those portions of the additional basic side information to the base layer, for transmission or storage).Each portion of the additional basic side information may correspond to a respective layer and may contain information specifying the decoding of one or more components assigned to a respective layer in dependence upon each other component assigned to that respective layer and any layers below that respective layer.

[0021] Implemented in this way, the proposed method ensures that for each layer, adequate additional basic side information is available to decode components contained in any layer up to that layer, without requiring effective reception or decoding (or generally knowledge) of higher layers. In the case of compressed HOA representations, the proposed method ensures that suitable V vectors are available for all components belonging to layers up to the highest available layer in vector coding mode. In particular, the proposed method eliminates cases where elements of the V vector corresponding to components in higher layers are not explicitly signaled. Thus, information contained in layers up to the highest available layer is sufficient to decode (e.g., decompress) any components belonging to layers up to the highest available layer. This ensures proper decompression of the respective reconstructed HOA representations for lower layers, even if the higher layers have not been effectively received by the decoder. On the other hand, the proposed method allows to fully take advantage of the reduction in required bandwidth that can be achieved when applying layer-based coding.

[0022] Embodiments of this aspect may be related to embodiments of the above aspect.

[0023] According to another aspect, a method for decoding a compressed sound representation of a sound or sound field is described. The compressed sound representation may be encoded in multiple hierarchical layers. The multiple hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The multiple layers may be assigned components of the basic compressed sound representation of the sound or sound field. In other words, the multiple layers may contain basic compressed side information components. These components may be assigned to each layer in respective component groups. The multiple components may be complementary components. The base layer may contain basic side information for decoding the basic compressed sound representation. Each layer may contain a portion of enhancement side information including parameters for improving the basic reconstructed sound representation obtained from data contained in the respective layer and any layers below the respective layer. The method may include receiving data payloads corresponding to the multiple hierarchical layers. The method may further include determining a first layer index indicating the highest available layer among the plurality of layers to be used for decoding the basic compressed sound representation into the basic reconstructed sound representation of the sound or sound field. The method may further include deriving the basic reconstructed sound representation using the basic side information from components assigned to the highest available layer and any layers lower than the highest available layer. The method may further include determining a second layer index indicating which portion of the enhancement side information should be used to improve (e.g., enhance) the basic reconstructed sound representation. The method may further include deriving a reconstructed sound representation of the sound or sound field from the basic reconstructed sound representation with reference to the second layer index.

[0024] So configured, the proposed method makes maximum use of the available (e.g. effectively received) information to ensure that the reconstructed sound representation has optimal quality.

[0025] In embodiments, the components of the elementary compressed sound representation may correspond to a mono signal (e.g., a mono transport signal), which may represent either a dominant sound signal or a coefficient sequence of an HOA representation, and which may be quantized.

[0026] In embodiments, the basic side information may include information that specifies the decoding (e.g., decompression) of one or more components of the plurality of components individually, independently of the other components. For example, the basic side information may represent side information related to an individual mono signal, independent of other mono signals. Hence, the basic side information may sometimes be referred to as independent basic side information.

[0027] In embodiments, the enhancement side information may represent enhancement side information, which may include prediction parameters for the basic compressed sound representation for improving (e.g. enhancing) the basic compressed sound representation and the basic reconstructed sound representation obtained from the basic side information.

[0028] In embodiments, the method may further include, for each tier, determining whether the respective tier was validly received. The method may further include determining the first tier index as the tier index of the tier immediately below the lowest tier that was not validly received.

[0029] In various embodiments, determining the second layer index may involve determining the second layer index to be equal to the first layer index, or determining as the second layer index an index value indicating that no enhancement side information is used when obtaining the reconstructed sound representation, in which case the reconstructed sound representation may be equal to the basic reconstructed sound representation.

[0030] In embodiments, the data payload may be received and processed for a series of time intervals, for example, for equal-sized time intervals. The series of time intervals may be frames. In this manner, the method may operate on a frame basis. The method may further determine the second layer index to be equal to the first layer index if the compressed sound representations for the series of time intervals can be decoded independently of each other.

[0031] In various embodiments, the data payload may be received and processed for a series of time intervals, e.g., equal-sized time intervals. The series of time intervals may be frames. In this manner, the method may operate on a frame basis. The method may further include, for a given time interval of the series of time intervals, determining for each layer whether the respective layer has been validly received if the compressed sound representations for the series of time intervals cannot be decoded independently of one another. The method may further include determining the first layer index for the given time interval as the smaller of the first layer index of a time interval preceding the given time interval and the layer index of the layer immediately below the lowest layer that has not been validly received.

[0032] In various embodiments, the method may include, for the given time interval, determining whether a first layer index for the given time interval is equal to a first layer index for a preceding time interval if the compressed sound representations for the successive time intervals cannot be decoded independently of one another. If the first layer index for the given time interval is equal to the first layer index for the preceding time interval, the method may further include determining the second layer index for the given time interval to be equal to the first layer index for the given time interval. If the first layer index for the given time interval is not equal to the first layer index for the preceding time interval, the method may further include determining as the second layer index an index value indicating that no enhancement side information is used when obtaining the reconstructed sound representation.

[0033] In embodiments, the base layer may include at least one portion of additional basic side information corresponding to a respective layer and including information specifying decoding of one or more components of the components assigned to the respective layer depending on other components assigned to the respective layer and any layers lower than the respective layer. The method may further comprise, for each portion of the additional basic side information, decoding said portion of the additional basic side information by referring to the components assigned to its respective layer and any layers lower than the respective layer. The method may further comprise correcting said portion of the additional basic side information by referring to the components assigned to the highest available layer and any layers between the highest available layer and the respective layer. A basic reconstructed sound representation may be obtained from the components assigned to the highest available layer and any layers lower than the highest available layer using the basic side information and corrected portions of the additional basic side information obtained from the portions of additional basic side information corresponding to layers up to the highest available layer.

[0034] In embodiments, the additional basic side information may include information specifying the decoding (e.g. decompression) of one or more of the components in dependence on other components. For example, the additional basic side information may represent side information related to an individual mono signal in dependence on the other mono signals. Thus, the additional basic side information may be referred to as dependent basic side information.

[0035] According to another aspect, a method for decoding a compressed sound representation of a sound or sound field is described. The compressed sound representation may be encoded in multiple hierarchical layers. The multiple hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The multiple layers may be assigned basic compressed sound representation components of the sound or sound field. In other words, the multiple layers may contain basic compressed side information components. These components may be assigned to each layer in respective component groups. The multiple components may be complementary components. The base layer may contain basic side information for decoding the basic compressed sound representation. The base layer may further include at least one portion of additional basic side information corresponding to a respective layer and including information specifying decoding of one or more of the components assigned to the respective layer depending on other components assigned to the respective layer and any layers below the respective layer. The method may further include receiving data payloads corresponding to each of the multiple hierarchical layers. The method may further comprise determining a first layer index indicating the highest usable layer of the plurality of layers to be used for decoding the basic compressed sound representation into the basic reconstructed sound representation of the sound or sound field. The method may further comprise decoding, for each portion of additional basic side information, said portion by referring to components assigned to its respective layer and any layers lower than said respective layer. The method may further comprise correcting, for each portion of additional basic side information, said portion by referring to components assigned to the highest usable layer and any layers between said highest usable layer and said respective layer. A basic reconstructed sound representation may be obtained from components assigned to the highest usable layer and any layers lower than said highest usable layer using the basic side information and corrected portions of additional basic side information obtained from portions of additional basic side information corresponding to layers up to the highest usable layer.The method may further include determining a second layer index that is equal to the first layer index or indicates omission of the enhancement side information during decoding.

[0036] So configured, the proposed method ensures that the additional basic side information that is ultimately used to decode the basic compressed sound representation does not contain redundant elements, thereby making the actual decoding of the basic compressed sound representation more efficient.

[0037] Embodiments of this aspect may be related to embodiments of the above aspect.

[0038] According to another aspect, an encoder for encoding a layered structure of a compressed sound representation of a sound or sound field is described. The compressed sound representation may include a basic compressed sound representation including a plurality of components, which may be complementary components. The compressed sound representation may further include basic side information for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The compressed sound representation may further include enhancement side information including parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The encoder may include a processor configured to perform some or all of the method steps of the methods according to the first-mentioned aspect and the second-mentioned aspect.

[0039] According to another aspect, a decoder for decoding a compressed sound representation of a sound or sound field is described. The compressed sound representation may be encoded in multiple hierarchical layers. The multiple hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The multiple layers may be assigned components of the basic compressed sound representation of the sound or sound field. In other words, the multiple layers may contain basic compressed side information components. These components may be assigned to each layer in respective component groups. The multiple components may be complementary components. The base layer may contain basic side information for decoding the basic compressed sound representation. Each layer may contain a portion of enhancement side information including parameters for improving (e.g., enhancing) the basic reconstructed sound representation obtained from data contained in the respective layer and any layers below the respective layer. The decoder may include a processor configured to perform some or all of the method steps of the methods according to the third and fourth aspects.

[0040] According to another aspect, methods, devices, and systems are directed to decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field. The device may, or the method may perform, a receiver configured to receive a bitstream including the compressed HOA representation corresponding to a plurality of hierarchical layers, including a base layer and one or more hierarchical enhancement layers. The layers are assigned components of the basic compressed sound representation of the sound or sound field, with the components being assigned to each layer in respective component groups. The device may, or the method may perform, a decoder configured to decode the compressed HOA representation based on basic side information associated with the base layer and based on enhancement side information associated with the one or more hierarchical enhancement layers. The basic side information may include basic independent side information related to first individual mono signals that are decoded independently of other mono signals. Each of the one or more hierarchical enhancement layers may include a portion of the enhancement side information including parameters for improving a basic reconstructed sound representation obtained from data included in the respective layer and any layers below the respective layer.

[0041] The basic independent side information may indicate that a first individual mono signal represents a directional signal having a certain direction of incidence. The basic side information may further include basic dependent side information related to a second individual mono signal that is decoded dependently on another mono signal. The basic dependent side information may include vector-based signals that are directionally distributed within the sound field, where the directional distribution is specified by a vector. The components of the vector are set to zero and are not part of the compressed vector representation.

[0042] The components of the basic compressed sound representation may correspond to a mono signal representing either a dominant sound signal or a coefficient sequence of an HOA representation. The bitstream includes data payloads corresponding to the plurality of hierarchical layers. The enhancement side information may include parameters related to at least one of: spatial prediction, subband directional signal synthesis, and parametric ambience replication. The enhancement side information may include information allowing prediction of missing parts of the sound or sound field from the directional signal. Furthermore, for each layer, it may be determined whether the respective layer was validly received, and the layer index of the layer immediately below the lowest layer that was not validly received may be determined.

[0043] According to another aspect, a software program is described that is adapted for execution on a processor and may be adapted to perform some or all of the method steps outlined herein when executed on a computing device.

[0044] According to yet another aspect, a storage medium is described that may include a software program adapted for execution on a processor and adapted to perform some or all of the method steps outlined herein when executed on a computing device.

[0045] As one skilled in the art will appreciate, statements made with respect to any of the above aspects or embodiments thereof also apply to other aspects or embodiments thereof, and repeating these statements for every single aspect or embodiment has been omitted for the sake of brevity.

[0046] The methods and apparatus, including the preferred embodiments outlined herein, may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and apparatus outlined herein may be combined in any manner. In particular, features of the claims may be combined in any manner with other features.

[0047] Method steps and apparatus features may be interchanged in many ways. In particular, those skilled in the art will appreciate that details of a disclosed method can be implemented as an apparatus adapted to perform some or all of the method steps, and vice versa. [Brief explanation of the drawings]

[0048] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1] 1 is a flowchart illustrating an example method for layer configuration encoding, according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram that schematically illustrates an example of an encoder stage according to an embodiment of the present disclosure. [Figure 3] 1 is a flowchart illustrating an example method for decoding a compressed sound representation of a sound or sound field encoded in multiple hierarchical layers, according to an embodiment of the present disclosure. [Figure 4A] FIG. 2 is a block diagram that schematically illustrates an example of a decoder stage according to an embodiment of the present disclosure. [Figure 4B] FIG. 2 is a block diagram that schematically illustrates an example of a decoder stage according to an embodiment of the present disclosure. [Figure 5] FIG. 2 is a block diagram that schematically illustrates an example of a hardware implementation of an encoder according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a block diagram that schematically illustrates an example of a hardware implementation of a decoder according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0049] First, we describe a compressed sound (or sound field) representation (hereinafter referred to as compressed sound representation for brevity) to which the methods and encoders / decoders according to the present disclosure are applicable. In general, a complete compressed sound (or sound field) representation (hereinafter referred to as complete compressed sound representation for brevity) may include (e.g. consist of) three components: a basic compressed sound (or sound field) representation (hereinafter referred to as basic compressed sound representation for brevity), basic side information and enhanced side information.

[0050] The basic compressed sound representation itself may contain (e.g., consist of) several components (e.g., complementary components). The basic compressed sound representation may constitute the largest proportion by far of the complete compressed sound representation. The basic compressed sound representation may consist of a mono transport signal representing the dominant sound signal or the coefficient sequence of the original HOA representation.

[0051] The basic side information is required to decode the basic compressed sound representation and may be assumed to be of a much smaller size than the basic compressed sound representation. It may furthermore consist of a majority of separate parts, each specifying the decompression of only one particular component of the basic compressed sound representation. The basic side information may consist of a first part, which may be known as independent basic side information, and a second part, which may be known as additional basic side information.

[0052] Both the first and second parts, i.e., the independent basic side information and the additional basic side information, may specify the decompression of particular components of the basic compressed sound representation. The second part is optional and may be omitted. In this case, the compressed sound representation may be said to include the first part (e.g., the basic side information).

[0053] The first part (e.g., basic side information) may include side information describing each (complementary) component of the basic compressed sound representation independently of the other (complementary) components. In particular, the first part (e.g., basic side information) may specify decoding of one or more components of said plurality of components individually and independently of the other components. Thus, the first part may be referred to as independent basic side information.

[0054] The second (optional) part, also known as additional basic side information, may describe each (complementary) component of the basic compressed sound representation as dependent on other (complementary) components. This second part may be called dependent basic side information. In particular, the dependency may have the following properties: · The dependent fundamental side information for each individual (complementary) component of the fundamental compressed sound representation achieves maximum range when the fundamental compressed sound representation does not contain certain other (complementary) components. If certain additional (complementary) components are added to the basic compressed sound representation, the dependent basic side information for each (complementary) component under consideration can become a subset of the original dependent basic side information, thereby reducing its size.

[0055] The enhancement side information is also optional. It can be used to improve or enhance (e.g., parametrically improve or enhance) the basic compressed sound representation. Its size is also expected to be much smaller than the size of the basic compressed sound representation.

[0056] Thus, in embodiments, the compressed sound representation may comprise a basic compressed sound representation comprising a plurality of components, basic side information for decoding (e.g., decompressing) the basic compressed sound representation into a basic reconstructed sound representation of said sound or sound field, and enhancement side information comprising parameters for improving or enhancing (e.g., parametrically improving or enhancing) the basic compressed sound representation. The compressed sound representation may further comprise additional basic side information for decoding (e.g., decompressing) the basic compressed sound representation into said basic reconstructed sound representation, which may include information specifying the decoding of one or more components of said plurality of components in dependence on each other component.

[0057] One example of such a type of complete compressed sound representation is given by the compressed Higher Order Ambisonics (HOA) sound field representation specified by Chapter 12 and Annex C.5 of the preliminary version of the MPEG-H 3D Audio standard (Non-Patent Document 1), i.e. the compressed sound representation may correspond to a compressed HOA sound (or sound field) representation of a sound or sound field.

[0058] For this example, the basic compressed sound field representation (basic compressed sound representation) may include (e.g., be identified with) several components. The components may be (e.g., correspond to) mono signals. The mono signals may be quantized mono signals. The mono signals may represent either a dominant sound signal or a coefficient sequence of an ambient sound HOA sound field component.

[0059] The basic side information may, among other things, describe for each of these mono signals how it spatially contributes to the sound field. For example, the basic side information may specify the dominant sound signal as a purely directional signal, i.e., a general plane wave with a certain direction of incidence. Alternatively, the basic side information may specify the mono signal as a coefficient sequence of the original HOA representation with a certain index. The basic side information may be further separated into a first part and a second part as described above.

[0060] The first part is side information (e.g., independent fundamental side information) related to a particular individual mono signal. This independent fundamental side information is independent of the presence of other mono signals. Such side information may, for example, specify a mono signal representing a directional signal (e.g., representing a general plane wave) with a certain direction of incidence. Alternatively, the mono signal may be specified as a coefficient sequence of the original HOA representation with a certain index. The first part may also be referred to as independent fundamental side information. In general, the first part (e.g., fundamental side information) may specify decoding of one or more mono signals of the plurality of mono signals individually and independently of the other mono signals.

[0061] The second part is side information related to a particular individual mono signal (e.g., additional elementary side information). This side information depends on the presence of other mono signals. Such side signals may be used, for example, when the mono signals are specified as vector-based signals (see, for example, Section 12.4.2.4.4 of [Illegible Text]). These signals are directionally distributed in the sound field, and the directional distribution can be specified by a vector. In certain modes (i.e., CodedVVecLength=1), certain components of this vector are implicitly set to 0 and are not part of the compressed vector representation. These components are the components with indices equal to the coefficient sequences of the original HOA representation that are part of the elementary compressed sound representation. That is, the total number of individual components of a vector to be coded depends on the elementary compressed sound representation. In particular, this total number depends on which coefficient sequences the original HOA representation contains.

[0062] If the coefficient sequence of the original HOA representation is not included in the basic compressed sound representation, the dependent basic side information for each vector-based signal consists of all vector components and has its maximum size. If the coefficient sequence of the original HOA representation with certain indices is added to the basic compressed sound representation, the vector components with those indices are removed from the side information for each vector-based signal, thereby reducing the size of the dependent basic side information for the vector-based signal.

[0063] The enhanced side information (e.g., enhanced side information) may include parameters related to (broadband) spatial prediction (see section 12.4.2.4.3 of non-patent document 1) and / or parameters related to subband directional signal synthesis and parametric ambient sound replication.

[0064] Parameters related to (broadband) spatial prediction can be used to (linearly) predict missing parts of the sound field from the directional signal.

[0065] Subband directional signal synthesis and parametric ambient sound replication are compression tools recently introduced in the MPEG-H 3D Audio standard through an amendment (see Section 1 of Non-Patent Document 2). These tools allow frequency-dependent parametric prediction of an additional mono signal to be spatially distributed to complement a spatially incomplete or missing compressed HOA representation. The prediction may be based on the coefficient sequence of the basic compressed sound representation.

[0066] It is important to note that these complementary contributions to the sound field are represented in the compressed HOA representation not by additional quantized signals, but by additional side information of a comparably much smaller size. Therefore, the two coding tools mentioned above are particularly suitable for the compression of HOA representations at low data rates.

[0067] A second example of a compressed representation of one or more mono signals having the structure described above may include: encoded spectral information for distinct frequency bands up to a certain upper frequency limit, which can be considered as a basic compressed representation; basic side information specifying the encoded spectral information (e.g., by the number and width of the encoded frequency bands); and enhanced side information including (e.g., consisting of) parameters for spectral band replication (SBR). The parameters of the enhanced side information describe how to parametrically reconstruct, from the basic compressed representation, spectral information for higher frequency bands not considered in the basic compressed representation.

[0068] The present disclosure proposes a method for layered coding of a complete compressed sound (or sound field) representation with the above-mentioned structure.

[0069] Compression may be frame-based, in the sense that it provides a compressed representation of a series of time intervals (e.g., in the form of data packets, or equivalently, frame payloads). The time intervals may have equal or different sizes. These data packets may be assumed to contain, in addition to the actual compressed representation data, a validity flag and a value indicating their size. In what follows, without any intention of limitation, it is assumed that compression is frame-based. Furthermore, unless otherwise stated, we will focus, without any intention of limitation, on the treatment of a single frame; therefore, the frame index is omitted.

[0070] Each frame payload of the considered complete compressed sound (or sound field) representation contains J data packets (or frame payloads), each of which is a BSRC j , j=1,…,J. Furthermore, each data packet is assumed to be for one component of the basic compressed sound representation. I It is assumed that the packet contains independent basic side information marked by BSI. I are the specific components of the basic compressed sound representation BSRC independently of other components. j Optionally, each data packet may further specify a BSI D It is assumed that the packet contains a dependent basic side information (additional basic side information) denoted BSI. D The specific components of the basic compressed sound representation BSRC depend on other components. j Specify.

[0071] Two data packets BSI I and BSI D The information contained within may optionally be grouped into single data packets of basic side information (BSI), each of which contains, among other things, one specific component of the basic compressed sound representation (BSRC). jEach of these portions may be said to include a portion of independent side information and optionally a portion of dependent side information.

[0072] Finally, each data packet may contain an enhancement side information payload, denoted ESI, with a description of how to improve or enhance the reconstructed sound (or sound field) from the complete underlying compressed sound representation.

[0073] The proposed solution for layered coding addresses the required steps to enable both the compressor, which involves packing data packets for transmission, and the receiver and decompressor, each of which is described in detail below.

[0074] First, we discuss compression and packing (e.g., for transmission), particularly the components and elements of a complete compressed sound (or sound field) representation in the case of layered coding.

[0075] Figure 1 shows a schematic flow chart of an example method for compression and packing (e.g., a method for encoding or layered encoding of a compressed sound representation of a sound or sound field). The allocation (e.g., allocation) of individual payloads to the base layer and (M-1) enhancement layers may be achieved by a transport layer packer. Figure 2 shows a schematic block diagram of an example of the allocation / allocation of individual payloads.

[0076] As indicated above, the complete compressed sound representation 2100 may relate to a compressed HOA representation, for example, including a basic compressed sound representation. The complete compressed sound representation 2100 may include multiple components (e.g., mono signals) 2110-1, ..., 2110-J, independent basic side information (basic side information) 2120, optional enhancement side information (enhancement side information) 2140, and optional dependent basic side information (additional basic side information) 2130. The basic side information 2120 may be information for decoding the basic compressed sound representation into a basic reconstructed sound representation of said sound or sound field. The basic side information 2120 may include information specifying the decoding of one or more components (e.g., mono signals) individually and independently of the other components. The enhancement side information 2140 may include parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The additional basic side information 2130 may be (further) information for decoding the basic compressed sound representation into said basic reconstructed sound representation, and may include information specifying the decoding of one or more components of said plurality of components individually and dependent on each other component.

[0077] 2 illustrates the underlying assumption that there are multiple hierarchical layers, including a base layer and one or more (hierarchical) enhancement layers. For example, there may be a total of M layers, i.e., one base layer and M-1 enhancement layers. The multiple hierarchical layers have increasing layer indices, with the lowest layer index (e.g., layer index 1) corresponding to the base layer. It is further understood that the layers are ordered from the base layer, through the enhancement layers, to the overall highest enhancement layer (i.e., the overall top layer).

[0078] The proposed method may be performed on a frame basis (i.e., in a frame-by-frame manner). In particular, the compressed sound representation 2100 may be compressed for a series of time intervals, for example time intervals of equal size. Each time interval may correspond to a frame. The following steps may be performed for each of the series of time intervals (e.g., frames):

[0079] In S1010 of FIG. 1 , the plurality of components 2110 are subdivided into a plurality of component groups. Each of the plurality of groups is then assigned (e.g., added or allocated) to a corresponding one of a plurality of hierarchical layers, where the number of groups corresponds to the number of layers. For example, the number of groups may be equal to the number of layers, such that there may be one group of components for each layer. As indicated above, the plurality of layers may include a base layer and one or more (e.g., M−1) hierarchical enhancement layers.

[0080] In other words, the basic compressed sound representation is subdivided into parts that are assigned to the individual layers. Without loss of generality, the grouping is made up of M+1 numbers J m , m=0,…,M, where J0=1, J M = J + 1, and the component BSRC j is J m-1 ≦j <J m is assigned to the m-th layer.

[0081] At S1020, groups of components are assigned to respective layers. At S1030, base side information 2120 is added (eg, allocated) to the base layer (ie, the lowest layer of the plurality of hierarchical layers).

[0082] That is, due to its small size, it is proposed to include the complete basic side information (basic side information and optional additional basic side information) in the base layer to avoid its unnecessary fragmentation.

[0083] If the compressed sound representation under consideration contains dependent basic side information (additional basic side information), the method may further comprise (not shown in Fig. 1) decomposing said additional basic side information into a plurality of portions 2130-1, ..., 2130-M of additional basic side information. The portions of the additional basic side information may then be added to (e.g. allocated to) the base layer. In other words, the portions of the additional basic side information may be included in the base layer. Each portion of the additional basic side information may correspond to a respective layer and may contain information specifying the decoding of one or more components assigned to the respective layer depending on other components assigned to the respective layer and to any layers lower than the respective layer.

[0084] Thus, the independent basic side information BSI I While the (basic side information) 2120 is left unchanged for allocation, the dependent basic side information needs to be treated specially for layered coding to allow correct decoding at the receiver and to reduce the size of the transmitted dependent basic side information. D,m It is proposed to decompose the audio signal into M parts, denoted by m=1,...,M, where the mth part is the component of the basic compressed sound representation BSRC assigned to the mth layer. j , J m-1 ≦j <J m This assumes that the optional dependent basic side information is present for the compressed sound representation under consideration. If the respective dependent side information is not present, then for that compressed sound representation of the parts, the BSI D,m is assumed to be empty. Each part of the dependent basic side information BSI D,m is the BSRC of all components included in all layers up to the m-th layer (i.e., included in all layers j=1,…,m). j , 1 ≤ j <J m may depend on

[0085] Independent Basic Side Information Packet BSII If is of negligibly small size, it is reasonable to keep it as a whole and add (allocate) it to the base layer. Optionally, the same decomposition can be performed on the independent basic side information as on the dependent basic side information, and the packet BSI I,m , m=1,...,M. This is useful to reduce the size of the base layer by adding (assigning) parts of the independent base side information to the layer with the corresponding components of the base compressed sound representation.

[0086] In S1040, multiple portions of enhancement side information 2140-1, ..., 2140-M may be determined, each portion of the enhancement side information may include parameters for improving (e.g., enhancing) the reconstructed sound representation obtained from the data contained in the respective layer and any layers below the respective layer.

[0087] The reason for this step is that in the case of layered coding, the enhancement side information needs to be calculated extra for each layer, since it is intended to enhance the preliminary decompressed sound (or sound field), but it depends on the layers available for decompression. Specifically, the preliminary decompressed sound (or sound field) for a given highest decodable layer (highest available layer) depends on the components contained in that highest decodable layer and any layers below that highest decodable layer. Thus, compression is performed using ESI m , m=1,...,M, where the enhancement side information ESI m is calculated to enhance the sound (or sound field) representation obtained from all the data contained in the base layer and in the enhancement layers with indices lower than m (e.g., all the data contained in the mth layer and any layers below the mth layer).

[0088] At S1050, the multiple portions 2140-1, ..., 2140-M of enhancement side information are assigned (e.g., added or allocated) to the multiple layers, with each portion of the multiple portions of enhancement side information assigned to a respective layer of the multiple layers, e.g., each layer of the multiple layers includes a respective portion of enhancement side information.

[0089] The allocation of the base and / or enhancement side information to the respective layers may be indicated in the configuration information generated by the encoding method. In other words, the correspondence between the base and / or enhancement side information and the respective layers may be indicated in the configuration information. Furthermore, the configuration information may indicate, for each layer, the components of the base compressed sound representation that are allocated to (e.g., included in) that layer. Portions of the additional base side information may be included in the base layer but correspond to different layers than the base layer.

[0090] In summary, the compression stage provides a frame data packet, denoted FRAME, with the following composition:

[0091]

number

[0092]

number

[0093] The individual data packets may then be grouped within a payload, which is defined as a special data packet that contains, in addition to the actual compressed representation data, a validity flag and a value indicating its size. The use of payloads allows simple demultiplexing at the receiver, and has the advantage that outdated payloads can be discarded without having to parse through them. One possible grouping is given by: ·Each BSRC j Packets, j=1,...,J, are individual payloads (BPs with j To assign (e.g., allocate) to a m-th Enhancement Side Information data packet ESI m and the m-th dependent side information data packet BSI D,m One improved payload (with  ̄ EP m Assign (e.g., allocate) to m=1,…,M, denoted by ·Independent Basic Side Information BSI I into a separate side information payload (denoted by BSIP with a ´).

[0094] Optionally, if the size of the independent basic side information is large, each m-th BSI of its components I,m , m=1,...,M is the enhanced payload (EP m In this case, the side information payload (BSIP with a ) is empty and can be ignored.

[0095] Another option is to include all subordinate Basic Side Information data packets (BSI) D,m This is reasonable when the size of the dependent basic side information is small.

[0096] Finally, a frame data packet, denoted FRAME, may be given, with the following composition:

[0097]

number

[0098] The method may further include (not shown in FIG. 1 ) generating, for each of the plurality of layers, a transport layer packet (e.g., base layer packet 2200 and M−1 enhancement layer packets 2300-1, ..., 2300-(M−1)) containing data for the respective layer (e.g., components, base side information, and enhancement side information for the base layer, or components and enhancement side information for the one or more enhancement layers).

[0099] Transport layer packets for different layers may have different transmission priorities. Thus, the method may further include (not shown in FIG. 1 ) generating a transport stream for transmission of data of the multiple layers, where the base layer has the highest transmission priority and successive enhancement layers have decreasing transmission priorities. Here, a higher transmission priority corresponds to a greater degree of error protection, and vice versa.

[0100] It will be understood that the steps described above may be performed in any order, and the exemplary order shown in FIG. 1 is not limiting, unless a step requires another step as a prerequisite.

[0101] Figure 3 shows a method for decoding a compressed sound representation of a sound or sound field for decoding or decompression (unpacking). Examples of corresponding receiver and decompression stages are shown schematically in the block diagrams of Figures 4A and 4B.

[0102] As can be seen from the above, the compressed sound representation may be encoded in the plurality of hierarchical layers. The plurality of layers may be assigned (e.g., may contain) components of the basic compressed sound representation, the components being assigned to each layer in respective component groups. The base layer may contain basic side information for decoding the basic compressed sound representation. Each layer may contain one of the above-mentioned pieces of enhancement side information containing parameters for improving the basic reconstructed sound representation obtained from data contained in the respective layer and any layers below the respective layer.

[0103] The proposed method may be performed on a frame basis (i.e., in a frame-by-frame manner). In particular, the reconstructed representation of the sound or sound field may be generated for a series of time intervals, for example time intervals of equal size. The time intervals may for example be frames. The following steps may be performed for each of the series of time intervals (e.g., frames):

[0104] At S3010, a data payload (e.g., transport layer packets) corresponding to the plurality of layers is received. The data payload may be received as part of a bitstream containing a compressed HOA representation of a sound or sound field corresponding to the plurality of hierarchical layers. The hierarchical layers may include a base layer and one or more enhancement layers. The plurality of layers may be assigned components of a basic compressed sound representation of the sound or sound field. The components are assigned to each layer in respective component groups.

[0105] Individual layer packets may be multiplexed to provide received frame packets of the complete compressed audio representation.

[0106]

number

[0107]

number

[0108] With payload, the received frame packet is

[0109]

number

[0110] The received frame packet may then be passed to the decompressor or decoder 4100. If the transmission of the individual layers was error-free, at least the enhancement side information payload (e.g., corresponding to a portion of the enhancement side information) contained therein may be

[0111]

number

[0112] In the decompressor 4100, the received frame packets may be demultiplexed. For this purpose, information about the size of each payload may be utilized to avoid unnecessary parsing through the data of the individual payloads.

[0113] In S3020, a first layer index indicating the highest layer (e.g., the highest available layer or the highest decodable layer) from the plurality of layers is determined to be used for decoding the basic compressed sound representation into the basic reconstructed sound representation of the sound or sound field.

[0114] Further, in S3020, a value (e.g., layer index) N of the highest layer (highest available layer) to be used for decompression of the basic sound representation is determined. B may be selected as the best one actually used for decompressing the basic sound representation. improvement Layer N B −1. Since each layer contains exactly one enhancement side information payload (part of the enhancement side information), it can be determined based on the enhancement side information payload whether the containing layer is valid (e.g., has been received validly). Thus, the selection is made by selecting all the enhancement side information payloads ESI m , m=1,…,M (or correspondingly

[0115]

number

[0116] In S3030, a basic reconstructed sound representation is obtained, which may be obtained using basic side information (or generally using basic side information) from components assigned to the highest available layer indicated by the first layer index and any layers lower than this highest available layer.

[0117] Basic compressed sound representation components BSRC1, ..., BSRC J The payload of the Basic Side Information payload (e.g., BSI or BSI I and BSI D,m , m=1,…,M) (all of them) and value N B, may be provided to the base representation decompression processing unit 4200. The base representation decompression processing unit 4200 (shown in FIGS. 4A and 4B) may be provided to the base representation decompression processing unit 4200 together with the lowest N B layers, namely the base layer and N B - Reconstruct the basic sound (or sound field) representation using only the basic compressed sound representation components contained in one enhancement layer (i.e., the layers up to the layer indicated by the first layer index). Alternatively, B Only the payloads of the basic compressed sound representation components contained in this layer may be provided to the basic representation decompression processing unit 4200 together with the respective basic side information payloads.

[0118] The required information about which components of the basic compressed sound (or sound field) representation are contained in each layer is assumed to be known to the decompressor 4100 from data packets with configuration information, which is assumed to be sent or received before the frame data packets.

[0119] Subordinate Side Information data packet BSI D,m , m=1,…,N B and Enhanced Side Information data packets ESI NE To provide all the boost payloads, the value N E and value N B , and may be input to the partial parser 4400 (see FIG. 4B) of the decompressor 4100. The parser may discard all payload and data packets that are not used in the actual decompression. E If the value of is equal to 0, all Enhancement Side Information data packets may be assumed to be empty.

[0120] If the base layer contains at least one subordinate basic side information payload (part of the additional basic side information) corresponding to each layer, each individual subordinate basic side information payload (e.g., BSI D,m , m=1,…,N BThe decoding of (a portion of the additional basic side information) may comprise (i) decoding said portion of the additional basic side information by referring to components assigned to the respective layer and to any layers below said respective layer (preliminary decoding), and (ii) correcting said portion of the additional basic side information by referring to components assigned to the highest available layer and to any layers between said highest available layer and said respective layer (correction), where the additional basic side information corresponding to each layer includes information specifying the decoding of one or more of the components assigned to said respective layer depending on other components assigned to said respective layer and to any layers below said respective layer.

[0121] A basic reconstructed sound representation can then be obtained (e.g., generated) from components assigned to the highest available layer and any layers lower than the highest available layer using basic side information and corrected portions of additional basic side information obtained from portions of additional basic side information corresponding to layers up to the highest available layer.

[0122] In particular, each payload BSI D,m , m=1,…,N B The preliminary decoding of is performed by the first J m - 1 basic compressed sound representation component BSRC1, ..., BSRC (Jm)-1 It may involve leveraging dependencies on

[0123] Each payload BSI D,m , m=1,…,N B The successive correction of the first N components is that the fundamental tonal components are more numerous than assumed for preliminary decoding. B > the first J in layer m NB - 1 basic compressed sound representation component BSRC1, ..., BSRC (JNB)-1This may involve taking into account that the final reconstruction from the original is achieved by discarding outdated information. This is possible due to an initially assumed property of the dependent basic side information, namely that if certain complementary components are added to the basic compressed sound representation, the dependent basic side information for each individual (complementary) component will be a subset of the original.

[0124] In S3040, a second layer index may be determined, which may indicate a portion or portions of the enhancement side information to be used to improve (e.g., enhance) the basic reconstructed sound representation.

[0125] In addition to the first layer index, an index (second layer index) N of the enhancement side information payload (part of the second enhancement information) to be used for decompression E may be determined. A second layer index N E is always the first layer index N B or may be equal to 0. Enhancement may always be achieved according to the basic sound representation obtained from the highest available layer, or may not be achieved at all.

[0126] In S3050, a reconstructed sound representation of said sound or sound field is obtained (eg, generated) from said basic reconstructed sound representation with reference to said second layer index.

[0127] That is, the reconstructed sound representation is obtained by (parametrically) improving or enhancing the basic reconstructed sound representation, for example by using the enhancement side information (part of the enhancement side information) indicated by the second layer index. As will be explained later, the second layer index may also indicate that no enhancement side information is used at this stage. The reconstructed sound representation will then correspond to the basic reconstructed sound representation.

[0128] For this purpose, the reconstructed basic sound representation is composed of all the enhancement side information payloads ESI1, ..., ESI M , Basic Side Information Payload (e.g., BSI or BSI I and BSI D,m , m=1,…,M) and value N E , and the enhanced representation decompression processing unit 4300 (shown in FIGS. 4A and 4B) which decompresses the enhanced representation together with the enhanced side information payload ESI NE Alternatively, the enhancement side information payload ESI may be used instead of all the enhancement side information payloads. NE Only N may be provided to the enhanced representation decompression processing unit 4300. E If the value of is equal to 0, all enhancement side information payloads are discarded (alternatively, no enhancement side information payloads are provided). Then, the reconstructed final enhanced sound representation 2100' is equal to the reconstructed basic sound representation. The enhancement side information payload ESI NE may be obtained by the partial parser 4400.

[0129] FIG. 3 also generally illustrates decoding of a compressed HOA representation based on base side information associated with the base layer and based on enhancement side information associated with one or more hierarchical enhancement layers.

[0130] It will be understood that the steps described above may be performed in any order, and the exemplary order shown in FIG. 3 is not limiting, unless a step requires another step as a prerequisite.

[0131] Next, details of layer selection (selection of first and second layer indexes) for decompression in steps S3020 and S3040 will be described.

[0132] Determining the first tier index may involve, for each layer, determining whether the layer was validly received. Determining the first tier index may further involve determining the first tier index as the tier index of the layer immediately below the lowest layer that was not validly received. Whether a layer was validly received may be determined by evaluating whether the enhancement side information payload for that layer was validly received. This may be done by evaluating a validity flag in the enhancement side information payload.

[0133] Determining the second layer index may generally involve determining the second layer index to be equal to the first layer index, or determining the second layer index to be an index value (e.g., index value 0) that indicates that no enhancement side information is used when obtaining the reconstructed sound representation.

[0134] If all frame data packets can be decompressed independently of each other, the highest layer number (highest available layer) actually used for decompressing the basic sound representation, N B and the index N of the enhanced side information payload used for decompression E may both be set to the highest numbered valid enhancement side information payload, L, which itself may be determined by evaluating the validity flag within the enhancement side information payload. By leveraging knowledge of the size of each enhancement side information payload, complex parsing through the actual data of the payload to determine validity can be avoided.

[0135] That is, if the compressed sound representations for a series of time intervals can be decoded independently, the second layer index may be determined to be equal to the first layer index, in which case the reconstructed basic sound representation may be enhanced based on the enhancement side information payload of the highest available layer.

[0136] If differential decompression is used, where there is inter-frame dependency, then decisions from previous frames must also be taken into account. In differential decompression, independent frame data packets are typically transmitted at regular time intervals to allow decompression to start from those points. For independent frame data packets, the value N B and N E The determination of is now frame independent and is performed as above.

[0137] To elaborate on the proposed frame-dependent decision, let L(k) denote the highest number (e.g., layer index) of the available enhancement side information payload for the kth frame, and let N denote the highest layer number (e.g., layer index) selected and used for decompression of the basic sound representation. B In (k), the number (e.g., layer index) of the enhancement side information payload used for decompression is N E This is expressed as (k).

[0138] Using this notation, the highest layer number N used for decompression of the basic sound representation B (k) is calculated according to the following formula:

[0139]

number

[0140] That is, if the compressed sound representations for a series of time intervals (e.g., frames) cannot be decoded independently of each other, determining the first layer index may include determining, for each layer, whether the respective layer was validly received, and determining the first layer index for the given time interval as the smaller of the first layer index of the time interval preceding the given time interval and the layer index of the layer immediately below the lowest layer that was not validly received.

[0141] Number N of enhanced side information payloads used for decompression E (k) may be determined according to the following formula:

[0142]

number

[0143] That is, specifically, the highest layer number N used for decompressing the basic sound representation. B As long as (k) remains the same, the same corresponding enhancement layer number is selected. However, B If (k) changes, N E Enhancement is disabled by setting (k) to 0. Due to the assumed differential decompression of the enhancement side information, N B The change based on (k) is not possible because it would require decompression of the corresponding enhancement side information layer in the previous frame, which is assumed not to have been performed.

[0144] That is, if compressed sound representations for a series of time intervals (e.g., frames) cannot be decoded independently of each other, determining the second layer index may include determining whether the first layer index for the given time interval is equal to the first layer index for a preceding time interval. If the first layer index for the given time interval is equal to the first layer index for a preceding time interval, the second layer index for the given time interval may be determined (e.g., selected) to be equal to the first layer index for the given time interval. On the other hand, if the first layer index for the given time interval is not equal to the first layer index for a preceding time interval, an index value indicating that no enhancement side information is used when obtaining the reconstructed sound representation may be determined (e.g., selected) as the second layer index.

[0145] Alternatively, in decompression, N E If all of the enhanced side information payloads numbered up to (k) are decompressed in parallel, the selection rule in equation (4) becomes N E (k)=N B (k) (9) is replaced by

[0146] Finally, for differential decompression, the number of the highest used layer, N B Note that can only be increased in independent frame data packets, while a decrease is possible in any frame.

[0147] It will be appreciated that the proposed method for encoding a layered structure of a compressed sound representation can be implemented by an encoder for encoding a layered structure of a compressed sound representation. Such an encoder may comprise respective units adapted to perform the respective steps described above. An example of such an encoder 5000 is shown schematically in Fig. 5. For instance, such an encoder 5000 may comprise a component subdivision unit 5010 adapted to perform S1010 described above, a component allocation unit 5020 adapted to perform S1020 described above, a basic side information allocation unit 5030 adapted to perform S1030 described above, an enhanced side information splitting unit 5040 adapted to perform S1040 described above, and an enhanced side information allocation unit 5050 adapted to perform S1050 described above. It will further be appreciated that each unit of such an encoder may be embodied by a processor 5100 of a computing device adapted to perform the processing performed by each of said units, i.e. adapted to perform some or all of the above steps and / or further steps of the proposed encoding method. The encoder or computing device may further include a memory 5200 accessible by the processor 5100 .

[0148] It will further be understood that the proposed method for decoding a compressed sound representation encoded in multiple hierarchical layers can be implemented by a decoder for decoding a compressed sound representation encoded in multiple hierarchical layers. Such a decoder may comprise respective units adapted to perform the respective steps described above. An example of such a decoder 6000 is shown schematically in FIG. 6. For example, such a decoder 6000 may comprise a receiving unit 6010 adapted to perform S3010 described above, a first layer index determination unit 6020 adapted to perform S3020 described above, a base reconstruction unit 6030 adapted to perform S3030 described above, a second layer index determination unit 6040 adapted to perform S3040 described above, and an enhancement reconstruction unit 6050 adapted to perform S3050 described above. It will further be understood that each unit of such a decoder may be embodied by a processor 6100 of a computing device adapted to perform the processing performed by each of said units, i.e. adapted to perform some or all of the above steps and / or further steps of the proposed decoding method. The decoder or computing device may further comprise a memory 6200 accessible by the processor 6100 .

[0149] It should be noted that the present description and drawings merely illustrate the principles of the proposed method and apparatus. It is therefore understood that those skilled in the art will be able to devise various configurations that embody the principles of the present invention and are within its spirit and scope, even if not explicitly described or shown herein. Furthermore, all examples described herein are expressly intended solely for educational purposes to aid the reader in understanding the principles of the proposed method and apparatus and the concepts contributed by the inventors to the advancement of the art, and are to be construed without limitation to such specifically described examples and conditions. Furthermore, all statements herein describing principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass equivalents thereof.

[0150] The methods and apparatus described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented, for example, as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or application-specific integrated circuits. Signals emerging from the described methods and apparatus may be stored on media such as random access memory or optical storage media, or transmitted over a network such as a radio, satellite, wireless, or wired network, e.g., the Internet.

[0151] Several aspects will be described. [Aspect 1] 1. A method for decoding a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field, the method comprising: receiving a bitstream containing the compressed HOA representation corresponding to a plurality of hierarchical layers including a base layer and one or more hierarchical enhancement layers, the plurality of layers being assigned components of the basic compressed sound representation of the sound or sound field, the components being assigned to each layer in respective component groups; decoding the compressed HOA representation based on base side information associated with a base layer and based on enhancement side information associated with the one or more hierarchical enhancement layers; each of said one or more hierarchical enhancement layers comprises a portion of said enhancement side information comprising parameters for improving a basic reconstructed sound representation obtained from data contained in said respective layer and any layers below said respective layer, method. [Aspect 2] said basic compressed sound representation components correspond to a monophonic signal; the mono signal represents either a dominant sound signal or a coefficient sequence of an HOA representation; The method of embodiment 1. Aspect 3 3. The method of claim 1 or 2, wherein the bitstream includes data payloads corresponding respectively to the one or more hierarchical layers. Aspect 4 4. The method of any one of aspects 1 to 3, wherein the enhancement side information includes parameters related to at least one of spatial prediction, subband directional signal synthesis, and parametric ambient sound replication. Aspect 5 5. The method of any one of aspects 1 to 4, wherein the enhanced side information includes information that allows prediction of missing parts of the sound or sound field from a directional signal. Aspect 6 for each layer, determining whether the respective layer was validly received; further comprising determining the tier index of the tier immediately below the lowest tier that has not been validly received; 6. The method of any one of embodiments 1 to 5. Aspect 7 7. The method of any one of aspects 1-6, further comprising determining a second tier index that is equal to the first tier index or indicates omission of enhancement side information upon decoding. Aspect 8 determining a first layer index indicating a highest available layer of the plurality of layers to be used for decoding the elementary compressed sound representation into an elementary reconstructed sound representation of the sound or sound field; and obtaining the basic reconstructed sound representation from components assigned to the highest available layer and any layers lower than the highest available layer using the first side information. 8. The method of any one of embodiments 1 to 7. Aspect 9 and wherein the base layer includes at least one portion of additional basic side information corresponding to each layer and including information specifying decoding of one or more of the components assigned to the respective layer depending on other components assigned to the respective layer and any layers lower than the respective layer, the method comprising, for each portion of additional basic side information: decoding said portion of additional basic side information by referencing components assigned to its respective layer and any layers below said respective layer; amending said portion of additional basic side information by reference to components assigned to said highest available layer and to any layers between said highest available layer and said respective layer; said basic reconstructed sound representation is obtained from components assigned to said highest available layer and any layers lower than said highest available layer using said basic side information and corrected portions of additional basic side information obtained from portions of additional basic side information corresponding to layers up to said highest available layer, 9. The method of any one of embodiments 1 to 8. Aspect 10 1. An apparatus for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, the apparatus comprising: a receiver for receiving a bitstream containing the compressed HOA representation corresponding to a plurality of hierarchical layers including a base layer and one or more hierarchical enhancement layers, the plurality of layers being assigned components of the basic compressed sound representation of the sound or sound field, the components being assigned to each layer in respective component groups; a decoder configured to decode the compressed HOA representation based on base side information associated with a base layer and based on enhancement side information associated with the one or more hierarchical enhancement layers, each of said one or more hierarchical enhancement layers comprises a portion of said enhancement side information comprising parameters for improving a basic reconstructed sound representation obtained from data contained in said respective layer and any layers below said respective layer, Device. Aspect 11 said basic compressed sound representation components correspond to a monophonic signal; the mono signal represents either a dominant sound signal or a coefficient sequence of an HOA representation; 11. The device according to embodiment 10. Aspect 12 12. The apparatus of claim 10 or 11, wherein the bitstream includes data payloads corresponding respectively to the one or more hierarchical layers. Aspect 13 13. The apparatus of any one of aspects 10 to 12, wherein the enhancement side information includes parameters related to at least one of spatial prediction, subband directional signal synthesis, and parametric ambient sound replication. Aspect 14 14. The apparatus of any one of aspects 10 to 13, wherein the enhanced side information includes information that allows prediction of missing parts of the sound or sound field from a directional signal. Aspect 15 for each layer, determining whether the respective layer was validly received; further comprising determining the tier index of the tier immediately below the lowest tier that has not been validly received; 15. The device of any one of embodiments 10 to 14. Aspect 16 16. The apparatus of any one of aspects 10-15, further comprising: determining a second tier index that is equal to the first tier index or that indicates omission of enhancement side information upon decoding. Aspect 17 determining a first layer index indicating a highest available layer of the plurality of layers to be used for decoding the elementary compressed sound representation into an elementary reconstructed sound representation of the sound or sound field; and obtaining the basic reconstructed sound representation from components assigned to the highest available layer and any layers lower than the highest available layer using the first side information. 17. The device of any one of embodiments 10 to 16. Aspect 18 and wherein the base layer includes at least one portion of additional basic side information corresponding to each layer and including information specifying decoding of one or more of the components assigned to the respective layer depending on other components assigned to the respective layer and any layers lower than the respective layer, the method comprising, for each portion of additional basic side information: decoding said portion of additional basic side information by referencing components assigned to its respective layer and any layers below said respective layer; amending said portion of additional basic side information by reference to components assigned to said highest available layer and to any layers between said highest available layer and said respective layer; said basic reconstructed sound representation is obtained from components assigned to said highest available layer and any layers lower than said highest available layer using said basic side information and corrected portions of additional basic side information obtained from portions of additional basic side information corresponding to layers up to said highest available layer, 18. The device of any one of embodiments 10 to 17.

Claims

1. 1. A method for decoding a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field, the method comprising: receiving a bitstream; extracting the compressed HOA representation from the bitstream, the bitstream comprising a plurality of hierarchical layers including a base layer and two or more hierarchical enhancement layers, the bitstream further comprising base side information associated with the base layer and enhancement side information associated with the two or more hierarchical enhancement layers, the two or more hierarchical enhancement layers comprising a highest available hierarchical enhancement layer; extracting from the bitstream independent side information indicative of a first individual mono signal representing a directional signal having an incidence direction; decoding the compressed HOA representation based on the independent side information, based on the base side information associated with the base layer, based on a portion of the enhancement side information associated with the highest available hierarchical enhancement layer, and not based on a second portion of the enhancement side information associated with any other layer of the two or more hierarchical enhancement layers. method.

2. A non-transitory carrier medium carrying computer executable code that, when executed on a processor, causes the processor to perform the method of claim 1.

3. 1. An apparatus for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, the apparatus comprising: a receiver for receiving a bitstream including the compressed HOA representation, the bitstream including a plurality of hierarchical layers including a base layer and two or more hierarchical enhancement layers, the bitstream further including base side information associated with the base layer and enhancement side information associated with the two or more hierarchical enhancement layers; a receiver, wherein the two or more tiered enhancement layers include a highest available tiered enhancement layer; a processor for determining from said bitstream independent side information indicative of a first individual mono signal representing a directional signal having a direction of incidence; a decoder for decoding the compressed HOA representation based on independent side information and the base side information associated with the base layer, and based on a portion of the enhancement side information associated with the highest available hierarchical enhancement layer, and not based on a second portion of the enhancement side information associated with any other layer of the two or more hierarchical enhancement layers. Device.

Citation Information

Patent Citations

  • Method for compressing a Higher Order Ambisonics (HOA) signal, method for decompressing a compressed HOA signal, apparatus for compressing a HOA signal, and apparatus for decompressing a compressed HOA signal

    EP2922057A1