Coding of layer structures for compressed sound or sound field representations

The method addresses the challenge of adapting layered audio coding to time-varying conditions by subdividing components into hierarchical layers and assigning side information, resulting in efficient bandwidth use and optimal audio quality for compressed HOA sound representations.

JP7690530B2Active Publication Date: 2025-06-10DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023144104
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-07-13
Filing Date
2023-09-06
Publication Date
2025-06-10
Estimated Expiration
2036-10-07

AI Technical Summary

Technical Problem

Existing layered audio coding methods struggle to efficiently adapt to time-varying transmission conditions, particularly for compressed higher-order ambisonics (HOA) sound representations, leading to signal dropouts and suboptimal quality.

Method used

A method for encoding a layer composition of a compressed sound representation, which subdivides components into hierarchical layers, assigns side information to each layer, and decodes the audio representation by utilizing the highest available layer and corresponding enhancement side information, ensuring optimal quality even with incomplete layer reception.

Benefits of technology

This approach enables efficient adaptation to varying transmission conditions, reduces bandwidth requirements, and ensures optimal audio quality by allowing decoding and enhancement using only the available layers, thereby minimizing signal dropouts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690530000011
    Figure 0007690530000011
  • Figure 0007690530000012
    Figure 0007690530000012
  • Figure 0007690530000013
    Figure 0007690530000013
Patent Text Reader

Abstract

To provide an encoding method of a layer structure of compressed sound representation of sound or a sound field including basic compressed sound representation including a plurality of components, basic side information for decoding the sound representation and making basic reconstructed sound representation of sound or sound field and improved side information including a parameter for improving the basic reconstructed sound representation.SOLUTION: A method subdivides a plurality of components into a plurality of component groups and allocates each subdivided group into one hierarchical layer in a plurality of hierarchical layers. The number of groups corresponds to the number of layers and the layer includes a basic layer and one or a plurality of hierarchical improved layers. The method adds basic side information to the basic layer, discriminates a plurality of portions from improved side information and allocates each of the plurality of portions to each of the plurality of layers. Each portion of the improved side information includes a parameter for improving reconstructed sound representation obtained from data included in the respective layers and arbitrary layers lower than the respective layers.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims the priority of European Patent Application No. 15306590.9 filed on October 15, 2015 and US Patent Application No. 62 / 361,809. The contents of these applications are hereby incorporated by reference in their entirety.

[0002] Technical Field This document relates to methods and apparatuses for layered audio coding. In particular, this document relates to methods and apparatuses for layered audio coding of compressed sound (or sound field) representations, such as higher - order ambisonics (HOA) sound (or sound field) representations.

Background Art

[0003] Regarding the streaming of sound (or sound field) representations through a transmission channel with time - varying conditions, layered coding is a means to adapt the quality of the received sound representation to the transmission conditions and in particular to avoid unwanted signal dropouts.

[0004] For layered coding, a sound (or sound field) representation is typically subdivided into a relatively small - sized high - priority base layer and additional enhancement layers with decreasing priority and arbitrary sizes. Each enhancement layer is typically assumed to contain incremental information to complement the information of all lower layers in order to improve the quality of the sound (or sound field) representation. The amount of error protection for the transmission of individual layers is controlled based on their priorities. In particular, the base layer is provided with high error protection, which is reasonable and acceptable due to its small size.

[0005] However, there is a need for a layered coding scheme for special types of compressed representations (extended versions) of sound or sound fields, such as compressed HOA sound or sound field representations.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] This paper addresses the above problems. In particular, methods and encoders / decoders for layer-constituted coding of compressed sound or sound field representations are described.

Means for Solving the Problems

[0008] According to one aspect, a method for encoding a layer composition of a compressed sound representation of a sound or sound field is described. The compressed sound representation may include a basic compressed sound representation including a plurality of components. The plurality of components may be complementary components. The compressed sound representation may further include basic side information for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The compressed sound representation may further include enhancement side information including parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The method may include subdividing (e.g., grouping) the plurality of components into a plurality of component groups. The method may further include assigning (e.g., adding) each of the plurality of groups to an individual one of a plurality of hierarchical layers. The assignment may indicate a correspondence between an individual group and a layer. The components assigned to each layer may be said to be included in that layer. The number of groups may correspond to (e.g., be equal to) the number of layers. The plurality of layers may include a base layer and one or more hierarchical enhancement layers. The plurality of hierarchical layers may be ordered from the base layer through a first enhancement layer, a second enhancement layer, etc. to an overall highest enhancement layer (overall top layer). The method may further include adding the basic side information to the base layer (e.g., including the basic side information in the base layer, or assigning the basic side information to the base layer, for purposes of transmission or storage). The method may further include discriminating a plurality of portions of the enhancement side information from the enhancement side information. The method may further include assigning (e.g., adding) each of the plurality of portions of the enhancement side information to an individual one of the plurality of layers. Each portion of the enhancement side information may include parameters for improving a reconstructed (e.g., decompressed) sound representation obtained from data included in (e.g., assigned or added to) that respective layer and any layers lower than that respective layer.The encoding of the layer structure may be performed for transmission through a transmission channel or for storage on a suitable storage medium such as, for example, a CD, DVD or Blu-ray Disc (trademark).

[0009] Configured as described above, the proposed method enables the efficient application of layer structure encoding to a compressed audio representation comprising a plurality of components and first and enhancement side information (such as, for example, independent base side information and enhancement side information) having the properties described above. In particular, the proposed method is such that each layer comprises suitable side information for reconstructing the reconstructed audio representation from the components included in any layer up to the layer in question. Here, the layers up to the layer in question are understood to include, for example, the base layer, the first enhancement layer, the second enhancement layer, etc. up to the layer in question. Thus, regardless of the actual highest usable layer (for example, the layer below the lowest layer not yet received effectively; all layers below the highest usable layer and the highest usable layer itself are received effectively), even if the reconstructed audio representation differs from the complete (for example, full) audio representation, the decoder is enabled to improve or enhance the reconstructed audio representation. In particular, regardless of the actual highest usable layer, in order to improve or enhance the reconstructed audio representation obtainable based on all the components included in the layers up to the actual highest usable layer, it is sufficient for the decoder to decode the payload of the enhancement side information for only a single layer (i.e., for the highest usable layer). That is, for each time interval (for example, frame), it may be sufficient to decode only a single payload of the enhancement side information. On the other hand, the proposed method allows to fully benefit from the advantage of bandwidth reduction achievable when applying layer structure encoding.

[0010] In some embodiments, the components of the basic compressed sound representation may correspond to a monaural signal (e.g., a transport signal or a monaural transport signal). The monaural signal may represent either a predominant sound signal or a sequence of coefficients of a HOA representation. The monaural signal may be quantized.

[0011] In some embodiments, the basic side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components individually and independently of the other components. For example, the basic side information may represent side information related to an individual monaural signal independently of other monaural signals. Thus, the basic side information may sometimes be referred to as independent basic side information.

[0012] In some embodiments, the enhancement side information may represent enhancement side information. The enhancement side information may include prediction parameters for the basic compressed sound representation for improving (e.g., enhancing) a basic reconstructed sound representation obtained from the basic compressed sound representation and the basic side information.

[0013] In some embodiments, the method may further include generating a transport stream for the transmission of the data of the plurality of layers (e.g., data assigned to or added to each layer or otherwise included in each layer). The base layer may have the highest priority for transmission, and the hierarchical enhancement layers may have a decreasing priority for transmission. That is, the priority for transmission may decrease from the base layer to the first enhancement layer, from the first enhancement layer to the second enhancement layer, and so on. The amount of error protection for the transmission of the data of the plurality of layers may be controlled according to the respective transmission priorities. Thereby, while reducing the overall required bandwidth by not applying excessive error protection to the upper layers, it can be ensured that at least some of the lower layers are transmitted in a reliable manner.

[0014] In some embodiments, the method may further include generating a transport layer packet for each of the plurality of layers, the transport layer packet including data for the respective layer. For example, for each time interval (e.g., frame), a transport layer packet may be generated for each of the plurality of layers.

[0015] In some embodiments, the compressed audio representation may further include additional base side information for decoding the basic compressed audio representation into a basic reconstructed audio representation. The additional base side information may include information specifying the decoding of one or more of the plurality of components depending on other components. The method may further include decomposing the additional base side information into a plurality of parts of the additional base side information. The method may further include adding those parts of the additional base side information to the base layer (e.g., including those parts of the additional base side information in the base layer for transmission or storage, or allocating those parts of the additional base side information to the base layer). Each part of the additional base side information may correspond to a respective layer and may include information specifying the decoding of one or more components assigned to the respective layer depending on (only) each other component assigned to the respective layer and any lower layer(s) below the respective layer. That is, each part of the additional base side information specifies the components in each respective layer to which that part of the additional base side information corresponds without referring to any other components assigned to layers above the respective layer.

[0016] Configured as such, the proposed method avoids fragmentation of additional basic side information by adding all parts to the base layer. In other words, all parts of the additional basic side information are included in the base layer. The decomposition of the additional basic side information ensures that, for each layer, a part of the additional basic side information that does not require knowledge of components of higher layers is available. Thus, regardless of the actual highest usable layer, it is sufficient for the decoder to decode the additional basic side information included in the layers up to the highest usable layer.

[0017] In some embodiments, the additional basic side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components depending on other components. For example, the additional basic side information may represent side information related to an individual monaural signal depending on other monaural signals. Thus, the additional basic side information may sometimes be referred to as dependent basic side information.

[0018] In some embodiments, the compressed audio representation may be processed for a series of time intervals, for example, time intervals of equal size. The series of time intervals may be frames. Thus, the method may operate on a frame-by-frame basis. That is, the compressed audio representation may be encoded frame by frame. The compressed audio representation may be available for each successive time interval (e.g., for each time frame). That is, the compression operation by which the compressed audio representation is obtained may also operate on a frame-by-frame basis.

[0019] In some embodiments, the method may further include generating, for each layer, configuration information indicating the components of the basic compressed audio representation assigned to that layer. Thus, the decoder can easily access the information necessary for decoding without performing unnecessary parsing through the received data payload.

[0020] According to another aspect, a method for encoding a layer composition of a compressed sound representation of a sound or sound field is described. The compressed sound representation may include a basic compressed sound representation including a plurality of components. The plurality of components may be complementary components. The compressed sound representation may further include basic side information (e.g., independent basic side information) and third information (e.g., dependent basic side information) for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The basic side information may include information for individually specifying the decoding of one or more of the plurality of components independently of other components. The additional basic side information may include information for specifying the decoding of one or more of the plurality of components depending on each other component. The method may include subdividing (e.g., grouping) the plurality of components into a plurality of component groups. The method may further include assigning (e.g., adding) each of the plurality of groups to an individual one of a plurality of hierarchical layers. The assignment may indicate a correspondence between an individual group and a layer. The components assigned to each layer may be said to be included in that layer. The number of groups may correspond to (e.g., be equal to) the number of layers. The plurality of layers may include a base layer and one or more hierarchical enhancement layers. The method may further include adding the basic side information to the base layer (e.g., including the basic side information in the base layer, or assigning the basic side information to the base layer, for purposes of transmission or storage). The method may further include decomposing the additional basic side information into a plurality of parts of the additional basic side information and adding those parts of the additional basic side information to the base layer (e.g., including those parts of the additional basic side information in the base layer, or allocating those parts of the additional basic side information to the base layer, for transmission or storage).Each part of the additional base side information may correspond to a respective layer and may include information specifying the decoding of one or more components assigned to that respective layer, depending on each other component assigned to that respective layer and any layer below that respective layer.

[0021] So configured, the proposed method ensures that for each layer, appropriate additional base side information is available to decode the components included in any layer up to that layer without requiring valid reception or decoding (or generally knowledge) of higher layers. In the case of a compressed HOA representation, the proposed method ensures that in the vector coding mode, suitable V-vectors are available for all components belonging to the layers up to the highest usable layer. In particular, the proposed method excludes cases where the elements of the V-vectors corresponding to components in higher layers are not explicitly signaled. Thus, the information included in the layers up to the highest usable layer is sufficient to decode (e.g., decompress) any component belonging to the layers up to the highest usable layer. Thereby, even if higher layers have not been validly received by the decoder, appropriate decompression of the respective reconstructed HOA representation for lower layers is guaranteed. On the other hand, the proposed method allows to fully benefit from the advantage of reduced bandwidth achievable when applying layer-configuration coding.

[0022] Embodiments of this aspect may be related to embodiments of the above aspect.

[0023] According to another aspect, a method for decoding a compressed audio representation of a sound or sound field is described. The compressed audio representation may be encoded in a plurality of hierarchical layers. The plurality of hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The plurality of layers may be assigned components of a basic compressed audio representation of the sound or sound field. In other words, the plurality of layers may include components of basic compressed side information. Those components may be assigned to respective layers within respective component groups. The plurality of components may be complementary components. The base layer may include basic side information for decoding the basic compressed audio representation. Each layer may include a portion of enhancement side information that includes parameters for improving a basic reconstructed audio representation obtained from data included in that respective layer and any layers below that respective layer. The method may include receiving data payloads corresponding respectively to the plurality of hierarchical layers. The method may further include determining a first layer index indicating the highest available layer to be used to decode the basic compressed audio representation into the basic reconstructed audio representation of the sound or sound field among the plurality of layers. The method may further include obtaining the basic reconstructed audio representation using the basic side information from components assigned to the highest available layer and any layers below the highest available layer. The method may further include determining a second layer index indicating which portion of the enhancement side information is to be used to improve (e.g., enhance) the basic reconstructed audio representation. The method may further include obtaining a reconstructed audio representation of the sound or sound field from the basic reconstructed audio representation with reference to the second layer index.

[0024] So configured, the proposed method guarantees that the reconstructed audio representation has optimal quality by making maximum use of the available (e.g., validly received) information.

[0025] In some embodiments, the components of the basic compressed sound representation may correspond to a monaural signal (e.g., a monaural transport signal). The monaural signal may represent either a predominant sound signal or a sequence of coefficients of a HOA representation. The monaural signal may be quantized.

[0026] In some embodiments, the basic side information may include information that individually specifies the decoding (e.g., decompression) of one or more of the plurality of components, independently of the other components. For example, the basic side information may represent side information related to an individual monaural signal, independently of other monaural signals. Thus, the basic side information may sometimes be referred to as independent basic side information.

[0027] In some embodiments, the enhancement side information may represent enhancement side information. The enhancement side information may include prediction parameters for the basic compressed sound representation for improving (e.g., enhancing) the basic reconstructed sound representation obtained from the basic compressed sound representation and the basic side information.

[0028] In some embodiments, the method may further include, for each layer, determining whether each layer has been successfully received. The method may further include determining the first layer index as the layer index of the layer immediately below the lowest layer that has not been successfully received.

[0029] In some embodiments, determining the second layer index may involve determining that the second layer index is equal to the first layer index, or determining an index value indicating that no enhancement side information is used when obtaining the reconstructed sound representation, as the second layer index. In the latter case, the reconstructed sound representation may be equal to the basic reconstructed sound representation.

[0030] In various embodiments, the data payload may be received and processed for a series of time intervals, such as time intervals of equal size. The series of time intervals may be frames. Thus, the method may operate in a frame-based manner. The method may further determine the second layer index to be equal to the first layer index if the compressed audio representations for the series of time intervals can be decoded independently of each other.

[0031] In various embodiments, the data payload may be received and processed for a series of time intervals, such as time intervals of equal size. The series of time intervals may be frames. Thus, the method may operate in a frame-based manner. The method may further include, for a given time interval of the series of time intervals, determining, for each layer, whether the respective layer was successfully received if the compressed audio representations for the series of time intervals cannot be decoded independently of each other. The method may further include determining the first layer index for the given time interval as the smaller of the first layer index of the time interval preceding the given time interval and the layer index of the layer immediately below the lowest layer that was not successfully received.

[0032] In some embodiments, the method may include determining, for the given time interval, whether a first layer index for the given time interval is equal to a first layer index for a preceding time interval if the compressed audio representations for the series of time intervals cannot be decoded independently of each other. The method may further include determining, if the first layer index for the given time interval is equal to the first layer index for a preceding time interval, the second layer index for the given time interval to be equal to the first layer index for the given time interval. The method may further include determining, if the first layer index for the given time interval is not equal to the first layer index for a preceding time interval, an index value indicating that no enhancement side information is used when obtaining the reconstructed audio representation as the second layer index.

[0033] In some embodiments, the base layer may include at least one portion of additional base side information corresponding to each layer and including information specifying decoding of one or more of the components assigned to that layer depending on other components assigned to that layer and any layers lower than that layer. The method may further include decoding the portion of the additional base side information by referring to the components assigned to that layer and any layers lower than that layer for each portion of the additional base side information. The method may further include correcting the portion of the additional base side information by referring to the components assigned to the highest available layer and any layers between the highest available layer and that layer. The basic reconstructed audio representation may be obtained using the base side information and the corrected portions of the additional base side information corresponding to the layers up to the highest available layer from the components assigned to the highest available layer and any layers lower than the highest available layer.

[0034] In some embodiments, the additional base side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components depending on other components. For example, the additional base side information may represent side information related to an individual monaural signal depending on other monaural signals. Thus, the additional base side information may sometimes be referred to as dependent base side information.

[0035] According to another aspect, a method for decoding a compressed audio representation of a sound or sound field is described. The compressed audio representation may be encoded in a plurality of hierarchical layers. The plurality of hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The plurality of layers may have components of a basic compressed audio representation of the sound or sound field assigned thereto. In other words, the plurality of layers may include components of basic compressed side information. Those components may be assigned to respective layers in respective component groups. The plurality of components may be complementary components. The base layer may include basic side information for decoding a basic compressed audio representation. The base layer may further include at least one portion of additional basic side information corresponding to each layer and specifying decoding of one or more of the components assigned to that layer depending on other components assigned to that layer and any layers lower than that layer. The method may further include receiving data payloads respectively corresponding to the plurality of hierarchical layers. The method may further include determining a first layer index indicating the highest available layer to be used for decoding the basic compressed audio representation into the basic reconstructed audio representation of the sound or sound field among the plurality of layers. The method may further include decoding the portion of the additional basic side information for each portion of the additional basic side information by referring to the components assigned to that layer and any layers lower than that layer. The method may further include correcting the portion of the additional basic side information for each portion of the additional basic side information by referring to the components assigned to the highest available layer and any layers between the highest available layer and that layer. The basic reconstructed audio representation may be obtained using the basic side information and the corrected portions of the additional basic side information from the portions of the additional basic side information corresponding to the layers up to the highest available layer from the components assigned to the highest available layer and any layers lower than the highest available layer.The method may further include determining a second layer index that is equal to the first layer index or indicates omission of enhancement side information during decoding.

[0036] Configured in such a way, the proposed method ensures that the additional basic side information ultimately used to decode the basic compressed audio representation does not contain redundant elements, thereby making the actual decoding of the basic compressed audio representation more efficient.

[0037] Embodiments of this aspect may be related to embodiments of the above aspect.

[0038] According to another aspect, an encoder for encoding a layer configuration of a compressed audio representation of an audio or an acoustic field is described. The compressed audio representation may include a basic compressed audio representation including a plurality of components. The plurality of components may be complementary components. The compressed audio representation may further include basic side information for decoding the basic compressed audio representation into a basic reconstructed audio representation of the audio or the acoustic field. The compressed audio representation may further include enhancement side information including parameters for improving (e.g., enhancing) the basic reconstructed audio representation. The encoder may include a processor configured to execute some or all of the method steps of the method based on the above-mentioned first-mentioned aspect and the above-mentioned second-mentioned aspect.

[0039] According to another aspect, a decoder for decoding a compressed sound representation of a sound or a sound field is described. The compressed sound representation may be encoded in a plurality of hierarchical layers. The plurality of hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The plurality of layers may be assigned components of a basic compressed sound representation of the sound or the sound field. In other words, the plurality of layers may include components of basic compressed side information. Those components may be assigned to respective layers in respective component groups. The plurality of components may be complementary components. The base layer may include basic side information for decoding a basic compressed sound representation. Each layer may include a part of enhancement side information including parameters for improving (e.g., enhancing) a basic reconstructed sound representation obtained from data included in that respective layer and any lower layers. The decoder may include a processor configured to execute some or all of the method steps of the method according to the third-mentioned aspect and the fourth-mentioned aspect above.

[0040] According to another aspect, a method, apparatus, and system are directed to decoding a compressed higher order ambisonics (HOA) sound representation of a sound or sound field. The apparatus may have a receiver configured to receive a bitstream including the compressed HOA representation corresponding to a plurality of hierarchical layers including a base layer and one or more hierarchical enhancement layers, or the method may perform the receiving. The plurality of layers are assigned components of a basic compressed sound representation of the sound or sound field, and the components are assigned to respective layers in respective component groups. The apparatus may have a decoder configured to decode the compressed HOA representation based on base side information associated with the base layer and enhancement side information associated with the one or more hierarchical enhancement layers, or the method may perform the decoding. The base side information may include base independent side information related to first individual monaural signals that are decoded independently of other monaural signals. Each of the one or more hierarchical enhancement layers may include a portion of the enhancement side information including parameters for improving a basic reconstructed sound representation obtained from data included in the respective layer and any layers lower than the respective layer.

[0041] The base independent side information may indicate that a first individual monaural signal represents a directional signal having an incident direction. The base side information may further include base dependent side information related to second individual monaural signals that are decoded depending on other monaural signals. The base dependent side information may include vector-based signals that are directionally distributed within the sound field. Here, the direction distribution is specified by a vector. The components of the vector are set to zero and are not part of the compressed vector representation.

[0042] The components of the basic compressed sound representation may correspond to a monaural signal representing either the dominant sound signal or the sequence of coefficients of the HOA representation. The bitstream includes data payloads respectively corresponding to the plurality of hierarchical layers. The enhancement side information may include parameters related to at least one of spatial prediction, subband directional signal synthesis, and parametric ambient sound replication. The enhancement side information may include information allowing prediction of sound from the directional signal or missing parts of the sound field. Further, for each layer, it may be determined whether the respective layer has been successfully received, and the layer index of the layer immediately below the lowest layer not successfully received may be determined.

[0043] According to another aspect, a software program is described. The software program is adapted for execution on a processor and may be adapted to perform some or all of the method steps outlined herein when executed on a computing device.

[0044] According to yet another aspect, a storage medium is described. The storage medium may include a software program adapted for execution on a processor and adapted to perform some or all of the method steps outlined herein when executed on a computing device.

[0045] Those skilled in the art will understand that statements made with respect to any of the above aspects or embodiments thereof also apply to other aspects or embodiments thereof. Repeating these statements for each individual aspect or embodiment has been omitted for the sake of brevity.

[0046] The methods and apparatuses including the preferred embodiments outlined herein may be used alone or in combination with other methods and systems disclosed herein. Further, all aspects of the methods and apparatuses outlined herein may be arbitrarily combined. In particular, the features of the claims may be combined with other features in any manner.

[0047] The method steps and apparatus features can be interchanged in many ways. In particular, as those skilled in the art will understand, the details of the disclosed method can be implemented as an apparatus adapted to perform some or all of the method steps, and vice versa.

Brief Description of the Drawings

[0048] The present invention will be described below in an exemplary manner with reference to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0049] First, a compressed audio (or sound field) representation (hereinafter referred to as a compressed audio representation for simplicity) to which the method and encoder / decoder according to the present disclosure are applicable will be described. In general, a complete compressed audio (or sound field) representation (hereinafter referred to as a complete compressed audio representation for simplicity) may include (e.g., consist of) the following three components: a basic compressed audio (or sound field) representation (hereinafter referred to as a basic compressed audio representation for simplicity), basic side information, and enhancement side information.

[0050] The basic compressed audio representation itself includes (e.g., consists of) several components (e.g., complementary components). The basic compressed audio representation may make up the largest and most prominent proportion of the complete compressed audio representation. The basic compressed audio representation may consist of a monaural transport signal representing a dominant audio signal or a coefficient sequence of the original HOA representation.

[0051] The basic side information is required to decode the basic compressed audio representation and may be assumed to be much smaller in size than the basic compressed audio representation. This may further consist of separate parts, most of which each specify the decompression of only one particular component of the basic compressed audio representation. The basic side information may consist of a first part, known as independent basic side information, and a second part, known as additional basic side information.

[0052] Both the first and second parts, i.e., the independent basic side information and the additional basic side information, may specify the decompression of a particular component of the basic compressed audio representation. The second part is optional and may be omitted. In this case, the compressed audio representation may be said to include the first part (e.g., the basic side information).

[0053] The first part (e.g., basic side information) may include side information that describes individual (complementary) components of a basic compressed audio representation independently of other (complementary) components. In particular, the first part (e.g., basic side information) may specify the decoding of one or more of the plurality of components individually, independently of the other components. Thus, the first part may be referred to as independent basic side information.

[0054] The second (optional) part, also known as additional basic side information, may describe individual (complementary) components of a basic compressed audio representation depending on other (complementary) components. This second part may be referred to as dependent basic side information. In particular, the dependency may have the following attributes: · The dependent basic side information for each individual (complementary) component of a basic compressed audio representation achieves its maximum extent when no other certain (complementary) components are included in the basic compressed audio representation. · When additional certain (complementary) components are added to the basic compressed audio representation, the dependent basic side information for the individual (complementary) component under consideration becomes a subset of the original dependent basic side information, thereby reducing its size.

[0055] Enhancement side information is also optional. This can be used to improve or enhance (e.g., parametrically improve or enhance) a basic compressed audio representation. Its size is also assumed to be much smaller than the size of the basic compressed audio representation.

[0056] Thus, in various embodiments, the compressed sound representation may include a basic compressed sound representation including a plurality of components, basic side information for decoding (e.g., decompressing) the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field, and enhancement side information including parameters for improving or enhancing (e.g., parametrically improving or enhancing) the basic compressed sound representation. The compressed sound representation may further include additional basic side information for decoding (e.g., decompressing) the basic compressed sound representation into the basic reconstructed sound representation, which may include information specifying decoding of one or more of the plurality of components depending on each other component.

[0057] One example of such a type of fully compressed sound representation is given by the compressed higher order ambisonics (HOA) sound field representation defined by Chapter 12 and Annex C.5 of the preliminary version of the MPEG-H 3D Audio standard (Non-Patent Document 1). That is, the compressed sound representation may correspond to a compressed HOA sound (or sound field) representation of the sound or sound field.

[0058] For this example, the basic compressed sound field representation (basic compressed sound representation) may include several components (e.g., may be identified by several components). Those components may be monaural signals (e.g., may correspond to monaural signals). Those monaural signals may be quantized monaural signals. Those monaural signals may represent either a dominant sound signal or a coefficient sequence of an ambient sound HOA sound field component.

[0059] Among other things, the basic side information can describe how each of these monaural signals contributes spatially to the sound field. For example, the basic side information may specify the dominant sound signal as a purely directional signal, i.e., a general plane wave having a certain incident direction. Alternatively, the basic side information may specify the monaural signal as a coefficient sequence of the original HOA representation having a certain index. The basic side information may further be separated into a first part and a second part as described above.

[0060] The first part is side information related to a particular individual monaural signal (e.g., independent basic side information). This independent basic side information is independent of the presence of other monaural signals. Such side information may, for example, specify a monaural signal representing a directional signal having a certain incident direction (e.g., meaning a general plane wave). Alternatively, the monaural signal may be specified as a coefficient sequence of the original HOA representation having a certain index. The first part may be referred to as independent basic side information. Generally, the first part (e.g., the basic side information) can specify the decoding of one or more of the plurality of monaural signals individually and independently of the other monaural signals.

[0061] The second part is side information related to a specific individual monaural signal (e.g., additional basic side information). This side information depends on the presence of other monaural signals. Such side signals may be used, for example, when the monaural signal is designated as a vector-based signal (see, e.g., Section 12.4.2.4.4 of Non-Patent Document 1). These signals are distributed directionally in the sound field, and the direction distribution can be specified by a vector. In certain modes (i.e., CodedVVecLength = 1), specific components of this vector are implicitly set to 0 and are not part of the compressed vector representation. These components have indices equal to those that are part of the basic compressed sound representation among the coefficient sequences of the original HOA representation. That is, when individual components of the vector are coded, the total number depends on the basic compressed sound representation. In particular, the total number depends on which coefficient sequence the original HOA representation contains.

[0062] If the coefficient sequence of the original HOA representation is not included in the basic compressed sound representation, the dependent basic side information for each vector-based signal consists of all vector components and has its maximum size. When coefficient sequences of the original HOA representation having certain indices are added to the basic compressed sound representation, the vector components having those indices are removed from the side information for each vector-based signal, thereby reducing the size of the dependent basic side information for the vector-based signal.

[0063] The enhancement side information (e.g., enhancement side information) may include parameters related to (broadband) spatial prediction (see Section 12.4.2.4.3 of Non-Patent Document 1) and / or parameters related to subband directional signal synthesis and parametric ambient sound replication.

[0064] The parameters related to (broadband) spatial prediction can be used to (linearly) predict the missing part of the sound field from the directional signal.

[0065] Sub-band directional signal synthesis and parametric ambient sound reproduction are compression tools recently introduced by the revision of the MPEG-H 3D audio standard (see Section 1 of Non-Patent Document 2). These tools allow for a frequency-dependent parametric prediction of additional monaural signals to be spatially distributed in order to complement a spatially incomplete or missing compressed HOA representation. The prediction may be based on the coefficient sequence of the basic compressed sound representation.

[0066] It is important to note that the above-mentioned complementary contribution to the sound field is represented within the compressed HOA representation by additional side information of a much smaller size, rather than by additional quantized signals. Thus, the two coding tools described above are particularly suitable for the compression of HOA representations at low data rates.

[0067] A second example of a compressed representation of one or more monaural signals having the structure described above includes encoded spectral information for separate frequency bands up to a certain upper frequency, which can be considered as the basic compressed representation; basic side information that specifies the encoded spectral information (e.g., by the number and width of the encoded frequency bands); and enhancement side information that includes parameters for spectral band replication (SBR) (e.g., consisting of). The parameters of the enhancement side information describe how to parametrically reconstruct spectral information for higher frequency bands that are not considered in the basic compressed representation from the basic compressed representation.

[0068] The present disclosure proposes a method for the coding of a layer structure of a complete compressed sound (or sound field) representation having the structure described above.

[0069] Compression may be frame-based in the sense that it provides a compressed representation (e.g., in the form of data packets, or equivalently frame payloads) for a series of time intervals. The time intervals may have equal or different sizes. These data packets may be assumed to contain, in addition to the data of the actual compressed representation, a validity flag, a value indicating its size. In the following, without intention of limitation, compression is assumed to be frame-based. Further, unless otherwise specified, without intention of limitation, the handling of a single frame is focused on. Thus, the frame index is omitted.

[0070] Each frame payload of the considered complete compressed sound (or sound field) representation is assumed to contain J data packets (or frame payloads), each data packet being for one component of the basic compressed sound representation denoted as BSRC j , where j = 1, …, J. Further, each data packet is assumed to contain a packet with independent basic side information denoted as BSI I . BSI I specifies certain components BSRC j of the basic compressed sound representation independently of the other components. Optionally, each data packet is further assumed to contain a packet with dependent basic side information (additional basic side information) denoted as BSI D . BSI D specifies certain components BSRC j of the basic compressed sound representation depending on the other components.

[0071] The information contained within two data packets BSI I and BSI D may optionally be grouped into a single data packet BSI of the basic side information. The single data packet BSI contains, among other things, each a specific component BSRC jIt may be said to include J parts that specify. Each of these parts may be said to include a part of independent side information and optionally a part of dependent side information.

[0072] Finally, each data packet may include an enhancement side information (ESI) payload that describes how to improve or enhance the reconstructed sound (or sound field) from the complete basic compressed sound representation.

[0073] The proposed solution for encoding the layer structure addresses the compression part including the packing of data packets for transmission and the necessary steps to enable both the receiver and the decompression part. Each part will be described in detail below.

[0074] First, compression and packing (for example, for transmission) will be described. In particular, the components and elements of the complete compressed sound (or sound field) representation in the case of encoding the layer structure will be described.

[0075] Figure 1 schematically shows a flowchart of an example of a method for compression and packing (for example, an encoding method of a compressed sound representation of sound or sound field, or an encoding method of a layer structure). The assignment (for example, allocation) of individual payloads to the base layer and (M - 1) enhancement layers may be achieved by a transport layer packer. Figure 2 schematically shows a block diagram of an example of the assignment / allocation of individual payloads.

[0076] As described above, the fully compressed sound representation 2100 may be related to, for example, a compressed HOA representation that includes a basic compressed sound representation. The fully compressed sound representation 2100 may include a plurality of components (e.g., a monaural signal) 2110-1, …, 2110-J, independent base side information (base side information) 2120, optional enhancement side information (enhancement side information) 2140, and optional dependent base side information (additional base side information) 2130. The base side information 2120 may be information for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The base side information 2120 may include information that individually specifies the decoding of one or more components (e.g., a monaural signal), independently of the other components. The enhancement side information 2140 may include parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The additional base side information 2130 may be (further) information for decoding the basic compressed sound representation into the basic reconstructed sound representation, and may include information that individually specifies the decoding of one or more of the plurality of components, depending on each of the other components.

[0077] Figure 2 shows a premise assumption that there are a plurality of hierarchical layers including one basic layer (basic layer) and one or more (hierarchical) enhancement layers. For example, there may be a total of M layers, i.e., one basic layer and M-1 enhancement layers. The plurality of hierarchical layers have successively increasing layer indices. The lowest value of the layer index (e.g., layer index 1) corresponds to the basic layer. Further, it is understood that the layers are ordered from the basic layer, through the various enhancement layers, to the overall highest enhancement layer (i.e., the overall top layer).

[0078] The proposed method may be performed on a frame-by-frame basis (i.e., on a per-frame basis). In particular, the compressed audio representation 2100 may be compressed for a series of time intervals, for example, time intervals of equal size. Each time interval may correspond to a frame. The following steps may be performed for each of the series of time intervals (e.g., frames).

[0079] In S1010 of FIG. 1, the plurality of components 2110 are subdivided into a plurality of component groups. Each of the plurality of groups is then assigned (e.g., added or allocated) to a corresponding one of the plurality of hierarchical layers. Here, the number of groups corresponds to the number of layers. For example, the number of groups may be equal to the number of layers, such that for each layer, there may be one group of components. As shown above, the plurality of layers may include a base layer and one or more (e.g., M - 1) hierarchical enhancement layers.

[0080] In other words, the basic compressed audio representation is subdivided into portions assigned to the individual layers. Without loss of generality, the grouping can be described by a number J of M + 1 m , m = 0,..., M. Here, J 0 = 1, J M = J + 1, and the component BSRC j is assigned to the m-th layer for J m-1 ≦ j < J m .

[0081] In S1020, the groups of components are assigned to the respective layers. In S1030, base side information 2120 is added (e.g., allocated) to the base layer (i.e., the lowest of the plurality of hierarchical layers).

[0082] That is, due to its small size, it is proposed to include the complete base side information (base side information and any additional base side information) in the base layer to avoid its useless fragmentation.

[0083] If the considered compressed audio representation includes dependent base side information (additional base side information), the method may further (not shown in FIG. 1) include decomposing the additional base side information into a plurality of parts 2130-1, …, 2130-M of the additional base side information. Those parts of the additional base side information may then be added to (e.g., assigned to) the base layer. In other words, those parts of the additional base side information may be included in the base layer. Each part of the additional base side information may correspond to a respective layer and may include information for specifying the decoding of one or more components assigned to the respective layer depending on other components assigned to the respective layer and any layers lower than the respective layer.

[0084] Thus, the independent base side information BSI I (base side information) 2120 remains unchanged for assignment, while the dependent base side information needs to be treated specially for the encoding of the layer structure to allow correct decoding at the receiver side and to reduce the size of the transmitted dependent base side information. It is proposed to decompose the dependent base side information into M parts (parts) denoted by BSI D,m , m = 1, …, M. Here, the m-th part includes the basic compressed audio representation component BSRC assigned to the m-th layer j , J m-1 ≦j < J m and the dependent base side information for each of them. This is under the assumption that the optional dependent base side information exists for the considered compressed audio representation. If the respective dependent side information does not exist, for that compressed audio representation of the parts, BSI D,m is assumed to be empty. Each part BSI of the dependent base side information D,m may depend on all components BSRC included in all layers up to the m-th layer (i.e., included in all layers j = 1, …, m), j 1 ≦ j < J m

[0085] Independent base side information packet BSI​I If it is of a size small enough to be negligible, it is reasonable to keep it as a whole and add (allocate) it to the base layer. Optionally, for the independent base side information as well, the same decomposition as for the dependent base side information can be performed, and the packet BSI I,m is given for m = 1, …, M. This is useful for reducing the size of the base layer by adding (allocating) the parts of the independent base side information to the layers having the corresponding components of the basic compressed sound representation.

[0086] In S1040, a plurality of parts 2140-1, …, 2140-M of the enhancement side information may be determined. Each part of the enhancement side information may include parameters for improving (e.g., enhancing) the reconstructed sound representation obtained from the data included in each layer and any layers lower than that respective layer.

[0087] The reason for performing this step is that in the case of encoding of the layer structure, it is important to recognize that since the enhancement side information needs to be calculated additionally for each layer with the intention of improving the preliminary decompressed sound (or sound field), it depends on the available layers for decompression. Specifically, the preliminary decompressed sound (or sound field) for a given highest decodable layer (highest usable layer) depends on the components included in the highest decodable layer and any layers lower than the highest decodable layer. Thus, the compression is ESI m needs to provide M individual enhancement side information data packets (parts of the enhancement side information) denoted by m = 1, …, M. Here, the enhancement side information ESI m in the m-th data packet is calculated to improve the sound (or sound field) representation obtained from all the data included in the base layer and the enhancement layers having indices lower than m (e.g., all the data included in the m-th layer and any layers lower than the m-th layer).

[0088] In S1050, the plurality of portions 2140-1, …, 2140-M of the enhancement side information are assigned (e.g., added or allocated) to the plurality of layers. Each of the plurality of portions of the enhancement side information is assigned to each of the plurality of layers. For example, each of the plurality of layers includes a respective portion of the enhancement side information.

[0089] The assignment of the base and / or enhancement side information to each layer may be indicated in the configuration setting information generated by the encoding method. In other words, the correspondence between the base and / or enhancement side information and each layer may be indicated in the configuration setting information. Further, the configuration setting information may indicate, for each layer, the components of the basic compressed audio representation that are assigned (e.g., included) in that layer. The portions of the additional base side information are included in the base layer but may correspond to layers different from the base layer.

[0090] In summary, in the compression stage, a frame data packet denoted as FRAME having the following composition is provided:

[0091]

Number

[0092]

Number

[0093] Individual data packets may then be grouped within the payload. The payload is defined as a special data packet that includes, in addition to the actual compressed representation data, a validity flag and a value indicating its size. The use of the payload allows for simple demultiplexing on the receiver side and has the advantage that old payloads can be discarded without having to parse through them. One possible grouping is given as follows. · Each BSRC j packet, j = 1, …, J is assigned (e.g., allocated) to an individual payload (denoted by BP with a bar). j ) · The m-th enhanced side information data packet ESI m and the m-th dependent side information data packet BSI D,m are assigned (e.g., allocated) to one enhanced payload (denoted by EP with a bar, m = 1, …, M). m · The independent basic side information BSI I is assigned to a separate side information payload (denoted by BSIP with a bar).

[0094] Optionally, if the size of the independent basic side information is large, each m-th BSI I,m among its components, m = 1, …, M may be assigned (e.g., allocated) to the aforementioned enhanced payload (denoted by EP with a bar). m In this case, the side information payload (denoted by BSIP with a bar) is empty and can be ignored.

[0095] Another option is to assign all the dependent basic side information data packets BSI D,m to the side information payload (denoted by BSIP with a bar). This is reasonable when the size of the dependent basic side information is small.

[0096] Finally, a frame data packet denoted by FRAME with the following composition may be provided.

[0097] ​

Number

[0098] The method may further (not shown in FIG. 1) include, for each of the plurality of layers, generating transport - layer packets (e.g., basic - layer packet 2200 and M - 1 enhancement - layer packets 2300 - 1, …, 2300-(M - 1)) that contain the data of the respective layer (e.g., for the basic layer, components, basic side - information, and enhancement side - information, or for the one or more enhancement layers, components and enhancement side - information).

[0099] The transport - layer packets for different layers may have different transmission priorities. Thus, the method may further (not shown in FIG. 1) include generating a transport stream for the transmission of the data of the plurality of layers. Here, the basic layer has the highest transmission priority, and the hierarchical enhancement layers have decreasing transmission priorities. Here, the higher the transmission priority, the greater the degree of error protection, and vice versa.

[0100] Unless another step that requires a pre - condition step is present, the above - mentioned steps may be executed in any order, and the exemplary order shown in FIG. 1 is not limiting.

[0101] FIG. 3 shows a method for decoding a compressed audio representation of sound or sound field for decoding or decompression (unpacking). Examples of corresponding receivers and decompression stages are schematically shown in the block diagrams of FIGS. 4A and 4B.

[0102] As can be seen from the above, the compressed sound representation may be encoded in the plurality of hierarchical layers. Basic compressed sound representation components may be assigned to the plurality of layers (for example, may include such components). These components are assigned to respective layers in respective component groups. The base layer may include basic side information for decoding the basic compressed sound representation. Each layer may include one of the above-described portions of enhancement side information that includes parameters for improving the basic reconstructed sound representation obtained from the data included in that respective layer and any layers lower than that respective layer.

[0103] The proposed method may be performed on a frame-by-frame basis (i.e., in a per-frame manner). In particular, the restored representation of the sound or sound field may be generated for a series of time intervals, such as time intervals of equal size. Those time intervals may be frames, for example. The following steps may be performed for each of a series of time intervals (such as frames).

[0104] In S3010, a data payload (for example, a transport layer packet) corresponding to the plurality of layers is received. The data payload may be received as part of a bitstream that includes a compressed HOA representation of the sound or sound field corresponding to the plurality of hierarchical layers. The hierarchical layers include a base layer and one or more enhancement layers. Basic compressed sound representation components of the sound or sound field may be assigned to the plurality of layers. The components are assigned to respective layers in respective component groups.

[0105] Individual layer packets may be multiplexed to provide a received frame packet of the complete compressed sound representation. The received frame packet may be

[0106]

Number

[0107]

Number

[0108] Using the payload, the received frame packets are

[0109]

Number

[0110] The received frame packets may then be passed to a decompressor or decoder 4100. If there are no errors in the transmission of the individual layers, at least the side information payload (for example, corresponding to the side information part of the enhancement side) included

[0111]

Number

[0112] In the decompressor 4100, the received frame packets may be demultiplexed. For this purpose, information about the size of each payload may be utilized to avoid unnecessary parsing through the data of the individual payloads.

[0113] In S3020, a first layer index indicating the highest layer (e.g., the highest available layer or the highest decodable layer) among the plurality of layers is determined to be used to decode the basic compressed sound representation into the basic reconstructed sound representation of the sound or sound field.

[0114] Furthermore, in S3020, a value (e.g., layer index) N of the highest layer (the highest available layer) to be used for decompressing the basic sound representation B may be selected. The highest layer actually used for decompressing the basic sound representation Upward is given by N B -1. Since each layer contains exactly one enhancement side information payload (a part of the enhancement side information), based on the enhancement side information payload, it can be determined whether the layer containing it is valid (e.g., validly received). Thus, the selection can be achieved using all the enhancement side information payloads ESI m , m = 1, …, M (or correspondingly

[0115]

Number

[0116] In S3030, a basic reconstructed sound representation is obtained. The basic reconstructed sound representation may be obtained using the basic side information (or generally using the basic side information) from the highest available layer indicated by the first layer index and components assigned to any layers lower than this highest available layer.

[0117] The payloads of the basic compressed sound representation components BSRC 1 , …, BSRC J are the basic side information payloads (e.g., BSI or BSI I and BSI D,m , m = 1, …, M) (all) and the value N BTogether with, it may be provided to the basic representation decompression processing unit 4200. The basic representation decompression processing unit 4200 (shown in FIGS. 4A and 4B) is the lowest N B layers, i.e., the base layer and N B -1 enhancement layers (i.e., the layers up to the layer indicated by the first layer index), reconstructs the basic sound (or sound field) representation using only the basic compressed sound representation components included therein. Alternatively, only the payloads of the basic compressed sound representation components included in the lowest N B layers may be provided to the basic representation decompression processing unit 4200 together with their respective basic side information payloads.

[0118] It is assumed that the information required about which components of the basic compressed sound (or sound field) representation are included in each individual layer is known to the decompressor 4100 from the data packet with the configuration setting information. The configuration setting information is assumed to be transmitted and received before the frame data packet.

[0119] Dependent side information data packet BSI D,m , m = 1, …, N B and the enhancement side information data packet ESI NE To provide, all enhancement payloads may be input to the partial parser 4400 (see FIG. 4B) of the decompressor 4100 together with the value N E and the value N B . The parser may discard all payloads and data packets that are not used for actual decompression. If the value of N E is equal to 0, it may be assumed that all enhancement side information data packets are empty.

[0120] If the base layer includes at least one dependent base side information payload (part of the additional base side information) corresponding to each layer, each individual dependent base side information payload (e.g., BSI D,m , m = 1, …, N BThe decoding of (a part of the additional basic side information) may include: (i) performing the decoding of the said part of the additional basic side information by referring to the components assigned to each of its layers and any layers lower than each of those layers (preliminary decoding); and (ii) performing the correction of the said part of the additional basic side information by referring to the components assigned to the highest available layer and any layers between the highest available layer and each of the layers (correction). Here, the additional basic side information corresponding to each layer includes information specifying the decoding of one or more of the components assigned to that layer depending on other components assigned to that layer and any layers lower than that layer.

[0121] Subsequently, the basic reconstructed sound representation can be obtained (e.g., generated) from the components assigned to the highest available layer and any layers lower than the highest available layer, using the basic side information and the corrected parts of the additional basic side information obtained from the parts of the additional basic side information corresponding to the layers up to the highest available layer.

[0122] In particular, the preliminary decoding of each payload BSI D,m , m = 1, …, N B may involve leveraging the dependencies on the first J m −1 basic compressed sound representation components BSRC 1 , …, BSRC (Jm)-1 included in the first m layers assumed in the encoding stage.

[0123] The sequential correction of each payload BSI D,m , m = 1, …, N B is such that the basic sound components are the first J B −1 basic compressed sound representation components BSRC NB included in the first N 1 > m layers, which are more components than assumed for the preliminary decoding, and …, BSRC (JNB)-1It may also be related to considering that it will finally be reconstructed. Thus, the correction may be achieved by discarding the outdated information. This is possible because if certain complementary components are added to the attributes assumed initially for the dependent basic side information, that is, in the basic compressed sound representation, the dependent basic side information for each individual (complementary) component will be a subset of the original.

[0124] In S3040, a second layer index may be determined. The second layer index may indicate the part(s) of the enhancement side information to be used to improve (e.g., enhance) the basic reconstructed sound representation.

[0125] In addition to the first layer index, an index (second layer index) N of the enhancement side information payload (part of the second enhancement information) to be used for decompression E may be determined. The second layer index N E may always be equal to the first layer index N B or equal to 0. The enhancement may always be achieved according to the basic sound representation obtained from the highest available layer or not achieved at all.

[0126] In S3050, the reconstructed sound representation of the sound or sound field is obtained (e.g., generated) from the basic reconstructed sound representation with reference to the second layer index.

[0127] That is, the reconstructed sound representation is obtained by (parametrically) improving or enhancing the basic reconstructed sound representation, for example, by using the enhancement side information (a part of the enhancement side information) indicated by the second layer index. As will be described later, the second layer index may indicate not to use any enhancement side information at this stage. Then, the reconstructed sound representation will correspond to the basic reconstructed sound representation.

[0128] For this purpose, the reconstructed basic sound representation, together with all enhancement side information payloads ESI 1 , …, ESI M , the basic side information payload (e.g., BSI or BSI I and BSI D,m , m = 1, …, M) and the value N E are provided to the enhancement representation decompression processing unit 4300 (shown in FIGS. 4A and 4B). The enhancement representation decompression processing unit 4300 uses only the enhancement side information payload ESI NE and discards all other enhancement side information payloads to calculate the final enhanced sound (or sound field) representation 2100'. Alternatively, only the enhancement side information payload ESI NE may be provided to the enhancement representation decompression processing unit 4300 instead of all enhancement side information payloads. If the value of N E is equal to 0, all enhancement side information payloads are discarded (alternatively, the enhancement side information payload is not provided). And the reconstructed final enhanced sound representation 2100' is equal to the reconstructed basic sound representation. The enhancement side information payload ESI NE may be obtained by the partial parser 4400.

[0129] FIG. 3 also generally shows decoding a compressed HOA representation based on basic side information associated with a base layer and enhancement side information associated with one or more hierarchical enhancement layers.

[0130] Unless another step that requires a prerequisite step is present, the above steps may be performed in any order, and the exemplary order shown in FIG. 3 is not limiting.

[0131] Next, details of the layer selection (selection of the first and second layer indices) for decompression in steps S3020 and S3040 are described.

[0132] The determination of the first layer index may involve, for each layer, determining whether the layer has been successfully received. The determination of the first layer index may further involve determining the first layer index as the layer index of the layer immediately below the lowest layer that has not been successfully received. Whether a layer has been successfully received may be determined by evaluating whether the enhancement side information payload of that layer has been successfully received. This may be done by evaluating the validity flag within the enhancement side information payload.

[0133] The determination of the second layer index may generally involve determining the second layer index to be equal to the first layer index, or determining an index value (e.g., index value 0) indicating that no enhancement side information is used when obtaining the reconstructed audio representation, as the second layer index.

[0134] If all frame data packets can be decompressed independently of each other, the number N of the highest layer (the highest available layer) actually used for decompressing the basic audio representation B and the index N of the enhancement side information payload used for decompression E may both be set to the highest number L of valid enhancement side information payloads. L itself may be determined by evaluating the validity flag within the enhancement side information payload. By utilizing knowledge of the size of each enhancement side information payload, complex parsing through the actual data of the payload for validity determination can be avoided.

[0135] That is, if the compressed audio representation for a series of time intervals can be decoded independently, the second layer index may be determined to be equal to the first layer index. In this case, the reconstructed basic audio representation may be enhanced based on the enhancement side information payload of the highest available layer.

[0136] When differential decompression with inter-frame dependencies is used, decisions from the previous frame need to be further considered. In differential decompression, typically, independent frame data packets are transmitted at regular time intervals. This is to allow decompression to start from those points. In independent frame data packets, the determination of values N B and N E becomes frame-independent and is executed as described above.

[0137] To explain the proposed frame-dependent determination in detail, let the highest number (e.g., layer index) of the valid improvement side information payload for the k-th frame be L(k), the highest layer number (e.g., layer index) selected and used for decompressing the basic audio representation be N B (k), and the number (e.g., layer index) of the improvement side information payload used for decompression be N E (k).

[0138] Using this notation, the highest layer number N B (k) used for decompressing the basic audio representation is calculated according to the following formula.

[0139]

Equation

[0140] That is, when the compressed audio representations for a series of time intervals (e.g., frames) cannot be decoded independently of each other, determining the first layer index may include, for each layer, determining whether the respective layer has been successfully received, and determining the first layer index for the given time interval as the smaller of the first layer index of the time interval preceding the given time interval and the layer index of the layer immediately below the lowest layer that was not successfully received.

[0141] The number N of the enhancement side information payload used for decompression E (k) may be determined according to the following equation.

[0142]

Equation

[0143] That is, specifically, as long as the highest layer number N B (k) does not change, the same corresponding enhancement layer number is selected. However, if N B (k) changes, then the enhancement is disabled by setting N E (k) to 0. Due to the assumed differential decompression of the enhancement side information, a change based on N B (k) is not possible. It would require decompression of the corresponding enhancement side information layer in the previous frame, but it is assumed that such decompression was not performed.

[0144] That is, if the compressed audio representations for a series of time intervals (e.g., frames) cannot be decoded independently of each other, the determination of the second layer index may include determining whether the first layer index for the given time interval is equal to the first layer index for the preceding time interval. If the first layer index for the given time interval is equal to the first layer index for the preceding time interval, the second layer index for the given time interval may be determined (e.g., selected) to be equal to the first layer index for the given time interval. On the other hand, if the first layer index for the given time interval is not equal to the first layer index for the preceding time interval, an index value indicating that no enhancement side information is used when obtaining the reconstructed audio representation may be determined (e.g., selected) as the second layer index.

[0145] Alternatively, in decompression, if all of the enhancement side information payloads numbered up to N E (k) are decompressed in parallel, the selection rule of Equation (4) is N E (k)=N B (k) (9) is replaced by.

[0146] Finally, for differential decompression, note that the number N B of the highest layer used can only increase in an independent frame - data - packet, while a decrease is possible in any frame.

[0147] It is understood that a proposed method for encoding a layer composition of a compressed sound representation can be implemented by an encoder for encoding a layer composition of a compressed sound representation. Such an encoder may have respective units adapted to perform each of the above steps. An example of such an encoder 5000 is schematically shown in FIG. 5. For example, such an encoder 5000 may include a component subdivision unit 5010 adapted to perform S1010 described above, a component assignment unit 5020 adapted to perform S1020 described above, a basic side information assignment unit 5030 adapted to perform S1030 described above, an enhanced side information splitting unit 5040 adapted to perform S1040 described above, and an enhanced side information assignment unit 5050 adapted to perform S1050 described above. Further, it is understood that each unit of such an encoder may be embodied by a processor 5100 of a computing device adapted to perform the processing performed by each of the units, i.e., adapted to perform some or all of the above steps of the proposed encoding method or further steps. The encoder or computing device may further have a memory 5200 accessible by the processor 5100.

[0148] Furthermore, it is understood that a proposed method for decoding a compressed sound representation encoded in a plurality of hierarchical layers can be implemented by a decoder for decoding the compressed sound representation encoded in the plurality of hierarchical layers. Such a decoder may have respective units adapted to perform each of the above steps. An example of such a decoder 6000 is schematically shown in FIG. 6. For example, such a decoder 6000 may have a reception unit 6010 adapted to perform S3010 described above, a first layer index determination unit 6020 adapted to perform S3020 described above, a basic reconstruction unit 6030 adapted to perform S3030 described above, a second layer index determination unit 6040 adapted to perform S3040 described above, and an improvement reconstruction unit 6050 adapted to perform S3050 described above. Furthermore, it is understood that each unit of such a decoder may be embodied by a processor 6100 of a computing device adapted to perform the processing executed by each of the units, i.e., adapted to perform some or all of the above steps of the proposed decoding method or further steps. The decoder or the computing device may further have a memory 6200 accessible by the processor 6100.

[0149] It should be noted that this document and the drawings merely show the principles of the proposed method and apparatus. Therefore, it is understood that those skilled in the art will be able to devise various configurations that embody the principles of the present invention and are within its spirit and scope, even if not explicitly described or illustrated in this document. Furthermore, all examples described in this document are clearly intended solely for the educational purpose of helping the reader understand the principles of the proposed method and apparatus and the concepts contributed by the inventors to the advancement of the art, and are to be interpreted without limitation to such individually described examples and conditions. Furthermore, any statements in this document describing the principles, aspects, and embodiments of the present invention, as well as its individual examples, are intended to include their equivalents.

[0150] The methods and apparatuses described in this document may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or a microprocessor. Other components may be implemented as hardware and / or as application-specific integrated circuits. The signals that occur in the described methods and apparatuses may be stored on a medium such as a random access memory or an optical storage medium, and may be transferred via a network such as a radio network, a satellite network, a wireless network, or a wired network, for example, the Internet.

[0151] Some aspects will be described. 〔Aspect 1〕 A method for decoding a compressed higher-order ambisonics (HOA) representation of a sound or a sound field, the method comprising: receiving a bitstream comprising the compressed HOA representation corresponding to a plurality of hierarchical layers including a base layer and one or more hierarchical enhancement layers, wherein the plurality of layers are assigned components of a basic compressed sound representation of the sound or the sound field, and the components are assigned to respective layers in respective component groups; decoding the compressed HOA representation based on base side information associated with the base layer and enhancement side information associated with the one or more hierarchical enhancement layers; wherein each of the one or more hierarchical enhancement layers comprises a portion of the enhancement side information that includes parameters for improving a basic reconstructed sound representation obtained from data included in each layer and any lower layers of that layer. Method. 〔Aspect 2〕 The components of the basic compressed sound representation correspond to monaural signals; The monaural signal represents either a dominant sound signal or a coefficient sequence of HOA representation. The method according to aspect 1. 〔Aspect 3〕 The method according to aspect 1 or 2, wherein the bitstream includes data payloads respectively corresponding to the one or more hierarchical layers. 〔Aspect 4〕 The method according to any one of aspects 1 to 3, wherein the enhancement side information includes parameters related to at least one of spatial prediction, subband directional signal synthesis, and parametric ambient sound replication. 〔Aspect 5〕 The method according to any one of aspects 1 to 4, wherein the enhancement side information includes information allowing prediction of sound from a directional signal or a missing part of a sound field. 〔Aspect 6〕 For each layer, determine whether the respective layer has been successfully received; further including determining the layer index of the layer immediately below the lowest layer that has not been successfully received. The method according to any one of aspects 1 to 5. 〔Aspect 7〕 The method according to any one of aspects 1 to 6, further including determining a second layer index equal to the first layer index or indicating omission of enhancement side information during decoding. 〔Aspect 8〕 Determine a first layer index indicating the highest available layer among the plurality of layers, which is used to decode the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field; further including obtaining the basic reconstructed sound representation using the first side information from components assigned to the highest available layer and any layers below the highest available layer. The method according to any one of aspects 1 to 7. 〔Aspect 9〕 wherein the base layer is at least one part corresponding to each layer of additional base side information, the part including information specifying decoding of one or more components of the components assigned to each layer depending on other components assigned to each of the respective layer and any layers lower than the respective layer, and the method for each part of the additional base side information: decoding the part of the additional base side information by referring to components assigned to each of the respective layer and any layers lower than the respective layer; including correcting the part of the additional base side information by referring to components assigned to the highest available layer and any layers between the highest available layer and each of the respective layer, wherein the basic reconstructed sound representation is obtained from the components assigned to the highest available layer and any layers lower than the highest available layer, using the base side information and the corrected parts of the additional base side information obtained from the parts of the additional base side information corresponding to the layers up to the highest available layer; The method according to any one of aspects 1 to 8. [Aspect 10] An apparatus for decoding a compressed higher order ambisonics (HOA) sound representation of a sound or sound field, the apparatus comprising: a receiver configured to receive a bitstream including the compressed HOA representation corresponding to a plurality of hierarchical layers including a base layer and one or more hierarchical enhancement layers, wherein components of a basic compressed sound representation of the sound or sound field are assigned to the plurality of layers, and the components are assigned to each layer in respective component groups; a decoder configured to decode the compressed HOA representation based on base side information associated with the base layer and enhancement side information associated with the one or more hierarchical enhancement layers; Each of the one or more hierarchical enhancement layers includes a portion of the enhancement side information that includes parameters for improving a basic reconstructed sound representation obtained from data included in each layer and any layers lower than that respective layer. Apparatus. [Aspect 11] The components of the basic compressed sound representation correspond to a monaural signal; The monaural signal represents either a dominant sound signal or a sequence of coefficients of a HOA representation. The apparatus according to aspect 10. [Aspect 12] The apparatus according to aspect 10 or 11, wherein the bitstream includes data payloads respectively corresponding to the one or more hierarchical layers. [Aspect 13] The apparatus according to any one of aspects 10 to 12, wherein the enhancement side information includes parameters related to at least one of spatial prediction, subband directional signal synthesis, and parametric ambient sound replication. [Aspect 14] The apparatus according to any one of aspects 10 to 13, wherein the enhancement side information includes information that allows prediction of sound or missing portions of a sound field from a directional signal. [Aspect 15] For each layer, determine whether that respective layer has been successfully received; Further including determining the layer index of the layer immediately below the lowest layer that has not been successfully received. The apparatus according to any one of aspects 10 to 14. [Aspect 16] The apparatus according to any one of aspects 10 to 15, further including determining a second layer index equal to the first layer index or indicating omission of enhancement side information during decoding. [Aspect 17] Determine a first layer index indicating the highest available layer among the plurality of layers used to decode the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field; further comprising obtaining the basic reconstructed audio representation from components assigned to the highest available layer and any layers below the highest available layer, using the first side information An apparatus according to any one of aspects 10 to 16. [Aspect 18] wherein the base layer comprises at least one portion of additional base side information corresponding to each layer, the portion specifying decoding of one or more components of the components assigned to the respective layer depending on the respective layer and other components assigned to any layers below the respective layer, and the method for each portion of the additional base side information: decoding the portion of the additional base side information by referring to the components assigned to the respective layer and any layers below the respective layer; comprising correcting the portion of the additional base side information by referring to the components assigned to the highest available layer and any layers between the highest available layer and the respective layer, wherein the basic reconstructed audio representation is obtained from the components assigned to the highest available layer and any layers below the highest available layer, using the basic side information and the corrected portions of the additional base side information obtained from the portions of the additional base side information corresponding to the layers up to the highest available layer An apparatus according to any one of aspects 10 to 17.

Claims

A method for decoding a compressed higher-order ambisonics (HOA) representation of sound or sound field, executed by a device having a receiver, a processor, and a decoder, the method comprising: receiving, by the receiver, a bitstream including the compressed HOA representation, the bitstream including a plurality of hierarchical layers including a base layer and two or more hierarchical enhancement layers, the bitstream including at least a data payload corresponding to the plurality of hierarchical layers, the bitstream further including base side information associated with the base layer and enhancement side information associated with the two or more hierarchical enhancement layers; assigning, to at least one of the plurality of hierarchical layers, a component of the compressed HOA representation of the sound or sound field, the component of the basic compressed sound representation corresponding to a monaural signal; the two or more hierarchical enhancement layers including a highest available hierarchical enhancement layer; each of the two or more hierarchical enhancement layers including a portion of the enhancement side information including parameters for improving a basic reconstructed sound representation obtained from data included in each layer and any layer lower than the respective layer, the bitstream further including a parameter CodedVVecLength; the method further comprising: determining, by the processor, that the parameter CodedVVecLength is equal to 1, where CodedVVecLength being equal to 1 indicates that at least some components of the vector corresponding to the compressed HOA representation are implicitly set to 0 and need not be processed by the decoder; decoding, by the decoder, the compressed HOA representation based on the base side information associated with the base layer, and based on the portion of the enhancement side information associated with the highest available hierarchical enhancement layer, and not based on any second portion of the enhancement side information associated with any other layer of the two or more hierarchical enhancement layers; A method. Claim 2 An apparatus for decoding a compressed higher-order ambisonics (HOA) sound representation of sound or sound field, the apparatus comprising: A receiver that receives a bitstream including the compressed HOA representation, the bitstream including a plurality of hierarchical layers including a base layer and two or more hierarchical enhancement layers, the bitstream including at least a data payload corresponding to the plurality of hierarchical layers, the bitstream further including a bitstream that includes base side information associated with the base layer and enhancement side information associated with the two or more hierarchical enhancement layers, including a receiver that receives the bitstream, At least one of the plurality of hierarchical layers has assigned thereto a component of the compressed HOA representation of the sound or sound field, and the component of the basic compressed sound representation corresponds to a monaural signal, The two or more hierarchical enhancement layers include the highest available hierarchical enhancement layer, Each of the two or more hierarchical enhancement layers includes a portion of the enhancement side information that includes parameters for improving a basic reconstructed sound representation obtained from data included in each layer and any layer lower than that layer, the bitstream further including a parameter CodedVVecLength, The apparatus further comprises A processor for determining that the parameter CodedVVecLength is equal to 1, where CodedVVecLength being equal to 1 indicates that at least some components of the vector corresponding to the compressed HOA representation are implicitly set to 0 and need not be processed by a decoder; A decoder for decoding the compressed HOA representation based on the basic side information associated with the base layer, and based on the portion of the enhancement side information associated with the highest available hierarchical enhancement layer, and not based on any second portion of the enhancement side information associated with any other layer of the two or more hierarchical enhancement layers, Apparatus.

Citation Information

Patent Citations

  • Method for compressing a Higher Order Ambisonics (HOA) signal, method for decompressing a compressed HOA signal, apparatus for compressing a HOA signal, and apparatus for decompressing a compressed HOA signal

    EP2922057A1

  • Speech encoding device, speech decoding device, and method thereof

    JP2006072026A

  • A method and apparatus for exploring and reproducing a hierarchical bitstream with a layered structure including a base layer and at least one improved layer.

    JP2013535023A

  • Speech encoding device, speech decoding device, speech encoding method, and speech decoding method

    WO2010103854A2