Encoding of layered structures for compressed sound or sound field representation
The method addresses inefficiencies in layered encoding by subdividing sound representations into hierarchical layers and optimizing side information assignment, ensuring efficient decoding and reduced bandwidth usage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2024-10-23
- Publication Date
- 2026-06-08
AI Technical Summary
Existing layered encoding schemes for compressed sound or sound field representations, such as Higher-Order Ambisonics (HOA), struggle to efficiently adapt to time-varying transmission conditions, leading to unwanted signal loss and inefficiencies in error protection.
A method for encoding a layered structure of compressed sound representations, where components are subdivided into hierarchical layers with base and enhancement layers, and side information is assigned and decomposed to ensure efficient decoding and reconstruction even with partial layer reception, reducing bandwidth requirements.
Ensures optimal sound quality and reduced bandwidth by allowing decoding with only a single layer's payload, even if higher layers are not received, while maintaining efficient error protection and minimizing redundant information.
Smart Images

Figure 0007871348000011 
Figure 0007871348000012 
Figure 0007871348000013
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims priority to European Patent Application No. 15306589.1 and No. 15306653.5, both filed on October 15, 2015, and to U.S. Patent Applications No. 62 / 361,461 and No. 62 / 361,416, both filed on October 15, 2015. The contents of these applications are incorporated herein by reference in their entirety.
[0002] Technical field This paper relates to a method and apparatus for layered audio coding. In particular, this paper relates to a method and apparatus for layered audio coding of compressed sound (or sound field) representations, such as Higher-Order Ambisonics (HOA) sound (or sound field) representations. [Background technology]
[0003] For streaming sound (or sound field) representations through transmission channels with time-varying conditions, layered encoding is a means of adapting the quality of the received sound representation to the transmission conditions and, in particular, avoiding unwanted signal loss.
[0004] For layered encoding, the sound (or sound field) representation is typically subdivided into a relatively small, high-priority base layer and additional enhancement layers with decrementing priority and arbitrary size. Each enhancement layer is typically assumed to contain incremental information to complement the information of all lower-level layers in order to improve the quality of the sound (or sound field) representation. The amount of error protection for the transmission of each layer is controlled based on their priority. In particular, the base layer is given high error protection, which is reasonable and acceptable given its small size.
[0005] However, there is a need for layered encoding schemes for compressed representations (or extended versions thereof) of special types of sounds or sound fields, such as compressed HOA sound or sound field representations. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] ISO / IEC JTC1 / SC29 / WG11 23008-3:2015(E), Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, February 2015 [Non-Patent Document 2] ISO / IEC JTC1 / SC29 / WG11 23008-3:2015 / PDAM3, Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, AMENDMENT 3: MPEG-H 3D Audio Phase 2, July 2015 [Overview of the project] [Problems that the invention aims to solve]
[0007] This paper addresses the above-mentioned problem. In particular, it describes a method and encoder / decoder for layered encoding of compressed sound or sound field representations. [Means for solving the problem]
[0008] In one aspect, a method for encoding a layered structure of a compressed sound representation of a sound or sound field is described. The compressed sound representation may include a basic compressed sound representation containing multiple components. The multiple components may be complementary components. The compressed sound representation may further include basic side information for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The compressed sound representation may further include enhancement side information containing parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The method may include subdividing (e.g., grouping) the multiple components into multiple component groups. The method may further include assigning (e.g., adding) each of the multiple groups to an individual of a multiple hierarchical layer. The assignment may indicate a correspondence between the individual group and the layer. The components assigned to each layer may be said to be included in that layer. The number of groups may correspond to (e.g., equal to) the number of layers. The multiple layers may include a basic layer and one or more hierarchical enhancement layers. The aforementioned hierarchical layers may be ordered from the base layer, through the first improvement layer, the second improvement layer, and so on, up to the overall highest improvement layer (the overall top layer). The method may further include adding base side information to the base layer (for example, including the base side information in the base layer for transmission or storage purposes, or assigning the base side information to the base layer). The method may further include determining multiple parts of the improvement side information from the improvement side information. The method may further include assigning (e.g., adding) each of the aforementioned multiple parts of the improvement side information to each of the layers. Each part of the improvement side information may include parameters for improving the reconstructed (e.g., decompressed) sound representation obtained from data contained in (e.g., assigned to or added to) the respective layer and any lower layers.Layered encoding may be performed for transmission through a transmission channel or for storage on a suitable storage medium such as a CD, DVD, or Blu-ray Disc (trademark).
[0009] The proposed method, configured as described above, enables the efficient application of layered encoding to a compressed sound representation containing multiple components as well as basic and enhanced side information (e.g., independent basic and enhanced side information) having the properties described above. In particular, the proposed method includes, in each layer, suitable side information for reconstructing the reconstructed sound representation from components contained in any layer up to the layer in question. Here, layers up to the layer in question are understood to include, for example, the basic layer, the first enhanced layer, the second enhanced layer, etc., up to the layer in question. Thus, regardless of the actual best available layer (e.g., the layer below the lowest layer that has not yet been effectively received; all layers below the best available layer and the best available layer itself are effectively received), the decoder is enabled to improve or enhance the reconstructed sound representation even if it differs from the complete (e.g., full) sound representation. In particular, in order to improve or enhance the reconstructed sound representation that can be obtained based on all components contained in the layers up to the actual best available layer, regardless of the actual best available layer, it is sufficient for the decoder to decode the payload of enhanced side information for only a single layer (i.e., the best available layer). In other words, for each time interval (e.g., frame), only a single payload of improved-side information may need to be decoded. On the other hand, the proposed method allows for the full benefit of reduced bandwidth requirements that can be achieved when applying layered coding.
[0010] In some embodiments, the components of the basic compressed sound representation may correspond to a monaural signal (e.g., a transport signal or a monaural transport signal). The monaural signal may represent either a dominant sound signal or a coefficient sequence of the HOA representation. The monaural signal may be quantized.
[0011] In various embodiments, the basic side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components individually and independently of the other components. For example, the basic side information may represent side information related to an individual monaural signal, independently of other monaural signals. Therefore, the basic side information may be referred to as independent basic side information.
[0012] In some embodiments, the improved side information may represent improved side information. The improved side information may include predictive parameters for the basic compressed sound representation to improve (e.g., enhance) the basic compressed sound representation and the basic reconstructed sound representation obtained from the basic side information.
[0013] In some embodiments, the method may further include generating transport streams for transmitting the data of the multiple layers (for example, data assigned to or added to each layer, or otherwise included in each layer). The base layer may have the highest transmission priority, and the hierarchical higher layers may have a decreasing transmission priority. That is, the transmission priority may decrease from the base layer to the first higher layer, from the first higher layer to the second higher layer, and so on. The amount of error protection for the transmission of the data of the multiple layers may be controlled according to the transmission priority of each layer. This reduces the overall required bandwidth by not applying excessive error protection to higher layers, while ensuring that at least some lower layers are transmitted reliably.
[0014] In some embodiments, the method may further include generating a transport layer packet for each of the plurality of layers, containing the data for that layer. For example, for each time interval (e.g., frame), a transport layer packet may be generated for each of the plurality of layers.
[0015] In some embodiments, the compressed sound representation may further include additional basic side information for decoding the basic compressed sound representation into a basic reconstructed sound representation. The additional basic side information may include information specifying the decoding of one or more of the plurality of components, depending on the other components. The method may further include decomposing the additional basic side information into a plurality of parts of the additional basic side information. The method may further include adding those parts of the additional basic side information to the basic layer (for example, including those parts of the additional basic side information in the basic layer for transmission or storage, or allocating those parts of the additional basic side information to the basic layer). Each part of the additional basic side information may correspond to a layer and may include information specifying the decoding of one or more components assigned to that layer, depending on (only) the other components assigned to that layer and any layers lower than that layer. That is, each part of the additional basic side information specifies the components in the layer to which that part of the additional basic side information corresponds, without referring to any other components assigned to layers higher than that layer.
[0016] The proposed method, thus constructed, avoids fragmentation of additional basic side information by adding all parts to the basic layer. In other words, all parts of the additional basic side information are included in the basic layer. The decomposition of the additional basic side information ensures that for each layer, a portion of the additional basic side information is available without requiring knowledge of the components of higher layers. Thus, regardless of the actual best usable layer, it is sufficient for the decoder to decode the additional basic side information contained in the layers up to the best usable layer.
[0017] In some embodiments, the additional basic side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components, depending on the other components. For example, the additional basic side information may represent side information related to individual mono signals, depending on other mono signals. Therefore, the additional basic side information may be referred to as dependent basic side information.
[0018] In some embodiments, the compressed sound representation may be processed over a series of time intervals, for example, time intervals of equal size. The series of time intervals may be frames. Thus, the method can operate on a frame basis. That is, the compressed sound representation may be encoded frame by frame. The compressed sound representation may be available for each successive time interval (for example, for each time frame). That is, the compression operation from which the compressed sound representation is obtained may operate on a frame basis.
[0019] In some embodiments, the method may further include generating configuration information for each layer that indicates the basic compressed sound representation components assigned to that layer. In this way, the decoder can easily access the information necessary for decoding without unnecessary parsing through the received data payload.
[0020] Another aspect describes a method for encoding a layered structure of a compressed sound representation of a sound or sound field. The compressed sound representation may include a basic compressed sound representation containing multiple components. The multiple components may be complementary components. The compressed sound representation may further include basic side information (e.g., independent basic side information) and third information (e.g., dependent basic side information) for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The basic side information may include information that specifies the decoding of one or more of the multiple components individually and independently of the other components. The additional basic side information may include information that specifies the decoding of one or more of the multiple components depending on each of the other components. The method may include subdividing the multiple components into multiple component groups (e.g., grouping). The method may further include assigning each of the multiple groups to an individual of a multiple hierarchical layer (e.g., adding). The assignment may indicate a correspondence between the individual group and the layer. The components assigned to each layer may be said to be included in that layer. The number of groups may correspond to the number of layers (for example, they may be equal). The plurality of layers may include a base layer and one or more hierarchical improvement layers. The method may further include adding base side information to the base layer (for example, including the base side information in the base layer for the purpose of transmission or storage, or assigning the base side information to the base layer). The method may further include decomposing additional base side information into multiple parts of the additional base side information and adding those parts of the additional base side information to the base layer (for example, including those parts of the additional base side information in the base layer for transmission or storage, or assigning those parts of the additional base side information to the base layer).Each part of the additional basic side information may correspond to a layer and may include information that specifies the decoding of one or more components assigned to each layer, depending on the other components assigned to each layer and any layers below each layer.
[0021] The proposed method, thus configured, ensures that for each layer, appropriate additional basic side information is available to decode the components contained in any layer up to that layer without requiring effective reception or decoding (or generally knowledge) of higher layers. In the case of a compressed HOA representation, the proposed method ensures that, in vector coding mode, a suitable V-vector is available for all components belonging to the layers up to the highest available layer. In particular, the proposed method eliminates cases where elements of the V-vector corresponding to components in higher layers are not explicitly signaled. Thus, the information contained in the layers up to the highest available layer is sufficient to decode (e.g., decompress) any components belonging to the layers up to the highest available layer. This ensures proper decompression of the reconstructed HOA representation for the lower layers, even if the higher layers have not been effectively received by the decoder. On the other hand, the proposed method allows for the full benefit of the reduced bandwidth requirement that can be achieved when applying layer-configured coding.
[0022] This side embodiment may be related to the side embodiment described above.
[0023] Another aspect describes a method for decoding a compressed sound representation of a sound or sound field. The compressed sound representation may be encoded in a plurality of hierarchical layers. The plurality of hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The plurality of layers may be assigned components of the basic compressed sound representation of the sound or sound field. In other words, the plurality of layers may include components of basic compressed side information. These components may be assigned to each layer in each component group. The plurality of components may be complementary components. The base layer may include basic side information for decoding the basic compressed sound representation. Each layer may include a portion of enhancement side information that includes parameters for improving the basic reconstructed sound representation obtained from the data contained in each layer and any layers lower than each layer. The method may include receiving data payloads corresponding to each of the plurality of hierarchical layers. The method may further include determining a first layer index indicating the best available layer among the plurality of layers to be used to decode the basic compressed sound representation into the basic reconstructed sound representation of the sound or sound field. The method may further include obtaining the basic reconstructed sound representation using the basic side information from the components assigned to the best available layer and any layers lower than the best available layer. The method may further include determining a second layer index indicating which parts of the enhancement side information should be used to improve (e.g., enhance) the basic reconstructed sound representation. The method may further include obtaining a reconstructed sound representation of the sound or sound field from the basic reconstructed sound representation by referring to the second layer index.
[0024] The method thus constructed and proposed ensures that the reconstructed sound representation has optimal quality by making maximum use of available (e.g., effectively received) information.
[0025] In some embodiments, the components of the basic compressed sound representation may correspond to a monaural signal (e.g., a monaural transport signal). The monaural signal may represent either a dominant sound signal or a coefficient sequence of the HOA representation. The monaural signal may be quantized.
[0026] In some embodiments, the basic side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components individually and independently of the other components. For example, the basic side information may represent side information related to an individual monaural signal, independently of other monaural signals. Therefore, the basic side information may be referred to as independent basic side information.
[0027] In some embodiments, the improved side information may represent improved side information. The improved side information may include predictive parameters for the basic compressed sound representation to improve (e.g., enhance) the basic compressed sound representation and the basic reconstructed sound representation obtained from the basic side information.
[0028] In some embodiments, the method may further include determining, for each layer, whether each layer has been effectively received. The method may further include determining the first layer index as the layer index of the layer immediately below the lowest layer that was not effectively received.
[0029] In some embodiments, determining the second layer index may involve determining that the second layer index is equal to the first layer index, or determining an index value as the second layer index that indicates that no enhancement-side information is used when obtaining the reconstructed sound representation. In the latter case, the reconstructed sound representation may be equal to the basic reconstructed sound representation.
[0030] In some embodiments, the data payload may be received and processed over a series of time intervals, for example, time intervals of equal size. The series of time intervals may be frames. Thus, the method can operate on a frame basis. The method may further determine the second layer index to be equal to the first layer index if the compressed sound representations over those time intervals can be decoded independently of each other.
[0031] In some embodiments, the data payload may be received and processed over a series of time intervals, for example, time intervals of equal size. The series of time intervals may be frames. Thus, the method can operate on a frame basis. The method may further include determining, for a given time interval among the series of time intervals, whether each layer has been effectively received if the compressed sound representations for those time intervals cannot be decoded independently of each other. The method may further include determining the first layer index for the given time interval as the smaller of the first layer index of the time interval preceding the given time interval and the layer index of the layer immediately below the lowest layer that was not effectively received.
[0032] In some embodiments, the method may include determining whether a first layer index for a given time interval is equal to a first layer index for a preceding time interval if the compressed sound representations for a series of time intervals cannot be decoded independently of each other. The method may further include determining a second layer index for a given time interval to be equal to the first layer index for a given time interval if the first layer index for a given time interval is equal to the first layer index for a preceding time interval. The method may further include determining a second layer index as an index value indicating that no enhancement-side information is used when obtaining the reconstructed sound representation if the first layer index for a given time interval is not equal to the first layer index for a preceding time interval.
[0033] In some embodiments, the basic layer may include at least one portion of additional basic side information, each corresponding to a layer, which includes information specifying the decoding of one or more components among those assigned to each layer, depending on other components assigned to each layer and any layers lower than each layer. The method may further include decoding the portion of the additional basic side information by referring to the components assigned to each layer and any layers lower than each layer. The method may further include correcting the portion of the additional basic side information by referring to the components assigned to the best available layer and any layers between the best available layer and each of the layers. The basic reconstructed sound representation may be obtained using the basic side information and corrected portions of the additional basic side information obtained from the portions of the additional basic side information corresponding to the layers up to the best available layer, with respect to the components assigned to the best available layer and any layers lower than the best available layer.
[0034] In some embodiments, the additional basic side information may include information that specifies the decoding (e.g., decompression) of one or more of the plurality of components, depending on the other components. For example, the additional basic side information may express side information related to individual monaural signals, depending on other monaural signals. Thus, the additional basic side information may be referred to as dependent basic side information.
[0035] Another aspect describes a method for decoding a compressed sound representation of a sound or sound field. The compressed sound representation may be encoded in a plurality of hierarchical layers. The plurality of hierarchical layers may include a base layer and one or more hierarchical higher layers. The plurality of layers may be assigned components of the basic compressed sound representation of the sound or sound field. In other words, the plurality of layers may include components of basic compressed side information. These components may be assigned to each layer in each component group. The plurality of components may be complementary components. The base layer may include basic side information for decoding the basic compressed sound representation. The base layer may further include at least one additional portion of basic side information, which corresponds to each layer and includes information specifying the decoding of one or more components among the components assigned to each layer, depending on other components assigned to each layer and any layers lower than each layer. The method may further include receiving data payloads corresponding to each of the plurality of hierarchical layers. The method may further include determining a first layer index indicating the best available layer among the plurality of layers to be used to decode the basic compressed sound representation to obtain the basic reconstructed sound representation of the sound or sound field. The method may further include decoding each portion of the additional basic side information by referring to the components assigned to its respective layer and any layers lower than each of those layers. The method may further include correcting each portion of the additional basic side information by referring to the components assigned to the best available layer and any layers between the best available layer and each of those layers. The basic reconstructed sound representation may be obtained using the basic side information and corrected portions of the additional basic side information obtained from the portions of the additional basic side information corresponding to the layers up to the best available layer, using the components assigned to the best available layer and any layers lower than the best available layer.The method may further include determining a second layer index that is equal to the first layer index or indicates the omission of improved side information during decoding.
[0036] The proposed method, thus structured, ensures that the additional basic side information ultimately used to decode the basic compressed sound representation does not contain redundant elements, thereby making the actual decoding of the basic compressed sound representation more efficient.
[0037] Embodiments of this aspect may relate to the embodiments of the aspect described above.
[0038] Another aspect describes an encoder for encoding a layered structure of compressed sound representations of a sound or sound field. The compressed sound representation may include a basic compressed sound representation containing multiple components, which may be complementary components. The compressed sound representation may further include basic side information for decoding the basic compressed sound representation into a basic reconstructed sound representation of the sound or sound field. The compressed sound representation may further include enhancement side information containing parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The encoder may include a processor configured to perform some or all of the method steps of the methods based on the first aspect and the second aspect mentioned above.
[0039] Another aspect describes a decoder for decoding a compressed sound representation of a sound or sound field. The compressed sound representation may be encoded in multiple hierarchical layers. The multiple hierarchical layers may include a base layer and one or more hierarchical enhancement layers. The multiple layers may be assigned components of the basic compressed sound representation of the sound or sound field. In other words, the multiple layers may include components of basic compressed side information. These components may be assigned to each layer in each component group. The multiple components may be complementary components. The base layer may include basic side information for decoding the basic compressed sound representation. Each layer may include a portion of enhancement side information that includes parameters for improving (e.g., enhancing) the basic reconstructed sound representation obtained from the data contained in each layer and any layers lower than each layer. The decoder may include a processor configured to perform some or all of the method steps of the methods based on the third aspect and the fourth aspect mentioned above.
[0040] In other words, methods, apparatus, and systems are directed toward decoding a compressed higher-order ambisonics (HOA) sound representation of sound or sound field. An apparatus may have a receiver configured to receive a bitstream containing the compressed HOA representation corresponding to a plurality of hierarchical layers, including a base layer and one or more hierarchical enhancement layers, or a method may perform such receiving. The plurality of layers are assigned components of the basic compressed sound representation of sound or sound field, and these components are assigned to each layer in their respective component groups. An apparatus may have a decoder configured to decode the compressed HOA representation based on base side information associated with the base layer and enhancement side information associated with one or more hierarchical enhancement layers, or a method may perform such decoding. The base side information may include base independent side information relating to a first set of separate monaural signals that are decoded independently of other monaural signals. Each of the one or more hierarchical enhancement layers may include a portion of the enhancement side information containing parameters for improving the basic reconstructed sound representation obtained from the data contained in each layer and any layers lower than each of them.
[0041] The basic independent side information may indicate that the first individual monaural signal represents a directional signal with a certain incident direction. The basic side information may further include basic dependent side information relating to a second individual monaural signal that is decoded depending on the other monaural signal. The basic dependent side information may include a vector-based signal that is directionally distributed within the sound field, where the directional distribution is specified by a vector. The components of the vector are set to 0 and are not part of the compressed vector representation.
[0042] The basic compressed sound representation components may correspond to a monaural signal representing either a dominant sound signal or a coefficient sequence of the HOA representation. The bitstream contains data payloads corresponding to each of the multiple hierarchical layers. The improved side information may include parameters related to at least one of spatial prediction, subband directional signal synthesis, and parametric ambient sound replication. The improved side information may include information that allows for the prediction of missing portions of sound or sound field from the directional signal. Furthermore, for each layer, it may be determined whether the respective layer was effectively received, and the layer index of the layer immediately below the lowest layer that was not effectively received may be determined.
[0043] From another perspective, a software program is described. This software program is adapted for execution on a processor and may be adapted to perform some or all of the method steps outlined in this paper when executed on a computing device.
[0044] Another aspect describes a storage medium, which may contain a software program adapted for execution on a processor and adapted to perform some or all of the method steps outlined in this paper when executed on a computing device.
[0045] Those skilled in the art will understand that any statement made with respect to any aspect or embodiment described above also applies to the other aspects or embodiments. For the sake of brevity, these statements have been omitted for each individual aspect or embodiment.
[0046] Methods and apparatuses including preferred embodiments outlined herein may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and apparatuses outlined herein may be combined in any way. In particular, features of the claims may be combined with other features in any way.
[0047] The steps and apparatus features can be replaced in many ways. In particular, as those skilled in the art will understand, the details of the disclosed method can be implemented as apparatus adapted to perform some or all of the steps of the method, and vice versa. [Brief explanation of the drawing]
[0048] The present invention is described below in an illustrative manner with reference to the accompanying drawings. [Figure 1] This flowchart shows an example of a layered encoding method based on embodiments of the present disclosure. [Figure 2] This is a schematic block diagram showing an example of an encoder stage based on an embodiment of the present disclosure. [Figure 3] This flowchart shows an example of a method for decoding a compressed sound representation of a sound or sound field encoded in multiple hierarchical layers, according to embodiments of the present disclosure. [Figure 4A] This is a schematic block diagram showing an example of a decoder stage based on an embodiment of the present disclosure. [Figure 4B] This is a schematic block diagram showing an example of a decoder stage based on an embodiment of the present disclosure. [Figure 5] This is a schematic block diagram illustrating an example of a hardware implementation of an encoder based on the embodiments of this disclosure. [Figure 6] This is a schematic block diagram illustrating an example of a hardware implementation of a decoder based on the embodiments of this disclosure. [Modes for carrying out the invention]
[0049] First, we will describe the compressed sound (or sound field) representation (hereinafter referred to as the compressed sound representation for brevity) to which the methods and encoder / decoder based on this disclosure are applicable. Generally, a fully compressed sound (or sound field) representation (hereinafter referred to as the fully compressed sound representation for brevity) may include (for example, consist of) the following three components: the basic compressed sound (or sound field) representation (hereinafter referred to as the basic compressed sound representation for brevity), basic side information, and enhanced side information.
[0050] The basic compressed sound representation itself contains several components (e.g., complementary components) (e.g., consisting of ). The basic compressed sound representation may constitute the overwhelmingly largest proportion of the complete compressed sound representation. The basic compressed sound representation may consist of a monaural transport signal representing the dominant sound signal or the coefficient sequence of the original HOA representation.
[0051] Basic side information is required to decode the basic compressed sound representation and can be assumed to be much smaller in size than the basic compressed sound representation. Furthermore, it may consist mostly of separate parts, each specifying the decompression of only one specific component of the basic compressed sound representation. Basic side information may consist of a first part, known as independent basic side information, and a second part, known as additional basic side information.
[0052] Both the first and second parts, namely the independent basic side information and the additional basic side information, can specify the decompression of certain components of the basic compressed sound representation. The second part is optional and may be omitted. In this case, the compressed sound representation is sometimes said to include the first part (e.g., the basic side information).
[0053] The first part (for example, basic side information) may include side information that describes the individual (complementary) components of the basic compressed sound representation independently of other (complementary) components. In particular, the first part (for example, basic side information) may specify the decoding of one or more of the plurality of components individually and independently of the other components. Thus, the first part may be called independent basic side information.
[0054] The second (optional) part, also known as additional basic side information, can describe the individual (complementary) components of the basic compressed sound representation in dependence on other (complementary) components. This second part may also be called dependent basic side information. In particular, the dependence may have the following attributes: The dependent basic side information for each individual (complementary) component of the basic compressed sound representation achieves its maximum range when the basic compressed sound representation does not contain any other kind of (complementary) component. When certain additional (complementary) components are added to a basic compressed sound representation, the dependent basic side information for each individual (complementary) component in question becomes a subset of the original dependent basic side information, thereby reducing its size.
[0055] The enhancement side information is also optional. This can be used to improve or enhance the basic compressed sound representation (for example, parametrically). Its size is also assumed to be much smaller than the size of the basic compressed sound representation.
[0056] Thus, in various embodiments, the compressed sound representation may include a basic compressed sound representation comprising a plurality of components, basic side information for decoding (e.g., decompressing) the basic compressed sound representation to obtain a basic reconstructed sound representation of the sound or sound field, and enhancement side information including parameters for improving or enhancing (e.g., parametrically improving or enhancing) the basic compressed sound representation. The compressed sound representation may further include additional basic side information for decoding (e.g., decompressing) the basic compressed sound representation to obtain the basic reconstructed sound representation, which may include information specifying the decoding of one or more of the plurality of components depending on the other components.
[0057] One example of such a fully compressed sound representation is given by the compressed higher-order ambisonics (HOA) sound field representation specified in Chapter 12 and Annex C.5 of the preliminary version of the MPEG-H 3D audio standard (Non-Patent Literature 1). That is, the compressed sound representation may correspond to a compressed HOA sound (or sound field) representation of sound or sound field.
[0058] In this example, the basic compressed sound field representation (basic compressed sound representation) may contain several components (for example, it may be identified by several components). These components may be monaural signals (for example, they may correspond to monaural signals). These monaural signals may be quantized monaural signals. These monaural signals may represent either a dominant sound signal or a coefficient sequence of ambient sound HOA sound field components.
[0059] The basic side information can, in particular, describe how each of these monaural signals contributes spatially to the sound field. For example, the basic side information may specify the dominant sound signal as a purely directional signal, i.e., a general plane wave with a certain direction of incidence. Alternatively, the basic side information may specify the monaural signal as a coefficient sequence of the original HOA representation with a certain index. The basic side information may further be separated into a first part and a second part as described above.
[0060] The first part is side information (e.g., independent fundamental side information) relating to a specific individual monaural signal. This independent fundamental side information is independent of the existence of other monaural signals. Such side information may, for example, specify a monaural signal representing a directional signal with a certain incident direction (e.g., meaning a general plane wave). Alternatively, the monaural signal may be specified as a coefficient sequence of the original HOA representation with a certain index. The first part may be called independent fundamental side information. In general, the first part (e.g., fundamental side information) can specify the decoding of one or more monaural signals from the plurality of monaural signals individually and independently of other monaural signals.
[0061] The second part is side information (e.g., additional basic side information) related to a specific individual monaural signal. This side information depends on the presence of other monaural signals. Such side information may be used, for example, when the monaural signal is specified to be a vector-based signal (see, for example, section 12.4.2.4.4 of Non-Patent Literature 1). These signals are distributed directionally in the sound field, and the directional distribution may be specified by a vector. In certain modes (i.e., CodedVVecLength=1), certain components of this vector are implicitly set to 0 and are not part of the compressed vector representation. These components are those that have indices equal to those in the coefficient sequence of the original HOA representation that are part of the basic compressed sound representation. In other words, when individual components of the vector are encoded, their total number depends on the basic compressed sound representation. In particular, the total number depends on which coefficient sequence the original HOA representation contains.
[0062] If the coefficient sequence of the original HOA representation is not included in the basic compressed sound representation, the dependent basic side information for each vector-based signal consists of all vector components and has its maximum size. If a coefficient sequence of the original HOA representation with a certain index is added to the basic compressed sound representation, the vector components with those indices are removed from the side information for each vector-based signal, thereby reducing the size of the dependent basic side information for the vector-based signal.
[0063] The improved side information (for example, improved side information) may include parameters related to (broadband) spatial prediction (see Section 12.4.2.4.3 of Non-Patent Document 1) and / or parameters related to subband directional signal synthesis and parametric ambient sound reproduction.
[0064] (Broadband) Parameters related to spatial prediction can be used to predict (linearly) missing parts of the sound field from the directional signal.
[0065] Subband directional signal synthesis and parametric ambient sound replication are compression tools recently introduced in revisions to the MPEG-H 3D audio standard (see Section 1 of Non-Patent Literature 2). These tools allow for frequency-dependent parametric predictions of additional monaural signals that should be spatially distributed to complement a spatially incomplete or missing compressed HOA representation. The predictions may be based on a coefficient sequence of the underlying compressed sound representation.
[0066] It is important to note that the aforementioned complementary contributions to the sound field are represented within the compressed HOA representation not by additional quantized signals, but by additional side information of a comparatively much smaller size. Therefore, the two encoding tools described above are particularly suitable for compressing HOA representations at low data rates.
[0067] A second example of a compressed representation of one or more monaural signals having the structure described above may include encoded spectral information for separate frequency bands up to a certain upper frequency, which can be considered a basic compressed representation; basic side information specifying the encoded spectral information (for example, by the number and width of the encoded frequency bands); and enhanced side information including parameters for spectral band duplication (SBR) (for example, consisting of ). The parameters of the enhanced side information describe how to parametrically reconstruct spectral information for higher frequency bands not considered in the basic compressed representation from the basic compressed representation.
[0068] This disclosure proposes a method for encoding a layered structure of a fully compressed sound (or sound field) representation having the structure described above.
[0069] Compression may be frame-based in the sense that it provides a compressed representation (e.g., in the form of data packets or equivalently frame payloads) for a series of time intervals. The time intervals may have equal or different sizes. These data packets may be assumed to contain, in addition to the data of the actual compressed representation, a validity flag and a value indicating its size. In the following, without intent to limit, compression is assumed to be frame-based. Further, unless otherwise specified and without intent to limit, the focus is on the handling of a single frame. Thus, the frame index is omitted.
[0070] Each frame payload of the considered complete compressed sound (or sound field) representation is assumed to contain J data packets (or frame payloads), each data packet being for one component of the basic compressed sound representation denoted as BSRC j , j = 1, …, J. Further, each data packet is assumed to contain a packet with independent basic side information denoted as BSI I . BSI I specifies certain components BSRC j of the basic compressed sound representation independently of the other components. Optionally, each data packet is further assumed to contain a packet with dependent basic side information (additional basic side information) denoted as BSI D . BSI D specifies certain components BSRC j of the basic compressed sound representation depending on the other components.
[0071] The information contained within two data packets BSI I and BSI D may optionally be grouped into a single data packet BSI of basic side information. The single data packet BSI, among other things, each being for one specific component BSRC jIt may also be said that it contains J parts that specify the following. Each of these parts may also be said to contain a part of independent-side information and optionally a part of dependent-side information.
[0072] Ultimately, each data packet may contain an enhancement side information payload, denoted as ESI, which describes how the reconstructed sound (or sound field) is improved or enhanced from the complete basic compressed sound representation.
[0073] The proposed solution for layered encoding addresses the necessary steps to enable both a compression unit, which includes packing data packets for transmission, and a receiver and decompression unit. Each unit is described in detail below.
[0074] First, we will discuss compression and packing (for example, for transmission). In particular, we will discuss the components and elements of the fully compressed sound (or sound field) representation in the case of layered encoding.
[0075] Figure 1 schematically shows a flowchart of an example of a method for compression and packing (e.g., an encoding method for a compressed sound representation of a sound or sound field, or an encoding method for a layered structure). The assignment (e.g., allocation) of individual payloads to the base layer and (M-1) higher layers may be achieved by a transport layer packer. Figure 2 schematically shows a block diagram of an example of the allocation / distribution of individual payloads.
[0076] As shown above, the fully compressed sound representation 2100 may relate to, for example, a compressed HOA representation that includes a basic compressed sound representation. The fully compressed sound representation 2100 may include a plurality of components (e.g., monaural signals) 2110-1, ..., 2110-J, independent basic side information (basic side information) 2120, optional enhancement side information (enhancement side information) 2140, and optional dependent basic side information (additional basic side information) 2130. The basic side information 2120 may be information for decoding the basic compressed sound representation to obtain a basic reconstructed sound representation of the sound or sound field. The basic side information 2120 may include information for specifying the decoding of one or more components (e.g., monaural signals) individually and independently of other components. The enhancement side information 2140 may include parameters for improving (e.g., enhancing) the basic reconstructed sound representation. The additional basic side information 2130 may be (further) information for decoding the basic compressed sound representation into the basic reconstructed sound representation, and may include information that specifies the decoding of one or more of the plurality of components individually, each depending on the other components.
[0077] Figure 2 illustrates the underlying assumption that there are multiple hierarchical layers, including one base layer and one or more (hierarchical) improvement layers. For example, there may be a total of M layers, i.e., one base layer and M-1 improvement layers. The multiple hierarchical layers have a sequentially increasing layer index. The lowest layer index (e.g., layer index 1) corresponds to the base layer. Furthermore, it is understood that the layers are ordered from the base layer through the improvement layers to the overall highest improvement layer (i.e., the overall top layer).
[0078] The proposed method may be carried out on a frame basis (i.e., on a frame-by-frame basis). In particular, the compressed sound representation 2100 may be compressed over a series of time intervals, for example, time intervals of equal size. Each time interval may correspond to a frame. The following steps may be carried out for each of the series of time intervals (e.g., a frame).
[0079] In S1010 of Figure 1, the plurality of components 2110 are subdivided into a plurality of component groups. Each of the plurality of groups is then assigned (e.g., added to or allocated to) a corresponding hierarchical layer, where the number of groups corresponds to the number of layers. For example, the number of groups may be equal to the number of layers, so that each layer may have one group of components. As shown above, the plurality of layers may include a base layer and one or more (e.g., M-1) hierarchical improvement layers.
[0080] In other words, the basic compressed sound representation is subdivided into parts that are assigned to individual layers. Without losing generality, the grouping is M+1 times J. m It can be described by m=0,...,M. Here, J0=1, J M =J+1, and the component BSRC j J m-1 ≤j <J m It is assigned to the mth layer.
[0081] In S1020, the component groups are assigned to their respective layers. In S1030, the basic side information 2120 is added to (for example, assigned to) the basic layer (i.e., the lowest layer among the multiple hierarchical layers).
[0082] In other words, due to its small size, it is proposed to include complete basic side information (basic side information and optional additional basic side information) in the basic layer to avoid unnecessary fragmentation.
[0083] If the compressed sound representation under consideration includes dependent basic side information (additional basic side information), the method may further (not shown in Figure 1) include decomposing the additional basic side information into a plurality of parts 2130-1, ..., 2130-M of the additional basic side information. These parts of the additional basic side information may then be added to (e.g., assigned to) the basic layer. In other words, these parts of the additional basic side information may be included in the basic layer. Each part of the additional basic side information may correspond to a layer and may include information specifying the decoding of one or more components assigned to each layer, depending on other components assigned to each layer and any layers lower than each layer.
[0084] Thus, the independent basic side information BSI I (Basic side information) 2120 is kept immutable for allocation purposes, while dependent basic side information needs to be treated specially for layered encoding to allow correct decoding at the receiver and to reduce the size of the transmitted dependent basic side information. D,m It is proposed to decompose it into M parts denoted by m=1, ..., M. Here, the m-th part is the component of the basic compressed sound representation BSRC assigned to the m-th layer. j , J m-1 ≤j <J m This includes the dependent basic side information for each of them. This assumes that the optional dependent basic side information exists for the compressed sound representation under consideration. If the dependent side information does not exist for each part, then for the compressed sound representation of each part, BSI D,m It is assumed to be empty. Each part of the dependent base side information BSI D,m This is all the components BSRC that are contained in all layers up to the mth layer (i.e., all layers j=1,...,m). j , 1≦j <J m It is also acceptable to depend on it.
[0085] Independent Basic Side Information Packet (BSI)I If the size is negligibly small, it is reasonable to keep it as a whole and add (allocate) it to the base layer. Optionally, independent base side information can also be decomposed in the same way as dependent base side information, and packet BSI I,m , gives m=1,...,M. This is useful for reducing the size of the base layer by adding (assigning) parts of the independent base side information to the layer that has the corresponding components of the basic compressed sound representation.
[0086] In S1040, several parts 2140-1, ..., 2140-M of the improved side information may be determined. Each part of the improved side information may include parameters for improving (e.g., enhancing) the reconstructed sound representation obtained from the data contained in each layer and any layers lower than each of those layers.
[0087] The reason for performing this step is that, in the case of layered encoding, it is important to recognize that improvement-side information needs to be calculated extra for each layer because the intention is to improve the preliminary decompressed sound (or sound field), but this depends on the available layers for decompression. Specifically, the preliminary decompressed sound (or sound field) for a given best decodeable layer (best available layer) depends on the components contained in that best decodeable layer and any layers below it. Thus, compression is ESI m It is necessary to provide M individual improved-side information data packets (parts of improved-side information) denoted by m=1,...,M. Here, the improved-side information ESI in the m-th data packet. m It is calculated to improve the sound (or sound field) representation obtained from all the data contained in the base layer and the improved layers with an index lower than m (for example, all the data contained in the m-th layer and any layers below the m-th layer).
[0088] In S1050, the plurality of parts 2140-1, ..., 2140-M of the improved side information are assigned (for example, added or allocated) to the plurality of layers. Each of the plurality of parts of the improved side information is assigned to each of the plurality of layers. For example, each of the plurality of layers contains each of the plurality of parts of the improved side information.
[0089] The assignment of basic and / or enhanced side information to each layer may be indicated in the configuration information generated by the encoding method. In other words, the correspondence between the basic and / or enhanced side information and each layer may be indicated in the configuration information. Furthermore, the configuration information may indicate, for each layer, the components of the basic compressed sound representation that are assigned to (e.g., included in) that layer. Additional parts of the basic side information may be included in the basic layer but correspond to layers other than the basic layer.
[0090] In summary, the compression stage provides a frame data packet labeled FRAME, with the following composition:
[0091]
number
[0092]
number
[0093] Individual data packets may then be grouped within a payload. This payload is defined as a special data packet containing, in addition to the actual compressed representation data, a validity flag and a value indicating its size. The use of payloads offers the advantage of allowing simple multiplexing on the receiver side, enabling the discarding of outdated payloads without the need to parse them. One possible grouping is given by: ·Each BSRC j Packets, j=1, ..., J are individual payloads (BP with  ̄) j Assign (for example, allocate) to the (indicated by). • mth Upside-Side Information Data Packet ESI m and the m-th dependent side information data packet BSI D,m One improved payload (EP with  ̄) m Assign (for example, allocate) to m=1, ..., M) as indicated by . • Independent Basic Side Information BSI I Assign this to a separate side information payload (indicated by BSIP with a  ̄).
[0094] Optionally, if the size of the independent basic side information is large, then each of its components, the m-th BSI I,m m=1,…,M is the aforementioned improved payload (EP with  ̄) m It may be assigned (for example, allocated) to (as indicated by ). In this case, the side information payload (BSIP with  ̄) is empty and can be ignored.
[0095] Another option is all dependent base side information data packets (BSI). D,m This involves assigning the aforementioned side information payload (BSIP with  ̄). This is reasonable when the size of the dependent basic side information is small.
[0096] Finally, a frame data packet, denoted as FRAME, may be given, having the following composition:
[0097]
number
[0098] The method may further (not shown in Figure 1) include generating transport layer packets for each of the plurality of layers that contain the data for each layer (for example, components, basic side information, and improved side information for the basic layer, or components and improved side information for one or more improved layers) (for example, a basic layer packet 2200 and M-1 improved layer packets 2300-1, ..., 2300-(M-1)).
[0099] Transport layer packets for different layers may have different transmission priorities. Thus, the method may further (not shown in Figure 1) include generating transport streams for transmitting data from the multiple layers. Here, the base layer has the highest transmission priority, and the hierarchical layers have progressively decreasing transmission priorities. Here, a higher transmission priority corresponds to a greater degree of error protection, and vice versa.
[0100] Unless a prerequisite is required for another stage, the above-mentioned stages may be performed in any order, and the exemplary order shown in Figure 1 is not considered limiting.
[0101] Figure 3 illustrates a method for decoding a compressed sound representation of a sound or sound field for decoding or unpacking. Examples of corresponding receivers and unpacking stages are schematically shown in the block diagrams of Figures 4A and 4B.
[0102] As can be seen from the above, the compressed sound representation may be encoded in the aforementioned hierarchical layers. The aforementioned layers may be assigned components of the basic compressed sound representation (for example, they may contain such components). These components are assigned to each layer in their respective component groups. The basic layer may contain basic side information for decoding the basic compressed sound representation. Each layer may contain one of the aforementioned portions of enhancement side information, which includes parameters for improving the basic reconstructed sound representation obtained from the data contained in each of its layers and any layers below it.
[0103] The proposed method may be carried out on a frame basis (i.e., on a frame-by-frame basis). In particular, the reconstructed representation of the sound or sound field may be generated for a series of time intervals, for example, time intervals of equal size. These time intervals may be, for example, frames. The following steps may be carried out for each of the series of time intervals (e.g., frames).
[0104] In S3010, a data payload (e.g., a transport layer packet) corresponding to the multiple layers is received. The data payload may be received as part of a bitstream containing a compressed HOA representation of sound or sound field corresponding to the multiple hierarchical layers. The hierarchical layers include a base layer and one or more enhancement layers. The multiple layers may be assigned components of the basic compressed sound representation of the sound or sound field. The components are assigned to each layer in each component group.
[0105] Individual layer packets may be multiplexed to provide a received frame packet of a fully compressed audio representation. The received frame packet is
[0106]
number
[0107]
number
[0108] Using the payload, the received frame packet is
[0109]
number
[0110] The received frame packet may then be passed to the decompressor or decoder 4100. If there are no errors in the transmission of the individual layers, at least the improved side information payload (corresponding to, for example, the improved side information portion) is included.
[0111]
number
[0112] In the decompressor 4100, the received frame packets may be multiplexed. For this purpose, information about the size of each payload may be used to avoid unnecessary parsing through the data of individual payloads.
[0113] In S3020, a first layer index is determined to represent the best layer among the multiple layers (for example, the best usable layer or the best decodeable layer) to be used to decode the basic compressed sound representation to obtain the basic reconstructed sound representation of the sound or sound field.
[0114] Furthermore, in S3020, the value (e.g., layer index) N of the best layer (best available layer) that will be used for decompressing the basic sound representation is the value of the best layer (the best available layer). B This may be selected. The best one actually used for decompressing basic sound representation. improvement The number of layers is N B This is given by -1. Since each layer contains exactly one improved-side information payload (part of the improved-side information), it can be determined whether the layer containing the improved-side information payload is valid (e.g., validly received) or not. Thus, the selection is given by all improved-side information payloads ESI m , m=1,...,M (or correspondingly
[0115]
number
[0116] In S3030, a basic reconstructed sound representation is obtained. The basic reconstructed sound representation may be obtained using basic side information (or generally using basic side information) from the components assigned to the highest available layer indicated by the first layer index and any layers lower than this highest available layer.
[0117] Basic compressed sound representation components BSRC1, ..., BSRC J The payload is a basic side information payload (for example, BSI or BSI) I and BSI D,m , m=1,...,M) (all of them) and value N BIt may also be provided to the basic representation decompression processing unit 4200. The basic representation decompression processing unit 4200 (shown in Figures 4A and 4B) is the lowest N B The number of layers, namely the basic layer and N B - Reconstruct the basic sound (or sound field) representation using only the fundamental compressed sound representation components contained within one improved layer (i.e., the layers up to the layer indicated by the first layer index). Alternatively, the lowest N B Only the payload of the basic compressed sound representation components contained in each layer may be provided to the basic representation decompression processing unit 4200, along with the respective basic side information payload.
[0118] The decompressor 4100 is assumed to know which components of the basic compressed sound (or sound field) representation are contained in each layer from the data packets containing the configuration information. The configuration information is assumed to be transmitted and received before the frame data packets.
[0119] Dependent side information data packet BSI D,m m=1,…,N B and improved side information data packet ESI NE To provide, all improvement payloads have a value of N E and value N B Along with this, it may be input to the partial parser 4400 of the decompressor 4100 (see Figure 4B). The parser may discard all payloads and data packets that are not used for actual decompression. E If the value is equal to 0, all upside information data packets may be assumed to be empty.
[0120] If the base layer contains at least one dependent base side information payload (part of the additional base side information) corresponding to each layer, then each individual dependent base side information payload (e.g., BSI) D,m m=1,…,N BDecoding (part of additional basic side information) may include (i) preliminary decoding of the part of the additional basic side information by referring to components assigned to each of its layers and any layers lower than each of them, and (ii) correction of the part of the additional basic side information by referring to components assigned to the highest available layer and any layers between the highest available layer and each of them (correction). Here, the additional basic side information corresponding to each layer includes information that specifies the decoding of one or more components among the components assigned to each of those layers, depending on other components assigned to each of those layers and any layers lower than them.
[0121] Next, a basic reconstructed sound representation can be obtained (for example, generated) using basic side information from the components assigned to the highest usable layer and any layers lower than the highest usable layer, and corrected parts of additional basic side information obtained from the parts of additional basic side information corresponding to the layers up to the highest usable layer.
[0122] In particular, each payload BSI D,m m=1,…,N B The preliminary decoding of the first J included in the first m layer assumed in the encoding stage is m -1 basic compressed sound representation component BSRC1, ..., BSRC (Jm)-1 It may also involve leveraging the dependency on [something].
[0123] Each payload BSI D,m m=1,…,N B The sequential correction is that the basic sound components are more components than assumed for preliminary decoding in the first N B >The first J contained in the m layer NB -1 basic compressed sound representation component BSRC1, ..., BSRC (JNB)-1It may also be necessary to consider that it will ultimately be reconstructed from. Therefore, correction may be achieved by discarding outdated information. This is possible because of the attribute initially assumed in the dependent basic side information, namely that if some complementary components are added to the basic compressed sound representation, the dependent basic side information for each individual (complementary) component becomes a subset of the original.
[0124] In S3040, a second layer index may be determined. The second layer index may indicate the portion(s) of improvement-side information that should be used to improve (e.g., enhance) the basic reconstructed sound representation.
[0125] In addition to the first layer index, there is an index (second layer index) N of the improved side information payload (the second improved information portion) that should be used for decompression. E The second layer index N may be determined. E The first layer index N is always B It may be equal to 0 or equal to 1. Improvement may always be achieved according to the basic sonic expression obtained from the best available layer, or it may not be achieved at all.
[0126] In S3050, the reconstructed sound representation of the sound or sound field is obtained (for example, generated) from the basic reconstructed sound representation by referring to the second layer index.
[0127] In other words, the reconstructed sound representation is obtained by (parametrically) improving or enhancing the basic reconstructed sound representation, for example by using the enhancement-side information (part of the enhancement-side information) indicated by the second layer index. As will be discussed later, the second layer index may also indicate that no enhancement-side information is used at all at this stage. In this case, the reconstructed sound representation will correspond to the basic reconstructed sound representation.
[0128] For this purpose, the reconstructed basic sound representation is all improved side information payload ESI1, ..., ESI M , basic side information payload (for example, BSI or BSI I and BSI D,m , m=1,…,M) and value N E This is then fed to the improved representation decompression processing unit 4300 (shown in Figures 4A and 4B). The improved representation decompression processing unit 4300 processes the improved side information payload ESI. NE Only the improved side information payload is used, and all other improved side information payloads are discarded to compute the final improved sound (or sound field) representation 2100'. Alternatively, the improved side information payload ESI is used instead of all improved side information payloads. NE Only the improved expression decompression processing unit 4300 may be provided. E If the value is equal to 0, all enhanced side information payloads are discarded (alternatively, no enhanced side information payloads are provided). The reconstructed final enhanced sound representation 2100' is then equal to the reconstructed base sound representation. Enhanced side information payload ESI NE This may be obtained by the partial parser 4400.
[0129] Figure 3 also provides an overview of how the compressed HOA representation is decoded based on the basic side information associated with the basic layer and the improved side information associated with one or more hierarchical improved layers.
[0130] Unless a prerequisite is required for another stage, the above-mentioned stages may be performed in any order, and the exemplary order shown in Figure 3 is not considered limiting.
[0131] Next, we will describe the details of the layer selection (selection of the first and second layer indices) for decompression in stages S3020 and S3040.
[0132] The determination of the first layer index may involve determining, for each layer, whether that layer was validly received. The determination of the first layer index may further involve determining the first layer index as the layer index of the layer immediately below the lowest layer that was not validly received. Whether a layer was validly received may be determined by evaluating whether the upside information payload of that layer was validly received. This may be done by evaluating the validity flag in the upside information payload.
[0133] Determining the second layer index may generally involve determining the second layer index to be equal to the first layer index, or determining an index value (e.g., index value 0) as the second layer index that indicates that no enhancement-side information is used when obtaining the reconstructed sound representation.
[0134] If all frame data packets can be decompressed independently of each other, then the number of the best layer (best available layer) actually used for decompressing the basic sound representation is N. B and index N of the improved side information payload used for decompression E In either case, L may be set to the highest number L of the valid improvement-side information payload. L itself can be determined by evaluating the effectiveness flags within the improvement-side information payload. By leveraging knowledge of the size of each improvement-side information payload, complex parsing through the actual data of the payload can be avoided to determine effectiveness.
[0135] In other words, if the compressed sound representations for a series of time intervals can be decoded independently, the second layer index may be determined to be equal to the first layer index. In this case, the reconstructed basic sound representation can be enhanced based on the enhanced side information payload of the best available layer.
[0136] When differential decompression with inter-frame dependencies is used, decisions from previous frames must also be considered. In differential decompression, independent frame data packets are typically transmitted at regular time intervals to allow decompression to begin from those points in time. In independent frame data packets, the value N B and N E The decision becomes frame-independent and is executed as described above.
[0137] To explain the proposed frame-dependent determination in detail, L(k) is the highest number (e.g., layer index) of the effective enhancement side information payload for the k-th frame, and N is the highest layer number (e.g., layer index) selected and used for decompression of the basic sound representation. B (k) The number of the improved side information payload used for decompression (e.g., layer index) is N E (k) is used to represent it.
[0138] This notation is used to find the best layer number N used for decompressing basic sound representations. B (k) is calculated according to the following formula.
[0139]
number
[0140] That is, when the compressed audio representations for a series of time intervals (e.g., frames) cannot be decoded independently of each other, determining the first layer index may include, for each layer, determining whether each respective layer was successfully received, and determining the first layer index for the given time interval as the smaller of the first layer index of the time interval preceding the given time interval and the layer index of the layer immediately below the lowest layer that was not successfully received.
[0141] The number N of the enhancement side information payload used for decompression E (k) may be determined according to the following formula.
[0142]
Equation
[0143] That is, specifically, as long as the highest layer number N B (k) does not change, the same corresponding enhancement layer number is selected. However, if N B (k) changes, the enhancement is disabled by setting N E (k) to 0. Due to the assumed differential decompression of the enhancement side information, a change based on N B (k) is not possible. It would require decompression of the corresponding enhancement side information layer in the previous frame, but it is assumed that such decompression was not performed.
[0144] That is, if the compressed audio representations for a series of time intervals (e.g., frames) cannot be decoded independently of each other, the determination of the second layer index may include determining whether the first layer index for the given time interval is equal to the first layer index for the preceding time interval. If the first layer index for the given time interval is equal to the first layer index for the preceding time interval, the second layer index for the given time interval may be determined (e.g., selected) to be equal to the first layer index for the given time interval. On the other hand, if the first layer index for the given time interval is not equal to the first layer index for the preceding time interval, an index value indicating that no enhancement side information is used when obtaining the reconstructed audio representation may be determined (e.g., selected) as the second layer index.
[0145] Alternatively, in decompression, if all of the enhancement side information payloads numbered up to N E (k) are decompressed in parallel, the selection rule of Equation (4) is N E (k)=N B (k) (9) is replaced by.
[0146] Finally, for differential decompression, note that the number N B of the highest layer used can only increase in independent frame - data packets, while it can decrease in any frame.
[0147] It is understood that the proposed method for encoding the layered structure of a compressed sound representation can be implemented by an encoder for encoding the layered structure of a compressed sound representation. Such an encoder may have units adapted to perform each of the above-described steps. An example of such an encoder 5000 is schematically shown in Figure 5. For example, such an encoder 5000 may have a component subdivision unit 5010 adapted to perform S1010 described above, a component assignment unit 5020 adapted to perform S1020 described above, a basic side information assignment unit 5030 adapted to perform S1030 described above, an improved side information division unit 5040 adapted to perform S1040 described above, and an improved side information assignment unit 5050 adapted to perform S1050 described above. Furthermore, it is understood that each unit of such an encoder may be embodied by a processor 5100 of a computing device adapted to perform the processing performed by each of the units, i.e., adapted to perform some or all of the above-described steps or further steps of the proposed encoding method. The encoder or computing device may further have memory 5200 accessible by the processor 5100.
[0148] Furthermore, it is understood that the proposed method for decoding a compressed sound representation encoded in multiple layers may be implemented by a decoder for decoding a compressed sound representation encoded in multiple layers. Such a decoder may have units adapted to perform each of the above-described steps. An example of such a decoder 6000 is schematically shown in Figure 6. For example, such a decoder 6000 may have a receiving unit 6010 adapted to perform S3010 described above, a first layer index determination unit 6020 adapted to perform S3020 described above, a basic reconstruction unit 6030 adapted to perform S3030 described above, a second layer index determination unit 6040 adapted to perform S3040 described above, and an improved reconstruction unit 6050 adapted to perform S3050 described above. Furthermore, it is understood that each unit of such a decoder may be embodied by a processor 6100 of a computing device adapted to perform the processing performed by each of the units, i.e., adapted to perform some or all of the above-described steps or further steps of the proposed decoding method. The decoder or computing device may further have memory 6200 accessible by the processor 6100.
[0149] It should be noted that this paper and its drawings merely illustrate the principles of the proposed method and apparatus. Therefore, it is understood that those skilled in the art may devise various configurations that embody the principles of the present invention and fall within its spirit and scope, even if not explicitly described or illustrated in this paper. Furthermore, all examples described in this paper are explicitly intended solely for educational purposes to assist the reader in understanding the principles of the proposed method and apparatus and the concepts to which the inventors contribute to the advancement of the art, and should be interpreted without limitation to such individually described examples and conditions. Moreover, all statements in this paper describing the principles, aspects, and embodiments of the present invention, as well as their individual examples, are intended to encompass their equivalents.
[0150] The methods and apparatus described in this paper may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or microprocessor. Other components may be implemented as, for example, hardware and / or application-specific integrated circuits. Signals emanating from the methods and apparatus described may be stored on a medium such as random-access memory or optical storage media and transmitted over a network such as a radio network, satellite network, wireless network, or wired network, such as the Internet.
Claims
1. A method for decoding a compressed higher-order ambisonics (HOA) sound representation of a sound or sound field encoded in multiple hierarchical layers using layered encoding, the method being: The stage of receiving a bitstream containing frame packets; The step of multiplexing the compressed HOA representation from the frame packet, the compressed HOA representation corresponding to the multiple hierarchical layers, which include a basic layer and at least one advanced layer, wherein at least one of the multiple hierarchical layers includes components of the basic compressed sound representation of the sound or sound field, and these components correspond to multiple monaural signals. If none of the coefficient sequences of the original HOA representation are included in the basic compressed sound representation of the basic layer, then the dependent basic side information for each vector-based signal consists of all the vector components of the vector-based signal and has the maximum size of the dependent basic side information, and is divided into steps; The step includes decoding the compressed HOA representation based on the dependent base side information, the base side information associated with the base layer, and the improved side information associated with the improved layer. The basic side information indicates that at least one individual monaural signal represents a directional signal with a certain incident direction, and the enhanced side information includes information that allows for the prediction of missing portions of the sound or sound field. method.
2. A non-temporary computer-readable storage medium containing instructions that, when executed by a processor, perform the method described in claim 1.
3. A device for decoding a compressed higher-order ambisonics (HOA) sound representation of a sound or sound field encoded in multiple hierarchical layers using layered encoding, the device comprising: A receiver that receives frame packets; A demultiplexer that multiplexes and demultiplexes from the frame packets a bitstream containing the compressed HOA representation corresponding to the multiple hierarchical layers, the multiple hierarchical layers containing components of the basic compressed sound representation of the sound or sound field, the components corresponding to multiple monaural signals, If none of the coefficient sequences of the original HOA representation are included in the basic compressed sound representation of the basic layer, then a demultiplexer is provided for each vector-based signal, the dependent basic side information for each vector-based signal consists of all the vector components of the vector-based signal and has the maximum size of the dependent basic side information; The decoder comprises a decoder that decodes the compressed HOA representation based on the dependent basic side information, the basic side information associated with the basic layer, and the improved side information associated with the improved layer. The basic side information indicates that at least one individual monaural signal represents a directional signal with a certain incident direction, and the enhanced side information includes information that allows for the prediction of missing portions of the sound or sound field. Device.