Method and apparatus for decoding compressed HOA signals
By incorporating dominant components at lower spatial resolution in the base layer, the HOA compression method is made scalable, allowing for a split into base and enhancement layers, ensuring robustness and quality in streaming environments.
Patent Information
- Application Number
- JP2025157658
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-03-21
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-21
AI Technical Summary
Existing HOA compression methods provide a monolithic compressed representation that is not scalable, making it difficult to split into lower-quality base and higher-quality enhancement layers, which is necessary for applications like broadcasting and Internet streaming.
Modify the HOA compression and decompression methods to include dominant components at lower spatial resolution in the base layer, allowing the compressed representation to be split into a self-contained base layer and an enhancement layer, with the use of a mode indication bit to signal this layered mode.
Enables the creation of a scalable compressed HOA representation that can maintain minimum quality under poor transmission conditions and improve quality with additional enhancement information, suitable for streaming applications.
Smart Images

Figure 2026009938000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for compressing a Higher Order Ambisonics (HOA) signal, a method for decompressing a compressed HOA signal, an apparatus for compressing an HOA signal, and an apparatus for decompressing a compressed HOA signal. [Background technology]
[0002] Higher Order Ambisonics (HOA) offers the possibility of representing three-dimensional sound. Other known techniques are wave field synthesis (WFS) or channel-based methods such as 22.2. However, in contrast to channel-based methods, HOA representations offer the advantage of being independent of a specific loudspeaker setup. However, this flexibility comes at the cost of the decoding process required for the reproduction of an HOA representation on a specific loudspeaker setup. Compared to WFS methods, which typically require a much larger number of loudspeakers, HOA may be rendered to setups consisting of only a few loudspeakers. A further advantage of HOA is that the same representation can also be used for binaural rendering to headphones without any modification.
[0003] HOA is based on the representation of the so-called spatial density of complex harmonic plane wave amplitudes by a truncated spherical harmonics (SH) expansion. Each expansion coefficient is a function of angular frequency, which can be equivalently represented by a time-domain function. Thus, without loss of generality, we can assume that the complete HOA sound field representation actually consists of O time-domain functions, where O represents the number of expansion coefficients. These time-domain functions are hereafter equivalently referred to as HOA coefficient sequences or HOA channels. Conventionally, a spherical coordinate system is used, with the x-axis pointing to the front position, the y-axis pointing to the left, and the z-axis pointing upward. At a position in space x=(r,θ,φ)T is expressed by the radius r>0 (i.e., the distance to the coordinate origin), the tilt angle θ∈[0,π] measured from the polar axis z, and the azimuthal angle φ∈[0,2π[ measured counterclockwise from the x-axis in the xy plane. Furthermore, (·) T represents the transposition.
[0004] A more detailed description of HOA encoding is provided below.
[0005] Fourier transform of sound pressure versus time F t (·), that is, where ω represents the angular frequency and i represents the imaginary unit,
number
number
[0006]
number
number
number
[0007]
number
number
[0008] The spatial resolution of the HOA representation improves with increasing maximum order of the expansion, N. Unfortunately, the number of expansion coefficients, O, is quadratic with the order, N, specifically, O = (N + 1). 2 For example, a typical HOA representation using order N=4 requires O=25 HOA (expansion) coefficients. According to these considerations, the total bit rate for transmission of the HOA representation grows as the desired single-channel sampling rate f s and the number of bits per sample, N b Given O·f s N b As a result, N per sample b = 16 bits for f s Transmitting an HOA representation of order N=4 at a sampling rate of N=48 kHz leads to a bit rate of 19.2 MBits / s, which is too high for many practical applications, e.g., streaming. Thus, compression of the HOA representation is highly desirable.
[0009] Compression of HOA sound field representations has been proposed in European Patent Applications EP 2743922A, EP 2665208A, and EP 2800401A. These approaches share the commonality of performing a sound field analysis and decomposing a given HOA representation into a directional component and a residual ambient component. On the one hand, the final compressed representation is assumed to contain several quantized signals resulting from perceptual coding of the directional signal and the associated coefficient sequences of the ambient HOA component. On the other hand, the final compressed representation is assumed to contain additional side information related to the quantized signals. This side information is necessary for the reconstruction of the HOA representation from its compressed version.
[0010] Furthermore, a similar method is described in Non-Patent Document 1, where the directional component is extended to the so-called predominant sound component, which is assumed to be partly represented by directional signals, i.e., monaural signals with corresponding directions assumed to be incident on the listener from that direction, together with some prediction parameters for predicting parts of the original HOA representation from those directional signals.
[0011] Furthermore, the dominant components are assumed to be represented by so-called vector-based signals, i.e., mono signals with corresponding vectors that define the directional distribution of the vector-based signals. The known compressed HOA representation consists of I quantized mono signals and some additional side information. Here, a fixed number O of these I quantized mono signals are MIN The surrounding HOA component C AMB First O of (k-2) MIN represents the spatially transformed version of the coefficient sequence. MIN The type of signal may change between successive frames and may be directional, vector-based, sky or ambient HOA component C. AMB representing (k-2) additional coefficient sequences.
[0012] One known method for compressing an HOA signal representation having an input time frame (C(k)) of HOA coding coefficient sequences includes spatial HOA encoding of the input time frame followed by perceptual and source encoding. Spatial HOA encoding involves performing direction and vector estimation processing of the HOA signal in direction and vector estimation block 101, as shown in FIG. 1a. Here, a first set of tuples M for the directional signal is DIR (k) and a second set of tuples M for vector-based signals VEC(k) are obtained. Each first set of tuples contains the index of a directional signal and its respective quantized direction, and each second set of tuples contains the index of a vector-based signal and a vector defining the signal's directional distribution. The next step is to convert each input time frame of the HOA coefficient sequence into a set of multiple dominant sound signals X PS Frame (k-1) and the surrounding HOA component C AMB (k-1) frames (103). Here, the dominant sound signal X PS (k-1) includes the directional sound signal and the vector-based sound signal. The decomposition further includes prediction parameters ξ(k-1) and a target assignment vector v A,T (k-1). The prediction parameter ξ(k-1) is used to estimate the dominant sound signal X PS We describe how to predict parts of the HOA signal representation from the directional signals in (k-1) to enrich the dominant HOA components. The target assignment vector v A,T (k-1) contains information on how to allocate the dominant sound signals to a given number I of channels. AMB (k-1) is the target allocation vector v A,T (k-1) (104). Here, it is determined which coefficient sequence of the ambient HOA components should be transmitted in a given number I of channels, depending on how many channels are occupied by the dominant sound signal. The modified ambient HOA components C M,A (k-2) and the temporally predicted corrected ambient HOA component C P,M,A (k-1) is obtained. Also, the target allocation vector v A,T From the information in (k-1), the final assignment vector v A (k-2) is also obtained. The dominant sound signal X obtained from the above decomposition PS (k-1) and the corrected ambient HOA component C M,A (k-2) and the temporally predicted corrected ambient HOA component C P,M,AThe determined coefficient sequence of (k-1) is the final assignment vector v A (k-2), where the transport signal y i (k-2), i=1,…,I and the predicted transport signal y P,i (k-2), i=1,...,I. Then, the transport signal y i (k-2) and the predicted transport signal y P,i (k-2), where gain control (or normalization) is performed on the gain-modified transport signal z i (k-2), index e i (k-2) and exception flag β i (k-2) is obtained.
[0013] As shown in Figure 1b, the perceptual and source encodings are performed on the gain-modified transport signal z i (k-2) perceptual encoding of a perceptually encoded transport signal
number
number
number
[0014] [Patent Document 1] EP12306569.0 [Patent Document 2] EP12305537.8 (published as EP2665208A) [Patent Document 3] EP133005558.2 [Non-patent literature]
[0015] [Non-Patent Document 1] ISO / IEC JTC1 / SC29 / WG11, N14264, "Working Draft 1-HOA Text of MPEG-H 3D audio", January 2014, San Jose Summary of the Invention [Problem to be solved by the invention]
[0016] One drawback of the proposed HOA compression method is that it provides a monolithic (i.e., non-scalable) compressed HOA representation. However, for certain applications, such as broadcasting or Internet streaming, it is desirable to be able to split the compressed representation into a lower-quality base layer (BL) and a higher-quality enhancement layer (EL). The base layer is said to provide a lower-quality compressed version of the HOA representation that can be decoded independently of the enhancement layer. Such a BL should typically be highly robust against transmission errors and transmitted at a low data rate to guarantee a certain minimum quality of the decompressed HOA representation even under poor transmission conditions. The EL contains additional information to improve the quality of the decompressed HOA representation. [Means for solving the problem]
[0017] The present invention provides a solution for modifying existing HOA compression methods so that they can provide a compressed representation that includes a (low quality) base layer and a (high quality) enhancement layer. Furthermore, the present invention provides a solution for modifying existing HOA decompression methods so that they can decode a compressed representation that includes at least a low quality base layer that has been compressed according to the present invention.
[0018] One improvement relates to obtaining a self-contained (low quality) base layer. According to the invention, the ambient HOA component C AMB The first O (without loss of generality) of (k-2) MIN It is assumed that O contains a spatially transformed version of the coefficient sequence MIN channels are used as the base layer. MIN The advantage of selecting these channels is their time-invariant nature. However, conventionally, each signal is completely devoid of the dominant components that are essential for the sound field. This means that the ambient HOA component C AMB It is clear from the previous calculation of (k-1). C AMB (k-1)=C(k-1)-C PS (k-1) (1) According to the original HOA representation C(k-1), the dominant HOA representation C PS This is done by subtracting (k-1).
[0019] Therefore, one improvement of the present invention relates to adding such dominant components. According to the present invention, the solution to this problem is to include the dominant components with low spatial resolution in the base layer. For this purpose, the ambient HOA components C output by the HOA decomposition process in the spatial HOA encoder according to the present invention are AMB (k-1) is replaced by its modified version. The modified ambient HOA components are the first O, which are assumed to always be transmitted in a spatially transformed form. MINIn this coefficient sequence, the coefficient sequence of the original HOA component is included. This improvement of the HOA decomposition process can be seen as an initial operation to make the HOA compression work in a layered mode (e.g., two-layer mode). This mode provides, for example, two bitstreams or a single bitstream that can be split into a base layer and an enhancement layer. The use or non-use of this mode is signaled by a mode indication bit (e.g., a single bit) in the access units of the overall bitstream.
[0020] In one embodiment, the base layer bitstream
number
number
number
number
number
number
[0021] A method for compressing a Higher Order Ambisonics (HOA) signal representation having a time frame of an HOA coefficient sequence is disclosed in claim 1. An apparatus for compressing a Higher Order Ambisonics (HOA) signal representation having a time frame of an HOA coefficient sequence is disclosed in claim 3.
[0022] A method for decompressing a Higher Order Ambisonics (HOA) signal representation having a time frame of a sequence of HOA coefficients is disclosed in claim 2. An apparatus for decompressing a Higher Order Ambisonics (HOA) signal representation having a time frame of a sequence of HOA coefficients is disclosed in claim 4.
[0023] A non-transitory computer-readable storage medium having executable instructions for causing a computer to perform a method for compressing a Higher Order Ambisonics (HOA) signal representation having a time frame of an HOA coefficient sequence is disclosed in claim 5. A non-transitory computer-readable storage medium having executable instructions for causing a computer to perform a method for decompressing a Higher Order Ambisonics (HOA) signal representation having a time frame of an HOA coefficient sequence is disclosed in claim 6.
[0024] Advantageous embodiments of the invention are disclosed in the dependent claims, the following description and the drawings. [Brief explanation of the drawings]
[0025] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1a] 1 shows the general architecture of the HOA compressor. [Figure 1b] 1 shows the general architecture of the HOA compressor. [Figure 2] 1 is a typical architecture of the HOA decompressor. [Figure 3] 1 is an architectural structure of the spatial HOA encoding and perceptual encoding portions of an HOA compressor according to one embodiment of the present invention. [Figure 4]1 is an architectural structure of the source encoder portion of the HOA compressor according to one embodiment of the present invention. [Figure 5] 1 is a diagram illustrating the architecture of the perceptual and source decoding of an HOA decompressor according to one embodiment of the present invention. [Figure 6] 1 is an architectural structure of the spatial HOA decoding portion of an HOA decompressor according to one embodiment of the present invention. [Figure 7] Frame transformation from ambient HOA signal to modified ambient HOA signal. [Figure 8] 1 is a flowchart of a method for compressing an HOA signal. [Figure 9] 1 is a flowchart of a method for decompressing a compressed HOA signal. [Figure 10] 10 details portions of the architecture of the spatial HOA decoding portion of an HOA decompressor according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] For ease of understanding, the prior art solutions of Figures 1 and 2 are reviewed below.
[0027] Figure 1 shows the general architecture of an HOA compressor. In the method described in [1], the directional component is expanded into a so-called dominant component. As a directional component, the dominant component is assumed to be represented in part by a directional signal, i.e., a monophonic signal with a corresponding direction assumed to be incident on the listener from that direction, together with some prediction parameters for predicting parts of the original HOA representation from those directional signals. Furthermore, the dominant component is assumed to be represented by a so-called vector-based signal, i.e., a monophonic signal with a corresponding vector defining the directional distribution of the vector-based signal. The overall architecture of the HOA compressor proposed in [1] is shown in Figure 1. It can be subdivided into a spatial HOA encoding part, depicted in Figure 1a, and a source encoding part, depicted in Figure 1b. The spatial HOA encoder provides a first compressed HOA representation consisting of I signals together with side information describing how to generate the HOA representation. In a perceptual and side source coder, the above mentioned I signals are perceptually encoded, the side information is subjected to source encoding, and then the two coded representations are multiplexed.
[0028] Typically, spatial encoding works as follows.
[0029] In the first stage, the kth frame C(k) of the original HOA representation is input to the direction and vector estimation processing block, which is a set of tuples M DIR (k) and M VEC (k) is given. The tuple set M DIR (k) consists of tuples whose first element represents the index of the directional signal and whose second element represents the respective quantized direction. VEC (k) consists of a tuple where the first element indicates the index of the vector-based signal and the second element represents the directional distribution of the signal, i.e., a vector that defines how the HOA representation of the vector-based signal is calculated.
[0030] Tuple set MDIR (k) and M VEC Using both (k) and (k), the initial HOA frame C(k) is the frame X of the total dominant (i.e., directional and vector-based) signal in the HOA decomposition. PS Frame (k-1) and frame C of the surrounding HOA component AMB (k-1). Note the delay of one frame each. This is due to the overlap-add process to avoid blocking artifacts. Furthermore, the HOA decomposition is assumed to output some prediction parameters ξ(k-1) that describe how to predict parts of the original HOA representation from the directional signal to enrich the dominant HOA component. Furthermore, a target assignment vector v is generated that contains information about the allocation of the dominant signal to the I available channels determined in the HOA decomposition processing block. A,T (k-1) is provided. The affected channels can be assumed to be occupied, i.e., they are not available to transmit any coefficient sequences of the ambient HOA components in each time frame.
[0031] In the ambient component correction processing block, the frame C of the ambient HOA component AMB (k-1) is the target allocation vector v A,T (k-1). In particular, which coefficient sequences of the ambient HOA components should be transmitted in the given I channels depends, among other aspects, on the information about which channels are available and not already occupied by the dominant sound signal (the target allocation vector v A,T (k-1)). Furthermore, if the index of the selected coefficient sequence changes between successive frames, a fade-in and fade-out of the coefficient sequence is performed.
[0032] Furthermore, the ambient HOA component C AMB First O of (k-2) MINA sequence of coefficients is always chosen to be perceptually coded and transmitted, where O MIN =(N MIN +1) 2 and N MIN ≦N is typically a smaller order than that of the original HOA representation. To decorrelate these HOA coefficient sequences, we map them into some predefined directions Ω MIN,d , d=1,…,O MIN It is proposed to convert the incident directional signal (i.e., a general plane wave function) from the corrected ambient HOA component C AMB (k-1) together with the time-predicted corrected ambient HOA component C for later use in the gain control processing block to allow reasonable look-ahead. P,M,A (k-1) is calculated.
[0033] The information about the correction of the ambient HOA components is directly related to the assignment of all possible types of signals to the available channels. The final information about the assignment is the final assignment vector v A (k-2). To calculate this vector, the target allocation vector v A,T The information contained in (k-1) is utilized.
[0034] The channel assignment is given by the assignment vector v A Using the information given by (k-2), X PS (k-2) and the appropriate signal C M,A (k-2) to the I available channels, and assign the appropriate signals in i (k-2), i=1,...,I. Furthermore, X PS (k-1) and the appropriate signal C P,AMB The appropriate signals in (k-1) are also assigned to the I available channels to generate the signal y P,i (k-2), i=1,...,I. Signal y iEach of the predicted signal frames y (k-2), i=1,...,I, is finally processed by gain control, where the signal gain is smoothly modified to achieve a range of values suitable for the perceptual encoder. P,i (k-2), i=1,...,I allows a kind of look-ahead to avoid drastic gain changes between successive blocks. The gain correction is performed in the spatial decoder by the index e i (k-2) and exception flag β i (k-2), i=1,...,I, are assumed to be inverted using gain control side information.
[0035] Figure 2 shows the general architectural structure of the HOA decompressor proposed in Non-Patent Document 1. Typically, the HOA decompressor consists of the counterparts of the HOA compressor components, which are naturally arranged in reverse order. The HOA decompressor is subdivided into a perceptual and source decoding part, depicted in Figure 2a), and a spatial HOA decoding part, depicted in Figure 2b).
[0036] In the perceptual and side source decoder, the bitstream is first demultiplexed into perceptually coded representations of the I signals and coded side information that describes how to generate the HOA representations. Subsequently, perceptual decoding of the I signals and decoding of the side information are performed. A spatial HOA decoder then generates reconstructed HOA representations from the I signals and the side information.
[0037] Typically, spatial HOA decoding works as follows.
[0038] In a spatial HOA decoder, the perceptually decoded signal
number
number
[0039] I gain-corrected signal frames
number
number
[0040] In dominant sound synthesis, the dominant sound component
number
[0041] In ambient synthesis, the ambient HOA component frame
number
[0042] As is clear from the above rough description of the HOA compression and decompression method, the compressed representation consists of I quantized mono signals and some additional side information. Of these I quantized mono signals, a fixed number O MIN The surrounding HOA component C AMB First O of (k-2) MIN represents the spatially transformed version of the coefficient sequence. MINThe type of signal may change between successive frames and may be directional, vector-based, sky- or ambient HOA component C. AMB It can either represent (k-2) additional coefficient sequences. As such, the compressed HOA representation is intended to be monolithic. In particular, one problem is how to split the described representation into a lower-quality base layer and an enhancement layer.
[0043] According to the disclosed invention, candidates for the low-quality base layer are the ambient HOA components C AMB First O of (k-2) MIN containing spatially transformed versions of the coefficient sequences MIN These are (without loss of generality) the first O MIN This channel is a good choice for the low-quality base layer due to its time-invariant nature. However, each signal lacks any dominant components that are essential for the sound field. This means that the ambient HOA component C AMB This can also be seen in the calculation of (k-1). C AMB (k-1)=C(k-1)-C PS (k-1) (1) According to the original HOA representation C(k-1), the dominant HOA representation C PS This is done by subtracting (k-1).
[0044] A solution to this problem is to include the dominant tonal components at lower spatial resolution in the base layer.
[0045] Proposed modifications to HOA compression are described below.
[0046] 3 shows the architecture of the spatial HOA encoding and perceptual encoding parts of the HOA compressor according to an embodiment of the present invention. In order to include dominant sound components at low spatial resolution in the base layer, the ambient HOA components C AMB(k-1) (see Figure 1a) is the modified version
number
[0047]
number
[0048] It is important to note that this change in the HOA decomposition process can be seen as the beginning of making HOA compression work in the so-called "dual layer" or "two-layer" mode. This mode provides a bitstream that can be split into a lower quality base layer and an enhancement layer. The use or non-use of this mode can be signaled by a single bit in the access units of the entire bitstream.
[0049] Possible resulting modifications of the bitstream multiplexing to provide bitstreams for the base layer and enhancement layer are shown in Figures 3 and 4 and are further described below.
[0050] Base Layer Bitstream
number
number
number
number
number
[0051] 3 and 4 show an apparatus for compressing an HOA signal, which is an input HOA representation having an input time frame (C(k)) of an HOA coefficient sequence. The apparatus comprises a spatial HOA encoding and perceptual encoding unit shown in Fig. 3 for spatial HOA encoding and subsequent perceptual encoding of the input time frame, and a source encoder unit shown in Fig. 4 for source encoding. The spatial HOA encoding and perceptual encoding unit comprises a direction and vector estimation block 301, an HOA decomposition block 303, an ambient component correction block 304, a channel allocation block 305, and multiple gain control blocks 306.
[0052] The direction and vector estimation block 301 is adapted to perform direction and vector estimation processing of the HOA signal, where a first set of tuples M DIR (k) and a second set of tuples M for vector-based signals VEC Data including (k) is obtained. Each first tuple set M DIR (k) contains the index of the directional signal and the respective quantized direction, and each second tuple set M VEC(k) contains a vector that defines the vector-based signal index and the directional distribution of the signal.
[0053] The HOA decomposition block 303 decomposes each input time frame of the HOA coefficient sequence into a plurality of dominant sound signals X PS Frame (k-1) and surrounding HOA components
number
number
[0054] The ambient component correction block 304 corrects the ambient HOA component C AMB (k-1) as the target allocation vector v A,T (k-1), where the ambient HOA component C AMB Which coefficient sequence of (k-1) should be transmitted in a given number I of channels is determined depending on how many channels are occupied by the dominant sound signal. M,A(k-2) and the temporally predicted corrected ambient HOA component C P,M,A (k-1) is obtained. Also, the target allocation vector v A,T From the information in (k-1), the final assignment vector v A (k-2) is obtained.
[0055] The channel allocation block 305 allocates the dominant signal X obtained from the above decomposition. PS (k-1) and the corrected ambient HOA component C M,A (k-2) and the temporally predicted corrected ambient HOA component C P,M,A The determined coefficient sequence of (k-1) and the final assignment vector v A (k-2) where the transport signal y is adapted to be assigned to the given number I of channels using the information given by i (k-2), i=1,…,I and the predicted transport signal y P,i (k-2), i=1,...,I is obtained.
[0056] The plurality of gain control blocks 306 are connected to the transport signal y i (k-2) and the predicted transport signal y P,i (k-2), where the gain-modified transport signal z i (k-2), index e i (k-2) and exception flag β i (k-2) is obtained.
[0057] Figure 4 shows the architecture of the source encoder part of the HOA compressor according to one embodiment of the present invention. The source encoder part shown in Figure 4 includes a perceptual encoder 310, a side source encoder block with two encoders 320, 330, namely, a base layer side source encoder 320 and an enhancement layer side information encoder 330, and two multiplexers 340, 350, namely, a base layer bitstream multiplexer 340 and an enhancement layer bitstream multiplexer 350. The side source encoders may also be a single side source encoder block.
[0058] The perceptual coder 310 converts the gain-modified transport signal z i perceptually encoding 806 (k-2)
number
[0059] The side source encoders 320, 330 use the index e i (k-2) and exception flag β i (k-2), the first tuple set M DIR (k) and a second tuple set M VEC (k), the prediction parameters ξ(k-1) and the final assignment vector v A (k-2), and the encoded side information
number
[0060] The multiplexers 340, 350 are connected to the perceptually encoded transport signal
number
number
number
number
number
number
number
[0061] Remaining IO MIN exponent e i (k-2), i=O MIN +1,…,I and exception flag β i (k-2), i=O MIN +1,...,I, the first tuple set M DIR (k-1) and a second set of tuples M VEC (k-1), the prediction parameters ξ(k-1) and the final assignment vector v A (k-2) is encoded in the enhancement layer side information encoder 330, where the encoded enhancement layer side information
number
[0062] Remaining IO MIN perceptually encoded transport signals
number
number
number
[0063] In an embodiment, the encoding device further comprises a mode selector adapted to select a mode, the mode being determined by a mode indicator LMF E and is one of the layered and non-layered modes. In the non-layered mode, the ambient HOA component (C AMB (k-1)] contains only the HOA coefficient sequence that represents the residual between the input HOA representation and the HOA representation of the dominant sound signal (ie, does not contain the coefficient sequence of the input HOA representation).
[0064] A proposed modification of the HOA decompression is described below.
[0065] In the layered mode, the ambient HOA component C in the HOA compression AMB The (k-1) modifications are taken into account in the HOA decompression by appropriately modifying the HOA synthesis.
[0066] In the HOA decompressor, demultiplexing and decoding of the base layer and enhancement layer bitstreams is performed according to Figure 5. Base Layer Bitstream
number
[0067] Specifically, the reconstructed HOA representation
number
number
[0068]
number
[0069] Below we will use a purely low-quality base layer bitstream.
number
[0070] The bit stream is first demultiplexed and decoded to produce the reconstructed signal ^z i (k) and the exponent e i (k) and exception flag β i (k), i=1,…,O MINand the corresponding gain control side information consisting of:
number
number
[0071] In the next step, in the spatial HOA decoder, the first O MIN Inverse gain control processing blocks generate gain corrected signal frames
number
[0072] Figures 5 and 6 show the architecture of an HOA decompressor according to one embodiment of the present invention. The device comprises a perceptual and source decoding unit shown in Figure 5, a spatial HOA decoding unit shown in Figure 6, and a base layer bitstream in which the compressed HOA signal is decoded.
number
[0073] 5 shows the architecture of the perceptual and source decoding unit of the HOA decompressor according to one embodiment of the present invention. The perceptual and source decoding unit includes a first demultiplexer 510, a second demultiplexer 520, a base layer perceptual decoder 540 and an enhancement layer perceptual decoder 550, a base layer side source decoder 530 and an enhancement layer side source decoder 560.
[0074] The first demultiplexer 510 receives the compressed base layer bitstream
number
number
number
number
number
number
[0075] The base layer perceptual decoder 540 and the enhancement layer perceptual decoder 550 receive the perceptually encoded transport signal
number
number
number
number
number
number
[0076] The base layer side source decoder 530 receives the first encoded side information
number
[0077] The enhancement layer side source decoder 560 generates the second encoded side information
number
[0078] 6 shows the architecture of the spatial HOA decoding section of the HOA decompressor according to one embodiment of the present invention. The spatial HOA decoding section includes multiple inverse gain control units 604, a channel reassignment block 605, a dominant sound synthesis block 606, an ambient synthesis block 607, and an HOA composition block 608.
[0079] A plurality of inverse gain control units 604 are adapted to perform inverse gain control, wherein the first perceptually decoded transport signal
number
number
[0080] The channel reassignment block 605 generates the first and second gain corrected signal frames ^y i (k), i=1,...,I, is adapted to redistribute the frame ^X of the dominant signal PS (k) is reconstructed, and the dominant sound signal includes directional and vector-based signals, and the corrected ambient HOA components
number
[0081] Furthermore, the channel reassignment block 605 calculates a first set I of indices of coefficient sequences of modified ambient HOA components that are active in the k-th frame. AMB,ACT(k) and a second set I of indices of coefficient sequences of modified ambient HOA components that need to be enabled, disabled or remain active in the (k-1)th frame. E (k-1), I D (k-1) and I U (k-1) and
[0082] The dominant sound synthesis block 606 synthesizes the dominant HOA sound component ^C PS The HOA representation of (k-1) is expressed as the dominant sound signal ^X PS (k) is adapted to be composed (912) from the first and second tuple sets M DIR (k+1), M VEC (k+1), the prediction parameters ζ(k+1) and the second set of indices I E (k-1), I D (k-1), I U (k-1) is used.
[0083] The ambient compound block 607 is the ambient HOA component.
number
number
[0084] Layering mode indication LMF D If shows a layered mode with at least two layers, the ambient HOA component is MINThe lowest positions (i.e., the positions with the lowest indices) contain the HOA coefficient sequences of the decompressed HOA signal ^C(k-1), and the remaining higher positions contain coefficient sequences that are part of the HOA representation of the residual, which is the decompressed HOA signal ^C(k-1) and the 914 dominant HOA tonal component ^C PS is the residual between the HOA representation of (k-1).
[0085] On the other hand, layered mode indication LMF D If indicates a single layer mode, the HOA coefficient sequence of the decompressed HOA signal ^C(k-1) is not included, and the ambient HOA component is the decompressed HOA signal ^C(k-1) and the dominant HOA tone component ^C PS is the residual between the HOA representation of (k-1).
[0086] The HOA synthesis block 608 is adapted to add the HOA representation of the dominant tonal component to the ambient HOA components.
[0087]
number
number
number
number
[0088] FIG. 7 shows the transformation of a frame from an ambient HOA signal to a modified ambient HOA signal.
[0089] FIG. 8 shows a flow chart of a method for compressing an HOA signal.
[0090] A method 800 for compressing a Higher Order Ambisonics (HOA) signal, which is an input HOA representation of order N having an input time frame C(k) of HOA coefficient sequences, includes spatial HOA encoding of the input time frame followed by perceptual encoding and source encoding.
[0091] Spatial HOA encoding is Executing a direction and vector estimation process 801 for HOA signals in the direction and vector estimation block 301, DIR (k) and a second set of tuples M for vector-based signals VEC (k) are obtained, and each first tuple set M DIR (k) contains the index of the directional signal and the respective quantized direction, and each second tuple set M VEC (k) is a vector-based signal including a vector defining the signal index and the signal directional distribution; In the HOA decomposition block 303, each input time frame of the HOA coefficient sequence is decomposed into a plurality of dominant sound signals X PS Frame (k-1) and surrounding HOA components
number
number
[0092] The perceptual encoding and source encoding are In the perceptual coder 310, the gain-modified transport signal z i 806, perceptually encoding (k-2),
number
number
number
number
number
[0093] The ambient HOA components (C with tilde) obtained in the decomposition step 802 AMB (k-1)] is the input HOA representation c n Let O be the first HOA coefficient sequence of (k-1). MIN In the lowest positions (i.e., the positions with the lowest indices), a second HOA coefficient sequence C AMB,n (k-1) in the remaining higher positions. The second HOA coefficient sequence is part of the HOA representation of the residual between the input HOA representation and the HOA representation of the dominant sound signal.
[0094] First O MIN exponent e i (k-2), i=1,…,O MIN and exception flag β i (k-2), i=1,…,O MIN is encoded in the base layer side source encoder 320, and the encoded base layer side information
number
[0095] First OMIN perceptually encoded transport signals
number
number
number
[0096] Remaining IO MIN exponent e i (k-2), i=O MIN +1,…,I and exception flag β i (k-2), i=O MIN +1,...,I, the first tuple set M DIR (k-1) and a second set of tuples M VEC (k-1), the prediction parameters ξ(k-1) and the final assignment vector v A (k-2) (v in the drawing) AMB,ASSIGN (k) is encoded in the enhancement layer side information encoder 330, where the encoded enhancement layer side information
number
[0097] Remaining IO MIN perceptually encoded transport signals
number
number
number
[0098] As described above, a mode indication signaling the use of layered mode is added 811. The mode indication is added by an indication insertion block or multiplexer.
[0099] In an embodiment, the method further comprises:
number
number
[0100] In one embodiment, the dominant direction estimation relies on the directional power distribution of the energetically dominant HOA components.
[0101] In one embodiment, if the HOA sequence index of the selected HOA coefficient sequence changes between successive frames, a fade-in and fade-out of the coefficient sequence is performed when modifying the ambient HOA components.
[0102] In one embodiment, when modifying the ambient HOA component, the ambient HOA component C AMB A (k-1) partial decorrelation is performed.
[0103] In one embodiment, the first set of tuples M DIR The quantized direction included in (k) is the dominant direction.
[0104] 9 shows a flowchart of a method for decompressing a compressed HOA signal. In this embodiment of the present invention, the method 900 for decompressing a compressed HOA signal includes perceptual and source decoding and subsequent spatial HOA decoding to obtain an output time frame ̂C(k−1) of the HOA coefficient sequence. The method further comprises:
number
number
[0105] The perceptual decoding and source decoding may include: Compressed Base Layer Bitstream
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0106] The spatial HOA decoding is performing 910 inverse gain control,
number
number
number
number
number
number
number
number
number
[0107] Layering mode indication LMF D The composition of the ambient HOA components depending on is as follows:
[0108] Layering mode indication LMF D If shows a layered mode with at least two layers, the ambient HOA component is MIN The lowest positions contain the HOA coefficient sequence of the decompressed HOA signal ^C(k-1), and the remaining higher positions contain the decompressed HOA signal ^C(k-1) and the dominant HOA tonal component ^C PS It contains a coefficient sequence that is part of the HOA representation of the residual between the (k-1) HOA representation.
[0109] On the other hand, layered mode indication LMF D If represents a single layer mode, the ambient HOA components are the decompressed HOA signal ^C(k-1) and the dominant HOA sound component ^C PS is the residual between the HOA representation of (k-1).
[0110] In one embodiment, the compressed HOA signal representation is in a multiplexed bitstream, and the method for decompressing the compressed HOA signal further comprises the initial step of demultiplexing the compressed HOA signal representation from the compressed base layer bitstream.
number
number
[0111] FIG. 10 shows details of portions of the architecture of the spatial HOA decoding portion of the HOA decompressor, according to one embodiment of the present invention.
[0112] Advantageously, it is possible to decode only the BL, for example if no EL is received or if the BL quality is sufficient. In this case, the signal of the EL can be set to 0 in the decoder. Then, the dominant tone signal ^X PS Since the frame of (k) is empty, in the channel reallocation block 605, the first and second gain-compensated signal frames ^y i (k), i=1,...,I, to I channels. It is very simple to redistribute 911 the second set I of indices of coefficient sequences of modified ambient HOA components that need to be enabled, disabled, or remain active in the (k-1)th frame. E (k-1), I D (k-1) and I U (k-1) is set to 0. Therefore, the dominant HOA sound signal ^X in the dominant sound synthesis block 606 PS Dominant HOA sound component ^C from (k) PS The synthesis 912 of the (k-1) HOA representation can be skipped and the modified ambient HOA components in the ambient synthesis block 607
number
number
[0113] The original (i.e., monolithic, non-scalable, non-layered) mode for HOA compression may still be useful for applications where a low-quality base layer is not required, e.g., for file-based compression. The ambient HOA component C, which is the difference between the original and directional HOA representations, AMB The spatially transformed first O MIN The advantage of perceptually encoding the coefficient sequence z instead of the spatially transformed coefficient sequence of the original HOA component C is that in the former case the cross-correlation between all signals to be perceptually encoded is reduced. i Any cross-correlation between i=1,…,I can cause constructive superposition of perceptual coding noise during the spatial decoding process, while at the same time, the noiseless HOA coefficient sequence is cancelled out in the superposition. This phenomenon is known as perceptual noise unmasking.
[0114] In layered mode, the signal z i , i=1,…,O MIN During each of these, the signal z i , i=1,…,O MIN and z i , i=O MIN There is a high cross-correlation between +1,...,I because the surrounding HOA components
number
[0115] While the basic novel features of the present invention have been shown, described, and pointed out as applied to its preferred embodiments, it will be understood that various omissions, substitutions, and changes in the described apparatus and methods may be made by those skilled in the art in the form and details of the disclosed devices and their operation without departing from the spirit of the invention. Any combination of elements that perform substantially the same function in substantially the same way to achieve the same results is expressly intended to be within the scope of the invention. The substitution of elements from one described embodiment for another described embodiment is also fully intended and contemplated.
[0116] It will be understood that the present invention has been described purely by way of example and modifications of detail can be made without departing from the scope of the invention.
[0117] Each feature disclosed in this description and (where appropriate) the claims and drawings may be provided independently or in any appropriate combination. Features may, where appropriate, be implemented in hardware, software or a combination of both. Connections, where applicable, may be implemented as wireless or wired, not necessarily direct or dedicated, connections.
[0118] Reference signs appearing in the claims are by way of example only and shall have no limiting effect on the scope of the claims. stomach.
[0119] Several aspects will be described. [Aspect 1] 1. A method (800) for compressing a Higher Order Ambisonics (HOA) signal, the HOA signal being an order-N input HOA representation having an input time frame (C(k)) of HOA coefficient sequences, the method comprising spatial HOA encoding of the input time frame followed by perceptual and source encoding; The spatial HOA encoding is A step of performing a direction and vector estimation process (801) of the HOA signal in a direction and vector estimation block (301), wherein a first set of tuples (M DIR (k)) and a second set of tuples for vector-based signals (M VEC (k)) is obtained, and the first set of tuples (M DIR (k)) each includes an index of a directional signal and a respective quantized direction, and the second set of tuples (M VEC (k)) each of the vector-based stages includes a vector defining a signal index and a signal directional distribution; In the HOA decomposition block (303), each input time frame of the HOA coefficient sequence is decomposed into a plurality of dominant sound signals (X PS (k-1)) frame and surrounding HOA components
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
Claims
1. 1. A method for decoding a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field, the method comprising: receiving a bitstream containing the compressed HOA representation; determining whether the bitstream includes a single layer or multiple layers; and decoding the compressed HOA representation from the bitstream based on a determination that the bitstream includes only a single layer to obtain a sequence of decoded HOA representations; decoding the compressed HOA representations includes determining a dominant HOA tonal component based on a dominant tonal synthesis and determining ambient HOA components based on ambient synthesis, wherein the sequence of decoded HOA representations is each based on a summation of the dominant HOA tonal component and the ambient HOA components; method.
2. A non-transitory computer-readable storage medium containing instructions that, when executed by a processor, perform the method of claim 1.
3. 1. An apparatus for decoding a compressed Higher Order Ambisonics (HOA) representation of a sound or sound field, the apparatus comprising: a receiver configured to receive a bitstream including the compressed HOA representation; a processor configured to determine whether the bitstream includes a single layer or multiple layers; a decoder configured to, based on a determination that the bitstream includes only a single layer, decode the compressed HOA representation from the bitstream to obtain a sequence of decoded HOA representations; decoding the compressed HOA representations includes determining a dominant HOA tonal component based on a dominant tonal synthesis and determining ambient HOA components based on ambient synthesis, wherein the sequence of decoded HOA representations is each based on a summation of the dominant HOA tonal component and the ambient HOA components; Device.
Citation Information
Patent Citations
EP133005558.2
Method and apparatus for compressing and decompressing a Higher Order Ambisonics signal representation
EP2665208A1
Method and apparatus for compressing and decompressing a higher order ambisonics representation for a sound field
EP2743922A1