Method and apparatus for improving the encoding of side information required to encode higher-order ambisonic representations of sound fields.

By optimizing the encoding of side information in HOA representations through predictive encoding and array omission, the method reduces bitrate from 54 to 46 bits, addressing inefficiencies in existing HOA encoding methods.

JP2026083325APending Publication Date: 2026-05-19DOLBY INTERNATIONAL AB
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2026-03-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The existing methods for encoding side information in higher-order ambisonic (HOA) representations are inefficient, leading to high bitrates that are impractical for applications like streaming, particularly due to the quadratic increase in the number of expansion coefficients with order N, resulting in bitrates like 19.2 Mbits/s for N=4, which is excessive for many practical uses.

Method used

The method involves encoding side information by adding a bit to indicate if prediction should be performed, transmitting the number of active predictions and their indices, and omitting arrays when no prediction is needed, thereby reducing the bitrate through efficient encoding of spatial prediction parameters.

Benefits of technology

This approach significantly reduces the average bitrate for data transmission by minimizing unnecessary data transmission, achieving a more efficient encoding of side information for HOA representations, potentially lowering the bitrate to 46 bits from 54 bits in current technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083325000001_ABST
    Figure 2026083325000001_ABST
Patent Text Reader

Abstract

This improves the encoding of side information required to encode higher-order ambisonic representations of sound fields. [Solution] Higher-order ambisonics represent three-dimensional sound independent of a specific speaker setup. However, transmitting HOA representations leads to very high bit rates. Therefore, compression using a fixed number of channels is employed, where directional and ambient signal components are processed in different ways. For encoding, parts of the original HOA representation are predicted from the directional signal components. This prediction provides the side information required for corresponding decoding. Known side information encoding processes are improved by using several additional purpose-specific bits, which reduces the number of bits required to encode that side information on average.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for improving the encoding of side information required to encode higher-order ambisonic representations of sound fields. [Background technology]

[0002] Higher-Order Ambisonics (HOA) offers one possibility for representing three-dimensional sound, alongside other techniques such as wave-field synthesis (WFS) or channel-based approaches like the 2.2 multi-channel audio format. In contrast to channel-based methods, HOA representation offers the advantage of being independent of a specific speaker setup. However, this flexibility comes at the cost of the decoding process required to reproduce the HOA representation in a particular speaker setup. Compared to the WFS approach, which typically requires a very large number of speakers, the HOA signal can be rendered to a setup with only a few speakers. A further advantage of HOA is that the same representation can also be used for binaural rendering to headphones without modification.

[0003] HOA is based on a truncated spherical harmonic (SH) expansion of the spatial density of complex harmonic plane wave amplitudes. Each expansion coefficient is a function of angular frequency, which can be equivalently expressed by a time-domain function. Thus, without loss of generality, a complete HOA sound field representation can be assumed to consist of O time-domain functions, where O represents the number of expansion coefficients. These time-domain functions are equivalent, but are referred to below as HOA coefficient sequences or HOA channels.

[0004] The spatial resolution of the HOA representation improves with increasing maximum order N of the expansion. Unfortunately, the number of expansion coefficients O is quadratic with increasing order N, and in particular O=(N+1) 2It increases in the form of . For example, a typical HOA representation using order N=4 requires 0=25 HOA (expansion) coefficients. According to previous considerations, the total bit rate for transmitting the HOA representation is the desired single-channel sampling rate f S and the number of bits per sample N b Given, O·f S ·N b This is determined by the following. As a result, the HOA representation of order N=4 is f S =48kHz sampling rate, N per sample b Transmitting using 16 bits results in a bitrate of 19.2 Mbits / s. This is extremely high for many practical applications, such as streaming. Thus, compression of the HOA representation is highly desirable.

[0005] Compression of HOA sound field representations has been proposed in WO2013 / 171083A1, EP13305558.2, and PCT / EP2013 / 075559. These processes share the common approach of performing a sound field analysis and decomposing a given HOA representation into a directional component and a residual ambient component. On the one hand, the final compressed representation is assumed to consist of several quantized signals, which result from perceptual coding of the directional signal and the associated coefficient sequence of the ambient HOA component. On the other hand, the final compressed representation is assumed to contain additional side information related to the quantized signals. This side information is necessary for the reconstruction of the HOA representation from its compressed version.

[0006] The important part of the side information is the description of the prediction of the various parts of the original HOA representation from the directional signals. For this prediction, since the original HOA representation is assumed to be equivalently represented by a number of spatially scattered general plane waves incident from spatially uniformly distributed directions, this prediction is hereinafter referred to as spatial prediction.

[0007] The encoding of such side information related to spatial prediction is described in Non-Patent Document 1. However, this prior art encoding of side information is quite inefficient.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] The problem to be solved by the present invention is to provide a more efficient method for encoding side information related to such spatial prediction.

Means for Solving the Problems

[0010] This problem is solved by the methods disclosed in claims 1 and 6. Apparatus using these methods are disclosed in claims 2 and 7.

[0011] Bits are added before the encoded side information representation data ζ COD This bit indicates whether any prediction should be performed. This feature is ζ CODThis reduces the average bitrate for data transmission over time. Furthermore, in specific situations, it is more efficient to transmit or transfer the number of active predictions and their respective indices rather than using a bit array indicating whether a prediction is performed in each direction. A single bit can be used to indicate how the index for the direction in which a prediction should be performed is encoded. On average, this behavior decreases over time, ζ COD Further reduce the bitrate for data transmission.

[0012] In principle, the method of the present invention is suitable for improving the encoding of side information required to encode a higher-order ambisonic representation (HOA) of a sound field, having an input time frame of HOA coefficient sequences. Here, a dominant directional signal and a residual ambient HOA component are determined, a prediction is used for the dominant directional signal, thereby providing side information data describing the prediction for the encoded frame of HOA coefficients, the side information data being: A bit array indicating whether or not a prediction is performed in a certain direction; A bit array where each bit indicates the type of prediction and the direction in which the prediction should be performed; A data array containing elements representing the indices of the directional signals to be used for the predictions to be made; It can include a data array with elements representing quantized scaling factors, The method in question is: • Provides a bit value indicating whether the prediction should be performed or not; If there is no prediction to be made, the bit array and data array in the side information data shall be omitted; - If the prediction should be performed, the step includes providing a bit value indicating whether the side information data includes the number of active predictions and a data array containing the index of the direction in which the prediction should be performed, instead of the bit array indicating whether the prediction should be performed in a certain direction.

[0013] In principle, the apparatus of the present invention is suitable for improving the encoding of side information required to encode a higher-order ambisonic representation (HOA) of a sound field, having an input time frame of HOA coefficient sequences. Here, a dominant directional signal and residual ambient HOA components are determined, a prediction is used for the dominant directional signal, and therefore, side information data describing the prediction is provided for the encoded frame of HOA coefficients, wherein the side information data is: A bit array indicating whether or not a prediction is performed in a certain direction; A bit array where each bit indicates the type of prediction and the direction in which the prediction should be performed; A data array containing elements representing the indices of the directional signals to be used for the predictions to be made; It can include a data array with elements representing quantized scaling factors, The device in question: • Provides a bit value indicating whether the prediction should be performed or not; If there is no prediction to be made, the bit array and data array in the side information data shall be omitted; The means includes, if the prediction should be performed, providing a bit value indicating whether the side information data includes the number of active predictions and a data array containing the index of the direction in which the prediction should be performed, instead of providing the bit array indicating whether the prediction should be performed in a certain direction.

[0014] Advantageous additional embodiments of the present invention are disclosed in their respective dependent claims. [Brief explanation of the drawing]

[0015] Exemplary embodiments of the present invention are described with reference to the accompanying drawings. [Figure 1] This figure shows an example of encoding of side information related to spatial prediction in the HOA compression process described in EP13305558.2. [Figure 2]This figure shows an example of decoding side information related to spatial prediction in the HOA decompression process described in patent application EP13305558.2. [Figure 3] This is a diagram showing the HOA decomposition described in patent application PCT / EP2013 / 075559. [Figure 4] This diagram shows the direction of the general plane wave representing the residual signal (indicated by ×) and the direction of the dominant sound source (indicated by ○). The directions are presented in a three-dimensional coordinate system as sampling positions on a unit sphere. [Figure 5] This figure shows the current coding techniques for side information in spatial prediction. [Figure 6] This figure shows the coding of side information for spatial prediction according to the present invention. [Figure 7] This figure shows the decoding of the encoded spatial prediction according to the present invention. [Figure 8] Continuation of Figure 7. [Modes for carrying out the invention]

[0016] Below, in order to provide context for the use of the encoding of side information related to spatial prediction according to the present invention, the HOA compression and decompression processes described in patent application EP13305558.2 are summarized.

[0017] <HOA compression> Figure 1 shows how encoding of side information related to spatial prediction can be embedded in the HOA compression process described in patent application EP13305558.2. For HOA representation compression, frame-by-frame processing is assumed using non-overlapping input frames C(k) of HOA coefficient sequences of length L, where k represents the frame index. The first stage or stage 11 / 12 in Figure 1 is arbitrary, and the non-overlapping k-th and (k-1)th frames of the HOA coefficient sequence C(k) are used as the long frame.

number

[0018] Long frames [C(k) with a tilde] are used successively in steps or stage 13 for estimation of the dominant sound source direction, as described in EP13305558.2. This estimation is based on a data set of indices of the detected relevant directional signals.

number

number

[0019] In stage 14, the current (long) frame of the HOA coefficient sequence [C(k) with a tilde] is set (as proposed in EP13305156.5)

number

number

[0020] In step or stage 15, the surrounding HOA component C AMB The number of coefficients in (k-2) is only O RED +DN DIR,ACT It is reduced to include (k-2) non-zero HOA coefficient sequences. Here,

number

number

[0021] The number of reduced numbers is O. RED +N DIR,ACTThe final surrounding HOA representation with (k-2) non-zero coefficient sequences is C AMB,RED It is represented by (k-2). The index of the selected surrounding HOA coefficient sequence is the data set

number

number

[0022] According to the present invention, after the decomposition of the original HOA representation in step 14, the spatial prediction parameters or side information data ζ(k-2) resulting from the decomposition of the HOA representation are encoded in step 19. COD To provide (k-2), an index set

number

[0023] <HOA compression removal> Figure 2 shows the received encoded side information data ζ related to spatial prediction. COD The decoding of (k-2) is illustrated in example how it is embedded in the HOA decompression process described in Figure 3 of patent application EP13305558.2 at stage 25. Encoded side information data ζ CODDecoding (k-2) is performed before inputting the decoded version ζ(k-2) into the synthesis of the HOA representation in step 23, based on the received index set.

number

[0024] In stage or stage 21,

number

number

[0025] In the signal redistribution stage or stage 22,

number

number

number

number

[0026] In the synthesis stage or stage 23, the current frame of the desired entire HOA representation

number

number

number

number

number

[0027] Number 22 is component in PCT / EP2013 / 075559.

number

number

number

number

number

number

number

[0028] <HOA decomposition> In relation to Figure 3, the HOA decomposition process will be described in detail to explain the meaning of spatial prediction therein. The process is derived from the process described in relation to Figure 3 of patent application PCT / EP2013 / 075559.

[0029] Firstly, the smoothed directional signal X DIR (k-1) and its HOA representation C DIR (k-1) is the long frame of the input HOA representation in stage 31 or stage 31.

number

number

number

number

[0030] In stage 33, the original HOA representation [C(k-1) with a tilde] and the dominant directional signal HOA representation C DIR The residual between (k-1) and O directional signals

number

[0031] In stage 34, these directional signals are dominant directional signal X DIR Predicted from (k-1). Predicted signal

number

number

[0032] In stage 35, the predicted directional signal

number

number

[0033] In stage 37, the original HOA representation [C(k-2) with a tilde] and the dominant directional signal HOA representation C DIR HOA representation of predicted directional signals from directions uniformly distributed in (k-2)

number

[0034] The required signal delay in the processing shown in Figure 3 is achieved by the corresponding delays 381 and 387.

[0035] Spatial prediction The goal of spatial prediction is to obtain 0 residual signals.

number

number

[0036] Each residual signal

number

[0037] Each direction signal

number

[0038] To illustrate the meaning of spatial prediction with an example, consider the decomposition of an HOA representation of order N=3. Here, the maximum number of directions to be extracted is equal to D=4. For simplicity, we further assume that only directional signals with indices 1 and 4 are active, while directional signals with indices 2 and 3 are inactive. Furthermore, for simplicity, we assume that the dominant source direction is constant for the frames considered, i.e., for d=1,4, Ω ACT,d (k-3) = Ω ACT,d (k-2) = Ω ACT,d (k-1) = Ω ACT,d (k=Ω) ACT,d (5) It is assumed that such a general plane wave exists, resulting in a spatially dispersed general plane wave with an order of N=3.

number

[0039] <Parameters of current technologies for describing spatial prediction> One method for describing spatial prediction is presented in the aforementioned ISO / IEC Non-Patent Document 1. Non-Patent Document 1 describes signals

number

[0040] ·Element p TYPE,qA vector p consisting of (k-1), q=1, ..., O TYPE (k-1) is the q-th direction Ω q This indicates whether a prediction will be made, and if so, what type of prediction it is. The meanings of the above elements are as follows: p TYPE,q (k-1)=0 direction Ω q If there is no prediction =1 directionΩ q For full bandwidth prediction (6) =2 directionΩ q For low-frequency predictions.

[0041] ·Element p IND,d,q (k-1), d=1,…,D PRED A matrix P consisting of q=1,...,O IND (k-1) is the direction Ω from the corresponding direction signal. q This represents the index for which a prediction must be made regarding the direction Ω. q If there is no prediction to be made about matrix P, IND The corresponding column of (k-1) consists of 0. Furthermore, direction Ω q The directional signal used for prediction is D PRED If the number is less than P IND The unnecessary element in the q-th column of (k-1) is also 0.

[0042] • Corresponding quantized predictor p Q,F,d,q (k-1), d=1,…,D PRED Matrix P containing q=1,...,O Q,F (k-1).

[0043] The following two parameters need to be known on the decoding side in order to enable proper interpretation of these parameters: ·General plane wave signal

number

[0044] These two parameters must be set to known fixed values ​​in the encoder and decoder, or transmitted additionally, but at a significantly lower frequency than the frame rate. The latter option may be used to adapt the two parameters to the HOA representation to be compressed. An example of the parameter set is O=16, D PRED =2, B SC Setting it to =8, something like this would also be fine.

[0045]

number

number

number

number

number

[0046] Given this side information, the prediction is assumed to be performed as follows:

[0047] First, the quantized predictor p Q,F,d,q (k - 1), d = 1, …, D PRED , q = 1, …, O are dequantized to give the actual predictors.

[0048]

Number

[0049] For the example described above, if B SC = 8, the following is obtained as a result of the dequantized predictor vector.

[0050]

Number

[0051] The signal predicted as a signal

Number

Number

Number

Number

[0052] has already been described and as can be seen from Equation (17) now, the signal

Number

[0053] <Encoding of side information related to spatial prediction using current technologies> In the above-mentioned ISO / IEC Non-Patent Document 1, the coding of side information for spatial prediction is dealt with. It is summarized in Algorithm 1 depicted in FIG. 5 and will be described below. To make the presentation clearer, the frame index k - 1 is ignored in all equations.

[0054] First, a bit array ActivePred consisting of O bits is generated. Here, the bit ActivePred[q] indicates whether prediction is to be performed for direction Ω q or not. The number of "1"s in this array is represented by NumActivePred.

[0055] Next, a bit array PredType of length NumActivePred is generated. Here, each bit indicates the type of prediction, i.e., full-band or low-pass, for the direction for which prediction is to be performed. At the same time, an unsigned integer array PredDirSigIds of length NumActivePred·D PRED is generated. Its elements represent the D PRED indices of the directional signals to be used for each active prediction. D REPDIf fewer than one directional signal is used for prediction, the index is assumed to be set to 0. Each element of the array PredDirSigIds is:

number

[0056] Finally, an integer array QuantPredGains of length NumNonZeroIds is generated. Its elements are the quantized scaling factors p that should be used in equation (17). Q,F,d,q It is assumed to represent (k-1). The corresponding dequantized scaling factor p F,d,q The dequantization to obtain (k-1) is given in equation (10). Each element of the array QuantPredGains is B SC It is assumed that it will be represented by bits.

[0057] Ultimately, the encoded representation of the side information ζ COD teeth, ζ COD =[ActivePred PredType PredDirSigIds QuantPredGains] (19) It consists of the above four sequences.

[0058] To illustrate this coding with an example, the coded representations of equations (7) through (9) are used:

number

[0059] <Encoding of side information related to spatial prediction according to the present invention> To improve the efficiency of encoding side information related to spatial prediction, the processing of current technologies is modified to be more advantageous.

[0060] A) When encoding a typical sound scene HOA representation, the inventors observed that there are often frames in which the decision is made not to perform any spatial prediction during the HOA compression process. However, in such frames, the bit array ActivePred consists only of 0s, and the number of 0s is equal to 0. Since such frame contents occur very frequently, the processing of the present invention is for the encoded representation ζ COD A single bit, PSPredictionActive, is prepended to indicate whether any prediction should be made. If the value of the bit PSPredictionActive is 0 (or "1" in the alternative example), the array ActivePred and any further data related to the prediction are encoded side information ζ COD It cannot be included. In practice, this process is ζ COD This reduces the average bitrate for transmission over time.

[0061] B) A further observation made when encoding the HOA representation of a typical sound scene is that the number of active predictions, NumActivePred, is often very small. In such situations, each direction Ω q Instead of using the bit array ActivePred to indicate whether a prediction is being made, it can sometimes be more efficient to transmit or transfer the number of active predictions and their respective indices. In particular, this variant encoding the active ones has a value of NumActivePred≦M M It is more efficient in this case. Here, M M is the largest integer that satisfies the following equation.

[0062]

number

[0063] In equation (25),

number

number

[0064] As explained above, a single bit KindOfCodedPredIds can be used to indicate how the index of the direction in which the prediction is to be performed is encoded. If the bit KindOfCodedPredIds has a value of "1" (or "0" in the alternative example), then the number NumActivePred and the array PredIds containing the index of the direction in which the prediction is to be performed are encoded side information ζ COD It is added to. Instead, if the bit KindOfCodedPredIds has the value "0" (or "1" in the alternative example), the array ActivePred is used to encode the same information. On average, this behavior is ζ COD The bitrate for transmission decreases over time.

[0065] C) To further improve the side information coding efficiency, the fact that the number of active directional signals actually available for prediction is often less than D is taken advantage of. This is because, for coding each element of the index array PredDirSigIds,

number

number

number

number

number

number

number

[0066] As a result of the above modifications A) to C) for known side information coding processes, the exemplary coding process shown in Figure 6 is obtained.

[0067] As a result, the encoded side information consists of the following components:

number

[0068] The encoded representations for the examples of formulas (7) through (9) are as follows:

[0069]

number

[0070] Advantageously, the representation encoded according to the present invention requires 8 bits less than the encoded representation of the current technology in equations (20) to (23).

[0071] It is also possible for the encoder not to provide the bit array PredType.

[0072] <Decoding of modified side information coding related to spatial prediction> The decoding of modified side information related to spatial prediction is summarized in the exemplary decoding processes shown in Figures 7 and 8 (the process shown in Figure 8 is a continuation of the process shown in Figure 7), and will be explained below.

[0073] First, vector p TYPE and matrix P IND and P Q,F All elements are initialized to 0. Then the bit PSPredictionActive is read. This indicates whether spatial prediction is performed at all. If spatial prediction is performed (i.e., PSPredictionActive=1), the bit KindOfCodedPredIds is read. This indicates the type of coding for the index in the direction in which the prediction should be performed.

[0074] If KindOfCodedPredIds=0, the bit array ActivePred of length O is read. The q-th element of this array is in direction Ω. qThis indicates whether a prediction will be made for a given direction. In the next step, the number of predictions, NumActivePred, is calculated from the array ActivePred, and a bit array PredType of length NumActivePred is read. The elements of this array indicate the type of prediction that should be made for each relevant direction. Using the information contained in ActivePred and PredType, the vector p TYPE The elements are calculated.

[0075] The bit array PredType is not provided on the encoder side; instead, a vector p is obtained from the bit array ActivePred. TYPE It is also possible to calculate the elements.

[0076] If KindOfCodedPredIds=1,

number

number

[0077] The encoder does not provide the bit array PredType, but instead uses a vector p from the NumActivePred and data array PredIds. TYPE It is also possible to calculate the elements.

[0078] In either case (i.e., KindOfCodedPredIds=0 and KindOfCodedPredIds=1), at the next step, NumActivePred·D PRED The array PredDirSigIds, which consists of 1 element, is read. Each element is

number

[0079]

number

[0080] Finally, each B SC The array QuantPredGains, consisting of NumNonZeroIds elements encoded by bits, is read. IND And using the information contained in QuantPredGains, matrix P Q,F The element is set.

[0081] The process of the present invention can be performed by a single processor or electronic circuit, or by several processors or electronic circuits operating in parallel and / or acting on different parts of the process of the present invention.

[0082] Several aspects are described below. [Aspect 1] A method for improving the encoding of side information required to encode a higher-order ambisonic representation (HOA) of a sound field, having an input time frame of HOA coefficient sequences, wherein a dominant directional signal and residual ambient HOA component are determined, a prediction is used for the dominant directional signal, thereby providing side information data (ζ(k-2)) describing the prediction for an encoded frame of HOA coefficients, wherein the side information data (ζ(k-2)) is: • A bit array (ActivePred) indicating whether or not a prediction will be made in a certain direction; • A data array (PredDirSigIds) containing elements representing the indices of the directional signals to be used for the prediction to be performed; It can include a data array (QuantPredGains) that has elements representing quantized scaling factors. The method in question is: • Provides a bit value (PSPredictionActive) indicating whether the prediction should be executed (19;34,384); If there is no prediction to be made, the bit array and data array in the side information data (ζ(k-2)) are omitted; If the prediction should be performed, instead of the bit array (ActivePred) indicating whether the prediction should be performed in a certain direction, a bit value (KindOfCodedPredIds) is provided indicating whether the number of active predictions (NumActivePred) and a data array (PredIds) containing the index of the direction in which the prediction should be performed should be included in the side information data (ζ(k-2)). A method including steps. [Aspect 2] An apparatus for improving the encoding of side information required to encode a higher-order ambisonic representation (HOA) of a sound field, having an input time frame of HOA coefficient sequences, wherein a dominant directional signal and residual ambient HOA component are determined, a prediction is used for the dominant directional signal, and thereby provides side information data (ζ(k-2)) describing the prediction for the encoded frame of HOA coefficients, wherein the side information data (ζ(k-2)) is: • A bit array (ActivePred) indicating whether or not a prediction will be made in a certain direction; • A data array (PredDirSigIds) containing elements representing the indices of the directional signals to be used for the prediction to be performed; It can include a data array (QuantPredGains) that has elements representing quantized scaling factors. The device in question: • Provide a bit value (PSPredictionActive) indicating whether the prediction should be performed; If there is no prediction to be made, the bit array and data array in the side information data (ζ(k-2)) are omitted; If the prediction should be performed, instead of the bit array (ActivePred) indicating whether the prediction should be performed in a certain direction, a bit value (KindOfCodedPredIds) is provided indicating whether the number of active predictions (NumActivePred) and a data array (PredIds) containing the index of the direction in which the prediction should be performed should be included in the side information data (ζ(k-2)). Apparatus including means (19;34,384). [Aspect 3] In the encoding of the HOA representation, estimation of the dominant sound source direction (13) is performed, and a data set of indices of the detected directional signals is obtained.

number

number

number

number

number

Number

number

number

number

Claims

1. A method for decoding a bitstream containing an encoded HOA representation, the method being: The stage of evaluating the value of the bit KindOfCodedPredIds; A step of evaluating a first array ActivePred based on the value of the bit KindOfCodedPredIds, wherein each element of the first array ActivePred indicates whether a prediction is made for the corresponding direction, and when there is an element of ActivePred for the corresponding direction, the variable NumActivePred is incremented, and the variable NumActivePred indicates how many 1s there are in the first array ActivePred; Based on the evaluation of the first sequence ActivePred, vector p type The stage of determining the elements; A step of evaluating a second sequence PredDirSigIds, wherein the elements of the second sequence PredDirSigIds represent indices of directional signals used for active prediction; The aforementioned vector p type And a matrix P representing the index on which a prediction for a certain direction is made from the corresponding directional signal, based on the elements of the second array PredDirSigIds. IND This includes the step of determining the elements of method.

2. A device having a decoder for decoding a bitstream containing an encoded HOA representation, the device: The stage of evaluating the value of the bit KindOfCodedPredIds; A step of evaluating a first array ActivePred based on the value of the bit KindOfCodedPredIds, wherein each element of the first array ActivePred indicates whether a prediction is made for the corresponding direction, and when there is an element of ActivePred for the corresponding direction, the variable NumActivePred is incremented, and the variable NumActivePred indicates how many 1s there are in the first array ActivePred; Based on the evaluation of the first sequence ActivePred, vector p type The stage of determining the elements; A step of evaluating a second sequence PredDirSigIds, wherein the elements of the second sequence PredDirSigIds represent indices of directional signals used for active prediction; The aforementioned vector p type And a matrix P representing the index on which a prediction for a certain direction is made from the corresponding directional signal, based on the elements of the second array PredDirSigIds. IND A processor configured to perform the step of determining the elements of Device.

3. A non-temporary computer-readable medium containing instructions that, when executed by a processor, perform the method described in claim 1.