Encoded HOA data frame representation comprising non-differential gain values associated with channel signals of individual ones of the data frames of the HOA data frame representation

By correlating the value range of HOA representations with maximum gain to determine encoding bits, the HOA data frame compression addresses high bit rate issues, enabling efficient and independent access in HOA decoding.

JP2026012702APending Publication Date: 2026-01-27DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025165095
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-06-27
Filing Date
2025-10-01
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

The high bit rate of Higher Order Ambisonics (HOA) representations poses challenges for practical applications like streaming, and existing compression methods require constraints on the value range of HOA representations that are not adequately defined, making efficient encoding of non-differential gain values difficult for independent access units.

Method used

Establish a correlation between the value range of input HOA representations and the maximum potential gain to determine the required bits for encoding non-differential gain values, ensuring correct compression and decompression by normalizing the HOA representation to a meaningful value range before gain control processing.

Benefits of technology

Enables efficient encoding and decoding of HOA data frames with non-differential gain values, allowing for independent access and reducing the bit rate while maintaining sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012702000001_ABST
    Figure 2026012702000001_ABST
Patent Text Reader

Abstract

To provide a method and apparatus for determining a minimum integer number of bits required to represent non-differential gain values for compression of a Higher Order Ambisonics (HOA) data frame representation.SOLUTION: When compressing the HOA data frame representation in the HOA compressor, a gain control 15, 151 is applied for each channel signal before it is perceptually encoded 16. Those gain values should be encoded with a minimum number of bits, and in order to determine such a lowest integer number of bits (β e), the HOA data frame representation (C (K)) is rendered in the spatial domain to virtual loudspeaker signals on a unit sphere, following which the HOA component is normalized to the directional signal data frame representation (C (K)). Then, the lowest integer number of bits is set to β e = [log2 ([log2 ({√ KMAX}·O)] + 1)].SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an encoded HOA data frame representation that includes non-differential gain values ​​associated with the channel signals of individual ones of the data frames of the HOA data frame representation. [Background technology]

[0002] Higher Order Ambisonics, denoted HOA, offers one possibility for representing three-dimensional sound. Other techniques are wave field synthesis (WFS) or channel-based approaches such as 22.2. In contrast to channel-based methods, HOA representations offer the advantage of being independent of a specific speaker setup. However, this flexibility comes at the cost of the decoding process required to reproduce the HOA representation on a specific speaker setup. Compared to WFS approaches, which typically require a much larger number of speakers, HOA can be rendered to setups with only a few speakers. An additional advantage of HOA is that the same representation can also be used for binaural rendering to headphones without any modification.

[0003] HOA is based on the representation of the spatial density of complex harmonic plane wave amplitudes by a truncated spherical harmonic (SH) expansion. Each expansion coefficient is a function of angular frequency, which can be equivalently represented by a time-domain function. Thus, without loss of generality, we can assume that the complete HOA sound field representation actually consists of O time-domain functions, where O represents the number of expansion coefficients. These time-domain functions are hereafter referred to as HOA coefficient sequences or HOA channels, although they are equivalent.

[0004] The spatial resolution of the HOA representation improves with increasing maximum order of the expansion, N. Unfortunately, the number of expansion coefficients, O, is quadratic with order N, in particular, O = (N + 1) 2For example, a typical HOA representation using order N=4 requires O=25 HOA (expansion) coefficients. The total bit rate for transmission of the HOA representation is multiplied by the desired single-channel sampling rate f S and the number of bits per sample, N b Given that, O·f S N b The HOA expression of degree N=4 is determined by f S = N per sample at a sampling rate of 48 kHz b = 16 bits leads to a bit rate of 19.2 MBits / s, which is too high for many practical applications, e.g. streaming. Thus, compression of the HOA representation is highly desirable.

[0005] Previously, compression of HOA sound field representations has been proposed in Patent Documents 1, 2, and 3 (see Non-Patent Document 1). These approaches have in common that they perform a sound field analysis and decompose the given HOA representation into a directional component and a residual ambient component. On the one hand, the final compressed representation is assumed to consist of several quantized signals resulting from the perceptual coding of directional and vector-based signals and associated coefficient sequences of the ambient HOA component. On the other hand, the final compressed representation includes additional side information related to the quantized signals. This side information is necessary for the reconstruction of the HOA representation from its compressed version.

[0006] Before being passed to the perceptual encoder, these intermediate time-domain signals are required to have a maximum amplitude within the value range [-1, 1[. This is a requirement arising from currently available implementations of perceptual encoders. To meet this requirement when compressing the HOA representation, a gain control processing unit (see Patent Document 4 and the above-mentioned non-patent document 1) is used prior to the perceptual encoder. It smoothly attenuates or amplifies the input signal. The resulting signal modification is assumed to be reversible and applied frame-by-frame. In particular, the change in signal amplitude between successive frames is assumed to be a power of 2. To facilitate reversing this signal modification in the HOA decompressor, corresponding normalized side information is included in the total side information. This normalized side information can consist of base-2 exponents that describe the relative amplitude change between two successive frames. These exponents are coded using run-length codes according to the above-mentioned non-patent document 1, since small amplitude changes between successive frames are more likely than larger changes. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] European Patent Application Publication No. 2665208 [Patent Document 2] European Patent Application Publication No. 2743922 [Patent Document 3] European Patent Application Publication No. 2800401 [Patent Document 4] European Patent Application Publication No. 2824661 [Non-patent literature]

[0008] [Non-Patent Document 1] ISO / IEC JTC1 / SC29 / WG11, N14264, WD1-HOA Text of MPEG-H 3D Audio, January 2014 [Non-patent document 2] J. Fliege, U. Maier, "A two-stage approach for computing cubature formula for the sphere", Technical report, Fachbereich Mathematik, University of Dortmund, 1999 [Non-patent document 3] EG Williams, "Fourier Acoustics", vol.93 of Applied Mathematical Sciences. Academic Press, 1999 [Non-patent document 4] B. Rafaely, "Plane-wave decomposition of the sound field on a sphere by spherical convolution", J. Acoust. Soc. Am., 4(116):2149-2157, October 2004 [Non-Patent Document 5] J. Daniel, "Repr´esentation de champs acoustiques, application `a la transmission et `a la reproduction de sc`enes sonores complexes dans un contexte multim´edia", PhD thesis, Universit´e Paris 6, 2001 Summary of the Invention [Problem to be solved by the invention]

[0009] Using differentially encoded amplitude changes to reconstruct the original signal amplitude in HOA decompression is feasible, for example, when a single file is decompressed from beginning to end without any time jumps. However, to facilitate random access, an independent access unit must exist in the encoded representation (which is typically a bitstream) to allow decompression to begin at a desired position (or at least nearby), independent of information from previous frames. Such an independent access unit must contain the total absolute amplitude change (i.e., the non-differential gain value) caused by the gain control processing unit from the first frame to the current frame. Given that the amplitude change between two consecutive frames is a power of two, it is sufficient to describe the total absolute amplitude change with a base-2 exponent as well. For efficient coding of this exponent, it is essential to know the maximum potential gain of the signal before applying the gain control processing unit. However, this knowledge strongly depends on specifying constraints on the value range of the HOA representation to be compressed. Unfortunately, the MPEG-H 3D Audio document in Non-Patent Document 1 only provides a description of the format for the input HOA representation, but does not set any constraints on the value range.

[0010] The problem to be solved by the present invention is to provide the minimum number of integer bits required to represent a non-differential gain value. This problem is solved in the coded HOA data frame representation disclosed in claim 1. Advantageous further embodiments of the invention are disclosed in the respective dependent claims. [Means for solving the problem]

[0011] The present invention establishes a correlation between the value range of the input HOA representation and the maximum potential gain of the signal before application of the gain control processing unit in the HOA compressor, based on which the amount of bits required is determined for efficient encoding of a base-2 exponent—for a given specification of the value range of the input HOA representation—to describe, within an access unit, the total absolute amplitude change (i.e., the non-differential gain value) of the modified signal caused by the gain control processing unit from the first frame to the current frame.

[0012] Furthermore, once the rules for calculating the amount of bits required for encoding the exponent are fixed, the present invention uses a process to verify whether a given HOA representation satisfies the required value range constraints so that it can be correctly compressed. [Brief explanation of the drawings]

[0013] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates an HOA compressor. [Figure 2] FIG. 1 illustrates an HOA decompressor. [Figure 3] FIG. 10 shows the scaling value K for the virtual directions Ωj(N), 1≦j≦0, for HOA orders N=1,...,29. [Figure 4] Figure 1 shows the Euclidean norm of the inverse modal matrix Ψ-1 for the virtual directions ΩMIN,d (N), d=1,...,OMIN, for HOA order NMIN=1,...,29. [Figure 5] FIG. 1 illustrates the determination of the maximum allowed magnitude γ dB of the signal of a virtual loudspeaker at position Ωj(N), 1≦j≦O, O=(N+1)2. [Figure 6] FIG. 1 is a diagram illustrating a spherical coordinate system. DETAILED DESCRIPTION OF THE INVENTION

[0014] Unless explicitly stated, the following embodiments can be used in any combination or subcombination.

[0015] In the following, the principles of HOA compression and decompression are presented to provide a more detailed context in which the above-mentioned challenges arise. The basis for this presentation is the process described in the MPEG-H 3D Audio document in Non-Patent Document 1. See also Patent Documents 1, 3, and 2. In Non-Patent Document 1, the term "directional component" is extended to "predominant sound component." As a directional component, the dominant sound component is assumed to be represented, in part, by a directional signal, i.e., a monophonic signal with a corresponding direction (from which the sound is assumed to be incident on the listener), together with some prediction parameters for predicting parts of the original HOA representation from the directional signal. In addition, the dominant sound component is assumed to be represented by a "vector-based signal," i.e., a monophonic signal with a corresponding vector that defines the directional distribution of the vector-based signal.

[0016] <HOA Compression> The overall architecture of the HOA compressor described in Patent Document 3 is shown in Figure 1. It has a spatial HOA encoding section depicted in Figure 1A and a perceptual and source encoding section depicted in Figure 1B. The spatial HOA encoder provides a first compressed HOA representation of I signals, along with side information describing how to generate the HOA representation. In the perceptual and source coder, the I signals are perceptually encoded, and the side information is source encoded. The two coded representations are then multiplexed.

[0017] Spatial HOA encoding In the first step, the current kth frame C(k) of the original HOA representation is input to the direction and vector estimation processing step or stage 11. This step is performed using a set of tuples M DIR (k) and M VEC(k) is assumed to provide a tuple set M DIR (k) consists of tuples whose first element represents the index of the directional signal and whose second element represents each quantized direction. VEC (k) consists of tuples whose first element represents the index of vector-based signals and whose second element represents the directional distribution of those signals, i.e., a vector that defines how the HOA representation of the vector-based signals is computed.

[0018] Both tuple sets M DIR (k) and M VEC (k), the initial HOA frame C(k) is the frame X of all dominant (i.e., directional and vector-based) signals in the HOA decomposition stage or stage 12. PS (k-1) and the frame C of the surrounding HOA component AMB (k-1) and (k-1). Note the one-frame delay. This is due to the overlap-add process to avoid blocking artifacts. Furthermore, the HOA decomposition stage / stage 12 is assumed to output some prediction parameters ζ(k-1) that describe how to predict parts of the original HOA representation from these directional signals in order to enrich the dominant HOA components. Furthermore, a target assignment vector v is generated containing information about the assignment of the dominant signals determined in the HOA decomposition processing stage / stage 12 to the I available channels. A,T (k-1) are assumed to be provided. The affected channels can be assumed to be occupied, i.e., they are not available to transmit any coefficient sequences of the ambient HOA components in each time frame.

[0019] In the ambient component correction processing stage or stage 13, the ambient HOA component frame C AMB (k-1) is the target allocation vector v A,T(k-1). In particular, which coefficient sequences of the ambient HOA components should be transmitted in the given I channels is determined (among other aspects) by the target allocation vector v A,T (contained in k-1). Furthermore, if the index of the selected coefficient sequence changes between successive frames, a fade-in and fade-out of the coefficient sequence is performed.

[0020] Furthermore, the ambient HOA component C AMB First O of (k-2) MIN It is assumed that a sequence of coefficients is always chosen to be perceptually coded and transmitted, where O MIN =(N MIN +1) 2 and N MIN ≦N is typically a smaller order than that of the original HOA representation. In order to decorrelate these HOA coefficient sequences, they are then correlated in step / stage 13 with some predefined directions Ω MIN,d , d=1,…,O MIN can be converted from the incident directional signal (i.e., a general plane wave function).

[0021] Corrected ambient HOA component C M,A (k-1) and the temporally predicted corrected ambient HOA component C at stage 13. P,M,A (k-1) is calculated and used in the gain control processing stage or stage 15, 151 to allow for reasonable look-ahead. Here, the information about the ambient HOA component correction is directly related to the assignment of all possible types of signals to the available channels in the channel assignment stage or stage 14. The final information about the assignment is the final assignment vector v A (k-2). To calculate this vector in step / stage 13, the target allocation vector v A,TThe information contained in (k-1) is utilized.

[0022] The channel allocation in phase / stage 14 is performed using the allocation vector

number

number

number

number

number

number

[0023] Signal Frame y i (k-2), i=1,...,I, are finally processed by the gain control 15, 151 to obtain the index e i (k-2) and exception flag β i (k-2), i=1,...,I ​​and signal z i(k-2), i=1,...,I, where the signal gain is smoothly modified to achieve a range of values ​​suitable for the perceptual encoder stage or stage 16. The stage / stage 16 generates the corresponding encoded signal frame

number

number

number

number

number

[0024] In the spatial HOA decoder, the gain correction in the stage 15, 151 is performed by the exponent e i (k-2) and exception flag β i(k-2), i=1, . . . , I.

[0025] <HOA decompression> The overall architecture of the HOA decompressor described in US Patent No. 6,239,999 is shown in Figure 2. It consists of the reversed counterparts of the components of the HOA compressor described above, including a perceptual and source decoding section depicted in Figure 2A, and a spatial HOA decoding section depicted in Figure 2B.

[0026] In the perceptual and source decoding section (which stands for perceptual and source decoder), a demultiplexing stage or stage 21 demultiplexes input frames from the bitstream.

number

number

number

number

number

number

[0027] Spatial HOA Decoding In the spatial HOA decoding section, the perceptually decoded signal

number

number

[0028] I gain-corrected signal frames

number

number

[0029] In the dominant synthesis stage or stage 26, the dominant component

number

[0030] In the ambient synthesis stage or stage 27, the ambient HOA component frame

number

[0031] A spatial HOA decoder then generates the reconstructed HOA representation from the I signals and the side information.

[0032] If the ambient HOA components are transformed into directional signals at the encoder side, the transformation is reversed in stage 27 at the decoder side.

[0033] The maximum potential gain of a signal before the gain control processing stage 15, 151 in the HOA compressor strongly depends on the value range of the input HOA expression. Therefore, first a meaningful value range for the input HOA expression is defined, and then the maximum potential gain of the signal before entering the gain control processing stage is concluded.

[0034] Normalization of input HOA expressions To use the inventive process, a normalization of the (entire) input HOA representation signal is performed beforehand. For HOA compression, a frame-by-frame process is performed, where the k-th frame C(k) of the original input HOA representation is expressed as

number

[0035] As stated in Patent Document 4, a meaningful normalization of the HOA representation from a practical point of view is to normalize the individual HOA coefficient sequences c n m This is not achieved by imposing constraints on the value range of (t), since these time-domain functions are not the signals that will actually be played by the speakers after rendering. Instead, we can map the HOA representation to O virtual speaker signals w j It is more convenient to think of the equivalent spatial domain representation obtained by rendering (t) into 1 ≤ j ≤ O. Each virtual speaker position is assumed to be represented by a spherical coordinate system, where each position is assumed to lie on a unit sphere and have a radius of 1. These positions are then bounded by order-dependent directions Ω j (N) =(θ j (N) ,φ j (N) ), 1≦j≦0, where θ j (N) and φ j (N) and represent the tilt angle and azimuth angle, respectively (see Fig. 6 and its explanation for the definition of the spherical coordinate system). These directions should be distributed as uniformly as possible on the unit sphere. For example, see Non-Patent Document 2. For the calculation of a specific direction, the number of nodes is fliege / nodes / nodes.html. Their location generally depends on the definition of "spherical uniform distribution" and is therefore not unambiguous.

[0036] The advantage of defining a value range for the virtual speaker signal over defining a value range for the HOA coefficient sequence is that the value range for the former can be intuitively set equal to the interval [-1, 1[, as is the case for normal speaker signals assuming a PCM representation. This leads to a spatially uniformly distributed quantization error, and so the quantization is advantageously applied in a region that is meaningful with respect to actual listening. An important aspect in this context is that the number of bits per sample can be chosen as low as is typically the case for normal speaker signals, e.g., 16, whereas normally a higher number of bits per sample (e.g., 24 or even 32) would be needed. This increases efficiency compared to direct quantization of the HOA coefficient sequence.

[0037] To describe the normalization process in detail in the spatial domain, all virtual speaker signals are expressed as w(t):=[w1(t) … w O (t)] T (2) Here, (·) T represents the transposition. The virtual direction Ω j (N) , the mode matrix for 1≦j≦O is

number

[0038] Using these definitions, reasonable requirements for a virtual speaker signal are:

number

[0039] As a result, the total power of the speaker signal is

number

[0040] Consequences for signal value range before gain control Assuming that the normalization of the input HOA expression is performed as described in the section "Normalization of the Input HOA Expression", the signal y i , i=1,...,l are considered below. These signals are the HOA coefficient sequences or the dominant tone signal x PS,d , d=1,…,D and / or the ambient HOA component c AMB,n , n=1,...,O, by assigning one or more of the specific coefficient sequences (some of which have undergone spatial transformations) to the available I channels. Therefore, it is necessary to analyze the possible value ranges of the different signal types listed here, under the normalization assumption in equation (6). Since all types of signals are intermediately calculated from the original HOA coefficient sequence, we will look at their possible value ranges.

[0041] The case where only one or more HOA coefficient sequences are included in I channels is not depicted in Figures 1A and 2B, i.e., in such cases, HOA decomposition, ambient correction and the corresponding synthesis block are not required.

[0042] Consequences for the value range of HOA expressions The time-continuous HOA representation is derived from the virtual speaker signal. c(t)=Ψw(t) (8) This is the inverse operation of equation (5). Therefore, the total power of all HOA coefficient sequences is bounded using equations (8) and (7) as follows:

[0043]

number

[0044] FIG. 3 shows the virtual direction Ω according to the paper of Non-Patent Document 2 mentioned above. j (N) , 1≦j≦0, and the values ​​of K are shown for HOA orders N=1,...,29.

[0045] Combining all previous discussions and considerations, an upper bound on the absolute value of the HOA coefficient sequence is given as follows:

[0046]

number

[0047] It is important to note that the condition in equation (6) implies the condition in equation (11), but the converse is not true, i.e., equation (11) does not imply equation (6).

[0048] A further important aspect is that, under the assumption of nearly uniformly distributed virtual speaker positions, the column vectors of the modal matrix Ψ, which represent the mode vectors with respect to the virtual speaker positions, are nearly orthogonal to each other and each have a Euclidean norm of N+1. This property means that the spatial transformation nearly preserves the Euclidean norm except for a multiplicative constant. That is,

number

[0049] Consequences for the value range of the dominant signal Both types of dominant sound signals (directional and vector-based) contribute to the HOA representation with Euclidean norm N+1, i.e. ||v1||2=N+1 (13) a single vector v1∈R such that O They have in common that they are described by

[0050] In the case of directional signals, this vector is directed to a certain source direction Ω S,1 corresponds to the mode vector with respect to

number

[0051] In the following, we consider D dominant sound signals x d The general case of (t), d=1,…,D is considered. These signals are x(t)=[x1(t) x2(t) … x D (t)] T (16) These signals can be collected into a vector x(t) according to the monophonic dominant signal x d (t), all vectors v representing directional distributions for d=1,…,D d , d=1,…,D V:=[v1v2… v D ] (17) The decision must be based on the following:

[0052] For meaningful extraction of the dominant sound signal x(t), the following constraints are formulated: a) Each dominant signal is obtained as a linear combination of the coefficient sequences of the original HOA representation, i.e. x(t)=A·c(t) (18) where A∈R D×O represents the mixing matrix. b) The mixing matrix A has a Euclidean norm that does not exceed the value 1, i.e.

number

number

[0053] Substituting equation (18) into equation (20), equation (20) becomes the constraint

number

[0054] From the constraints in equations (18) and (19) and the consistency of the Euclidean matrix and vector norm, the upper bound on the absolute value of the dominant sound signal can be obtained by using equations (18), (19) and (11):

number

number

[0055] Example for choosing a confusion matrix An example of how to determine the mixing matrix that satisfies the constraint (20) is the one that minimizes the Euclidean norm of the residuals after extraction, i.e.

number

[0056] Nevertheless, the matrix V still satisfies the constraint (19), i.e.

number

[0057] In the case of directional signals only, the matrix V is a function of several source signal directions Ω S,d , the modal matrix for d=1,…,D, i.e.

number

[0058] Consequences for the range of coefficient sequences of the surrounding HOA components The ambient HOA component is calculated by subtracting the HOA representation of the dominant sound signal from the original HOA representation, i.e.,

number

number

[0059] <Value range of the spatially transformed coefficient sequence of the ambient HOA components> A further aspect of the HOA compression process proposed in the MPEG documents of Patent Document 2 and the above-mentioned Non-Patent Document 1 is that the first O MIN coefficient sequences are always chosen to be assigned to a transport channel, where O MIN =(N MIN +1) 2 and N MIN≦N is typically a smaller order than the order of the original HOA representation. To decorrelate these HOA coefficient sequences, they are normalized to some predefined direction Ω (similar to the concept described in the section Normalizing the Input HOA Representation). MIN,d , d=1,…,O MIN can be converted into an incident virtual speaker signal from the order index n≦N MIN Let c be the vector of all coefficient sequences of the surrounding HOA components. AMB,MIN (t) and the virtual direction Ω MIN,d , d=1,…,O MIN The mode matrix for MIN When defined by MIN The vector of all virtual speaker signals (defined by) (t) is

number

[0060] Therefore, using the consistency of Euclidean matrices and vector norms,

number

[0061] In the MPEG document of the above-mentioned non-patent document 1, the virtual direction Ω MIN,d , d=1,…,O MIN is chosen according to the paper in Non-Patent Document 2 mentioned above. The modal matrix Ψ MIN The Euclidean norm of each inverse matrix of MIN 4 for =1,...,9.

[0062]

number

[0063] However, N MIN>9, this is not true in general. In this case, ||Ψ MIN -1 The value of ||2 is typically much larger than 1. Nevertheless, at least 1 ≤ N MIN For ≦9, the amplitude of the virtual speaker signal is limited by

[0064]

number

[0065] Maximum degree of interest N MAX Any degree N up to 1≦N≦N MAX For the amplitude of the signal before gain control, the value (√K MAX )·O. Here,

number

[0066] K MAX is the maximum order of interest, N MAX and the virtual speaker direction Ω j (N) , 1≦j≦O, and can be expressed as follows:

[0067]

number

number

[0068] If the amplitude of the signals before gain control is too small, the MPEG document in Non-Patent Document 1 states that these amplitudes are

number

[0069] Thus, each exponent to base 2 describing the absolute amplitude change of the total modified signal caused by the gain control processing unit from the beginning to the current frame in access units is in the interval [e MIN ,e MAX ]. As a result, the (lowest integer) number of bits required to encode it, β e is given by the following equation:

[0070]

number

[0071]

number

[0072] This number of bits for the exponent, β e Using ,guarantees that all possible absolute amplitude changes caused by the HOA compressor gain control processing units 15,...,151 can be captured, and allows decompression to begin at some predefined entry point within the compressed representation.

[0073] At the HOA decompressor, when starting to decompress the compressed HOA representation, it represents the total absolute amplitude change allocated to the side information for several data frames and is used to decompress the received data stream.

number

[0074] Further Embodiments When implementing a specific HOA compression / decompression system as described in sections HOA Compression, Spatial HOA Encoding, HOA Decompression, and Spatial HOA Decoding, the amount of bits β for encoding the exponent is e is the scaling factor K MAX,DESThis K MAX,DES The desired maximum degree N of the HOA representation to be compressed MAX,DES and some kind of virtual speaker direction

number

[0075] For example, N MAX,DES = 29, and when choosing the virtual speaker direction according to the paper in Non-Patent Document 2, the rational choice is √K MAX,DES = 1.5. In this situation, the same virtual speaker direction Ω DES,1 (N) ,…,Ω DES,O (N) are normalized according to the section on Normalizing Input HOA Expressions, where 1≦N≦N MAX Correct compression is guaranteed for an HOA representation of order N such that: However, this guarantee does not apply to the virtual speaker direction Ω, which is equivalently represented by a virtual speaker signal, also in PCM format (for efficiency reasons). j (N) , 1≦j≦O is the virtual speaker direction Ω assumed in the system design stage. DES,1 (N) ,…,Ω DES,O (N) In the case of HOA expressions chosen differently from

[0076] Due to this different choice of virtual speaker positions, even if these virtual speaker signals are in the interval [1,1[, the amplitude of the signal before gain control will be different from the value (√K MAX,DES )·O. Therefore, it cannot be guaranteed that this HOA representation has the correct normalization for compression according to the process described in the MPEG document [1].

[0077] In this situation, it is advantageous to have a system that provides the maximum allowable amplitude of the virtual speaker signals based on knowledge of the virtual speaker positions, in order to ensure that each HOA representation is suitable for compression according to the process described in the MPEG document [1]. Such a system is shown in Figure 5. This requires O=(N+1) 2 , where N∈N0, and the input is the virtual speaker position Ω j (N) , 1≦j≦O, and output the maximum allowed amplitude (measured in decibels) of the virtual speaker signal γ dB In step or stage 51, the modal matrix Ψ for the virtual speaker positions is calculated according to equation (3). In the following step or stage 52, the Euclidean norm ||Ψ||2 of the modal matrix is ​​calculated. In the third step or stage 53, the amplitude γ is calculated as 1 and the square root of the number of virtual speaker positions and K MAX,DES and the Euclidean norm of the modal matrix.

number

[0078] To illustrate: from the above derivation, the magnitude of the HOA coefficient sequence is given by the value (√K MAX,DES )·O, i.e.

number

[0079] From equation (9), the magnitude of the HOA coefficient sequence is

number

number

number

[0080] That is, the maximum magnitude value 1 in equation (6) is replaced by the maximum magnitude value γ in equation (47).

[0081] <Higher-Order Ambisonics Basics> Higher Order Ambisonics (HOA) is based on the description of the sound field in a compact region of interest where no sound sources are assumed. In that case, the spatiotemporal behavior p(t,x) of the sound pressure at position x and time t in the region of interest is completely determined physically by the homogeneous wave equation. In the following, we consider the spherical coordinate system shown in Figure 6. In this coordinate system used, the x-axis points to the forward position, the y-axis points to the left, and the z-axis points up. At a position x=(r,θ,φ) in space T is expressed by the radius r>0 (i.e., the distance to the coordinate origin), the tilt angle θ∈[0,π] measured from the polar axis z, and the azimuthal angle φ∈[0,2π[ measured counterclockwise from the x-axis in the xy plane. Furthermore, (·) T represents the transposition.

[0082] Then, assuming that ω represents the angular frequency and i represents the imaginary unit, from the textbook in Non-Patent Document 3, F t The Fourier transform of the sound pressure with respect to time is represented by (·), i.e.

number

number

[0083] If a sound field is represented by a superposition of infinite harmonic plane waves of different angular frequencies ω arriving from all possible directions specified by the angle tuple (θ,φ), it can be shown that each plane wave complex amplitude function C(ω,θ,φ) can be expressed by the following spherical harmonic expansion (Non-Patent Document 4):

[0084]

number

number

number

[0085] HOA coefficient sequence c in vector c(t) n m The position index of (t) is n(n+1)+1+m The total number of elements in the vector c(t) is given by O=(N+1) 2 is given by The final Ambisonics format is a sampled version of c(t) using the sampling frequency fs,

number

[0086] Definition of real-valued spherical harmonics real-valued spherical harmonic functions S n m (θ,φ) (assuming SN3D normalization according to Non-Patent Document 5, Chapter 3.1) is given by

[0087]

number

[0088]

number

[0089] The present invention can be carried out by a single processor or electronic circuit, or by several processors or electronic circuits operating in parallel and / or operating on different parts of the processing of the invention.

[0090] Instructions for operating such processor(s) may be stored in one or more memories.

[0091] Several aspects will be described. [Aspect 1] The HOA data frame representation (C(k)) is the non-differential gain value (2 e ) encoded HOA data representation

number

number

number

number

number

number

number

number

Claims

1. 1. A method of decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, comprising: The compressed HOA representation is expressed as side information and bit number β e and decoding the data based on the number of bits β e teeth [Equation 1] is determined based on [Equation 2] , N is the order of the compressed HOA representation, and N MAX is the maximum degree of interest of the compressed HOA representation, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between O and e MAX >0, √K MAX = 1.5, and rendering the decoded HOA representation; A method comprising:

2. 1. An apparatus for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, the apparatus comprising: The compressed HOA representation is expressed as side information and bit number β e a processor configured to decode based on the number of bits β e teeth [Equation 3] is determined based on [Equation 4] where N is the order of the compressed HOA representation, and N MAX is the maximum degree of interest of the compressed HOA representation, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between O and e MAX >0, √K MAX = 1.5; and A renderer that renders the decompressed HOA representation An apparatus having:

3. A storage device containing instructions that, when executed by a processor, perform the method of claim 1.

4. 1. A device for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, comprising: a processor; a storage device containing instructions that, when executed by said processor, perform the method of claim 1; A device having

5. 10. A computer program product having instructions which, when executed by a computing device or system, cause the computing device or system to perform the method of claim 1.

6. To compress the HOA data frame representation (C(k)), we use an exponent of 2 (2 e The minimum number of integer bits β to describe the non-differential gain value representation for amplitude change as e A method for determining The HOA data frame representation (C(k)) is expressed in the spatial domain as O virtual speaker signals w j (t), where the positions of the virtual speakers lie on a unit sphere and are targeted to be uniformly distributed on the unit sphere, and the rendering is performed by matrix multiplication w(t) = (Ψ) -1 represented by c(t), where w(t) is a vector containing all virtual speaker signals, Ψ is the virtual speaker position mode matrix, and c(t) is a vector of the corresponding HOA coefficient sequence in said HOA data frame representation; The HOA data frame representation (C(k)) is [Equation 5] and the method is normalized so that: forming a channel signal, the forming comprising: a) multiplying a vector c(t) of HOA coefficient sequences by a mixing matrix A to represent a dominant sound signal (x(t)) in the channel signal; b) ambient components c in the channel signal AMB (t) by subtracting the dominant sound signal from the normalized HOA data frame representation and subtracting the resulting minimum ambient component c AMB,MIN (t) to w MIN (t)=Ψ MIN -1 ・c AMB,MIN (t) by calculating ||Ψ MIN -1 || 2 < 1, and Ψ MIN is the minimum surrounding component c AMB,MIN (t) is the mode matrix for the stage; c) selecting a portion of the HOA coefficient sequence c(t) that relates to the coefficient sequence of the surrounding HOA components to which a spatial transformation is applied, The method further comprises: The minimum number of integer bits β e of [Equation 6] including making a determination based on [Equation 7] where N is the order and N MAX is the maximum order of interest, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between and , method.

7. 1. A method for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, the method comprising: receiving a bitstream including the compressed HOA representation and decoding the compressed HOA representation to generate a perceptually decoded signal; [Equation 8] Associated gain correction exponent e i (k) and gain correction exception flag β i decoding (k); The perceptually decoded signal ̂z i (k), i = 1,...,I, the associated gain correction exponent e i (k) and the gain correction exception flag β i (k) is a signal frame that has been gain corrected by performing an inverse gain control process on the signal frame. [Equation 9] providing a Frame of dominant signal [Equation 10] and frame C of the intermediate representation of the surrounding HOA components I,AMB (k), the gain-compensated signal frame ^y i (k), i = 1, ..., I, and A method comprising:

Citation Information

Patent Citations

  • Method and apparatus for compressing and decompressing a Higher Order Ambisonics signal representation

    EP2665208A1

  • Method and apparatus for compressing and decompressing a higher order ambisonics representation for a sound field

    EP2743922A1

  • Method and Apparatus for compressing and decompressing a Higher Order Ambisonics representation

    EP2800401A1

  • Method and Apparatus for generating from a coefficient domain representation of HOA signals a mixed spatial / coefficient domain representation of said HOA signals

    EP2824661A1