Apparatus for determining the minimum number of integer bits required to represent a non-differential gain value for compression of an HOA data frame representation

By correlating the value range of HOA representations with maximum gain to determine bit requirements, the method addresses inefficiencies in HOA compression, enabling efficient encoding and decompression of non-differential gain values for independent access.

JP7757471B2Active Publication Date: 2025-10-21DOLBY INTERNATIONAL AB
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2024107100
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-06-27
Filing Date
2024-07-03
Publication Date
2025-10-21
Estimated Expiration
2035-06-22

AI Technical Summary

Technical Problem

Existing HOA compression methods lack efficient encoding of non-differential gain values, which are necessary for independent access to HOA data frames, due to the absence of constraints on the value range of HOA representations in MPEG-H 3D Audio, leading to inefficiencies in bit rate management.

Method used

Establish a correlation between the value range of input HOA representations and the maximum potential gain to determine the minimum number of bits required for encoding non-differential gain values, ensuring correct compression by normalizing the HOA data frames and applying gain control processing units.

Benefits of technology

Enables efficient encoding of non-differential gain values, allowing for independent access to HOA data frames and reducing bit rate requirements, thereby facilitating effective compression and decompression of HOA representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007757471000090
    Figure 0007757471000090
  • Figure 0007757471000091
    Figure 0007757471000091
  • Figure 0007757471000092
    Figure 0007757471000092
Patent Text Reader

Abstract

To provide a device that determines the minimum integer bit number required to express non-differential gain values for compression of higher-order ambisonics (HOA) data frame expressions.SOLUTION: In an HOA compressor, when compressing HOA data frame expressions, before a perceptual encoder stage or stage 16, which perceptually encodes each channel signal, gain control 15, 151 are applied to each channel signal, and absolute gain values necessary to start decoding streamed and the compressed HOA data frame expressions are encoded using the minimum number of bits. In order to determine such a minimum integer bit number (βe), the HOA data frame expressions (C(k)) are rendered in a spatial domain to a virtual speaker signal on a unit sphere, followed by the normalization of directional signal data frame expressions (C(k)) of the HOA component. Then, the minimum integer bit number is set.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus for determining the minimum number of integer bits required to represent a non-differential gain value associated with a channel signal of a particular one of the HOA data frames for compression of the HOA data frame representation. [Background technology]

[0002] Higher Order Ambisonics, denoted HOA, offers one possibility for representing three-dimensional sound. Other techniques are wave field synthesis (WFS) or channel-based approaches such as 22.2. In contrast to channel-based methods, HOA representations offer the advantage of being independent of a specific speaker setup. However, this flexibility comes at the cost of the decoding process required to reproduce the HOA representation on a specific speaker setup. Compared to WFS approaches, which typically require a much larger number of speakers, HOA can be rendered to setups with only a few speakers. An additional advantage of HOA is that the same representation can also be used for binaural rendering to headphones without any modification.

[0003] HOA is based on the representation of the spatial density of complex harmonic plane wave amplitudes by a truncated spherical harmonic (SH) expansion. Each expansion coefficient is a function of angular frequency, which can be equivalently represented by a time-domain function. Thus, without loss of generality, we can assume that the complete HOA sound field representation actually consists of O time-domain functions, where O represents the number of expansion coefficients. These time-domain functions are hereafter referred to as HOA coefficient sequences or HOA channels, although they are equivalent.

[0004] The spatial resolution of the HOA representation improves with increasing maximum order of the expansion, N. Unfortunately, the number of expansion coefficients, O, is quadratic with order N, in particular, O = (N + 1) 2For example, a typical HOA representation using order N=4 requires O=25 HOA (expansion) coefficients. The total bit rate for transmission of the HOA representation is multiplied by the desired single-channel sampling rate f S and the number of bits per sample, N b Given that, O·f S N b The HOA expression of degree N=4 is determined by f S = N per sample at a sampling rate of 48 kHz b = 16 bits leads to a bit rate of 19.2 MBits / s, which is too high for many practical applications, e.g. streaming. Thus, compression of the HOA representation is highly desirable.

[0005] Previously, compression of HOA sound field representations has been proposed in Patent Documents 1, 2, and 3 (see Non-Patent Document 1). These approaches have in common that they perform a sound field analysis and decompose the given HOA representation into a directional component and a residual ambient component. On the one hand, the final compressed representation is assumed to consist of several quantized signals resulting from the perceptual coding of directional and vector-based signals and associated coefficient sequences of the ambient HOA component. On the other hand, the final compressed representation includes additional side information related to the quantized signals. This side information is necessary for the reconstruction of the HOA representation from its compressed version.

[0006] Before being passed to the perceptual encoder, these intermediate time-domain signals are required to have a maximum amplitude within the value range [-1, 1[. This is a requirement arising from currently available implementations of perceptual encoders. To meet this requirement when compressing the HOA representation, a gain control processing unit (see Patent Document 4 and the above-mentioned non-patent document 1) is used prior to the perceptual encoder. It smoothly attenuates or amplifies the input signal. The resulting signal modification is assumed to be reversible and applied frame-by-frame. In particular, the change in signal amplitude between successive frames is assumed to be a power of 2. To facilitate reversing this signal modification in the HOA decompressor, corresponding normalized side information is included in the total side information. This normalized side information can consist of base-2 exponents that describe the relative amplitude change between two successive frames. These exponents are coded using run-length codes according to the above-mentioned non-patent document 1, since small amplitude changes between successive frames are more likely than larger changes. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] European Patent Application Publication No. 2665208 [Patent Document 2] European Patent Application Publication No. 2743922 [Patent Document 3] European Patent Application Publication No. 2800401 [Patent Document 4] European Patent Application Publication No. 2824661 [Non-patent literature]

[0008] [Non-Patent Document 1] ISO / IEC JTC1 / SC29 / WG11, N14264, WD1-HOA Text of MPEG-H 3D Audio, January 2014 [Non-patent document 2] J. Fliege, U. Maier, "A two-stage approach for computing cubature formula for the sphere", Technical report, Fachbereich Mathematik, University of Dortmund, 1999 [Non-patent document 3] EG Williams, "Fourier Acoustics", vol.93 of Applied Mathematical Sciences. Academic Press, 1999 [Non-patent document 4] B. Rafaely, "Plane-wave decomposition of the sound field on a sphere by spherical convolution", J. Acoust. Soc. Am., 4(116):2149-2157, October 2004 [Non-patent document 5] J. Daniel, "Repr´esentation de champs acoustiques, application `a la transmission et `a la reproduction de sc`enes sonores complexes dans un contexte multim´edia", PhD thesis, Universit´e Paris 6, 2001 Summary of the Invention [Problem to be solved by the invention]

[0009] Using differentially encoded amplitude changes to reconstruct the original signal amplitude in HOA decompression is feasible, for example, when a single file is decompressed from beginning to end without any time jumps. However, to facilitate random access, an independent access unit must exist in the encoded representation (which is typically a bitstream) to allow decompression to begin at a desired position (or at least nearby), independent of information from previous frames. Such an independent access unit must contain the total absolute amplitude change (i.e., the non-differential gain value) caused by the gain control processing unit from the first frame to the current frame. Given that the amplitude change between two consecutive frames is a power of two, it is sufficient to describe the total absolute amplitude change with a base-2 exponent as well. For efficient coding of this exponent, it is essential to know the maximum potential gain of the signal before applying the gain control processing unit. However, this knowledge strongly depends on specifying constraints on the value range of the HOA representation to be compressed. Unfortunately, the MPEG-H 3D Audio document in Non-Patent Document 1 only provides a description of the format for the input HOA representation, but does not set any constraints on the value range.

[0010] The problem to be solved by the present invention is to provide the minimum number of integer bits required to represent a non-differential gain value. This problem is solved by the method disclosed in claim 1. Advantageous further embodiments of the invention are disclosed in the respective dependent claims. [Means for solving the problem]

[0011] The present invention establishes a correlation between the value range of the input HOA representation and the maximum potential gain of the signal before application of the gain control processing unit in the HOA compressor, based on which the amount of bits required is determined for efficient encoding of a base-2 exponent—for a given specification of the value range of the input HOA representation—to describe, within an access unit, the total absolute amplitude change (i.e., the non-differential gain value) of the modified signal caused by the gain control processing unit from the first frame to the current frame.

[0012] Furthermore, once the rules for calculating the amount of bits required for encoding the exponent are fixed, the present invention uses a process to verify whether a given HOA representation satisfies the required value range constraints so that it can be correctly compressed.

[0013] In principle, the method of the present invention allows for the compression of HOA data frame representations by determining the minimum integer number of bits β required to represent the non-differential gain value for the channel signal of a particular one of said HOA data frames. e wherein each channel signal in each frame comprises a group of sample values, a differential gain value is assigned to each channel signal in each frame of the HOA data frame, such differential gain value causing a change in amplitude of the sample values ​​of the channel signal in a current HOA data frame relative to the sample values ​​of that channel signal in a immediately preceding HOA data frame, and such gain-adapted channel signals are encoded in an encoder; The HOA data frame representation is expressed as O virtual speaker signals w in the spatial domain. j (t), the positions of the O virtual speakers are on the unit sphere, and β e does not match the position assumed for the calculation of the matrix multiplication w(t)=(Ψ) -1where w(t) is a vector containing all virtual speaker signals, Ψ is the mode matrix calculated for these virtual speaker positions, c(t) is a vector of the corresponding HOA coefficient sequence in said HOA data frame representation, and

number

number

number

number

number

number

[0014] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates an HOA compressor. [Figure 2] FIG. 1 illustrates an HOA decompressor. [Figure 3] FIG. 10 shows the scaling value K for the virtual directions Ωj(N), 1≦j≦0, for HOA orders N=1,...,29. [Figure 4] Figure 1 shows the Euclidean norm of the inverse modal matrix Ψ-1 for the virtual directions ΩMIN,d (N), d=1,...,OMIN, for HOA order NMIN=1,...,29. [Figure 5] FIG. 1 illustrates the determination of the maximum allowed magnitude γ dB of the signal of a virtual loudspeaker at position Ωj(N), 1≦j≦O, O=(N+1)2. [Figure 6] FIG. 1 is a diagram illustrating a spherical coordinate system. DETAILED DESCRIPTION OF THE INVENTION

[0015] Unless explicitly stated, the following embodiments can be used in any combination or subcombination.

[0016] In the following, the principles of HOA compression and decompression are presented to provide a more detailed context in which the above-mentioned challenges arise. The basis for this presentation is the process described in the MPEG-H 3D Audio document in Non-Patent Document 1. See also Patent Documents 1, 3, and 2. In Non-Patent Document 1, the term "directional component" is extended to "predominant sound component." As a directional component, the dominant sound component is assumed to be represented, in part, by a directional signal, i.e., a monophonic signal with a corresponding direction (from which the sound is assumed to be incident on the listener), together with some prediction parameters for predicting parts of the original HOA representation from the directional signal. In addition, the dominant sound component is assumed to be represented by a "vector-based signal," i.e., a monophonic signal with a corresponding vector that defines the directional distribution of the vector-based signal.

[0017] <HOA Compression> The overall architecture of the HOA compressor described in Patent Document 3 is shown in Figure 1. It has a spatial HOA encoding section depicted in Figure 1A and a perceptual and source encoding section depicted in Figure 1B. The spatial HOA encoder provides a first compressed HOA representation of I signals, along with side information describing how to generate the HOA representation. In the perceptual and source coder, the I signals are perceptually encoded, and the side information is source encoded. The two coded representations are then multiplexed.

[0018] Spatial HOA encoding In the first step, the current kth frame C(k) of the original HOA representation is input to the direction and vector estimation processing step or stage 11. This step is performed using a set of tuples M DIR (k) and M VEC (k) is assumed to provide a tuple set M DIR (k) consists of tuples whose first element represents the index of the directional signal and whose second element represents each quantized direction. VEC(k) consists of tuples whose first element represents the index of vector-based signals and whose second element represents the directional distribution of those signals, i.e., a vector that defines how the HOA representation of the vector-based signals is computed.

[0019] Both tuple sets M DIR (k) and M VEC (k), the initial HOA frame C(k) is the frame X of all dominant (i.e., directional and vector-based) signals in the HOA decomposition stage or stage 12. PS (k-1) and the frame C of the surrounding HOA component AMB (k-1) and (k-1). Note the one-frame delay. This is due to the overlap-add process to avoid blocking artifacts. Furthermore, the HOA decomposition stage / stage 12 is assumed to output some prediction parameters ζ(k-1) that describe how to predict parts of the original HOA representation from these directional signals in order to enrich the dominant HOA components. Furthermore, a target assignment vector v is generated containing information about the assignment of the dominant signals determined in the HOA decomposition processing stage / stage 12 to the I available channels. A,T (k-1) are assumed to be provided. The affected channels can be assumed to be occupied, i.e., they are not available to transmit any coefficient sequences of the ambient HOA components in each time frame.

[0020] In the ambient component correction processing stage or stage 13, the ambient HOA component frame C AMB (k-1) is the target allocation vector v A,T (k-1). In particular, which coefficient sequences of the ambient HOA components should be transmitted in the given I channels is determined (among other aspects) by the target allocation vector v A,T(k-1)). Furthermore, if the index of the selected coefficient sequence changes between successive frames, a fade-in and fade-out of the coefficient sequence is performed.

[0021] Furthermore, the ambient HOA component C AMB First O of (k-2) MIN It is assumed that a sequence of coefficients is always chosen to be perceptually coded and transmitted, where O MIN =(N MIN +1) 2 and N MIN ≦N is typically a smaller order than that of the original HOA representation. In order to decorrelate these HOA coefficient sequences, they are then correlated in step / stage 13 with some predefined directions Ω MIN,d , d=1,…,O MIN can be converted from the incident directional signal (i.e., a general plane wave function).

[0022] Corrected ambient HOA component C M,A (k-1) and the temporally predicted corrected ambient HOA component C at stage 13. P,M,A (k-1) is calculated and used in the gain control processing stage or stage 15, 151 to allow for reasonable look-ahead. Here, the information about the ambient HOA component correction is directly related to the assignment of all possible types of signals to the available channels in the channel assignment stage or stage 14. The final information about the assignment is the final assignment vector v A (k-2). To calculate this vector in step / stage 13, the target allocation vector v A,T The information contained in (k-1) is utilized.

[0023] The channel allocation in phase / stage 14 is calculated using the allocation vector v A Using the information given by (k-2), frame X PS(k-2) contains the appropriate signal and frame C M,A (k-2) to the I available channels, and i (k-2), i=1,...,I. Furthermore, frame X PS (k-1) and frame C P,AMB The appropriate signals contained in (k-1) are also assigned to the I available channels to generate the predicted signal frame y P,i (k-2), i=1,…,I.

[0024] Signal Frame y i (k-2), i=1,...,I, are finally processed by the gain control 15, 151 to obtain the index e i (k-2) and exception flag β i (k-2), i=1,...,I ​​and signal z i (k-2), i=1,...,I, where the signal gain is smoothly modified to achieve a range of values ​​suitable for the perceptual encoder stage or stage 16. The stage / stage 16 generates the corresponding encoded signal frame

number

number

number

number

number

[0025] In the spatial HOA decoder, the gain correction in the stage 15, 151 is performed by the exponent e i (k-2) and exception flag β i (k-2), i=1, . . . , I.

[0026] <HOA decompression> The overall architecture of the HOA decompressor described in US Patent No. 6,239,999 is shown in Figure 2. It consists of the reversed counterparts of the components of the HOA compressor described above, including a perceptual and source decoding section depicted in Figure 2A, and a spatial HOA decoding section depicted in Figure 2B.

[0027] In the perceptual and source decoding section (which stands for perceptual and source decoder), a demultiplexing stage or stage 21 demultiplexes input frames from the bitstream.

number

number

number

number

number

number

[0028] Spatial HOA Decoding In the spatial HOA decoding section, the perceptually decoded signal

number

number

[0029] I gain-corrected signal frames

number

number

[0030] In the dominant synthesis stage or stage 26, the dominant component

number

[0031] In the ambient synthesis stage or stage 27, the ambient HOA component frame

number

[0032] A spatial HOA decoder then generates the reconstructed HOA representation from the I signals and the side information.

[0033] If the ambient HOA components are transformed into directional signals at the encoder side, the transformation is reversed in stage 27 at the decoder side.

[0034] The maximum potential gain of a signal before the gain control processing stage 15, 151 in the HOA compressor strongly depends on the value range of the input HOA expression. Therefore, first a meaningful value range for the input HOA expression is defined, and then the maximum potential gain of the signal before entering the gain control processing stage is concluded.

[0035] Normalization of input HOA expressions To use the inventive process, a normalization of the (entire) input HOA representation signal is performed beforehand. For HOA compression, a frame-by-frame process is performed, where the k-th frame C(k) of the original input HOA representation is expressed as

number

[0036] As stated in Patent Document 4, a meaningful normalization of the HOA representation from a practical point of view is to normalize the individual HOA coefficient sequences c n m This is not achieved by imposing constraints on the value range of (t), since these time-domain functions are not the signals that will actually be played by the speakers after rendering. Instead, we can map the HOA representation to O virtual speaker signals w j It is more convenient to think of the equivalent spatial domain representation obtained by rendering (t) into 1 ≤ j ≤ O. Each virtual speaker position is assumed to be represented by a spherical coordinate system, where each position is assumed to lie on a unit sphere and have a radius of 1. These positions are then bounded by order-dependent directions Ω j (N) =(θ j (N) ,φ j (N) ), 1≦j≦0, where θ j (N) and φ j (N)and represent the tilt angle and azimuth angle, respectively (see Fig. 6 and its explanation for the definition of the spherical coordinate system). These directions should be distributed as uniformly as possible on the unit sphere. For example, see Non-Patent Document 2. For the calculation of a specific direction, the number of nodes is fliege / nodes / nodes.html. Their location generally depends on the kind of definition of "spherical uniform distribution" and is therefore not unambiguous.

[0037] The advantage of defining a value range for the virtual speaker signal over defining a value range for the HOA coefficient sequence is that the value range for the former can be intuitively set equal to the interval [-1, 1[, as is the case for normal speaker signals assuming a PCM representation. This leads to a spatially uniformly distributed quantization error, and so the quantization is advantageously applied in a region that is meaningful with respect to actual listening. An important aspect in this context is that the number of bits per sample can be chosen as low as is typically the case for normal speaker signals, e.g., 16, whereas normally a higher number of bits per sample (e.g., 24 or even 32) would be needed. This increases efficiency compared to direct quantization of the HOA coefficient sequence.

[0038] To describe the normalization process in detail in the spatial domain, all virtual speaker signals are expressed as w(t):=[w1(t) … w O (t)] T (2) Here, (·) T represents the transposition. The virtual direction Ω j (N) , the mode matrix for 1≦j≦O is

number

[0039] Using these definitions, reasonable requirements for a virtual speaker signal are:

number

[0040] As a result, the total power of the speaker signal is

number

[0041] Consequences for signal value range before gain control Assuming that the normalization of the input HOA expression is performed as described in the section "Normalization of the Input HOA Expression", the signal y i , i=1,...,l are considered below. These signals are the HOA coefficient sequences or the dominant tone signal x PS,d , d=1,…,D and / or the ambient HOA component c AMB,n , n=1,...,O, by assigning one or more of the specific coefficient sequences (some of which have undergone spatial transformations) to the available I channels. Therefore, it is necessary to analyze the possible value ranges of the different signal types listed here, under the normalization assumption in equation (6). Since all types of signals are intermediately calculated from the original HOA coefficient sequence, we will look at their possible value ranges.

[0042] The case where only one or more HOA coefficient sequences are included in I channels is not depicted in Figures 1A and 2B, i.e., in such cases, HOA decomposition, ambient correction and the corresponding synthesis block are not required.

[0043] Consequences for the value range of HOA expressions The time-continuous HOA representation is derived from the virtual speaker signal. c(t)=Ψw(t) (8) This is the inverse operation of equation (5). Therefore, the total power of all HOA coefficient sequences is bounded using equations (8) and (7) as follows:

[0044]

number

[0045] FIG. 3 shows the virtual direction Ω according to the paper of Non-Patent Document 2 mentioned above. j (N) , 1≦j≦0, and the values ​​of K are shown for HOA orders N=1,...,29.

[0046] Combining all previous discussions and considerations, an upper bound on the absolute value of the HOA coefficient sequence is given as follows:

[0047]

number

[0048] It is important to note that the condition in equation (6) implies the condition in equation (11), but the converse is not true, i.e., equation (11) does not imply equation (6).

[0049] A further important aspect is that, under the assumption of nearly uniformly distributed virtual speaker positions, the column vectors of the modal matrix Ψ, which represent the mode vectors with respect to the virtual speaker positions, are nearly orthogonal to each other and each have a Euclidean norm of N+1. This property means that the spatial transformation nearly preserves the Euclidean norm except for a multiplicative constant. That is,

number

[0050] Consequences for the value range of the dominant signal Both types of dominant sound signals (directional and vector-based) contribute to the HOA representation with Euclidean norm N+1, i.e. ||v1||2=N+1 (13) a single vector v1∈R such that O They have in common that they are described by

[0051] In the case of directional signals, this vector is directed to a certain source direction Ω S,1 corresponds to the mode vector with respect to

number

[0052] In the following, we consider D dominant sound signals x d The general case of (t), d=1,…,D is considered. These signals are x(t)=[x1(t) x2(t) … x D (t)] T (16) These signals can be collected into a vector x(t) according to the monophonic dominant signal x d (t), all vectors v representing directional distributions for d=1,…,D d , d=1,…,D V:=[v1v2… v D ] (17) The decision must be based on the following:

[0053] For meaningful extraction of the dominant sound signal x(t), the following constraints are formulated: a) Each dominant signal is obtained as a linear combination of the coefficient sequences of the original HOA representation, i.e. x(t)=A·c(t) (18) where A∈R D×O represents the mixing matrix. b) The mixing matrix A has a Euclidean norm that does not exceed the value 1, i.e.

number

number

[0054] Substituting equation (18) into equation (20), equation (20) becomes the constraint

number

[0055] From the constraints in equations (18) and (19) and the consistency of the Euclidean matrix and vector norm, the upper bound on the absolute value of the dominant sound signal can be obtained by using equations (18), (19) and (11):

number

number

[0056] Example for choosing a confusion matrix An example of how to determine the mixing matrix that satisfies the constraint (20) is the one that minimizes the Euclidean norm of the residuals after extraction, i.e.

number

[0057] Nevertheless, the matrix V still satisfies the constraint (19), i.e.

number

[0058] In the case of directional signals only, the matrix V is a function of several source signal directions Ω S,d , the modal matrix for d=1,…,D, i.e.

number

[0059] Consequences for the range of coefficient sequences of the surrounding HOA components The ambient HOA component is calculated by subtracting the HOA representation of the dominant sound signal from the original HOA representation, i.e.,

number

number

[0060] <Value range of the spatially transformed coefficient sequence of the ambient HOA components> A further aspect of the HOA compression process proposed in the MPEG documents of Patent Document 2 and the above-mentioned Non-Patent Document 1 is that the first O MIN coefficient sequences are always chosen to be assigned to a transport channel, where O MIN =(N MIN +1) 2 and N MIN≦N is typically a smaller order than the order of the original HOA representation. To decorrelate these HOA coefficient sequences, they are normalized to some predefined direction Ω (similar to the concept described in the section Normalizing the Input HOA Representation). MIN,d , d=1,…,O MIN can be converted into an incident virtual speaker signal from the order index n≦N MIN Let c be the vector of all coefficient sequences of the surrounding HOA components. AMB,MIN (t) and the virtual direction Ω MIN,d , d=1,…,O MIN The mode matrix for MIN When defined by w MIN The vector of all virtual speaker signals (defined by) (t) is

number

[0061] Therefore, using the consistency of Euclidean matrices and vector norms,

number

[0062] In the MPEG document of the above-mentioned non-patent document 1, the virtual direction Ω MIN,d , d=1,…,O MIN is chosen according to the paper in Non-Patent Document 2 mentioned above. The modal matrix Ψ MIN The Euclidean norm of each inverse matrix of MIN 4 for =1,...,9.

[0063]

number

[0064] However, N MIN >9, this is not true in general. In this case, ||Ψ MIN-1 The value of ||2 is typically much larger than 1. Nevertheless, at least 1 ≤ N MIN For ≦9, the amplitude of the virtual speaker signal is limited by

[0065]

number

[0066] Maximum degree of interest N MAX Any degree N up to 1≦N≦N MAX For the amplitude of the signal before gain control, the value (√K MAX )·O. Here,

number

[0067] K MAX is the maximum order of interest, N MAX and the virtual speaker direction Ω j (N) , 1≦j≦O, and can be expressed as follows:

[0068]

number

number

[0069] If the amplitude of the signals before gain control is too small, the MPEG document in Non-Patent Document 1 states that these amplitudes are

number

[0070] Thus, each exponent to base 2 describing the absolute amplitude change of the total modified signal caused by the gain control processing unit from the beginning to the current frame in access units is in the interval [e MIN ,e MAX ]. As a result, the (lowest integer) number of bits required to encode it, β e is given by the following equation:

[0071]

number

[0072]

number

[0073] The number of bits for the exponent, β e Using ,guarantees that all possible absolute amplitude changes caused by the HOA compressor gain control processing units 15,...,151 can be captured, and allows decompression to begin at some predefined entry point within the compressed representation.

[0074] At the HOA decompressor, when starting to decompress the compressed HOA representation, it represents the total absolute amplitude change allocated to the side information for several data frames and is used to decompress the received data stream.

number

[0075] Further Embodiments When implementing a specific HOA compression / decompression system as described in sections HOA Compression, Spatial HOA Encoding, HOA Decompression, and Spatial HOA Decoding, the amount of bits β for encoding the exponent is e is the scaling factor K MAX,DES This K MAX,DES The desired maximum degree N of the HOA representation to be compressed MAX,DES and some kind of virtual speaker direction

number

[0076] For example, N MAX,DES = 29, and when choosing the virtual speaker direction according to the paper in Non-Patent Document 2, the rational choice is √K MAX,DES = 1.5. In this situation, the same virtual speaker direction Ω DES,1 (N) ,…,Ω DES,O (N) are normalized according to the section on Normalizing Input HOA Expressions, where 1≦N≦N MAX Correct compression is guaranteed for an HOA representation of order N such that: However, this guarantee is not guaranteed if the virtual speaker signal is equivalently represented (for efficiency reasons) also in PCM format, but with a virtual speaker direction Ω j (N) , 1≦j≦O is the virtual speaker direction Ω assumed in the system design stage. DES,1 (N) ,…,Ω DES,O (N) In the case of HOA expressions chosen differently from

[0077] Due to this different choice of virtual speaker positions, even if these virtual speaker signals are in the interval [1,1[, the amplitude of the signal before gain control will be different from the value (√K MAX,DES )·O. Therefore, it cannot be guaranteed that this HOA representation has the correct normalization for compression according to the process described in the MPEG document [1].

[0078] In this situation, it is advantageous to have a system that provides the maximum allowable amplitude of the virtual speaker signals based on knowledge of the virtual speaker positions, in order to ensure that each HOA representation is suitable for compression according to the process described in the MPEG document [1]. Such a system is shown in Figure 5. This requires O=(N+1) 2 , where N∈N0, and the input is the virtual speaker position Ω j (N), 1≦j≦O, and output the maximum allowed amplitude (measured in decibels) of the virtual speaker signal γ dB In step or stage 51, the modal matrix Ψ for the virtual speaker positions is calculated according to equation (3). In the following step or stage 52, the Euclidean norm ||Ψ||2 of the modal matrix is ​​calculated. In the third step or stage 53, the amplitude γ is calculated as 1 and the square root of the number of virtual speaker positions and K MAX,DES and the Euclidean norm of the modal matrix.

number

[0079] To illustrate: from the above derivation, the magnitude of the HOA coefficient sequence is given by the value (√K MAX,DES )·O, i.e.

number

[0080] From equation (9), the magnitude of the HOA coefficient sequence is

number

number

number

[0081] That is, the maximum magnitude value 1 in equation (6) is replaced by the maximum magnitude value γ in equation (47).

[0082] <Higher-Order Ambisonics Basics> Higher Order Ambisonics (HOA) is based on the description of the sound field in a compact region of interest where no sound sources are assumed. In that case, the spatiotemporal behavior p(t,x) of the sound pressure at position x and time t in the region of interest is completely determined physically by the homogeneous wave equation. In the following, we consider the spherical coordinate system shown in Figure 6. In this coordinate system used, the x-axis points to the forward position, the y-axis points to the left, and the z-axis points up. At a position x=(r,θ,φ) in space T is expressed by the radius r>0 (i.e., the distance to the coordinate origin), the tilt angle θ∈[0,π] measured from the polar axis z, and the azimuthal angle φ∈[0,2π[ measured counterclockwise from the x-axis in the xy plane. Furthermore, (·) T represents the transposition.

[0083] Then, assuming that ω represents the angular frequency and i represents the imaginary unit, from the textbook in Non-Patent Document 3, F t The Fourier transform of the sound pressure with respect to time is represented by (·), i.e.

number

number

[0084] If a sound field is represented by a superposition of infinite harmonic plane waves of different angular frequencies ω arriving from all possible directions specified by the angle tuple (θ,φ), it can be shown that each plane wave complex amplitude function C(ω,θ,φ) can be expressed by the following spherical harmonic expansion (Non-Patent Document 4):

[0085]

number

number

number

[0086] HOA coefficient sequence c in vector c(t) n m The position index of (t) is n(n+1)+1+m The total number of elements in the vector c(t) is given by O=(N+1) 2 is given by The final Ambisonics format is a sampled version of c(t) using the sampling frequency fs,

number

[0087] Definition of real-valued spherical harmonics real-valued spherical harmonic functions S n m (θ,φ) (assuming SN3D normalization according to Non-Patent Document 5, Chapter 3.1) is given by

[0088]

number

[0089]

number

[0090] The present invention can be carried out by a single processor or electronic circuit, or by several processors or electronic circuits operating in parallel and / or operating on different parts of the processing of the invention.

[0091] Instructions for operating such processor(s) may be stored in one or more memories.

[0092] Several aspects will be described. [Aspect 1] For the compression of the HOA data frame representation (C(k)), the non-differential gain values ​​(2 e ) the minimum number of integer bits required to represent e wherein each channel signal in each frame comprises a group of sample values, and each channel signal (y1(k-2),...,y I a differential gain value is assigned to the channel signal in the current HOA data frame ((k-2)), such differential gain value causing a change in amplitude of the sample values ​​of the channel signal in the previous HOA data frame ((k-3)), and such gain-adapted channel signal is encoded in an encoder (16); The HOA data frame representation (C(k)) is expressed as O virtual speaker signals w in the spatial domain. j (t), the positions of the O virtual speakers are on the unit sphere, and β e does not match the position assumed for the calculation of the matrix multiplication w(t)=(Ψ) -1where w(t) is a vector containing all virtual speaker signals, Ψ is the modal matrix (51) calculated for these virtual speaker positions, and c(t) is a vector of the corresponding HOA coefficient sequence in the HOA data frame representation (C(k)), Maximum allowed amplitude value

number

number

number

number

number

number

number

number

Claims

1. 1. A method of decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, comprising: demultiplexing the compressed HOA representation from a bitstream, where some HOA coefficients correspond to the compressed HOA representation; When there are independent access units in the bitstream, the lowest integer β e and decoding the compressed HOA representation based on the lowest integer β e teeth [Equation 1] is determined based on [Equation 2] where N is the order and N MAX is the maximum order of interest, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between and √K MAX = 1.5, method.

2. 1. An apparatus for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, comprising: a demultiplexer configured to demultiplex the compressed HOA representation from a bitstream, wherein a number of HOA coefficients correspond to the compressed HOA representation; and When there are independent access units in the bitstream, the lowest integer β e and a processor for decoding the compressed HOA representation based on the lowest integer β e teeth [Equation 3] is determined based on [Equation 4] where N is the order and N MAX is the maximum order of interest, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between and √K MAX = 1.5, Device.

3. A non-transitory computer-readable medium storing executable instructions that cause a computer to perform the steps of the method of claim 1.

4. 1. A method of decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, comprising: receiving a bitstream including the compressed HOA representation, the bitstream including a number of HOA coefficients corresponding to the compressed HOA representation; When there are independent access units in the bitstream, the lowest integer β e and decoding the compressed HOA representation based on the lowest integer β e teeth [Equation 7] is determined based on [Equation 8] where N is the order and N MAX is the maximum order of interest, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between O and e MAX >0, when the amplitude of the sample value of the channel signal before the gain control is lower than the threshold, the lowest integer bit number β e It acts to increase [Equation 9] is the maximum gain applied to the channel signal by the gain control; method.

5. To compress the HOA data frame representation, the minimum integer number of bits β is used to express the representation of the non-differential gain value corresponding to the amplitude change of the channel signal of the HOA data frame as an exponent of two. e wherein each channel signal in each frame comprises a group of sample values, a differential gain value is assigned to each channel signal in each frame of the HOA data frame, the differential gain value causing a change in amplitude of a first sample value of the channel signal in a current HOA data frame relative to a second sample value of the channel signal in a immediately preceding HOA data frame, and the resulting gain-adapted channel signal is encoded in an encoder (16); The HOA data frame representation is expressed in the spatial domain as O virtual speaker signals w j (t), the positions of the virtual speakers lie on a unit sphere and are targeted to be uniformly distributed on the unit sphere, and the rendering is performed by matrix multiplication w(t) = (Ψ) -1 represented by c(t), where w(t) is a vector containing all virtual speaker signals, Ψ is the virtual speaker position mode matrix, and c(t) is a vector of the corresponding HOA coefficient sequence in said HOA data frame representation; The HOA data frame representation is [Equation 10] and the method is normalized so that: - Channel signal, a) multiplying a vector c(t) of HOA coefficient sequences by a mixing matrix A to represent a dominant sound signal in the channel signal, the mixing matrix A representing a linear combination of coefficient sequences of normalized HOA data frame representations; b) ambient components c in the channel signal AMB To represent (t), subtract the dominant sound signal from the normalized HOA data frame representation and subtract the resulting minimum ambient component c AMB,MIN (t) to w MIN (t)=Ψ MIN -1 ・c AMB,MIN (t) by calculating w MIN (t) is the vector of all virtual speaker signals, and ||Ψ MIN -1 || 2 < 1, and Ψ MIN is the minimum surrounding component c AMB,MIN (t) is the modal matrix for the substep; c) selecting a portion of the HOA coefficient sequence vector c(t) that relates to the coefficient sequence of the smallest surrounding component to which a spatial transformation is applied; forming the ion beam by performing When an independent access unit exists in the bitstream, the minimum integer bit number β e of [0011] and determining based on [0012] where N is the order and N MAX is the maximum order of interest, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the modal matrix ||Ψ|| 2 2 is the ratio between O and e MAX >0, when the amplitude of the sample value of the channel signal before the gain control is lower than the threshold, the lowest integer bit number β e It acts to increase [0013] is the maximum gain applied to the channel signal by the gain control; method.

6. 1. A method of decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or sound field, comprising: receiving a bitstream including the compressed HOA representation, the bitstream including a number of HOA coefficients corresponding to the compressed HOA representation; When there are independent access units in the bitstream, the lowest integer β e and decoding the compressed HOA representation based on the lowest integer β e teeth [0014] and [Equation 15] where N is the order and N MAX is the maximum order of interest, and Ω 1 (N) ,…,Ω O (N) is the direction of the virtual speaker, and O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the squared Euclidean norm of the virtual speaker position mode matrix ||Ψ|| 2 2 is the ratio between O and e MAX >0, when the amplitude of the sample value of the channel signal before the gain control is lower than the threshold, the lowest integer bit number β e It acts to increase [0016] is the maximum gain applied to the channel signal by the gain control; method.

Citation Information

Patent Citations

  • Method and apparatus for compressing and decompressing a Higher Order Ambisonics signal representation

    EP2665208A1

  • Method and apparatus for compressing and decompressing a higher order ambisonics representation for a sound field

    EP2743922A1

  • Method and Apparatus for compressing and decompressing a Higher Order Ambisonics representation

    EP2800401A1

  • Method and Apparatus for generating from a coefficient domain representation of HOA signals a mixed spatial / coefficient domain representation of said HOA signals

    EP2824661A1

  • Method and apparatus for encoding and decoding successive frames of ambisonics representation of two-dimensional or three-dimensional sound field

    JP2012133366A