Method for determining minimum integer number of bits required to represent non-differential gain values for compression of HOA data frame representations

By establishing the relationship between the value range represented by HOA and the maximum gain of the signal, the minimum number of integer bits required for the non-differential gain value of the channel signal in the HOA data frame representation is determined, and the problem of low compression efficiency of HOA representation is solved, and efficient compression of HOA representation is achieved.

CN120032652APending Publication Date: 2025-05-23DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510186602.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2014-06-27
Filing Date
2015-06-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively determine the minimum number of integer bits required for the non-differential gain value of the channel signal in the HOA data frame representation, resulting in inefficient compression of the HOA representation.

Method used

By establishing the relationship between the value range represented by the HOA and the maximum gain before the signal is applied in the HOA compressor, the required number of bits is determined to describe the total absolute amplitude change of the signal caused by the gain control processing unit.

Benefits of technology

The minimum number of integer bits required to determine the non-differential gain value in HOA compression is realized, which improves the compression efficiency of the HOA representation and ensures the correct compression of the HOA representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032652A_ABST
    Figure CN120032652A_ABST
Patent Text Reader

Abstract

A method of determining a minimum integer number of bits required to represent a non-differential gain value for compression of an HOA data frame representation is disclosed. When compressing HOA data frame representations, gain control (15, 151) is performed on each channel signal before it is perceptually encoded (16). The gain value is transmitted as side information in a differential manner. However, in order to start decoding such a streaming compressed HOA data frame representation, an absolute gain value is required, which should be encoded with a minimum number of bits. In order to determine such a minimum integer number of bits {[beta] e), the HOA data frame representation (C (k)) is rendered in the spatial domain into a virtual loudspeaker signal located on a unit sphere, then the HOA data frame representation (C (k)) is normalized, and then the minimum integer number of bits (AA) is set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application filed in respect of a divisional application with application number 202111089841.0 (the application number of the parent application is 201580035127.X). Technical Field

[0002] The present invention relates to a method for determining the minimum number of integer bits required to represent a non-differential gain value associated with a channel signal of a specific data frame in an HOA data frame for compression of the HOA data frame representation. Background Art

[0003] Higher-order high-fidelity stereophonic sound reproduction represented as HOA provides a possibility to represent three-dimensional sound. Other techniques are wave field synthesis (WFS) or channel-based methods such as 22.2. Compared with channel-based methods, HOA representation has the advantage of being independent of a specific speaker setup. However, this flexibility comes at the cost of a decoding process required to play back the HOA representation on a specific speaker setup. Compared with the WFS method where the number of required speakers is usually large, HOA can also be presented in a setup including only a few speakers. Another advantage of HOA is that the same representation can also be adopted without any modification to the binaural rendering of headphones.

[0004] HOA is based on representing the spatial density of the complex harmonic plane wave amplitude by truncated spherical harmonic function (SH) expansion. Each expansion coefficient is a function of the angular frequency, which can be equivalently represented by a time-domain function. Therefore, without loss of generality, the complete HOA sound field representation can actually be assumed to be composed of O time-domain functions, where O represents the number of expansion coefficients. These time-domain functions will be equivalently referred to as HOA coefficient sequences or HOA channels hereinafter.

[0005] The spatial resolution of the HOA representation increases with the growth of the maximum order N of the expansion. Unfortunately, the number of expansion coefficients O grows quadratically with the order N, specifically, O=(N + 1) 2 . For example, using a typical HOA representation with order N = 4 requires O = 25 HOA (expansion) coefficients. Assuming the desired mono sampling rate is f S and the number of bits per sample is N b , then the total bit rate for transmitting the HOA representation is determined by O·f S ·N b . Taking f b = 16 bits per sample S=48kHz sampling rate to transmit the HOA representation of order N=4, resulting in a bit rate of 19.2MBits / s, which is very high for many practical applications (such as streaming). Therefore, it is very desirable to compress the HOA representation.

[0006] Previously, compression of HOA sound field representations was proposed in EP 2665208 A1, EP 2743922 A1, EP 2800401 A1, see ISO / IEC JTC1 / SC29 / WG11, N14264, WD1-HOA text for MPEG-H 3D Audio, January 2014. These methods have in common that they all perform a sound field analysis and decompose a given HOA representation into a directional component and a residual ambient component. On the one hand, the final compressed representation is assumed to consist of several quantized signals, which are generated by the perceptual encoding of the directional signal and the vector-based signal and the correlation coefficient sequence of the ambient HOA components. On the other hand, the final compressed representation includes additional side information related to the quantized signal, which is required to reconstruct the HOA representation from its compressed version.

[0007] Before being passed to the perceptual encoder, these intermediate time domain signals are required to have a maximum amplitude in the value range of [-1,1], which is a requirement arising from the implementation of currently available perceptual encoders. In order to meet this requirement when compressing the HOA representation, a gain control processing unit that smoothly attenuates or amplifies the input signal is used before the perceptual encoder (see EP 2824661 A1 and the ISO / IEC JTC1 / SC29 / WG11 N14264 document mentioned above). The resulting signal modification is assumed to be reversible and is applied frame by frame, wherein in particular the change in the signal amplitude between consecutive frames is assumed to be a power of "2". In order to facilitate the inversion of this signal modification in the HOA decompressor, the corresponding normalized side information is included in the total side information. This normalized side information can be composed of exponentials with a base of "2" that describe the relative amplitude change between two consecutive frames. Since smaller amplitude changes between consecutive frames are more likely to occur than larger amplitude changes, these indices are encoded using run-length coding in accordance with the ISO / IEC JTC1 / SC29 / WG11 N14264 document mentioned above. Summary of the invention

[0008] For example, in the case of decompressing a single file without any time jumps from the beginning to the end, it is feasible to use differentially encoded amplitude changes in HOA decompression to reconstruct the original signal amplitude. However, in order to facilitate random access, independent access units must exist in the encoded representation (which is usually a bitstream) to enable decompression to start from the desired position (or at least near it) independently of the information from the previous frame. Such independent access units must contain the total absolute amplitude change (i.e., non-differential gain value) caused by the gain control processing unit from the first frame to the current frame. Assuming that the amplitude change between two consecutive frames is a power of "2", it is sufficient to describe the total absolute amplitude change by an exponential with a base of "2". In order to efficiently encode this exponent, it is necessary to know the possible maximum gain of the signal before applying the gain control processing unit. However, this knowledge is highly dependent on the specification of constraints on the value range of the HOA representation to be compressed. Unfortunately, the MPEG-H 3D Audio document ISO / IEC JTC1 / SC29 / WG11 N14264 only provides a description of the format for inputting HOA representations without setting any constraints on the value range.

[0009] The problem to be solved by the present invention is to provide the minimum number of integer bits required to represent non-differential gain values.

[0010] The present invention establishes a correlation between the range of values ​​of the input HOA representation and the maximum possible gain of the signal before the gain control processing unit is applied in the HOA compressor.

[0011] Based on this relationship, the number of required bits is determined for a given specification of the range of values ​​represented by the input HOA and for the efficient encoding of an exponent with base "2" to describe within the access unit the total absolute amplitude change (i.e., the non-differential gain value) of the modified signal caused by the gain control processing unit from the first frame to the current frame.

[0012] Furthermore, once the rules for calculating the required amount of bits for encoding the exponent are determined, the present invention uses a process for verifying whether a given HOA representation satisfies the required value range constraints so that the given HOA representation can be correctly compressed.

[0013] In principle, the method of the present invention is suitable for determining the minimum number of integer bits β required for representing the non-differential gain value of the channel signal of a specific HOA data frame in the HOA data frame for compression of the HOA data frame representation. e, wherein each channel signal in each frame comprises a set of sample values, and wherein a differential gain value is assigned to each channel signal of each of said HOA data frames, and such differential gain value causes a change in the amplitude of the sample values ​​of the channel signal in a current HOA data frame relative to the sample values ​​of the channel signal in a previous HOA data frame, and wherein such gain-adjusted channel signals are encoded in an encoder,

[0014] And wherein the HOA data frame representation is rendered as O virtual speaker signals w in the spatial domain j (t), where the positions of the O virtual speakers are located on the unit sphere and are aligned with respect to β e The calculation of the assumed position does not match the rendering by matrix multiplication w(t) = (Ψ) -1 c(t), where w(t) is a vector containing all virtual loudspeaker signals, Ψ is a modulus matrix calculated for the virtual loudspeaker positions, and c(t) is a vector of corresponding HOA coefficient sequences represented by the HOA data frame,

[0015] And among them, calculate the maximum allowed amplitude value And the HOA data frame representation is normalized so that

[0016] The method comprises the following steps:

[0017] - forming the channel signal from the normalized HOA data frame representation by one or more of the following sub-steps a), b), c):

[0018] a) in order to represent the main sound signal in the channel signal, multiplying the vector of the HOA coefficient sequence c(t) by a mixing matrix A, the Euclidean norm of the mixing matrix A is not greater than "1", wherein the mixing matrix A represents a linear combination of the coefficient sequence represented by the normalized HOA data frame;

[0019] b) To represent the ambient component c in the channel signal AMB (t), subtracting the primary sound signal from the normalized HOA data frame representation and selecting the ambient component c AMB (t), where ||c AMB (t)|| 2 2 ≤||c(t)|| 2 2 , and by calculating For the minimum environmental component c AMB,MIN (t) is transformed, where AndMIN is the minimum environmental component c AMB,MIN (t) modulus matrix;

[0020] c) selecting a portion of the HOA coefficient sequence c(t), wherein the selected coefficient sequence is related to the coefficient sequence of the ambient HOA component on which the spatial transform is performed, and a minimum order N describing the number of selected coefficient sequences MIN For B MIN ≤9;

[0021] - the minimum number of integer bits β required to represent the non-differential gain value of the channel signal e Set to

[0022] in, N is the order, O = (N + 1) 2 is the number of HOA coefficient sequences, K is the ratio between the square of the Euclidean norm of the modulus matrix and O, and where N MAX,DES is the order of interest, and is the direction of the virtual loudspeaker for each order, wherein the direction is assumed to achieve the compression of the HOA data frame representation, so that To select β e , thereby encoding the exponent of the non-differential gain value with a base of "2",

[0023] And among them, for the calculation ||Ψ|| 2 is the Euclidean norm of the modular matrix Ψ, N is the order, N MAX is the maximum order of interest, is the direction of the virtual speaker, O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the square of the Euclidean norm of the module matrix ||Ψ|| 2 2 The ratio between and O. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Exemplary embodiments of the present invention are described with reference to the accompanying drawings, in which:

[0025] Figure 1 HOA compressor;

[0026] Figure 2 HOA decompressor;

[0027] Figure 3 Virtual direction Ω j (N)(1≤j≤O) scaling value K with respect to HOA order (N=1, ..., 29);

[0028] Figure 4 For HOA order (N MIN =1, ..., 9), the inverse matrix Ψ -1 About the virtual direction Ω MIN,d (d=1,...,O MIN )’s Euclidean norm;

[0029] Figure 5 The virtual speaker is at position Ω j (N) (1≤j≤O, where O=(N+1) 2 ) is the maximum allowable amplitude of the signal at dB determination;

[0030] Figure 6 Spherical coordinate system. DETAILED DESCRIPTION

[0031] The following embodiments may be used in any combination or sub-combination even if not explicitly described.

[0032] In the following, the principles of HOA compression and decompression are introduced to provide a more detailed background to the above-mentioned problems. The basis for this introduction is the processing described in the MPEG-H 3D Audio document ISO / IEC JTC1 / SC29 / WG11 N14264 (see also EP 2665208A1, EP 2800401 A1 and EP 2743922 A1). In N14264, the "directional component" is extended to the "primary sound component". As a directional component, the primary sound component is assumed to be represented in part by a directional signal together with some prediction parameters for predicting multiple parts of the original HOA representation based on the directional signal. The directional signal refers to a monophonic signal with a corresponding direction from which the listener is assumed to be impacted. In addition, the primary sound component is assumed to be represented by a "vector-based signal", which refers to a monophonic signal with a corresponding vector that defines the directional distribution of the vector-based signal.

[0033] HOA Compression

[0034] Figure 1 The overall architecture of the HOA compressor described in EP 2800401 A1 is shown. The overall architecture of the HOA compressor has Figure 1 The spatial HOA encoding unit shown in A and Figure 1B. The spatial HOA encoder provides a first compressed HOA representation consisting of the I signal together with side information describing how its HOA representation was created. The I signal is perceptually encoded in the perceptual encoder and the side information source encoder, and the side information is source encoded, before multiplexing the two encoded representations.

[0035] Spatial HOA Coding

[0036] In the first step, the current k-th frame C(k) of the original HOA representation is input to the direction and vector estimation processing step or stage 11, which is assumed to provide the tuple set and Tuple Set It consists of a tuple whose first element represents the index of the direction signal and whose second element represents the corresponding quantization direction. consists of a tuple whose first element represents the index of the vector-based signal and the second element represents the vector defining the directional distribution of the signal (ie, how to calculate the HOA representation of the vector-based signal).

[0037] Using two tuple sets and In the HOA decomposition step or stage 12, the initial HOA frame C(k) is decomposed into a frame X of all dominant sound (ie, directional and vector-based) signals. PS (k-1) and frame C of the ambient HOA component AMB (k-1). Note the one frame delay caused by the overlap-add process to avoid blocking artifacts. Furthermore, the HOA decomposition step / stage 12 is assumed to output some prediction parameters ζ(k-1) describing how to predict parts of the original HOA representation from the directional signals to enrich the dominant sound HOA components. In addition, it is assumed that a target allocation vector v containing information about the allocation of the dominant sound signal determined in the HOA decomposition processing step or stage 12 to the I available channels is provided. A,T (k-1). The affected channels may be assumed to be occupied, which means that the affected channels cannot be used for transmitting any coefficient sequence of the ambient HOA component in the corresponding time frame.

[0038] In the environmental component modification process step or stage 13, the target allocation vector v A,T (k-1) provides information to modify the frame C of the ambient HOA component AMB (k-1). In particular, (among other aspects) according to which channels are available and not yet occupied by the main sound signal (contained in the target allocation vector v A,TThe information in (k-1) determines which coefficient sequences of the ambient HOA components to transmit in the given I channels.

[0039] Additionally, if the index of the selected coefficient sequence changes between consecutive frames, a cross-fading of the coefficient sequence is performed.

[0040] In addition, assuming that the ambient HOA component C AMB The first O of (k-2) MIN The coefficient sequence is always chosen to be perceptually coded and transmitted, where O MIN =(N MIN +1) 2 (N MIN ≤N) is usually of a smaller order than the original HOA representation. In order to decorrelate these HOA coefficient sequences, they can be transformed in step / stage 13 to be from some predefined direction Ω MIN,d (d=1,...,O MIN ) is the directional signal of the impact (i.e., a general plane wave function).

[0041] Temporarily predicted modified ambient HOA component C P,M,A (k-1) together with the modified ambient HOA component C M,A (k-1) are calculated together in step / stage 13 and used in the gain control processing steps / stages 15, ..., 151 to achieve reasonable foresight, where the modified information about the ambient HOA components is directly related to the allocation of all possible types of signals to the available channels in the channel allocation step or stage 14. The final information about this allocation is assumed to be contained in the final allocation vector v A (k-2). To calculate this vector in step / stage 13, use the vector contained in the target allocation vector v A,T The information in (k-1).

[0042] The channel allocation in step / stage 14 utilizes the allocation vector v A The information provided by (k-2) will be included in frame X PS (k-2) and contained in frame C M,A The appropriate signals in (k-2) are assigned to I available channels, thus obtaining signal frame y i (k-2), i=1, ..., I. In addition, it will also be included in the frame X PS (k-1) and frame C P,AMB The appropriate signals in (k-1) are assigned to I available channels, thereby obtaining the predicted signal frame y P,i (k-1), i=1, ..., I.

[0043] Signal frame y iEach of (k-2), i=1, ..., I is finally processed by a gain control processing step / stage 15, ..., 151 to obtain an index e i (k-2) and abnormal marker β i (k-2), i=1, ..., I and signal z i (k-2), i=1, ..., I, where the signal gain is smoothly modified to achieve a value range suitable for the perceptual encoder step or stage 16. Step / stage 16 outputs the corresponding encoded signal frame Predicted signal frame y P,i (k-1), i=1, ..., I achieves reasonable foresight to avoid large gain variations between consecutive blocks. In the side information source encoder step or stage 17, the side information data e i (k-2), β i (k-2), ζ(k-1), and v A (k-2) Perform source coding to obtain the coded side information frame In the multiplexer 18, the coded signal of frame (k-2) is The encoded side information data of the frame Combine to get the output frame

[0044] In the spatial HOA decoder, the gain modification in the gain control processing steps / stages 15, ..., 151 is assumed to be performed by using an index e i (k-2) and abnormal marker β i (k-2), i=1, ..., I constitute the gain control side information to restore.

[0045] HOA Unpacked

[0046] Figure 2 The overall architecture of the HOA decompressor described in EP 2800401 A1 is shown. The overall architecture consists of paired components of the HOA compressor components, which are arranged in reverse order and include Figure 2 The perceptual decoding unit and the source decoding unit shown in A and Figure 2 B shows the spatial HOA decoding unit.

[0047] In the perceptual decoding section and the source decoding section (representing the perceptual decoder and the side information source decoder), the demultiplexing step or stage 21 receives the input frame from the bitstream And provide a perceptual coded representation of I signals and the encoded side information data describing how to create its HOA representation In the perceptual decoder step or stage 22 The signal is perceptually decoded to obtain a decoded signal In the side information source decoder step or stage 23, the encoded side information data Decode to get the data set Index e i (k), abnormal marker β i (k), prediction parameter ζ(k+1) and allocation vector v AMB,ASSIGN (k). About v A With v AMB,ASSIGN For the differences between them, see the MPEG document N14264 mentioned above.

[0048] Spatial HOA decoding

[0049] In the spatial HOA decoding unit, the decoded signal is sensed Each of the above together with its associated gain correction index e i (k) and gain correction abnormality flag β i (k) are input to the inverse gain control processing step or stage 24, 241. The i-th inverse gain control processing step / stage provides a gain-corrected signal frame

[0050] Total I gain-corrected signal frames Together with the allocation vector v AMB,ASSIGN (k) and tuple sets and are fed together to the channel re-splitting

[0051] Matching step or stage 25, see Tuple Set and Assign vector v AMB,ASSIGN (k) is composed of I components, which indicate for each transmission channel whether it contains a coefficient sequence of the ambient HOA component and which coefficient sequence it contains. In the channel reallocation step / stage 25, the gain-corrected signal frame The frames are redistributed to reconstruct all primary sound signals (i.e., all directional signals and vector-based signals) and the intermediate representation frame C of the ambient HOA component I,AMB (k). In addition, a set of indices of coefficient sequences of the ambient HOA components active in the kth frame is provided and a dataset of coefficient indices of ambient HOA components that must be enabled, disabled, and kept active in the (k-1)th frame and

[0052] In the main sound synthesis step or stage 26, the tuple set is used The set of prediction parameters ζ(k+1), the set of tuples And the dataset and According to the frames of all main sound signals To calculate the main sound components The HOA represents.

[0053] In the ambient synthesis step or stage 27, the set of indices of the coefficient sequences of the ambient HOA components active in the kth frame is used According to the intermediate representation of the ambient HOA component frame C I,AMB (k) to create the ambient HOA component frame A one frame delay is introduced due to synchronization with the main sound HOA component.

[0054] Finally, in the HOA composition step or stage 28, the ambient HOA component frame Frame with main sound HOA component Overlay to provide decoded HOA frames

[0055] Thereafter, the spatial HOA decoder creates a reconstructed HOA representation based on the I signals and the side information.

[0056] In the case at the encoding side, the ambient HOA component is transformed into a directional signal, the inverse of which is performed at the decoder side in step / stage 27 .

[0057] The possible maximum gain of the signal before the gain control processing step / stage 15, ..., 151 in the HOA compressor strongly depends on the value range of the input HOA representation. Therefore, a meaningful value range of the input HOA representation is first limited and then a conclusion is drawn about the possible maximum gain of the signal before entering the gain control processing step / stage.

[0058] Normalization of input HOA representation

[0059] In order to use the processing of the present invention, a normalization of the (total) input HOA representation signal is first performed. For HOA compression, a frame-by-frame process is performed, where the k-th frame C(k) of the original input HOA representation is defined as, with respect to the vector c(t) of the time-continuous HOA coefficient sequence specified in formula (54) in the chapter Basics of High-Order Ambisonics,

[0060]

[0061] Where k is the frame index, L is the frame length (in samples), and O = (N + 1)2 is the number of HOA coefficient sequences, and T S Indicates the sampling period.

[0062] As mentioned in EP 2824661 A1, from a practical point of view, a meaningful normalization of the HOA representation is not achieved by normalizing the individual HOA coefficient sequences. This is achieved by imposing constraints on the value range of , because these time domain functions are not the signals actually played by the loudspeakers after rendering. Instead, it is more convenient to consider the HOA representation rendered as O virtual loudspeaker signals w j (t), 1≤j≤O. It is assumed that the corresponding virtual speaker positions are represented by means of a spherical coordinate system, where each position is assumed to be located on a unit sphere with a radius of "1". Therefore, the order-dependent direction Ω j (N) =(θ j (N) ,φ j (N) ), 1≤j≤O equivalently expresses the position, where θ j (N) and φ j (N) denote the inclination and azimuth, respectively (see also Figure 6 and its description of the definition of the spherical coordinate system). See, for example, the technical report "A two-stage approach for computing cubature formulae for the sphere" by J. Fliege and U. Maier in the course of mathematics at the University of Dortmund in 1999. These directions should be distributed as evenly as possible on the unit sphere. The number of nodes used for the calculation in a specific direction can be found at the following website: http: / / www.mathematik.uni-dortmund.de / lsx / research / projects / fliege / nodes / nodes.html. These positions usually depend on the kind of definition of "uniform distribution on the sphere" and are therefore ambiguous.

[0063] The advantage of limiting the value range of the virtual loudspeaker signal by limiting the value range of the HOA coefficient sequence is that the value range of the virtual loudspeaker signal can be intuitively set equal to the interval [-1,1], as is the case with conventional loudspeaker signals assuming a PCM representation. This results in spatially uniformly distributed quantization errors, allowing quantization to be advantageously applied in a domain relevant for actual listening. An important aspect in this context is that the number of bits per sample can be chosen to be as low as the number of bits typically used for conventional loudspeaker signals (i.e., 16), which improves efficiency compared to direct quantization of the HOA coefficient sequence, which typically requires a higher number of bits per sample (e.g., 24 or even 32).

[0064] To describe the normalization process in the spatial domain in detail, all virtual loudspeaker signals are summarized as vectors w(t): = [w 1 (t)...w O (t)] T , (2)

[0065] in,(·) T Denotes transposition. Ψ denotes the virtual direction Ω j (N) , a modular matrix with 1≤j≤O, Ψ is defined as in,

[0066]

[0067] , the rendering process can be formulated as a matrix product

[0068] w(t)=(Ψ) -1 ·c(t). (5)

[0069] Using these definitions, reasonable requirements for the virtual loudspeaker signal are:

[0070]

[0071] This means that the amplitude of each virtual loudspeaker signal needs to fall within the range [-1, 1]. The time instant t is determined by the sampling index l of the sample value of the HOA data frame and the sampling period T S To express.

[0072] The total power of the loudspeaker signal therefore satisfies the condition

[0073]

[0074] HOA data frame representation rendering and normalization in Figure 1 The upstream execution of A’s input C(k).

[0075] Resulting signal value range before gain control

[0076] Assume that the normalization represented by the input HOA is performed according to the description in the normalization section represented by the input HOA. Consider next the signal y input to the gain control processing unit in the HOA compressor i , for values of i = 1, ..., I. These signals are created by allocating the available I channels to one or more of a specific coefficient sequence of the HOA coefficient sequence or the main sound signal x PS,d , for d = 1, ..., D and / or the environmental HOA component c AMB,n , for n = 1, ..., O, and a spatial transform is applied to a part of these signals. Thus, under the normalization assumption in Equation (6), it is necessary to analyze the possible value ranges of these different signal types mentioned. Since all types of signals are calculated intermmediately based on the original HOA coefficient sequence, their possible value ranges are examined

[0077] Figure 1 A and Figure 2 B do not depict the case where only one or more HOA coefficient sequences are included in the I channels, i.e., in this case, no HOA decomposition, environmental component modification block, and corresponding synthesis block are required

[0078] HOA representation value range results

[0079] The time - continuous HOA representation is obtained from the virtual loudspeaker signal by c(t) = Ψw(t), (8)

[0080] Equation (8) is the inverse operation of Equation (5)

[0081] Therefore, Equations (8) and (7) are used to limit the total power of all HOA coefficient sequences as follows

[0082] ||c(lT S )|| 2 2 ≤||Ψ|| 2 2 ·||W(lT S )|| 2 2 ≤||Ψ|| 2 2 ·O (9)

[0083] Under the assumption of N3D normalization of the spherical harmonic function, the square of the Euclidean norm of the modulus matrix can be written as: ||Ψ|| 2 2 = K·O, (10a)

[0084] where represents the ratio between the square of the Euclidean norm of the modulus matrix and the number of HOA coefficient sequences, O. This ratio depends on the specific HOA order N and the specific virtual loudspeaker directions It can be expressed as follows by appending the corresponding parameter list to the ratio:

[0085]

[0086] Figure 3 shows the virtual orientation according to the article by Fliege et al. mentioned above Value of K with respect to HOA order (N=1, ..., 29).

[0087] Combining all previous arguments and considerations, the following upper bound on the magnitude of the HOA coefficient sequence is provided:

[0088]

[0089] Among them, the first inequality follows directly from the definition of the norm.

[0090] It is important to note that the conditions in equation (6) imply the conditions in equation (11), but the reverse is not true, ie, equation (11) does not imply equation (6).

[0091] Another important aspect is that, under the assumption that the virtual speaker positions are approximately uniformly distributed, the column vectors of the modulus matrix Ψ representing the modulus vectors about the virtual speaker positions are almost orthogonal to each other and each has a Euclidean norm of N+1. This property means that, except for the multiplication constant, the spatial transformation almost preserves the Euclidean norm, i.e.,

[0092] ||c(lT S )|| 2 ≈(N+1)||w(lT S )|| 2 (12)

[0093] The true norm ||c(lT S )|| 2 The greater the deviation from the approximation in equation (12), the more the assumption of orthogonality of the modulus vectors is violated.

[0094] Results of the value range of the main sound signal

[0095] Both types of (directional and vector-based) primary sound signals have in common that their contribution to the HOA representation consists of a single vector with Euclidean norm N+1 To describe, that is, ||v 1 || 2 =N+1. (13)

[0096] In the case of a directional signal, this vector is related to the direction Ω of a certain signal source. S,1 The modulus vector of , that is,

[0097]

[0098] This vector describes the directional beam as the signal source direction Ω using the HOA representation S,1 In the case of vector-based signals, the vector v 1 It is not limited to modulus vectors about any direction, and thus can describe more general directional distributions of vector-based monophonic signals.

[0099] Next, consider D main sound signals x d (t), d = 1, ..., D, the D main sound signals can be concentrated in the vector x(t) according to the following formula:

[0100] x(t)=[x 1 (t) x 2 (t)...x D (t)] T (16)

[0101] These signals must be determined based on the following matrix:

[0102] V:=[v 1 v 2 ...v D ] (17)

[0103] The matrix is ​​represented by the monophonic main sound signal x d (t), d = 1, ..., all vectors v in the direction distribution of D d , d=1,...,D is composed.

[0104] For the meaningful extraction of the main sound signal x(t), the following constraints are specified:

[0105] a) Each dominant sound signal is obtained as a linear combination of the coefficient sequence of the original HOA representation, i.e.

[0106] x(t)=A·c(t), (18)

[0107] in, represents the mixing matrix.

[0108] b) The mixing matrix A should be chosen so that its Euclidean norm does not exceed the value "1", i.e.,

[0109]

[0110] And the square (or power) of the Euclidean norm of the residual between the original HOA representation and the HOA representation of the main sound signal is not greater than the square (or power) of the Euclidean norm of the original HOA representation, that is,

[0111]

[0112] By substituting formula (18) into formula (20), it can be seen that formula (20) is equivalent to the following constraint:

[0113]

[0114] Where I represents the identity matrix.

[0115] Using formula (18), formula (19) and formula (11), according to the constraints in formula (18) and formula (19) and according to the compatibility of the Euclidean matrix and the vector norm, the upper limit of the amplitude of the main sound signal is limited by the following formula:

[0116]

[0117] Therefore, it is ensured that the main sound signal remains in the same range as the original HOA coefficient sequence (compare with formula (11)), that is,

[0118] Example of selecting a mixing matrix

[0119] An example of how to determine a mixing matrix that satisfies constraint (20) is obtained by calculating the main sound signal so that the Euclidean norm of the residual after extraction is minimized, that is,

[0120] x(t)=argmin x(t) ||V·x(t)-c(t)|| 2 (26)

[0121] The solution to the minimization problem in equation (26) is given by:

[0122] x(t)=V + c(t), (27)

[0123] in,(·) + Denotes the Moore-Penrose generalized inverse. By comparing formula (27) with formula (18), it can be concluded that in this case, the mixing matrix is ​​equal to the Moore-Penrose generalized inverse of the matrix V, that is, A = V + .

[0124] However, the matrix V must still be chosen to satisfy constraint (19), i.e.,

[0125] In the case of directional signals only, the matrix V is related to some source signal direction Ω S,d, =1,..., the modular matrix of D, that is

[0126] V=[S(Ω S,1 S(Ω S,2 )...S(Ω S,D )], (29)

[0127] Constraint (28) can be satisfied by selecting the source signal directions ΩS,d, d=1, ..., D so that the distance between any two adjacent directions is not too small.

[0128] The value range of the coefficient sequence of the ambient HOA component results

[0129] The ambient HOA component is calculated by subtracting the HOA representation of the primary sound signal from the original HOA representation, i.e., c AMB (t) = c(t) - V·x(t). (30)

[0130] If the vector of the main sound signal x(t) is determined according to criterion (20), it can be concluded that:

[0131]

[0132]

[0133] The value range of the spatial transformation coefficient sequence of the ambient HOA component

[0134] Another aspect of the HOA compression process proposed in EP 2743922 A1 and the above-mentioned MPEG document N14264 is that the first O of the ambient HOA component MIN The coefficient sequence is always selected to be assigned to the transmission channel, where O MIN (N MIN +1) 2 , N MIN ≤N is usually a smaller order than the original HOA representation. In order to decorrelate these HOA coefficient sequences, they can be transformed to be from some predefined direction Ω MIN,d , d = 1, ..., O MIN (Similar to the concept described in the Normalization of the Input HOA Representation section) The virtual loudspeaker signal of the impact.

[0135] Use c AMB,MIN (t) to define the order index as n≤N MIN The vector of all coefficient sequences of the ambient HOA components and Ψ MIN To define the virtual direction Ω MIN,d, d = 1, ..., O MIN The vector of all virtual loudspeaker signals is defined as MIN (t) is obtained by the following formula:

[0136]

[0137] Therefore, using the compatibility of Euclidean matrices with vector norms,

[0138]

[0139] In the above-mentioned MPEG document N14264, the virtual direction Ω is selected according to the above-mentioned article by Fliege et al. MIN,d , d = 1, ..., O MIN . Figure 4 The modular matrix Ψ MIN The inverse matrix of the order (N MIN =1, ..., 9). It can be seen that for N MIN =1,...,9, (39)

[0140] However, this does not usually apply to The value of is usually much larger than N of "1". MIN >9. However, at least for 1≤N MIN ≤9, the amplitude of the virtual speaker signal is limited by the following formula:

[0141]

[0142] By constraining the input HOA representation to satisfy condition (6), where condition (6) requires that the amplitude of the virtual loudspeaker signal created from the HOA representation does not exceed the value "1", it is guaranteed that the amplitude of the signal before gain control will not exceed the value under the following conditions (See formula (25), formula (34) and formula (40)):

[0143] a) The vector of all dominant sound signals x(t) is calculated according to formulas / constraints (18), (19) and (20);

[0144] b) If virtual loudspeaker positions as defined in the above-mentioned article by Fliege et al. are used, the number of first coefficient sequences of ambient HOA components to which the spatial transformation is applied is determined. MIN The minimum order N MIN Must be less than "9".

[0145] It can be further concluded that for up to the maximum order N of interest MAXAny order N, that is, 1≤N≤N MAX , the amplitude of the signal before gain control will not exceed the value in,

[0146]

[0147] In particular, from Figure 3 It can be concluded that if the virtual speaker directions used for the initial spatial transformation are assumed is chosen based on the distribution in Fliege et al., and if we additionally assume that the maximum order of interest is N MAX = 29 (see, for example, MPEG document N14264), the amplitude of the signal before gain control will not exceed the value 1.50, because in this special case That is, you can choose

[0148] K MAX Depends on the maximum order N of interest MAX and virtual speaker directions It can be expressed by the following formula:

[0149]

[0150] Therefore, the minimum gain applied by gain control to ensure that the signal before perceptual coding is within the interval [-1, 1] is given by Given, where

[0151]

[0152] In the case where the amplitude of the signal before gain control is too small, MPEG document N14264 proposes that up to A factor of 2 is used to smoothly amplify them, where e MAX ≥0 is transmitted as side information in the encoded HOA representation.

[0153] Therefore, each index with a base of "2" describing the total absolute amplitude change of the modified signal from the first frame to the current frame caused by the gain control processing unit within the access unit can be assumed to be in the interval [e MIN , e MAX ] any integer value within ]. Therefore, the (minimum integer) number of bits required for encoding is β e Given by:

[0154]

[0155] When the amplitude of the signal before gain control is not too small, formula (42) can be simplified to:

[0156]

[0157] This number of bits β can be calculated at the input of the gain control processing step / stage 15, ..., 151 e .

[0158] Use this number of bits β for the exponent e It ensures that all possible absolute amplitude changes caused by the HOA compressor gain control processing unit can be captured, allowing decompression to start at some predefined entry points in the compressed representation.

[0159] When the compressed HOA representation is decompressed in the HOA decompressor, the side information assigned to some data frames and in addition to the received data stream In addition, the non-differential gain value received from the demultiplexer 21 and representing the total absolute amplitude change is used in the inverse gain control step or stage 24,…,241 to implement correct gain control in a manner opposite to the processing performed in the gain control processing step / stage 15,…,151.

[0160] Other Embodiments

[0161] When implementing a specific HOA compression / decompression system as described in the sections HOA compression, spatial HOA encoding, HOA decompression and spatial HOA decoding, the number of bits β used to encode the exponent e Must depend on the scaling factor K MAX,DES According to formula (42), the scaling factor K is set MAX,DES itself depends on the desired maximum order N of the HOA representation to be compressed MAX,DES and specific virtual speaker directions

[0162] For example, when assuming that N MAX,DES = 29 and according to Fliege et al., a reasonable choice for choosing the virtual speaker direction is In this case, the order of the pair is guaranteed to be N (1≤N≤N MAX ) is correctly compressed using the HOA representation of the same virtual speaker direction Normalized according to the normalization of the chapter input HOA representation. However, this guarantee cannot be given in the case of an HOA representation that is also (for efficiency reasons) equivalently represented by a virtual loudspeaker signal in PCM format, but where the directions of the virtual loudspeakers are The virtual loudspeaker orientations are chosen to correspond to those assumed during the system design phase. different.

[0163] Due to this different choice of virtual loudspeaker positions, it is no longer guaranteed that the amplitude of the signal before gain control will not exceed the value , even if the amplitude of these virtual loudspeaker signals is in the interval [-1,1]. Therefore, there is no guarantee that the HOA representation has proper normalization for compression according to the process described in MPEG document N14264.

[0164] In this case, it is advantageous to have a system that provides the maximum allowed amplitude of the virtual loudspeaker signal based on the knowledge of the virtual loudspeaker position to ensure that the corresponding HOA representation is suitable for compression according to the process described in MPEG document N14264. Figure 5 Such a system is shown in . It uses virtual speaker positions As input, And provides the maximum allowed amplitude of the virtual loudspeaker signal γ dB (which is measured in decibels) as output. In step or stage 51, the modulus matrix Ψ is calculated for the virtual loudspeaker positions according to formula (3). In a subsequent step or stage 52, the Euclidean norm of the modulus matrix ‖Ψ‖ is calculated 2 In a third step or stage 53, the amplitude γ is calculated as the minimum of "1" and the square root of the number of virtual loudspeaker positions and K MAX,DES The quotient of the square root of and the Euclidean norm of the modular matrix,

[0165] Right now

[0166] The value in decibels is obtained by the following formula: γ dB =20log 10 (γ). (44)

[0167] To illustrate: From the above derivation, it can be seen that if the amplitude of the HOA coefficient sequence does not exceed the value That is, if

[0168]

[0169] All signals before the gain control processing unit will accordingly not exceed this value, which is a requirement for proper HOA compression.

[0170] From formula (9), it is found that the amplitude of the HOA coefficient sequence is limited by

[0171] ||c(lT S )|| ∞ ≤||c(lT S )|| 2 ≤||Ψ|| 2 ·||W(lTS )|| 2 (46)

[0172] Therefore, if γ is set according to formula (43) and the virtual speaker signal in PCM format satisfies

[0173] ||w(lT S )|| ∞ ≤γ, (47)

[0174] Then from formula (7) we can get

[0175] And meets requirement (45).

[0176] That is, the maximum amplitude value "1" in formula (6) is replaced by the maximum amplitude value γ in formula (47).

[0177] The basis of high-order Ambisonics

[0178] Higher-order Ambisonics (HOA) is based on the description of the sound field within a dense region of interest, which is assumed to be free of sound sources. In this case, the spatiotemporal behavior of the sound pressure p(t, x) at time t and position x within the region of interest is physically completely determined by the homogeneous wave equation. In the following, it is assumed that Figure 6 The spherical coordinate system shown. In the coordinate system used, the x-axis points forward, the y-axis points to the left, and the z-axis points to the top. The position x in space = (r, θ, φ) T It is represented by the radius r>0 (i.e., the distance to the origin of the coordinate system), the inclination angle θ∈[0,π] measured from the polar axis z, and the azimuth angle φ∈[0,2π[ measured counterclockwise from the x-axis in the xy plane. In addition, (·) T Indicates transpose.

[0179] Then, from the textbook "Fourier Acoustics" it can be seen that the Fourier transform of the sound pressure with respect to time is given by Indicates, that is,

[0180]

[0181] Where ω represents the angular frequency and i represents the imaginary unit. The Fourier transform of the above sound pressure with respect to time can be expanded into a series of spherical harmonic functions according to the following formula:

[0182]

[0183] Among them, c s represents the speed of sound, k represents the angular wave number, which is And is related to the angular frequency ω. In addition, j n(·) represents the spherical Bessel function of the first kind, and represents the real-valued spherical harmonic functions of order n and degree m, which are defined in the section Definitions of Real-Valued Spherical Harmonic Functions. depends only on the angular wave number k. Note that it has been implicitly assumed that the sound pressure is spatially band-limited. Therefore, the order is truncated with respect to the order index n at an upper limit N of the order called the HOA representation.

[0184] If the sound field is represented by the superposition of an infinite number of harmonic plane waves with different angular frequencies ω arriving from all possible directions specified by the angle tuple (θ, φ), it can be seen (see B. Rafaely, "Plane-wave decomposition of the sound field on a sphere by spherical convolution", J. Acoust. Soc. Am, Vol. 4 (116), pp. 2149-2157, October 2004) that the corresponding plane wave complex amplitude function C(ω, θ, φ) can be represented by the following spherical harmonic function expansion:

[0185]

[0186] Among them, the expansion coefficient Through the following formula and expansion coefficient Related:

[0187]

[0188] Assuming that the coefficients is a function of the angular frequency ω, then the inverse Fourier transform (given by The application of ) provides the following time domain functions for each order n and degree m

[0189]

[0190] These time domain functions are referred to here as the continuous-time HOA coefficient sequence, which can be concentrated into a single vector c(t) by

[0191]

[0192] The HOA coefficient sequence in the vector c(t) The position index is given by n(n+1)+1+m. The total number of elements in the vector c(t) is given by O=(N+1) 2 Given.

[0193] The final Ambisonics format uses a sampling frequency of f S Provide the following sampled version of c(t)

[0194]

[0195] Among them, T S =1 / f S Represents the sampling period. Element c(lT S ) is called the discrete-time HOA coefficient sequence, which can always be real-valued. This property also applies to the continuous-time version Definition of real-valued spherical harmonic functions

[0196] Real-valued spherical harmonic functions (assuming SN3D normalization according to J. Daniel, “Report of champs acoustiques, application to transmission and reproduction of scenes sonores complexes in a multimédia context”, PhD thesis, University of Paris, June 2001, chapter 3.1) is given by

[0197]

[0198] in,

[0199]

[0200] The associated Legendre function P n,m (x) is defined as

[0201]

[0202] It has Legendre polynomial P n (x), and unlike the one in E.G. Williams' "Fourier Acoustics" in Applied Mathematical Sciences, vol. 93, Academic Press, 1999, there is no Condon-Shortley phase term (-1) m .

[0203] The process of the present invention may be performed by a single processor or electronic circuit, or by several processors or electronic circuits working in parallel and / or working in different parts of the process of the present invention.

[0204] Instructions for operating the one or more processors may be stored in one or more memories.

Claims

1. A method for decoding a compressed higher-order Ambisonics (HOA) sound representation of a sound or sound field, the method include: receiving a bitstream comprising a compressed HOA representation, wherein the bitstream comprises a plurality of HOA coefficients corresponding to the compressed HOA representation, demultiplexing the compressed HOA representation from the bitstream, and The compressed HOA representation is decoded to determine a decoded HOA representation, wherein the decoding is based on a minimum integer number of bits β e and wherein the minimum integer bit number β e based on Sure, in, N is the order of HOA representation, N MAX is the maximum order of the HOA representation of interest, is the direction of the virtual speaker, O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the square of the Euclidean norm of the modular matrix ||Ψ|| 2 2 The ratio of O, and in, and The decoding at least includes main sound synthesis and environment synthesis.

2. A device for decoding a compressed Higher Order Ambisonics (HOA) sound representation of a sound or a sound field, the device receiving a bitstream comprising a compressed HOA representation, in, The bitstream comprises a plurality of HOA coefficients corresponding to the compressed HOA representation, and the apparatus comprises: a demultiplexer configured to demultiplex the compressed HOA representation from the bitstream, and A processor configured to decode the compressed HOA representation to determine a decoded HOA representation, wherein the decoding is based on a minimum integer bit number β e Conducted, Among them, the minimum integer bit number β e based on Sure, in, N is the order of HOA representation, N MAX is the maximum order of the HOA representation of interest, is the direction of the virtual speaker, O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the square of the Euclidean norm of the modular matrix ||Ψ|| 2 2 The ratio of O, and in, and The decoding at least includes main sound synthesis and environment synthesis.

3. A non-transitory computer-readable medium having executable instructions stored thereon for causing a computer to perform the steps of the method according to claim 1.

4. A method for decoding a compressed higher-order ambisonics (HOA) sound representation of a sound or sound field, the method include: demultiplexing a compressed HOA representation from the bitstream, wherein a plurality of HOA coefficients correspond to the compressed HOA representation, and When independent access units exist in the bitstream, based on the minimum integer bit number β e The compressed HOA representation is decoded, wherein the minimum integer bit number β e based on Sure, in, N is the order of HOA representation, N MAX is the maximum order of the HOA representation of interest, is the direction of the virtual speaker, O = (N + 1) 2 is the number of HOA coefficient sequences, and K is the square of the Euclidean norm of the modular matrix ||Ψ|| 2 2 The ratio of O and e MAX >0 is used to increase the minimum integer bit number β when the sampling value amplitude of the channel signal before gain control is less than the threshold e , is the maximum gain applied to the channel signal by the gain control.

5. A method for decoding a compressed higher-order Ambisonics (HOA) sound representation of a sound or sound field, the method include: receiving a bitstream containing the compressed HOA representation and decoding the compressed HOA representation to determine a perceptually decoded signal The associated gain correction index e i (k) and gain correction abnormality flag β i (k); By perceiving the decoding signal The associated gain correction index e i (k) and gain correction abnormality flag β i (k) performing inverse gain control processing to provide a gain-corrected signal frame Redistribute gain-corrected signal frames during channel reallocation In order to reconstruct the frame of the main sound signal and the intermediate representation frame C of the ambient HOA component I,AMB (k).

Citation Information

Patent Citations

  • Method for determining the minimum number of integer bits required to represent non-differential gain values ​​for compression of HOA data frame representation

    CN113808600B

  • Method and apparatus for compressing and decompressing a Higher Order Ambisonics signal representation

    EP2665208A1

  • Method and apparatus for compressing and decompressing a higher order ambisonics representation for a sound field

    EP2743922A1

  • Method and Apparatus for compressing and decompressing a Higher Order Ambisonics representation

    EP2800401A1

  • Method and Apparatus for generating from a coefficient domain representation of HOA signals a mixed spatial / coefficient domain representation of said HOA signals

    EP2824661A1