Method or apparatus for compressing or decompressing higher-order ambisonic signal representations

By decomposing HOA signals into directional and ambient components and applying perceptual coding, the method achieves efficient compression with reduced data rates and noise exposure, maintaining high spatial resolution.

JP2026069501APending Publication Date: 2026-04-23DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2025-12-24
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for compressing Higher-Order Ambisonics (HOA) representations suffer from high data rates and perceptual coding noise due to cross-correlations between HOA coefficient sequences, leading to inefficient compression and reduced spatial resolution.

Method used

The method decomposes the HOA signal into dominant directional signals and ambient components, reduces the order of the ambient component, and applies perceptual coding to both, followed by inverse transformation to reconstruct the HOA representation.

Benefits of technology

This approach maintains high spatial resolution while significantly reducing data rates, minimizing perceptual coding noise exposure, and improving compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069501000001_ABST
    Figure 2026069501000001_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for compressing and decompressing higher-order ambisonics (HOA) representations, which process directional and ambient components in different formats. [Solution] The compression method involves a dominant direction estimation unit 22 that estimates the dominant direction and performs a process to decompose the ambisonic signal C(l) into a directional component and an ambient component. The directional component is calculated in stage 23, thereby determining the direction of the ambisonic representation. TIFF2026069501000138.tif7153 The signal is converted into a time-domain signal represented by the above, and the residual ambient component is calculated in stage 24, and the HOA domain coefficient C is calculated. A Expressed by (l).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for compressing and decompressing higher-order ambisonic representations, in which the directional and ambient components are processed in different ways. [Background technology]

[0002] Higher-Order Ambisonics (HOA) offers the advantage of obtaining a complete sound field near a specific location in three-dimensional space (a location called the "sweet spot"). Such HOA representations are independent of specific speaker configurations, and in this respect, they differ from channel-based technologies such as stereo or surround sound. This flexibility comes at the cost of the decoding process having to reproduce the HOA representation for a specific speaker configuration.

[0003] HOA is based on the complex amplitude representation of air pressure for individual angular wavenumbers k at a location x near the desired listener position, and without loss of generality, we may assume that the listener position is the origin of the spherical coordinate system, and HOA is expressed using a truncated spherical harmonics (SH) expansion. The spatial resolution of this representation is improved by increasing the maximum order N of the expansion. Unfortunately, the number of expansion coefficients O increases quadratically with respect to the order N, specifically O=(N+1) 2 For example, a typical HOA representation using order N=4 requires 0=25 coefficients. If the desired sampling rate is f s The number of bits per sample is N b If so, the overall bitrate for transmitting the HOA signal representation is O·f s ·N b Determined by, the order N=4, and the sampling rate f s =48kHz, and the number of bits per sample is N bWhen the HOA signal representation is set to =16, the transmission rate becomes as high as 19.2 Mbit / s. Therefore, compression of the HOA signal representation is highly desirable.

[0004] An overview of existing spatial audio compression methods is described in Patent Document 1 or Non-Patent Document 1, etc.

[0005] The following technologies are relevant to the background technology of this invention.

[0006] The B-format signal is equivalent to a first-order ambisonic representation, and the B-format signal can be compressed using Directional Audio Coding (DirAC), as described in Non-Patent Document 2.

[0007] In one proposed application for video conferencing, the B-format signal is encoded in the form of a single omnidirectional signal and side information, with a single directional and frequency band-specific dispersion parameter. However, this results in a significant reduction in data rate, at the cost of only minor signal quality degradation during playback. Furthermore, DirAC is limited to first-order ambisonic compression, suffering from the disadvantage of very low spatial resolution.

[0008] Few existing methods are known for compressing HOA representations when N>1. One method involves using the Perceptual Advanced Audio Coding (AAC) codec to perform direct encoding of individual HOA coefficient sequences, as described, for example, in Non-Patent Document 3. However, an essential problem with such methods is that they perform perceptual coding of signals that are never actually heard. The reconstructed playback signal is typically obtained by weighted summation of HOA coefficient sequences. When the decompressed HOA representation is expressed in relation to a specific speaker arrangement, there is a high probability that perceptual coding noise will be exposed. More precisely, the main problem associated with identifying perceptual coding noise is the high cross-correlation between individual HOA coefficient sequences. The coded noise signals in individual HOA coefficient sequences are usually uncorrelated or low-correlated with each other, so a constructive superposition of perceptual coding noise occurs, while noise-free HOA coefficient sequences are canceled out by the superposition. Another problem is that the above cross-correlation leads to a decrease in the efficiency of perceptual coding.

[0009] To minimize the extent of such effects, Patent Document 1 proposes converting the HOA representation to an equivalent representation in the spatial domain before perceptual coding. In addition to corresponding to conventional directional signals, the spatial domain signal will also correspond to speaker signals if the (multiple) speakers are positioned in exactly the same direction as assumed in the spatial domain conversion.

[0010] The conversion to the spatial domain reduces the cross-correlation between individual spatial domain signals. However, the cross-correlation is not completely eliminated. A specific example of a directional signal that yields relatively high cross-correlation is when the direction of the directional signal lies between adjacent directions covered by (multiple) spatial domain signals. Another drawback of Patent Document 1 and Non-Patent Document 3 is that the number of perceptual coding signals is (N+1) 2 This means that N is the order of the HOA representation, and therefore the data rate of the compressed HOA representation increases quadratically with respect to the order of the ambisonics.

[0011] As described later, the compression process according to the present invention decomposes the HOA sound field representation into a directional component and an ambient component. In particular, a new process for estimating multiple dominant sound directions is described herein in relation to the calculation of the directional sound field component.

[0012] Regarding existing direction estimation methods based on ambisonics, the method described in Non-Patent Document 2 above relates to DirAC coding for direction estimation based on B-format sound field representation. Direction is obtained from an average intensity vector indicating the direction in which sound field energy flows. An alternative example based on B-format is described, for example, in Non-Patent Document 4. Direction estimation is performed iteratively by searching for the direction in which a beamformer output signal directed in a particular direction yields maximum power.

[0013] However, both methods are constrained by the B-format of direction estimation and suffer from the disadvantage of relatively small spatial resolution. Another drawback is that such estimations are limited to a single dominant direction.

[0014] HOA representations result in improved spatial resolution and enable improved estimation of multiple dominant directions. Few existing methods are known for performing estimation of multiple directions based on HOA sound field representations. Methods based on compression detection have been proposed in Non-Patent Documents 5 and 6. The main idea is to estimate a spatially sparse sound field, i.e., to construct only a small number of directional signals. After setting a number of test directions on a sphere, an optimal algorithm is executed to find as few test signals as possible for the corresponding directional signals, so that the test directions are adequately described by a given HOA representation. This method results in improved spatial resolution compared to the spatial resolution actually provided by a given HOA representation because it avoids the spatial dispersion caused by the limited order of a given HOA representation. However, the performance of the algorithm is highly dependent on whether the sparsity assumption is met. In particular, this method becomes problematic when the sound field contains some minor additional ambient components or when the HOA representation is affected by noise generated when calculated by multi-channel recording.

[0015] Furthermore, an intuitive method, as described in Non-Patent Document 7, is to convert a given HOA representation into a spatial domain and then search for the maximum value of the directional power. The drawback of this method is that the presence of ambient components obscures the directional power distribution and distorts the maximum directional power compared to the case where no ambient components are present. [Prior art documents] [Patent Documents]

[0016] [Patent Document 1] European Patent Application Publication No. 10306472.1 Specification [Non-patent literature]

[0017] [Non-Patent Document 1] I. Elfitri, B. Gunel, AM Kondoz, “Multichannel Audio Coding Based on Analysis by Synthesis”, Proceedings of the IEEE, vol.99, no.4, pp.657-670, April 2011 [Non-Patent Document 2] V. Pulkki, “Spatial Sound Reproduction with Directional Audio Coding”, Journal of Audio Eng. Society, vol.55(6), pp.503-5 16, 2007 [Non-Patent Document 3] E. Hellerud, I. Burnett, A. Solvang, U. Peter Svensson, “Encoding Higher Order Ambisonics with AAC”, 124th AES Conven tion, Amsterdam, 2008 [Non-Patent Document 4] D. Levin, S. Gannot, EAP Habets, “Direction-of-Arrival Estimation using Acoustic Vector Sensors, in the Presence of Noise”, IEEE Proc. of the ICASSP, pp.105-108, 2011 [Non-Patent Document 5] N. Epain, C. Jin, A. van Schaik, “The Application of Compressive Sampling to the Analysis and Synthesis of Spatial Sound Fields”, 127th Convention of the Audio Eng. Soc, New York, 2009, [Non-Patent Document 6] A. Wabnitz, N. Epain, A. van Schaik, C Jin, “Time Domain Reconstruction of Spatial Sound Fields Using Compressed Sensing”, IEEE Proc. of the ICASSP, pp.465-468, 2011 [Non-Patent Document 7] B. Rafaely, “Plane-wave decomposition of the sound field on a sphere by spherical convolution”, J. Acoust. Soc. Am., vol.4, no.116, pp .2149-2157, October 2004 [Overview of the Initiative]

[0018] The problem to be solved by the embodiment is to compress the HOA signal while maintaining high spatial resolution of the HOA signal representation. This problem is solved by the method described in the claims. The present application also discloses an apparatus utilizing such a method.

[0019] This invention relates to compressing the Higher-Order Ambisonics (HOA) representation of a sound field. In this application, "HOA" refers not only to the higher-order ambisonic representation but also to the associated encoded or represented audio signal. A dominant sound direction is estimated, and the HOA signal representation is decomposed into a plurality of dominant directional signals and associated directional information in the time domain and an ambient component in the HOA domain, the ambient component then being compressed to reduce its order. After the decomposition, the reduced-order ambient component is transformed into the spatial domain and, together with the directional signals, is subjected to perceptual coding processing.

[0020] At the receiver or decoder, the encoded directional signal and the de-ordered and encoded ambient component are subjected to perceptual decompression processing. The perceptually decompressed ambient signal is converted to a de-ordered HOA domain representation and then subjected to order expansion processing. From the directional signal and corresponding directional information, as well as the original order ambient HOA component, the complete or final HOA representation is reconstructed.

[0021] Advantageously, ambient sound field components can be represented with sufficient accuracy by HOA representation at a lower order than the original, and the extraction of dominant directional signals ensures that high spatial resolution is achieved after compression and decompression.

[0022] In principle, the method of the present invention is suitable for compressing higher-order ambisonic (HOA) signal representations. Steps include: estimating a dominant direction, wherein the dominant direction depends on the directional power distribution of an energetically dominant HOA signal component; and decomposing or decoding the HOA signal component into a plurality of dominant directional signals and associated directional information in the time domain and a residual ambient component in the HOA domain, wherein the residual ambient component represents the difference between the HOA signal representation and the representation of the dominant directional signal. The steps include compressing the residual ambient component by reducing its order from its original order, The steps include converting the reduced-order residual ambient component into a spatial domain, The steps include perceptual encoding the converted residual ambient component and the dominant directional signal, This is a method that has [something].

[0023] In principle, the method of the present invention is suitable for decompressing a compressed higher-order ambisonics (HOA) signal representation, wherein the compression is A step of estimating the dominant direction, wherein the dominant direction depends on the directional power distribution of the energetically dominant HOA signal component, A step of decomposing or decoding the HOA signal component into a plurality of dominant directional signals and associated directional information in the time domain and a residual ambient component in the HOA domain, wherein the residual ambient component represents the difference between the HOA signal representation and the representation of the dominant directional signals. The steps include compressing the residual ambient component by reducing its order from its original order, The steps include converting the reduced-order residual ambient component into a spatial domain, The method comprises the steps of perceptually encoding the converted residual ambient component and the dominant directional signal, and the method is The steps include perceptual decoding of the perceptually encoded dominant directional signal and the perceptually encoded transformed residual ambient component, The steps include: inversely transforming the perceptually decoded and transformed residual ambient components to obtain a representation of the HOA region; The steps include performing an order extension process on the inversely transformed residual ambient component to obtain the ambient HOA component of the original order, The steps include: synthesizing the perceptually decoded dominant directional signal, the directional information, and the original order ambient HOA component to obtain an HOA signal representation; This is a method that has [something].

[0024] In principle, the device of the present invention is suitable for compressing higher-order ambisonic (HOA) signal representations. Means adapted to estimate a dominant direction, wherein the dominant direction depends on the directional power distribution of the energetically dominant HOA signal component, Means adapted to decompose or decode the HOA signal component into a plurality of dominant directional signals and associated directional information in the time domain and a residual ambient component in the HOA domain, wherein the residual ambient component represents the difference between the HOA signal representation and the representation of the dominant directional signals. A means adapted to compress the residual ambient component by reducing the order of the residual ambient component from its original order, Means adapted to convert the reduced-order residual ambient components into a spatial domain, The apparatus comprises means adapted to perceptually encode the converted residual ambient component and the dominant directional signal.

[0025] In principle, the apparatus of the present invention is suitable for decompressing a compressed higher-order ambisonics (HOA) signal representation, and the above compression is A step of estimating the dominant direction, wherein the dominant direction depends on the directional power distribution of the energetically dominant HOA signal component, A step of decomposing or decoding the HOA signal component into a plurality of dominant directional signals and associated directional information in the time domain and a residual ambient component in the HOA domain, wherein the residual ambient component represents the difference between the HOA signal representation and the representation of the dominant directional signals. The steps include compressing the residual ambient component by reducing its order from its original order, The steps include converting the reduced-order residual ambient component into a spatial domain, The apparatus comprises a step formed to perceptually encode the converted residual ambient component and the dominant directional signal, and the apparatus is configured to perform A means configured to perceptually decode a perceptually encoded dominant directional signal and a perceptually encoded transformed residual ambient component, Means formed to invert the perceptually decoded transformed residual ambient component and obtain a representation in the HOA domain, Means formed to perform a degree expansion process on the inverted residual ambient component and obtain an ambient HOA component of the original degree, Means formed to synthesize the perceptually decoded dominant directional signal, the direction information, and the ambient HOA component of the original degree to obtain an HOA signal representation, and an apparatus having the same.

Brief Description of the Drawings

[0026] [Figure 1] A diagram showing a normalized variance function for various ambisonics degrees N and angles Θ ∈ [0, π]. [Figure 2] A block diagram related to the compression process according to the present invention. [Figure 3] A block diagram related to the decompression process according to the present invention.

Mode for Carrying Out the Invention

[0027] <Detailed Description of the Embodiment> Ambisonics signals describe the sound field in a source-free region using spherical harmonic (SH) expansion. The feasibility of this theory is due to the physical property that the temporal and spatial behavior of sound pressure is essentially determined by the wave equation.

[0028] <Wave Equation and Spherical Harmonic Expansion> To provide a detailed description of ambisonics, in the following, a spherical coordinate system or a polar coordinate system is assumed, and a point x = (r, θ, φ) in space T is represented by a radius r > 0 (i.e., the distance to the origin of the coordinate system), an inclination angle θ ∈ [0, π] with respect to the z-axis which is the original line or the polar axis, and an azimuth angle φ ∈ [0, 2π] measured from the x-axis in the xy plane. In this spherical coordinate system, the wave equation for the sound pressure p(t, x) in a connected source-free area is given as follows.

number

[0029] The Fourier transform of sound pressure with respect to time is given by the following equation.

number

number

[0030] In equation (4), k represents the angular wavenumber defined by the following equation.

number

[0031] Furthermore, Y n m (θ,φ) is an SH function with order n and degree m.

number

[0032] The associated Legendre function with respect to non-negative order m is the Legendre polynomial P n m This is determined by (x).

number

[0033] For negative orders (i.e., m < 0), associated Legendre functions are defined as follows:

number

[0034] Furthermore, Legendre polynomials P n (x)(n≧0) may be defined using Rodrigues' Formula.

number

[0035] Alternatively, the Fourier transform of a sound wave with respect to time is the real number SH function S n m It may also be expressed using (θ,φ). The real SH function may also be referred to as the real SH function, real SH function, etc.

[0036]

number

number

number

[0037] The real-valued SH function, by definition, takes real values, but the corresponding expansion coefficient q n m This is not generally true for (kr).

[0038] The complex SH function has the following relationship with the real SH function:

number

[0039] Direction vector Ω :=(θ,φ) T The complex SH function Y n m (θ,φ) and the real number SH function S n m (θ,φ) is the unit sphere S in three-dimensional space. 2 It forms an orthogonal basis for squared integrable complex valued functions on the above.

number

[0040] <Internal Problems and Ambisonics Coefficients> The goal of ambisonics is to represent the sound field near the origin of a coordinate system. Without loss of generality, the region of study is assumed to be a sphere or ball with radius R from the center of the coordinate system, mathematically specified by the set {x|0≦r≦R}. A key assumption regarding this representation is that this ball contains no sound sources. The problem of finding a representation of the sound field within this ball is referred to as the "internal problem" (e.g., Williams' book mentioned above).

[0041] Regarding internal problems, the SH function expansion coefficient P n m It is understood that (kr) can be expressed as follows:

number

number

number

[0042] <Plane wave decomposition> The sound field inside a ball without a sound source, centered at the origin of the coordinate system, can be represented as an infinite superposition of plane waves of various angular wavenumbers k incident on the ball from all possible directions (for more on this point, see, for example, "Plane-wave decomposition..." in Williams' book mentioned above). Assuming that the complex amplitude of the plane wave of angular wavenumber k from the direction of Ω0 is given by D(k,Ω0), the corresponding ambisonic coefficients for the order SH function expansion are given by the following equation, similar to the derivation method using equations (11) and (19).

number

[0043] Therefore, the ambisonics coefficient for a sound field obtained by superimposing an infinite number of plane waves of angular wavenumber k is given by all possible directions Ω0∈S in equation (20). 2 It is obtained from the integral with respect to .

number

[0044] The function D(k,Ω) is referred to as the "amplitude density" and is located on a unit sphere S. 2 It is assumed that the function is square-integrable. This can be expanded into a series of real SH functions as follows:

number

number

[0045] By substituting equation (24) into equation (22), we obtain the ambisonic coefficient b. n m (k) is the expansion coefficient cn m It can be seen that this is a version of (k) with a different scale. That is, it can be written as follows: b n m (k) = 4πi n c n m (k) (25)

[0046] Scaled Ambisonics coefficient c n m Applying the inverse Fourier transform with respect to time to (k) and the amplitude density function D(k,Ω), we obtain the following equation as the corresponding time-domain representation.

number

number

[0047] The time-domain directional signal d(t,Ω) may also be expressed by a real-valued SH function expansion according to the following equation.

number

[0048] SH function S n m Using the knowledge that (Ω) takes real values, the complex conjugate of d(t,Ω) can be expressed as follows:

number

[0049] Below, c~ n m (t) is sometimes referred to as the scaled time-domain ambisonics coefficient. Furthermore, in the following explanation, it is assumed that the sound field representation is described by these coefficients, which will be explained in detail in the following section on compression.

[0050] Coefficient c~ used in the processing according to the present invention n m The time domain is represented by the HOA representation c of the corresponding frequency domain. n m It should be noted that this is equivalent to (k). Therefore, the compression and decompression described can be equivalently realized in the frequency domain with some modification of the formulas.

[0051] <Spatial resolution of finite order> In reality, the sound field near the origin of the coordinate system has a finite number of ambisonic coefficients c of order n ≤ N. n m It is described using only (k). Calculating the amplitude density function from the truncated SH function series according to the following equation introduces a certain spatial dispersion component to the true amplitude density function D(k,Ω) (see, for example, "Plane-wave decompression..." in the above literature).

number

number

[0052] In equation (34), the ambisonics coefficient for plane waves in equation (20) is used, and in equations (35) and (36), several mathematical theories are used (see, for example, "Plane-wave decompression..." in the above-mentioned literature). The properties of equation (33) can be shown using equation (14).

[0053] Comparing equation (37) with the true amplitude density function, we obtain the following equation.

number

[0054] ν N The first zero point of (Θ) is approximately at π / N for N≧4 (see, for example, "Plane-wave decompression..." in the above literature), and the effect of variance decreases as the ambisonic order N increases (and the spatial resolution also improves).

[0055] As N→∞, the variance function ν N (Θ) converges to the scaled Dirac delta function. This is obtained by using the completeness relation of Legendre polynomials (equation (41)) together with equation (35) for ν as N→∞. N This can be understood by expressing the limit of (Θ).

number

number

[0056] The following equation defines a vector of real SH functions of degree n ≤ N:

number

[0057] In the time domain, variance can be expressed equivalently as follows:

number

[0058] <Sampling> In a certain application, there are a finite number of J discrete directions Ω j From a sample of the time-domain amplitude density function, the scaled time-domain ambisonic coefficient C~ n m It is desirable to determine (t). The integral in equation (28) can be approximated by a finite sum of terms as follows, according to B. Rafaely, "Analysis and Design of Spherical Microphone Arrays", IEEE Transactions on Speech and Audio Processing, vol.13, no.1, pp.135-143, January 2005.

number

[0059] If this condition is not met, equation (50) will be affected by spatial aliasing errors. This point is discussed, for example, in B. Rafaely, "Spatial Aliasing in Spherical Microphone Arrays", IEEE Transactions on Signal Processing, vol.55, no.3, pp.1003-1010, March 2007.

[0060] The next necessary condition is the sampling point Ω j Furthermore, the corresponding weight coefficients must satisfy the conditions described in "Analysis and Design..." of the aforementioned book.

number

[0061] Conditions (51) and (52) are sufficient for accurate sampling.

[0062] The sampling conditions (52) form a set of linear equations, which can be compactly expressed using a single matrix equation as shown below. ΨGΨ H =I (53) Here, Ψ represents the mode matrix defined by the following equation.

number

[0063] According to Equation (53), it can be seen that the condition necessary for Equation (52) to hold is that the number J of sampling points satisfies J ≧ O. The values of the amplitude density in the time domain at J sampling points are summarized in vector form as follows,

Number

Number

[0064] Using the introduced vector notation, calculating the scaled time-domain ambisonics coefficients from the values of the amplitude density function samples in the time domain can be expressed as follows. c(t) ≒ ΨGw(t) (59)

[0065] For a given fixed ambisonics order N, it is often not possible to calculate the number J ≧ O of sampling points Ω and the corresponding weight coefficients so that the sampling condition Equation (52) holds. However, when the sampling points are selected so that the sampling condition is sufficiently approximated, the rank of the mode matrix Ψ becomes O and the number of conditions decreases. In that case, there exists a pseudo-inverse matrix Ψ j of the mode matrix Ψ, and + Ψ Ψ + := (ΨΨ H ) -1 ΨΨ H (60) A reasonable approximation of the scaled time-domain ambisonics coefficient vector c(t) from the time-domain amplitude density function sample vector is: c(t) ≈ Ψ + w(t) (61) It is given by.

[0066] If J=O and the rank of the mode matrix is ​​O, the pseudo-inverse matrix is ​​equal to the inverse matrix because the following equation holds. Ψ + =(ΨΨ H ) -1 Ψ=Ψ -H Ψ -1 Ψ=Ψ -H (62)

[0067] Furthermore, if the sampling condition formula (52) is satisfied, Ψ -H =ΨG (63) The equation holds true, and the approximate formulas (59) and (61) are equivalent and coincide.

[0068] The vector w(t) can be interpreted as a vector of time-domain signals with respect to space. The transformation from the HOA domain to the spatial domain can be performed, for example, by equation (58). This type of transformation is referred to in this application as the "spherical harmonic transformation (SHT)" and is used when the de-ordered ambient HOA components are transformed into the spatial domain. The spatial sampling point Ω with respect to the SHT. j is g j It is implicitly assumed that the sampling conditions of equation (52) are approximately satisfied along with ≈4π / O(j=1,...,J), and that J=O. Under these assumptions, the SHT matrix is ​​Ψ H ≈(4π / O)Ψ -1 The following relationship is satisfied. If the scaling of the absolute value with respect to SHT is not important, (4π / O) may be ignored.

[0069] <Compression> This invention relates to the compression of a given HOA signal representation. As described above, the HOA signal representation is decomposed into a predetermined number of dominant directional signals in the time domain and ambient components in the HOA domain, followed by a process of compressing the HOA representation of the ambient components by lowering the order. This process assumes that the test is being monitored and utilizes the assumption that the surrounding sound field components can be represented sufficiently accurately by the low-order HOA representation. By extracting the dominant directional signals, it is possible to ensure that high spatial resolution is maintained after the compression and corresponding decompression processes.

[0070] After decompression, the reduced-order ambient HOA component is converted to a spatial domain and perceptually encoded along with a directional signal as shown in Patent Document 1.

[0071] The compression process involves two consecutive steps, as shown in Figure 2. The precise definitions of each signal are explained in detail in the following description of compression.

[0072] In the first step or stage of Figure 2(a), the dominant direction estimation unit 22 estimates the dominant direction and decomposes the ambisonics signal C(l) into a directional component and an ambient component, where "l" represents the frame index. The directional component is calculated in the directional signal calculation step or stage 23, thereby determining the ambisonics representation as a group of D normal directional signals X(l) and the corresponding direction.

number

[0073] In the second step shown in Figure 2(b), the perceptual coding process for the directional signal X(l) and the ambient HOA component is performed as follows: A typical time-domain directional signal X(l) can be individually compressed in the perceptual encoder 27 using some known perceptual compression technique. Ambient HOA region component C A The compression of (l) is performed in two substeps or stages.

[0074] The first substep or stage 25 is the original ambisonics order N. RED (For example, N RED The process is performed to reduce the ambient HOA component C to =2). A,RED (l) is obtained. It is assumed that the components of the ambient sound field can be represented sufficiently accurately by a low-order HOA. The second substep or stage 26 is based on compression as described in Patent Document 1. O with respect to the components of the ambient sound field RED :=(N RED +1) 2 Individual HOA signals C A,RED (l) is calculated in substep / stage 25, and these signals are obtained by applying a spherical harmonic transform to O in the spatial domain. RED Individual equivalent signals W A,RED This is converted to (l) and becomes a normal time-domain signal that can be input into a bank of parallel perceptual encoders 27. Some existing perceptual coding or compression technique can be applied. Encoded directional signal

number

number

[0075] Advantageously, all time-domain signals X(l) and W A,REDThe perceptual compression of (l) can be performed together in the perceptual encoder 27 and improves the overall coding efficiency by utilizing potentially remaining inter-channel correlations.

[0076] <Decompression> Figure 3 shows the decompression process for a received or regenerated signal. Similar to the compression process, it involves two steps.

[0077] In the first step or stage shown in Figure 3(a), the perceptual decoding unit 31 processes the encoded directional signal

number

number

number

number

number

number

number

number

[0078] In the second step or stage shown in Figure 3(b), the HOA signal construction unit 34 generates a directional signal.

number

number

number

number

[0079] <Achievable reduction in required data rate> The problem solved by embodiments of the present invention is to achieve a significant reduction in data rate compared to existing compression methods for HOA representations. Below, we discuss the achievable compression ratio for an uncompressed HOA representation. The compression ratio is obtained from the ratio of the data rate required to transmit an uncompressed HOA signal C(l) of order N to the data rate required to transmit a compressed signal representation, where the compressed signal representation consists of D perceptually encoded directional signals X(l) and corresponding directional information.

number

[0080] When transmitting an uncompressed HOA signal C(l), O·f s ·N bA data rate of is required. In contrast, to transmit D encoded directional signals X(l), D·f b,COD A data rate of f is required. b,COD This indicates the bitrate of the perceptually encoded signal. Similarly, N RED Individual perceptually encoded spatial domain signals W A,RED (l) Signal transmission is O RED ·f b,COD Requires a bitrate. Direction

number

[0081] Therefore, the transmission of the compressed representation is approximately (D+O RED )·f b,COD This requires a data rate of r. Therefore, the compression ratio r COMPR This can be expressed as follows:

number

[0082] For example, if the order is N=4 and the sampling rate is f s =48kHz, N per sample b = 16 bits, the number of dominant directions D=3, and the reduced HOA order is N RED When = 2 and the bitrate is 64 kbits / s, the compression ratio of the HOA representation is r COMPR This results in a compression ratio of approximately 25. Transmission of the compressed representation requires a data rate of approximately 768 kbits / s.

[0083] <Reducing the probability of unmasked encoded noise occurring> As explained in the background technology section, the perceptual compression of spatial domain signals described in Patent Document 1 is susceptible to the influence of residual cross-correlations between signals, raising concerns about the exposure (unmasking) of perceptual coding noise. According to the present invention, the dominant directional signal is first extracted from the HOA sound field representation before perceptual coding. This means that when constructing the HOA representation, after perceptual decoding, the coding noise has spatial directivity that precisely matches its directional signal. In particular, the influence of not only the coding noise but also the directional signal on any direction is deterministically described by the spatial dispersion function, as explained in the section on finite-order spatial resolution. In other words, at any given time, the HOA coefficient vector representing the coding noise is exactly a multiple of the HOA coefficient vector representing the directional signal. Therefore, any weighted addition of HOA coefficients including noise will not lead to any exposure of perceptual coding noise.

[0084] Furthermore, although low-order ambient components are also described in Patent Document 1, by definition, the spatial domain signals of ambient components show only low correlation with each other, thus reducing the likelihood of perceptual noise being exposed.

[0085] <Improved direction estimation> The direction estimation according to the present invention relies on the directional power distribution of the energetically dominant HOA component. The directional power distribution is calculated from a rank-reduced correlation matrix of the HOA representation, which is obtained from the eigenvalue decomposition of the correlation matrix of the HOA representation.

[0086] Compared with the direction estimation used in "Plane-wave decomposition..." of the above-mentioned book, this embodiment brings the advantage of high precision. The reason is that instead of using all HOA representations for direction estimation, by focusing on the dominant HOA components from the perspective of energy, the spatial ambiguity of the directional power distribution can be reduced.

[0087] Compared with the direction estimation proposed in the above-mentioned literature "The Application of Compressive Sampling to the Analysis and Synthesis of Spatial Sound Fields" and "Time Domain Reconstruction of Spatial Sound Fields Using Compressed Sensing", the present invention brings the advantage of excellent robustness. This is because the decomposition of HOA representation into directional components and ambient components is rarely achieved completely, and a small amount of ambient components remain in the directional components (and yet the direction estimation can still be continued appropriately). Compressive sampling methods such as those in the above two literatures are feared to fail to provide proper direction estimation results due to being very sensitive to the presence of ambient signals.

[0088] Advantageously, the direction estimation according to the present invention is not affected by such concerns.

[0089] <Alternative example of decomposing HOA representation> The technique of decomposing the HOA representation into a plurality of directional signals and related direction information and the ambient component of the HOA region can be used for signal-adaptive DirAC like rendering of the HOA representation according to the method shown in Pulkki's literature "Spatial Sound Reproduction with Directional Audio Coding".

[0090] Since the physical properties of the two components are different, each of the HOA components can be rendered separately. For example, the directional signals can be rendered to the speakers using a signal panning technique such as Vector Based Amplitude Panning (VBAP). VBAP is described, for example, in the following literature: Pulkki, "Virtual Sound Source Positioning Using Vector Base Amplitude Panning", Journal of Audio Eng. Society, vol.45, no.6, pp.456- 466, 1997. The ambient HOA components can be processed using existing standard HOA rendering techniques.

[0091] Such rendering is not limited to an ambisonics representation with degree "1", but can be understood as an extension of DirAC-like rendering for HOA representations with degree N>1.

[0092] Estimation of multiple directions based on the HOA signal representation can be used for any associated sound field analysis.

[0093] The signal processing steps are described in more detail below.

[0094] <Compression> <Determination of input format> As input, the scaled time-domain HOA coefficients determined by Equation (26)

Number

Number

[0095] <Framing> The arriving vector c(j) of the scaled HOA coefficients is framed in a non - overlapping frame group of length B as follows in the framing step or stage 21: [Number] Assuming that the sampling rate is fs = 48 kHz and the appropriate frame length is B = 1200 samples, the duration of the frame corresponds to 25 ms.

[0096] <Estimation of the dominant direction> To estimate the dominant direction, the following correlation matrix is calculated: [Number] The sum over the current sample l and the L - 1 past frames (l' = 0 to L - 1) indicates that the direction analysis is based on a long overlapping frame group of L·B samples, i.e., for each current frame, the content of the adjacent frames is considered. This contributes to the stability of the direction analysis for two reasons: (1) longer frames result in a larger number of observations, and (2) the direction estimation is smoothed due to the overlapping frames.

[0097] f s Assuming f = 48 kHz and B = 1200, an appropriate value of L is, for example, 4, which corresponds to a total frame duration of 100 ms.

[0098] Next, the eigenvalue decomposition of the correlation matrix B(l) is performed as B(l)=V(l)Λ(l)V T (l) (68) where the matrix V(l) is formed by the eigen - vectors v i (l)(1≦i≦O) as follows: [Number] The matrix Λ(l) corresponds to the eigenvalues λ as follows: i It is a diagonal matrix with (1 ≤ i ≤ O):

Number

[0099] And the index group {1,..., I^(l)} of the dominant eigenvalues is obtained. One possible way to do this is to calculate the desired minimum value DAR of the ratio of the broadband directional power to the ambient power MIN and determine I^(l) according to the following formula:

Number

[0100] An appropriate value of DAR MIN 15 dB may be selected. The number of dominant eigenvalues is limited so as not to exceed D, concentrating on at most D dominant directions. This is achieved by replacing the index group {1,..., I^(l)} with {1,..., I(l)}, where in this case, I(l): = max(I^(l), D) (73).

[0101] Next, an I(l)-rank approximation of B(l) is performed:

Number

[0102] And then, vectors like the following formula are calculated:

Number

[0103] The mode matrix Ξ is defined as follows:

number

[0104] σ 2 (l) element σ 2 q (l) is Ω q This approximately represents the power of the plane wave corresponding to the dominant direction signal arriving from that direction. A theoretical explanation for this point will be provided in the section on <Explanation of the Direction Search Algorithm>.

[0105] To determine the directional signal component, σ 2 (l)

number

number

number

[0106]

number

[0107]

number

number

number

[0108] The overall processing of calculations for all dominant directions can be performed by "Algorithm 1 for searching for dominant directions based on the power distribution on a sphere," as follows:

number

[0109] Next, the direction obtained for the current frame

number

number

[0110] (a) Current dominant direction

number

number

number

number

number

number

number

number

number

[0111] Note: If more time can be spent on the entire compression algorithm, the allocation of direction estimates may be performed in a way that provides greater robustness. For example, sudden changes in direction may be appropriately judged as outliers resulting from estimation errors and therefore not taken into consideration.

[0112] (b) Smoothing direction

number

number

number

number

number

[0113] With respect to the azimuth angle, smoothing needs to be modified to achieve appropriate smoothing in the transition from π-ε to -π (ε>0) and the reverse transition. This can be taken into account by performing the following process. First, the angle difference modulo 2π is calculated as shown in the following equation (modulo 2π is an operation modulo 2π):

number

number

[0114] The smoothed dominant azimuth angle (modulo 2π) is determined as follows:

number

number

[0115]

number

number

number

number

[0116] From now on, M ACT The set of active-direction indices shown by (l) is calculated. The key point is D ACT (l):=|M ACT (l)| is used to represent this.

[0117] All smoothed directions are concatenated into a single direction matrix:

number

[0118] <Calculation of Directional Signals> The calculation of the directional signal is based on mode matching. Specifically, a search is performed to find the directional signal, and the HOA representation of that directional signal provides the best approximation of a given HOA signal. Since changes in direction between consecutive frames can lead to discontinuities in the directional signal, after performing the estimation calculation of the directional signal for overlapping frames, the results for consecutive overlapping frames are smoothed using an appropriate window function. However, this smoothing introduces a one-frame delay.

[0119] The following describes the detailed estimation method for directional signals.

[0120] First, the mode matrix based on the smoothed active direction is calculated according to the following equation:

number

[0121] Next, matrix X contains the unsmoothed estimation results of all directional signals for the (l-1)th and (l)th frames. INST (l) is calculated:

number

[0122] This is performed in two steps. In the first step, the directional signal samples belonging to the row corresponding to the inactive direction are set to zero, as shown in the following equation:

number

[0123] In the second step, directional signal samples corresponding to the active direction are obtained by arranging a matrix according to the following equation.

number

number

[0124] Directional signal estimation result x INST,d (l,j)(1≦d≦D) can be reshaped by an appropriate window function w(j): x INST,WIN,d (l,j):=x INST,d (l,j)·w(j), 1≦j≦2B (99)

[0125] A concrete example of a window function is given by a periodic Hamming window, as shown in the following equation:

number

[0126] All smoothed directional signal samples for the (l-1)th frame are placed in matrix X(l-1) as follows:

number

[0127] <Calculation of Ambient HOA Components> Ambient HOA component C A (l-1) is derived from the overall HOA representation C(l-1) as shown in the following equation, and the overall directional HOA component C DIR Obtained by subtracting (l-1):

number

number

number

[0128] <Lowering the order of ambient HOA components> C A (l-1) can be expressed in terms of components as follows:

number

number

[0129] <Spherical Harmonic Transform of Ambient HOA Components> The spherical harmonic transform is applied to the lower-order ambient HOA component C. A,RED This is done by multiplying (l) by the inverse of the mode matrix:

number

[0130] <Decompression> <Inverse Spherical Harmonic Transformation> Spatial domain signals that have undergone perceptual compression decompression

number

number

number

[0131] <Degree expansion> HOA expression

Number

Number

[0132] <HOA coefficient construction> The final decompressed HOA coefficients are calculated by adding the directional component and the ambient HOA component as follows:

Number

[0133] To calculate the smoothed directional HOA component, two consecutive frames containing the estimation results of all individual directional signals are concatenated into one long frame according to the following formula:

Number

[0134] Each individual signal contained in this long frame is multiplied by a window function such as formula (100).

Number

number

number

number

[0135] Note that the overall direction HOA component C DIR (l-1) is obtained by encoding all windowed directional signals in the appropriate direction and superimposing them in an overlapping manner:

number

[0136] <Explanation of Direction Search Algorithms> The following explains the direction-finding algorithm mentioned in the section on <estimating the dominant direction>. First, it is based on several assumptions.

[0137] <Assumption> The HOA coefficient vector c(j) is generally related to the time-domain amplitude density function d(j,Ω) as shown in the following equation:

number

number

[0138] In this model, the HOA coefficient vector c(j) is in direction Ω in the l-th frame. xi (l) I dominant directional source signals x iWe show that (j) is formed by (1≦i≦I). In particular, the direction is assumed to be invariant for the duration of one frame. The number of dominant source signals I is assumed to be clearly smaller than the total number of HOA coefficients O. Furthermore, the frame length B is assumed to be clearly larger than O. Also, the vector c(j) is a residual component c that can represent an ideal isotropic peripheral sound field. A (Includes j).

[0139] The individual HOA coefficient vector components are assumed to have the following properties: • It is assumed that the dominant source signals are, on average, zero:

number

number

number

number

number

number

[0140] <Supplementary explanation regarding direction finding> For the sake of explanation, we consider a situation where the correlation matrix B(l) (equation (67)) is calculated based only on the sample of the l-th frame, without considering the samples of L-1 preceding frames. This process is equivalent to setting L to 1 (L=1). Therefore, the correlation matrix can be expressed as follows:

number

[0141] By substituting the model assumed in equation (120) into equation (128), and using equations (122), (123), and definition (124), the correlation matrix B(l) can be approximated as follows:

number

[0142] According to equation (131), B(l) can be found to consist of two additive components: an additive component belonging to the directional component and an additive component belonging to the ambient component. I(l) rank approximation B I (l) provides an approximation of the directional HOA component, that is, it can be written as follows:

number

[0143] However, the first item

number

number

[0144] In equation (135), the properties of spherical harmonics mentioned in equation (47) are used:

number

[0145] Formula (136) is σ 2 (l) element σ 2 q (l) is the test direction Ω q This shows that it approximates the power of the signal arriving from (1≦q≦Q).

Claims

1. A method for decompressing a compressed higher-order ambisonics (HOA) signal, which includes an encoded directional signal and an encoded ambient signal, Receiving the compressed HOA signal, The compressed HOA signal is perceptually decoded to generate a decoded directional HOA signal and a decoded ambient HOA signal, wherein an inverse spatial transform is applied to determine the decoded ambient HOA signal. The process involves performing order extension on the decoded ambient HOA signal to obtain a representation of the decoded ambient HOA signal, wherein the order extension is performed by adding a signal with zero-value samples to the decoded ambient HOA signal. Reconstructing the decoded HOA representation from the decoded ambient HOA signal and the decoded directional HOA signal, including, method.

2. The method according to claim 1, wherein the decoded HOA representation has a first order greater than 1.

3. A non-temporary computer-readable medium storing instructions for performing the method described in claim 1 when executed by a processor.

4. A device for decompressing compressed higher-order ambisonics (HOA) signals, which include encoded directional signals and encoded ambient signals, The compressed HOA signal is received by an input interface, An audio decoder that perceptually decodes the compressed HOA signal and generates a decoded directional HOA signal and a decoded ambient HOA signal, wherein the audio decoder includes an inverse converter that applies an inverse spatial transform to determine the decoded ambient HOA signal. A processor that performs order extension on the decoded ambient HOA signal to obtain a representation of the decoded ambient HOA signal, wherein the order extension is performed by adding a signal having zero-value samples to the decoded ambient HOA signal. A combiner that reconstructs a decoded HOA signal from the decoded ambient HOA signal and the decoded directional HOA signal, including, Device.

Citation Information

Patent Citations

  • Method and apparatus for encoding and decoding successive frames of an ambisonics representation of a 2- or 3-dimensional sound field

    EP2469741A1