Method or apparatus for compressing or decompressing a higher-order Ambisonics signal representation

By decomposing HOA signals into directional and ambient components and encoding them efficiently, the method achieves reduced data rates and improved spatial resolution with minimized noise exposure.

JP7797560B2Active Publication Date: 2026-01-13DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024062459
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-05-14
Filing Date
2024-04-09
Publication Date
2026-01-13
Estimated Expiration
2033-05-06

AI Technical Summary

Technical Problem

Existing methods for compressing Higher Order Ambisonics (HOA) representations suffer from high data rates due to quadratic growth with order N, poor signal quality, perceptual coding noise exposure, and limited spatial resolution, especially when transforming into the spatial domain.

Method used

Decompose the HOA signal into dominant directional and ambient components, reduce the order of ambient components, transform them into the spatial domain, and perceptually encode both for efficient compression and decompression.

Benefits of technology

Maintains high spatial resolution while significantly reducing data rate by accurately representing ambient components with lower-order HOA and minimizing perceptual coding noise exposure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797560000137
    Figure 0007797560000137
  • Figure 0007797560000138
    Figure 0007797560000138
  • Figure 0007797560000139
    Figure 0007797560000139
Patent Text Reader

Abstract

To provide a method and device for compressing and decompressing higher-order ambisonic representations that process directional and ambient components in different formats.SOLUTION: A compression method performs the process of estimating a dominant direction and decomposing an ambisonics signal C(l) into directional and ambient components in a dominant direction estimation section 22. I(L) indicates a frame index. The directional component is calculated in the directional signal calculation step or stage 23, the ambisonics representation is transformed into a time domain signal represented by a set of D normal directional signals X(l) and the corresponding direction, and the residual ambient component is calculated in an ambient HOA component calculation step or stage 24, is expressed by the HOA domain coefficients CA(l), and compression is performed. After that, it is order-extended to reconstruct the complete HOA representation from the direction signal, corresponding direction information and the ambient HOA components of an original order.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to methods and apparatus for compressing and decompressing higher-order Ambisonics representations, in which the directional and ambient components are processed differently. [Background technology]

[0002] Higher Order Ambisonics (HOA) offers the advantage of capturing the complete sound field near a specific location in three-dimensional space (called the "sweet spot"). Such an HOA representation is independent of the specific speaker setup, which distinguishes it from channel-based techniques such as stereo or surround. This flexibility comes at the cost of the decoding process having to reproduce the HOA representation for a specific speaker setup.

[0003] The HOA is based on a complex amplitude representation of air pressure for a particular angular wavenumber k at a location x near the desired listener position; without loss of generality, the listener position may be assumed to be the origin of a spherical coordinate system, and the HOA is expressed using a truncated spherical harmonics (SH) expansion. The spatial resolution of this representation is improved by increasing the maximum order N of the expansion. Unfortunately, the number of expansion coefficients O(O) grows quadratically with the order N, specifically, O=(N+1) 2 For example, a typical HOA representation using order N=4 requires O=25 coefficients. If the desired sampling rate is f s and the number of bits per sample is N b If , the overall bit rate for the transmission of the HOA signal representation is O f s N b is determined by the order N=4 and the sampling rate is f s = 48 kHz and the number of bits per sample is N bTransmission of the HOA signal representation with .DELTA..times ...

[0004] An overview of existing spatial audio compression methods is described in Patent Document 1 or Non-Patent Document 1, etc.

[0005] The following techniques qualify as background to the present invention.

[0006] A B-format signal is equivalent to a first-order Ambisonics representation, and B-format signals can be compressed using Directional Audio Coding (DirAC) as described in Non-Patent Document 2.

[0007] In one proposed implementation for videoconferencing applications, B-format signals are encoded into an omnidirectional signal and side information in the form of dispersion parameters for one direction and each frequency band. However, this significant reduction in data rate comes at the expense of poor signal quality being obtained upon playback. Furthermore, DirAC is limited to compressing a first-order Ambisonics representation, which suffers from very low spatial resolution.

[0008] Few existing methods for compressing HOA representations for N>1 are known. One method utilizes a perceptual advanced audio coding (AAC) codec to perform direct encoding of individual HOA coefficient sequences, as described, for example, in [Non-Patent Document 3]. However, an essential problem with such methods is perceptual coding of a signal that will never be heard. The reconstructed playback signal is typically obtained by weighted addition of HOA coefficient sequences. If the decompressed HOA representation is represented for a specific speaker arrangement, there is a high probability that perceptual coding noise will be exposed. More precisely, the main problem with identifying perceptual coding noise is the high cross-correlation between individual HOA coefficient sequences. Since the coding noise signals in the individual HOA coefficient sequences are typically uncorrelated or low with each other, constructive superposition of the perceptual coding noise occurs, while noise-free HOA coefficient sequences are cancelled by the superposition. Another problem is that this cross-correlation reduces the efficiency of perceptual coding.

[0009] To minimize the extent of such effects, Patent Document 1 proposes to transform the HOA representation into an equivalent representation in the spatial domain before perceptual encoding. In addition to corresponding to the traditional directional signal, the spatial domain signal would also correspond to the speaker signal if the loudspeakers were positioned in exactly the same direction as assumed in the spatial domain transformation.

[0010] Transforming to the spatial domain reduces the cross-correlation between the individual spatial domain signals. However, the cross-correlation is not completely eliminated. A specific example of a directional signal that results in a relatively high cross-correlation is when the direction of the directional signal is between adjacent directions covered by the spatial domain signals. Another drawback of Patent Document 1 and Non-Patent Document 3 is that the number of perceptually coded signals is (N+1) 2 where N is the order of the HOA representation. Therefore, the data rate of the compressed HOA representation grows quadratically with the Ambisonics order.

[0011] As will be described below, the compression process according to the present invention performs a process of decomposing the HOA sound field representation into directional and ambient components. In particular, for the calculation of the directional sound field components, a new process for estimating multiple dominant sound directions is described herein.

[0012] Regarding existing direction estimation methods based on Ambisonics, the method described in the above-mentioned non-patent document 2 relates to DirAC coding for direction estimation based on a B-format sound field representation. The direction is obtained from a mean intensity vector pointing in the direction in which the sound field energy flows. An alternative based on the B-format is described for example in non-patent document 4. Direction estimation is performed iteratively by searching for the direction in which the beamformer output signal directed in a particular direction yields the maximum power.

[0013] However, both methods are limited by the B-format of direction estimation and suffer from a relatively small spatial resolution. Another drawback is that such estimation is limited to a single dominant direction.

[0014] The HOA representation provides improved spatial resolution and enables improved estimation of multiple dominant directions. Few existing methods are known for performing multiple direction estimation based on the HOA sound field representation. A method based on compressive detection has been proposed in [5] and [6]. The main idea is to estimate a spatially sparse sound field, i.e., to construct only a small number of directional signals. After defining a large number of inspection directions on a sphere, an optimization algorithm is implemented to find as few inspection signals as possible with respect to the corresponding directional signals so that the inspection direction is adequately described by the given HOA representation. This method provides improved spatial resolution compared to the spatial resolution actually provided by the given HOA representation, because it avoids spatial dispersion due to the limited order of the given HOA representation. However, the performance of the algorithm strongly depends on whether the sparsity assumption is met. This method is particularly disadvantageous when the sound field contains some minor additional ambient components or when the HOA representation is affected by noise that occurs when it is calculated from multi-channel recordings.

[0015] Furthermore, an intuitive method is to transform the given HOA representation into the spatial domain and then search for the maximum of the directional power, as described in Non-Patent Document 7. The drawback of this method is that the presence of ambient components can obscure the directional power distribution and displace the directional power maximum compared to the case where no ambient components are present. [Prior art documents] [Patent documents]

[0016] [Patent Document 1] European Patent Application Publication No. 10306472.1 [Non-patent literature]

[0017] [Non-Patent Document 1] I. Elfitri, B. Gunel, AM Kondoz, “Multichannel Audio Coding Based on Analysis by Synthesis”, Proceedings of the IEEE, vol.99, no.4, pp.657-670, April 2011 [Non-patent document 2] V. Pulkki, “Spatial Sound Reproduction with Directional Audio Coding”, Journal of Audio Eng. Society, vol.55(6), pp.503-5 16, 2007 [Non-patent document 3] E. Hellerud, I. Burnett, A. Solvang, U. Peter Svensson, “Encoding Higher Order Ambisonics with AAC”, 124th AES Conven tion, Amsterdam, 2008 [Non-patent document 4] D. Levin, S. Gannot, EAP Habets, “Direction-of-Arrival Estimation using Acoustic Vector Sensors, in the Presence of Noise”, IEEE Proc. of the ICASSP, pp.105-108, 2011 [Non-Patent Document 5] N. Epain, C. Jin, A. van Schaik, “The Application of Compressive Sampling to the Analysis and Synthesis of Spatial Sound Fields”, 127th Convention of the Audio Eng. Soc, New York, 2009, [Non-patent document 6] A. Wabnitz, N. Epain, A. van Schaik, C Jin, “Time Domain Reconstruction of Spatial Sound Fields Using Compressed Sensing”, IEEE Proc. of the ICASSP, pp.465-468, 2011 [Non-Patent Document 7] B. Rafaely, “Plane-wave decomposition of the sound field on a sphere by spherical convolution”, J. Acoust. Soc. Am., vol.4, no.116, pp .2149-2157, October 2004 Summary of the Invention

[0018] The problem solved by the embodiments is to compress HOA signals while maintaining high spatial resolution of the HOA signal representation. This problem is solved by the methods set forth in the claims. The present application also discloses devices utilizing such methods.

[0019] The present invention relates to compressing a high-order Ambisonics HOA representation of a sound field. In this application, "HOA" refers not only to the high-order Ambisonics representation but also to the associated encoded or represented audio signal. The dominant sound direction is estimated, and the HOA signal representation is decomposed into a plurality of dominant directional signals and associated directional information in the time domain and ambient components in the HOA domain, after which the ambient components are compressed to reduce their order. After the decomposition, the reduced-order ambient components are transformed into the spatial domain and subjected to perceptual coding together with the directional signals.

[0020] At the receiver or decoder side, the encoded directional signal and the reduced-order encoded ambient components are subjected to a perceptual decompression process. The perceptually decompressed ambient signal is converted into a reduced-order HOA domain representation and then subjected to an order expansion process. A complete or final HOA representation is reconstructed from the directional signal and corresponding direction information, as well as the original-order ambient HOA components.

[0021] Advantageously, the ambient sound field components can be represented with sufficient accuracy by a lower-order HOA representation than the original, and the extraction of the dominant directional signals ensures that high spatial resolution is achieved after compression and decompression.

[0022] In principle, the method of the invention is suitable for compressing a Higher Order Ambisonics (HOA) signal representation, comprising: estimating a dominant direction, the dominant direction depending on the directional power distribution of the energetically dominant HOA signal components; decomposing or decoding the HOA signal components into a plurality of dominant directional signals and associated directional information in the time domain and residual ambient components in the HOA domain, the residual ambient components representing the difference between the HOA signal representation and a representation of the dominant directional signals; compressing the residual ambient components by reducing the order of the residual ambient components below their original order; transforming the reduced order residual ambient components into the spatial domain; perceptually encoding the transformed residual ambient components and the dominant directional signal; It is a method having the following.

[0023] In principle, the method of the invention is suitable for decompressing a compressed Higher Order Ambisonics (HOA) signal representation, said compression comprising: estimating a dominant direction, said dominant direction depending on the directional power distribution of the energetically dominant HOA signal components; decomposing or decoding the HOA signal components into a plurality of dominant directional signals and associated directional information in the time domain and residual ambient components in the HOA domain, the residual ambient components representing differences between the HOA signal representation and a representation of the dominant directional signals; compressing the residual ambient components by reducing the order of the residual ambient components below their original order; transforming the reduced order residual ambient components into the spatial domain; and perceptually encoding the transformed residual ambient components and the dominant directional signal, the method comprising: perceptually decoding the perceptually coded dominant directional signal and the perceptually coded transformed residual ambient component; inverse transforming the perceptually decoded transformed residual ambient components to obtain a representation in the HOA domain; performing an order expansion process on the inverse transformed residual ambient components to obtain original order ambient HOA components; combining the perceptually decoded dominant directional signal, the directional information, and the original-order ambient HOA components to obtain an HOA signal representation; It is a method having the following.

[0024] In principle, the device of the invention is suitable for compressing a Higher Order Ambisonics (HOA) signal representation, comprising: means adapted to estimate a dominant direction, said dominant direction depending on the directional power distribution of the energetically dominant HOA signal component; means adapted to decompose or decode the HOA signal components into a plurality of dominant directional signals and associated directional information in the time domain and residual ambient components in the HOA domain, the residual ambient components representing differences between the HOA signal representation and a representation of the dominant directional signals; means adapted to compress the residual ambient components by reducing the order of the residual ambient components below their original order; means adapted to transform the reduced order residual ambient components into the spatial domain; and means adapted to perceptually encode the transformed residual ambient component and the dominant directional signal.

[0025] In principle, the device of the invention is suitable for decompressing a compressed Higher Order Ambisonics (HOA) signal representation, said compression comprising: estimating a dominant direction, said dominant direction depending on the directional power distribution of the energetically dominant HOA signal components; decomposing or decoding the HOA signal components into a plurality of dominant directional signals and associated directional information in the time domain and residual ambient components in the HOA domain, the residual ambient components representing differences between the HOA signal representation and a representation of the dominant directional signals; compressing the residual ambient components by reducing the order of the residual ambient components below their original order; transforming the reduced order residual ambient components into the spatial domain; and a step configured to perceptually encode the transformed residual ambient component and the dominant directional signal, the apparatus comprising: means arranged to perceptually decode the perceptually coded dominant directional signal and the perceptually coded transformed residual ambient component; means configured to inverse transform the perceptually decoded transformed residual ambient components to obtain a representation in the HOA domain; means configured to perform an order expansion process on the inverse transformed residual ambient components to obtain original order ambient HOA components; a perceptually decoded dominant directional signal and means configured to combine said directional information with said original-order ambient HOA components to obtain an HOA signal representation. [Brief explanation of the drawings]

[0026] [Figure 1] Normalized dispersion functions for various Ambisonics orders N and angles Θ∈[0,π]. [Figure 2] FIG. 2 is a block diagram of a compression process according to the present invention. [Figure 3] FIG. 2 is a block diagram of a decompression process according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] <Detailed Description of Embodiments> Ambisonics signals describe the sound field in source-free areas using the spherical harmonic (SH) expansion. The feasibility of this theory stems from the physical property that the time and spatial behavior of sound pressure is essentially governed by the wave equation.

[0028] <Wave equation and spherical harmonic expansion> For the purposes of a detailed description of Ambisonics, a spherical or polar coordinate system is assumed below, where a point in space x=(r,θ,φ) T is expressed by the radius r>0 (i.e., the distance to the origin of the coordinate system), the tilt angle θ∈[0,π] relative to the z-axis, which is the primitive or polar axis, and the azimuthal angle φ∈[0,2π] measured from the x-axis in the xy-plane. In this spherical coordinate system, the wave equation for the sound pressure p(t,x) in a connected source-free area is given as follows:

number

[0029] The Fourier transform of sound pressure versus time is given by:

number

number

[0030] In the formula (4), k represents the angular wave number defined by the following formula:

number

[0031] Furthermore, Y n m (θ,φ) is an SH function of order n and degree m.

number

[0032] The associated Legendre function for non-negative order m is the Legendre polynomial P n m (x) is specified.

number

[0033] For negative orders (i.e., m<0), the associated Legendre functions are defined as follows:

number

[0034] Also, the Legendre polynomials P n (x)(n≧0) may be defined using Rodrigues' Formula.

number

[0035] Alternatively, the Fourier transform of a sound wave with respect to time is the real SH function S n m It may be expressed using (θ, φ). A real SH function may be referred to as a real SH function, a real SH function, or the like.

[0036]

number

number

number

[0037] Although real SH functions are real-valued by definition, the corresponding expansion coefficients q n m This does not hold generally for (kr).

[0038] The complex SH functions have the following relationship to the real SH functions:

number

[0039] Direction vector Ω:=(θ,φ) T along with the complex SH function Y n m (θ,φ) and real SH function S n m (θ,φ) is a unit sphere S in three-dimensional space 2 form an orthogonal basis for squared integrable complex valued functions on

number

[0040] <Internal Problems and Ambisonics Coefficients> The goal of Ambisonics is to represent the sound field near the origin of a coordinate system. Without loss of generality, the region of interest is assumed to be a sphere or ball of radius R from the center of the coordinate system, specified mathematically by the set {x|0≦r≦R}. An important assumption about this representation is that this ball is assumed not to contain any sound sources. The problem of finding a representation of the sound field inside this ball is referred to as the "interior problem" (e.g., in Williams' book mentioned above).

[0041] Regarding the internal problem, the SH function expansion coefficient P n m It is understood that (kr) can be expressed as follows:

number

number

number

[0042] <Plane wave decomposition> The sound field inside a ball with no sound source centered at the origin of the coordinate system can be expressed as an infinite superposition of plane waves of various angular wavenumbers k incident on the ball from all possible directions (for more on this, see, for example, "Plane-wave decomposition..." in the above-mentioned book by Williams). Assuming that the complex amplitude of a plane wave of angular wavenumber k from the direction Ω0 is given by D(k,Ω0), the corresponding Ambisonics coefficients for the order SH function expansion are given by the following equations, similar to the derivation performed using equations (11) and (19):

number

[0043] Therefore, the Ambisonics coefficients for the sound field obtained by superposing an infinite number of plane waves with angular wave number k are given by all possible directions Ω0∈S in equation (20). 2 is obtained from the integral with respect to

number

[0044] The function D(k,Ω) is referred to as the "amplitude density" and is expressed as a function of the unit sphere S 2 is assumed to be square integrable in , which can be expanded into a series of real SH functions as follows:

number

number

[0045] By substituting equation (24) into equation (22), the Ambisonics coefficient b n m (k) is the expansion coefficient cn m (k) is a scaled version of (k), i.e., b n m (k)=4πi n c n m (k) (25)

[0046] Scaled Ambisonics coefficients c n m Applying an inverse Fourier transform with respect to time to (k) and the amplitude density function D(k,Ω) gives the corresponding time domain expression:

number

number

[0047] The time domain directional signal d(t,Ω) may be expressed by a real SH function expansion according to the following equation:

number

[0048] SH function S n m Using the knowledge that (Ω) is real-valued, the complex conjugate of d(t,Ω) can be expressed as follows:

number

[0049] Below, c~ n m (t) are sometimes referred to as scaled time-domain Ambisonics coefficients, and in the following discussion it is assumed that the sound field representation is described by these coefficients, which are explained in more detail below in the section on compression.

[0050] The coefficient c~ used in the processing according to the invention n m The time domain is expressed by the corresponding frequency domain HOA representation c n m (k) Therefore, the compression and decompression described can be equivalently achieved in the frequency domain with some modification of the mathematical formulas.

[0051] <Spatial resolution of finite order> In reality, the sound field near the origin of the coordinate system is expressed as a finite number of Ambisonics coefficients c n m The calculation of the amplitude density function from the series of truncated SH functions according to the following formula introduces a certain spatial dispersion component to the true amplitude density function D(k,Ω) (see, for example, "Plane-wave decompression..." in the above reference).

number

number

[0052] In equation (34), the plane wave Ambisonics coefficients of equation (20) are used, and some mathematical theory is used in equations (35) and (36) (see, for example, "Plane-wave decompression..." in the above-mentioned publication). The properties of equation (33) can be shown using equation (14).

[0053] Comparing Equation (37) with the true amplitude density function, we obtain

number

[0054] ν N The first zero of (Θ) is approximately at π / N for N≧4 (see, for example, "Plane-wave decompression..." in the reference cited above), and the effect of dispersion decreases (and the spatial resolution improves) as the Ambisonics order N increases.

[0055] As N→∞, the dispersion function ν N (Θ) converges to the scaled Dirac delta function, which can be seen by using the completeness relation of the Legendre polynomials (Eq. (41)) together with Eq. (35) to find the ν N This can be understood by expressing the limit of (Θ).

number

number

[0056] If we define a vector of real SH functions of degree n≦N by the following equation,

number

[0057] Variance can be expressed equivalently in the time domain as

number

[0058] <Sampling> For some applications, there are a finite number J of discrete directions Ω j From the sample of the time domain amplitude density function at n m It is desirable to determine (t). The integral in equation (28) can be approximated by a finite sum according to B. Rafaely, "Analysis and Design of Spherical Microphone Arrays", IEEE Transactions on Speech and Audio Processing, vol. 13, no. 1, pp. 135-143, January 2005, as follows:

number

[0059] If this condition is not satisfied, Equation (50) will be affected by spatial aliasing errors, as described, for example, in B. Rafaely, "Spatial Aliasing in Spherical Microphone Arrays," IEEE Transactions on Signal Processing, vol. 55, no. 3, pp. 1003-1010, March 2007.

[0060] The next necessary condition is the sampling point Ω j and the corresponding weighting coefficients must satisfy the conditions described in "Analysis and Design..." of the above book.

number

[0061] Conditions (51) and (52) are sufficient for accurate sampling.

[0062] The sampling condition (52) forms a set of linear equations and can be compactly expressed using a single matrix equation as follows: ΨGΨ H =I (53) Here, Ψ denotes the modal matrix defined by the following equation:

number

[0063] According to Equation (53), it can be seen that the necessary condition for Equation (52) to hold is that the number of sampling points J satisfies J≧0. The amplitude density values ​​in the time domain at J sampling points are organized in vector form as follows:

number

number

[0064] Using the introduced vector notation, the calculation of scaled time-domain Ambisonics coefficients from the values ​​of the time-domain amplitude density function samples can be expressed as: c(t)≒ΨGw(t) (59)

[0065] For a given fixed Ambisonics order N, the sampling points Ω are chosen so that the sampling condition (52) holds. j It is often not possible to calculate the number J≧O of modal matrix Ψ and the corresponding weighting coefficients. However, if the sampling points are chosen so that the sampling conditions are well approximated, the rank of the modal matrix Ψ is O and the number of conditions is small. In that case, the pseudo-inverse of the modal matrix Ψ is + exists, Ψ + :=(ΨΨ H ) -1 ΨΨ H (60) A reasonable approximation to the scaled time-domain Ambisonics coefficient vector c(t) from the vector of time-domain amplitude density function samples is c(t)≒Ψ + w(t) (61) is given by

[0066] If J=0 and the rank of the modal matrix is ​​0, then the pseudo-inverse matrix is ​​equal to its inverse since the following equation holds: Ψ + =(ΨΨ H ) -1 Ψ=Ψ -H Ψ -1 Ψ=Ψ -H (62)

[0067] Furthermore, if the sampling condition (52) is satisfied, Ψ -H =ΨG (63) holds, and the approximate formulas (59) and (61) are equivalent and coincident.

[0068] The vector w(t) can be interpreted as a vector of time domain signals with respect to space. The transformation from the HOA domain to the spatial domain can be performed, for example, by Equation (58). This type of transformation is referred to herein as the "spherical harmonic transform (SHT)" and is used when the reduced order ambient HOA components are transformed into the spatial domain. The spatial sampling points Ω with respect to the SHT j is g j It is implicitly assumed that the sampling condition in (52) is approximately satisfied with ≈4π / O (j=1,...,J) and that J=O. Under these assumptions, the SHT matrix is H ≒(4π / O)Ψ -1 If the scaling of the absolute value with respect to SHT is not important, (4π / O) may be ignored.

[0069] <Compression> The present invention relates to the compression of a given HOA signal representation. As described above, the HOA signal representation is decomposed into a predetermined number of dominant directional signals in the time domain and ambient components in the HOA domain, followed by a process of compressing the HOA representation of the ambient components by order reduction. This process is based on the assumption that the ambient sound field components can be represented accurately enough by a low-order HOA representation, subject to monitoring tests. Extracting the dominant directional signals ensures that high spatial resolution is maintained after the compression and corresponding decompression processes.

[0070] After decompression, the reduced order ambient HOA components are transformed into the spatial domain and perceptually coded together with the directional signal as shown in US Pat. No. 6,213,999.

[0071] The compression process involves two successive steps as shown in Figure 2. The exact definition of the individual signals will be explained in detail in the following description of compression.

[0072] In the first step or stage of Figure 2(a), the dominant direction is estimated in the dominant direction estimation unit 22 and a process is carried out to decompose the Ambisonics signal C(l) into directional and ambient components, where "l" denotes the frame index. The directional components are calculated in a directional signal calculation step or stage 23, so that the Ambisonics representation is calculated as a set of D common directional signals X(l) and the corresponding directional components.

number

[0073] In the second step shown in FIG. 2(b), the process of perceptual coding for the directional signal X(l) and the ambient HOA components is performed as follows: The normal time domain directional signal X(l) can be separately compressed in a perceptual coder 27 using any known perceptual compression technique. Ambient HOA area component C A The compression of (l) is performed in two sub-steps or stages.

[0074] The first substep or stage 25 converts the original Ambisonics order N to N RED to (e.g., N RED = 2), and the ambient HOA component C A,RED (l) is obtained. It is assumed that the ambient sound field components can be represented accurately enough by low-order HOAs. The second substep or stage 26 is based on compression as described in US Pat. No. 5,619,239. O for the ambient sound field components is obtained. RED :=(N RED +1) 2 HOA signal C A,RED (l) are calculated in substep / stage 25, and these signals are transformed into O in the spatial domain by applying a spherical harmonic transform. RED Equivalent signals W A,RED (l) and becomes a regular time domain signal that can be input to a bank of parallel perceptual coders 27. Any existing perceptual coding or compression technique can be applied. The coded directional signal

number

number

[0075] Advantageously, all the time domain signals X(l) and W A,RED(l) Perceptual compression can be performed jointly in the perceptual coder 27, potentially improving overall coding efficiency by exploiting remaining inter-channel correlation.

[0076] <Decompress> The decompression process for a received or reproduced signal is shown in Figure 3. As with the compression process, two steps are involved.

[0077] In the first step or stage shown in FIG. 3(a), a perceptual decoder 31 decodes the encoded directional signal

number

number

number

number

number

number

number

number

[0078] In the second step or stage shown in FIG. 3(b), the HOA signal constructor 34 generates a directional signal

number

number

number

number

[0079] <Achievable reduction in required data rate> The problem solved by embodiments of the present invention is to achieve a significant reduction in data rate compared to existing compression methods for HOA representations. In the following, we discuss the achievable compression ratio for the uncompressed HOA representation. The compression ratio is obtained from the ratio of the data rate required to transmit the uncompressed HOA signal C(l) of order N to the data rate required to transmit the compressed signal representation, which is composed of D perceptually coded directional signals X(l) and the corresponding directional information

number

[0080] When transmitting the uncompressed HOA signal C(l), O·f s N bIn contrast, transmitting D coded directional signals X(l) requires a data rate of D f b,COD requires a data rate of f b,COD denotes the bit rate of the signal to be perceptually coded. Similarly, N RED The number of perceptually encoded spatial domain signals W A,RED (l) The transmission of signals is RED f b,COD Requires a bit rate of .

number

[0081] Therefore, transmission of the compressed representation is approximately (D+O RED )·f b,COD Therefore, the data rate required is r COMPR can be expressed as follows:

number

[0082] For example, if the order is N=4 and the sampling rate is f s = 48 kHz, N per sample b = 16 bits, the number of dominant directions is D = 3, and the reduced HOA order is N RED = 2, and the compression ratio of the HOA representation when the bit rate is 64 kbits / s is r COMPR This results in a compression ratio of approximately 25. Transmission of the compressed representation requires a data rate of approximately 768 kbits / s.

[0083] <Reducing the probability of unmasked coding noise> As explained in the Background Art section, the perceptual compression of spatial domain signals described in Patent Document 1 is subject to residual cross-correlation between signals, which may result in the unmasking of perceptual coding noise. According to the present invention, the dominant directional signal is first extracted from the HOA sound field representation before perceptual coding. This means that when constructing the HOA representation, after perceptual decoding, the coding noise will have a spatial directionality that closely matches the directional signal. In particular, the influence of the directional signal as well as the coding noise on any given direction is deterministically described by the spatial variance function, as explained in the section on finite-order spatial resolution. In other words, at any given time, the HOA coefficient vector representing the coding noise is exactly a multiple of the HOA coefficient vector representing the directional signal. Therefore, any weighted addition of noisy HOA coefficients will not result in any unmasking of perceptual coding noise.

[0084] Furthermore, reduced order ambient components are also described in Patent Document 1, but by definition, the spatial domain signals of the ambient components exhibit low correlation with each other, reducing the likelihood of exposing perceptual noise.

[0085] Improved direction estimation Our direction estimation relies on the directional power distribution of the energetically dominant HOA components, which is calculated from a rank-reduced correlation matrix for the HOA representation, obtained from an eigenvalue decomposition of the correlation matrix of the HOA representation.

[0086] When compared with the direction estimation used in "Plane-wave decomposition..." in the above-mentioned book, this embodiment brings the advantage of high precision. The reason is that instead of using all HOA representations for direction estimation, by focusing on the dominant HOA components from the perspective of energy, the spatial ambiguity of the directional power distribution can be reduced.

[0087] When compared with the direction estimation proposed in the above-mentioned literature "The Application of Compressive Sampling to the Analysis and Synthesis of Spatial Sound Fields" and "Time Domain Reconstruction of Spatial Sound Fields Using Compressed Sensing", the present invention brings the advantage of excellent robustness. This is because the decomposition of the HOA representation into directional components and ambient components is rarely achieved completely, and a small amount of ambient components remain in the directional components (and yet the direction estimation can still be continued appropriately). Compressive sampling methods such as those in the above two literatures are feared to fail to provide appropriate direction estimation results due to being very sensitive to the presence of ambient signals.

[0088] Advantageously, the direction estimation according to the present invention is not affected by such concerns.

[0089] <Alternative example of decomposing HOA representation> The technique of decomposing the HOA representation into a plurality of directional signals and related direction information and the ambient components in the HOA region can be used for signal-adaptive DirAC like rendering of the HOA representation according to the method shown in Pulkki's literature "Spatial Sound Reproduction with Directional Audio Coding".

[0090] Because the physical properties of the two components are different, each of the HOA components can be rendered separately. For example, the directional signal can be rendered to the speakers using a signal panning technique such as Vector Based Amplitude Panning (VBAP), which is described, for example, in Pulkki, "Virtual Sound Source Positioning Using Vector Based Amplitude Panning," Journal of Audio Eng. Society, vol. 45, no. 6, pp. 456-466, 1997. The ambient HOA component can be processed using existing standard HOA rendering techniques.

[0091] Such rendering is not limited to Ambisonics representations of order "1", but can be understood as an extension of DirAC-like rendering to HOA representations of order N>1.

[0092] The multiple direction estimation based on the HOA signal representation can be used for any relevant sound field analysis.

[0093] The signal processing steps are described in more detail below.

[0094] <Compression> <Determining the input format> As input, the scaled time-domain HOA coefficients determined by Eq. (26)

number

number

[0095] <Frame> The incoming vector c(j) of scaled HOA coefficients is framed in a framing step or stage 21 into non-overlapping frames of length B as follows:

number

[0096] <Estimation of dominant direction> To estimate the dominant direction, the following correlation matrix is ​​calculated:

number

[0097] f s = 48 kHz and B = 1200, a suitable value of L is, for example, 4, which corresponds to a total frame duration of 100 ms.

[0098] Next, the eigenvalue decomposition of the correlation matrix B(l) is B(l)=V(l)Λ(l)V T (l) (68) where the matrix V(l) is the eigenvalue vector v as follows: i (l) (1≦i≦O) is formed by:

number

number

[0099] Then, the set of indices {1,...,I^(l)} of the dominant eigenvalues ​​is found. One possible way to do this is to find the desired minimum value of the ratio of broadband directional power to ambient power, DAR MIN and determine I^(l) according to the following formula:

number

[0100] Appropriate DAR MIN may be chosen to be 15 dB. The number of dominant eigenvalues ​​is constrained to not exceed D, so as to concentrate on at most D dominant directions. This is achieved by replacing the index set {1,...,Î(l)} with {1,...,I(l)}, where I(l):=max(Î(l),D) (73).

[0101] Next, a rank I(l) approximation of B(l) is performed:

number

[0102] Then the following vector is calculated:

number

[0103] The modal matrix Ξ is defined as:

number

[0104] σ 2 σ, an element of (l) 2 q (l) is Ω q It approximately represents the power of a plane wave corresponding to a signal of a dominant direction coming from the direction of . The theoretical explanation of this point is explained in the section <Description of the Direction Search Algorithm>.

[0105] To determine the directional signal component, σ 2 (l)

number

number

number

[0106]

number

[0107]

number

number

number

[0108] The overall process of calculating all dominant directions can be performed by the following "Algorithm 1: Searching for dominant directions by power distribution on a sphere":

number

[0109] Then the orientation obtained for the current frame

number

number

[0110] (a) Current dominant direction

number

number

number

number

number

number

number

number

number

[0111] Note: If the overall compression algorithm is allowed to take longer, the allocation of the sequence of direction estimates may be performed to provide greater robustness, e.g., sudden direction changes may be appropriately discarded as outliers due to estimation errors.

[0112] (b) Smoothing direction

number

number

number

number

number

[0113] For azimuth angles, the smoothing needs to be modified to achieve proper smoothing at the transition from π-ε to -π (ε>0) and back. This can be taken into account by doing the following: First, the angular difference is calculated modulo 2π (modulo 2π is the modulo 2π operation):

number

number

[0114] The smoothed dominant azimuth angle (modulo 2π) is determined as follows:

number

number

[0115]

number

number

number

number

[0116] From now on, M ACT The indexes of the active directions indicated by (l) are calculated. ACT (l):=|M ACT (l) is expressed by |

[0117] All smoothed directions are concatenated into one direction matrix:

number

[0118] <Directional signal calculation> The calculation of the directional signal is based on mode matching. In particular, a search is performed to find a directional signal whose HOA representation provides the best approximation of the given HOA signal. Since changes in direction between successive frames may cause discontinuities in the directional signal, after performing the directional signal estimation calculation for overlapping frames, an appropriate window function is used to smooth the results for successive overlapping frames. However, the smoothing introduces a one-frame delay.

[0119] A detailed method for estimating the directionality signal will be described below.

[0120] First, a modal matrix based on the smoothed active directions is calculated according to the following formula:

number

[0121] Next, a matrix X containing the unsmoothed estimates of all directional signals for the (l-1)th and (l)th frames is INST (l) is calculated:

number

[0122] This is done in two steps: In the first step, the directional signal samples belonging to the rows corresponding to the inactive directions are set to zero, as shown in the following equation:

number

[0123] In a second step, the directional signal samples corresponding to the active directions are obtained by arranging the matrix according to

number

number

[0124] Directional signal estimation result x INST,d (l,j) (1≦d≦D) is shaped by an appropriate window function w(j): x INST,WIN,d (l,j):=x INST,d (l,j)·w(j), 1≦j≦2B (99)

[0125] An example of a window function is given by the periodic Hamming window:

number

[0126] All smoothed directional signal samples for the (l-1)th frame are arranged in a matrix X(l-1) as follows:

number

[0127] <Calculation of ambient HOA components> Ambient HOA Component C A (l-1) is calculated from the overall HOA expression C(l-1) as follows: DIR By subtracting (l-1) we get:

number

number

number

[0128] <Reducing the order of the ambient HOA component> C A (l-1) can be expressed in components as follows:

number

number

[0129] <Spherical harmonic transformation of ambient HOA components> The spherical harmonic transformation is performed on the reduced-order ambient HOA components C A,RED This is done by multiplying (l) with the inverse of the modal matrix:

number

[0130] <Decompress> <Inverse spherical harmonic transform> Perceptually decompressed spatial domain signals

number

number

number

[0131] <Degree expansion> HOA representation

Number

Number

[0132] <HOA coefficient construction> The final decompressed HOA coefficients are calculated by adding the directional component and the ambient HOA component as follows:

Number

number

number

number

[0135] In addition, the overall directional HOA component C DIR (l-1) is obtained by encoding all of the windowed directional signals in the appropriate direction and superimposing them in an overlapping fashion:

number

[0136] <Explanation of the direction search algorithm> The following explains the direction search algorithm mentioned in the section on "Estimation of Dominant Direction." First, this is based on several assumptions.

[0137] <Assumptions> The HOA coefficient vector c(j) is generally related to the time-domain amplitude density function d(j,Ω) as follows:

number

number

[0138] This model assumes that the HOA coefficient vector c(j) is in the lth frame in the direction Ω xi (l) I dominant directional source signals x i(j) (1≦i≦I). In particular, the direction is assumed to be constant for the duration of one frame. The number of dominant source signals I is assumed to be significantly smaller than the total number of HOA coefficients O. Furthermore, the frame length B is assumed to be significantly larger than O. Also, the vector c(j) is composed of residual components c that can represent an ideal isotropic ambient sound field. A Includes (j).

[0139] The individual HOA coefficient vector components are assumed to have the following properties: The dominant source signal(s) are assumed to be zero on average:

number

number

number

number

number

number

[0140] <Supplementary explanation regarding direction finding> For ease of explanation, consider the situation where the correlation matrix B(l) (Equation (67)) is calculated based only on the samples of the lth frame, without considering the samples of the L-1 previous frames. This process is equivalent to setting L to 1 (L=1). Therefore, the correlation matrix can be expressed as follows:

number

[0141] By substituting the model assumed in (120) into (128) and using (122), (123) and definition (124), the correlation matrix B(l) can be approximated as follows:

number

[0142] According to Equation (131), it can be seen that B(l) is approximately composed of two additive components: one attributed to the directional component and the other attributed to the ambient component. I(l) rank approximation B I (l) provides an approximation of the directional HOA component, i.e., it can be written as:

number

[0143] However, the first term

number

number

[0144] In (135), the properties of spherical harmonics mentioned in (47) are used:

number

[0145] Equation (136) is σ 2 Element σ of (l) 2 q (l) is the test direction Ω q It is shown that the power of the signal arriving from (1≦q≦Q) is approximated.

Claims

1. 1. A method for decompressing a high-order Ambisonics signal representation, the method comprising: receiving an encoded directional signal and an encoded ambient signal; perceptually decoding the encoded directional signals and the encoded ambient signals to generate decoded directional signals and decoded ambient signals, respectively; transforming the decoded ambient signal from a spatial domain to a higher-order Ambisonics domain representation of the ambient signal, the transforming comprising extending an order of the higher-order Ambisonics domain representation of the ambient signal; reconstructing a higher-order Ambisonics signal from the higher-order Ambisonics domain representation of the ambient signal and the decoded directional signal; Including, method.

2. The method of claim 1 , wherein the higher-order Ambisonics signal representation has an order greater than one.

3. 1. An apparatus for decompressing a higher-order Ambisonics signal representation, the apparatus comprising: an input interface for receiving the encoded directional signal and the encoded ambient signal; an audio decoder that perceptually decodes the encoded directional signals and the encoded ambient signals to generate decoded directional signals and decoded ambient signals, respectively; an inverse transformer that converts the decoded ambient signal from a spatial domain to a higher-order Ambisonics domain representation of the ambient signal, the conversion comprising extending the order of the higher-order Ambisonics domain representation of the ambient signal; and a synthesizer for reconstructing a higher-order Ambisonics signal from the higher-order Ambisonics domain representation of the ambient signal and the decoded directional signal; Including, Device.

4. The apparatus of claim 3 , wherein the higher-order Ambisonics signal representation has an order greater than one.

5. 1. A method for decompressing a high-order Ambisonics signal representation, the method comprising: receiving an encoded directional signal and an encoded ambient signal; perceptually decoding the encoded directional signals and the encoded ambient signals to generate decoded directional signals and decoded ambient signals, respectively; transforming the decoded ambient signal from a spatial domain to a higher-order Ambisonics domain representation of the ambient signal, the transforming comprising extending an order of the higher-order Ambisonics domain representation of the ambient signal; reconstructing a higher-order Ambisonics signal from the higher-order Ambisonics domain representation of the ambient signal and the decoded directional signal; smoothing the reconstructed higher-order Ambisonics signal, said smoothing being based on two consecutive frames of the reconstructed higher-order Ambisonics signal; Including, method.

6. 1. An apparatus for decompressing a higher-order Ambisonics signal representation, the apparatus comprising: an input interface for receiving the encoded directional signal and the encoded ambient signal; an audio decoder that perceptually decodes the encoded directional signals and the encoded ambient signals to generate decoded directional signals and decoded ambient signals, respectively; an inverse transformer that converts the decoded ambient signal from a spatial domain to a higher-order Ambisonics domain representation of the ambient signal, the conversion comprising extending the order of the higher-order Ambisonics domain representation of the ambient signal; and a synthesizer for reconstructing a higher-order Ambisonics signal from the higher-order Ambisonics domain representation of the ambient signal and the decoded directional signal; a smoother for smoothing a reconstructed higher-order Ambisonics signal, said smoothing being based on two consecutive frames of the reconstructed higher-order Ambisonics signal; Including, Device.

7. A non-transitory computer readable medium containing instructions that, when executed by a processor, perform the method of any one of claims 1, 2 and 5.

8. one or more processors; and one or more storage media storing instructions that, when executed by said one or more processors, cause said instructions to perform the method of any one of claims 1, 2 or 5. A device that decompresses higher-order Ambisonics signal representations.

9. Apparatus for decompressing a higher-order Ambisonics signal representation, comprising means for carrying out the method according to any one of claims 1, 2 and 5.

Citation Information

Patent Citations

  • Method and apparatus for encoding and decoding successive frames of an ambisonics representation of a 2- or 3-dimensional sound field

    EP2469741A1

  • Method or apparatus for compressing or decompressing higher-order ambisonic signal representations

    JP2015520411A

  • Method and apparatus for three-dimensional acoustic field encoding and optimal reconstruction

    US20110305344A1

  • Sound system

    US20120014527A1

  • Method and apparatus for encoding and decoding successive frames of an ambisonics representation of a 2- or 3-dimensional sound field

    US20120155653A1