Method and device for compressing and decompressing higher-order ambisonics representation for sound field
By decomposing and transforming HOA representations into dominant directional and residual components, the method reduces cross-correlations and data rate, enhancing compression efficiency and noise reduction in HOA sound field processing.
Patent Information
- Application Number
- JP2025060899
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-12-12
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-23
AI Technical Summary
Existing methods for compressing higher-order ambisonics (HOA) representations face high bit rates and inefficiencies due to cross-correlations between HOA coefficient sequences, leading to perceptual coding noise and increased data rates, especially when assumptions about sound fields are not met.
The method decomposes the HOA representation into dominant directional signals and residual components, transforms the residual into a discrete spatial region, predicts the residual from the dominant signals, and reduces the order of the residual components, followed by perceptual coding to minimize cross-correlations and data rate.
This approach reduces the number of signals to be coded, minimizes perceptual coding noise, and decreases the data rate while maintaining high spatial resolution, effectively addressing the inefficiencies of previous methods.
Smart Images

Figure 2025108471000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for compressing and decompressing a higher-order ambisonics representation for a sound field.
Background Art
[0002] The higher-order ambisonics representation, called HOA, is one way to represent three-dimensional sound. Other techniques are wave field synthesis (WFS) and channel-based methods such as 22.2. Compared with channel-based methods, HOA representations have the advantage of being independent of a specific loudspeaker setup. However, to obtain this flexibility, a decoding process is required to reproduce the HOA representation with a specific loudspeaker setup. Usually, HOA can be configured with a very small number of loudspeakers compared to the WFS approach where the number of required loudspeakers becomes very large. A further advantage of HOA is that the same representation can be used without the need to modify the binaural rendering to headphones.
[0003] HOA is based on the representation of the spatial density of complex harmonic plane wave amplitudes by means of a cut-off spherical harmonic (SH) expansion. Each expansion coefficient is a function of the angular frequency, and this can be equivalently represented by a time-domain function. Therefore, without loss of generality, a complete HOA sound field representation can actually be considered to be composed of "Ο" time-domain functions. Here, Ο represents the number of expansion coefficients. Refer to the following HOA coefficient sequence as having the same meaning as these time-domain functions.
[0004] The spatial resolution of the HOA representation improves with an increase in the maximum order N of the expansion. Unfortunately, the number "Ο" of expansion coefficients increases quadratically with respect to the order N, especially Ο = (N + 1) 2 and becomes. For example, a general HOA representation using order N = 4 requires Ο = 25 numbers of HOA (expansion) coefficients. Considering the above points, the total bit rate for the transmission of the HOA representation is the sampling rate f of the desired single channels and the number of bits N per sample b is given, Ο·f s ·N b can be obtained by. For each sample, N b = 16 bits are used to transmit the HOA representation of degree N = 4 at a sampling rate of f s = 48 kHz. As a result, the bit rate becomes 19.2 megabits / second, which is extremely high for many practical applications, such as streaming. Therefore, it is highly desirable to compress the HOA representation.
Summary of the Invention
[0005] There are few existing methods that deal with the compression of HOA representations higher than the first order. The most direct approach explored by E. Hellerud, I. Burnett, A. Solvang, and U. P. Svensson, "Encoding Higher Order Ambisonics with AAC" (Encoding Higher Order Ambisonics with AAC), 124th AES Convention, Amsterdam, 2008, is to directly encode individual HOA coefficient sequences using the perceptual coding algorithm AAC (Advanced Audio Coding). However, the inherent problem with this approach is the perceptual coding of signals that are never heard at all. The reconstructed playback signal is usually obtained by the weighted sum of the HOA coefficient sequences, and when the HOA representation decoded at a specific loudspeaker setting is rendered, it is highly likely to mask out the perceptual coding noise. The main problem with masking out the perceptual coding noise is the high cross-correlation between individual HOA coefficient sequences. Since the coding noise signals in individual HOA coefficient sequences are not correlated with each other, a structural superposition of the perceptual coding noise may occur, and at the same time, the HOA coefficient sequences without noise are canceled out by the superposition. Another problem is that these cross-correlations lead to a decrease in the efficiency of the perceptual coder.
[0006] In order to minimize the influence of both sides, in European Patent Application No. 2469742 (EP2469742A2), it is proposed to convert the HOA representation into an equivalent representation in a discrete spatial domain before perceptual coding. Formally, the discrete spatial domain is a time domain equivalent to the spatial density of complex harmonic plane wave amplitudes sampled in some discrete direction. Therefore, the discrete spatial domain is represented by "Ο" conventional time domain signals. This signal can be interpreted as a general plane wave coming from the sampling direction, and would correspond to the loudspeaker signal if the loudspeaker were located in exactly the same direction as assumed for the spatial domain conversion.
[0007] The conversion to the discrete spatial domain reduces the cross-correlation between individual spatial domain signals, but these cross-correlations are not completely removed. An example of a relatively high cross-correlation is a directional signal that is directed between a plurality of adjacent directions included by the spatial domain signal.
[0008] The main drawback of both approaches is that the number of signals to be perceptually coded is (N + 1) 2 and the data rate of the compressed HOA representation increases with the square of the Ambisonics order N.
[0009] In order to reduce the number of signals to be perceptually coded, European Patent Application Publication No. 2665208 proposes to decompose the HOA representation into a given maximum number of dominant directional signals and a residual ambient component. The reduction in the number of signals to be perceptually coded can be achieved by reducing the order of the residual ambient component. The theoretical basis behind this approach is to maintain a high spatial resolution for the dominant directional signals while representing the residual with sufficient accuracy by a lower order HOA representation.
[0010] This approach works very well as long as the assumptions regarding the sound field are met, i.e., as long as the assumption is met that the sound field consists of a small number of dominant directional signals (which represent a general plane wave function encoded at full order N) and a non-directional residual ambient component. However, if after decomposition the residual ambient component still contains some dominant directional components, dimensionality reduction will cause errors that are significantly perceptible during post-decomposition rendering. A common example of an HOA representation where that assumption is not met is a general plane wave encoded at an order lower than N. Such general plane waves of an order lower than N can arise as a result of artistic creations that make the sound source seem to have a wide extent, or can occur in connection with the recording of an HOA sound field representation using a spherical microphone. In both examples, the sound field is represented by a large number of highly correlated spatial region signals (see the section on high-order ambisonics spatial resolution for an explanation).
[0011] The problem solved by the present invention is to avoid the above-mentioned disadvantages of other prior arts by eliminating the disadvantages resulting from the process described in European Patent Application Publication No. 2665208. This problem is solved by the methods disclosed in claims 1 and 3. The corresponding devices using these methods are disclosed in claims 2 and 4.
[0012] The present invention improves the HOA sound field representation compression process described in European Patent Application Publication No. 2665208. First, similar to European Patent Application Publication No. 2665208, the HOA representation is analyzed for the presence of a dominant sound source, and its direction is estimated. Using the information on the direction of the dominant sound source, the HOA representation is decomposed into a plurality of dominant directional signals representing general plane waves and a residual component. However, instead of immediately reducing the order of this residual HOA component, the residual HOA component is transformed into a discrete spatial region in order to obtain a general plane wave function in a uniform sampling direction representing the residual HOA component. After that, these plane wave functions are predicted from the dominant directional signals. The reason for performing this process is that the part of the residual HOA component may have a high correlation with the dominant directional signal.
[0013] The prediction can be made simple, generating only a small amount of side information. In the simplest case, the prediction consists of appropriate scaling and delay. Finally, the prediction error is transformed back into the HOA domain and becomes the residual ambient HOA component for which dimensionality reduction is performed.
[0014] Advantageously, the effect of subtracting the signal predictable from the residual HOA component is to reduce its overall order and the remaining amount of the dominant directional signal, and in this way, to reduce the decomposition error resulting from dimensionality reduction.
[0015] In principle, the compression method of the present invention is suitable for compressing a high-order ambisonics representation called HOA for a sound field. This method includes -estimating the direction of the dominant sound source from the current time frame of the HOA coefficients, and - A step of decomposing the HOA representation into a dominant directional signal in the time domain and residual HOA components depending on the HOA coefficient and the dominant sound source direction, wherein in order to obtain a plane wave function in a uniform sampling direction representing the residual HOA components, the residual HOA components are converted into a discrete spatial region, the plane wave function is predicted from the dominant directional signal, parameters describing the prediction are provided, and the corresponding prediction error is converted back into the HOA region, the step of decomposing; - A step of reducing the current order of the residual HOA components to a lower order, resulting in reduced-dimensional residual HOA components, the step of reducing; - A step of performing correlation removal on the reduced-dimensional residual HOA components to obtain a corresponding residual HOA component time domain signal; - A step of perceptually coding the dominant directional signal and the residual HOA component time domain signal so as to supply a compressed dominant directional signal and a compressed residual component signal.
[0016] In principle, the compression device of the present invention is suitable for compressing a higher-order ambisonics representation called HOA for a sound field. This device - Means configured to estimate a dominant sound source direction from a current time frame of HOA coefficients; - Means configured to decompose the HOA representation into a dominant directional signal in the time domain and residual HOA components depending on the HOA coefficient and the dominant sound source direction, wherein in order to obtain a plane wave function in a uniform sampling direction representing the residual HOA components, the residual HOA components are converted into a discrete spatial region, the plane wave function is predicted from the dominant directional signal, parameters describing the prediction are provided, and the corresponding prediction error is converted back into the HOA region; - Means configured to reduce the current order of the residual HOA components to a lower order, resulting in reduced-dimensional residual HOA components being generated; - means configured to obtain the HOA component time-domain signal of the corresponding residual by performing decorrelation on the HOA components of the down-dimensioned residual; - means configured to perform perceptual coding on the dominant directional signal and the HOA component time-domain signal of the residual so as to supply the compressed dominant directional signal and the component signal of the compressed residual.
[0017] In principle, the decompression method of the present invention is suitable for decompressing the high-order ambisonics representation compressed according to the above-described compression method. This method - a step of performing perceptual decoding on the compressed dominant directional signal and the component signal of the compressed residual so as to supply a decompressed dominant directional signal and a decompressed time-domain signal representing the HOA components of the residual in the spatial domain; - a step of re-correlating the decompressed time-domain signal to obtain the HOA components of the corresponding down-dimensioned residual; - a step of expanding the order of the HOA components of the down-dimensioned residual to the original order, the expanding step supplying the corresponding decompressed HOA components of the residual; - a step of synthesizing a corresponding decompressed and resynthesized frame of HOA coefficients using the decompressed dominant directional signal, the decompressed HOA components of the residual of the original order, the estimated dominant sound source direction, and the parameters describing the prediction.
[0018] In principle, the decompression apparatus of the present invention is suitable for decompressing the high-order ambisonics representation compressed according to the above-described compression method. This apparatus - means configured to perform perceptual decoding on the compressed dominant directional signal and the component signal of the compressed residual so as to supply a decompressed dominant directional signal and a decompressed time-domain signal representing the HOA components of the residual in the spatial domain; - Means configured to re-correlate the decompressed time-domain signal, the means for obtaining the HOA components of the corresponding low-dimensionalized residual, - Means configured to extend the order of the HOA components of the low-dimensionalized residual to the original order, the means for supplying the HOA components of the corresponding decompressed residual, - By using the decompressed dominant directional signal, the HOA components of the decompressed residual of the original order, the estimated dominant sound source direction, and the parameters describing the prediction, means configured to synthesize the corresponding decompressed and resynthesized frames of the HOA coefficients.
[0019] Advantageous additional embodiments of the present invention are disclosed in each of the dependent claims.
[0020] Exemplary embodiments of the present invention are described with reference to the accompanying drawings.
Brief Description of the Drawings
[0021]
Figure 1a
Figure 1b
Figure 2a
Figure 2b
Figure 3
Figure 4
Figure 5
Figure 6
Best Mode for Carrying Out the Invention
[0022] Compression Processing The compression processing according to the present invention includes two consecutive steps which are the steps illustrated in each of FIGS. 1a and 1b. The exact definition of each signal is described in the section of the detailed description of HOA decomposition and resynthesis. Frame unit processing for compression using non-overlapping input frames D(k) of an HOA coefficient sequence of length B is used. Here, k represents the frame index. The frame is defined with respect to the HOA coefficient sequence specified by the following equation (1).
Equation
[0023] In FIG. 1a, the frame D(k) of the HOA coefficient sequence is input to the dominant sound source direction estimation step or stage 11, and in this step 11, the HOA representation is analyzed for the presence of the dominant directional signal, and its direction is estimated. The estimation of the direction can be performed, for example, by the process described in European Patent Application Publication No. 2665208. The estimated direction is
Equation
Equation
Equation
[0024] Implicitly, it is assumed that the direction estimates are properly ordered by assigning them to the direction estimates from the previous frame. Thus, the temporal sequence of the individual direction estimates is assumed to describe the direction trajectory of the dominant sound source. In particular, if the d-th dominant sound source is assumed to be inactive,
Number
Number
[0025] In Figure 1b, the perceptual coding of the directional signal X DIR (k - 1) and the perceptual coding of the residual ambient HOA component D A (k - 2) are shown. The directional signal X DIR (k - 1) is a conventional time-domain signal, and this signal can be individually compressed using any existing perceptual compression technique. The compression of the ambient HOA region component D A (k - 2) can be performed in two consecutive steps or stages. In the dimensionality reduction step or stage 13, the reduction of the ambisonics order N RED is performed. Here, for example, N RED = 1. As a result, the ambient HOA component D A,RED (k - 2) is obtained. Such dimensionality reduction is in D A (k - 2), N REDThis is done by retaining only the HOA coefficients and discarding the other coefficients. On the decoder side, corresponding zero values are added to the omitted values, as will be described below.
[0026] Note that, compared to the approach of European Patent Application Publication No. 2665208, the reduced order N RED may generally be selected to be smaller. This is because the overall order, and further, the remaining amount of the directional residual of the ambient HOA components of the residual is smaller. Therefore, due to the dimensionality reduction, the error becomes smaller compared to the case of European Patent Application Publication No. 2665208.
[0027] In the following correlation removal step or stage 14, the dimensionality-reduced ambient HOA component D A,RED (k - 2) is represented by a sequence of HOA coefficients that is correlation removed, and the time-domain signal W A,RED (k - 2) is obtained. This time-domain signal is input to a (bank of) parallel perceptual encoders or compressors 15 that operate by any perceptual compression technique. This correlation removal is performed to avoid masking of the perceptual coding noise when rendering the HOA representation after decompression (see European Patent Application No. 12305860 for an explanation). Approximate correlation removal can be achieved by applying a spherical harmonic transform to D A,RED (k - 2) to convert it to an equivalent signal in the spatial domain within Ο RED as described in European Patent Application Publication No. 2469742.
[0028] Alternatively, the adaptive spherical harmonic transform proposed in European Patent Application No. 12305861 can be used. Here, the grid in the sampling direction is rotated to obtain the maximum correlation removal effect. Another alternative correlation removal technique is the Karhunen - Loeve transform (KLT) described in European Patent Application No. 12305860. Note that for these last two types of correlation removal, some side information represented by α(k - 2) is supplied to enable the inverse processing of the correlation removal at the HOA decompression stage.
[0029] In one embodiment, to improve the coding efficiency, perceptual compression of all time-domain signals X DIR (k - 1) and W A,RED (k - 2) is performed together.
[0030] The output of the perceptual coding is the compressed directional signal
Number
Number
[0031] The decompression process The decompression process is shown in FIGS. 2a and 2b. Similar to the compression process, the decompression process consists of two consecutive steps. In FIG. 2a, the directional signal
Number
Number
Number
Number
Number
Number
Number
[0032] In FIG. 2b, all HOA representations are decoded dominant directional signals
Number
Number
Number
Number
[0033] To improve the coding efficiency, all time-domain signals XDIR (k - 1) and W A,RED When the perceptual compression of (k - 2) is performed together, the compressed directional signal
Number
Number
[0034] A detailed description of the resynthesis exists in the item of HOA resynthesis.
[0035] HOA decomposition A block diagram illustrating the process executed for HOA decomposition is given in FIG. 3. This process is summarized as follows. First, the smoothed dominant directional signal X DIR (k - 1) is calculated and output for perceptual compression. Next, the residual between the HOA representation D DIR (k - 1) of the dominant directional signal and the initial HOA representation D(k - 1) is represented by the "Ο" number of directional signals
Number
Number
[0036] Before describing in detail, it is described that the change in direction between consecutive frames may cause discontinuities in all the calculated signals during synthesis. Therefore, first, the instantaneous estimated value of the signal of each overlapping frame having a length of 2B is calculated. Second, the results of consecutive overlapping frames are smoothed using an appropriate window function. However, each smoothing involves a waiting time for one frame.
[0037] Calculation of the instantaneous dominant direction signal For the current frame D(k) of the HOA coefficient sequence
Equation
[0038] Furthermore, without loss of generality, according to the following formula, the tilt angle θ DOM,d (k) ∈ [0, π] and the azimuth angle φ DOM,d (k) ∈ [0, 2π] (see the content shown in Figure 5). The estimated value of each direction of the active dominant sound source is
Equation
Equation
[0039] First, the mode matrix based on the estimated value of the active sound source direction is calculated according to the following formula, [Number] Here, [Number] In Equation (4), D ACT (k) represents the number of active directions for the k-th frame, and d ACT,j (k), where 1 ≤ j ≤ D ACT (k) indicates their subscripts. Also, [Number] represents a real-valued spherical harmonic function, which is defined in the item of the definition of real-valued spherical harmonic functions.
[0040] Second, the matrix [Number] is calculated according to the following equation, which includes the instantaneous estimated values of all dominant directional signals for the (k - 1)-th and k-th frames. [Number] Here, [Number] This calculation can be performed in two steps. In the first step, the directional signal samples in the columns corresponding to the non-active directions are set to zero, that is, it becomes as follows. [Number] Here, M ACT (k) is a set of active directions. In the second step, the directional signal samples corresponding to the active directions can first be obtained by arranging them into a matrix according to the following. [Number] This matrix is then calculated to minimize the Euclidean norm of the following error.
Number
Number
[0041] Temporal smoothing Regarding step or stage 31, smoothing is only described for the directional signal
Number
Number
Number
Number
Number
Number
[0042] Calculation of the HOA representation of the smoothed dominant directional signal X DIR (k - 1) and [Number] From, in step or stage 32, depending on the continuous signal x DIR,d (l), to mimic the process similar to that performed for HOA synthesis, the HOA representation of the smoothed dominant directional signal is calculated. Since the change in the direction estimation value between consecutive frames may cause discontinuity, the instantaneous HOA representation of overlapping frames of length 2B is calculated again, and the results of consecutive overlapping frames are smoothed by using an appropriate window function. Thus, the HOA representation D DIR (k - 1) is obtained by the following equation. [Number] Here, [Number] Furthermore, [Number]
[0043] Representing the residual HOA representation by a directional signal on a uniform grid D DIR (k - 1) and D(k - 1) (i.e., D(k) delayed by frame delay 381), the residual HOA representation by a directional signal on a uniform grid is calculated at step or stage 33. The purpose of this process is to obtain a directional signal (i.e., a general plane wave function) arriving from some fixed, approximately uniformly distributed direction (also referred to as the grid direction) to represent the residual [D(k - 2)D(k - 1)] - [D DIR (k - 2)D DIR (k - 1)].
Number
[0044] First, with respect to the grid direction, the mode matrix Ξ GRID is calculated as follows.
Number
Number
[0045] The directional signal on each grid is obtained by the following formula.
Number
[0046] Prediction of the directional signal on a uniform grid from the dominant directional signal
Number
Number
Number
Number
[0047] First,[[]]
Number
Number
Number
Number
Number
[0048] Next, each grid signal
Number
Number
Number
Number
Number
[0049] It is also possible to make other types of predictions. For example, instead of calculating the scaling coefficient for the entire band, it is also reasonable to obtain the scaling coefficient for the frequency band oriented to perception. However, in this process, although the prediction is improved, the amount of side information increases.
[0050] All prediction parameters can be arranged in a parameter matrix as follows.
[0051]
Number
Number
Num
Num
[0052] Calculation of the HOA representation of the predicted directional signal on a uniform grid The HOA representation of the predicted grid signal is calculated according to the following formula in step or stage 35
Num
Num
[0053] Calculation of the HOA representation of the residual ambient sound field component
Num
Num
Num
[0054] HOA resynthesis Before explaining the processing of the individual steps or stages in FIG. 4 in detail, an overview will be described. For the uniformly distributed direction, the directional signal
Number
Number
Number
Number
Number
Number
Number
[0055] Calculation of the HOA representation of the dominant directional signal
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0056] Prediction of the directional signal on a uniform grid from the dominant directional signal
Number
Number
Number
Number
Number
[0057] Calculation of the HOA representation of the predicted directional signal on a uniform grid In step or stage 44 of calculating the HOA representation of the predicted directional signal on a uniform grid, the HOA representation of the predicted grid directional signal is obtained by the following equation.
Number
[0058] Synthesis of the HOA sound field representation
Number
Number
Number
Number
Number
Number
[0059] Fundamentals of Higher-Order Ambisonics Higher-order ambisonics is based on the description of the sound field within a compact region of interest, assuming no sound sources are present. In that case, the spatio-temporal behavior of the sound pressure p(t,x) at time t and position x within the region of interest is fully physically determined by the wave equation for a homogeneous medium. The following content is based on the spherical coordinate system shown in Fig. 5. The x-axis points to the front, the y-axis points to the left, and the z-axis points upward. The position x = (r, θ, φ) in space T is represented by a radius r > 0 (i.e., the distance from the coordinate origin), an inclination angle θ ∈ [0, π] measured from the polar axis z, and an azimuth angle φ ∈ [0, 2π] measured counterclockwise in the x - y plane from the x-axis. (·) T represents the transpose.
[0060] F t the Fourier transform of the sound pressure with respect to the time represented by (·), i.e.,
Number
Number
Number
Number
[0061] When the sound field is represented by the superposition of an infinite number of harmonic plane waves ω with different angular frequencies and arrives from all possible directions specified by the angular pair (θ, φ), it can be seen that each plane-wave complex amplitude function D(ω, θ, φ) can be represented by the following spherical harmonic expansion (see "Plane-wave Decomposition of the Sound Field on a Sphere by Spherical Convolution" by B. Rafaely, Journal of the Acoustical Society of America 4(116), pp. 2149 - 2157, 2004). [Number] Here, the expansion coefficient [Number] is [Number] related by the following equation. [Number]
[0062] Each coefficient [Number] Assuming that is a function of the angular frequency ω, applying the inverse Fourier transform ( [Number] represented by ) gives the following time-domain functions for each order n and degree m.
Number
Number
Number
[0063] The final ambisonics format results in the following sampled version of d(t) using the sampling frequency f s Here, T
Number
Number
[0064] Definition of real-valued spherical harmonic functions Real-valued spherical harmonic functions
Number
Number
Number
Number
[0065] Spatial resolution of high-order ambisonics Direction Ω0 = (θ0, φ0) T A general plane wave function x(t) arriving from is expressed in HOA by the following equation.
Number
Number
Number
Number
Number
[0066] Discrete spatial region The spatial density of the plane wave amplitude is zero for the spatial direction Ω o (When discretized with 1 ≤ ο ≤ Ο, the spatial directions Ω o are approximately uniformly distributed on the unit sphere, and Ο directional signals d(t, Ω o ) are obtained. When these signals are grouped into a vector, it is represented by the following formula, [Number] It can be verified that using Eq. (47), this vector can be calculated from the continuous ambisonics representation d(t) defined in Eq. (41) by a simple matrix multiplication as follows. d SPAT (t) = Ψ H d(t) (52) Here, (·) H denotes the complex conjugate transpose, and Ψ represents the mode matrix defined by the following formula. [Number] Here, [Number] Direction Ω o is approximately uniformly distributed on the unit sphere. Generally, the mode matrix is invertible. Therefore, the continuous ambisonics representation can be calculated from the directional signals d(t, Ω o ) by the following formula. d(t) = Ψ -H d SPAT (t) (55) The equations of both sides constitute the conversion and inverse conversion between the ambisonic representation and the spatial domain. In the present application, these conversions are called the spherical harmonic function conversion and the inverse spherical harmonic function conversion.
[0067] Direction Ω o is distributed almost uniformly on the unit sphere, so
Number
[0068] On both the encoding side and the decoding side, the processing of the present invention can be executed by a single processor or electronic circuit, or by several processors or electronic circuits operating in parallel and / or operating on a plurality of different parts of the processing of the present invention.
[0069] The present invention can be applied to the processing corresponding to an audio signal that can be rendered and reproduced on a loudspeaker configuration in a home environment or on a loudspeaker configuration in a theater.
[0070] Some aspects will be described. 〔Aspect 1〕 A method for compressing a higher-order ambisonic representation called HOA for a sound field, comprising: - estimating (step 11) the dominant sound source direction (
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Claims
1. A method for decompressing a compressed higher-order ambisonics (HOA) representation, the method comprising: Perceptually decoding the compressed HOA representation to determine a decompressed time-domain signal representing a decompressed dominant directional signal and HOA components of a residual within a spatial region, the decompressed time-domain signal corresponding to downsampled residual HOA components; Determining a predicted directional signal based on the decompressed dominant directional signal, the predicted directional signal being determined based on smoothing using a window function; Determining decompressed HOA components of the residual based on the decompressed time-domain signal, the decompressed HOA components being based on expanding the order of the downsampled residual HOA components, the expanding including adding zero values to the downsampled residual HOA components; Determining an HOA sound field representation based on the predicted directional signal and the decompressed HOA components of the residual; Outputting the HOA sound field representation for rendering to a speaker feed and a method.
2. A non-transitory computer-readable medium including instructions that, when executed, cause a computer to perform the steps of the method of claim 1.
3. An apparatus for decompressing a higher-order ambisonics (HOA) representation, the apparatus comprising: A decoder that perceptually decodes the compressed HOA representation to determine a decompressed time-domain signal representing a decompressed dominant directional signal and HOA components of a residual within a spatial region, the decompressed time-domain signal corresponding to downsampled residual HOA components; A first processor that determines a predicted directional signal based on the decompressed dominant directional signal, the first processor being configured to determine the predicted directional signal based on smoothing using a window function; A second processor that determines the HOA components of the decompressed residual based on the decompressed time-domain signal, wherein the decompressed HOA components are based on extending the order of the downsampled residual HOA components, and the extending includes adding zero values to the downsampled residual HOA components, and the second processor; A third processor that determines a HOA sound field representation based on the predicted directional signal and the HOA components of the decompressed residual; A further processor that outputs the HOA sound field representation for rendering to speaker feeds; having; device.
Citation Information
Patent Citations
Data structure for Higher Order Ambisonics audio data
EP2450880A1
Method and apparatus for encoding and decoding successive frames of ambisonics representation of two-dimensional or three-dimensional sound field
JP2012133366A
Data structure for high-order ambisonic audio data
JP2013545391A