Method and apparatus for compressing and decompressing high-order ambisonics representation
The method optimizes HOA compression by adaptively reallocating HOA coefficient sequences between directional and ambient components, addressing inefficiencies in existing methods to reduce bit rates and enhance streaming suitability.
Patent Information
- Application Number
- JP2024101601
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-04-29
- Filing Date
- 2024-06-25
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2034-04-24
AI Technical Summary
Existing high-order ambisonics (HOA) compression methods fail to optimally determine the number of dominant directional signals, leading to inefficient utilization of bandwidth and suboptimal direction-ambient decomposition, resulting in high bit rates that are impractical for streaming applications.
A method and apparatus that dynamically allocate HOA coefficient sequences between directional and ambient components based on perceptual relevance, minimizing perceived error by adaptively determining the number of directional signals and utilizing additional channels for ambient components where necessary, while ensuring smooth transitions and efficient encoding.
This approach optimizes bandwidth usage and reduces perceived errors in HOA compression, making it suitable for practical streaming applications by intelligently reallocating channels for directional and ambient components based on perceptual criteria.
Smart Images

Figure 0007717911000210 
Figure 0007717911000211 
Figure 0007717911000212
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for compressing and decompressing high-order ambisonics representation by separately processing a directional signal component and an ambient signal component.
Background Art
[0002] High-order ambisonics (HOA) provides one possibility for representing three-dimensional sound while other techniques such as wave field synthesis (WFS) and channel-based approaches like 22.2 exist. In contrast to channel-based methods, HOA representations have the advantage of being independent of specific loudspeaker setups. However, to obtain this flexibility, decoding processing is required to reproduce the HOA representation with a specific loudspeaker setup. Usually, HOA can be configured with a very small number of loudspeakers compared to the WFS approach where the number of required loudspeakers is very large. A further advantage of HOA is that the same representation can be utilized without the need to change for binaural rendering to headphones.
[0003] HOA is based on the representation of the spatial density of complex harmonic plane wave amplitudes by means of a cut spherical harmonic (SH) expansion. Each expansion coefficient is a function of the angular frequency, and this can be equivalently represented by a time-domain function. Therefore, without loss of generality, a complete HOA sound field representation can actually be considered to be composed of "Ο" time-domain functions. Here, Ο represents the number of expansion coefficients. These time-domain functions are referred to as the following HOA coefficient sequence or HOA channels as having the same meaning.
[0004] The spatial resolution of the HOA representation improves with an increase in the maximum order N of the expansion. Unfortunately, the number "Ο" of expansion coefficients increases quadratically with respect to the order N, especially Ο = (N + 1) 2This results in, for example, a general HOA representation using an order N = 4 requiring a number of HOA (expansion) coefficients of Ο = 25. Considering the above, the total bit rate for the transmission of the HOA representation is given by the sampling rate f s of the desired single channel and the number of bits per sample N b and is obtained by Ο·f s ·N b Thus, using N b = 16 bits per sample to transmit an HOA representation of order N = 4 at a sampling rate f s = 48 kHz results in a bit rate of 19.2 megabits per second, which is extremely high for many practical applications, such as streaming.
[0005] The compression of HOA sound field representations has been proposed in European patent applications EP 12306569 and EP 12305537. For example, instead of individually perceptually coding the HOA coefficient series, as is done in "Coding of Higher-Order Ambisonics Using AAC" by E. Hellerud, I. Burnett, A. Solvang, and U. P. Svensson, 124th AES Convention, Amsterdam, 2008, attempts have been made to reduce the number of signals to be perceptually coded by performing a sound field analysis in particular and decomposing a given HOA representation into directional components and residual ambient components. Generally, the directional components are assumed to be represented by a few dominant directional signals that can be regarded as general plane wave functions. The order of the residual ambient HOA components is reduced. The reason is that after extracting the dominant directional signals, it is considered that the lower-order HOA coefficients retain the most relevant information. SUMMARY OF THE INVENTION
[0006] In summary, by performing such a process, the initial number (N + 1) of the HOA coefficient series to be perceptually coded 2 is a predetermined number of D dominant directional signals and the truncation order N RED<N is used to represent the ambient HOA component of the residual (N RED +1) 2 The number of HOA coefficient sequences is reduced to. Thereby, the number of signals to be encoded is determined, that is, D+(N RED +1) 2 becomes. In particular, this number is independent of the actually detected number D ACT of the active dominant directional sound sources in time frame k. This means that when the actually detected number D ACT (k) of the active dominant directional sound sources in time frame k is less than the maximum allowable number D of directional signals, some or even all of the perceptually encoded dominant directional signals will become zero. That is, this means that these multiple channels are not used at all to capture the relevant information of the sound field.
[0007] In this situation, another assumed weakness of the processing in European Patent Application No. 12306569 and European Patent Application No. 12305537 is the criterion for determining the number of dominant directional signals within each time frame. The reason is that no attempt has been made to determine the optimal number of active dominant directional signals for the continuous perceptual coding of the sound field. For example, in European Patent Application No. 12305537, the number of dominant sound sources is estimated using a simple power criterion, that is, by obtaining the dimension of the subspace of the correlation matrix of the coefficients belonging to the largest eigenvalue. In European Patent Application No. 12306569, an incremental detection of dominant directional sound sources has been proposed. Here, when the power of the plane wave function from each direction is sufficiently high with respect to the first directional signal, the directional sound source is considered to be dominant. Using a power-based criterion as in the cases of European Patent Application No. 12306569 and European Patent Application No. 12305537 may result in a direction-ambient decomposition that is not optimal for the perceptual coding of the sound field.
[0008] The problem to be solved by the present invention is to improve HOA compression by determining how to allocate coefficients for the directional signal and the ambient HOA components to a predetermined reduced number of channels for current HOA audio signal content. This problem is solved by the respective methods disclosed in claims 1 and 2. An apparatus using these methods is disclosed in claim 4.
[0009] The present invention improves the compression process proposed in European patent application No. 12306569 in two aspects. First, the bandwidth provided by a given number of channels to be perceptually encoded is well utilized. In time frames where the dominant sound source signal is not detected, the channels initially reserved for the dominant directional signal are used in the form of an additional HOA coefficient sequence of the residual ambient HOA component to capture additional information about the ambient component. Second, with the aim of using a given number of channels to perceptually encode a given HOA sound field representation in mind, the criterion for determining the number of directional signals extracted from the HOA representation is adapted to that purpose. The number of directional signals is determined such that the error perceived by the decoded and reconstructed HOA representation is minimized. The criterion compares the modeling error resulting from extracting the directional signals and using fewer HOA coefficient sequences to describe the residual ambient HOA component, with the modeling error resulting from not extracting the directional signals and instead using additional HOA coefficient sequences to describe the residual ambient HOA component. The criterion further takes into account, for both cases, the spatial power distribution of the quantization noise resulting from the perceptual encoding of the HOA coefficient sequences of the directional signals and the residual ambient HOA component.
[0010] To carry out the above-described process, the total number I of signals (channels) is determined before starting HOA compression. This total number I is reduced compared to the number of the initial Ο HOA coefficient sequences. The ambient HOA component has a minimum number of Ο REDIt is assumed to be represented by a set of HOA coefficient sequences. In some cases, the minimum number may be zero. The remaining D = I - Ο RED The RED channels are assumed to include either the directional signal or an additional set of coefficients of the ambient HOA component, depending on what is more perceptually meaningful as determined by the directional signal extraction process. The assignment to the remaining D channels of either the directional signal or the ambient HOA component coefficient set is assumed to be changeable on a frame-by-frame basis. For the reconstruction of the sound field on the receiver side, information about this assignment is transmitted as additional side information.
[0011] In principle, the compression method of the present invention is suitable for compressing a higher-order ambisonics representation of a sound field called HOA using a predetermined number of perceptual encoding processes with the input time frames of the HOA coefficient sequences. This method is performed on a frame-by-frame basis, - For the current frame, estimating a set of dominant directions and a data set of indices of the corresponding detected directional signals; - Decomposing the HOA coefficient sequence of the current frame, using each of the non-predetermined number of directional signals, each direction included in the set of dominant direction estimates, and each data set of indices of the directional signals, where the non-predetermined number is smaller than the predetermined number, into the non-predetermined number of directional signals, a residual ambient HOA component represented by a reduced number of HOA coefficient sequences corresponding to the difference between the predetermined number and the non-predetermined number, and a data set of indices of the corresponding reduced number of residual ambient HOA coefficient sequences; the step of decomposing; - Assigning the HOA coefficient sequences of the directional signals and the residual ambient HOA components to a number of channels corresponding to the predetermined number, where the data set of indices of the directional signals and the data set of indices of the reduced number of residual ambient HOA coefficient sequences are used for the assignment; the step of assigning; - a step of perceptually encoding the channels of the relevant frame, wherein a coded and compressed frame is obtained, the step of perceptually encoding; and the method includes the step.
[0012] In principle, the compression device of the present invention is suitable for compressing a high-order ambisonics representation called HOA of a sound field using a predetermined number of perceptual encoding processes with a time frame in which an HOA coefficient sequence is input. The device executes processing in units of frames, - means configured to estimate, for the current frame, a set of dominant directions and a data set of corresponding detected directivity signal indices; - means configured to decompose the HOA coefficient sequence of the current frame, wherein the non-predetermined number of directivity signals, each direction included in the set of dominant direction estimates, and each data set of the indices of the directivity signals are used, the non-predetermined number being smaller than the predetermined number, the non-predetermined number of directivity signals, and an ambient HOA component of a residual represented by a reduced number of HOA coefficient sequences corresponding to the difference between the predetermined number and the non-predetermined number, and a corresponding data set of indices of the corresponding reduced number of residual ambient HOA coefficient sequences; the means configured to decompose; - means configured to assign the HOA coefficient sequences of the directivity signals and the ambient HOA components of the residuals to a number of channels corresponding to the predetermined number, wherein, for the assignment, the data set of the indices of the directivity signals and the data set of the indices of the reduced number of residual ambient HOA coefficient sequences are used; the means; - means configured to perceptually encode the channels of the relevant frame, wherein a coded and compressed frame is obtained; the means; and the method includes the means.
[0013] In principle, the decompression method of the present invention is suitable for decompressing a high-order ambisonics representation compressed according to the above-described compression method. This decompression method is, - Decoding the current encoded and compressed frame to obtain the perceptually decoded frame of the channel; - Redistributing the perceptually decoded frame of the channel to reform the corresponding frame of the directional signal and the corresponding frame of the residual ambient HOA components using the data set of the indices of the detected directional signals and the data set of the indices of the selected ambient HOA coefficient sequences; - Resynthesizing the current decompressed frame of the HOA representation from the frame of the directional signal and the frame of the residual ambient HOA components using the data set of the indices of the detected directional signals and the set of dominant direction estimates, The directional signal for uniformly distributed directions is predicted from the directional signal, and then the current decompressed frame is resynthesized from the frame of the directional signal, the predicted signal, and the residual ambient HOA components.
[0014] In principle, the decompression device of the present invention is suitable for decompressing the high-order ambisonics representation compressed according to the above-described compression method. This device - Means configured to decode the current encoded and compressed frame to obtain the perceptually decoded frame of the channel; - Means configured to redistribute the perceptually decoded frame of the channel to reform the corresponding frame of the directional signal and the corresponding frame of the residual ambient HOA components using the data set of the indices of the detected directional signals and the data set of the indices of the selected ambient HOA coefficient sequences; - Means configured to resynthesize the current decompressed frame of the HOA representation from the frame of the directional signal and the frame of the residual ambient HOA components using the data set of the indices of the detected directional signals and the set of dominant direction estimates, A directional signal for a uniformly distributed direction is predicted from the directional signal, and then the current decompressed frame is resynthesized from the frame of the directional signal, the predicted signal, and the ambient HOA components of the residual.
[0015] Additional embodiments of the present invention are disclosed in each of the dependent claims and are advantageous.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0017] Exemplary embodiments of the present invention are described with reference to the accompanying drawings. A. Improved HOA Compression The compression process according to the present invention is based on European Patent Application No. 12306569 and is shown in FIG. 1. Here, the signal processing blocks have been modified or newly introduced with respect to European Patent Application No. 12306569, and the signal processing blocks are shown in bold boxes. In this application, "
Number
[0018] For HOA compression, frame-by-frame processing using non-overlapping input frames C(k) of an HOA coefficient sequence of length L is used. Here, k represents the frame index. A frame is defined with respect to the HOA coefficient sequence specified by the following equation (1).
Equation
[0019] Step or stage 11 / 12 in FIG. 1 is optionally performed, and the non-overlapping k-th and (k - 1)-th frames of the HOA coefficient sequence are concatenated according to the following equation to form a long frame
Equation
Equation
Equation
[0020] In principle, the dominant sound source estimation step or stage 13 is performed as proposed in European patent application No. 13305156, but with important modifications. This modification relates to the determination of the number of detected directions, i.e., how many directional signals are to be extracted from the HOA representation. This is achieved by the idea of extracting directional signals instead of using additional HOA coefficient sequences only when it is perceptually more relevant to extract directional signals than to use additional HOA coefficient sequences for a good approximate calculation of the ambient HOA components. A detailed description of this technique will be given in item A.2.
[0021] By estimating the dominant sound source, a dataset of the indices of the detected directional signals
Number
Number
[0022] In step or stage 14, the current (long) frame of the HOA coefficient sequence
Number
Number
Number
Number
Number
[0023] In particular, the following three cases should be distinguished.
[0024] a) N DIR,ACT (k - 2)=N DIR,ACT(k - 3): In this case, it is assumed that the same set of HOA coefficients is selected as in the case of frame k - 3.
[0025] b) N DIR,ACT (k - 2) < N DIR,ACT (k - 3): In this case, to represent the ambient HOA components within the current frame, more sets of HOA coefficients than those in the previous frame k - 3 can be used. The set of HOA coefficients selected at k - 3 is assumed to be selected within the current frame as well. Additional sets of HOA coefficients can be selected according to different criteria. For example, select the set of HOA coefficients with the highest average power as C AMB within (k - 2), or select the set of HOA coefficients regarding their respective perceptual importance.
[0026] c) N DIR,ACT (k - 2) > N DIR,ACT (k - 3): In this case, to represent the ambient HOA components within the current frame, fewer sets of HOA coefficients than those present in the last frame k - 3 can be used. The problem to be solved here is which of the already selected sets of HOA coefficients should be deactivated. A reasonable solution is to deactivate the set of HOA coefficients assigned to the channel at step or stage 16 of assigning the signal in frame k - 3
Number
[0027] To avoid discontinuities at the frame boundaries when additional sets of HOA coefficients are activated or deactivated, each signal may be smoothly faded - in or faded - out.
[0028] Ο RED +N DIR,ACT (k - 2) The reduced number of final ambient HOA representations is indicated by C AMB,RED (k - 2). The indices of the selected ambient coefficient sets are in the dataset [Number] is output therein.
[0029] In step / stage 16, X DIR the active directional signals included in (k - 2) and C AMB,RED the HOA coefficient sequences included in (k - 2) are assigned to the frame Y(k - 2) of I channels for individual perceptual encoding. To describe the signal assignment in more detail, the frame X DIR (k - 2), Y(k - 2) and C AMD,RED (k - 2) are assumed to be composed of the individual signals x DIR,d (k - 2) (d ∈ {1, …, D}), y i (k - 2) (i ∈ {1, …, D}) and c AMB, RED, ο (k - 2) (ο = 1, …, Ο). [Number]
[0030] To obtain continuous signals for continuous perceptual encoding, the active directional signals are assigned to hold the index of each channel. This can be expressed as in the following formula. [Number]
[0031] The HOA coefficient sequences of the ambient components are assigned such that the minimum number of Ο RED coefficient sequences are always included in the last Ο RED signals of Y(k - 2), that is, are assigned according to the following formula. [Number]
[0032] Additional D - N DIR,ACTFor the HOA coefficient sequences of the (k-2) ambient components, it should be distinguished whether they were also selected in the previous frame. a) Additional D-N DIR,ACT If the HOA coefficient sequences of the (k-2) ambient components were selected within the previous frame as being transmitted, i.e., each index was also within the dataset
Number
Number
Number
Number
[0033] This particular assignment provides the advantage that during the HOA decompression process, the rearrangement and synthesis of the signals can be performed without information on which ambient HOA coefficient sequences are included in which channels of the Y(k-2). Instead, the datasets
Number
Number
[0034] This assignment process advantageously results in the assignment vector
Number
Number
[0035] For frames in which the vector γ(k) is not transmitted in step / stage 16, on the decompression side, the data parameter set
Number
Number
[0036] A.1 Estimation of the Dominant Sound Source Direction The estimation step / stage 13 for the dominant sound source direction in FIG. 1 is depicted in more detail in FIG. 2. This is essentially carried out according to the content described in European Patent Application No. 13305156, but there is a decisive difference. The decisive difference is the method of determining the number of dominant sound sources. The number of dominant sound sources corresponds to the number of directional signals extracted from a given HOA representation. This number is important because it is used to control whether a given HOA representation is better represented either by using more directional signals or, alternatively, by using more HOA coefficient columns to better model the ambient HOA components.
[0037] The estimation of the dominant sound source direction starts at step or stage 21 in a preliminary search for the dominant sound source direction using a long frame of the input HOA coefficient columns.
Number
Number
Number
Number
[0038] At step or stage 22, the preliminary direction estimate values, the directional signals, and the HOA sound field components are the frame of the HOA coefficient columns input to determine the number
Number
Number
Number
Number
Number
Number
Number
Number
[0039] In step or stage 23, the resulting direction trajectory is smoothed according to the sound source movement model, and it is determined which sound source is considered to be active (see European Patent Application No. 13305156). By this last process, a set of indices of the active directional sound sources
Number
Number
[0040] A.2 Determination of the number of directional signals to be extracted In step / stage 22, a situation is assumed where there are a given total number of I channels that are used to capture the perceptually most relevant sound field information in order to determine the number of directional signals. Therefore, considering the issue of whether the current HOA representation is better represented either by using more directional signals or by using more HOA coefficient sequences for better modeling of the ambient HOA components for the overall HOA compression / decompression quality, the number of directional signals to be extracted is determined. In order to derive the criteria for determining the number of directional sound sources to be extracted in step / stage 22, it is considered which criteria are relevant to human perception, and that HOA compression is performed, in particular, by the following two processes. - Reduction of the HOA coefficient sequence for representing the ambient HOA components (this means reduction of the number of relevant channels) - Perceptual coding of the HOA coefficient sequences for representing the directional signals and the ambient HOA components
[0041] Depending on the number M (0 ≦ M ≦ D) of the extracted directional signals, by the first process, an approximate calculation is performed according to the following formula.
Number
Number
Number
Number
Number
[0042] The approximate calculation from the second process can be expressed by the following formula.
Number
Number
Number
Number
[0043] Formation of a reference The number of directional signals to be extracted
Number
Number
Number
Number
Number
[0044] Subtract "1", and the process of finding the continuous maximum value is performed so that the perceptual level surely becomes zero as long as the error power is less than the masking threshold. Finally, the number of extracted directional signals [Number] It is selected such that the average value for all test directions of the maximum value of the error perception level across all critical bands is minimized, i.e., it is selected according to the following equation.
Number
[0045] Alternatively, in Equation (15), the maximum value of the error perception level can be replaced by an averaging process.
[0046] Calculation of the directional perception masking power distribution Original HOA representation
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0047] Calculation of directional power distribution In the following description, two alternative methods for calculating the directional power distribution
Number
[0048] a. One possibility is to actually calculate the approximation of the desired HOA representation
Number
Number
[0049] b. An alternative solution is [Number] instead of the approximation
Number
Number
Number
Number
Number
Number
Number
Number
[0050] Hereinafter, how to calculate the directional power distributions of the three errors for each Bark scale critical band will be described.
[0051] a. Error
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0052] b. Error
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0053] Rotated direction [Number] For [Number] is defined as the mode matrix, [Number] All scaling parameters in the vector according to [Number] are arranged, and the HOA components [Number] can be described as in the following formula. [Number]
[0054] As a result, the true directional HOA component [Number] and [Number] [Number] The directional signal perceptually decoded by [Number] (d = 1, …, M) is synthesized [Number] (Refer to Equation (23)) is the perceptual coding error represented by the following equation
Number
Number
[0055] Test direction Ω q For (q = 1, …, Q), the error within the spatial region
Number
Number
[0056] Element β of the vector (d) (k) is
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0057] c. Error obtained as a result of perceptual coding of the HOA coefficient sequence of the ambient HOA component
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0058] B. Improved HOA decompression The corresponding HOA decompression process is shown in FIG. 3, and this HOA decompression process includes the following steps or stages.
[0059] In step or stage 31,
Number
Number
[0060] In the signal redistribution stage or stage 32,
Number
Number
Number
Number
Number
[0061] In the synthesis step or stage 33, (according to the process described in FIGS. 2b and 4 of European Patent Application No. 12306569,) the frame of the directional signal
Number
Number
Number
Number
Number
Number
Number
[0062] C. Basis of Higher-Order Ambisonics Higher-order Ambisonics (HOA) is based on the description of the sound field within a compact region of interest, assuming that there are no sound sources present. In that case, the spatio-temporal behavior of the sound pressure p(t,x) at time t and position x within the region of interest is physically fully determined by the wave equation for a homogeneous medium. The following content is based on the spherical coordinate system shown in Figure 4. In the coordinate system being used, the x-axis points to the front position, the y-axis points to the left, and the z-axis points upward. A position x = (r,θ,φ) in space T is represented by a radius r > 0 (i.e., the distance to the coordinate origin), an inclination angle θ ∈ [0,π] measured from the polar axis z, and an azimuth angle φ ∈ [0,2π] measured counterclockwise in the x-y plane from the x-axis. Further, (·) T denotes the transpose.
[0063] F t the Fourier transform of the sound pressure with respect to time represented by (·), i.e.,
Number
Number
Number
Number
[0064] When the sound field is represented by the superposition of an infinite number of harmonic plane waves with different angular frequencies ω and arrives from all possible directions specified by the angular pair (θ, φ), it can be seen that each plane wave complex amplitude function C(ω, θ, φ) can be represented by the following spherical harmonic expansion (see "Plane-wave Decomposition of the Sound Field on a Sphere by Spherical Convolution" by B. Rafaely, Journal of the Acoustical Society of America 4(116), pp. 2149-2157, 2004).
Number
Number
Number
Number
[0065] Each coefficient
Number
Number
Number
Number
Number
[0066] The final ambisonics format results in a sampled version of the following c(t) using the sampling frequency f s .
Number
Number
[0067] C.1 Definition of Real-Valued Spherical Harmonic Functions Real-valued spherical harmonic functions
Number
Number
Number
Number
[0068] C.2 Spatial resolution of higher-order ambisonics The general plane wave function x(t) arriving from the direction Ω0=(θ0,φ0) T is expressed in HOA by the following formula.
Number
Number
Number
Number
[0069] As understood from Equation (51), this is the product of the general plane wave function x(t) and the spatial dispersion function ν N (Θ), and the spatial dispersion function ν N (Θ) is shown to depend only on the angle Θ between Ω and Ω0 with the characteristics of the following formula.
Number
Number
[0070] However, for finite dimensions N, the contribution of a general plane wave from direction Ω bleeds into neighboring directions, and the degree of this bleed decreases with increasing order. The normalized function ν for several different values of N N A plot of (Θ) is shown in FIG.
[0071] It is pointed out that the time domain behavior of the spatial density of the plane wave amplitude in any direction Ω is a multiple of the time domain behavior of the spatial density of the plane wave amplitude in any other direction. In particular, the functions c(t,Ω1) and c(t,Ω2) for some given directions Ω1 and Ω2 with respect to time t are highly correlated.
[0072] C.3 Spherical Harmonic Transform The spatial density of plane wave amplitudes is O spatial directions Ω o When discretized in (1≦ο≦O), the spatial direction Ω o are distributed almost uniformly on the unit sphere, but O directional signals c(t,Ω o ) is obtained. These signals are collected into a vector, which is expressed as follows:
number
number
number
[0073] Direction Ω o is almost uniformly distributed on the unit sphere. Therefore, generally, the mode matrix is invertible. Thus, the continuous ambisonic representation can be calculated from the directional signal c(t, Ω o ) by the following equation.
Equation
[0074] Both equations constitute the conversion and inverse conversion between the ambisonic representation and the spatial domain. In the present application, these conversions are called spherical harmonic function conversion and inverse spherical harmonic function conversion.
[0075] Note that since the direction Ω o is almost uniformly distributed on the unit sphere, approximate calculation
Equation
[0076] It is advantageous that all of the above-described relationships are also valid in the discrete time domain.
[0077] The processing of the present invention can be executed by a single processor or electronic circuit, or a plurality of processors or electronic circuits operating in parallel, and / or a plurality of processors or electronic circuits operating on a plurality of different parts of the processing of the present invention.
Claims
1. A method for decompressing a compressed high-order ambisonics (HOA) representation, comprising: receiving the compressed HOA representation and side information corresponding to the compressed HOA representation; decoding the compressed HOA representation to determine a decoded frame of the signal; determining, from the side information, an assignment vector indicating a first index of a sequence of coefficients that may include ambient HOA components related to non-zero ambient HOA components; re-distributing the decoded frame of the signal based on the assignment vector, the re-distribution determining a frame of ambient HOA components; re-synthesizing a currently decompressed frame of the compressed HOA representation from the frame of the ambient HOA components and a method including the above.
2. The frame of the ambient HOA components is generated based on the assignment vector, The method according to claim 1.
3. A program that causes a processor to execute the method according to claim 1 or 2 when executed by the processor.
4. A non-transitory computer-readable storage medium storing the program according to claim 3.
5. An apparatus for decompressing a compressed high-order ambisonics (HOA) representation, comprising: a receiving unit that receives the compressed HOA representation and side information corresponding to the compressed HOA representation; a decoding unit that decodes the compressed HOA representation to determine a decoded frame of the signal; a first processing unit that determines, from the side information, an assignment vector indicating a first index of a sequence of coefficients that may include ambient HOA components related to non-zero ambient HOA components; a second processing unit that re-distributes the decoded frame of the signal based on the assignment vector, the re-distribution determining a frame of ambient HOA components; a third processing unit that re-synthesizes a currently decompressed frame of the compressed HOA representation from the frame of the ambient HOA components and an apparatus having the above.
Citation Information
Patent Citations
Data structure for Higher Order Ambisonics audio data
EP2450880A1
Method and apparatus for encoding and decoding successive frames of ambisonics representation of two-dimensional or three-dimensional sound field
JP2012133366A
Method and apparatus for encoding and optimally reproducing a three-dimensional sound field
JP2012514358A