Method and apparatus for compressing and decompressing higher order ambisonics representation
The method optimizes HOA compression by dynamically allocating HOA coefficient sequences to directional and ambient components based on perceptual relevance, addressing high bit rate issues and optimizing channel utilization for efficient sound field reconstruction.
Patent Information
- Application Number
- JP2025123005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-04-29
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2034-04-24
AI Technical Summary
Higher-Order Ambisonics (HOA) representations require high bit rates for transmission due to the quadratic growth of expansion coefficients with order N, making practical applications like streaming challenging, and existing compression methods fail to optimally determine the number of dominant directional signals for efficient perceptual coding.
A method and apparatus that dynamically allocate HOA coefficient sequences to directional and ambient components based on perceptual relevance, using a criterion that minimizes perceived error by adjusting the number of channels for directional signals and ambient components, and transmitting side information for reconstruction.
Improves HOA compression by optimizing channel utilization and reducing the number of signals to encode, resulting in more efficient bandwidth use and lower perceived errors in reconstructed sound fields.
Smart Images

Figure 2025157488000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for compressing and decompressing higher-order Ambisonics representations by separately processing the directional and ambient signal components. [Background technology]
[0002] Higher-Order Ambisonics (HOA) offers one possibility for representing three-dimensional sound, while other techniques exist, such as wave field synthesis (WFS) and channel-based approaches like 22.2. In contrast to channel-based approaches, HOA representations have the advantage of being independent of a specific loudspeaker configuration. However, this flexibility requires a decoding process to reproduce the HOA representation in a specific loudspeaker configuration. Compared to WFS approaches, which typically require a much larger number of loudspeakers, HOA allows for configurations with only a very small number of loudspeakers. An additional advantage of HOA is that the same representation can be used, without modification, for binaural rendering to headphones.
[0003] HOA is based on the representation of the spatial density of complex harmonic plane wave amplitudes via a truncated spherical harmonic (SH) expansion. Each expansion coefficient is a function of angular frequency, which can be equivalently represented by a time-domain function. Therefore, without loss of generality, the complete HOA sound field representation can actually be considered to consist of "O" time-domain functions, where O represents the number of expansion coefficients. These time-domain functions are equivalently referred to below as HOA coefficient sequences or HOA channels.
[0004] The spatial resolution of the HOA representation improves with increasing maximum order of the expansion, N. Unfortunately, the number of expansion coefficients, O, grows quadratically with order N, in particular O=(N+1) 2For example, a general HOA representation with order N=4 requires O=25 HOA (expansion) coefficients. Taking the above into account, the total bit rate for transmission of the HOA representation is calculated as follows: s and the number of bits per sample, N b Given, O·f s N b Therefore, for each sample, N b = 16 bits are used to s Transmitting an order N=4 HOA representation at a sampling rate of N=48 kHz results in a bit rate of 19.2 Mbit / s, which is quite high for many practical applications, e.g., streaming.
[0005] Compression of HOA sound field representations has been proposed in European Patent Applications No. 12306569 and No. 12305537. Instead of individually perceptually encoding the HOA coefficient sequences, as is done, for example, in E. Hellerud, I. Burnett, A. Solvang, and U.S. Pensson, "Coding Higher-Order Ambisonics with AAC," 124th AES Convention, Amsterdam, 2008, attempts have been made to reduce the number of signals to be perceptually encoded by specifically analyzing the sound field and decomposing a given HOA representation into directional and residual ambient components. Generally, the directional components are assumed to be represented by a small number of dominant directional signals that can be viewed as general plane wave functions. The order of the residual ambient HOA components is reduced because, after extracting the dominant directional signals, the lower-order HOA coefficients are believed to retain the most relevant information. Summary of the Invention
[0006] In summary, by performing such processing, the initial number of HOA coefficient sequences to be perceptually encoded (N+1) 2 is a given number of D dominant directional signals and a cut order N REDThe ambient HOA components of the residual are represented using N (N RED + 1) 2 the number of HOA coefficient sequences is reduced to. Thereby, the number of signals to be encoded is determined, that is, D+(N RED + 1) 2 becomes. In particular, this number is independent of the actually detected number D ACT (k) ≤ D of the active dominant directional sound sources in time frame k. This means that in time frame k, if the actually detected number D ACT (k) of the active dominant directional sound sources is less than the maximum allowable number D of the directional signals, some or even all of the perceptually encoded dominant directional signals will become zero. That is, this means that these multiple channels are not used at all to capture the relevant information of the sound field.
[0007] In this situation, another assumed weakness of the processing in European Patent Application No. 12306569 and European Patent Application No. 12305537 is the criterion for determining the number of dominant directional signals within each time frame. The reason is that no attempt has been made to determine the optimal number of active dominant directional signals for the continuous perceptual coding of the sound field. For example, in European Patent Application No. 12305537, the number of dominant sound sources is estimated using a simple power criterion, that is, by obtaining the dimension of the subspace of the correlation matrix between the coefficients belonging to the largest eigenvalue. In European Patent Application No. 12306569, an incremental detection of dominant directional sound sources has been proposed. Here, when the power of the plane wave function from each direction is sufficiently high with respect to the first directional signal, the directional sound source is considered to be dominant. Using a power-based criterion as in the cases of European Patent Application No. 12306569 and European Patent Application No. 12305537 may result in a direction-ambient decomposition that is not optimal for the perceptual coding of the sound field.
[0008] The problem solved by the present invention is to improve HOA compression by determining how to allocate coefficients for directional signals and ambient HOA components to a predetermined reduced number of channels for the current HOA audio signal content. This problem is solved by the respective methods disclosed in claims 1 and 2. An apparatus utilizing these methods is disclosed in claim 4.
[0009] The present invention improves the compression process proposed in EP 12306569 in two aspects. First, the bandwidth provided by a given number of channels to be perceptually coded is better utilized. In time frames where no dominant sound source signal is detected, the channels originally reserved for the dominant directional signals are used to capture additional information about the ambient components in the form of additional HOA coefficient sequences of the residual ambient HOA components. Second, keeping in mind the goal of utilizing a given number of channels to perceptually code a given HOA sound field representation, the criterion for determining the number of directional signals to be extracted from the HOA representation is adapted to that goal. The number of directional signals is determined so that the decoded and reconstructed HOA representation produces the smallest perceived error. The criterion compares the modeling error resulting from extracting a directional signal and using fewer HOA coefficient sequences to describe the residual ambient HOA component with the modeling error resulting from not extracting a directional signal and instead using additional HOA coefficient sequences to describe the residual ambient HOA component. The criterion further takes into account, for both cases, the spatial power distribution of the quantization noise introduced by the perceptual coding of the HOA coefficient sequences of the directional signal and the residual ambient HOA component.
[0010] To perform the above process, the total number of signals (channels) I is determined before starting the HOA compression. This total number I is reduced compared to the number of the original O HOA coefficient sequences. The ambient HOA components are reduced to the minimum number O. REDIt is assumed that the HOA coefficients are expressed by a sequence of HOA coefficients. In some cases, the minimum number may be zero. The remaining D=I-O RED These channels are assumed to contain either directional signals or additional coefficient sequences of ambient HOA components, depending on which is more perceptually meaningful as determined by the directional signal extraction process. The allocation of either directional signals or ambient HOA component coefficient sequences to the remaining D channels is assumed to be changeable on a frame-by-frame basis. For sound field reconstruction at the receiver side, information about this allocation is transmitted as additional side information.
[0011] In principle, the compression method of the present invention is suitable for compressing a high-order Ambisonics representation of a sound field, called HOA, using a given number of perceptual coding processes, with a time frame of input HOA coefficients. The method is performed frame-by-frame, - estimating for a current frame a set of dominant directions and a data set of indices of the corresponding detected directional signals; - decomposing the HOA coefficient sequence of the current frame into a non-predetermined number of directional signals, using each direction included in the set of dominant direction estimates and a data set of indices of each of the directional signals, the non-predetermined number being smaller than the predetermined number, residual ambient HOA components represented by a reduced number of HOA coefficient sequences corresponding to the difference between the predetermined number and the non-predetermined number, and a data set of indices of the corresponding reduced number of residual ambient HOA coefficient sequences; - allocating HOA coefficient sequences of the directional signals and the residual ambient HOA components to a number of channels corresponding to the predetermined number, wherein for said allocation, the data set of indices of the directional signals and the data set of indices of the reduced number of residual ambient HOA coefficient sequences are used; - perceptually encoding said channels of relevant frames, said perceptual encoding resulting in encoded compressed frames.
[0012] In principle, the compressor of the invention is suitable for compressing a high-order Ambisonics representation, called HOA, of a sound field using a number of perceptual coding processes, with an input time frame of a sequence of HOA coefficients. The device performs frame-by-frame processing; - means configured to estimate, for a current frame, a dataset of a set of dominant directions and corresponding indices of detected directional signals; - means configured to decompose the HOA coefficient sequence of the current frame into a non-predetermined number of directional signals, the non-predetermined number being smaller than the predetermined number, using each direction included in the set of dominant direction estimates and each data set of indices of the directional signals, and residual ambient HOA components represented by a reduced number of HOA coefficient sequences corresponding to the difference between the predetermined number and the non-predetermined number, and a corresponding data set of indices of the corresponding reduced number of residual ambient HOA coefficient sequences; - means adapted to allocate HOA coefficient sequences of the directional signals and the residual ambient HOA components to a number of channels corresponding to the predetermined number, wherein for said allocation, said data set of indices of the directional signals and said data set of indices of the reduced number of residual ambient HOA coefficient sequences are used; and - means adapted to perceptually encode said channels of associated frames, such that encoded compressed frames are obtained.
[0013] In principle, the decompression method of the present invention is suitable for decompressing higher order Ambisonics representations compressed according to the compression methods described above. - decoding the current encoded compressed frame to obtain a perceptually decoded frame of the channel; - reallocating the perceptually decoded frames of channels to reconstruct the corresponding frames of directional signals and the corresponding frames of residual ambient HOA components using the data set of indices of detected directional signals and the data set of indices of the selected ambient HOA coefficient sequences; - resynthesizing the current decompressed frame of HOA representation from the frame of directional signals and the frame of residual ambient HOA components using the data set of indices of detected directional signals and the set of dominant direction estimates, A directional signal for uniformly distributed directions is predicted from the directional signal, and then the current decompressed frame is resynthesized from the frame of directional signals, the predicted signal, and the ambient HOA component of the residual.
[0014] In principle, the decompression device of the invention is suitable for decompressing higher-order Ambisonics representations compressed according to the compression method described above. - means configured to decode the current encoded compressed frame to obtain a perceptually decoded frame of the channel; means configured to reallocate the perceptually decoded frames of the channels to reconstruct the corresponding frames of the directional signals and the corresponding frames of the residual ambient HOA components using the data set of indices of the detected directional signals and the data set of indices of the selected ambient HOA coefficient sequences; means configured to resynthesize a current decompressed frame of the HOA representation from the frame of directional signals and the frame of residual ambient HOA components using the data set of indices of detected directional signals and the set of dominant direction estimates, A directional signal for uniformly distributed directions is predicted from the directional signal, and then the current decompressed frame is resynthesized from the frame of directional signals, the predicted signal, and the ambient HOA component of the residual.
[0015] Further embodiments of the invention are advantageously disclosed in the respective dependent claims. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a block diagram of HOA compression. [Figure 2] FIG. 1 is a block diagram of dominant sound source direction estimation. [Figure 3] FIG. 1 is a block diagram of HOA decompression. [Figure 4] FIG. 1 illustrates a spherical coordinate system. [Figure 5] FIG. 10 illustrates the normalized dispersion function vN(Θ) for different Ambisonics orders N and angles θ∈[0,π]. DETAILED DESCRIPTION OF THE INVENTION
[0017] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. A. Improved HOA Compaction The compression process according to the present invention is based on European Patent Application No. 12306569 and is shown in Figure 1, where signal processing blocks that have been modified or newly introduced with respect to European Patent Application No. 12306569 are shown in bold boxes and are referred to as "in the present application."
number
[0018] For HOA compression, a frame-wise processing is used with non-overlapping input frames C(k) of HOA coefficient sequences of length L, where k represents the frame index, and the frame is defined in terms of the HOA coefficient sequence specified in equation (1) below.
number
[0019] Step or stage 11 / 12 in Figure 1 is performed arbitrarily and involves concatenating the kth and (k-1)th non-overlapping frames of the HOA coefficient sequence to form a long frame according to the following formula:
number
number
number
[0020] In principle, the dominant sound source estimation step or stage 13 is performed as proposed in European Patent Application No. 13305156, but with an important modification. This modification concerns the determination of the number of directions to be detected, i.e., how many directional signals are to be extracted from the HOA representation. This is achieved by the idea of extracting directional signals instead of using additional HOA coefficient sequences only when this is more perceptually relevant than using additional HOA coefficient sequences for a good approximation of the ambient HOA component. Section A.2 provides a detailed explanation of this technique.
[0021] A dataset of indices of detected directional signals with estimation of dominant sound sources
number
number
[0022] In step or stage 14, the current (long) frame of the HOA coefficient sequence
number
number
number
number
number
[0023] In particular, three cases should be distinguished:
[0024] a)N DIR,ACT (k-2)=N DIR,ACT(k-3): In this case, it is assumed that the same HOA coefficient sequence is selected as in frame k-3.
[0025] b)N DIR,ACT (k-2) <N DIR,ACT (k-3): In this case, more HOA coefficient sequences can be used to represent the ambient HOA components in the current frame than in the previous frame k-3. The HOA coefficient sequences selected in k-3 are assumed to be selected in the current frame as well. The additional HOA coefficient sequences can be selected according to different criteria. For example, the HOA coefficient sequence with the highest average power is designated as C AMB Select within (k-2) or HOA coefficient sequences with respect to their respective perceptual importance.
[0026] c)N DIR,ACT (k-2)>N DIR,ACT (k-3): In this case, fewer HOA coefficient sequences can be used to represent the ambient HOA components in the current frame than those present in the last frame k-3. The problem to be solved here is which of the already selected HOA coefficient sequences must be deactivated. A reasonable solution is to allocate the signal or channel in stage 16 in frame k-3.
number
[0027] To avoid discontinuities at frame boundaries when additional HOA coefficient sequences are activated or deactivated, it is advisable to smoothly fade in or out each signal.
[0028] O RED +N DIR,ACT The final (k-2) reduced number of ambient HOA representations is C AMB,RED The index of the selected ambient coefficient sequence is denoted by (k-2).
number
[0029] In Step / Stage 16, X DIR (k-2) contains the active directional signal and C AMB,RED The HOA coefficient sequence contained in (k-2) is assigned to frame Y(k-2) of I channels for individual perceptual coding. DIR (k-2), Y(k-2) and C AMD,RED (k-2) is the individual signal x DIR,d (k-2)(d∈{1,… ,D}), y i (k-2)(i∈{1,… ,D}) and c AMB, RED, ο It is assumed to consist of (k-2)(ο=1,… ,O).
number
[0030] To obtain a continuous signal for successive perceptual coding, an active directional signal is assigned to hold the index of each channel, which can be expressed as:
number
[0031] The HOA coefficient sequence for the ambient component is the minimum number of O RED The coefficient sequence is the last O of Y(k-2). RED are assigned so as to be always included in the signal, i.e., according to the following formula:
number
[0032] Additional DNs DIR,ACTFor the HOA coefficient sequences of the (k-2) ambient components, it should be distinguished whether they were also selected in the previous frame. a) Additional DNs DIR,ACT If the HOA coefficient sequences of the (k-2) ambient components have also been selected to be transmitted in the previous frame, i.e., each index is also included in the data set
number
number
number
number
[0033] This particular assignment provides the advantage that signal redistribution and synthesis during the HOA decompression process can be performed without knowledge of which ambient HOA coefficient sequences are contained in which of the Y(k-2) channels.
number
number
[0034] This allocation process creates an allocation vector
number
number
[0035] For frames where the vector γ(k) is not transmitted in step / stage 16, the decompressor uses the data parameter set
number
number
[0036] A.1 Estimation of dominant sound source direction The estimation step / stage 13 for the dominant sound source directions of FIG. 1 is depicted in more detail in FIG. 2. This is essentially done according to what is described in European Patent Application No. 13305156, with a crucial difference. The crucial difference is the way in which the number of dominant sound sources is determined. The number of dominant sound sources corresponds to the number of directional signals extracted from a given HOA representation. This number is important because it is used to control whether a given HOA representation is better represented, either by using more directional signals or, alternatively, by using more HOA coefficient sequences to better model the ambient HOA components.
[0037] The dominant sound source direction estimation is performed by a long frame of input HOA coefficients.
number
number
number
number
[0038] In step or stage 22, the preliminary direction estimates, directional signals, and HOA sound field components are calculated based on the number of directional signals to be extracted.
number
number
number
number
number
number
number
number
[0039] In step or stage 23, the resulting directional trajectories are smoothed according to a source motion model to determine which of the sources are considered active (see European Patent Application No. 13305156). This final process results in a set of indices of active directional sources.
number
number
[0040] A.2 Determining the number of directional signals to be extracted To determine the number of directional signals in step / stage 22, a situation is assumed in which there is a given total number of I channels utilized to capture the most perceptually relevant sound field information. Therefore, the number of directional signals to be extracted is determined taking into consideration whether the current HOA representation is better represented by either using more directional signals for the overall HOA compression / decompression quality or by using more HOA coefficient sequences for better modeling of the ambient HOA components. To derive a criterion for determining the number of directional sound sources to be extracted in step / stage 22, it is taken into consideration which criterion is related to human perception and that HOA compression is performed by, among other things, the following two processes: - Reduction of the HOA coefficient sequence for representing the ambient HOA components (which means a reduction in the number of channels involved) -Perceptual coding of HOA coefficient sequences to represent directional signals and ambient HOA components
[0041] Depending on the number M (0≦M≦D) of extracted directional signals, the first process performs an approximation according to the following formula:
number
number
number
number
number
[0042] The approximate calculation from the second process can be expressed by the following formula:
number
number
number
number
[0043] Formation of standards Number of directional signals extracted
number
number
number
number
number
number
number
number
number
number
number
number
number
[0044] A process of subtracting "1" and finding successive maxima is performed to ensure that the perceived level is zero as long as the error power is below the masking threshold. Finally, the number of extracted directional signals is
number
number
[0045] Alternatively, the maximum value of the error perception level in equation (15) can be replaced by an averaging process.
[0046] Calculation of directional perceptual masking power distribution Original HOA Representation
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0047] Calculation of directional power distribution In the following description, directional power distribution
number
[0048] a. One possibility is to create the desired HOA expression by performing the first two steps described in section A.2.
number
number
number
number
number
number
number
number
number
number
number
number
[0049] b.An alternative solution is
number
number
number
number
number
number
number
number
number
[0050] Below we describe how to calculate the directional power distributions of the three errors for each Bark scale critical band.
[0051] a.Error
number
number
number
number
number
number
number
number
number
number
[0052] b.Error
number
number
number
number
number
number
number
number
number
number
number
number
number
[0053] Rotated Direction
number
number
number
number
number
number
[0054] As a result, the true directional HOA component
number
number
number
number
number
number
number
[0055] Test direction Ω q For (q=1,… ,Q), the error in the spatial domain
number
number
[0056] Element β of the vector (d) (k)
number
number
number
number
number
number
number
number
number
[0057] c. Error resulting from perceptual coding of the HOA coefficient sequence of the ambient HOA component
number
number
number
number
number
number
number
number
number
[0058] B. Improved HOA decompression The corresponding HOA decompression process is shown in FIG. 3 and includes the following steps or stages:
[0059] In step or stage 31,
number
number
[0060] In the signal redistribution stage or stage 32,
number
number
number
number
number
[0061] In a synthesis step or stage 33, a frame of directional signals is generated (in accordance with the process described in connection with Figures 2b and 4 of European Patent Application No. 12306569).
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0062] C. Higher-Order Ambisonics Fundamentals Higher Order Ambisonics (HOA) is based on the description of the sound field within a compact region of interest, where the absence of sound sources is assumed. In that case, the spatiotemporal behavior of the sound pressure p(t,x) at time t and position x within the region of interest is completely determined physically by the wave equation for a homogeneous medium. The following is based on the spherical coordinate system shown in Figure 4. In the coordinate system used, the x-axis points to the front position, the y-axis points to the left, and the z-axis points upward. Position x=(r,θ,φ) in space T is expressed in terms of a radius r>0 (i.e., the distance to the coordinate origin), a tilt angle θ∈[0,π] measured from the polar axis z, and an azimuth angle φ∈[0,2π] measured counterclockwise in the xy plane from the x-axis. T represents the transposition.
[0063] F t The Fourier transform of the sound pressure versus time is represented by (·), i.e.,
number
number
number
number
[0064] If a sound field is represented by an infinite superposition of harmonic plane waves of different angular frequencies ω, arriving from all possible directions specified by the pair of angles (θ,φ), it can be shown that each plane-wave complex amplitude function C(ω,θ,φ) can be expressed by the following spherical harmonic expansion (see B. Rafaely, "Plane-wave Decomposition of the Sound Field on a Sphere by Spherical Convolution", Journal of the Acoustical Society of America 4(116), pp. 2149-2157, 2004):
number
number
number
number
[0065] Individual coefficients
number
number
number
number
number
[0066] The final Ambisonics format is s to yield a sampled version of c(t) below:
number
number
[0067] C.1 Definition of real-valued spherical harmonics real-valued spherical harmonics
number
number
number
number
[0068] C.2 Spatial resolution of higher-order Ambisonics Direction Ω0=(θ0,φ0) T A general plane wave function x(t) coming from is expressed in the HOA by the following equation:
number
number
number
number
[0069] As can be seen from equation (51), this is a general plane wave function x(t) and a spatial dispersion function ν N (Θ) and the spatial dispersion function ν N (Θ) can be shown to depend only on the angle Θ between Ω and Ω with the property
number
number
[0070] However, for finite dimensions N, the contribution of a general plane wave from direction Ω bleeds into neighboring directions, and the degree of this bleed decreases with increasing order. The normalized function ν for several different values of N N A plot of (Θ) is shown in FIG.
[0071] It is pointed out that the time domain behavior of the spatial density of the plane wave amplitude in any direction Ω is a multiple of the time domain behavior of the spatial density of the plane wave amplitude in any other direction. In particular, the functions c(t,Ω1) and c(t,Ω2) for some given directions Ω1 and Ω2 with respect to time t are highly correlated.
[0072] C.3 Spherical Harmonic Transform The spatial density of plane wave amplitudes is O spatial directions Ω o When discretized in (1≦ο≦O), the spatial direction Ω o are distributed almost uniformly on the unit sphere, but O directional signals c(t,Ω o ) is obtained. These signals are collected into a vector, which is expressed as follows:
number
number
number
[0073] direction Ω o In general, the modal matrix is invertible because the θ is approximately uniformly distributed on the unit sphere. Therefore, the continuous Ambisonics representation is o ) can be calculated using the following formula:
number
[0074] Both formulas constitute the transform and inverse transform between the Ambisonics representation and the spatial domain, which in this application are called the spherical harmonics transform and the inverse spherical harmonics transform.
[0075] In addition, the direction Ω o Since is distributed almost uniformly on the unit sphere, an approximate calculation
number
[0076] Advantageously, all of the above relationships are also valid in the discrete time domain.
[0077] The processing of the present invention can be carried out by a single processor or electronic circuit, or by multiple processors or electronic circuits operating in parallel and / or operating on different parts of the processing of the present invention.
Claims
1. 1. A method for decompressing a compressed Higher Order Ambisonics (HOA) representation, comprising: decoding the compressed HOA representation to provide a decoded frame of the signal and an assignment vector indicating a first index of a coefficient sequence that may include an ambient HOA component, the assignment vector indicating a first index of a coefficient sequence that may include an ambient HOA component that is related to a non-zero ambient HOA component; determining a first set of indices indicating active directional signals of a decoded frame of the signal and a direction of each of the active directional signals; redistributing the decoded frames of the signal based on the first set of indexes and the respective directions, the redistribution determining frames of HOA directional signals and frames of ambient HOA components; outputting a frame of the HOA directionality signal; outputting the frame of the ambient HOA component; A method comprising:
2. the frame of the HOA directional signal is generated based on the first set of indexes; The method of claim 1.
3. A program that, when executed by a processor, causes the processor to carry out the method of claim 1 or 2.
4. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 or 2.
5. 1. An apparatus for decompressing a compressed Higher Order Ambisonics (HOA) representation, comprising: a decoder that decodes the compressed HOA representation to provide a decoded frame of the signal and an assignment vector indicating a first index of a coefficient sequence that may contain an ambient HOA component, the assignment vector indicating a first index of a coefficient sequence that may contain an ambient HOA component that is related to a non-zero ambient HOA component; a first processing unit for determining a first set of indices indicating active directional signals of a decoded frame of the signal and a direction of each of the active directional signals; a second processing unit that reallocates the decoded frames of the signal based on the first set of indexes and the respective directions, the reallocation determining frames of an HOA directional signal and frames of an ambient HOA component; The second processing unit outputs a frame of the HOA directionality signal and outputs a frame of the ambient HOA component. Device.
6. the first set of indices corresponds to a corresponding set of direction estimates; The method of claim 1.
7. the corresponding set of direction estimates corresponds to energetically dominant components of the compressed HOA representation. The method of claim 6.
Citation Information
Patent Citations
Method and apparatus for encoding and decoding successive frames of ambisonics representation of two-dimensional or three-dimensional sound field
JP2012133366A
Method and apparatus for encoding and optimally reproducing a three-dimensional sound field
JP2012514358A