Method and device for decoding audio soundfield representation for audio playback

The method addresses decoding challenges in Ambisonics by using VBAP to calculate panning functions and a pseudo-inverse mode matrix, improving spatial localization and timbre accuracy in 3D audio reproduction with irregular speaker setups.

JP2025163200APending Publication Date: 2025-10-28DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025131171
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2010-03-26
Filing Date
2025-08-06
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing speaker setups for 3D audio reproduction, such as Ambisonics, face challenges in decoding sound field representations due to non-normal spatial distributions, leading to incorrect localization and timbre issues, particularly with irregular speaker arrangements.

Method used

A method using a geometric approach based on Vector Basis Amplitude Panning (VBAP) to derive panning functions and a pseudo-inverse mode matrix for calculating the decoding matrix, eliminating the need for source direction knowledge and addressing matrix inversion difficulties.

Benefits of technology

Improves spatial localization and reduces timbre alteration by deriving a decoding matrix that accurately reproduces sound fields with irregular speaker setups, enhancing the quality of 3D audio reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025163200000001_ABST
    Figure 2025163200000001_ABST
Patent Text Reader

Abstract

To provide a method, system, and nontransitory computer readable medium for decoding an audio soundfield representation for audio playback.SOLUTION: An improved method for decoding an audio soundfield representation for audio playback comprises: a step (110) of calculating a panning function (W) using a geometrical method based on positions of a plurality of loudspeakers and a plurality of source directions; a step (120) of calculating a mode matrix (Ξ) from the loudspeaker positions, a step (130) of calculating a pseudo-inverse mode matrix (Ξ+); and a step (130) of decoding the audio soundfield representation. The decoding is based on a decode matrix (D) that is obtained from the panning function (W) and the pseudo-inverse mode matrix (Ξ+).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for decoding an audio sound field representation, and more particularly to an Ambisonics formatted audio representation for audio reproduction. [Background technology]

[0002] This section is intended to introduce the reader to aspects of art that may be related to various aspects of the present invention, which are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present invention. As such, it should be understood that these statements are to be read in this light, and not as admissions of prior art, unless expressly stated as source.

[0003] Accurate localization is a primary goal of any spatial audio playback system. Such playback systems are highly practical for conferencing systems, games, or other virtual environments that benefit from 3D sound. A 3D sound scene can be synthesized or captured as a natural sound field. Sound field signals, such as Ambisonics, represent the desired sound field. The Ambisonics format is based on a spherical harmonic decomposition of the sound field. The basic Ambisonics format, or B-format, uses spherical harmonics of orders 0 and 1, while so-called Higher Order Ambisonics (HOA) also uses additional spherical harmonics of at least order 2. A decoding process is required to obtain the individual speaker signals. To synthesize an audio scene, panning functions related to the spatial speaker arrangement are required to obtain the spatial localization of a given sound source. If a natural sound field is recorded, a microphone array is required to capture the spatial information. The well-known Ambisonics technique is a very suitable tool for achieving this. An Ambisonics-formatted signal carries a representation of the desired sound field. A decoding process is required to obtain the individual speaker signals from such an Ambisonics-formatted signal. Again, the panning function is the main issue for describing the task of spatial localization, since it can be derived from the decoding function. The spatial arrangement of the speakers is referred to as the speaker setup in this paper.

[0004] Commonly used speaker setups are the stereo setup with two speakers, the standard surround setup with five speakers, and the extended surround setup with more than five speakers. Although these setups are well known, they are limited to two dimensions (2D), for example, they do not reproduce height information.

[0005] Speaker setups for three-dimensional (3D) reproduction are described in, for example, the 2+2+2 configuration of NHK Ultra High Definition TV in 22.2 format or the 2+2+2 configuration of the Dubbing House (mdg-musikproduction dabringhaus und grimm, www.mdg.de) and the proposal for a 10.2 setup in Non-Patent Document 2 (NPL 1). One of the few known systems that mention spatial reproduction and panning strategies is the vector base amplitude panning (VBAP) technique in Non-Patent Document 3. VBAP (Vector Base Amplitude Panning) was used by Non-Patent Document 3 to reproduce virtual acoustic sources with any speaker setup. To place a virtual source in a 2D plane, a pair of speakers is required. On the other hand, for 3D, a triplet of speakers is required. For each virtual source, a monophonic signal with a different gain (depending on the virtual source's position) is applied to selected speakers from the full setup. The speaker signals for all virtual sources are then summed. VBAP applies a geometric approach to calculate the gain of speaker signals for panning between speakers.

[0006] The newly proposed exemplary 3D speaker setup considered in this paper has 16 speakers positioned as shown in Figure 2. This positioning was chosen for practical considerations: there are four pillars with three speakers each, with additional speakers between the pillars. More specifically, eight speakers are evenly distributed on a circle around the listener's head, with a 45-degree angle between them. The four additional speakers are positioned above and below, enclosing a 90-degree azimuth angle. For Ambisonics, this setup is irregular, leading to problems in decoder design, as discussed in [4].

[0007] Conventional Ambisonics decoding, as described in [5], uses a commonly known mode mapping process. Modes are described by mode vectors containing spherical harmonic function values ​​for distinct incident directions. The combination of all directions provided by individual loudspeakers results in a mode matrix for the loudspeaker setup. The mode matrix thus represents the loudspeaker positions. To reproduce distinct source signal modes, the loudspeaker modes are weighted so that the superimposed modes of the individual loudspeakers sum to the desired mode. To obtain the necessary weights, an inverse matrix representation of the speaker mode matrix must be calculated. For signal decoding, the weights represent the loudspeaker drive signals, and the inverse speaker mode matrix, called the "decoding matrix," is applied to decode the Ambisonics-formatted signal representation. In particular, for many loudspeaker setups, such as the setup shown in Figure 2, it is difficult to invert the mode matrix.

[0008] As mentioned above, commonly used speaker sets are constrained to 2D, i.e., height information is not reproduced. Mathematically decoding the sound field representation of a speaker setup with a non-regular spatial distribution leads to problems with localization and coloration using commonly known techniques. To decode Ambisonics signals, a decoding matrix (i.e., a matrix of decoding coefficients) is used. Conventional decoding of Ambisonics signals, especially HOA signals, presents at least two problems. First, correct decoding requires knowledge of the signal source direction in order to derive the decoding matrix. Second, mapping to existing speaker setups is systematically incorrect due to the following mathematical problem: mathematically correct decoding yields not only positive speaker amplitudes but also some negative speaker amplitudes. However, these are erroneously reproduced as positive signals, thereby resulting in the problems mentioned above. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] K. Hamasaki, T. Nishiguchi, R. Okumaura, and Y. Nakayama, "Wide listening area with exceptional spatial sound quality of a 22.2 multichannel sound system", Audio Engineering Society Preprints, Vienna, Austria, May 2007. [Non-patent document 2] T. Holman, "Sound for Film and Television", 2nd ed., Boston, Focal Press, 2002. [Non-patent document 3] Pulkki, "Virtual sound source positioning using vector base amplitude panning", Journal of Audio Engineering Society, vol.45, no.6, pp.456-466, June 1997. [Non-patent document 4] H. Pomberger and F. Zotter, "An ambisonics format for flexible playback layouts," Proceedings of the 1st Ambisonics Symposium, Graz, Austria, July 2009 [Non-patent document 5] M. Poletti, "Three-dimensional surround sound systems based on spherical harmonics", J. Audio Eng. Soc, vol.53, no.11, pp.1004-1025, Nov. 2005 Summary of the Invention [Problem to be solved by the invention]

[0010] The present invention describes a method for decoding sound field representations for non-normal spatial distributions with significantly improved localization and timbre attributes. [Means for solving the problem]

[0011] This method represents an alternative way of obtaining a decoding matrix for sound field data, e.g., data in Ambisonics format, and uses a process in a system estimation manner. Given a set of possible incidence directions, panning functions related to the desired loudspeakers are calculated. The panning functions are taken as the output of the Ambisonics decoding process. The required input signals are the mode matrices for all possible directions. Thus, the decoding matrix is ​​obtained by post-multiplying the weighting matrix by the inverse version of the mode matrix of the input signal, as shown below:

[0012] Regarding the second problem mentioned above, it has been discovered that it is also possible to derive the decoding matrix from the inverse of a so-called mode matrix, which represents the speaker positions, and a position-dependent weighting function ("panning function") W. One aspect of the present invention is that these panning functions W can be derived using a different method than commonly used. Advantageously, a simple geometric method is used. Such a method does not require knowledge of any signal source direction, thus solving the first problem mentioned above. One such method is known as "vector-based amplitude panning" (VBAP). According to the present invention, VBAP is used to calculate the required panning function, which is then used to calculate the Ambisonics decoding matrix. Another problem arises in that the inverse of the mode matrix (representing the speaker setup) is required. However, the exact inverse is difficult to calculate, which also leads to erroneous audio reproduction. Therefore, an additional aspect is that to obtain the decoding matrix, a pseudo-inverse mode matrix, which is much easier to calculate, is calculated.

[0013] The present invention uses a two-stage approach: the first stage is the derivation of panning functions that depend on the speaker setup used for playback. In the second stage, the Ambisonics decoding matrix is ​​calculated from these panning functions for all speakers.

[0014] One advantage of the present invention is that a parametric description of the sound source is not required, and sound field descriptions such as Ambisonics can be used.

[0015] According to the present invention, a method for decoding an audio sound field representation for audio reproduction comprises the steps of: calculating, for each of a plurality of speakers, a panning function using a geometric method based on the positions of the speakers and a plurality of source directions; calculating a mode matrix from the source directions; calculating a pseudo-inverse mode matrix of the mode matrix; and decoding the audio sound field representation, wherein the decoding is based on a decoding matrix obtained from at least the panning function and the pseudo-inverse mode matrix.

[0016] According to another aspect, an apparatus for decoding an audio sound field representation for audio reproduction includes: first computing means for computing, for each of a plurality of loudspeakers, a panning function using a geometric method based on positions of the loudspeakers and a plurality of source directions; second computing means for computing a modal matrix from the source directions; third computing means for computing a pseudo-inverse modal matrix of the modal matrix; and decoder means for decoding the sound field representation, the decoding being based on a decoding matrix, the decoder means obtaining the decoding matrix using at least the panning function and the pseudo-inverse modal matrix. The first, second, and third computing means may be a single processor or two or more separate processors.

[0017] According to yet another aspect, a computer-readable medium has stored executable instructions for causing a computer to execute a method for decoding an audio sound field representation for audio reproduction, the method including: calculating, for each of a plurality of speakers, a panning function using a geometric method based on positions of the speakers and a plurality of source directions; calculating a modal matrix from the source directions; calculating a pseudo-inverse of the modal matrix; and decoding the audio sound field representation, the decoding based on a decoding matrix obtained from at least the panning function and the pseudo-inverse modal matrix.

[0018] Advantageous embodiments of the invention are disclosed in the dependent claims, the following description and the drawings. [Brief explanation of the drawings]

[0019] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] 2 is a flowchart of the method. [Figure 2] FIG. 1 illustrates an exemplary 3D setup with 16 speakers. [Figure 3] Figure 1 shows the beam patterns resulting from decoding using non-regularized mode matching. [Figure 4] FIG. 1 shows the beam pattern resulting from decoding using a regularized mode matrix. [Figure 5] This figure shows the beam pattern resulting from decoding using a decoding matrix derived from VBAP. [Figure 6] FIG. 10 shows the results of a listening test. [Figure 7] FIG. 1 is a block diagram of the device. DETAILED DESCRIPTION OF THE INVENTION

[0020] As shown in Figure 1, the audio sound field representation SF for audio playbackc The method for decoding comprises a step 110 of calculating, for each of a plurality of speakers, a panning function W using a geometric method based on the positions 102 of the speakers (L is the number of speakers) and a plurality of source directions 103 (S is the number of source directions), a step 120 of calculating a modal matrix Ξ from the source directions and a given order N of the sound field representation, and a step 130 of calculating a pseudo-inverse modal matrix Ξ of the modal matrix Ξ. + and a step 130 of calculating said audio sound field representation SF c Decode the decoded sound data AU dec and obtaining the panning function W and the pseudo-inverse modal matrix Ξ. + In one embodiment, the pseudo-inverse mode matrix is ​​based on the (135) decoding matrix D obtained from + =Ξ H [ΞΞ H ] -1 The order N of the sound field representation may be predefined or may be obtained by the input signal SF c may be extracted 105 from

[0021] As shown in FIG. 7 , the apparatus for decoding an audio sound field representation for audio reproduction includes a first calculation means 210 for calculating, for each of a plurality of speakers, a panning function W using a geometric method based on the positions 102 of the speakers and a plurality of source directions 103, a second calculation means 220 for calculating a modal matrix Ξ from the source directions, and a pseudo-inverse modal matrix Ξ of the modal matrix Ξ. + and a decoder means 240 for decoding said sound field representation, said decoding being based on a decoding matrix D, which comprises at least said panning function W and said pseudo-inverse modal matrix Ξ. + , by a decoding matrix calculation means 235 (e.g., a multiplier). The decoder means 240 uses the decoding matrix D to generate the decoded audio signal AU decThe first, second and third calculation means 220, 230, 240 may be a single processor or two or more separate processors. The order N of the sound field representation may be predefined or may be calculated based on the input signal SF c The order may be obtained by a means 205 for extracting the order from

[0022] A particularly useful 3D speaker setup has 16 speakers. As shown in Figure 2, there are four pillars with three speakers each, with additional speakers between the pillars. Eight speakers are evenly distributed on a circle around the listener's head, at a 45-degree angle. Four additional speakers are positioned above and below, at a 90-degree azimuth angle. For Ambisonics, this setup is irregular, leading to problems in decoder design.

[0023] Below, we describe Vector Basis Amplitude Panning (VBAP) in detail. In one embodiment, VBAP is used in this application to position virtual acoustic sources with an arbitrary speaker setup, where the same distance of the speakers from the listening position is assumed. VBAP uses three speakers to position one virtual source in 3D space. For each virtual source, monophonic signals with different gains are provided to the speakers to be used. The gains for different speakers depend on the position of the virtual source. VBAP is a geometric approach to calculate the gain of speaker signals for panning between speakers. In the 3D case, three speakers arranged in a triangle construct a vector basis. Each vector basis is defined by speaker numbers k, m, n and a speaker position vector l given in Cartesian coordinates normalized to length 1. k ,l m ,l n The vector basis for speakers k, m, and n is L kmn ={l k ,l m ,l n} (1) is defined by

[0024] The desired direction Ω = (θ, φ) of the virtual source must be given as an azimuth angle φ and a tilt angle θ. Therefore, the position vector p(Ω) of the virtual source of length 1 in Cartesian coordinates is given by p(Ω)={cosφsinθ,sinφsinθ,cosθ} T (2) is defined by

[0025] The virtual source locations are given by the vector basis and the gain factor g(Ω) = ( ~ g k , ~ g m , ~ g n ) T Using p(Ω)=L kmn g(Ω)= ~ g k l k + ~ g m l m + ~ g n l n (3) It can be expressed by:

[0026] By inverting the vector basis matrix, the required gain factor is g(Ω)=L -1 kmn p(Ω) (4) It can be calculated by:

[0027] The vector basis used is determined according to [3]: first, the gain is calculated for all vector bases according to [3]. Then, for each vector basis, the minimum over their gain factors is calculated as ~ g min =min{ ~ g k , ~ g m , ~ g n}. Finally, ~ gmin The vector basis with the highest value of is used. The resulting gain factors must not be negative. Depending on the acoustic properties of the listening room, the gain factors may be normalized for energy conservation.

[0028] In the following, an exemplary sound field format, the Ambisonics format, is described. The Ambisonics representation is a method of describing a sound field using a mathematical approximation of the sound field at a single position. Using a spherical coordinate system, the pressure at a point r = (r,θ,φ) in space can be calculated using the spherical Fourier transform

number

[0029] For simplicity, plane waves are often assumed to represent the sound field. s The Ambisonics coefficients describing a plane wave as an acoustic source from

[0030]

number

[0031]

number

number

[0032] Mode matching is a commonly used approach to compute loudspeaker signals from an Ambisonics representation of a sound field. The basic idea is to find the s ) into the speaker sound field description A(Ω l ) weighted sum

number

[0033] Y(Ω s ) * =Ψw(Ω s ) (9) where Ψ is the modal matrix of the speaker setup Ψ=[Y(Ω1) * ,Y(Ω2) * ,…,Y(Ω L ) * ] (10) and has O×L elements. To obtain the desired weighting vector w, various strategies are known to achieve this. If M=3 is chosen, Ψ is square and can be invertible. However, due to non-normal speaker setups, the matrix scales poorly. In such cases, the pseudo-inverse matrix is ​​often chosen. D=[Ψ H Ψ] -1 Ψ H (11) gives the L×O decoding matrix D. Finally, w(Ω s )=DY(Ω s )* (12) where the weight w(Ω s ) is the minimum energy solution for equation (9). We will discuss the consequences of using the pseudoinverse later.

[0034] Below we explain the connection between the panning functions and the Ambisonics decoding matrix. Starting from Ambisonics, the panning functions for the individual speakers can be calculated using equation (12).

[0035] Ξ=[Y(Ω1) * ,Y(Ω2) * ,…,Y(Ω S ) * ] (13) S input signal directions (Ω s ) mode matrix. The input signal directions are, for example, a spherical grid with tilt angles running from 1° to 180° in 1 degree increments and azimuth angles running from 1° to 360°. This mode matrix has O×S elements. Using equation (12), the resulting matrix W has L×S elements. Row l has S pan weights for each speaker.

[0036] W=DΞ (14) As a representative example, the panning function for a single loudspeaker 2 is shown as a beam pattern in Figure 3. In this example, the decoding matrix D has order M=3. As can be seen, the panning function values ​​are completely unrelated to the physical positioning of the loudspeakers. This is due to the mathematically non-normal positioning of the loudspeakers, which does not provide sufficient spatial sampling for the chosen order. Therefore, the decoding matrix is ​​referred to as a non-normalized modal matrix. This problem can be overcome by normalizing the loudspeaker mode matrix Ψ in Equation (11). This solution works at the expense of the spatial resolution of the decoding matrix, which can be expressed as a reduction in Ambisonics order. Figure 4 shows an example beam pattern resulting from decoding using a normalized modal matrix, specifically using the average of the eigenvalues ​​of the modal matrix for normalization. Compared to Figure 3, the direction of the intended loudspeaker is now clearly discernible.

[0037] As outlined in the introduction, another way to obtain the decoding matrix D for the reproduction of Ambisonics signals is possible if the panning function is known. The panning function W is viewed as the desired signal defined on a set of virtual source directions Ω, and the mode matrix Ξ of these directions acts as the input signal. The decoding matrix can then be calculated using the following formula:

[0038] D=WΞ H [ΞΞ H ] -1 =WΞ + (15) Here, Ξ H [ΞΞ H ] -1 or simply Ξ + is the pseudo-inverse of the mode matrix Ξ. In this new approach, we take the panning function in W from VBAP and compute the Ambisonics decoding matrix from this.

[0039] The panning function for W is taken as the gain value g(Ω) calculated using equation (4), where Ω is chosen according to equation (13). The resulting decoding matrix using equation (15) is the Ambisonics decoding matrix that facilitates the VBAP panning function. An example showing the beam pattern resulting from decoding using the VBAP-derived decoding matrix is ​​depicted in Figure 5. Advantageously, the sidelobes SL are much smaller than the sidelobes SL of the normalized mode-matching result in Figure 4. reg Furthermore, the VBAP-derived beam patterns for individual speakers follow the geometry of the speaker setup, since the VBAP panning function depends on the vector basis of the targeted direction. As a result, our new approach produces better results across all orientations of the speaker setup.

[0040] The source directions 103 can be defined quite freely. The condition for the number of source directions S is that there must be at least (N+1) 2 Therefore, the sound field signal SF c For a given degree N, S≧(N+1) 2 and distribute the S source directions evenly over the unit sphere. As mentioned above, the result can be a spherical grid with tilt angles running in regular increments of x degrees (e.g., x=1...5 or x=10,20, etc.) from 1°...180° and azimuth angles from 1...360°. Each source direction Ω=(θ,φ) can be given by an azimuth angle φ and a tilt angle θ.

[0041] The beneficial effects were confirmed in listening tests. For the evaluation of single-source localization, a virtual source is compared against a real source as a reference. For the real source, a loudspeaker at the desired position is used. The playback methods used are VBAP, Ambisonics mode matching decoding, and the newly proposed Ambisonics decoding using the VBAP panning function according to the present invention. For the second and third methods, a third-order Ambisonics signal is generated for each tested position and each tested input signal. This synthesized Ambisonics signal is then decoded using the corresponding decoding matrix. The test signals used were wideband pink noise and a male speech signal. The tested positions were located in the front region with the following orientations:

[0042] Ω1=(76.1°,-23.2°), Ω2=(63.3°,-4.3°) (16) The listening test was conducted in an acoustic chamber with an average reverberation time of approximately 0.2 seconds. Nine people participated in the listening test. The subjects were asked to rate the spatial reproduction performance of all reproduction methods compared to the reference. A single rating value had to be found to represent the changes in localization and timbre of the virtual sources. Figure 5 shows the results of the listening test.

[0043] As the results show, unnormalized Ambisonics mode matching decoding was rated perceptually worse than the other methods tested. This result corresponds to Figure 3. The Ambisonics mode matching method acts as an anchor in this listening test. Another advantage is that the confidence interval for the noise signal is larger for VBAP than for the other methods. The mean value is highest for Ambisonics decoding using the VBAP panning function. Thus, although spatial resolution is reduced—due to the Ambisonics order used—this method demonstrates advantages over the parametric VBAP approach. Compared to VBAP, both Ambisonics decoding using the robust panning function and the VBAP panning function have the advantage that not only three speakers are used to render the virtual source. VBAP single-speaker decoding can be dominant when the virtual source location is close to one of the speaker's physical locations. Most subjects reported less timbre alteration with Ambisonics-driven VBAP than with directly applied VBAP. The problem of timbre alteration for VBAP is already known from [3]. Contrary to VBAP, the newly proposed method uses more than three speakers for the reproduction of one virtual source, but surprisingly, results in less timbre coloration.

[0044] In conclusion, a new method for deriving Ambisonics decoding matrices from VBAP panning functions is disclosed. For various loudspeaker setups, this approach has advantages over matrices from the mode-matching approach. The attributes and consequences of these decoding matrices are discussed above. In summary, the newly proposed Ambisonics decoding using VBAP panning functions avoids the typical problems of well-known mode-matching methods. Listening tests have shown that VBAP-derived Ambisonics decoding can produce better spatial reproduction quality than direct use of VBAP can. While VBAP requires a parametric description of the virtual sources to be rendered, the proposed method requires only a sound field description.

[0045] While the fundamental novel features of the present invention as applied to the preferred embodiments thereof have been illustrated, described, and pointed out, it will be understood that various omissions, substitutions, and changes may be made in the described apparatus and methods in the form and details of the disclosed apparatus and in its operation by those skilled in the art without departing from the spirit of the invention. Any combination of elements that perform substantially the same function in substantially the same way to achieve the same result is expressly intended to be within the scope of the present invention. The transfer of elements from one described embodiment to another is fully intended and contemplated. It will be understood that modifications of detail can be made without departing from the scope of the invention. Each feature disclosed in this document and (where appropriate) in the claims and drawings may be provided independently or in any suitable combination. Features may, where appropriate, be implemented in hardware, software, or a combination of both. Any reference signs appearing in the claims are for illustrative purposes only and shall have no limiting effect on the scope of the claims.

[0046] Several aspects will be described. [Aspect 1] 1. A method of decoding an audio sound field representation for audio reproduction, comprising: · calculating a panning function for each of a plurality of loudspeakers using a geometric method based on the positions of the loudspeakers and a plurality of source directions; calculating a modal matrix from the source directions; calculating a pseudo-inverse modal matrix of the modal matrix; and decoding the audio sound field representation, the decoding being based on a decoding matrix derived from at least the panning function and the pseudo-inverse modal matrix. method. [Aspect 2] 2. The method of embodiment 1, wherein the geometric method used in the step of calculating a panning function is vector basis amplitude panning (VBAP). Aspect 3 3. The method of embodiment 1 or 2, wherein the sound field representation is in at least second-order Ambisonics format. Aspect 4 Ξ is the modal matrix of the multiple source directions, and the pseudo-inverse modal matrix (Ξ + ) is Ξ H [ΞΞ H ] -1 4. The method of any one of embodiments 1 to 3, obtained according to Aspect 5 The decoding matrix is ​​D=WΞ, where W is the set of panning functions for each speaker. H [ΞΞ H ] -1 =WΞ + 5. The method of embodiment 4, wherein the compound is obtained according to the following formula: Aspect 6 1. An apparatus for decoding an audio sound field representation for audio reproduction, comprising: · first computing means for computing a panning function for each of a plurality of loudspeakers using a geometric method based on the positions of the loudspeakers and a plurality of source directions; second calculation means for calculating a modal matrix from the source directions; a third calculation means for calculating a pseudo-inverse modal matrix of the modal matrix; decoder means for decoding the sound field representation, the decoding being based on a decoding matrix, the decoder means using at least the panning function and the pseudo-inverse mode matrix to obtain the decoding matrix; Device. Aspect 7 7. The apparatus of claim 6, wherein the decoding apparatus further comprises: means for calculating the decoding matrix from the panning function and the pseudo-inverse mode matrix; Device. Aspect 8 8. The apparatus of embodiment 6 or 7, wherein the geometric method used in calculating the panning function is Vector Basis Amplitude Panning (VBAP). Aspect 9 9. The apparatus of any one of aspects 6 to 8, wherein the sound field representation is in at least second-order Ambisonics format. Aspect 10 where Ξ is the modal matrix of the plurality of source directions, and the pseudo-inverse modal matrix Ξ + Ξ + =Ξ H [ΞΞ H ] -1 10. The device according to any one of embodiments 6 to 9, obtained according to Aspect 11 The decoding matrix is ​​D=WΞ, where W is the set of panning functions for each speaker. H [ΞΞ H ] -1 =WΞ + 11. The apparatus of embodiment 10, wherein the means for calculating a decoding matrix according to Aspect 12 1. A computer-readable medium storing executable instructions for causing a computer to perform a method for decoding an audio sound field representation for audio reproduction, the method comprising: · calculating a panning function for each of a plurality of loudspeakers using a geometric method based on the positions of the loudspeakers and a plurality of source directions; calculating a modal matrix from the source directions; calculating a pseudo-inverse modal matrix of the modal matrix; and decoding the audio sound field representation, the decoding being based on a decoding matrix derived from at least the panning function and the pseudo-inverse modal matrix. Computer-readable medium. Aspect 13 13. The computer-readable medium of embodiment 12, wherein the geometric method used in calculating the panning function is vector basis amplitude panning (VBAP). Aspect 14 14. The computer-readable medium of aspect 12 or 13, wherein the sound field representation is in at least second-order Ambisonics format. Aspect 15 where Ξ is the modal matrix of the plurality of source directions, and the pseudo-inverse modal matrix Ξ + Ξ + =Ξ H [ΞΞ H ] -1 15. The computer-readable medium of any one of embodiments 12 to 14, obtained according to

Claims

1. 1. A method of decoding an Ambisonics audio sound field representation for reproduction, the Ambisonics audio sound field representation having order N, the method comprising: receiving the audio sound field representation by a processor configured to decode the audio sound field representation; receiving, by the processor, a decoding matrix for decoding the audio sound field representation to determine a decoded audio signal, the decoding matrix is ​​based on a mode matrix determined based on a source direction and an order of the Ambisonics audio sound field representation; the decoding matrix is ​​further based on a second matrix including panning functions for a first plurality of L speaker positions and a second plurality of S source directions, the second matrix having a size of L×S, the plurality of S source directions being distributed on a unit sphere, each direction of the plurality of S source directions including a corresponding azimuth angle and a corresponding tilt angle, S≧(N+1)^2, and the panning function being indicated by a gain value; determining the decoded audio signal based on multiplication of the decoding matrix and the audio sound field representation. method.

2. The method of claim 1 , wherein the decoding matrix is ​​predetermined.

3. 10. A non-transitory computer-readable medium having stored thereon executable instructions for causing a computer to perform the method of decoding an Ambisonics audio sound field representation for audio reproduction according to claim 1.

4. 1. A system for decoding an Ambisonics audio sound field representation for reproduction, the Ambisonics audio sound field representation having order N, the system comprising: a receiver for receiving the audio sound field representation; a processor receiving a decoding matrix for decoding the audio sound field representation to determine a decoded audio signal, the decoding matrix is ​​based on a mode matrix determined based on a source direction and an order of the Ambisonics audio sound field representation; a processor, wherein the decoding matrix is ​​further based on a second matrix including panning functions for a first plurality of L speaker positions and a second plurality of S source directions, the second matrix having a size of L×S, the plurality of S source directions being distributed on a unit sphere, each direction of the plurality of S source directions including a corresponding azimuth angle and a corresponding tilt angle, S≧(N+1)^2, and the panning function being indicated by a gain value; a decoder for determining the decoded audio signal based on multiplication of the decoding matrix and the audio sound field representation. system.