Method and apparatus for decoding stereo loudspeaker signals from higher order ambisonics audio signals
The method and apparatus for decoding stereo signals from higher-order Ambisonics audio signals improve localization and reduce sidelobes by using panning functions and attenuation, addressing issues in first-order decoders and enhancing sound quality for multi-speaker setups.
Patent Information
- Application Number
- JP2025145858
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-03-28
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-23
AI Technical Summary
First-order Ambisonics decoders, such as those based on Blumlein Stereo, suffer from high negative sidelobes and poor localization in the front direction, leading to issues like reproducing sound objects from the rear right direction on the left stereo loudspeaker.
A method and apparatus for decoding a stereo loudspeaker signal from higher-order Ambisonics audio signals using panning functions defined on a circle, employing circular harmonic functions and attenuation in specific directions to improve localization and reduce negative sidelobes, particularly for loudspeaker setups with three or more speakers arranged on a semicircle.
Enhances localization in the front direction, minimizes negative sidelobes, and provides attenuation in the rear direction, resulting in an improved direct-to-diffuse sound ratio and better dialogue intelligibility.
Smart Images

Figure 2025186291000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and apparatus for decoding a stereo loudspeaker signal from a high-order Ambisonics audio signal using a panning function for sampling points on a circle. [Background technology]
[0002] Decoding of Ambisonics representations for stereo loudspeaker or headphone setups is known for first order Ambisonics, e.g. from equation (10) in [1], and from [2]. These approaches are based on Blumlein stereo, as disclosed in [1]. Another approach uses mode matching: [3]. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] British Patent No. 394325 [Patent Document 2] International Publication No. 2011 / 117399 [Non-patent literature]
[0004] [Non-Patent Document 1] J.S. Bamford, J. Vender-kooy, "Ambisonic sound for us", Audio Engineering Society Preprints, Convention paper 4138, presented at the 99th Convention, October 1995, New York [Non-patent document 2] XiphWiki-Ambisonics http: / / wiki.xiph.org / index.php / Ambisonics#Default_channel_conversions_from_B-Format [Non-patent document 3] MA Poletti, "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics", J. Audio Eng. Soc., vol.53(11), pp.1004-1025, November 2005 [Non-patent document 4] S. Weinzierl, "Handbuch der Audiotechnik", Springer, Berlin, 2008, section 3.3.4.1 [Non-patent document 5] JM Batke, F. Keiler, "Using VBAP-derived panning functions for 3D Ambisonics decoding", Proc. of the 2nd International Symposium on Ambisonics and Spherical Acoustics, May 6-7 2010, Paris, France, URL http: / / ambisonics10.ircam.fr / drupal / files / proceedings / presentations / O14_47.pdf [Non-patent document 6] V. Pulkki, "Virtual sound source positioning using vector base amplitude panning", J. Audio Eng. Society, 45(6), pp.456-466, June 1997 [Non-Patent Document 7] Earl G. Williams, "Fourier Acoustics", vol.93 of Applied Mathematical Sciences, Academic Press, 1999 Summary of the Invention [Problem to be solved by the invention]
[0005] Such first-order Ambisonics approaches, like Ambisonics decoders based on Blumlein Stereo (Patent Document 1) with virtual microphones with a figure-of-eight pattern, either have high negative sidelobes or result in poor localization in the front direction, where, for example, a sound object from the rear right direction is reproduced on the left stereo loudspeaker.
[0006] The problem to be solved by the present invention is to provide an Ambisonics signal decoding with an improved stereo signal output. [Means for solving the problem]
[0007] This problem is solved by the methods disclosed in claims 1 and 2. An apparatus utilizing these methods is disclosed in claim 3.
[0008] This invention describes a process for a stereo decoder for higher-order Ambisonics (HOA) audio signals. Desired panning functions can be derived from panning laws for the placement of virtual sources among the loudspeakers. For each loudspeaker, desired panning functions for all possible input directions are defined. The Ambisonics decoding matrix is calculated similarly to the corresponding descriptions in [5] and [2]. The panning functions are approximated by circular harmonic functions; the approximation matches the desired panning function with less error as the Ambisonics order increases. For areas in front of the loudspeakers, particularly those between the loudspeakers, panning laws such as tangent law or vector base amplitude panning (VBAP) can be used. For directions behind the loudspeakers beyond the loudspeaker positions, panning functions with a slight attenuation of sounds from these directions are used.
[0009] A special case is to use half of the cardioid pattern for the rear direction, pointing towards the loudspeaker.
[0010] In this invention, the higher spatial resolution of higher-order Ambisonics is exploited, especially in the front region, and the attenuation of negative sidelobes in the rear direction increases with increasing Ambisonics order. The invention can also be used for loudspeaker setups with three or more loudspeakers arranged on a semicircle or a smaller arc (circle segment). The invention also facilitates more artistic downmixing to stereo, where some spatial regions receive greater attenuation. This is beneficial for producing an improved direct-to-diffuse sound ratio, which can improve dialogue intelligibility.
[0011] The stereo decoder according to the invention has several important attributes: good localization in the front direction between the loudspeakers, only small negative sidelobes in the resulting panning function and slight attenuation in the rear direction, and also allows attenuation or masking of spatial regions that would otherwise be perceived as noisy or annoying when listening to the two-channel version.
[0012] In comparison to Patent Document 2, the desired panning function is defined arc by arc, allowing well-known panning (e.g., VBAP or tangent law) to be used in the front region midway between loudspeaker positions, while the rear direction can be slightly attenuated, an attribute that is not achievable when using a first-order Ambisonics decoder.
[0013] In principle, the method of the invention is suitable for decoding a stereo loudspeaker signal l(t) from a high-order Ambisonics audio signal a(t), said method comprising: calculating a matrix G containing desired panning functions for all virtual sampling points from the azimuth angle values of the left and right loudspeakers and from the number S of virtual sampling points on the circle,
number
[0014] In principle, the method of the invention is suitable for determining a decoding matrix D that can be used to decode a stereo loudspeaker signal l(t)=Da(t) from a 2D higher-order Ambisonics audio signal a(t), said method comprising: receiving an order N of the Ambisonics audio signal a(t); The desired azimuth angle values of the left and right loudspeakers (φ L ,φ R ) and from the number S of virtual sampling points on the circle, calculating a matrix G containing desired panning functions for all virtual sampling points,
number
[0015] In principle, the device of the invention is suitable for decoding a stereo loudspeaker signal l(t) from a high-order Ambisonics audio signal a(t), said device comprising: means adapted to calculate a matrix G containing desired panning functions for all virtual sampling points from the azimuth angle values of the left and right loudspeakers and from the number S of virtual sampling points on the circle,
number
[0016] Advantageous further embodiments of the invention are disclosed in the respective dependent claims. [Brief explanation of the drawings]
[0017] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] Desired panning function, loudspeaker positions φL=30°, φR=-30°. [Figure 2] Desired panning function in polar coordinate diagram, loudspeaker positions φL=30°, φR=-30°. [Figure 3] The resulting panning function for N=4, loudspeaker positions φL=30°, φR=-30°. [Figure 4] Resulting panning function for N=4 in polar coordinate diagram, loudspeaker positions φL=30°, φR=-30°. [Figure 5] FIG. 2 is a block diagram of a process according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] In the first stage of the decoding process, the loudspeaker positions need to be defined. They are assumed to have the same distance from the listening position, so the loudspeaker positions are defined by their azimuth angles. The azimuth angle is denoted by φ and is measured counterclockwise. The azimuth angle of the left and right loudspeakers is φ L and φ R and in a symmetric setup φ R =-φ L A typical value is φ L = 30°. In the following description, all angle values can be interpreted with an offset of 2π (radians) or an integer multiple of 360°.
[0019] Virtual sampling points on a circle should be defined. These are the directions of the virtual sources used in the Ambisonics decoding process, and for these directions the desired panning function values for, for example, two real loudspeaker positions are defined. The number of virtual sampling points is denoted by S, and the corresponding directions are evenly distributed around the circle. Thus,
number
[0020] The desired panning function g for the left and right loudspeakers L (φ) and g R(φ) needs to be defined. In contrast to the approaches of Patent Document 2 and Non-Patent Document 5, the panning function is defined for multiple segments, and different panning functions are used for those segments. For example, for the desired panning function, three segments are used: a) For the forward direction between the two loudspeakers, well-known panning laws are used, such as the tangent law or equivalently the vector-based amplitude panning (VBAP) as described in [6]. b) For directions beyond the loudspeaker circular section position, a slight attenuation in the rear direction is defined, so that this part of the panning function approaches a value of zero at an angle roughly opposite the loudspeaker position. c) The remainder of the desired panning function is set to 0 to prevent sounds from the right from being played on the left loudspeaker and sounds from the left from being played on the right loudspeaker.
[0021] The point or angle value at which the desired pan function approaches 0 is φ for the left loudspeaker. L,0 For the right loudspeaker, R,0 The desired panning functions for the left and right loudspeakers can be expressed as:
[0022]
number
[0023]
number
number
number
[0024]
number
[0025] Circular harmonic functions are combined into vectors.
[0026] y(φ)=[Y -N (φ),…,Y0(φ),…,Y N (φ)] T (11) (·) * The complex conjugate, denoted by, gives:
[0027] y * (φ)=[Y * -N (φ),…,Y * 0(φ),…,Y * N (φ)] T (12) The mode matrix for these virtual sampling points is Ξ=[y * (φ1),y * (φ2),…,y* (φ S )] (13) The resulting 2D decoding matrix is D=GΞ + (14) where Ξ + is the pseudo-inverse of the matrix Ξ. For uniformly distributed virtual sampling points as given in equation (1), the pseudo-inverse is Ξ H can be replaced by a scaled version of H is the adjoint (conjugate transpose) of Ξ. In this case, the decoding matrix is D=αGΞ H (15) where the scaling factor α depends on the normalization method of the circular harmonic function and the number of design directions S.
[0028] The vector l(t) representing the loudspeaker sample signal at time t is l(t)=Da(t) (16) It is calculated by:
[0029] When a three-dimensional higher-order Ambisonics signal a(t) is used as the input signal, an appropriate transformation to two-dimensional space is applied to give the transformed Ambisonics coefficients a'(t). In this case, equation (16) changes to l(t) = Da'(t).
[0030] The matrix D already contains the 3D / 2D transformation and is applied directly to the 3D Ambisonics signal a(t). 3D It is also possible to define
[0031] Below we describe an example of a panning function for a stereo loudspeaker setup. At the middle of the loudspeaker position, the panning function g from equations (2) and (3) L,1 (φ) and g R,1 Panning gains based on (φ) and VBAP are used. These panning functions are followed by half of a cardioid pattern with its maximum at the loudspeaker position. L,0 and φR,0 is defined to have a position opposite the loudspeaker position: φ L,0 =φ L +π (17) φ R,0 =φ R +π (18) The normalized panning gain is g L,1 (φ L )=1 and g R,1 (φ R )=1. L and φ R The cardioid pattern is g L,2 (φ)=(1 / 2)(1+cos(φ-φ L )) (19) g R,2 (φ)=(1 / 2)(1+cos(φ-φ R )) (20) is defined by
[0032] For decoding evaluation, the resulting panning function for any input direction is W=DΥ (21) where Υ is the mode matrix for the input direction under consideration, and W is a matrix containing the panning weights for the input direction and loudspeaker position used when applying the Ambisonics decoding process.
[0033] Figures 1 and 2 depict the desired (i.e., theoretical or perfect) panning functions, respectively, on a linear angular scale and in polar coordinate form. The resulting panning weights for Ambisonics decoding are calculated using equation (21) for the input directions used. Figures 3 and 4 depict the corresponding resulting panning functions, respectively, calculated for Ambisonics order N=4, on a linear angular scale and in polar coordinate form.
[0034] Comparing Figures 3 and 4 with Figures 1 and 2, it can be seen that the desired panning functions are well matched and the resulting negative sidelobes are very small.
[0035] Below, examples of 3D to 2D conversions are provided for complex-valued spherical and circular harmonics (real-valued basis functions can be performed in a similar manner). The spherical harmonics for 3D Ambisonics are
number
number
number
number
[0036] In FIG. 5, the step or stage 51 of calculating the desired panning function is the azimuth angle φ of the left and right loudspeakers. L and φ R and the number of virtual sampling points S, and then calculates—as described above—a matrix G containing the desired panning function values for all the virtual sampling points. From the Ambisonics signal a(t), the order N is derived in step / stage 52. From S and N, the mode matrix Ξ is calculated in step / stage 53 based on equations (11) to (13).
[0037] Step or stage 54 is the pseudo-inverse matrix Ξ of the matrix Ξ + Calculate the matrices G and Ξ. + , a decoding matrix D is calculated in step / stage 55 according to equation (15). In step / stage 56, the loudspeaker signal l(t) is calculated from the Ambisonics signal a(t) using the decoding matrix D. If the Ambisonics input signal a(t) is a three-dimensional spatial signal, a 3D to 2D conversion can be performed in step or stage 57, and step / stage 56 receives the 2D Ambisonics signal a'(t).
[0038] Several aspects will be described. [Aspect 1] 1. A method for decoding a stereo loudspeaker signal l(t) from a three-dimensional spatial higher-order Ambisonics audio signal a(t), the method comprising: calculating a matrix G containing desired panning functions for all virtual sampling points from the azimuth angle values of the left and right loudspeakers and from the number S of virtual sampling points on the circle,
number
number
number
Claims
1. 1. A method for decoding a Higher Order Ambisonics (HOA) audio signal, the method comprising: determining an order N of the HOA signal; determining a matrix G of panning function values, said matrix G being a function of a gain vector g for each of S virtual sampling points on the circle; 1 …g s wherein at least a first panning function value for a first virtual sampling point located opposite a loudspeaker position approaches zero and at least a second panning function value for a source located near said loudspeaker position has a value that does not approach zero; receiving a mode matrix, the mode matrix being determined based on a number S of virtual sampling points and an order N of the HOA audio signal; determining a decoding matrix D based on the matrix G and the mode matrix; and rendering, by at least one processor, the HOA audio signals into stereo loudspeaker signals based on the decoding matrix D. method.
2. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 1.
3. 1. An apparatus for decoding a Higher Order Ambisonics (HOA) audio signal, the apparatus comprising: a first receiver configured to receive the HOA audio signal; a first processor for determining an order N of the HOA signal; a second processor configured to determine a matrix G of panning function values, said matrix G comprising a gain vector g for each of S virtual sampling points on the circle; 1 …g s a second processor, wherein at least a first panning function value for a first virtual sampling point located opposite a loudspeaker position approaches zero and at least a second panning function value for a source located near said loudspeaker position has a value that does not approach zero; a second receiver for receiving a mode matrix, the mode matrix being determined based on a number S of virtual sampling points and an order N of the HOA audio signal; a third processor for determining a decoding matrix D based on the matrix G and the mode matrix; a renderer for rendering the HOA audio signal into a stereo loudspeaker signal based on the decoding matrix D. Device.
Citation Information
Patent Citations
Improvements in and relating to sound-transmission, sound-recording and sound-reproducing systems
GB394325A
Method and device for decoding an audio soundfield representation for audio playback
WO2011117399A1