Method and apparatus for decoding a stereo loudspeaker signal from a high-order Ambisonics audio signal

The method addresses poor localization and sidelobe issues in first-order Ambisonics by using higher-order Ambisonics with defined panning functions and pseudo-inverse matrices, improving spatial resolution and reducing noise in stereo output.

JP7739546B2Active Publication Date: 2025-09-16DOLBY INTERNATIONAL AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024117388
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-03-28
Filing Date
2024-07-23
Publication Date
2025-09-16
Estimated Expiration
2033-03-20

AI Technical Summary

Technical Problem

First-order Ambisonics decoders based on Blumlein Stereo result in high negative sidelobes or poor localization, particularly in the front direction, leading to issues like sound objects being reproduced incorrectly on the wrong stereo loudspeaker.

Method used

A method and apparatus for decoding a stereo loudspeaker signal from higher-order Ambisonics audio signals using panning functions defined on a circle, employing circular harmonic functions and pseudo-inverse matrices to improve localization and reduce negative sidelobes, with specific panning laws applied for different regions around the loudspeakers.

Benefits of technology

Enhances spatial resolution in the front region, reduces negative sidelobes, and provides improved direct-to-diffuse sound ratio, enhancing dialogue intelligibility and reducing noise in the stereo output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739546000017
    Figure 0007739546000017
  • Figure 0007739546000018
    Figure 0007739546000018
  • Figure 0007739546000019
    Figure 0007739546000019
Patent Text Reader

Abstract

To provide a method and an apparatus for decoding stereo loudspeaker signals from a higher-order Ambisonics audio signal.SOLUTION: A method includes: in Step 51 of calculating a desired panning function, receiving the values of azimuth angles φL and φR of left and right loudspeakers as well as a number S of virtual sampling points, and calculating a matrix G containing desired panning function values for all virtual sampling points; deriving, from an Ambisonics signal a(t), an order N in Step 52; calculating, from the number S and the order N, a mode matrix Ξ in Step 53; computing a pseudo-inverse matrix Ξ+ of the matrix Ξ in Step 54; calculating, from matrices G and Ξ+, a decoding matrix D in Step 55; and calculating loudspeaker signals l(t) from the Ambisonics signal a(t) using the decoding matrix D in Step 56.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for decoding a stereo loudspeaker signal from a high-order Ambisonics audio signal using a panning function for sampling points on a circle. [Background technology]

[0002] Decoding of Ambisonics representations for stereo loudspeaker or headphone setups is known for first order Ambisonics, e.g. from equation (10) in [1], and from [2]. These approaches are based on Blumlein stereo, as disclosed in [1]. Another approach uses mode matching: [3]. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] British Patent No. 394325 [Patent Document 2] International Publication No. 2011 / 117399 [Non-patent literature]

[0004] [Non-Patent Document 1] J.S. Bamford, J. Vender-kooy, "Ambisonic sound for us", Audio Engineering Society Preprints, Convention paper 4138, presented at the 99th Convention, October 1995, New York [Non-patent document 2] XiphWiki-Ambisonics http: / / wiki.xiph.org / index.php / Ambisonics#Default_channel_conversions_from_B-Format [Non-patent document 3] MA Poletti, "Three-Dimensional Surround Sound Systems Based on Spherical Harmonics", J. Audio Eng. Soc., vol.53(11), pp.1004-1025, November 2005 [Non-patent document 4] S. Weinzierl, "Handbuch der Audiotechnik", Springer, Berlin, 2008, section 3.3.4.1 [Non-Patent Document 5] JM Batke, F. Keiler, "Using VBAP-derived panning functions for 3D Ambisonics decoding", Proc. of the 2nd International Symposium on Ambisonics and Spherical Acoustics, May 6-7 2010, Paris, France, URL http: / / ambisonics10.ircam.fr / drupal / files / proceedings / presentations / O14_47.pdf [Non-patent document 6] V. Pulkki, "Virtual sound source positioning using vector base amplitude panning", J. Audio Eng. Society, 45(6), pp.456-466, June 1997 [Non-Patent Document 7] Earl G. Williams, "Fourier Acoustics", vol.93 of Applied Mathematical Sciences, Academic Press, 1999 Summary of the Invention [Problem to be solved by the invention]

[0005] Such first-order Ambisonics approaches, like Ambisonics decoders based on Blumlein Stereo (Patent Document 1) with virtual microphones with a figure-of-eight pattern, either have high negative sidelobes or result in poor localization in the front direction, where, for example, a sound object from the rear right direction is reproduced on the left stereo loudspeaker.

[0006] The problem to be solved by the present invention is to provide an Ambisonics signal decoding with an improved stereo signal output. [Means for solving the problem]

[0007] This problem is solved by the methods disclosed in claims 1 and 2. An apparatus utilizing these methods is disclosed in claim 3.

[0008] This invention describes a process for a stereo decoder for higher-order Ambisonics (HOA) audio signals. Desired panning functions can be derived from panning laws for the placement of virtual sources among the loudspeakers. For each loudspeaker, desired panning functions for all possible input directions are defined. The Ambisonics decoding matrix is ​​calculated similarly to the corresponding descriptions in [5] and [2]. The panning functions are approximated by circular harmonic functions; the approximation matches the desired panning function with less error as the Ambisonics order increases. For areas in front of the loudspeakers, particularly those between the loudspeakers, panning laws such as tangent law or vector base amplitude panning (VBAP) can be used. For directions behind the loudspeakers beyond the loudspeaker positions, panning functions with a slight attenuation of sounds from these directions are used.

[0009] A special case is to use half of the cardioid pattern for the rear direction, pointing towards the loudspeaker.

[0010] In this invention, the higher spatial resolution of higher-order Ambisonics is exploited, especially in the front region, and the attenuation of negative sidelobes in the rear direction increases with increasing Ambisonics order. The invention can also be used for loudspeaker setups with three or more loudspeakers arranged on a semicircle or a smaller arc (circle segment). The invention also facilitates more artistic downmixing to stereo, where some spatial regions receive greater attenuation. This is beneficial for producing an improved direct-to-diffuse sound ratio, which can improve dialogue intelligibility.

[0011] The stereo decoder according to the invention has several important attributes: good localization in the front direction between the loudspeakers, only small negative sidelobes in the resulting panning function and slight attenuation in the rear direction, and also allows attenuation or masking of spatial regions that would otherwise be perceived as noisy or annoying when listening to the two-channel version.

[0012] In comparison to Patent Document 2, the desired panning function is defined arc by arc, allowing well-known panning (e.g., VBAP or tangent law) to be used in the front region midway between loudspeaker positions, while the rear direction can be slightly attenuated, an attribute that is not achievable when using a first-order Ambisonics decoder.

[0013] In principle, the method of the invention is suitable for decoding a stereo loudspeaker signal l(t) from a high-order Ambisonics audio signal a(t), said method comprising: calculating a matrix G containing desired panning functions for all virtual sampling points from the azimuth angle values ​​of the left and right loudspeakers and from the number S of virtual sampling points on the circle,

number

[0014] In principle, the method of the invention is suitable for determining a decoding matrix D that can be used to decode a stereo loudspeaker signal l(t)=Da(t) from a 2D higher-order Ambisonics audio signal a(t), said method comprising: receiving an order N of the Ambisonics audio signal a(t); The desired azimuth angle values ​​of the left and right loudspeakers (φ L ,φ R ) and from the number S of virtual sampling points on the circle, calculating a matrix G containing desired panning functions for all virtual sampling points,

number

[0015] In principle, the device of the invention is suitable for decoding a stereo loudspeaker signal l(t) from a high-order Ambisonics audio signal a(t), said device comprising: means adapted to calculate a matrix G containing desired panning functions for all virtual sampling points from the azimuth angle values ​​of the left and right loudspeakers and from the number S of virtual sampling points on the circle,

number

[0016] Advantageous further embodiments of the invention are disclosed in the respective dependent claims. [Brief explanation of the drawings]

[0017] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. [Figure 1] Desired panning function, loudspeaker positions φL=30°, φR=-30°. [Figure 2] Desired panning function in polar coordinate diagram, loudspeaker positions φL=30°, φR=-30°. [Figure 3] The resulting panning function for N=4, loudspeaker positions φL=30°, φR=-30°. [Figure 4] Resulting panning function for N=4 in polar coordinate diagram, loudspeaker positions φL=30°, φR=-30°. [Figure 5] FIG. 2 is a block diagram of a process according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] In the first stage of the decoding process, the loudspeaker positions need to be defined. They are assumed to have the same distance from the listening position, so the loudspeaker positions are defined by their azimuth angles. The azimuth angle is denoted by φ and is measured counterclockwise. The azimuth angle of the left and right loudspeakers is φ L and φ R and in a symmetric setup φ R =-φ L A typical value is φ L = 30°. In the following description, all angle values ​​can be interpreted with an offset of 2π (radians) or an integer multiple of 360°.

[0019] Virtual sampling points on a circle should be defined. These are the directions of the virtual sources used in the Ambisonics decoding process, and for these directions the desired panning function values ​​for, for example, two real loudspeaker positions are defined. The number of virtual sampling points is denoted by S, and the corresponding directions are evenly distributed around the circle. Thus,

number

[0020] The desired panning function g for the left and right loudspeakers L (φ) and g R(φ) needs to be defined. In contrast to the approaches of Patent Document 2 and Non-Patent Document 5, the panning function is defined for multiple segments, and different panning functions are used for those segments. For example, for the desired panning function, three segments are used: a) For the forward direction between the two loudspeakers, well-known panning laws are used, such as the tangent law or equivalently the vector-based amplitude panning (VBAP) as described in [6]. b) For directions beyond the loudspeaker circular section position, a slight attenuation in the rear direction is defined, so that this part of the panning function approaches a value of zero at an angle roughly opposite the loudspeaker position. c) The remainder of the desired panning function is set to 0 to prevent sounds from the right from being played on the left loudspeaker and sounds from the left from being played on the right loudspeaker.

[0021] The point or angle value at which the desired pan function approaches 0 is φ for the left loudspeaker. L,0 For the right loudspeaker, R,0 The desired panning functions for the left and right loudspeakers can be expressed as:

[0022]

number

[0023]

number

number

number

[0024]

number

[0025] Circular harmonic functions are combined into vectors.

[0026] y(φ)=[Y -N (φ),…,Y0(φ),…,Y N (φ)] T (11) (·) * The complex conjugate, denoted by, gives:

[0027] y * (φ)=[Y * -N (φ),…,Y * 0(φ),…,Y * N (φ)] T (12) The mode matrix for these virtual sampling points is Ξ=[y * (φ1),y * (φ2),…,y* (φ S )] (13) The resulting 2D decoding matrix is D=GΞ + (14) where Ξ + is the pseudo-inverse of the matrix Ξ. For uniformly distributed virtual sampling points as given in equation (1), the pseudo-inverse is Ξ H can be replaced by a scaled version of H is the adjoint (conjugate transpose) of Ξ. In this case, the decoding matrix is D=αGΞ H (15) where the scaling factor α depends on the normalization method of the circular harmonic function and the number of design directions S.

[0028] The vector l(t) representing the loudspeaker sample signal at time t is l(t)=Da(t) (16) It is calculated by:

[0029] When a three-dimensional higher-order Ambisonics signal a(t) is used as the input signal, an appropriate transformation to two-dimensional space is applied to give the transformed Ambisonics coefficients a'(t). In this case, equation (16) changes to l(t) = Da'(t).

[0030] The matrix D already contains the 3D / 2D transformation and is applied directly to the 3D Ambisonics signal a(t). 3D It is also possible to define

[0031] Below we describe an example of a panning function for a stereo loudspeaker setup. At the middle of the loudspeaker position, the panning function g from equations (2) and (3) L,1 (φ) and g R,1 Panning gains based on (φ) and VBAP are used. These panning functions are followed by half of a cardioid pattern with its maximum at the loudspeaker position. L,0 and φR,0 is defined to have a position opposite the loudspeaker position: φ L,0 =φ L +π (17) φ R,0 =φ R +π (18) The normalized panning gain is g L,1 (φ L )=1 and g R,1 (φ R )=1. L and φ R The cardioid pattern is g L,2 (φ)=(1 / 2)(1+cos(φ-φ L )) (19) g R,2 (φ)=(1 / 2)(1+cos(φ-φ R )) (20) is defined by

[0032] For decoding evaluation, the resulting panning function for any input direction is W=DΥ (21) where Y is the mode matrix for the input direction under consideration, and W is a matrix containing the panning weights for the input direction and loudspeaker position used when applying the Ambisonics decoding process.

[0033] Figures 1 and 2 depict the desired (i.e., theoretical or perfect) panning functions, respectively, on a linear angular scale and in polar coordinate form. The resulting panning weights for Ambisonics decoding are calculated using equation (21) for the input directions used. Figures 3 and 4 depict the corresponding resulting panning functions, respectively, calculated for Ambisonics order N=4, on a linear angular scale and in polar coordinate form.

[0034] Comparing Figures 3 and 4 with Figures 1 and 2, it can be seen that the desired panning functions are well matched and the resulting negative sidelobes are very small.

[0035] Below, examples of 3D to 2D conversions are provided for complex-valued spherical and circular harmonics (real-valued basis functions can be performed in a similar manner). The spherical harmonics for 3D Ambisonics are

number

number

number

number

[0036] In FIG. 5, the step or stage 51 of calculating the desired panning function is the azimuth angle φ of the left and right loudspeakers. L and φ R and the number of virtual sampling points S, and then calculates—as described above—a matrix G containing the desired panning function values ​​for all the virtual sampling points. From the Ambisonics signal a(t), the order N is derived in step / stage 52. From S and N, the mode matrix Ξ is calculated in step / stage 53 based on equations (11) to (13).

[0037] Step or stage 54 is the pseudo-inverse matrix Ξ of the matrix Ξ + Calculate the matrices G and Ξ. + , a decoding matrix D is calculated in step / stage 55 according to equation (15). In step / stage 56, the loudspeaker signal l(t) is calculated from the Ambisonics signal a(t) using the decoding matrix D. If the Ambisonics input signal a(t) is a three-dimensional spatial signal, a 3D to 2D conversion can be performed in step or stage 57, and step / stage 56 receives the 2D Ambisonics signal a'(t).

[0038] Several aspects will be described. [Aspect 1] 1. A method for decoding a stereo loudspeaker signal l(t) from a three-dimensional spatial higher-order Ambisonics audio signal a(t), the method comprising: calculating a matrix G containing desired panning functions for all virtual sampling points from the azimuth angle values ​​of the left and right loudspeakers and from the number S of virtual sampling points on the circle,

number

number

number

Claims

1. 1. A method for decoding a Higher Order Ambisonics (HOA) audio signal, the method comprising: receiving the HOA audio signal; determining a matrix G of panning function values, said matrix G being a function of a gain vector g for each of S virtual sampling points on the circle; 1 …g s wherein at least a first panning function value for a first virtual sampling point located opposite a loudspeaker position approaches zero, and at least a second panning function value for a source located near the loudspeaker position has a value that does not approach zero; and determining a decoding matrix D based on the matrix G and a mode matrix, the mode matrix being determined based on a number S of virtual sampling points and an order of the HOA audio signal; applying a 3D to 2D transformation based on the matrix D to determine a 2D rendering matrix; and rendering, by at least one processor, the HOA audio signal into a stereo loudspeaker signal based on the 2D rendering matrix. method.

2. The method of claim 1 , wherein the matrix G has size L×S, where L corresponds to a number of loudspeakers.

3. The gain vector g 1 …g s is adapted to achieve a panned mix in S directions of L loudspeakers.

4. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 1.

5. 1. An apparatus for decoding a Higher Order Ambisonics (HOA) audio signal, the apparatus comprising: a first receiver configured to receive the HOA audio signal; a first processor configured to determine a matrix G of panning function values, said matrix G comprising a gain vector g for each of S virtual sampling points on the circle; 1 …g s a first processor, wherein at least a first panning function value for a first virtual sampling point located opposite a loudspeaker position approaches zero, and at least a second panning function value for a source located near the loudspeaker position has a value that does not approach zero; a second processor for determining a decoding matrix D based on the matrix G and a mode matrix, the mode matrix being determined based on a number S of virtual sampling points and an order of the HOA audio signal; a third processor for determining a 3D / 2D matrix by applying a 3D to 2D transformation based on the matrix D to determine a 2D rendering matrix; a renderer for rendering the HOA audio signal into a stereo loudspeaker signal based on the 2D rendering matrix. Device.

Citation Information

Patent Citations

  • Audio signal, method and apparatus for encoding or transmitting the same and method and apparatus for processing the same

    EP2094032A1

  • Improvements in and relating to sound-transmission, sound-recording and sound-reproducing systems

    GB394325A

  • Method and device for decoding an audio soundfield representation for audio playback

    WO2011117399A1