Method and device for decoding an ambisonics audio soundfield representation for audio reproduction using a 2d setup

The method addresses sound reproduction issues in 2D setups by adding virtual loudspeakers and using an energy-conserving matrix design to enhance sound localization and timbre, ensuring accurate signal energy reproduction from all directions.

JP2026032132APending Publication Date: 2026-02-25DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025203357
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-10-23
Filing Date
2025-11-26
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing 2D loudspeaker setups fail to accurately reproduce sound sources from directions where no loudspeakers are located, leading to attenuated sound and uneven loudness in spatial panning, particularly for 2D setups with limited elevation angles.

Method used

A method is introduced to generate a decoding matrix for 2D loudspeaker setups by adding virtual loudspeakers at specific angles (e.g., +90° and -90°) and using an energy-conserving 3D matrix design, followed by downmixing and normalization to distribute coefficients, ensuring accurate signal energy reproduction from all directions.

Benefits of technology

The method enhances sound localization and timbre characteristics in 2D setups by reproducing sounds from directions without real loudspeakers with consistent energy and loudness, improving overall audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026032132000001_ABST
    Figure 2026032132000001_ABST
Patent Text Reader

Abstract

To provide a method and apparatus for decoding an Ambisonics audio soundfield representation for audio playback using a 2D setup.SOLUTION: For decoding, a decode matrix specific to the given loudspeaker setup is needed, which is generated using the known loudspeaker positions. A method for decoding an encoded audio signal in a sound field format for L loudspeakers at known positions comprises the steps of adding 10 at least one virtual loudspeaker position to the L loudspeaker positions, generating 11 a 3D decoding matrix D ', downmixing 12 said 3D decoding matrix D ', and decoding 14 the encoded audio signal 3D using the downscaled i14 decoding matrix to obtain a plurality of decoded loudspeaker signals q14.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an Ambisonics audio sound field representation for audio reproduction using a 2D or near-2D setup, in particular to a method and apparatus for decoding an audio representation in Ambisonics format. [Background technology]

[0002] Accurate localization is a primary goal of any spatial audio playback system. Such playback systems are highly practical for conferencing systems, games, or other virtual environments that benefit from 3D sound. A 3D sound scene can be synthesized or captured as a natural sound field. For example, an Ambisonics-like sound field signal carries a representation of the desired sound field. A decoding process is required to obtain individual speaker signals from the sound field representation. Decoding an Ambisonics-formatted signal is also called "rendering." Synthesizing an audio scene requires a panning function that references the spatial speaker placement to obtain the spatial localization of a given sound source. Recording a natural sound field requires a microphone array to capture the spatial information. The Ambisonics method is a well-suited tool for achieving this. An Ambisonics-formatted signal carries a representation of the desired sound field based on spherical harmonic decomposition of the sound field. The basic Ambisonics format, or B-Form, uses spherical harmonics of orders 0 and 1, while so-called Higher Order Ambisonics (HOA) also uses further spherical harmonics of order at least 2. The spatial arrangement of the loudspeakers is called the loudspeaker setup. For the decoding process, a decoding matrix (also called rendering matrix) is required, which is specific to a given loudspeaker setup and is generated using the known loudspeaker positions.

[0003] Commonly used speaker setups include stereo setups using two speakers, standard surround setups using five speakers, and extended surround setups using more than five speakers. However, while these setups are well-known, they are limited to two dimensions (2D) and, for example, do not reproduce height information. Renderings for known speaker setups that can reproduce height information have drawbacks in sound localization and timbre. These drawbacks include the perception of highly uneven loudness in spatially vertical panning or the speaker signals having strong sidelobes, which are particularly noticeable when listening from off-center positions. Therefore, when rendering HOA sound field descriptions for speakers, so-called energy-conserving rendering designs are preferred. This means that rendering a single sound source results in speaker signals with constant energy, independent of the source's direction. In other words, the input energy preserved by the Ambisonics representation is preserved by the speaker renderer. In our WO2014 / 012945 [Reference 1], we describe an HOA renderer design with good energy conservation and localization properties for 3D loudspeaker setups. However, while this approach works extremely well for omnidirectional 3D loudspeaker setups, some sound source directions are attenuated in 2D loudspeaker setups (e.g., 5.1 surround). This is especially true for directions from above where there are no loudspeakers.

[0004] In "All-Round Ambisonic Panning and Decoding" by F. Zotter and M. Frank [Reference 2], "fictitious" loudspeakers are added when there are holes in the convex hull constructed by the loudspeakers. However, the resulting signals for the fictitious loudspeakers are omitted from playback on the real loudspeakers. Therefore, source signals from those directions (i.e., directions where no real loudspeakers are located) are still attenuated. Furthermore, this document only discloses the use of fictitious loudspeakers in conjunction with VBAP (Vector-Based Amplitude Panning). Summary of the Invention

[0005] Therefore, the remaining challenge is to design an energy-conserving Ambisonics renderer for 2D (two-dimensional) loudspeaker setups, such that sound sources from directions where no loudspeakers are located are attenuated less or not at all. 2D loudspeaker setups can be classified as those in which the loudspeaker elevation angles are close to the horizontal plane within a small range (e.g., less than 10°).

[0006] This specification describes a solution for rendering / decoding Ambisonics-style sound field representations for regular or irregular spatial loudspeaker arrangements, which results in highly improved localization and timbre characteristics, is energy conserving, and also renders sounds from directions where no loudspeakers are available. Advantageously, sounds from directions where no loudspeakers are available are rendered with roughly similar energy and perceived loudness as if loudspeakers were available in each direction. Of course, accurate localization of these sound sources is not possible because no loudspeakers are available in those directions.

[0007] In particular, at least some described embodiments provide a novel method for obtaining a decoding matrix for decoding sound field data in HOA format. Because at least the HOA format describes a sound field that is not directly related to loudspeaker positions and the obtained loudspeaker signals are necessarily channel-based audio formats, decoding of HOA signals is always closely related to rendering of audio signals. In principle, the same applies to other audio sound field formats. Therefore, the present disclosure relates to both decoding and rendering of sound field-related audio formats. The terms decoding matrix and rendering matrix are used synonymously.

[0008] To obtain a decoding matrix for a given setup with good energy conservation properties, one or more virtual loudspeakers are added where no loudspeakers are available. For example, to obtain an improved decoding matrix for a 2D setup, two virtual loudspeakers are added at the top and bottom (which correspond to elevation angles of +90° and -90° for a 2D loudspeaker placed at approximately 0° elevation). A decoding matrix that satisfies the energy conservation properties is designed for this virtual 3D loudspeaker setup. Finally, the weighting coefficients from the decoding matrix for the virtual loudspeakers are mixed with a constant gain for the real loudspeakers of the 2D setup.

[0009] According to one embodiment, a decoding matrix (or rendering matrix) for rendering or decoding an Ambisonics-formatted audio signal for a given set of loudspeakers is generated by: generating a first pre-decoding matrix using modified loudspeaker positions using a conventional method, where the modified loudspeaker positions include the loudspeaker positions of the given set of loudspeakers and at least one additional virtual loudspeaker position; and downmixing the first pre-decoding matrix, where coefficients associated with the at least one additional virtual loudspeaker are removed and distributed to the loudspeaker-related coefficients of the given set of loudspeakers. In one embodiment, this is followed by a subsequent step of normalizing the decoding matrix. The resulting decoding matrix is ​​suitable for rendering or decoding Ambisonics signals for the given set of loudspeakers, and even sounds from positions where no loudspeakers are present are reproduced with accurate signal energy. This is due to the construction of an improved decoding matrix. Preferably, the first pre-decoding matrix is ​​energy conserving.

[0010] In one embodiment, the decoding matrix has L rows and O 3D The number of rows corresponds to the number of speakers in the 2D loudspeaker setup, and the number of columns corresponds to O 3D =(N+1) 2 Ambisonics coefficients O depending on HOA order N according to 3D Each of the coefficients of the decoding matrix for the 2D loudspeaker setup is the sum of at least a first intermediate coefficient and a second intermediate coefficient. The first intermediate coefficient is obtained by an energy-conserving 3D matrix design method for the current loudspeaker positions of the 2D loudspeaker setup, and the energy-conserving 3D matrix design method uses the position of at least one virtual loudspeaker. The second intermediate coefficient is obtained by multiplying the coefficient by a weighting factor g obtained from the energy-conserving 3D matrix design method for the at least one virtual loudspeaker position. In one embodiment, the weighting factor g is

number

[0011] In one embodiment, the present invention relates to a computer-readable medium having stored thereon executable instructions for causing a computer to perform a method including the steps set forth above or claimed in the claims. An apparatus for utilizing this method is disclosed in claim 9.

[0012] Advantageous embodiments are disclosed in the dependent claims, the following description and the drawings.

[0013] Exemplary embodiments of the present invention are described with reference to the accompanying drawings. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a flowchart of a method according to an embodiment. [Figure 2] FIG. 10 illustrates an exemplary configuration of a downmixed HOA decoding matrix. [Figure 3] 10 is a flowchart for obtaining and changing the position of a speaker. [Figure 4] 1 is a block diagram illustrating an apparatus according to an embodiment. [Figure 5] FIG. 1 illustrates the energy distribution resulting from a conventional decoding matrix. [Figure 6] FIG. 10 illustrates an energy distribution resulting from a decoding matrix according to an embodiment. [Figure 7] FIG. 1 illustrates the use of separately optimized decoding matrices for different frequency bands. DETAILED DESCRIPTION OF THE INVENTION

[0015] 1 shows a flow chart of a method for decoding an audio signal, in particular a sound field signal, according to an embodiment of the present invention. The decoding of a sound field signal generally requires the positions of the loudspeakers at which the audio signal is rendered. Such loudspeaker positions for the L loudspeakers are

number

number

number

[0016] The 3D decoding matrix step 11 implements any known method for generating a 3D decoding matrix. Preferably, the 3D decoding matrix is ​​suitable for energy-conserving decoding / rendering. For example, the method described in International Patent Application No. EP2013 / 065034 can be used. As a result of the 3D decoding matrix design step 11, L'=L+L virt A decoding or rendering matrix D' suitable for rendering the L loudspeaker signals is obtained, where L virt is the number of virtual speaker positions added in step 10 “Add virtual speaker positions.”

[0017] Since only L loudspeakers are physically available, the decoding matrix D' resulting from the 3D decoding matrix design step 11 needs to be adapted to the L loudspeakers in the downmixing step 12. In step 12, the decoding matrix D' is downmixed, in which coefficients associated with virtual loudspeakers are weighted and distributed to coefficients associated with existing loudspeakers. Preferably, coefficients of any particular HOA order (i.e., columns of the decoding matrix D') are weighted and added to coefficients of the same HOA order (i.e., the same columns of the decoding matrix D'). An example is downmixing according to equation (8) below. The result of the downmixing step 12 is a downmixed 3D decoding matrix D' having L rows, i.e., fewer rows than the decoding matrix D' but the same number of columns as the decoding matrix D'.

number

number

[0018] Figure 2 shows the downmixed HOA decoding matrix D'.

number

number

number

number

number

number

number

[0019] Usually, the downmixed HOA decoding matrix

number

number

number

[0020] The normalized downmixed HOA decoding matrix D is then used in a sound field decoding step 14, where the input sound field signal i14 is decoded into L loudspeaker signals q14. Typically, the normalized downmixed HOA decoding matrix D does not need to be changed unless the loudspeaker setup is changed. Therefore, in one embodiment, the normalized downmixed HOA decoding matrix D is stored in a decoding matrix storage.

[0021] Figure 3 shows details of how the speaker positions are obtained and modified in one embodiment. This embodiment uses L speaker positions.

number

number

[0022] In one embodiment, the at least one virtual location

number

number

number

[0023] In one embodiment, in step 103, two virtual positions corresponding to two virtual speakers are selected.

number

number

number

number

[0024] According to one embodiment, a method for decoding encoded audio signals for L loudspeakers at known positions includes: determining the positions of the L loudspeakers;

number

number

number

number

number

number

[0025] In one embodiment, the encoded audio signal is a sound field signal, for example a sound field signal in HOA format.

[0026] In one embodiment, at least one virtual position of said virtual speaker

number

number

number

[0027] In one embodiment, the coefficient for the position of the virtual speaker is a weighting coefficient

number

[0028] In one embodiment, the method comprises:

number

number

[0029] According to one embodiment, a decoding matrix for rendering or decoding a sound field signal for a given set of loudspeakers is generated by: generating a first pre-decoding matrix using modified loudspeaker positions using a conventional method, the modified loudspeaker positions including the loudspeaker positions of the given set of loudspeakers and at least one additional virtual loudspeaker; and downmixing the first pre-decoding matrix, excluding coefficients associated with the at least one additional virtual loudspeaker and allocating them to the coefficients associated with the loudspeakers of the given set of loudspeakers. In one embodiment, this is followed by the following step of normalizing the decoding matrix. The resulting decoding matrix is ​​suitable for rendering or decoding a sound field signal for the given set of loudspeakers, and even sounds from positions where no loudspeakers are present are reproduced with accurate signal energy. This is due to the construction of the improved decoding matrix. Preferably, the first pre-decoding matrix is ​​energy-conserving.

[0030] Fig. 4a) shows a block diagram of an apparatus 400 for decoding coded audio signals in sound field format for L loudspeakers at known positions, comprising an adding unit 410 for adding at least one position of at least one virtual loudspeaker to the positions of the L loudspeakers, and a decoding matrix generating unit 411 for generating a 3D decoding matrix D', where D' is the position of the L loudspeakers.

number

number

number

number

[0031] In one embodiment, the device comprises a downscaled 3D decoding matrix

number

[0032] In one embodiment shown in FIG. 4b, the device calculates the positions of L loudspeakers (Ω L ) and a first specifying unit 4101 that specifies an order N of the coefficients of the sound field signal, a second specifying unit 4102 that specifies that the L speakers are substantially on a 2D plane from the positions of the L speakers, and at least one virtual position of the virtual speaker.

number

[0033] In one embodiment, the apparatus includes a bandpass filter 715b that separates the encoded audio signal into multiple frequency bands, and multiple separated 3D decoding matrices Db' (one for each frequency band) are generated in 711b, and each 3D decoding matrix Db' is downmixed and may be separately normalized in 712b, with a decoder 714b decoding each frequency band separately. In this embodiment, the apparatus further includes multiple adders 716b, one for each speaker. Each adder sums the frequency bands associated with a respective speaker.

[0034] The functions of each of the adding unit 410, the decoding matrix generating unit 411, the matrix downmixing unit 412, the normalizing unit 413, the decoding unit 414, the first identifying unit 4101, the second identifying unit 4102, and the virtual speaker position generating unit 4103 are implemented by one or more processors, and each of these units may share the same processor with other units among these units or with other units other than these units.

[0035] FIG. 7 illustrates an embodiment using decoding matrices optimized separately for different frequency bands of an input signal. In this embodiment, the decoding method includes separating an encoded audio signal into multiple frequency bands using band-pass filters. At 711b, multiple separated 3D decoding matrices Db′ (one for each frequency band) are generated, and at 712b, each 3D decoding matrix Db′ is downmixed. Separate normalization may also be performed. At 714b, the encoded audio signal is decoded separately for each frequency band. This has the advantage of taking into account frequency-dependent differences in human perception, resulting in different decoding matrices for different frequency bands. In one embodiment, only one or several (but not all) of the decoding matrices are generated by adding virtual loudspeaker positions as described above, and then weighting and distributing the coefficients for each of the virtual loudspeaker positions with the coefficients for the existing loudspeaker positions. In another embodiment, each encoding matrix is ​​generated by adding virtual speaker positions as described above, then weighting and distributing the coefficients for each virtual speaker position to the coefficients for the existing speaker positions. Finally, in a process that reverses the frequency band division, one frequency band adder 716b sums all frequency bands associated with the same speaker for each speaker.

[0036] Each of the adding unit 410, the decoding matrix generating unit 711b, the matrix downmixing unit 712b, the normalizing unit 713b, the decoding unit 714b, the frequency band summing unit 716b, and the band pass filter unit 715b is implemented by one or more processors, and each of these functional units may share the same processor with other functional units among these functional units or with other functional units other than these functional units.

[0037] One aspect of the present disclosure is to obtain a rendering matrix for a 2D setup with good energy conservation properties. In one embodiment, two loudspeakers are added at the top and bottom (at elevation angles of +90° and -90° for 2D loudspeakers placed at approximately 0° elevation). A rendering matrix that satisfies the energy conservation properties is designed for this virtual 3D loudspeaker setup. Finally, the weighting coefficients from the rendering matrix for the virtual loudspeakers are mixed with constant gains for the real loudspeakers in the 2D setup.

[0038] In the following, we will explain the rendering of Ambisonics (particularly HOA).

[0039] Ambisonics rendering is the process of computing loudspeaker signals from an Ambisonics sound field description. It is sometimes called Ambisonics decoding. A 3D Ambisonics sound field representation of order N is considered, where the number of coefficients is given by equation (1) below: O 3D =(N+1) 2 (1)

[0040] The coefficients of this time sample t are O 3D a vector with elements

number

number

number

number

[0041] The position of the speakers is determined by the tilt angle θ l and azimuth angle φ l These tilt angles θ l and azimuth angle φ l Combine vectors

number

[0042] The signal energy in the HOA region is given by equation (3) below. E=b H b (3) where: H represents the complex conjugate transpose. The corresponding energy of the loudspeaker signal is calculated by the following equation (4):

number

[0043] To achieve energy-conserving decoding / rendering, the ratio of energy-conserving decoding / rendering matrices is

number

[0044] In principle, the following extension for improved 2D rendering is proposed: To design a rendering matrix for a 2D loudspeaker setup, one or more virtual loudspeakers are added. The 2D setup is considered as a situation where the elevation angles of the loudspeakers are within a small predetermined range and close to the horizontal plane. This can be expressed as Equation (5) below.

number

[0045] Usually, the threshold θ thres2d is chosen in one embodiment to correspond to a value in the range 5° to 10°.

[0046] For the rendering design, the speaker angle of the modified pair is shown.

number

number

[0047] Then the new number of speakers used for the rendering design is L' = L + 2. From these modified speaker positions, we use an energy-conserving approach to calculate the rendering matrix

number

number

[0048] intermediate matrix

number

number

number

number

number

[0049] Figures 5 and 6 show the energy distribution for a 5.0 surround speaker setup. In both figures, the energy values ​​are shown as a grayscale, and the circles indicate the speaker locations. It is clear that the attenuation is reduced, especially at the top (and also at the bottom, not shown here), using the disclosed method.

[0050] Figure 5 shows the resulting energy distribution from a conventional decoding matrix. The small circles around the z=0 plane represent the loudspeaker locations. We can see that the energy range of [-3.9,...,2.1] decibels (dB) is covered, resulting in an energy difference of 6 dB. Furthermore, signals from the top of the unit sphere (and also signals on the bottom, not shown) are reproduced with very low energy, i.e., inaudible, since no loudspeakers are available here.

[0051] FIG. 6 shows the energy distribution resulting from a decoding matrix according to one or more embodiments. The same number of loudspeakers are present in the same locations as in FIG. 5. At least the following advantages are achieved: First, a smaller energy range of [-1.6, . . . , 0.8] decibels (dB) is covered, resulting in a smaller energy difference of only 2.4 dB. Second, signals from all directions on the unit sphere are reproduced with their correct energy, even if no loudspeakers are present. Because these signals are reproduced through the available loudspeakers, their localization is not accurate. However, the signals are audible with the correct loudness. In this example, the signals from the top and the signals on the bottom (not shown) become audible after decoding using the improved decoding matrix.

[0052] In one embodiment, a method for decoding an Ambisonics-format encoded audio signal for L loudspeakers at known positions includes the steps of adding at least one position of at least one virtual loudspeaker to the positions of the L loudspeakers, and generating a 3D decoding matrix D', where D' is the position of the L loudspeakers.

number

number

number

number

[0053] In another embodiment, an apparatus for decoding Ambisonics-format encoded audio signals for L loudspeakers at known positions includes an adding unit 410 for adding at least one position of at least one virtual loudspeaker to the positions of the L loudspeakers, and a decoding matrix generating unit 411 for generating a 3D decoding matrix D', the decoding matrix D' being added to the positions of the L loudspeakers.

number

number

number

number

[0054] In yet another embodiment, an apparatus for decoding Ambisonics-format encoded audio signals for L speakers at known positions includes at least one processor and at least one memory, the memory storing instructions that, when executed on the processor, cause the processor to include an adding unit 410 for adding at least one position of at least one virtual speaker to the positions of the L speakers; and a decoding matrix generating unit 411 for generating a 3D decoding matrix D′, the decoding matrix D′ being added to the positions of the L speakers.

number

number

number

number

[0055] In yet another embodiment, a computer-readable storage medium stores executable instructions for causing a computer to perform a method for decoding Ambisonics-format encoded audio signals for L speakers at known positions, the method including adding at least one position of at least one virtual speaker to the positions of the L speakers; and generating a 3D decoding matrix D′, wherein the positions of the L speakers are added to the positions of the L speakers.

number

number

number

number

[0056] The present invention has been described purely for illustrative purposes and modifications of detail may be made without departing from the scope of the invention, for example, although described solely in relation to HOA, the invention may also be applied to other sound field audio formats.

[0057] Each element disclosed in the specification, claims (if applicable), and drawings may be provided independently or in any suitable combination. Elements may be implemented in hardware, software, or a combination of both hardware and software, as appropriate. Reference signs in the claims are for illustrative purposes only and shall have no limiting effect on the scope of the claims.

[0058] The references cited are as follows: [Reference 1] International Patent Publication No. 2014 / 012945 (PD120032) [Reference 2] F. Zotter and M. Frank, "All-Round Ambisonic Panning and Decoding," Journal of the Society of Audio Engineers, 2012, Vol. 60, pp. 807-820

[0059] Several aspects will be described. [Aspect 1] 1. A method for decoding Ambisonics encoded audio signals for L loudspeakers at known positions, comprising: - adding (10) at least one position of at least one virtual loudspeaker to the positions of said L loudspeakers; a step (11) of generating a 3D decoding matrix (D'),

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

Claims

1. 1. A method for rendering an Ambisonics audio signal, the method comprising: One or more virtual speaker positions for a set of L speaker positions [Number 104] Add L 2 determining a new set of speaker positions; Said L 2 determining a first decoding matrix for the new set of speaker positions; determining a second decoding matrix for the set of L speaker positions, the second decoding matrix being determined based on at least one coefficient of the first decoding matrix, the second decoding matrix further comprising a coefficient for each of the one or more virtual speaker positions; [Number 105] and weighting and distributing at least one coefficient for determining a rendering matrix based on a normalization of the second decoding matrix, the normalization being based on a Frobenius norm; rendering the Ambisonics audio signal based on the rendering matrix; Including, method.

2. A non-transitory computer-readable storage medium storing executable instructions for causing a computer to perform the method of claim 1.

3. 1. An apparatus for rendering an Ambisonics audio signal, comprising: One or more virtual speaker positions for a set of L speaker positions [Number 106] Add L 2 a first processor that determines a new set of speaker positions; Said L 2 a second processor that determines a first decoding matrix for the new set of speaker positions; a third processor for determining a second decoding matrix for the set of L speaker positions, the second decoding matrix being determined based on at least one coefficient of the first decoding matrix, the second decoding matrix further comprising: [Number 107] and a third processor further based on weighting and distributing at least one coefficient for a fourth processor for determining a rendering matrix based on a normalization of the second decoding matrix, the normalization being based on a Frobenius norm; a fifth processor for rendering the Ambisonics audio signal based on the rendering matrix; and An apparatus having:

Citation Information

Patent Citations

  • Audio signal, method and apparatus for encoding or transmitting the same and method and apparatus for processing the same

    EP2094032A1

  • Method and apparatus for decoding audio sound field representation for audio playback.

    JP2013524564A

  • Method and apparatus for decoding an Ambisonics audio sound field representation for audio reproduction using a 2D setup

    JP6660493B2

  • Enabling 3D sound reproduction using a 2d speaker arrangement

    US20110216906A1

  • Method and apparatus for decoding stereo loudspeaker signals from a higher-order ambisonics audio signal

    WO2013143934A1