Method and apparatus for rendering an ambisonics audio signal

JP7793803B2Active Publication Date: 2026-01-05DOLBY LABORATORIES LICENSING CORP +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024546005
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-13
Filing Date
2023-02-03
Publication Date
2026-01-05
Estimated Expiration
2043-02-03

Smart Images

  • Figure 0007793803000009
    Figure 0007793803000009
  • Figure 0007793803000010
    Figure 0007793803000010
  • Figure 0007793803000011
    Figure 0007793803000011
Patent Text Reader

Abstract

The present document describes a method (400) for rendering an Ambisonics signal using a loudspeaker arrangement including S loudspeakers. The method (400) includes converting (401) a set of N Ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), where N>1 and S>1. The method (400) further includes performing near-field compensation (referred to as NFC) filtering (402) of M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using the corresponding S loudspeakers.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Patent Application No. 63 / 330,687 (Docket No. D21151USP1) filed April 13, 2022, European Patent Application No. 22168180.2 (Docket No. D21151EP) filed April 13, 2022, and Indian Patent Application No. 202241005922 (Docket No. D21151IN) filed February 03, 2022.

[0002] Technical Field This paper concerns the efficient rendering of Ambisonics audio signals. [Background technology]

[0003] The sound or sound field in a listening environment of a listener located at a listening position can be described using Ambisonics (audio) signals. Ambisonics signals may be considered multi-channel audio signals, with each channel corresponding to a specific directional pattern of the sound field at the listener's listening position. Ambisonics signals may be described using a three-dimensional (3D) Cartesian coordinate system, with the origin of the coordinate system corresponding to the listening position, the x-axis pointing forward, the y-axis pointing left, and the z-axis pointing up.

[0004] By increasing the number of audio signals or channels and the number of corresponding directional patterns (and corresponding panning functions), the precision with which the sound field is described can be increased. As an example, a first-order Ambisonics signal has N=4 channels or waveforms: a W channel representing the omnidirectional component of the sound field, an X channel describing the sound field with a dipole directional pattern corresponding to the x-axis, a Y channel describing the sound field with a dipole directional pattern corresponding to the y-axis, and a Z channel describing the sound field with a dipole directional pattern corresponding to the z-axis. A second-order Ambisonics signal has N=9 channels, including the four channels of a first-order Ambisonics signal (also called B-format) plus five additional channels for different directional patterns. Generally, an nth-order Ambisonics signal has N=(n+1) 2 channels, which is the (n-1)th order Ambisonics signal plus (n+1) for additional directional patterns (when using the 3D Ambisonics format). 2 -n 2 ] additional channels. An nth-order Ambisonics signal, for n>1, is sometimes called a higher order Ambisonics (HOA) signal.

[0005] The HOA signal may be used to describe a 3D sound field independent of the arrangement of speakers used to render the HOA signal. Example arrangements of speakers include headphones, or one or more arrangements of speakers, or a virtual reality rendering environment. Thus, it may be beneficial to provide the HOA signal to an audio renderer to allow the audio renderer to flexibly adapt the rendering of the HOA signal to different arrangements of speakers. Summary of the Invention [Problem to be solved by the invention]

[0006] This paper addresses the technical problem of rendering Ambisonics audio signals in an efficient manner. This technical problem is solved by the independent claims. Preferred examples are set out in the dependent claims. [Means for solving the problem]

[0007] According to a first aspect, a method for rendering Ambisonics signals using a loudspeaker arrangement having S loudspeakers is described. The method includes converting a set of N Ambisonics channel signals into a set of unfiltered pre-rendered signals, where N>1 and S>1, and N may be different from S. The method further includes performing near-field compensation (NFC) filtering of M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals to provide a set of S filtered loudspeaker channel signals for rendering using the corresponding S loudspeakers. M may depend on the number of loudspeakers, S, and the order n of a Higher-Order Ambisonics (HOA) signal corresponding to the set of N Ambisonics channel signals. More specifically, M=S*(n+1-m1), where m1 is the number of Ambisonics channel modes from the order n of the HOA signal filtered by the all-pass NFC filter. If no all-pass NFC filter is used, m1=0. The set of N Ambisonics channel signals may additionally or alternatively be Higher Order Ambisonics (HOA) signals. 2 where n is the order of the HOA signal, and n>1. Furthermore, S may be greater than 2. In specific examples, S may be equal to 6 or 16.

[0008] In some embodiments, performing NFC filtering of M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals to provide a set of S filtered loudspeaker channel signals for rendering using a corresponding S loudspeakers includes determining a set of M filtered pre-rendered signals from the M unfiltered pre-rendered signals based on NFC coefficients. Specifically, this determining may include multiplying each of S unfiltered pre-rendered signals of the M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals by an NFC coefficient d(m) for each m in the frequency domain, where 0≦m≦n. Alternatively, this determining may include multiplying each of S unfiltered pre-rendered signals of the M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals by an NFC filter d(m) for each m in the frequency domain, where 0≦m≦n. m The method may include convolution in the time domain with m*m, where 0≦m≦n. The method may further include summing the filtered pre-rendered signals and the remaining unfiltered signals corresponding to the loudspeakers to provide S filtered loudspeaker channel signals for each loudspeaker S. The number of filtered pre-rendered signals may be n+1−m1. Note that in the end, the summation includes m1 unfiltered pre-rendered signals corresponding to the loudspeakers. This is because the remaining S*m1 unfiltered pre-rendered signals still need to be made available and taken into account (although no filtering is actually applied due to their all-pass characteristics) to properly obtain the final S filtered loudspeaker channel signals.

[0009] In some embodiments, the method may further include determining whether (N-m0) is less than M, where m0≧0 and depends on the number of Ambisonics channel modes from order n of the HOA signal filtered by the all-pass NFC filter. m0 may be the number of Ambisonics channel indices corresponding to the number of Ambisonics channel modes filtered by the all-pass NFC filter. If no all-pass NFC filter is used, m0=0. Further, performing near-field compensation (referred to as NFC) filtering of M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals to provide a set of S filtered loudspeaker channel signals for rendering with the corresponding S loudspeakers includes performing NFC filtering on the M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals if, and particularly only if, (N-m0)>M or (N-m0)≧M.

[0010] In some embodiments, before transforming the set of N Ambisonics channel signals into a set of unfiltered pre-rendered signals, the method may further include determining whether (N-m0) is less than M, where m0≧0 and depends on the number of Ambisonics channel modes from order n of the HOA signal filtered by the all-pass NFC filter. m0 may be the number of Ambisonics channel indices corresponding to the number of Ambisonics channel modes filtered by the all-pass NFC filter. If an all-pass NFC filter is not used, m0=0. Depending on whether (N-m0) is less than M, NFC filtering may be performed on M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals, or the order of the transformation and NFC filtering may be reversed. Specifically, if (N-m0)>M, NFC filtering may be performed on M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals. Otherwise, the order of the transformation and NFC filtering may be reversed. Reversing the order may be understood as performing NFC filtering on a set of N Ambisonics channel signals to generate a set of N filtered Ambisonics channel signals, and using that set of N filtered Ambisonics channel signals to generate a set of S filtered loudspeaker channel signals. In other words, NFC filtering, if necessary, is performed on the HOA signals, and transformation is performed on the already filtered HOA signals.

[0011] In some embodiments, performing NFC filtering on M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals may include performing time-domain filtering using a digital finite impulse response filter or a digital infinite impulse response filter on each one of the M unfiltered pre-rendered signals individually.

[0012] In some embodiments, the method may further include determining a reference distance for the set of N Ambisonics channel signals, particularly based on the bitstream of the Ambisonics signals, and further determining a filter, particularly filter coefficients, for performing NFC filtering based on the reference distance.

[0013] In some embodiments, converting the set of N Ambisonics channel signals into a set of M unfiltered pre-rendered signals includes, for a given loudspeaker k, converting an Ambisonics signal matrix C representing a frame of the set of N Ambisonics channel signals into a set of M unfiltered pre-rendered signals, and a renderer matrix R k The renderer matrix R for a given loudspeaker k may be k Specifically, is an (n+1) × N matrix.

[0014] According to a second aspect, a rendering device for rendering an Ambisonics signal using a loudspeaker arrangement comprising S loudspeakers is described, the rendering device being configured to perform the method of the first aspect and any optional embodiment referred to in the first aspect.

[0015] According to a third aspect, a software program is described, which may be adapted for execution on a processor and for performing the method of the first aspect and any optional embodiments referred to in the first aspect.

[0016] According to a fourth aspect, a storage medium is described, which may comprise a software program adapted for execution on a processor and for implementing the method of the first aspect and any optional embodiments referred to in the first aspect.

[0017] According to a fifth aspect, a computer program product is described, which may include executable instructions for carrying out the method of the first aspect and any optional embodiments referred to in the first aspect.

[0018] According to a sixth aspect, a decoder configured to decode a bitstream representative of an Ambisonics signal to be rendered by a loudspeaker arrangement comprising S loudspeakers is described, the decoder comprising a rendering device according to the second aspect.

[0019] It should be noted that the methods, devices, and systems, including preferred embodiments thereof, as outlined in this patent application may be used independently or in combination with other methods, devices, and systems disclosed herein. Furthermore, all aspects of the methods, devices, and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief explanation of the drawings]

[0020] The invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0021] [Figure 1] FIG. 1 illustrates an exemplary rendering device for rendering an Ambisonics audio signal.

[0022] [Figure 2a] An example of a rendering device with modified NFC processing is shown.

[0023] [Figure 2b] 1 shows an exemplary rendering device with flexible NFC processing before or after "Ambisonics-to-Loudspeaker" conversion.

[0024] [Figure 3] 1 shows a flowchart of an exemplary method for rendering an Ambisonics audio signal.

[0025] [Figure 4] 1 shows a flowchart of an exemplary method for performing NFC filtering on an unfiltered pre-rendered signal. DETAILED DESCRIPTION OF THE INVENTION

[0026] As outlined above, this paper is concerned with the efficient rendering of Ambisonics, specifically HOA signals, as used, for example, within the MPEG-H 3D Audio standard, which supports channel-based, object-based, and scene-based audio coding to provide an improved and immersive 3D sound experience.

[0027] The MPEG-H 3D Audio Low Complexity Profile decoder supports all formats, including channel-based audio, object-based audio, and scene-based audio via Higher Order Ambisonics (HOA). MPEG-H 3D Audio Low Complexity Profile decoders compliant with ISO / IEC23008-3:2019 / AMD2:2020 also support the MPEG-H 3D Audio Baseline Profile, a subset of the MPEG-H 3D Audio Low Complexity Profile. ISO / IEC23008-3:2019 / AMD2:2020 is incorporated herein by reference. [Non-Patent Document 1] ISO / IEC23008-3:2019 / AMD2:2020

[0028] In accordance with Section 12.4.3.4 NFC Processing (Near-Field Compensation Processing) of Specification ISO / IEC 23008-3:2019(E) (incorporated by reference), and as shown in the Ambisonics (particularly HOA) renderer 100 of FIG. 1 , an NFC filtering block 110 may be used before the “HOA to Loudspeaker Conversion” block 120. In particular, the Ambisonics renderer 100 may be configured to receive N Ambisonics (particularly HOA) channel signals 111 from a decode block, which is configured to generate the Ambisonics channel signals 111 from an encoded bitstream. In the case of headphone rendering, the Ambisonics channel signals 111 may be provided to an “Ambisonics to Headphone Conversion (H2B)” block 130, which may be configured to perform binaural rendering of the N Ambisonics channel signals 111. For loudspeaker rendering, the Ambisonics channel signals 111 may be provided to an "Ambisonics to Loudspeaker Transform" block 120. Block 120 is configured to generate S loudspeaker signals 114 for corresponding S loudspeakers of a loudspeaker arrangement based on the N Ambisonics channel signals 111. The loudspeaker signals 114 may be generated using an Ambisonics rendering matrix R 113, which may be determined based on the arrangement of the S loudspeakers.

[0029] The processing in the "HOA to Loudspeaker Conversion" block 120 may include matrix multiplication, in particular the matrix multiplication: P=R×C where R is the HOA renderer matrix 113, C is the matrix of HOA decoded output channels 111 (also referred to herein as Ambisonics channel signals), and P is the renderer output 114 (also referred to herein as loudspeaker channel signals), as shown below.

Number

[0030] The definitions of the variables L, S, and N are given in Table 1.

Table 1

[0031] The elements of the HOA rendering matrix 113 are r i,j For all 0 ≤ i < S and 0 ≤ j < N, they are typically the scalar weights for a given speaker layout. Thus, the process of obtaining the final rendered output for a particular channel (i.e., the loudspeaker channel for a particular loudspeaker) is the sum of the weighted samples of the HOA channels c j, k for all 0 ≤ j < N and 0 ≤ k < L can be regarded as the weighted sum. That is, the elements p of the matrix P i,k for all 0 ≤ i < S and 0 ≤ k < L can be regarded as the weighted sum of c j,k for all 0 ≤ j < N. Here, the weights are the rows of the renderer matrix R.

[0032] As shown in FIG. 1, if it is determined in the determination block 101 that the ambisonics channel signal 111 should be applied with NFC processing, it can be submitted to the NFC processing in the NFC processing block 110 before rendering. The bitstream received from the corresponding ambisonics encoder can indicate whether NFC processing should be applied in the variable or flag "UsesNfc". Alternatively or additionally, the bitstream may indicate the variable "NfcReferenceDistance" indicating the reference distance assumed to be where the sound source and / or loudspeaker was located during encoding.

[0033] The decision block 101 may be configured to determine whether to apply an NFC process based on the variables "UsesNfc" and / or "NfcReferenceDistance." In particular, the NFC process may be applied only if the variable "UsesNfc" indicates that the NFC process should be used (e.g., with a value of "1"). Also, the NFC process may be applied only if the variable "NfcReferenceDistance" has a value r max This can only be applied if the max is the maximum distance at which the loudspeakers of the actual loudspeaker arrangement are located from the listener's position. The loudspeakers may, for example, be arranged on one or more circles around the listener's position.

[0034] If NFC processing is applied, the Ambisonics channel signals 111 may be filtered in an NFC processing block 110 to provide filtered channel signals 112. The filtered channel signals 112 may then be processed in a transform block 120 to provide (filtered) loudspeaker channel signals 114. The Ambisonics channel signals C 111 may then be replaced by the corresponding filtered channel signals 112 in the matrix multiplication described above.

[0035] The NFC processing may include, and may consist of, the application of digital filters, particularly finite impulse response (FIR) and / or infinite impulse response (IIR) filters, to the individual Ambisonics channel signals 111. The filter coefficients may be determined based on the variable "NfcReferenceDistance." For a given HOA order n, the filter coefficients may be the same for all Ambisonics channels corresponding to HOA "modes" m within the given HOA order n, where 0≦m≦n. Thus, for a given HOA order n, a total of (n+1) different NFC filter sets (corresponding to the number of HOA "modes") are used. The grouping of Ambisonics channels for different HOA orders from 1 to 6 is presented in Table 2. [Table 2]

[0036] The NFC processing block 110 may be configured to apply an NFC filter in the time domain. Further details regarding NFC processing are described in ISO / IEC 23008-3:2019I, in particular section 12.4.3.4, which is incorporated herein by reference.

[0037] It can be shown that the filtering operations performed in the NFC processing block 110 and the matrix multiplication operations performed in the transformation block 120 can be reordered (without affecting the final result). As a result, an alternative setup 200 of the Ambisonics renderer 100 may be provided, as shown in FIG. 2a. In the setup 200, the conversion of Ambisonics signals to loudspeaker signals is performed in two stages: in a first stage 220, (n+1) sets of S unfiltered pre-rendered loudspeaker signals 211 are obtained from the Ambisonics channel signals 111. Parallel processing may be applied to obtain each of the (n+1) sets of S unfiltered pre-rendered loudspeaker signals. For each of the (n+1) sets of S unfiltered pre-rendered loudspeaker signals, the transformation is performed by a weighted sum of the Ambisonics channel signals 111. The weights used to process each set are taken from a subset of the rendering matrix R 113. For example, for each set m, 0≦m≦n, a subset R(m) of the rendering matrix R 113 may be multiplied with the matrix C 111 of Ambisonics channel signals.

number

[0038] For example, the weights for the first set (m=0) of S pre-rendered loudspeaker signals 211 correspond to the first column of the rendering matrix R 113, and the weights for the second set (m=1) of S unfiltered pre-rendered loudspeaker signals 211 correspond to the matrices in the second, third, and fourth columns of the rendering matrix R 113. This assignment is derived from the grouping as shown in Table 2, where it should be noted that for the first set, only one Ambisonics channel is involved, namely the first channel (0), and thus the first column of the rendering matrix R 113. On the other hand, for the second set, the group of three subsequent Ambisonics channels, namely channels (1, 2, 3), and thus the next three columns of the rendering matrix R 113, are involved.

[0039] In a second stage, NFC filtering 210 is applied to sets or subsets of pre-rendered loudspeaker signals 211. In particular, for each of the (n+1) sets or subsets of unfiltered pre-rendered S loudspeaker signals 211, NFC filtering is applied individually to provide (n+1) sets or subsets of S filtered pre-rendered loudspeaker signals 213. For example, one of the (n+1) sets of S filtered pre-rendered loudspeaker signals 213 is calculated as follows:

number

[0040] The resulting (n+1) sets of S filtered pre-rendered loudspeaker signals 213 are summed in block 212 to obtain the corresponding loudspeaker channel signals 114. If only a subset of the (n+1) sets is filtered, the subset and the remaining unfiltered signal are summed, i.e., the sum of the (n+1) signals for each loudspeaker S. For example, the loudspeaker channel signals 114 are calculated as follows:

number

[0041] In other words, in block 212, for each loudspeaker, the (n+1) signals corresponding to the same particular loudspeaker are summed to provide a total of S filtered and rendered loudspeaker signals 114. Notably, the loudspeaker channel signals 114 in Figure 2a are identical to the loudspeaker channel signals 114 in Figure 1.

[0042] If NFC processing is not required in this setup, a direct conversion to loudspeaker channel signals 114 can also be performed by processing the Ambisonics channel signals 111 with a rendering matrix R 113.

[0043] In the setup of FIG. 1, the NFC processing is applied to N ambisonics channel signals 111, while in the setup of FIG. 2a, the NFC processing is applied to S*(n + 1) signals. Thus, the setup of FIG. 2a is computationally more efficient than the setup of FIG. 1 when the following holds, i.e., S*(n + 1) < N, i.e., S < (n + 1) (in the case of a 3D setup). This paper mainly describes the 3D setup, but it should be noted that it can be easily extended to a 2D setup. Therefore, for example, when a "global pass" NFC filter is applied to one of the modes, as is generally done for m = 0, the above conditions should be adjusted, for example, to S*n < N - 1. More generally, the condition can be expressed as S*(n + 1 - m1) < N - m0. Here, m1 is the number of ambisonics channel modes from the order n of the HOA signal filtered by the global pass NFC filter, and m0 is the number of ambisonics channel indices corresponding to the number of ambisonics channel modes m1.

[0044] Table 3 shows the reduction of filtering operations achievable for various scenarios.

Table 3

[0045] Thus, the setup of FIG. 2a is efficient for a relatively common speaker layout with S = 6 speakers supported by the MPEG-H LC decoder at a relatively high ambisonics order n. Generally, the setup of FIG. 2a is more efficient when the number N of HOA channels is higher than the number S*(n + 1) of speakers in the target loudspeaker layout.

[0046] 2b shows an Ambisonics renderer 300 that utilizes a decision unit 201 configured to change the order of the transformation and NFC processing depending on the number of Ambisonics channels N and the number of loudspeaker channels S. If S≧(n+1), the NFC processing (in block 110) may be performed before the transformation processing (in block 120). On the other hand, if S<(n+1), the NFC processing transformation (in block 210) may be performed after the processing (in block 220). As a result, a particularly efficient processing may be achieved, as shown in Table 4. [Table 4]

[0047] FIG. 3 shows a flowchart of an exemplary (computer-implemented) method 400 for rendering an Ambisonics audio signal using a loudspeaker arrangement including S loudspeakers. The Ambisonics audio signal may be provided in a bitstream. The method 400 may be performed by a decoder configured to decode the bitstream. In particular, a set of N Ambisonics channel signals 111 may be derived from the bitstream. The set of N Ambisonics channel signals 111 may be Higher Order Ambisonics (HOA) signals. The number N of Ambisonics channel signals 111 is N=(n+1). 2 where n is the order of the Ambisonics signal, for example n>1.

[0048] The method 400 may include converting 401 the N Ambisonics channel signals 111 into a set of unfiltered pre-rendered signals 211. For example, the size of the set of unfiltered pre-rendered signals 211 is equal to S*(n+1). Typically, N>1, and typically S>1. In particular, S>2, e.g., S=6 or S=16. Thus, the loudspeaker channel signals 114 may be derived from the HOA signals.

[0049] Transforming the set of N Ambisonics channel signals 111 into a set of unfiltered pre-rendered signals 211 may be and / or may be performed using matrix multiplication of the Ambisonics signal matrix C (which represents a frame of the set of N Ambisonics channel signals 111) and the renderer matrix R(m) for each "mode" m, given an Ambisonics order n, where 0≦m≦n. For each "mode" m, the renderer matrix R(m) may be an S×N matrix filled with zeros except for the non-zero elements of a subset of the column vectors of the renderer matrix R. The indices of those column vectors are taken from Table 2. For example, for n=2, the non-zero elements of the renderer matrix R(m) for m=2 are the (5th, 6th, 7th, 8th, 9th) column vectors of the renderer matrix R. These correspond to the Ambisonics channel indices in the group (4, 5, 6, 7, 8) of Table 2. Note that the column vector index is simply the offset (+1) of the Ambisonics channel index, which is located at the same position in R(m). In a practical implementation, there is no need to construct such a redundant matrix; element-wise multiplication is sufficient to perform this operation.

[0050] Thus, the transformation from the HOA signal to the different loudspeaker channel signals can be performed by computing linear combinations of the different Ambisonics channel signals using different sets of weights (from the renderer matrix R).

[0051] Further, the method 400 includes performing 402 near-field compensation (NFC) filtering of M unfiltered pre-rendered signals 211 of the set of unfiltered pre-rendered signals 211. For example, M is equal to S*(n+1). An example of step 402 is shown in FIG.

[0052] FIG. 4 shows a flowchart of an exemplary (computer-implemented) method 500 for NFC filtering of a set of M unfiltered pre-rendered signals.

[0053] In step 501, a set of M filtered pre-rendered signals 213 is determined from the set of unfiltered pre-rendered signals 211 based on the NFC coefficients. In particular, the unfiltered pre-rendered signals 211 of the set of unfiltered pre-rendered signals 211 may be multiplied by corresponding NFC filter coefficients to provide the set of M filtered pre-rendered signals.

[0054] In step 502, for each loudspeaker, the filtered pre-rendered signals 213 corresponding to the loudspeaker are summed to provide a set of S filtered loudspeaker channel signals 114. The set of S filtered loudspeaker channel signals 114 may be rendered to the corresponding loudspeaker.

[0055] Method 400 may include providing the S filtered loudspeaker channel signals 114 to the corresponding S loudspeakers, respectively. Alternatively or additionally, method 400 may include rendering the S filtered loudspeaker channel signals 114 with the corresponding S loudspeakers, respectively.

[0056] Performing NFC filtering on the set of unfiltered pre-rendered signals 211 may include performing time-domain filtering using a digital finite impulse response (FIR) filter and / or a digital infinite impulse response (IIR) filter individually on each of the M unfiltered pre-rendered signals 211. The filter for the NFC processing may be determined based on data provided in a bitstream related to the Ambisonics audio signal. In particular, method 400 may include determining a reference distance for the set of N Ambisonics channel signals 111, particularly based on the bitstream of the Ambisonics signal. Further, method 400 may include determining a filter, particularly a filter coefficient, for performing the NFC processing based on the reference distance.

[0057] NFC processing may be used to compensate for the fact that loudspeakers located at a limited distance from the listener position do not emit ideal plane sound waves, and the sound field representations used for Ambisonics typically assume that the emitted sound waves are plane waves.

[0058] Thus, a method 400 is described for applying NFC processing to loudspeaker channel signals (as opposed to applying NFC processing to Ambisonics channel signals), which can lead to a substantial reduction in computational complexity without affecting perceptual quality.

[0059] The method may include determining whether the number S of loudspeakers is less than (and equal to) the number (n+1), where n is the HOA order. M may be S*(n+1). NFC filtering may be performed on the S*(n+1) unfiltered pre-rendered signals 211 if, and especially only if, (n+1)>S or (n+1)≧S. This may result in reduced computational complexity.

[0060] Therefore, the method includes determining whether the number S of loudspeakers is less than (n + 1), where n is the HOA order. Depending on this, NFC filtering can be performed either on S*(n + 1) unfiltered pre-rendered signals 211 or on a set of N ambisonics channel signals 111. In particular, NFC filtering can be performed on a set of S*(n + 1) unfiltered pre-rendered signals 211 when (n + 1) < S. On the other hand, if (n + 1) < S, NFC filtering can be performed on a set of N ambisonics channel signals 111. In particular, when (n + 1) < S, method 400 can include performing NFC filtering on a set of N ambisonics channel signals 111 to generate a set of N filtered ambisonics channel signals 112 and converting the set of N filtered ambisonics channel signals 112 into S filtered loudspeaker channel signals 114. Thus, when (n + 1) < S, NFC filtering can be selectively performed before the "ambisonics to loudspeaker" conversion. As a result, the computational complexity can be significantly reduced in particular.

[0061] As a result, method 400 can perform NFC filtering flexibly either before or after the "ambisonics to loudspeaker" conversion. By doing so, the computational complexity can be reduced.

[0062] In this paper, methods and rendering devices are described that allow ambisonics, particularly HOA signals, to be rendered in a particularly efficient manner by a loudspeaker layout.

[0063] It should be noted that the description and drawings merely illustrate the principles of the proposed method and system. Those skilled in the art will be able to implement various configurations not explicitly described or shown herein, but which embody the principles of the present invention and are within the spirit and scope of the present invention. Moreover, all examples and embodiments outlined herein are expressly and primarily intended for illustrative purposes only, to aid the reader in understanding the principles of the proposed method and system. Furthermore, all statements herein providing principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass equivalents thereof.

[0064] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE).

[0065] [EEE1] 1. A method (400) for rendering an Ambisonics signal using a loudspeaker arrangement having S loudspeakers, the method (400) comprising: converting (401) a set of N Ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), where N>1 and S>1; performing near-field compensation (referred to as NFC) filtering (402) of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeaker; method. [EEE2] performing NFC filtering (402) of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeakers; · from said set of M unfiltered pre-rendered signals, determining a set of M filtered pre-rendered signals based on NFC coefficients; for each loudspeaker S, summing the filtered pre-rendered signal and the remaining unfiltered pre-rendered signal corresponding to the loudspeaker to provide the set of S filtered loudspeaker channel signals (114); Method described in EEE2. [EEE3] 3. The method according to claim 1 or 2, wherein M depends on the number S of loudspeakers and the order n of a Higher Order Ambisonics (HOA) signal corresponding to said set of N Ambisonics channel signals (111). [EEE4] M=S*(n+1-m1), where m1 is the number of Ambisonics channel modes from order n of the HOA signal filtered by the all-pass NFC filter; The method according to any one of EEE1 to EEE3. [EEE5] The method according to EEE4 when citing EEE2, wherein the filtered pre-rendered signals and the remaining unfiltered pre-rendered signals corresponding to the loudspeakers are numbered n+1-m1 and m1 signals, respectively. [EEE6] The method of any one of claims 8 to 10, wherein m1=0. [EEE7] determining a set of M filtered pre-rendered signals from said set of unfiltered pre-rendered signals based on NFC coefficients; multiplication in the frequency domain of each of the S unfiltered pre-rendered signals of the M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals by an NFC coefficient d(m) for each m, where 0≦m≦n; or for each of the S unfiltered pre-rendered signals among the M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals, and an NFC filter d for each m m and m, where 0≦m≦n. The method according to any one of EEE3 to EEE6 when referencing EEE2. [EEE8] The method (400) Determining whether (N-m0) is less than M, where m0≧0 and depends on the number of Ambisonics channel modes from order n of the HOA signal filtered by the all-pass NFC filter; performing (402) near-field compensation (referred to as NFC) filtering of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeaker; performing (402) NFC filtering on M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) if and only if (N-m0)>M or (N-m0)≧M, The method according to any one of EEE3 to 7. [EEE9] Before converting (401) a set of N ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), the method (400) further · determining whether (N - m0) is less than M, where m0 ≥ 0 and depends on the number of ambisonics channel modes from the order n of the HOA signal filtered by the all-pass NFC filter; · depending thereon, performing NFC filtering (402) on M unfiltered pre-rendered signals of the set of unfiltered pre-rendered signals (211), or reversing the order of conversion and NFC filtering, The method according to any one of EEE3 to 8. [EEE10] · when (N - m0) > M, performing NFC filtering (402) on M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals (211); · when (N - m0) < M, reversing the order of conversion and NFC filtering, The method according to EEE9. [EEE11] Reversing the order of conversion and NFC filtering means · performing NFC filtering on the set of N ambisonics channel signals (111) to generate a set of N filtered ambisonics channel signals (112); · converting the set of N filtered ambisonics channel signals into the set of S filtered loudspeaker channel signals (114). The method according to EEE9 or 10. ​ The method according to any one of EEE8 to 12, wherein m0 is the number of Ambisonics channel indices corresponding to the number of Ambisonics channel modes filtered by the all-pass NFC filter. [EEE13] The method according to EEE12, wherein m0=0. [EEE14] 14. The method of any one of EEE1 to EEE13, wherein performing NFC filtering (402) on M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals (211) comprises performing time-domain filtering using a digital finite impulse response filter or a digital infinite impulse response filter on each one of the M unfiltered pre-rendered signals (211) individually. [EEE15] The method comprises: determining a reference distance of the set of N Ambisonics channel signals, in particular based on a bitstream of the Ambisonics signals; determining a filter, in particular a filter coefficient, for performing the NFC filtering based on the reference distance, The method of any one of EEE1 to EEE14. [EEE16] the set of N Ambisonics channel signals (111) are Higher Order Ambisonics (HOA) signals, and / or N=(n+1) 2 where n is the order of the HOA signal, and n>1. The method of any one of EEE1 to EEE15. [EEE17] S>2; and / or S=6; or S=16 17. The method of any one of EEE1 to 16, wherein: [EEE18] Transforming (401) the set of N Ambisonics channel signals (111) into the set of M unfiltered pre-rendered signals includes, for a given loudspeaker k, an Ambisonics signal matrix C representing a frame of the set of N Ambisonics channel signals (111) and a renderer matrix R k This can be done using matrix multiplication with The renderer matrix R for a given loudspeaker k k Specifically, is a (n+1) × N matrix, The method of any one of EEE1 to 17. [EEE19] A computer program product having instructions which, when executed by a computer, cause the computer to perform the method (400) of any one of EEE1 to EEE18. [EEE20] A rendering device (100) for rendering an Ambisonics signal using a loudspeaker arrangement having S loudspeakers, the rendering device (100) comprising: converting (401) a set of N Ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), where N>1 and S>1; performing near-field compensation (referred to as NFC) filtering (402) of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeaker; Rendering device. [EEE21] 1. A method (400) for rendering an Ambisonics signal using a loudspeaker arrangement having S loudspeakers, the method (400) comprising: converting (401) a set of N Ambisonics channel signals (111) into a set of S unfiltered loudspeaker channel signals (214), where N>1 and S>1; performing near-field compensation (referred to as NFC) filtering (402) of said set of S unfiltered loudspeaker channel signals (214) to provide a set of S filtered loudspeaker channel signals (114) for rendering using the corresponding S loudspeakers; method. [EEE22] The method (400) determining whether the number N of Ambisonics channel signals (111) is greater than the number S of loudspeakers; performing NFC filtering (402) on said set of S unfiltered loudspeaker channel signals (214) if and only if N>S or N≧S; Method according to EEE21. [EEE23] The method (400) determining whether the number N of Ambisonics channel signals (111) is greater than the number S of loudspeakers; accordingly performing NFC filtering (402) on said set of S unfiltered loudspeaker channel signals (214) or performing NFC filtering on said set of N Ambisonics channel signals (111); The method according to claim EEE21 or 22. [EEE24] The method (400) · When N > S, perform NFC filtering on the set of S unfiltered loudspeaker channel signals (214); · When N < S, including performing NFC filtering on the set of N ambisonics channel signals (111), The method according to EEE23. 〔EEE25〕 When the method (400) is such that N < S, · Perform NFC filtering on the set of N ambisonics channel signals (111) to generate a set of N filtered ambisonics channel signals (112); · including converting the set of N filtered ambisonics channel signals (112) into the set of S filtered loudspeaker channel signals (114), The method according to any one of EEE23 to 25. 〔EEE26〕 Performing NFC filtering on the set of S unfiltered loudspeaker channel signals (214) includes performing time-domain filtering on each one of the S unfiltered loudspeaker channel signals (214) individually using a digital finite impulse response filter or a digital infinite impulse response filter, The method according to any one of EEE21 to 25. 〔EEE27〕 When the method is · determining the reference distance of the set of N ambisonics channel signals (111), particularly based on the bitstream of the ambisonics signal; and · determining a filter for performing NFC processing, particularly the coefficients of the filter, based on the reference distance, The method according to any one of EEE21 to 26. 〔EEE28〕 the set of N Ambisonics channel signals (111) are higher-order Ambisonics signals, and / or N=(n+1) 2 where n is the order of the Ambisonics signal, and n>1. 8. The method of any one of EEE21 to 27. [EEE29] S>2; and / or S=6; or S=16 29. The method of any one of EEE21 to 28, wherein: [EEE30] Transforming (401) the set of N Ambisonics channel signals (111) into the set of S unfiltered loudspeaker channel signals comprises: an Ambisonics signal matrix C representing a frame of the set of N Ambisonics channel signals (111) and a renderer matrix R k giving a loudspeaker signal matrix P representing a frame of said set of S unfiltered loudspeaker channel signals (214), which can be performed using matrix multiplication with The renderer matrix R is specifically an S×N matrix. 8. The method of any one of EEE21 to 29. [EEE31] The method (400) providing the S filtered loudspeaker channel signals (114) to corresponding S loudspeakers, respectively; and / or Rendering the S filtered loudspeaker channel signals (114) using the S corresponding loudspeakers, respectively; 8. The method according to any one of claims EEE21 to 30. [EEE32] 1. A method (500) for rendering an Ambisonics signal using a loudspeaker arrangement having S loudspeakers, the method (500) comprising: performing joint synthesis and rendering (501) on the set of Ambisonics ambient signals (301) to determine a set of S ambient loudspeaker signals; performing joint synthesis and rendering (501) on the set of dominant sound signals (302) to determine a set of S dominant loudspeaker signals; combining (503) the set of S ambient loudspeaker signals with the set of S dominant loudspeaker signals to provide a set of S unfiltered loudspeaker channel signals (214), where S>1; performing near-field compensation (referred to as NFC) filtering (504) of said set of S unfiltered loudspeaker channel signals (214) to provide a set of S filtered loudspeaker channel signals (114) for rendering using the corresponding S loudspeakers; method. [EEE33] A computer program product having instructions which, when executed by a computer, cause the computer to carry out a method (300, 400) according to any one of EEE21 to EEE32. [EEE34] A rendering device (100) for rendering an Ambisonics signal using a loudspeaker arrangement comprising S loudspeakers, the rendering device (100) comprising: converting a set of N Ambisonics channel signals (111) into a set of S unfiltered loudspeaker channel signals (214), where N>1 and S>1; performing near-field compensation (referred to as NFC) filtering of the set of S unfiltered loudspeaker channel signals (214) to provide a set of S filtered loudspeaker channel signals (114) for rendering using the corresponding S loudspeakers; configured to run Rendering device. [EEE35] 1. A rendering device for rendering an Ambisonics signal using a loudspeaker arrangement comprising S loudspeakers, the rendering device comprising: performing joint synthesis and rendering (301) on a set of Ambisonics ambient signals to determine a set of S ambient loudspeaker signals; · performing joint synthesis and rendering (302) on the set of dominant sound signals to determine a set of S dominant loudspeaker signals; combining the set of S ambient loudspeaker signals and the set of S dominant loudspeaker signals to provide a set of S unfiltered loudspeaker channel signals (214), where S>1; performing near-field compensation (referred to as NFC) filtering of said set of S unfiltered loudspeaker channel signals (214) to provide a set of S filtered loudspeaker channel signals (114) for rendering using said corresponding S loudspeakers; The rendering device that is configured to run [EEE36] 1. A decoder (300) configured to decode a bitstream representing an Ambisonics signal to be rendered by a loudspeaker arrangement comprising S loudspeakers, the decoder (300) comprising a rendering device (100) according to any one of claims EEE34 to EEE35.

Claims

1. 1. A method (400) for rendering an Ambisonics signal using a loudspeaker arrangement having S loudspeakers, the method (400) comprising: converting (401) a set of N Ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), where N>1 and S>1; performing near-field compensation (referred to as NFC) filtering (402) of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeaker; method.

2. performing NFC filtering (402) of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeakers; - determining a set of M filtered pre-rendered signals from said set of M unfiltered pre-rendered signals based on NFC coefficients; for each loudspeaker S, summing the filtered pre-rendered signal and the remaining unfiltered pre-rendered signal corresponding to the loudspeaker to provide the set of S filtered loudspeaker channel signals (114); The method of claim 1.

3. 2. The method of claim 1, wherein M depends on the number S of loudspeakers and the order n of a Higher Order Ambisonics (HOA) signal corresponding to said set of N Ambisonics channel signals (111).

4. M = S*(n+1-m 1 ) and m 1 is the number of Ambisonics channel modes filtered by the all-pass NFC filter from order n of a Higher Order Ambisonics (HOA) signal corresponding to said set of N Ambisonics channel signals (111); The method of claim 1.

5. The filtered pre-rendered signals and the remaining unfiltered pre-rendered signals corresponding to the loudspeakers are each a number n+1-m 1 and m 1 The individual signals are The method of claim 4.

6. m 1 5. The method of claim 4, wherein: =0.

7. M depends on the number S of loudspeakers and the order n of the Higher Order Ambisonics (HOA) signal corresponding to said set of N Ambisonics channel signals (111); determining a set of M filtered pre-rendered signals from said set of M unfiltered pre-rendered signals based on NFC coefficients; - multiplication in the frequency domain of each of the S unfiltered pre-rendered signals of M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals by an NFC coefficient d(m) for each m, where 0≦m≦n; or For each of the S unfiltered pre-rendered signals out of M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals, an NFC filter d for each m m and m, where 0≦m≦n. The method of claim 2.

8. The method (400) ・(N-m 0 ) is less than M, where m 0 ≧0 and depends on the number of Ambisonics channel modes from order n of the HOA signal filtered by the all-pass NFC filter; performing (402) near-field compensation (referred to as NFC) filtering of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeaker; ・(N-m 0 )>M or (N-m 0 performing NFC filtering (402) on M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) if and only if The method of claim 3.

9. Before converting (401) the set of N Ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), the method (400) further comprises: ・(N-m 0 ) is less than M, where m 0 ≧0 and depends on the number of Ambisonics channel modes from order n of the HOA signal filtered by the all-pass NFC filter; performing NFC filtering (402) on M unfiltered pre-rendered signals of said set of unfiltered pre-rendered signals (211) depending thereon, or reversing the order of transformation and NFC filtering, The method of claim 3.

10. ・(N-m 0 )>M, performing NFC filtering (402) on M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211); ・(N-m 0 ) < M, reverse the order of conversion and NFC filtering; 10. The method of claim 9.

11. Reversing the order of conversion and NFC filtering performing NFC filtering on said set of N Ambisonics channel signals (111) to generate a set of N filtered Ambisonics channel signals (112); converting the set of N filtered Ambisonics channel signals into the set of S filtered loudspeaker channel signals (114); 10. The method of claim 9.

12. m 0 9. The method of claim 8, wherein Λ is the number of Ambisonics channel indices corresponding to the number of Ambisonics channel modes filtered by the all-pass NFC filter.

13. m 0 13. The method of claim 12, wherein: = 0.

14. 2. The method of claim 1, wherein performing NFC filtering (402) on M unfiltered pre-rendered signals (211) of the set of unfiltered pre-rendered signals (211) comprises performing time-domain filtering using a digital finite impulse response filter or a digital infinite impulse response filter on each one of the M unfiltered pre-rendered signals (211) individually.

15. The method comprises: determining a reference distance of said set of N Ambisonics channel signals, in particular based on a bitstream of said Ambisonics signals; determining a filter, in particular the coefficients of the filter, for performing the NFC filtering based on the reference distance, The method of claim 1.

16. said set of N Ambisonics channel signals (111) are Higher Order Ambisonics (HOA) signals; and / or ・N=(n+1) 2 where n is the order of the HOA signal, and n>1. The method of claim 1.

17. S>2; and / or S = 6; or ・S=16 The method of claim 1, wherein

18. Transforming (401) the set of N Ambisonics channel signals (111) into the set of M unfiltered pre-rendered signals comprises, for a given loudspeaker k, an Ambisonics signal matrix C representing a frame of the set of N Ambisonics channel signals (111) and a renderer matrix R k This can be done using matrix multiplication with the renderer matrix R for a given loudspeaker k k Specifically, is a (n+1) × N matrix, The method of claim 1.

19. A computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the method (400) of any one of claims 1 to 18.

20. A rendering device (100) for rendering an Ambisonics signal using a loudspeaker arrangement having S loudspeakers, the rendering device (100) comprising: converting (401) a set of N Ambisonics channel signals (111) into a set of unfiltered pre-rendered signals (211), where N>1 and S>1; performing near-field compensation (referred to as NFC) filtering (402) of M unfiltered pre-rendered signals (211) of said set of unfiltered pre-rendered signals (211) to provide a set of S filtered loudspeaker channel signals (114) for rendering using a corresponding S loudspeaker; Rendering device.

Citation Information

Patent Citations

  • AUDIO DATA PROCESSING METHOD AND SOUND COLLECTOR FOR IMPLEMENTING THIS METHOD

    JP2006506918A

  • Apparatus and method for surround audio signal processing

    JP2018196133A

  • Method for processing audio data and sound acquisition device implementing this method

    US20060045275A1

  • Near field compensation for decomposed representations of a sound field

    US20150127354A1

  • Rendering audio objects with multiple types of renderers

    WO2020227140A1