Spatial spread modeling for volumetric audio sources

A parametric model for determining the effective spatial extent of volumetric audio sources addresses the challenges of unnatural rendering and computational inefficiency by focusing on acoustically relevant portions, improving spatial realism and efficiency in XR environments.

JP2026041760APending Publication Date: 2026-03-10TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing methods for rendering volumetric audio sources in XR environments face challenges due to the direct use of geometric extent, leading to unnatural spatial perception and excessive computational complexity, especially for large sources, which are not acoustically relevant to the listener.

Method used

A parametric model is developed to determine the effective spatial extent of volumetric audio sources based on geometric and distance parameters, coherence properties, and frequency characteristics, allowing for more natural rendering and reduced computational effort.

Benefits of technology

The model provides a computationally efficient and perceptually accurate rendering of large volumetric audio sources by focusing on acoustically relevant portions, reducing computational load and enhancing spatial realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041760000001_ABST
    Figure 2026041760000001_ABST
Patent Text Reader

Abstract

A spatial extent modeling method, carrier and apparatus for volumetric audio sources are provided for rendering the audio sources for a listener. [Solution] Method 800 includes obtaining a spatial extent value indicating the spatial extent of an audio source (s802), obtaining a distance value specifying the distance between the audio source and a listener (s804), determining whether the distance value is less than a threshold distance value (s806), and, as a result of determining that the distance value is less than the threshold distance value, rendering the audio source to the listener using the effective spatial extent value (s808).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] SUMMARY OF THE INVENTION Embodiments are disclosed that relate to methods and systems for spatial extent modeling for volumetric audio sources. [Background technology]

[0002] An XR (virtual reality, augmented reality, or mixed reality) scene may contain many audio sources spatially distributed within the space of the scene. Many of these audio sources have specific, well-defined locations in space and can be considered point-like sources. These audio sources are generally rendered as point-like audio objects to the listener.

[0003] However, XR scenes also often contain audio sources that are volumetric rather than point-like in nature, meaning that the audio sources have a certain spatial extent in one or more spatial dimensions.

[0004] In some cases, such a volumetric audio source may correspond to a single physical entity in a scene (e.g., an airplane, a piano, a train, a transport pipe in a factory, etc.) Some of these volumetric audio sources may radiate audio as a single coherent audio source, while others may radiate audio more like a spatially extended diffuse audio source.

[0005] In other cases, rather than corresponding to a single physical entity, a volumetric audio source may represent an area in a scene containing multiple (perhaps even a continuum) independent audio sources that together can be considered a compound volumetric audio source. Examples of this type of volumetric audio source are a seaside beach and a busy highway. In the busy highway example, each car is essentially an independent audio source, but the highway with many cars on it can be considered a compound volumetric audio source.

[0006] As in the beach and highway example described above, in many cases the spatial extent of a volumetric audio source can be extremely large in one or more of its spatial dimensions, and in some cases this can even become virtually infinitely large (e.g., with respect to the distance from the listener to the volumetric audio source). Summary of the Invention

[0007] Generally, scene description data for an XR scene specifies the extent of a volumetric audio source in terms of the source's physical geometry (e.g., the source's physical size in one or more dimensions, or a geometric mesh structure that describes the source's physical geometry). This specified physical geometry of the source generally relates directly to the physical geometry of some corresponding physical (and often visual) entity in the XR scene (e.g., a car, a piano, etc.).

[0008] However, as described above, a volumetric audio source may be physically very large in one or more dimensions, such as the beach and busy highway described above. In such cases, the physical size or geometry of the volumetric audio source as typically specified in its spread data is often not well suited to being used directly to render the audio source to a listener.

[0009] In particular, in many cases, only a limited portion of the geometric extent of a volumetric audio source contributes significantly to the audio energy received by a listener at a given listening position. This is true for extremely large (especially "infinitely" large) volumetric sources, where the outer portions of the geometric extent are so far away from the listener that distance and intermediate attenuation prevent significant audio energy from reaching the listener from these outer portions.

[0010] It may also be true when a listener is near a moderately sized volumetric audio source, where the audio energy reaching the listener from portions of the source that are close to the listener essentially dominates over the audio energy coming from portions that are farther away from the listener. Thus, the "acoustically relevant" portions of a given volumetric audio source may depend on the listener's position relative to the source.

[0011] Thus, for large volumetric audio sources, it is often not very appropriate or convenient to simply use the specified geometric extent as a direct measure of how wide or tall the audio source should be rendered to a listener at a given listening position. Indeed, doing so can create various problems.

[0012] One problem with using the specified geometric extent of a volumetric audio source directly for audio rendering is that the resulting subjective spatial extent of the source (e.g., the size of the source as perceived by the listener) can be unnatural (e.g., unnaturally wide, i.e., the spatial extent can be perceived as wider than is perceived in real life). This problem can arise, for example, in rendering scenarios where the audio of a volumetric audio source is rendered to a listener using virtual loudspeakers located at the edges of the specified geometric extent of the source. These virtual loudspeakers, as described above, are often spaced too widely apart.

[0013] Specifying the intended perceived spatial extent of a source instead of its geometric extent would also be problematic, as the intended perceived spatial extent is valid only for one particular listening position, and deriving the intended perceived spatial extent for other listening positions (as would be required in a 6-DOF XR use case) may not be straightforward or even possible.

[0014] Furthermore, in rendering scenarios in which advanced physical modeling techniques are used to accurately render audio radiated by volumetric audio sources, the computational complexity required for rendering generally increases rapidly as the physical size of the source increases. For large volumetric sources (e.g., the beaches and busy highways and passing trains described above), using the specified geometric extent of the source directly to render the source to the listener can easily require excessive computational effort, especially in real-time interactive XR applications. Furthermore, a significant portion of this computational effort may even be spent unnecessarily, since that portion is used to render audio radiated by portions of the volumetric source that do not even significantly contribute to the audio at a particular listener position.

[0015] Therefore, it would be extremely beneficial to have a method for modeling the acoustically relevant spatial extent of a volumetric audio source at a given listening position based on the specified geometric extent and possibly other properties of that source, in order to be able to render large volumetric audio sources in a perceptually appropriate and computationally efficient manner. It would be particularly desirable if the model were extremely simple so that it could be implemented as a lightweight add-on to existing real-time renderer architectures.

[0016] Embodiments of the present disclosure are directed to methods and systems for providing extremely low-complexity parametric models for determining the effective spatial extent of a volumetric audio source, i.e., the portion of the geometric extent of the volumetric audio source that significantly contributes to the audio received by a listener at a given listening position.

[0017] The parameters of the model include (i) a size parameter indicating the size of the geometric extent of the volumetric audio source in one or more dimensions and / or (ii) a distance parameter indicating the distance from the listener to the volumetric audio source.

[0018] The parameters of the model may also include (i) parameters indicating the coherence properties of the volumetric audio source (e.g., coherent, diffuse, or something in between) and / or (ii) frequency parameters (in the case of a (partially) coherent source).

[0019] The determined effective spatial extent may be used in rendering the volumetric audio source to a listener. For example, the determined effective spatial extent may be used to (i) determine a target auditory rendering size for the volumetric audio source at a given listening position and / or (ii) select only an acoustically relevant sub-portion of the geometric extent of the volumetric audio source for rendering the audio of the volumetric audio source to a listener, while discarding other acoustically irrelevant portions of the geometric spatial extent for rendering.

[0020] In one aspect, a method is provided for rendering an audio source for a listener. The method includes obtaining a spatial extent value indicating the spatial extent of the audio source and obtaining a distance value specifying a distance between the audio source and the listener (also known as an "observation distance"). The method also includes determining whether the distance value is less than a threshold distance value. The method further includes, as a result of determining that the distance value is less than the threshold distance value, rendering the audio source for the listener using the effective spatial extent value.

[0021] In another aspect, a computer program is provided, the computer program comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform a method according to any one of the embodiments disclosed herein. In another aspect, a carrier containing the computer program is provided, the carrier being one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.

[0022] In another aspect, there is provided an apparatus adapted to perform a method according to any one of the embodiments disclosed herein, in one embodiment, the apparatus comprises a processing circuit and a memory, the memory including instructions executable by the processing circuit, whereby the apparatus is adapted to perform a method according to any one of the embodiments disclosed herein.

[0023] advantage

[0024] For large volumetric audio sources, using the effective spatial extent according to embodiments of the present disclosure allows for a more natural and realistic spatial rendering of the source than using the geometric extent of the source directly.

[0025] Also, compared to directly specifying the intended perceived spatial extent of a volumetric audio source, which is only valid for a specific listening position, the modeled effective spatial extent according to embodiments of the present disclosure is valid at any listening position.

[0026] Furthermore, in some rendering scenarios, methods and systems according to embodiments of the present disclosure allow for greater computational efficiency when rendering audio from large volumetric audio sources, since only that portion of the acoustically relevant geometric extent at a given listening position is considered in rendering.

[0027] Also, in embodiments of the present disclosure, the parametric model for determining effective spatial extent is extremely simple and can be easily implemented as a lightweight add-on to existing render architectures.

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various embodiments. [Brief explanation of the drawings]

[0029] [Figure 1] FIG. 2 illustrates different parameters used for audio rendering. [Figure 2] FIG. 1 illustrates the sound pressure level (SPL) behavior of a fully incoherent or diffuse one-dimensional volumetric audio source according to one embodiment. [Figure 3] FIG. 1 illustrates the SPL behavior of a coherent one-dimensional volumetric audio source according to one embodiment. [Figure 4] FIG. 1 illustrates the SPL behavior of a coherent one-dimensional volumetric audio source according to one embodiment. [Figure 5] FIG. 1 illustrates a system, according to some embodiments. [Figure 6] FIG. 1 illustrates a system, according to some embodiments. [Figure 7A] FIG. 1 illustrates a system, according to some embodiments. [Figure 7B] FIG. 1 illustrates a system, according to some embodiments. [Figure 8] FIG. 1 is a diagram of a process, according to one embodiment. [Figure 9] 1 is a diagram of an apparatus, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0030] Effective spatial extent: A qualitative model

[0031] As explained above, it is desirable to find a model that can infer the acoustically relevant, or effective, spatial extent of a volumetric audio source for a listener at a given listening position from some of the physical properties of that source.

[0032] A reasonable general starting assumption for developing such a model is that portions of a volumetric audio source that do not contribute significantly to the perceived loudness of the source at a particular listening position (in the sense that the listener will not be able to perceive any difference in loudness regardless of whether audio coming from those portions is included in the rendering) will also not contribute significantly to the perceived spatial properties of the source (including the perceived spatial extent of the source) at that same listening position.

[0033] The starting point for building the model is the simple case of a one-dimensional volumetric audio source (i.e., a line source) with a variable geometric spatial extent of size L in a single dimension (i.e., an acoustic line source with variable length L). For this line source, the behavior of the sound pressure level (SPL) at an observation point O located at a perpendicular distance D relative to the midpoint of the line source can be evaluated as a function of the source length L (see Figure 1).

[0034] If the length L of a line source is very small, the line source is essentially a point source. As the length L steadily increases on either side of the line source, with a constant source intensity density along the length (e.g., at length 2L, the source radiates twice the amount of acoustic energy than at length L), the SPL at observation point O is expected to increase as well.

[0035] However, as the source length L increases, the contribution from the outer edges of the source to the SPL at the observation point becomes smaller and smaller with increasing distance from these outer portions to the observation point. Thus, the rate at which the SPL increases as a function of increasing length L may decrease. At some point, when the source becomes very long, the rate at which the SPL increases may become insignificant, such that the SPL no longer increases with further increases in source length. In other words, the SPL decreases with increasing length of the line source (L eff ), it will saturate.

[0036] For an infinitely long line source with a constant source intensity density, the size L of the infinitely long line source at different observation points O eff Different segments of the source can be acoustically significant parts of the source. In other words, as the listener moves along a line parallel to an infinitely long line source, the listener effectively moves along with the source of size L. eff The source is perceived through a spatial window.

[0037] In some embodiments of the present disclosure, the effective spatial extent of a one-dimensional volumetric audio source (i.e., a line source) is defined as the size of the smallest source segment whose sound level at a given listening position is less than a threshold sound level difference value below the sound level of the complete source. In other words, adding portions of the line source beyond the edge of the effective spatial extent does not add more than the threshold sound level difference value to the sound level at the listening position.

[0038] The threshold sound level difference value used in determining the effective spatial extent can be chosen in different ways, but is most conveniently defined relatively (e.g., as a percentage of the linear sound pressure due to the full source, or as a number of decibels below the SPL of the full source).

[0039] Since the objective is to link the physical size of a source to the corresponding perceived auditory size, it is desirable to select the threshold sound level difference value such that it bears a relationship to this perceived auditory size. In particular, it may be desirable to select the threshold sound level difference value such that the resulting effective spatial extent corresponds to the size of the smallest source segment whose loudness (a perceptual measure) is indistinguishable from that of the complete source. In this context, therefore, a perceptually relevant criterion for setting the threshold sound level difference value is the just-noticeable-difference (JND) for loudness, which is known from the acoustic perception literature to be approximately 1 dB SPL.

[0040] Effective spatial extent: A quantitative model

[0041] Volumetric audio sources can be physically modeled as a dense distribution of point sources. For a one-dimensional volumetric audio source, the total pressure response P linecan be expressed as follows:

[0042] TIFF2026041760000002.tif14170

[0043] where N is the total number of point sources used to model the one-dimensional volumetric audio source, and A i (ω) is the complex amplitude of the i-th point source at radial frequency ω, k is the wave number ω / c, c is the speed of sound in air, and r i is the distance from the i-th point source to the observation point The distance to TIFF2026041760000003.tif10170.

[0044] The sound pressure level (SPL) of a one-dimensional volumetric source can then be expressed as:

[0045] TIFF2026041760000004.tif11170

[0046] When modeling continuous volumetric sources in this manner, a sufficiently small spacing between individual point sources should be used to obtain accurate results over the entire frequency range of interest (e.g., 0-20 kHz).

[0047] There are various types of one-dimensional volumetric audio sources, including coherent one-dimensional volumetric audio sources (all of whose points coherently radiate the same acoustic signal) and diffuse one-dimensional volumetric sources (all of whose points radiate independent, fully uncorrelated signals). These two extreme types of one-dimensional volumetric audio sources behave significantly differently in various ways, as will be explained in more detail below.

[0048] Diffuse 1D volumetric audio source

[0049] Since diffuse volumetric audio sources can be viewed as a dense distribution of independent point sources that all have frequency-independent SPL-vs-distance behavior, the behavior of diffuse volumetric audio sources is also frequency-independent. Therefore, the following results are valid for any individual frequency as well as for the broadband.

[0050] Figure 2 shows simulation results for the total SPL of a fully incoherent or diffuse one-dimensional volumetric audio source as a function of line source length L for observation distances D of 0.1 m, 1 m, 10 m, and 100 m.

[0051] One common characteristic of all four curves is that they have two distinct regions: at small line source lengths (or at large observation distances), the SPL increases at a constant rate of 3 dB per doubling of the line source length L (note the logarithmic scale on the horizontal axis), and at long line source lengths (or at small observation distances), the SPL becomes constant as a function of L.

[0052] The 3 dB SPL increase per doubling of length for a small line source length (or for a long observation distance) is expressed in terms of total pressure (Equation 1), where the complex amplitude A in Equation 1 is used to determine the total pressure. i The distance r from each point source to the observation point is the only relevant i If the source powers for the individual point sources are assumed to be equal, it can be easily shown from equations 1 and 2 that the following relationship holds:

[0053] TIFF2026041760000005.tif12170

[0054] Equation 3 implies a 3 dB increase in SPL when doubling the number of point sources, N. If it is further assumed that the spacing between individual point sources is uniform, Equation 3 implies a 3 dB SPL increase for doubling the line source length, L.

[0055] The saturation of SPL to a constant value at large line source lengths (or at small observation distances) is consistent with the qualitative model described above. As explained above, the saturation of SPL can be explained by the fact that as the length increases, the contribution of newly added point sources at the outer edges becomes less and less significant, and eventually becomes completely insignificant. In other words, once a diffuse one-dimensional volumetric audio source reaches a certain line source length, further increases in length do not lead to a further increase in SPL.

[0056] Comparing the curves for different observation distances shows that the line source length L at which the transition between the two regions occurs, i.e., between the region where the SPL increases and the region where the SPL remains at a substantially constant value, depends on the observation distance, with the transition length being larger for larger observation distances.

[0057] The effective spatial extent of a very long diffuse one-dimensional volumetric audio source can be estimated from the curve shown in Figure 2 by finding the line source length L, a threshold sound level difference value at which the SPL falls below the saturation SPL observed at very large source lengths. The line source length at which the SPL saturates is found to be proportional to the observation distance D.

[0058] If the JND for loudness difference is chosen as the criterion for setting the threshold sound level difference value, which is equal to about 1 dB SPL, then the line source length at which the SPL saturates (i.e., in this particular case, the -1 dB point) is found to have a value of about 6 times the observation distance D.

[0059] Thus, the effective spatial extent of a source with a length of 6D or greater is equal to 6D (i.e., it is proportional to the observation distance D). Equivalently, at observation distances less than L / 6, the spatial extent is equal to 6D.

[0060] For line source lengths less than 6D, or equivalently, observation distances greater than L / 6, the effective spatial extent is simply equal to the line source length L (i.e., every part of the one-dimensional volumetric source contributes significantly to the sound received at the observation point). This characterization allows for more efficient and realistic rendering of audio sources that behave like one-dimensional diffuse volumetric sources.

[0061] The effective spatial extent can be expressed either (i) in terms of length or (ii) in terms of angular span ("opening angle"). When the effective spatial extent is expressed in terms of length, the effective spatial extent for line source lengths greater than the saturation length is proportional to the observation distance D (e.g., the effective spatial extent is equal to 6D for some audio sources). In contrast, when the effective spatial extent is expressed in terms of angular span ("opening angle (OA)"), the effective spatial extent has a constant value (i.e., is independent of the observation distance). A common expression for opening angle (OA) is OA = 2 * atan((0.5 * effective spatial extent in units of length) / D), where atan is the arctangent function. For a diffuse source, the effective spatial extent in units of length is 6D, and therefore the expression for OA is OA = 2 * atan((0.5 * 6D) / D) = 2 * atan(3) = 143 degrees. So if the renderer has OA and D, it can calculate the effective spatial extent in units of length.

[0062] Another common characteristic of the curves shown in Figure 2 is that in the region where the SPL increases by 3 dB per doubling of length, the SPL decreases by 20 dB for a 10-fold increase in observation distance, or equivalently, by 6 dB per doubling of distance. This means that in this region, a one-dimensional diffuse volumetric audio source behaves like a point source with respect to SPL (i.e., p ∝ 1 / r). In contrast, in the region where the SPL is constant as a function of line source length, the SPL decreases by only 10 dB for each 10-fold increase in distance, or by 3 dB per doubling of distance. This means that in this region, a one-dimensional diffuse volumetric source behaves like a theoretical line source. This distance-dependent SPL behavior of a finite-length line source is described in U.S. Provisional Patent Application No. 62 / 950,272, filed December 19, 2019.

[0063] Coherent 1D volumetric audio source

[0064] For a fully coherent uniform one-dimensional volumetric audio source, all amplitudes A in Eq. i are equivalent (i.e., A i =A∀i). Phase terms of individual point sources of volumetric audio sources The frequency dependence and coherency of TIFF2026041760000007.tif10170 means that the total pressure response of a volumetric audio source is also frequency dependent, and therefore it is necessary to analyze the effective spatial extent for individual frequencies as well as for broadband.

[0065] Figure 3 shows the SPL response as a function of linear source length for various frequencies and one observation distance. Like the diffuse one-dimensional volumetric audio source, for a coherent one-dimensional volumetric audio source, there is an expected saturation of SPL beyond a certain linear source length for each individual frequency, but unlike the curves shown in Figure 2, this time the saturation length is frequency dependent.

[0066] The SPL at small line source lengths increases at a rate of 6 dB per doubling of length, instead of the 3 dB shown in Figure 2 for a diffuse one-dimensional volumetric audio source. Following the same reasoning as for diffuse sources, for small line source lengths (or large observation distances), the distance r from the individual point source to the observation point increases. i is essentially equal to the pressure P for each point source of the volumetric audio source. i are all equivalent to a common pressure P, so (assuming equal power and equal spacing for the individual point sources as above) Equation 1 reduces to:

[0067] TIFF2026041760000008.tif12170

[0068] This leads to the following:

[0069] TIFF2026041760000009.tif11170

[0070] This indeed corresponds to the observed 6 dB increase for each doubling of the length L.

[0071] Based on an analysis of corresponding simulation results for many other observation distances, and using the same procedure for determining the effective spatial extent as was done for diffuse sources, it was found that for small observation distances (or large line source lengths) the effective spatial extent is frequency dependent, TIFF2026041760000010.tif12170 (where f is the frequency and c1 is a constant), but the effective spatial extent is again simply equal to the line source length L for large observation distances (or small line source lengths).

[0072] The transition distance between the two regions is L 2 It is found to be proportional to f, and the proportionality coefficient is TIFF2026041760000011.tif11170. As mentioned above, for the particular choice of using the JND for loudness difference (1 dB SPL) as the threshold sound level difference value for finding the saturation length, it has been empirically found that c1 is equal to approximately 18.4.

[0073] As shown above, the behavior of a coherent one-dimensional volumetric audio source is frequency-dependent and will generally be rendered in a frequency-dependent manner. To observe the wideband behavior of a coherent one-dimensional volumetric audio source, simulations can be performed for 128 uniformly spaced frequencies from 20 Hz to 20 kHz, and the results for all individual frequencies can be summed to obtain a wideband result, which may imply a white source spectrum assumption.

[0074] Figure 4 shows the broadband SPL as a function of line-source length L for several observation distances D. Comparing Figure 4 with Figure 2 (which shows the SPL of a frequency-independent diffuse source), the overall broadband behavior of the coherent and diffuse sources is quite similar, especially for very small and very large line-source lengths. The main differences are: (1) the transition region for the coherent source is much wider, and (2) there is some ripple within the transition region for the coherent source. Furthermore, at large observation distances, the SPL for the coherent source increases by 6 dB per doubling of line-source length (as observed for each frequency), instead of 3 dB per doubling of line-source length for the diffuse source.

[0075] For small observation distances (or large line source lengths), the broadband effective spatial extent is proportional to the square root of the observation distance D, i.e. TIFF2026041760000012.tif11170, but the broadband effective spatial extent is again simply equal to the line source length L for large observation distances (or small source lengths). The transition distance between the two regions is proportional to the square of the line source length L, with a proportionality factor of TIFF2026041760000013.tif11170. For the particular choice of using the JND for loudness difference (1 dB SPL) as the threshold sound level difference value for finding the saturation length, c2 was found to be approximately equal to 3.5.

[0076] Summary of results from the simulation

[0077] The effective spatial extent of a one-dimensional volumetric audio source is equal to the linear source length L for observation distances greater than some transition distance (also known as a threshold distance value) (or for linear source lengths less than some transition length).

[0078] For observation distances smaller than the transition distance (or for line source lengths larger than the transition length), the effective spatial extent is proportional to the observation distance D for diffuse one-dimensional volumetric audio sources, while the effective spatial extent is proportional to the square root of D for coherent one-dimensional volumetric audio sources.

[0079] The transition distance is proportional to the source length L for diffuse one-dimensional volumetric audio sources, whereas for coherent one-dimensional volumetric audio sources the transition distance is proportional to L squared.

[0080] Parametric model for effective spatial extent

[0081] The simulation results described above show the effective spatial extent L of a one-dimensional volumetric audio source of length L as a function of observation distance D. eff This leads to the following parametric model for

[0082] Diffuse 1D volumetric source:

[0083] TIFF2026041760000014.tif17170

[0084] Coherent 1D volumetric sources, frequency dependent:

[0085] TIFF2026041760000015.tif21170

[0086] Coherent 1D volumetric source, broadband:

[0087] TIFF2026041760000016.tif19170

[0088] For the particular choice of threshold sound level difference value of 1 dB (JND for loudness), the constants have been empirically found to have the following approximate values: c0≈6, c1≈18.4, and c2≈3.5.

[0089] Numerical example

[0090] Below is an example showing how a parametric model for effective spatial extent affects the rendering of a one-dimensional volumetric audio source.

[0091] Example 1:

[0092] For an "infinitely" long diffuse source (e.g., a coastline at a beach), the effective spatial extent will be in the "c0D" range for a listener at any practically relevant observation distance. Thus, the effective spatial extent will be 6 m (143 degrees) at a 1 m observation distance, 60 m (143 degrees) at a 10 m observation distance, and 600 m (143 degrees) at a 100 m observation distance.

[0093] Example 2:

[0094] For a diffuse one-dimensional volumetric source with length L = 10 m, the effective spatial extent will be 0.6 m (143 degrees) at an observation distance of 0.1 m, 6 m (143 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 1.7 m (which yields 53 degrees at a 10 m observation distance and 6 degrees at a 100 m observation distance).

[0095] Example 3:

[0096] For a coherent one-dimensional volumetric source with length L = 10 m, the broadband effective spatial extent will be 1.1 m (160 degrees) at an observation distance of 0.1 m, 3.5 m (121 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 8.2 m.

[0097] Example 4:

[0098] For a coherent one-dimensional volumetric source with length L=10 m, the effective spatial extent is:

[0099] At f = 100 Hz: 1.8 m (85 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 30 m;

[0100] At f = 1000 Hz: 0.6 m (32 degrees) at a 1 m observation distance, and 10 m at any observation distance greater than 300 m.

[0101] Taking advantage of effective spatial extent when rendering volumetric audio sources

[0102] The renderer may use the derived effective spatial extent in a variety of ways.

[0103] Set the target spatial extent:

[0104] The derived effective spatial extent may be used to set a target spatial extent for rendering a long volumetric audio source to a listener at a particular listening position. This results in delivering a more appropriate rendered source width to the listener compared to simply using the received geometric extent data. For example, in one scenario, the derived effective spatial extent may be used to determine the optimal position of virtual stereo loudspeakers used to render the source to a listener at a particular listening position. In another scenario, the derived effective spatial extent may be used to set a target spatial width in a spatial widening algorithm used to render a volumetric audio source to a listener at a particular listening position.

[0105] Determine the spatial window:

[0106] For very long volumetric audio sources, the derived effective spatial extent can be used to determine which part of the source should be rendered at any instant in time to a listener moving along, away from, and / or towards the source. This is like applying a spatial window that slides with the listener, with the derived effective spatial extent (dynamically updated as the listener's position changes) determining the width of the spatial window.

[0107] Save computing power:

[0108] In use cases where audio from a volumetric audio source is rendered using some form of physical modeling, computational power can be saved by using the derived effective spatial extent to limit the portion of the source that needs to be rendered to a listener at a particular listening position.

[0109] Extension to 2D and 3D volumetric audio sources

[0110] The one-dimensional quantitative parametric model described above is valid at least for volumetric audio sources that have significant spatial extent in only one spatial dimension, meaning that the extent in the other two dimensions is small enough relative to the observation distance so that the extent in these dimensions does not significantly affect the effective spatial extent in the major (longer) dimension. In particular, this will be true if, at a particular observation distance, the source behaves essentially like a point source in the other two dimensions.

[0111] The above-identified provisional patent application describes a model for determining when a one-dimensional audio source behaves like a point source as a function of source length and observation distance. In particular, this document explains that a diffuse one-dimensional audio source behaves like a point source at observation distances that exceed the source length.

[0112] Thus, if a diffuse 2D volumetric audio source has two sizes in two dimensions, dimension 1 and dimension 2, where the size in dimension 1 is longer than the size in dimension 2, then the 1D quantitative effective spatial extent model of Equation 6 can be applied to dimension 1 of this 2D source if the size in dimension 2 is smaller than the observation distance D (or if the observation distance D is larger than L2, the size in dimension 2).

[0113] Similar criteria can be obtained for the validity of one-dimensional Equations 7 and 8 for coherent 2D volumetric sources. In that case, the one-dimensional model is of size in dimension 2. For the frequency-dependent model of Equation 7, TIFF2026041760000017.tif14170 or (2) Regarding the broadband model of Eq. TIFF2026041760000018.tif11170 (or equivalently, if the observation distance is less than f(L2) for the frequency-dependent model) 2 / 339, or 23(L2) for wideband models 2 is valid when the

[0114] For example, if we have a 2D diffuse source with a width of 10 m and a height of 1 m, its effective spatial extent (in the long dimension) can be calculated from Equation 6 (for L=10 m) if the observation distance is greater than 1 m. For a fully coherent 2D source of the same size, its effective spatial extent (in the long dimension) at 500 Hz can be calculated from Equation 8 if the observation distance is greater than 1.5 m.

[0115] These examples show that the requirements regarding the "1-dimensionality" of a volumetric source for the 1-D quantitative model of Eqs. 6-8 to be applicable are fairly relaxed, and that the 1-D model can in fact be applied to a wide range of "long" 2D (and 3D) volumetric sources.

[0116] It should be noted that the quantitative criteria described above for the validity of a 1D model should be understood as an indication, rather than a strict boundary between areas where a 1D model applies and areas where it does not apply, for effective spatial extent. This provides a means for identifying types of 2D sources for which a 1D model may be applicable and / or the conditions under which a given 2D source may be modeled by a 1D model.

[0117] Thus, an additional feature of embodiments of the present disclosure is that the renderer may determine, based on the above criteria, whether to apply a one-dimensional model to a 2D or 3D volumetric source and / or when to switch between rendering the 2D or 3D volumetric source according to a simplified one-dimensional model and rendering it using a more complex 2D or 3D model.

[0118] For volumetric audio sources that have significant extent in two or more dimensions (and therefore do not meet the criteria described above for being considered a one-dimensional audio source), the same qualitative models and principles described above still apply.

[0119] In general, expanding the geometric spatial extent in one dimension has the effect of increasing the effective spatial extent in the other dimension(s). For example, for a given observation distance D, comparing a purely one-dimensional source (i.e., a line source) having length L with a two-dimensional source (i.e., an area source) having length L and height H, the effective extent of the two-dimensional source along the dimension for length L will be larger than that of the one-dimensional source. Also, the transition distance between the region where the effective spatial extent is a function of observation distance and the region where the effective spatial extent is simply equal to the physical size of the geometric extent will be smaller for an area source than for a line source. In other words, for a 2D area source, the entire width of the source already needs to be taken into account in rendering at shorter distances than for a line source of equal width.

[0120] In the above example, the two-dimensional surface source can be thought of as consisting of a continuum of vertical line sources of size H distributed along a length L (instead of point sources as in the one-dimensional model described above). Each of these vertical line sources has a distance attenuation that is slower than the 1 / r attenuation of a point source, and therefore points along the horizontal extent of the source must be farther away from the observation position to become insignificant (for a given SPL-based significance criterion, e.g., the 1 dB SPL loudness JND criterion described above). The consequence of this is that for a two-dimensional surface source, the spatial extent in each dimension will be greater than for each of the two dimensions individually.

[0121] Note that the effective spatial extent in each dimension is still limited by the geometric size of the extent in that dimension (i.e., the effective spatial extent will never exceed the geometric size).

[0122] The above description holds for rectangular 2D surface sources where the source power is distributed approximately evenly across the surface. For this class of 2D sources, a simple extension of the 1D quantitative model of Equations 6-8 can be constructed as described below.

[0123] The provisional application discloses a parametric model for the distance-dependent SPL decay function for a finite-length one-dimensional source that essentially identifies three different observation distance regions: a region where the source behaves like a point source (at small source lengths and / or large observation distances), a region where the source behaves like a line source (at large source lengths and / or small observation distances), and a transition region with intermediate behavior.

[0124] Thus, for a uniform rectangular 2D surface source with width L1 and height L2, which can be considered as consisting of a continuous distribution of vertical line sources of length L2 as described above, the SPL of each of these vertical line sources (at observation distance D) can be determined from Equation 3 disclosed in the provisional application.

[0125] The total pressure response of the 2D surface source at observation distance D is now calculated as 1 / r (corresponding to the point source pressure response) in Equation 1 provided above. i The distance dependence can be simulated by substituting the distance-dependent attenuation model in Equation 3 from the provisional application. In other words, a 2D surface source can be modeled as a one-dimensional distribution of point-like sources of size L1, each with a distance attenuation function corresponding to a finite-length line source of size L2.

[0126] Performing such simulations for uniform 2D rectangular sources of various sizes leads to the conclusion that the 1D model of Eqs. 6-8 is also valid for these sources when the observation distance is smaller than the height L, by simply applying a simple scaling factor α to the obtained effective spatial extent, and therefore Eq. 6 is modified to: TIFF2026041760000019.tif27170

[0127] For a diffuse 2D rectangular source, the scaling factor α is a monotonic function of the ratio between the source height L2 and the observation distance D. The table below provides values ​​for α as a function of L2 / D obtained from simulations.

[0128] Table: Scaling factor α as a function of the ratio of source height L2 to observation distance D. TIFF2026041760000020.tif14170

[0129] For more arbitrarily shaped 2D and 3D extents and / or non-uniform power distributions, the same qualitative concept of effective spatial extent according to embodiments of the present disclosure still applies.

[0130] For basic 2D and 3D geometric spreading shapes with uniform power distribution (e.g., circle, sphere, cylinder, rectangle, box), it is entirely feasible to create specific parametric models for the effective spatial spread as a function of observation distance similar to the models above.

[0131] Parametric Model Implementation

[0132] In some embodiments, the audio renderer determines the effective spatial extent of a volumetric audio source based on received information about the geometric extent (e.g., physical size), shape, and / or other characteristics of the source. In such embodiments, the parametric models described above may be implemented in the audio renderer, and the renderer determines the transition distance and effective spatial extent from the parametric model(s), the received source information, and the listener distance.

[0133] In some embodiments, the renderer may receive parameters for configuring the parametric model(s), for example, in the bitstream. In particular, parameters c0, c1, and c2 related to selected SPL threshold sound level difference values ​​used by the model may be received by the renderer.

[0134] Other model parameters that can be sent to the audio renderer as source-specific metadata in the bitstream are:

[0135] (1) Coherence data for volumetric sources, which instructs the renderer which version of the parametric model (diffuse, coherent broadband, or coherent frequency-dependent) to use, or specifies a mix of model versions (possibly frequency-dependent) (e.g., using a coherent frequency-dependent model for low frequencies, a diffuse model for high frequencies, and a mix of these two models at mid-frequencies).

[0136] (2) A flag indicating that the source should be considered "infinitely long." In this case, the renderer may ignore the source's geometric extent data to determine the effective spatial extent, and may always use the formula for distances smaller than the transition distance to determine the effective spatial extent for the source.

[0137] (3) A flag that instructs the renderer whether to use the effective spatial extent model for the source. Using a model for a particular volumetric source is not always appropriate or desirable. This can be the case, for example, for a volumetric source that does not radiate sound from its full extent, but is simply a conceptual volume containing a limited number of individual sound sources.

[0138] In other embodiments, the parametric model(s) may be implemented outside the renderer, for example, in the encoder. In such scenarios, the transition distance and / or the effective spatial extent for an observation distance smaller than the transition distance (which is constant with respect to aperture angle for diffuse one-dimensional audio sources) is transmitted to the renderer. In these embodiments, the renderer therefore does not need to implement the parametric model(s), but simply needs to be able to switch between two "spatial extent modes" for rendering the source: one mode in which the renderer uses the received geometric extent (constant with respect to absolute size) for rendering, and an alternative mode in which the renderer uses the received effective spatial extent (constant with respect to angle), with the received transition distance used as the selection criterion between the two modes.

[0139] As briefly mentioned above, quantitative parametric models assume uniform source power distribution across the expanse. This limits the applicability of quantitative models to sources that are "reasonably" uniform, but there are many relevant types of sources that meet this criterion (e.g., busy highways, ocean coastlines, high-speed trains, etc.).

[0140] In the above disclosure, only a central observation position was considered. However, for non-central observation positions, the same qualitative conceptual model still applies. For very long sources (specifically infinitely long sources), the lateral observation position is irrelevant for the effective spatial extent, and therefore the quantitative parametric model applies to any observation position.

[0141] The provisional application describes how the coherence properties of volumetric sources and the processing of partially coherent volumetric sources are determined.

[0142] Exemplary Systems and / or Methods

[0143] 5 shows an example system 500 for rendering an audio source according to some embodiments of the present disclosure. The system 500 includes an encoder 501 and an audio renderer 502. The audio renderer 502 includes an effective spatial extent calculation module 526 and an audio rendering module 528. Optionally, the audio renderer 502 may also include a spatial extent calculation module 522 and a threshold distance value calculation module 524.

[0144] In system 500, renderer 502 receives audio input signal 512 and audio source metadata 514 from encoder 501. Metadata 514 may include any one or combination of (i) coherence information associated with the audio source, (ii) spatial extent data or geometry information associated with the audio source, and / or (iii) threshold distance values ​​required to calculate the effective spatial extent of the audio source. The coherence information indicates coherence properties of the audio source, for example, indicating that the audio source is a coherent source or a diffuse source. The geometry information indicates the geometry of the audio source.

[0145] If the metadata 514 includes spatial extent data and a threshold distance value, the effective spatial extent calculation module 526 calculates the effective spatial extent of the audio source based on the spatial extent data, the threshold distance value, and the distance value specifying the distance between the audio source and the listener, and the rendering module 528 renders the audio source using the effective spatial extent of the audio source.

[0146] If the metadata 514 includes spatial extent data but not a threshold distance value, the threshold distance value calculation module 524 calculates a threshold distance value based on the received spatial extent data. The effective spatial extent calculation module 526 then calculates an effective spatial extent of the audio source based on the spatial extent data, the threshold distance value, and the distance value specifying the distance between the audio source and the listener, and the audio rendering module 528 renders the audio source using the effective spatial extent of the audio source.

[0147] If the metadata 514 does not include spatial extent data or a threshold distance value but does include geometry information, the spatial extent calculation module 522 calculates the spatial extent based on the geometry information, and the threshold distance value calculation unit 524 calculates a threshold distance value based on the calculated spatial extent. Then, the effective spatial extent calculation module 526 calculates an effective spatial extent of the audio source based on the calculated spatial extent data, the calculated threshold distance value, and a distance value specifying the distance between the audio source and the listener, and the audio rendering module 528 renders the audio source using the effective spatial extent of the audio source.

[0148] FIG. 6 shows an example renderer 502 for creating sound for an XR scene. The system 600 includes a controller 601, a signal modifier 602 for modifying an audio signal 651 (e.g., a multi-channel audio signal), a left speaker 604, and a right speaker 605. While one audio signal and two speakers are shown in FIG. 6 , this is merely for illustrative purposes and does not limit the embodiments of the present disclosure in any way. The controller 601 may be configured to receive one or more parameters and trigger the signal modifier 602 to implement modifications to the audio signal 651 (e.g., increase or decrease the volume level) based on the received parameters. The received parameters include (1) information 653 regarding the listener's position (e.g., direction and distance to the audio source) and (2) metadata 514 regarding the audio objects described herein.

[0149] In some embodiments of the present disclosure, information 653 may be provided from one or more sensors included in XR system 700 shown in FIG. 7A. As shown in FIG. 7A, XR system 700 is configured to be worn by a user. As shown in FIG. 7B, XR system 700 may include an orientation sensing unit 701, a position sensing unit 702, and a processing unit 703 coupled to controller 601 of system 600. Orientation sensing unit 701 is configured to detect changes in the listener's orientation and provide information regarding the detected changes to processing unit 703. In some embodiments, processing unit 703 determines an absolute orientation (with respect to some coordinate system) given the detected change in orientation detected by orientation sensing unit 701. There may also be different systems for determining orientation and position, e.g., the HTC Vive system using a lighthouse tracker (lidar). In one embodiment, orientation sensing unit 701 may determine an absolute orientation (with respect to some coordinate system) given the detected change in orientation. In this case, the processing unit 703 may simply multiplex the absolute orientation data from the orientation sensing unit 701 and the absolute position data from the position sensing unit 702. In some embodiments, the orientation sensing unit 701 may comprise one or more accelerometers and / or one or more gyroscopes.

[0150] 8 is a flowchart illustrating a process 800, according to one embodiment, for rendering an audio source for a listener. Process 800 may begin at step s802 and may be performed by renderer 502. Step s802 includes obtaining at least a first spatial extent value indicating a first spatial extent of the audio source. Step s804 includes obtaining a distance value specifying a distance between the audio source and the listener. Step s806 includes determining whether the distance value is less than a threshold distance value. Step s808 includes rendering the audio source for the listener using the effective spatial extent value as a result of determining that the distance value is less than the threshold distance value.

[0151] In some embodiments, the threshold distance value is a function of the first spatial extent value. In some embodiments, the effective spatial extent value is a function of the distance value.

[0152] In some embodiments, the effective spatial extent value is proportional to a power of the distance value, where the power has a value between 0.5 and 1.

[0153] In some embodiments, the process 800 further includes obtaining coherence property information. The coherence property information indicates a degree of coherence for the audio source. Thus, the coherence property information can be used to make a decision as to whether the audio source is a coherent source, a diffuse source, or a mixture thereof.

[0154] In some embodiments, the process 800 further includes calculating an effective spatial extent value based on the obtained coherence property information.

[0155] In some embodiments, the process 800 further includes determining whether the audio source is a diffuse source or a coherent source based on the obtained coherence property information.

[0156] If the source is a diffuse source, calculating the effective spatial spread value includes calculating the effective spatial spread value based on C0×D, where C0 is a constant and D is the obtained distance value.

[0157] If the source is a coherent source, calculating the effective spatial spread value is TIFF2026041760000021.tif13170, where C1 is a constant and D is the obtained distance value.

[0158] In some embodiments, the effective spatial extent value is used to identify segments of the audio source, the identified segments of the audio source being acoustically relevant segments of the audio source for a listener.

[0159] In some embodiments, obtaining the first spatial extent value includes receiving, from the encoder, metadata associated with the audio source, the metadata including geometry information associated with the audio source, and obtaining the first spatial extent value further includes deriving the first spatial extent value based on the geometry information included in the metadata.

[0160] In some embodiments, process 800 further includes receiving metadata associated with the audio source, the metadata including (i) a flag indicating that the size of the audio source is essentially infinite, and / or (ii) a flag instructing whether an effective spatial extent model should be used to render the audio source.

[0161] In some embodiments, rendering the audio source includes determining positions for one or more virtual loudspeakers based on the effective spatial spread value, and using the one or more virtual loudspeakers to render the audio source.

[0162] In some embodiments, the audio source is essentially a one-dimensional (1D) audio source.

[0163] In some embodiments, the audio source is a two-dimensional (2D) audio source or a three-dimensional (3D) audio source, and process 800 includes receiving metadata from the encoder that includes a flag indicating whether a 1D effective spatial extent model should be used to render the 2D audio source or the 3D audio source.

[0164] In some embodiments, the audio sources are two-dimensional (2D) audio sources (i.e., the audio sources have a first spatial extent in a first spatial dimension (e.g., width) and the audio sources have a second spatial extent in a second spatial dimension (e.g., height)), or three-dimensional (3D) audio sources (i.e., the audio sources have a first spatial extent in a first spatial dimension (e.g., width), a second spatial extent in a second spatial dimension (e.g., height), and a third spatial extent in a third spatial dimension (e.g., depth)), and process 800 includes determining whether a 1D effective spatial extent model described herein can be used to render the 2D or 3D audio sources and / or when to switch between (i) rendering the 2D or 3D audio sources according to the 1D model and (ii) using a more complex 2D or 3D model. As described above, the determination of whether a 2D or 3D audio source can be rendered using the 1D effective spatial extent of the audio source may be based on the size of one or two other dimensions and the observation distance. For example, given a 2D audio source with a width (L) of 50 meters and a height (H) of 1 meter, the render may be configured such that, based on H and the observation distance (e.g., based on determining that the observation distance > H), the render determines that the audio source can be rendered as a 1D audio source with an effective length of Leff, where Leff <Lである。

[0165] Thus, in some embodiments, the first spatial extent of the audio source is spatial extent in a first spatial dimension, and the method further includes: i) obtaining a second spatial extent value indicative of a second spatial extent of the audio source, the second spatial extent being spatial extent in the second spatial dimension; and ii) determining whether to derive the effective spatial extent value as if the audio source had spatial extent in only one spatial dimension. In some embodiments, determining whether to derive the effective spatial extent value as if the audio source had spatial extent in only one spatial dimension includes receiving a flag indicating that the effective spatial extent value may be derived as if the audio source had spatial extent in only one spatial dimension. In some embodiments, determining whether to derive an effective spatial extent value as if the audio source had spatial extent in only one spatial dimension comprises determining whether i) a difference between the first or second spatial extent value and a distance value is greater than a threshold value, or ii) a difference between the first or second spatial extent value and a value that is a function of the distance value is greater than a threshold value. In some embodiments, if the audio source is a diffuse audio source, the method comprises determining whether a difference between the first or second spatial extent value and the distance value is greater than a threshold value; and if the audio source is not a diffuse audio source, the method comprises determining whether a difference between the first or second spatial extent value and a value that is a function of the distance value is greater than a threshold value. In some embodiments, determining whether a difference between the first or second spatial extent value and the distance value is greater than a threshold value comprises determining whether the distance value is greater than the first or second spatial extent value.

[0166] 9 is a block diagram of an apparatus 900, according to some embodiments, for implementing system 500 or a portion of system 500 (e.g., renderer 502) and / or system 600. As shown in FIG. 9, apparatus 900 includes a processing circuit (PC) 902 that may include one or more processors (P) 955 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.), which may be co-sited in a single housing or in a single data center or may be geographically distributed (i.e., apparatus 900 may be a distributed computing device), and at least one network interface 948, each network interface 948 configured to enable apparatus 900 to communicate with the network interface The PC 902 may comprise at least one network interface 948, a transmitter (Tx) 945, and a receiver (Rx) 947 for enabling the network interface 948 to transmit data to and receive data from other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 948 is connected (directly or indirectly) (e.g., the network interface 948 may be wirelessly connected to the network 110, in which case the network interface 948 is connected in an antenna configuration), and one or more storage units (a.k.a., "data storage system") 908, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 902 includes a programmable processor, a computer program product (CPP) 941 may be provided. The CPP 941 includes a computer-readable medium (CRM) 942, which stores a computer program (CP) 943 comprising computer-readable instructions (CRI) 944. The CRM 942 may be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, or a memory device (e.g., a random access memory, a flash memory).In some embodiments, the CRI 944 of the computer program 943, when executed by the PC 902, is configured such that the CRI causes the device 900 to perform the steps described herein (e.g., steps described herein with reference to flowcharts). In other embodiments, the device 900 may be configured to perform the steps described herein without the need for code. That is, for example, the PC 902 may simply consist of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.

[0167] The above-described embodiments provide at least some advantages. For example, using an effective spatial extent according to embodiments of the present disclosure for a large volumetric audio source allows for a more natural and realistic spatial rendering of the source than using the geometric extent of the source directly. Also, compared to directly specifying the intended perceived spatial extent of a volumetric audio source, which is valid only for a specific listening position, the modeled effective spatial extent according to embodiments of the present disclosure is valid at any listening position. Furthermore, in some rendering scenarios, methods and systems according to embodiments of the present disclosure allow for better computational efficiency in rendering audio for large volumetric audio sources, because only the portion of the geometric extent that is acoustically relevant at a given listening position is considered in rendering. As another example, in embodiments of the present disclosure, the parametric model for determining the effective spatial extent is quite simple and can be easily implemented as a lightweight add-on to existing render architectures.

[0168] While various embodiments have been described herein, it should be understood that these embodiments have been presented by way of example only, and not limitation. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments. Moreover, unless otherwise indicated herein or clearly contradicted by context, any combination of the above-described elements in all possible variations thereof is encompassed by the present disclosure.

[0169] Additionally, while the process and message flows described above and illustrated in the figures have been depicted as a sequence of steps, this has been done for illustrative purposes only. As such, it is contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.

Claims

1. A method (800) for rendering an audio source for a listener, the method comprising: obtaining (s802) at least a first spatial extent value indicative of a first spatial extent of the audio source; obtaining (s804) a distance value specifying a distance between the audio source and the listener; determining whether the distance value is less than a threshold distance value (s806); rendering the audio source to the listener using an effective spatial extent value as a result of determining that the distance value is less than the threshold distance value (s808); and The method (800) includes:

2. The method of claim 1 , wherein the effective spatial extent value is a function of the distance value.

3. The method of claim 1 , further comprising receiving metadata comprising the effective spatial extent value.

4. The method of claim 3 , wherein the effective spatial extent value is an aperture angle value.

5. The method of claim 1 , wherein the threshold distance value is a function of the first spatial extent value.

6. 6. The method of claim 1, wherein the effective spatial extent value is proportional to a power of the distance value, the power having a value between 0.5 and 1, inclusive.

7. The method of claim 1 , further comprising obtaining coherence property information, the coherence property information indicating a degree of coherence for the audio source.

8. The method of claim 7 , further comprising calculating the effective spatial extent value based on the obtained coherence property information.

9. The method comprises: determining whether the audio source is a diffuse source, a coherent source, or a mixture of diffuse and coherent sources based on the coherence factor for the audio source; The method of claim 8 further comprising:

10. If the source is a diffuse source, calculating the effective spatial extent value may be performed by 0 ×D, wherein C 0 is a constant and D is the distance value, 10. The method according to any one of claims 1 to 9.

11. If the source is a coherent source, Calculating the effective spatial extent value comprises: wherein C 1 is a constant and D is the distance value, 10. The method according to any one of claims 1 to 9.

12. 12. The method of claim 1, wherein the effective spatial extent values ​​are used to identify segments of the audio source, the segments of the audio source being acoustically relevant segments of the audio source for the listener.

13. The method of claim 12 , wherein rendering the audio source comprises rendering only the identified segment of the audio source.

14. 13. The method of claim 1, wherein obtaining the first spatial extent value comprises: (i) receiving, from an encoder, metadata associated with the audio source, the metadata including geometry information associated with the audio source; and (ii) deriving the first spatial extent value based on the geometry information included in the metadata.

15. 15. The method of claim 1, further comprising receiving metadata associated with the audio source, the metadata including (i) a flag indicating that the size of the audio source is essentially infinite, and / or (ii) a flag instructing whether to use an effective spatial extent model to render the audio source, and / or (iii) the threshold distance.

16. Rendering the audio source comprises: determining positions for one or more virtual loudspeakers based on the effective spatial extent values; using the one or more virtual loudspeakers to render the audio source; 16. The method of any one of claims 1 to 15, comprising:

17. 17. The method of claim 1, wherein the audio source is an essentially one-dimensional (1D) audio source.

18. the first spatial extent of the audio source is a spatial extent in a first spatial dimension; The method comprises: obtaining a second spatial extent value indicative of a second spatial extent of the audio source, the second spatial extent being a spatial extent in a second spatial dimension; determining whether the effective spatial extent value should be derived as if the audio source had spatial extent in only one spatial dimension; 17. The method of any one of claims 1 to 16, further comprising:

19. 20. The method of claim 18, wherein determining whether to derive the effective spatial extent value as if the audio source had spatial extent in only one spatial dimension comprises receiving a flag indicating that the effective spatial extent value can be derived as if the audio source had spatial extent in only one spatial dimension.

20. Determining whether to derive the effective spatial extent value as if the audio source had spatial extent in only one spatial dimension comprises: i) whether the difference between the first spatial extent value or the second spatial extent value and the distance value is greater than a threshold value; or ii) whether the difference between the first spatial extent value or the second spatial extent value and a value that is a function of the distance value is greater than a threshold value; 20. The method of claim 18, comprising determining:

21. If the audio source is a diffuse audio source, the method includes determining whether the difference between the first spatial extent value or the second spatial extent value and the distance value is greater than a threshold; If the audio source is not a diffuse audio source, the method includes determining whether the difference between the first spatial extent value or the second spatial extent value and the value that is a function of the distance value is greater than a threshold.

21. The method of claim 20.

22. 22. The method of claim 20 or 21, wherein determining whether the difference between the first spatial extent value or the second spatial extent value and the distance value is greater than a threshold value comprises determining whether the distance value is greater than the first spatial extent value or the second spatial extent value.

23. A computer program (943) comprising instructions (944) which, when executed by a processing circuit (902), cause said processing circuit to perform the method of any one of claims 1 to 22.

24. 24. A carrier containing the computer program of claim 23, said carrier being one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (942).

25. An apparatus (900) for rendering an audio source for a listener, said apparatus comprising: obtaining a spatial extent value indicative of the spatial extent of the audio source (s802); obtaining (s804) a distance value specifying a distance between the audio source and the listener; determining whether the distance value is less than a threshold distance value (s806); rendering the audio source to the listener using an effective spatial extent value as a result of determining that the distance value is less than the threshold distance value (s808); and An apparatus (900) configured to:

26. 26. The apparatus of claim 25, wherein the apparatus is further configured to perform the method of any one of claims 2 to 22.

27. An apparatus (900) for rendering an audio source for a listener, said apparatus comprising: A memory (942); a processing circuit (902) coupled to said memory; wherein the processing circuitry causes the device to: obtaining a spatial extent value indicative of the spatial extent of the audio source (s802); obtaining (s804) a distance value specifying a distance between the audio source and the listener; determining whether the distance value is less than a threshold distance value (s806); rendering the audio source to the listener using an effective spatial extent value as a result of determining that the distance value is less than the threshold distance value (s808); and An apparatus (900) configured to: