Spatial range modeling for volumetric audio sources
The effective spatial range of the volume audio source is determined through the parameter model, which solves the problems of unnatural rendering of large volume audio sources in virtual reality scenes and achieves a natural and efficient rendering effect.
Patent Information
- Application Number
- CN202510296452.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-22
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to efficiently and naturally render the spatial range of large-volume audio sources in virtual reality, augmented reality or mixed reality scenarios, resulting in high computational complexity and unnatural rendering effects.
A simple parameter model is used to determine the effective spatial range of the volume audio source through size, distance, coherence characteristics and frequency parameters, and only the acoustic related parts are rendered to reduce the computational burden.
It realizes natural rendering effect at any listening position, and improves the computing efficiency of rendering large-volume audio sources, reducing rendering complexity.
Smart Images

Figure CN120144085A_ABST
Abstract
Description
Division Case Explanation
[0001] This application is a divisional application of the invention patent application with the application date of July 22, 2020, application number 202080104750.7, and invention title "Spatial Extent Modeling for Volumetric Audio Sources". Technical Field
[0002] Embodiments related to methods and systems for spatial extent modeling of volumetric audio sources are disclosed. Background Art
[0003] XR (virtual reality, augmented reality, or mixed reality) scenes can contain many audio sources spatially distributed within the space of the scene. Many of these audio sources have specific, well-defined positions in space and can be considered point sources. These audio sources are typically rendered as point audio objects to the listener.
[0004] However, XR scenes typically also contain audio sources with volumetric characteristics rather than point-like characteristics, meaning they have a specific spatial extent in one or more spatial dimensions.
[0005] In some cases, such volumetric audio sources can correspond to a single physical entity in the scene (e.g., an airplane, a piano, a train, a transportation pipeline in a factory, etc.). Some of these volumetric audio sources can radiate audio as a single coherent audio source, while other volumetric audio sources can radiate audio more like a spatially extended diffuse audio source.
[0006] In other cases, instead of corresponding to a single physical entity, a volumetric audio source can represent a region in the scene that contains a large (and possibly even continuous) number of independent audio sources that can be considered together as a composite volumetric audio source. Examples of this type of volumetric audio source are the seashore at a beach and a busy highway. In the example of a busy highway, although each vehicle is in principle an independent audio source, the highway with many cars on it can be considered a composite volumetric audio source.
[0007] Similar to the seashore and highway examples discussed above, in many cases, the spatial extent of a volumetric audio source can be very large in one or more of its spatial dimensions, and in some cases, it may even be effectively infinite (e.g., related to the distance from the listener to the volumetric audio source). Summary of the Invention
[0008] Typically, scene description data for XR scenarios specify the extent of a volumetric audio source based on the physical geometry of the source (e.g., the physical dimensions of the source in one or more dimensions, or a geometric mesh structure that describes the physical geometry of the source). This specified physical geometry of the source is typically directly related to the physical geometry of some corresponding physical (and usually also visual) entity in the XR scenario (e.g., a car, a piano, etc.).
[0009] However, as noted above, a volumetric audio source can be physically very large in one or more dimensions, such as the seashore and a busy highway discussed above. In such cases, the physical size or geometry of the volumetric audio source as typically specified in its extent data is generally not very suitable for directly rendering the audio source to a listener.
[0010] Specifically, in many cases, only a limited portion of the geometric extent of the volumetric audio source significantly contributes to the audio energy received by a listener at a given listening position. This is the case for very large (especially for "infinitely" large) volumetric sources, where the outer portions of the geometric extent are so far from the listener that no significant audio energy reaches the listener from these outer portions due to distance and medium attenuation.
[0011] It can also be the case that if a listener is close to a medium-sized volumetric audio source, the audio energy reaching the listener from the portion of the source close to the listener substantially overwhelms the audio energy from the portion far from the listener. Thus, the "acoustically relevant" portion of a given volumetric audio source can depend on the position of the listener relative to the source.
[0012] Therefore, for large volumetric audio sources, simply using the specified geometric extent as a direct measure of how wide or tall the audio source should be rendered to a listener at a given listening position is generally not very appropriate or convenient. In fact, doing so can lead to various problems.
[0013] One problem with directly using the specified geometric extent of a volumetric audio source for audio rendering is that the final subjective spatial extent of the source (e.g., the size of the source perceived by the listener) can be unnatural (e.g., unnaturally wide - the spatial extent may be perceived as wider than in real-life situations). For example, this problem can occur in a rendering scenario where virtual speakers located at the edges of the specified geometric extent of the source are used to render the audio of the volumetric audio source to a listener. As noted above, these virtual speakers are spaced too wide in many cases.
[0014] Specifying the expected perceived spatial extent of the source instead of its geometric extent is also problematic because the expected perceived spatial extent is only valid for one specific listening position, and deriving the expected perceived spatial extent for other listening positions (as required in 6-degree-of-freedom XR use cases) may not be straightforward or even possible.
[0015] In addition, in a rendering scenario where advanced physical modeling techniques are used to accurately render audio radiated by a volumetric audio source, the computational complexity required for rendering typically increases rapidly as the physical size of the source increases. For large volumetric sources (e.g., the seashore and busy highway discussed above, and a passing train), using the specified geometric extent of the source directly for rendering the source to a listener may easily require a large amount of computational work, especially in real-time interactive XR applications. In addition, a large portion of this computational work may even be spent unnecessarily because it is used to render audio radiated by parts of the volumetric source that do not significantly contribute to the audio at a particular listener location.
[0016] Accordingly, in order to be able to render large volumetric audio sources in a perceptually appropriate and computationally efficient manner, it would be highly beneficial to have a method for modeling the acoustically relevant spatial extent of a volumetric audio source at a given listening location based on the specified geometric extent of the source and other possible attributes. It is particularly desirable if the model is a very simple model such that it can be implemented as a lightweight add-on to an existing real-time renderer architecture.
[0017] Embodiments of the present disclosure relate to methods and systems for providing a very low-complexity parametric model to determine the effective spatial extent of a volumetric audio source (i.e., a portion of the geometric extent of the volumetric audio source that significantly contributes to the audio received by a listener at a given listening location).
[0018] The parameters of the model include: (i) dimension parameters indicating the dimensions of the geometric extent of the volumetric audio source in one or more dimensions, and / or (ii) distance parameters indicating the distance from the listener to the volumetric audio source.
[0019] The parameters of the model may further include: (i) a parameter indicating the coherence characteristics (e.g., coherent, diffuse, or in between) of the volumetric audio source, and / or (ii) a frequency parameter (in the case of a (partially) coherent source).
[0020] The determined effective spatial extent can be used in rendering the volumetric audio source to a listener. For example, the determined effective spatial extent can be used for: (i) determining the target auditory rendering size of the volumetric audio source at a given listening location, and / or (ii) only selecting the acoustically relevant sub-portions of the geometric extent of the volumetric audio source for rendering the volumetric source to a listener, while discarding other acoustically non-relevant portions of the geometric spatial extent for rendering.
[0021] In one aspect, a method for rendering an audio source for a listener is provided. The method includes: obtaining a spatial extent value indicative of a spatial extent of the audio source and obtaining a distance value specifying a distance (also referred to as an "observation distance") between the audio source and the listener. The method further includes: determining whether the distance value is less than a threshold distance value. The method further includes: as a result of determining that the distance value is less than the threshold distance value, rendering the audio source to the listener using an effective spatial extent value.
[0022] In another aspect, a computer program is provided. The computer program includes instructions which, when executed by a processing circuit, cause the processing circuit to perform the method of any one of the embodiments disclosed herein. In another aspect, a carrier containing the computer program is provided, wherein the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0023] In another aspect, a device is provided that is adapted to perform the method of any one of the embodiments disclosed herein. In one embodiment, the device includes: a processing circuit; and a memory containing instructions executable by the processing circuit, whereby the device is adapted to perform the method of any one of the embodiments disclosed herein.
[0024] Advantages
[0025] For large-volume audio sources, using the effective spatial extent according to embodiments of the present disclosure achieves a more natural and realistic spatial rendering of the source than directly using the geometric extent of the source.
[0026] Furthermore, compared to directly specifying the expected perceived spatial extent of a volume audio source, which is only valid for a specific listening position, the modeled effective spatial extent according to embodiments of the present disclosure is valid at any listening position.
[0027] In addition, in some rendering scenarios, the methods and systems according to embodiments of the present disclosure achieve better computational efficiency in rendering the audio of large-volume audio sources because only the acoustically relevant portion of the geometric extent at a given listening position is considered in the rendering.
[0028] In addition, in embodiments of the present disclosure, the parametric model for determining the effective spatial extent is very simple and can be easily implemented as a lightweight add-on to existing rendering architectures. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings incorporated herein and forming a part of the specification illustrate various embodiments.
[0030] Figure 1 Different parameters for audio rendering are shown.
[0031] Figure 2Shows the behavior of the sound pressure level (SPL) of a completely incoherent or diffuse one-dimensional volume audio source according to an embodiment.
[0032] Figure 3 Shows the behavior of the SPL of a coherent one-dimensional volume audio source according to an embodiment.
[0033] Figure 4 Shows the behavior of the SPL of a coherent one-dimensional volume audio source according to an embodiment.
[0034] Figures 5 to 7B Shows a system according to some embodiments.
[0035] Figure 8 Is a process according to an embodiment.
[0036] Figure 9 Is a device according to an embodiment. Detailed Description
[0037] Effective Spatial Range: Qualitative Model
[0038] As described above, it is desirable to find a model that can infer the acoustically relevant or effective spatial range of a volume audio source for a listener at a given listening position from some physical characteristics of the source.
[0039] A reasonable general starting assumption for developing such a model is that portions of the volume audio source that do not significantly contribute to the perceived loudness of the source at a particular listening position (in the sense that the listener cannot perceive any difference in loudness whether the audio from these portions is included in the rendering or not) will also not significantly contribute to the perceived spatial characteristics of the source (including the perceived spatial range of the source) at the same listening position.
[0040] The starting point for constructing the model is the simple case of a one-dimensional volume audio source (i.e., a line source) that has a variable geometric spatial range of size L in a single dimension (i.e., an acoustic line source with variable length L). For this line source, the behavior of the sound pressure level (SPL) at an observation point O located at a perpendicular distance D from the midpoint of the line source can be evaluated according to the source length L (see Figure 1 )
[0041] If the length L of the line source is very small, the line source is substantially a point source. As the length L steadily increases on both sides of the line source with a constant source intensity density along the length (e.g., at length 2L, the source radiates twice the acoustic energy radiated at length L), it is expected that the SPL at the observation point O also increases.
[0042] However, as the source length L increases, the contribution to the SPL at the observation point from the outer edges of the source becomes smaller and smaller, as the distance from these outer portions to the observation point increases. Thus, the rate at which the SPL increases with increasing length L may decrease. At some point, when the source becomes very long, the rate at which the SPL increases may become negligible, such that the SPL will no longer increase with further increase in the source length. In other words, the SPL will saturate when exceeding a specific length (Leff) of the line source.
[0043] For an infinitely long line source with a constant source strength density, at different observation points O, different segments of the infinitely long line source with a size of Leff can be the acoustically significant parts of the source. In other words, when a listener moves along a line parallel to the infinitely long line source, the listener effectively perceives the source through a spatial window with a size of Leff that moves with the listener.
[0044] In some embodiments of the present disclosure, the effective spatial extent of a one-dimensional volume audio source (i.e., a line source) is defined as the size of the smallest source segment, such that the sound level at a given listening position caused by this segment is less than a threshold sound level difference below the sound level of the complete source. In other words, adding portions of the line source beyond the effective spatial extent does not add more than the threshold sound level difference to the sound level at the listening position.
[0045] The threshold sound level difference used in determining the effective spatial extent can be selected in different ways, but can most conveniently be defined in relative terms (e.g., as a specific percentage of the linear sound pressure caused by the complete source or a specific number of decibels below the SPL of the complete source).
[0046] Since the goal is to link the physical size of the source to the corresponding perceived auditory size, it is desirable to select the threshold sound level difference such that it is related to this perceived auditory size. Specifically, it may be desirable to select the threshold sound level difference such that the resulting effective spatial extent corresponds to the size of the smallest source segment for which the loudness (as a perceived metric) is indistinguishable from the loudness of the complete source. In this context, therefore, the perception-related criterion for setting the threshold sound level difference is the just noticeable difference (JND) of loudness, which is known from the acoustic perception literature to be approximately 1 dB SPL.
[0047] Effective Spatial Extent: Quantitative Model
[0048] A volume audio source can be physically modeled as a dense distribution of point sources. For a one-dimensional volume audio source, the total pressure response can be expressed as:
[0049] (1)
[0050] where is the total number of point sources used to model the one-dimensional volume audio source, is the complex amplitude of the i-th point source at the radial frequency and is the wave number , where c is the speed of sound in air, and is the distance from the i-th point source to the observation point .
[0051] Then the sound pressure level (SPL) of a one-dimensional volume source can be expressed as:
[0052] (2)
[0053] When modeling a continuous volume source in this way, a small enough spacing between individual point sources should be used to obtain accurate results over the entire frequency range of interest (e.g., 0 - 20 kHz).
[0054] There are various types of one-dimensional volume audio sources, including coherent one-dimensional volume audio sources (where all points radiate the same acoustic signal coherently) and diffuse one-dimensional volume audio sources (where all points radiate independent, completely uncorrelated signals). These two extreme types of one-dimensional volume audio sources behave significantly the same in various aspects, as described in more detail below.
[0055] Diffuse one-dimensional volume audio source
[0056] Since a diffuse volume audio source can be regarded as a dense distribution of independent point sources, and all independent point sources have SPL - distance behavior independent of frequency, the behavior of a diffuse volume audio source is also independent of frequency. Therefore, the following results are valid for any single frequency and for broadband.
[0057] Figure 2 shows simulation results of the total SPL of a completely incoherent or diffuse one-dimensional volume audio source according to the line source length for observation distances of 0.1, 1, 10, and 100 m .
[0058] A common characteristic of all four curves is that they have two different regions: at small line source lengths (or at large observation distances), for every doubling of the line source length L, the SPL increases at a constant rate of 3 dB (note the logarithmic scale on the horizontal axis), while at large line source lengths (or at small observation distances), the SPL becomes constant according to L.
[0059] At small line source lengths (or at large observation distances), the increase in SPL by 3 dB for every doubling of the length can be explained by the fact that in the representation of the total pressure (Equation 1), the distances from the individual point sources to the observation point are substantially equal, such that only the complex amplitude Relates to determining the total pressure. If it is assumed that the source powers of individual point sources are equal, then the following relationship can be easily shown to hold from Equations 1 and 2:
[0060]
[0061] Equation 3 implies that when the number of point sources N is doubled, the SPL increases by 3 dB. If it is further assumed that the spacing between individual point sources is uniform, then Equation 3 implies that for a doubling of the line source length L, the SPL increases by 3 dB.
[0062] At large line source lengths (or small observation distances), the SPL saturates to a constant value, which is consistent with the qualitative model discussed above. As mentioned above, the saturation of the SPL can be explained by the fact that as the length increases, the contribution of newly added point sources at the outer edges becomes increasingly unimportant and eventually completely negligible. In other words, once the diffused one-dimensional volume audio source has reached a specific line source length, further increasing the length does not result in a further increase in the SPL.
[0063] Comparing the curves for different observation distances indicates that the line source length L at which the transition occurs between the two regions (the region where the SPL increases and the region where the SPL remains essentially at a constant value) depends on the observation distance, where for larger observation distances, the transition length is larger.
[0064] The effective spatial range of a very long diffused one-dimensional volume audio source can be estimated from the curves shown by finding the line source length L at which the SPL is a specific threshold sound level difference below the saturated SPL observed at very large source lengths. The line source length at which the SPL is found to saturate is proportional to the observation distance Figure 2
[0065] If the JND of loudness difference is chosen as the criterion for setting the threshold sound level difference, then the threshold sound level difference is equal to approximately 1 dB SPL. Then, the line source length at which the SPL is found to saturate (i.e., the -1 dB point in this specific case) has a value of approximately 6 times the observation distance
[0066] Therefore, the effective spatial range of a source with a length of 6D or longer is equal to 6D (i.e., it is proportional to the observation distance D). Equivalently, at observation distances less than
[0067] For line source lengths less than 6D or equivalent observation distances greater than the effective spatial range is simply equal to the line source length L (i.e., each part of the one-dimensional volume source contributes significantly to the sound received at the observation point). This property enables more effective and realistic audio source rendering, where the audio source behaves like a one-dimensional diffused volume source.
[0068] The effective spatial extent can be expressed (i) in terms of length or (ii) in terms of angular span ("opening angle"). If the effective spatial extent is expressed in terms of length, the effective spatial extent for a line source length greater than the saturation length is proportional to the observation distance D (e.g., for some audio sources, the effective spatial extent is equal to 6D). In contrast, when the effective spatial extent is expressed in terms of angular span ("opening angle (OA)"), the effective spatial extent has a constant value (i.e., is independent of the observation distance). The general expression for the opening angle (OA) is: OA = 2 * atan ((0.5 * effective spatial extent in units of length) / D), where atan is the arctangent function. For a diffuse source, the effective spatial extent in units of length is 6D, so the expression for OA becomes: OA = 2 * atan ((0.5 * 6D) / D) = 2 * atan (3) = 143 degrees. Thus, if the renderer is given OA and D, the renderer can calculate the effective spatial extent in units of length.
[0069] Figure 2 Another common characteristic of the curves shown is that in the region where the SPL increases by 3 dB for every doubling of the length, for a 10-fold increase in the observation distance, the SPL decreases by 20 dB, or equivalently, the SPL decreases by 6 dB for every doubling of the distance. This means that in this region, the one-dimensional diffuse volume audio source behaves similarly to a point source in terms of SPL (i.e., ). In contrast, in the region where the SPL is constant with respect to the line source length, for every 10-fold increase in distance, the SPL only decreases by 10 dB, or the SPL decreases by 3 dB for every doubling of the distance. This means that in this region, the one-dimensional diffuse volume source behaves similarly to a theoretical line source (i.e., ). This distance-dependent SPL behavior of a finite-length line source is described in U.S. Provisional Patent Application No. 62 / 950,272, filed December 19, 2019.
[0070] Coherent one-dimensional volume audio source
[0071] For a fully coherent uniform one-dimensional volume audio source, all the amplitudes in Equation 1 are the same (i.e., . Due to the frequency dependence and coherence of the phase terms of the individual point sources of the volume audio source, the total pressure response of the volume audio source will also be frequency dependent, and thus it is necessary to analyze the effective spatial extent at individual frequencies as well as broadband.
[0072] Figure 3Shows the SPL response according to the line source length for various frequencies and an observation distance. Similar to the diffuse one-dimensional volume audio source, for the coherent one-dimensional volume audio source, for each of the individual frequencies, there is an expected saturation of the SPL beyond a specific line source length, but unlike the Figure 2 curve shown, the saturation length now depends on the frequency.
[0073] At small line source lengths, for each doubling of the length, the SPL increases at a rate of 6 dB, rather than 3 dB as for the diffuse one-dimensional volume audio source as Figure 2 shown. Following the same reasoning as in the case of the diffuse source, for small line source lengths (or large observation distances), the distances from the individual point sources to the observation point are substantially equal, and the pressure of each point source of the volume audio source is all the same as the common pressure such that (as mentioned before, assuming equal power and spacing of the individual point sources) Equation 1 simplifies to:
[0074]
[0075] Resulting in:
[0076]
[0077] This is indeed consistent with the observed increase of 6 dB for each doubling of the length L.
[0078] Based on the analysis of the corresponding simulation results for many other observation distances and using the same method as for the diffuse source to determine the effective spatial range, for small observation distances (or large line source lengths), the effective spatial range is frequency-dependent and can be expressed as equal to (where f is the frequency and is a constant), while for large observation distances (or small line source lengths), the effective spatial range is again simply equal to the line source length L.
[0079] The transition distance between the two regions is found to be proportional to where the proportionality factor is equal to . For a specific choice, as mentioned before, using the JND of loudness difference (1 dB SPL) as the threshold sound level difference for finding the saturation length, it is found empirically that is equal to approximately 18.4.
[0080] As shown above, the behavior of a coherent one-dimensional volumetric audio source is frequency-dependent and will generally also be rendered in a frequency-dependent manner. To observe the broadband behavior of a coherent one-dimensional volumetric audio source, simulations can be performed at 128 evenly-spaced frequencies from 20 Hz to 20 kHz, and the results for all individual frequencies can be summed to obtain a broadband result, which implies a white source spectrum assumption.
[0081] Figure 4 Shows the broadband SPL versus line source length for several observation distances Figure 4 and Figure 2 (which shows the SPL of a diffuse source that is independent of frequency). The overall broadband behavior of the coherent and diffuse sources is quite similar, especially for very small and very large line source lengths. The main differences are: (1) the transition region of the coherent source is much wider, and (2) there are some ripples within the transition region of the coherent source. Additionally, at large observation distances, the SPL of the coherent source increases by 6 dB for every doubling of the line source length (as observed for each individual frequency), rather than the 3 dB increase in the SPL of the diffuse source for every doubling of the line source length.
[0082] For small observation distances (or large line source lengths), the broadband effective spatial extent is proportional to the square root of the observation distance D, i.e., , while for large observation distances (or small source lengths), it simply equals the line source length L again. The transition distance between the two regions is proportional to the square of the line source length L, where the proportionality factor equals . For a specific choice of using a JND of loudness difference (1 dB SPL) as the threshold sound level difference for finding the saturation length, it was found that is approximately equal to 3.5.
[0083] Summary of the results from the simulations
[0084] For observation distances greater than a specific transition distance (also known as the threshold distance value) (or line source lengths less than a specific transition length), the effective spatial extent of the one-dimensional volumetric audio source equals the line source length L.
[0085] For observation distances less than the transition distance (or for line source lengths greater than the transition length), in the case of a diffuse one-dimensional volumetric audio source, the effective spatial extent is proportional to the observation distance D, while for a coherent one-dimensional volumetric audio source, it is proportional to the square root of D.
[0086] For a diffuse one-dimensional volumetric audio source, the transition distance is proportional to the source length L, while for a coherent one-dimensional volumetric audio source, the transition distance is proportional to the square of L.
[0087] Parametric model of the effective spatial extent
[0088] The simulation results discussed above lead to the following parametric models for the effective spatial extent of a one-dimensional volumetric audio source as a function of the observation distance :
[0089] Diffuse one-dimensional volume source:
[0090] (6)
[0091] Coherent one-dimensional volume source, frequency-dependent:
[0092] (7)
[0093] Coherent one-dimensional volume source, wideband:
[0094] (8)
[0095] For a specific choice of a threshold sound level difference of 1 dB (JND of loudness), the constants have been found empirically to have the following approximate values: , and .
[0096] Numerical examples
[0097] The following are examples showing how the parametric models of the effective spatial extent affect the rendering of a one-dimensional volumetric audio source.
[0098] Example 1:
[0099] For an "infinitely" long diffuse source (e.g., the shoreline of a beach), the effective spatial extent will be within " " for a listener at any practically relevant observation distance. Thus, the effective spatial extent will be: 6 m (143 degrees) at an observation distance of 1 m; 60 m (143 degrees) at an observation distance of 10 m; and 600 m (143 degrees) at an observation distance of 100 m.
[0100] Example 2:
[0101] For a diffuse one-dimensional volume source with a length L = 10 m, the effective spatial extent will be: 0.6 m (143 degrees) at an observation distance of 0.1 m, 6 m (143 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 1.7 m (which results in 53 degrees at a 10 m observation distance and 6 degrees at a 100 m observation distance).
[0102] Example 3:
[0103] For a coherent one-dimensional volume source with length L = 10 m, the broadband effective spatial extent will be: 1.1 m (160 degrees) at an observation distance of 0.1 m, 3.5 m (121 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 8.2 m.
[0104] Example 4:
[0105] For a coherent one-dimensional volume source with length L = 10 m, the effective spatial extent will be:
[0106] At f = 100 Hz: 1.8 m (85 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 30 m;
[0107] At f = 1000 Hz: 0.6 m (32 degrees) at an observation distance of 1 m, and 10 m at any observation distance greater than 300 m.
[0108] Utilizing the effective spatial extent in rendering volumetric audio sources
[0109] The renderer can use the exported effective spatial extent in various ways.
[0110] Setting the target spatial extent:
[0111] The exported effective spatial extent can be used to set the target spatial extent for rendering a long volumetric audio source to a listener at a specific listening position. This will convey a more appropriate rendered source width to the listener compared to simply using the received geometric extent data. For example, in a scenario, the exported effective spatial extent can be used to determine the optimal position of a virtual stereo speaker for rendering a source to a listener at a specific listening position. In another scenario, the exported effective spatial extent can be used to set the target spatial width in a spatial widening algorithm for rendering a volumetric audio source to a listener at a specific listening position.
[0112] Determining the spatial window:
[0113] For very long volumetric audio sources, the exported effective spatial extent can be used to determine at what times to render which part of the source to a listener moving along, away from, and / or towards the source. This is analogous to applying a sliding spatial window with the listener, where the exported effective spatial extent (which is dynamically updated based on the listener's position) determines the width of the spatial window.
[0114] Saving computational power:
[0115] In use cases where some form of physical modeling is used to render audio from a volumetric audio source, computational power can be saved by using the exported valid spatial extent to limit the portion of the source that needs to be rendered to a listener at a specific listening location.
[0116] Extension to 2D and 3D volumetric audio sources
[0117] The one-dimensional quantitative parameter model discussed above is effective at least for volumetric audio sources that have a significant spatial extent in no more than one spatial dimension, meaning that the extent in the other two dimensions is small enough relative to the observation distance such that the extent in these dimensions does not significantly affect the effective spatial extent in the primary (long) dimension. In particular, this will be the case if, at a particular observation distance, the source behaves substantially like a point source in the other two dimensions.
[0118] The above provisional patent application describes a model for determining when a one-dimensional audio source behaves like a point source based on the source length and the observation distance. Specifically, the document describes that a diffuse one-dimensional audio source behaves like a point source at observation distances that exceed the length of the source.
[0119] Thus, if a diffuse 2D volumetric audio source has two dimensions (dimension 1 and dimension 2) with two sizes, where the size in dimension 1 is longer than the size in dimension 2, then if the size in dimension 2 is less than the observation distance D (or if the observation distance D is greater than ), the one-dimensional quantitative effective spatial extent model of Equation 6 can be applied to dimension 1 of the 2D source.
[0120] Similar criteria can also be obtained for the validity of the one-dimensional Equations 7 and 8 for a coherent 2D volumetric source. In this case, if the size in dimension 2 is less than (1) (for the frequency-dependent model of Equation 7) or (2) (for the broadband model of Equation 8) (or equivalently, if the observation distance is greater than (for the frequency-dependent model) or (for the broadband model)), then the one-dimensional model is valid.
[0121] For example, if we have a 2D diffuse source with a width of 10m and a height of 1m, then if the observation distance is greater than 1m, its effective spatial extent (in the long dimension) can be calculated according to Equation 6 (where L = 10m). For a fully coherent two-dimensional source of the same dimensions, if the observation distance is greater than 1.5m, then the effective spatial extent (in the long dimension) at 500Hz can be calculated according to Equation 8.
[0122] These examples illustrate that in order for the one-dimensional quantitative model of equations 6-8 to be applicable, the requirement for the "one-dimensionality" of the volume source is rather loose, and that the one-dimensional model can actually be applied to a wide range of "long" 2D (and 3D) volume sources.
[0123] It should be noted that the quantitative criteria for the validity of the one-dimensional model described above should be understood as indicative rather than as an exact boundary between regions where the one-dimensional model is applicable and inapplicable in terms of the effective spatial range. It provides a means of identifying the types of 2D sources to which the one-dimensional model can be applied, and / or the conditions under which a given 2D source can be modeled by the one-dimensional model.
[0124] Accordingly, an additional feature of embodiments of the present disclosure is that the renderer can determine whether to apply the one-dimensional model to a 2D or 3D volume source based on the above criteria, and / or when to switch between rendering a 2D or 3D volume source according to the simplified one-dimensional model and using a more complex 2D or 3D model.
[0125] For volume audio sources that have significant extent in more than one dimension (and thus do not meet the criteria for being considered a one-dimensional audio source as described above), the same qualitative models and principles apply.
[0126] Generally speaking, increasing the geometric spatial extent in one dimension has the effect of increasing the effective spatial extent in other dimensions. For example, in a comparison between a purely one-dimensional source with length L (i.e., a line source) and a two-dimensional source with length L and height H (i.e., a surface source), for a given observation distance D, the effective extent along the dimension of length L of the two-dimensional source will be greater than that of the one-dimensional source. Additionally, for the surface source, the transition distance between the region where the effective spatial extent is a function of the observation distance and the region where it simply equals the physical dimensions of the geometric extent is less than that of the line source. In other words, for a 2D surface source, the entire width of the source needs to be considered in the rendering at a shorter distance than for an equally wide line source.
[0127] In the above example, the two-dimensional surface source can be considered to be composed of a continuum of vertical line sources of size H distributed along length L (as opposed to point sources in the one-dimensional model discussed above). Each of these vertical line sources has a less steep distance attenuation than that of a point source, and thus points along the horizontal extent of the source need to be further away from the observation position to become insignificant (according to a given SPL-based significance criterion, e.g., the 1 dB SPL loudness JND criterion discussed above). As a result, for a two-dimensional surface source, the spatial extent in each dimension will be greater than in each of the individual two dimensions. It should be noted that the effective spatial extent in each dimension is still limited by the geometric dimensions of the extent in that dimension (i.e., the effective spatial extent will never exceed the geometric dimensions).
[0128] Note that the effective spatial extent in each dimension is still limited by the geometric dimensions of the extent in that dimension (i.e., the effective spatial extent will never exceed the geometric dimensions).
[0129] The above description applies to a rectangular 2D surface source where the source power is distributed more or less uniformly over the surface. For such 2D sources, a simple extension of the one-dimensional quantitative model of equations 6 - 8 can be constructed as described below.
[0130] The provisional application discloses a parametric model of the distance-dependent SPL attenuation function for a one-dimensional sound source of finite length, essentially identifying three different regions of observation distance where the source behaves respectively like a point source (at small source length and / or large observation distance), a line source (at large source length and / or small observation distance), and a transition region with intermediate behavior.
[0131] Then, for a uniform rectangular 2D surface source with width L 1 and height L 2 as described above, it can be considered to be composed of a continuous distribution of vertical line sources of length L 2 and the SPL of each of these vertical line sources (at the observation distance D) can be determined according to equation 3 disclosed in the provisional application.
[0132] The total pressure response of the 2D surface source at the observation distance D can now be simulated by replacing the distance dependence (corresponding to the point source pressure response) in equation 1 provided above with the distance-dependent attenuation model of equation 3 from the provisional application. In other words, the 2D surface source can be modeled as a one-dimensional distribution of point sources of size L 1 each with a distance attenuation function corresponding to a finite length line source of size L 2
[0133] Running this simulation for uniform 2D rectangular sources of various sizes leads to the conclusion that the one-dimensional model of equations 6 - 8 is also valid for these sources, provided that in the case where the observation distance is less than the height L 2 a simple scaling factor is applied to the resulting effective spatial range such that equation 6 is modified to:
[0134]
[0135]
[0136] For a diffusive 2D rectangular source, the scaling factor is a monotonic function of the ratio between the height L 2 of the source and the observation distance D. The following table provides the values of as obtained from the simulation according to :
[0137] Table 1: According to the source height L 2 Scale factor as a ratio to the observation distance D .
[0138]
[0139] For more arbitrarily shaped 2D and 3D extents and / or non-uniform power distributions, the same qualitative concept of the effective spatial extent according to embodiments of the present disclosure still applies.
[0140] For basic 2D and 3D geometric extent shapes with uniform power distributions (e.g., circles, spheres, cylinders, rectangles, boxes), it is entirely feasible to have a specific parametric model of the effective spatial extent as a function of the observation distance similar to the models described above.
[0141] Implementation of the parametric model
[0142] In some embodiments, the audio renderer determines the effective spatial extent of a volumetric audio source based on the received information about the geometric extent (e.g., physical dimensions), shape, and / or other characteristics of the source. In such embodiments, the above parametric model can be implemented in the audio renderer, and the renderer determines the transition distance and the effective spatial extent based on the parametric model, the received source information, and the listener distance.
[0143] In some embodiments, the renderer can receive, for example in a bitstream, parameters for configuring the parametric model. Specifically, the parameter c related to the selected SPL threshold sound level difference used by the model 0 , c 1 and c 2 can be received by the renderer.
[0144] Other model parameters that can be sent to the audio renderer as source-specific metadata in a bitstream are:
[0145] (1) Coherence data of the volumetric source indicating which version of the parametric model (diffuse, coherent broadband, or coherent frequency-dependent) the renderer should use or specifying a mixture of versions of the model (possibly frequency-dependent) (e.g., using a coherent frequency-dependent model for low frequencies, a diffuse model for high frequencies, and a mixture of both models for mid frequencies).
[0146] (2) A flag indicating that the source should be considered "infinitely long". In this case, the renderer can ignore the geometric extent data of the source to determine the effective spatial extent and can always use the equation for distances less than the transition distance to determine the effective spatial extent of the source.
[0147] (3) A flag indicating whether the renderer uses the effective spatial extent model of the source. It may not always be appropriate or necessary to use a model for a specific volume source. For example, this may be the case for a volume source that represents a conceptual volume containing a finite number of individual sound sources and does not radiate sound from the entire extent of the volume source.
[0148] In other embodiments, the parametric model may be implemented outside the renderer (e.g., in the encoder). In such a scenario, the transition distance and / or the effective spatial extent for viewing distances less than the transition distance (which is constant in terms of the angular spread in the case of a diffused one-dimensional audio source) are sent to the renderer. In these embodiments, the renderer does not need to implement the parametric model as such, but only needs to be able to switch between the following two "spatial extent modes" for rendering the source: one mode that uses the received geometric extent for rendering (which is constant in absolute dimensions), and an alternative mode that uses the received effective spatial extent (which is constant with respect to the angle), where the received transition distance is used as a selection criterion between the two modes.
[0149] As briefly mentioned above, the quantitative parametric model assumes a spatially uniform source power distribution. Although this limits the application of the quantitative model to "reasonably" uniform sources, there are many relevant types of sources that meet this criterion (e.g., a busy highway, a coastline, a high-speed train, etc.).
[0150] In the above disclosure, only the central viewing position was considered. However, in the case of a non-central viewing position, the same qualitative conceptual model still applies. For very long sources (specifically, infinitely long sources), the lateral viewing position is independent of the effective spatial extent, and thus the quantitative parametric model applies to any viewing position.
[0151] This provisional application describes how to determine the coherence characteristics of a volume source and the processing of a partially coherent volume source.
[0152] Exemplary system and / or method
[0153] Figure 5 An example system 500 for rendering an audio source in accordance with some embodiments of the present disclosure is shown. System 500 includes an encoder 501 and an audio renderer 502. The audio renderer 502 includes an effective spatial extent calculation module 526 and an audio rendering module 528. Optionally, the audio renderer 502 may also include a spatial extent calculation module 522 and a threshold distance value calculation module 524.
[0154] In system 500, a renderer 502 receives an audio input signal 512 and audio source metadata 514 from an encoder 501. The metadata 514 can include any one or a combination of the following: (i) coherence information associated with the audio source, (ii) spatial extent data or geometric information associated with the audio source, and / or (iii) a threshold distance value required to calculate the effective spatial extent of the audio source. The coherence information indicates the coherence characteristics of the audio source, which for example indicates whether the audio source is a coherent source or a diffuse source. The geometric information indicates the geometry of the audio source.
[0155] If the metadata 514 includes spatial extent data and a threshold distance value, an effective spatial extent calculation module 526 calculates the effective spatial extent of the audio source based on the spatial extent data, the threshold distance value, and a distance value specifying the distance between the audio source and the listener, and a rendering module 528 uses the effective spatial extent of the audio source to render the audio source.
[0156] If the metadata 514 includes spatial extent data but does not include a threshold distance value, a threshold distance value calculation module 524 calculates a threshold distance value based on the received spatial extent data. Then, the effective spatial extent calculation module 526 calculates the effective spatial extent of the audio source based on the spatial extent data, the threshold distance value, and a distance value specifying the distance between the audio source and the listener, and an audio rendering module 528 uses the effective spatial extent of the audio source to render the audio source.
[0157] If the metadata 514 includes neither spatial extent data nor a threshold distance value but includes geometric information, a spatial extent calculation module 522 calculates a spatial extent based on the geometric information, and a threshold distance value calculation unit 524 calculates a threshold distance value based on the calculated spatial extent. Then, the effective spatial extent calculation module 526 calculates the effective spatial extent of the audio source based on the calculated spatial extent data, the calculated threshold distance value, and a distance value specifying the distance between the audio source and the listener, and an audio rendering module 528 uses the effective spatial extent of the audio source to render the audio source.
[0158] Figure 6 An example renderer 502 for generating sound for an XR scene is shown. System 600 includes a controller 601, a signal modifier 602 for modifying an audio signal 651 (e.g., a multi-channel audio signal), a left speaker 604, and a right speaker 605. Although Figure 6An audio signal and two speakers are shown, but this is for illustrative purposes only and does not limit the embodiments of the present disclosure in any way. The controller 601 may be configured to receive one or more parameters and trigger the signal modifier 602 to perform modifications on the audio signal 651 based on the received parameters (e.g., increasing or decreasing the volume level). The received parameters include (1) information 653 regarding the position of the listener (e.g., the direction and distance to the audio source) and (2) metadata 514 regarding the audio object as described herein.
[0159] In some embodiments of the present disclosure, the information 653 may be provided from Figure 7A one or more sensors included in the XR system 700 shown. As Figure 7A shown, the XR system 700 is configured to be worn by a user. As Figure 7B shown, the XR system 700 may include an orientation sensing unit 701, a position sensing unit 702, and a processing unit 703 coupled to the controller 601 of the system 600. The orientation sensing unit 701 is configured to detect changes in the orientation of the listener and provide information regarding the detected changes to the processing unit 703. In some embodiments, given a detected change in the orientation detected by the orientation sensing unit 701, the processing unit 703 determines the absolute orientation (relative to a certain coordinate system). There may also be different systems for determining orientation and position, e.g., the HTC Vive system using a lighthouse tracker (lidar). In one embodiment, given a detected change in the orientation, the orientation sensing unit 701 may determine the absolute orientation (relative to a certain coordinate system). In this case, the processing unit 703 may simply multiplex the absolute orientation data from the orientation sensing unit 701 and the absolute position data from the position sensing unit 702. In some embodiments, the orientation sensing unit 701 may include one or more accelerometers and / or one or more gyroscopes.
[0160] Figure 8 is a flowchart showing a process 800 for rendering an audio source for a listener according to one embodiment. The process 800 may start at step s802 and may be executed by the renderer 502. Step s802 includes: obtaining at least a first spatial range value indicating a first spatial range of the audio source. Step s804 includes: obtaining a distance value specifying the distance between the audio source and the listener. Step s806 includes: determining whether the distance value is less than a threshold distance value. Step s808 includes: as a result of determining that the distance value is less than the threshold distance value, rendering the audio source to the listener using an effective spatial range value.
[0161] In some embodiments, the threshold distance value is a function of the first spatial range value. In some embodiments, the effective spatial range value is a function of the distance value.
[0162] In some embodiments, the effective spatial extent value is proportional to a power of the distance value, where the value of the power is between 0.5 and 1.
[0163] In some embodiments, process 800 further includes: obtaining coherence characteristic information. The coherence characteristic information indicates the degree of coherence of the audio source. Thus, the coherence characteristic information can be used to determine whether the audio source is a coherent source, a diffuse source, or a mixture thereof.
[0164] In some embodiments, process 800 further includes: calculating an effective spatial extent value based on the obtained coherence characteristic information.
[0165] In some embodiments, process 800 further includes: determining whether the audio source is one of a diffuse source or a coherent source based on the obtained coherence characteristic information.
[0166] In the case where the source is a diffuse source, calculating the effective spatial extent value includes: based on to calculate the effective spatial extent value, where is a constant, and D is the obtained distance value.
[0167] In the case where the source is a coherent source, calculating the effective spatial extent value includes: based on to calculate the effective spatial extent value, where is a constant, and D is the obtained distance value.
[0168] In some embodiments, the effective spatial extent value is used to identify a segment of the audio source, where the identified segment of the audio source is the acoustically relevant segment of the audio source for the listener.
[0169] In some embodiments, obtaining a first spatial extent value includes: receiving metadata associated with the audio source from an encoder. The metadata includes geometric information associated with the audio source. Obtaining the first spatial extent value further includes: deriving the first spatial extent value based on the geometric information included in the metadata.
[0170] In some embodiments, process 800 further includes receiving metadata associated with the audio source, where the metadata includes: (i) a flag indicating that the size of the audio source is substantially infinite, and / or (ii) a flag indicating whether to use the effective spatial extent model to render the audio source.
[0171] In some embodiments, rendering the audio source includes: determining the positions of one or more virtual speakers based on the effective spatial extent value; and using the one or more virtual speakers to render the audio source.
[0172] In some embodiments, the audio source is substantially a one-dimensional (1D) audio source.
[0173] In some embodiments, the audio source is a two-dimensional (2D) audio source or a three-dimensional (3D) audio source, and process 800 includes: receiving metadata from an encoder, the metadata including a flag indicating whether to use a 1D effective spatial extent to render the 2D or 3D audio source.
[0174] In some embodiments, the audio source is a two-dimensional (2D) audio source (i.e., the audio source has a first spatial extent in a first spatial dimension (e.g., width) and the audio source has a second spatial extent in a second spatial dimension (e.g., height)) or a three-dimensional (3D) audio source (i.e., the audio source has a first spatial extent in a first spatial dimension (e.g., width), a second spatial extent in a second spatial dimension (e.g., height), and a third spatial extent in a third spatial dimension (e.g., depth)), and process 800 includes: determining whether the 1D effective spatial extent model described herein can be used to render the 2D or 3D audio source, and / or when to switch between (i) rendering the 2D or 3D audio source according to the 1D model and (ii) using a more complex 2D or 3D model. As described above, it can be determined whether the 1D effective spatial extent of the audio source can be used to render the 2D or 3D audio source based on the dimensions of one or two other dimensions and the viewing distance. For example, given a 2D audio source having a width (L) of 50 meters and a height (H) of 1 meter, the rendering can be configured such that based on H and the viewing distance (e.g., based on determining that the viewing distance > H), the rendering determines that the audio source can be rendered as a 1D audio source having an effective length of Leff, where Leff < L.
[0175] Thus, in some embodiments, the first spatial extent of the audio source is a spatial extent in a first spatial dimension, and the method further comprises: i) obtaining a second spatial extent value indicative of a second spatial extent of the audio source, the second spatial extent being a spatial extent in a second spatial dimension, and ii) determining whether to derive a valid spatial extent value as if the audio source had a spatial extent in only one spatial dimension. In some embodiments, determining whether to derive a valid spatial extent value as if the audio source had a spatial extent in only one spatial dimension includes: receiving a flag indicating that the valid spatial extent value can be derived as if the audio source had a spatial extent in only one spatial dimension. In some embodiments, determining whether to derive a valid spatial extent value as if the audio source had a spatial extent in only one spatial dimension includes determining: i) whether a difference between the first spatial extent value or the second spatial extent value and a distance value is greater than a threshold, or ii) whether a difference between the first spatial extent value or the second spatial extent value and a value that is a function of the distance value is greater than a threshold. In some embodiments, if the audio source is a diffuse audio source, the method includes determining whether a difference between the first spatial extent value or the second spatial extent value and the distance value is greater than the threshold, and if the audio source is not a diffuse audio source, the method includes determining whether a difference between the first spatial extent value or the second spatial extent value and the value that is a function of the distance value is greater than the threshold. In some embodiments, determining whether a difference between the first spatial extent value or the second spatial extent value and the distance value is greater than the threshold includes: determining whether the distance value is greater than the first spatial extent value or the second spatial extent value.
[0176] Figure 9 is a block diagram of an apparatus 900 for implementing the system 500 or a part of the system 500 (e.g., the renderer 502) and / or the system 600 according to some embodiments. As Figure 9As shown, apparatus 900 may include: processing circuitry (PC) 902, which may include one or more processors (P) 955 (e.g., general purpose microprocessors and / or one or more other processors such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.), which may be co-located in a single housing or a single data center, or may be geographically distributed (i.e., apparatus 900 may be a distributed computing apparatus); at least one network interface 948, each network interface 948 including a transmitter (Tx) 945 and a receiver (Rx) 947 for enabling apparatus 900 to send data to and receive data from other nodes connected to network 110 (e.g., an Internet Protocol (IP) network), the network interface 948 being (directly or indirectly) connected to network 110 (e.g., the network interface 948 may be wirelessly connected to network 110, in which case the network interface 948 is connected to an antenna arrangement); and one or more storage units (also referred to as "data storage systems") 908, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 902 includes a programmable processor, a computer program product (CPP) 941 may be provided. CPP 941 includes a computer readable medium (CRM) 942 that stores a computer program (CP) 943 including computer readable instructions (CRI) 944. CRM 942 may be a non-transitory computer readable medium, e.g., a magnetic medium (e.g., a hard disk), an optical medium, a storage device (e.g., random access memory, flash memory), etc. In some embodiments, the CRI 944 of computer program 943 is configured such that when executed by PC 902, the CRI causes apparatus 900 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, apparatus 900 may be configured to perform the steps described herein without code. That is, for example, PC 902 may consist of only one or more ASICs. Thus, the features of the embodiments described herein may be implemented in hardware and / or software fashion.
[0177] The above embodiments provide at least several advantages. For example, for a large-volume audio source, using the effective spatial extent according to embodiments of the present disclosure enables a more natural and realistic spatial rendering of the source than directly using the geometric extent of the source. Additionally, compared to directly specifying the expected perceived spatial extent of a volume audio source, which is only valid for a specific listening position, the modeled effective spatial extent according to embodiments of the present disclosure is valid at any listening position. Further, in some rendering scenarios, the methods and systems according to embodiments of the present disclosure achieve better computational efficiency in rendering the audio of a large-volume audio source because only the acoustically relevant portion of the geometric extent at a given listening position is considered in the rendering. As another example, in embodiments of the present disclosure, the parametric model for determining the effective spatial extent is very simple and can be easily implemented as a lightweight add-on to existing rendering architectures.
[0178] Although various embodiments are described herein, it should be understood that they are presented by way of example and not limitation. Accordingly, the width and scope of the present disclosure should not be limited by any of the above exemplary embodiments. Additionally, any combination of the above elements in all possible variations is included in the present disclosure unless otherwise indicated or clearly conflicts with the context in some other way.
[0179] Additionally, although the processes and message flows described above and shown in the figures are shown as a series of steps, they are for illustrative purposes only. Thus, it is contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be executed in parallel.
Claims
1. A method (800) for rendering an audio source for a listener, the method comprises: obtaining at least (s802) a first spatial extent value indicating a first spatial extent of the audio source; obtaining (s804) a distance value specifying a distance between the audio source and the listener; determining (s806) whether the distance value is less than a threshold distance value; and as a result of determining that the distance value is less than the threshold distance value, rendering (s808) the audio source to the listener using an effective spatial extent value.
2. The method according to claim 1, wherein, the effective spatial extent value is a function of the distance value.
3. The method according to claim 1, further comprises: receiving metadata including the effective spatial extent value.
4. The method according to claim 3, wherein, the effective spatial extent value is an angular extent value.
5. The method according to any one of claims 1 to 4, wherein, the threshold distance value is a function of the first spatial extent value.
6. The method according to any one of claims 1 to 5, wherein, the effective spatial extent value is proportional to a power of the distance value, where the value of the power is between 0.5 and 1 and includes 0.5 and 1.
7. The method according to any one of claims 1 to 6, the method further comprises: obtaining coherence characteristic information, wherein the coherence characteristic information indicates a degree of coherence of the audio source.
8. The method according to claim 7, the method further comprises: calculating the effective spatial extent value based on the obtained coherence characteristic information.
9. The method according to claim 8, the method further comprises: determining, based on the degree of coherence of the audio source, whether the audio source is a diffuse source, a coherent source, or a mixture of a diffuse source and a coherent source.
10. The method according to any one of claims 1 to 9, wherein, If the source is a diffusion source, calculating the effective spatial range value includes: based on to calculate the effective spatial range value, where is a constant, and D is the distance value.
11. The method according to any one of claims 1 to 9, wherein, if the source is a coherent source, Then calculating the effective space range value includes: Based on to calculate the effective space range value, where is a constant, and D is the distance value.
12. The method according to any one of claims 1 to 11, wherein, the effective spatial extent value is used to identify a segment of the audio source, wherein the segment of the audio source is an acoustically relevant segment of the audio source for the listener.
13. The method according to claim 12, wherein, rendering the audio source comprises: rendering only the identified segment of the audio source.
14. The method according to any one of claims 1 to 12, wherein, obtaining the first spatial extent value comprises: (i) receiving metadata associated with the audio source from an encoder, wherein the metadata includes geometric information associated with the audio source; and (ii) deriving the first spatial extent value based on the geometric information included in the metadata.
15. The method according to any one of claims 1 to 14, the method further comprises receiving metadata associated with the audio source, wherein, the metadata includes: (i) a flag indicating that the dimensions of the audio source are substantially infinite, and / or (ii) a flag indicating whether an effective spatial extent model is used to render the audio source, and / or (iii) the threshold distance.
16. The method according to any one of claims 1 to 15, wherein, rendering the audio source includes: determining the positions of one or more virtual loudspeakers based on the effective spatial extent value, and using the one or more virtual loudspeakers to render the audio source.
17. The method according to any one of claims 1 to 16, wherein, the audio source is substantially a one-dimensional 1D audio source.
18. The method according to any one of claims 1 to 16, wherein, a first spatial extent of the audio source is a spatial extent in a first spatial dimension, and the method further includes: obtaining a second spatial extent value indicating a second spatial extent of the audio source, the second spatial extent being a spatial extent in a second spatial dimension; and determining whether to derive the effective spatial extent value as if the audio source had a spatial extent only in one spatial dimension.
19. The method according to claim 18, wherein, determining whether to derive the effective spatial extent value as if the audio source had a spatial extent only in one spatial dimension includes: receiving a flag indicating that the effective spatial extent value can be derived as if the audio source had a spatial extent only in one spatial dimension.
20. The method according to claim 18, wherein, determining whether to derive the effective spatial extent value as if the audio source had a spatial extent only in one spatial dimension includes determining: i) whether a difference between the first spatial extent value or the second spatial extent value and the distance value is greater than a threshold, or ii) whether a difference between the first spatial extent value or the second spatial extent value and a value that is a function of the distance value is greater than a threshold.
21. The method according to claim 20, wherein, if the audio source is a diffuse audio source, the method includes: determining whether a difference between the first spatial extent value or the second spatial extent value and the distance value is greater than a threshold, and if the audio source is not a diffuse audio source, the method includes: determining whether a difference between the first spatial extent value or the second spatial extent value and a value that is a function of the distance value is greater than a threshold.
22. The method according to claim 20 or 21, wherein, determining whether a difference between the first spatial extent value or the second spatial extent value and the distance value is greater than a threshold includes: determining whether the distance value is greater than the first spatial extent value or the second spatial extent value.
23. A computer program (943) comprising instructions (944) which, when executed by a processing circuit (902), cause the processing circuit to perform the method according to any one of claims 1 to 22.
24. A carrier comprising the computer program according to claim 23, wherein, the carrier is one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium (942).
25. An apparatus (900) for rendering an audio source for a listener, the apparatus being configured to: obtain (s802) a spatial extent value indicating the spatial extent of the audio source; Obtain (s804) a distance value specifying a distance between the audio source and the listener; Determine (s806) whether the distance value is less than a threshold distance value; and As a result of determining that the distance value is less than the threshold distance value, render (s808) the audio source to the listener using an effective spatial extent value.
26. The apparatus according to claim 25, wherein, the apparatus is further configured to perform the method according to any one of claims 2 to 22.
27. An apparatus (900) for rendering an audio source to a listener, the apparatus comprising: a memory (942); and a processing circuit (902), coupled to the memory, wherein the processing circuit is configured to cause the apparatus to: Obtain (s802) a spatial extent value indicating a spatial extent of the audio source; Obtain (s804) a distance value specifying a distance between the audio source and the listener; Determine (s806) whether the distance value is less than a threshold distance value; and As a result of determining that the distance value is less than the threshold distance value, render (s808) the audio source to the listener using an effective spatial extent value.