Apparatus and method for path length compensation for shortest paths on 8-connected grid maps

A lightweight post-processing step for path finding algorithms in VR/AR systems compensates for path length overestimations in 8-connected grid maps, addressing audible discontinuities and maintaining seamless audio transitions.

WO2025157406A1PCT designated stage Publication Date: 2025-07-31FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/051759
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing path finding algorithms for sound propagation in VR/AR systems using 8-connected grid maps result in audible discontinuities due to overestimation of path lengths, leading to perceptual implausibilities and audio artifacts, particularly when transitioning between direct and occluded areas.

Method used

A lightweight post-processing step for path finding algorithms like Dijkstra's or jump point search, which compensates for the overestimation of path lengths by calculating an offset based on the difference between Euclidean and 8-connected grid distances, ensuring seamless transitions and reducing audible discontinuities.

Benefits of technology

The proposed solution effectively minimizes audible discontinuities by compensating for path length overestimations, maintaining smooth audio transitions without significantly increasing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024051759_31072025_PF_FP_ABST
    Figure EP2024051759_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Fig. 1a illustrates an apparatus according to an embodiment. The apparatus comprises an audio metadata determiner (110) configured for determining audio metadata information depending on a first path and depending on a second path, wherein sound waves propagate from a first position to a second position, wherein the first path represents a shortest path from a first position to a second position connecting neighbouring grid pixels of a connected grid, wherein the second path represents a straight line from the first positon to the second position. The first position represents a sound source position where a sound source emits sound waves of the audio source signal and the second position represents a listener position where the sound waves are perceived; or wherein sound waves from the sound source position to the listener position propagate passing the first position and passing the second position.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Apparatus and Method for Path Length Compensation for Shortest Paths on 8-Connected Grid Maps Description The present invention relates to path length compensation of path lengths of paths along which of sound waves propagate. Auralization, e.g., real-time and offline audio rendering auditory scenes and environments, becomes increasingly important. In this context, Virtual Reality (VR) and Augmented Reality (AR) systems like the MPEG-I 6-DoF Immersive Audio renderer (see [1]) are of particular interest. AR / VR systems generate a visual and auditory virtual environment by visualizing and auralizing a virtual scene. The auralization process simulates the sound propagation in the virtual environment accounting for acoustic effects like occlusion, reflection, and diffraction of sound waves interacting with the surrounding geometric objects. An approach for simplifying the geometry of the acoustic environment is to use voxels (volumetric pixels), where the environment is discretized by a uniform grid of three- dimensional blocks. Some approaches may only consider a two-dimensional space, while not considering or ignoring a third dimension. For example, only two dimensions may be considered, while a third dimension (e.g., a height dimension) may, e.g., be assumed to a have a particular fixed value. Or, it may be assumed that all two-dimensional considerations apply to any (e.g., height) value of the third dimension, or at least to all height values in a particular range. E.g., it may be assumed that all two-dimensional considerations apply to a usual listener position as far as the third (height) dimension is considered (e.g., apply for height values in a range between 1.50 m – 2.10 m above ground level). Shortest path search algorithms like Dijkstra's algorithm, A*, or jump point search (JPS) may, e.g., be employed to determine the direction from which the first wave front is reaching the listener at a discretized position on a 2D map (see [2], [3], [4]). These algorithms determine the shortest path on an 8-connected uniform grid where the individual steps along the path are limited to 8 discrete directions given by the 8 neighbors on the (two-dimensional) grid map. Consequently, the computed path distance from the source to the listener deviates from the Euclidean distance. If the AR / VR renderer uses different distance metrics, for example an Euclidean distance for the direct sound and an 8-connected uniform grid distance for diffracted sound, then a transition between these methods can result in audible discontinuities. In the current draft version of the MPEG-I 6-DoF, Immersive Audio renderer, the Euclidean distance is used for rendering the direct sound while an 8-connected uniform grid distance is used for diffracted sound (see [1]). When the listener and / or the sound source moves between an area where a sound source is directly visible and an occluded area where the sound source is not directly visible, the different distance metrics result in a discontinuity of the estimated sound propagation path length. In the downstream auralization, these transitions can lead to perceptual implausibilities and audio artifacts. In case of the current draft version of the MPEG-I renderer, the audibility of this discontinuity is further amplified by the algorithm that is used for the determination of the diffraction angle as it is very susceptible to small differences between direct sound distance and diffracted path length:      If the diffraction path length ^^^^^^^^^^^^^^^^^^^^^^is overestimated for example by only 5%, then the estimated diffraction angle ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^deviates by 35.5°. Regarding multi-path voxel diffraction, diffraction EQs H(α) are derived from diffraction path lengths. Problem arise from path length discontinuities, which amount to audible discontinuities. The problems are caused from diffraction path lengths which quantized to their voxel centers, and are caused from that jump point search overestimates the path length. Fig. 2 illustrates an Euclidean distance (line 210) and an 8-connected grid distance (line 215) between two points on the uniform grid map, in particular, between source S and listener L. Fig. 3a illustrates how these distances change when the listener moves behind an occluding object (x-axis: audio frame index; y-axis: distance). In particular, Fig. 3a illustrates the direct sound distances 310 and the diffraction propagation path lengths 320 when the listener moves for a plurality of frames. The overestimation of the path length is visible by the offset between the Euclidean distance for the direct sound 310 and the 8- connected grid distance for the diffracted sound 320. Fig. 3b shows the diffraction angle 330, which is estimated from the ratio between the direct sound distance 310 and the diffraction propagation path length 320 according to (1). In particular, Fig.3b illustrates diffraction angles when a listener moves, which result from the direct sound distances and diffraction propagation path lengths of Fig.3a for a plurality of frames. It should be noted that when the listener just moves behind an occluding object at audio frame index 0, the diffraction angle is expected to be slightly less than 180° (“straight line”) and decrease with further listener movement. In Fig.3b, the angle at frame 0 is underestimated to be approximately 135.5°. As the estimated diffraction angle is used for selecting a tabulated equalizer-filter (EQ-filter) for the diffracted sound, the estimation error can result in clearly audible discontinuities. Fig. 3c illustrates the direct sound distances 310 and the diffraction propagation path lengths 320 of Fig.3a and the diffraction angles 330 of Fig.3b when a listener moves for a plurality of frames. In particular, Fig.3c illustrates quantized listener positions with jump point search resulting in EQs with quantization steps. The particular example of Fig. 3c depicts a chart for a downtown drummer; the voxel size is 25cm; and the scenario relates to a listener which moves behind the bus from right to center. There are path finding algorithms, which do not exhibit this overestimation. “Any-angle” path finding algorithms like for example Theta* do not suffer from the described problem (see [5]). But they are significantly more computational complex as they involve a large number of additional line-of-sight visibility checks. Visibility graphs can be used to speed up this process, but re-generating the visibility graph on any scene change, can be a computational burden for dynamic scenes. The same is true if a path smoothing step is applied to the result of one of the previously mentioned path finding algorithms. This post-processing method removes path nodes and hence reduces the path to larger segments, if the line-of-sight between 2 path nodes is not blocked (see [5]). This requires additional line-of-sight visibility tests and it is not guaranteed that the result is actually the shortest path in the “any-angle” sense (see [5]). The object of the present invention is to provide improved concepts for path length compensation of path lengths, in particular, of path lengths of paths along which of sound waves propagate. In particular, the present invention relates to an apparatus and a method for path length compensation, for example, for shortest paths on 8-connected grid maps. The object of the present invention is solved by the subject-matter of the independent claims. Particular embodiments are provided in the dependent claims. An apparatus according to an embodiment is provided. The apparatus comprises an audio metadata determiner configured for determining audio metadata information depending on a first path and depending on a second path, wherein sound waves propagate from a first position to a second position, wherein the first path represents a shortest path from a first position to a second position connecting neighbouring grid pixels of a connected grid, wherein the second path represents a straight line from the first positon to the second position. The first position represents a sound source position where a sound source emits sound waves of the audio source signal and the second position represents a listener position where the sound waves are perceived; or wherein sound waves from the sound source position to the listener position propagate passing the first position and passing the second position. Moreover, a method according to an embodiment is provided. The method comprises determining audio metadata information depending on a first path and depending on a second path, wherein sound waves propagate from a first position to a second position, wherein the first path represents a shortest path from a first position to a second position connecting neighbouring grid pixels of a connected grid, wherein the second path represents a straight line from the first positon to the second position. The first position represents a sound source position where a sound source emits sound waves of the audio source signal and the second position represents a listener position where the sound waves are perceived; or wherein sound waves from the sound source position to the listener position propagate passing the first position and passing the second position. Furthermore, a computer program is provided which configured to implement the above- described method when being executed on a computer or signal processor. In the state of the art, either no compensation or computationally more complex algorithms are used. Embodiments avoid these disadvantages. According to embodiments, concepts for reducing the audible discontinuities caused by the overestimation of the path length on the 8-connected grid map are provided. In embodiments, an offset is provided, which compensates the quantization error from Multi-Path CE in existing solutions. For example, the offset may, e.g., compensate an overestimation of the JPS path length. In embodiments, a lightweight post-processing step for path finding algorithms, for example, for Dijkstra's algorithm, A*, or jump point search, are provided. In the following, embodiments of the present invention are described in more detail with reference to the figures, in which: Fig.1a illustrates an apparatus according to an embodiment comprising an audio metadata determiner. Fig.1b illustrates an apparatus according to an embodiment further comprising a signal generator. Fig.2 illustrates an Euclidean distance and an 8-connected grid distance between a source and a listener. Fig.3a illustrates direct sound distances and diffraction propagation path lengths when a listener moves for a plurality of frames. Fig.3b illustrates diffraction angles when a listener moves, which result from the direct sound distances and diffraction propagation path lengths of Fig. 3a for a plurality of frames. Fig.3c illustrates the direct sound distances 310 and the diffraction propagation path lengths 320 of Fig.3a and the diffraction angles 330 of Fig.3b when a listener moves for a plurality of frames. Fig.4a illustrates a distance compensation concept according to an embodiment applied to the previously used example of Fig. 3a to Fig. 3c, where the listener moves behind an occluding object. Fig.4b illustrates the resulting diffraction angle 430 resulting from the embodiment of Fig.4a. Fig.4c illustrates the distance compensation concept according to the embodiment of Fig.4a and the resulting diffraction angles 430 of Fig.4b when a listener moves for a plurality of frames. Fig.5a illustrates quantized listener positions resulting in EQs with quantization steps. Fig.5b illustrates unquantized listener positions resulting in EQs without quantization steps according to an embodiment. Fig.6 illustrates an Euclidean distance, an 8-connected grid distance, and a diffracted sound path between a source and a listener. Fig.7 illustrates a schematic illustration of the data processing on a decoder side. Fig.1a illustrates an apparatus according to an embodiment. The apparatus comprises an audio metadata determiner 110 configured for determining audio metadata information depending on a first path and depending on a second path, wherein sound waves propagate from a first position to a second position, wherein the first path represents a shortest path from a first position to a second position connecting neighbouring grid pixels of a connected grid, wherein the second path represents a straight line from the first positon to the second position. The first position represents a sound source position where a sound source emits sound waves of the audio source signal and the second position represents a listener position where the sound waves are perceived; or wherein sound waves from the sound source position to the listener position propagate passing the first position and passing the second position. According to an embodiment, the audio metadata information may, e.g., depend on a difference between the first path and the second path. Fig. 1b illustrates an embodiment, wherein the apparatus further comprises a signal generator 120 configured for generating one or more audio output signals depending on an audio source signal and depending on the audio metadata information. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine audio metadata information indicating a difference between a length of the first path and a length of the second path, the length of the second path representing an Euclidean distance between the first position and the second position. In an embodiment, the first position represents the sound source position, wherein the second position represents the listener position, wherein the length of the second path represents the Euclidean distance between the sound source position and the listener position. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the length of the first path. The audio metadata determiner 110 may, e.g., be configured to determine the length of the second path. Moreover, the audio metadata determiner 110 may, e.g., be configured to determine the difference between the length of the first path and the length of the second path to determine the audio metadata information. In an embodiment, the connected grid may, e.g., be an 8-connected grid. One or more grid pixels of a plurality of grid pixels of the 8-connected grid exhibits exactly 8 neighbouring grid pixels of the plurality of grid pixels of the 8-connected grid. None of the plurality of grid pixels of the 8-connected grid comprises more than 8 neighbouring grid pixels of the plurality of grid pixels of the 8-connected grid. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the length of the second path ^^^^^ௗaccording to wherein ^^ௗ^^^^^^^ ൌ min^Δx, Δy^^^^௧^^^^^௧ ൌ max^Δx, Δy^ െ min^Δx, Δy^wherein Δx ൌ | lx- sx| Δy ൌ | ly- sy| wherein S indicates the sound source position, with S ൌ ^sx, sy^T, wherein sx indicates a first coordinate of the sound source position, wherein syindicates a second coordinate of the sound source position, wherein L indicates the listener position, with L ൌ ^lx, ly^T, wherein lxindicates a first coordinate of the listener position, wherein lyindicates a second coordinate of the listener position. In an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the difference between the length of the first path and the length of the second path according to or according to wherein wherein ^^^^^ௗindicates a length of a pixel grid of the 8-connected grid. According to an embodiment, the connected grid may, e.g., be a 26-connected grid. One or more grid pixels of a plurality of grid pixels of the 26-connected grid exhibits exactly 26 neighbouring grid pixels of the plurality of grid pixels of the 26-connected grid. None of the plurality of grid pixels of the 26-connected grid comprises more than 26 neighbouring grid pixels of the plurality of grid pixels of the 26-connected grid. In an embodiment, the signal generator 120 may, e.g., be configured to generate the one or more audio output signals depending on the audio source signal and depending on the compensated length of the diffraction path. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine a compensated length of a diffraction path from an uncompensated length of the diffraction path depending on the audio metadata information, wherein the diffraction path indicates an estimation of a length of a path along which diffracted sound waves propagate from the sound source position to the listener position along neighbouring grid pixels of the connected grid. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the compensated length of the diffraction path from the uncompensated length of the diffraction path by subtracting the difference between a length of the first path and a length of the second path from the uncompensated length of the diffraction path. In an embodiment, the audio metadata determiner 110 may, e.g., be configured to conduct unquantization of the compensated length of the diffraction path, such that a length difference of the compensated length of the diffraction path, which is caused by a movement of the sound source position and / or of the listener position, is smoothed. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the diffraction path, which exhibits the uncompensated length, using a shortest search path algorithm for determining the estimation of the path along which diffracted sound waves propagate from the sound source position to the listener position along neighbouring grid pixels of the connected grid. In an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the diffraction path, which exhibits the uncompensated length, using a Dijkstra's algorithm or an A* algorithm, or a jump point search algorithm as the shortest search path algorithm. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine a diffraction angle depending on the compensated length of the diffraction path between the sound source position and the listener position and depending on an Euclidean distance between the sound source position and the listener position. The signal generator 120 may, e.g., be configured to generate the one or more audio output signals depending on the audio source signal and depending on the compensated length of the diffraction path. In an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine the diffraction angle ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^according to: ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^ ൌ 2 ⋅ ^^^^^^^^^^^^^^^^^^^^^^^^^^^ / ^^^^^^^^^^^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^^ wherein ^^^^^^^^^^^^^^indicates the Euclidean distance between the sound source position and the listener position, wherein ^^^^^^^^^^^^^^^^^^^^^^indicates the diffraction path between the sound source position and the listener position, which exhibits the uncompensated length,wherein ^^^^^^^^^^^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^ indicates the diffraction path between the sound sourceposition and the listener position, which exhibits the compensated length, and wherein ^^^^^^^^^^^^ indicates the difference between a length of the first path and a length of the second path, wherein the length of the first path indicates the length of the shortest path from the sound source position to the listener position connecting neighbouring grid pixels of the connected grid, wherein the length of the second path indicates the Euclidean distance ^^^^^^^^^^^^^^between the sound source position and the listener position. According to an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine one or more gains depending on the diffraction angle. The signal generator 120 may, e.g., be configured to generate the one or more audio output signals depending on the one or more gains and depending on the audio source signal. In an embodiment, the audio metadata determiner 110 may, e.g., be configured to determine one or more equalizer filters and / or a delay depending on the diffraction angle. Moreover, the signal generator 120 may, e.g., be configured to generate the one or more audio output signals depending on the audio source signal and depending on the one or more equalizer filters and / or the delay. In an embodiment, the signal generator 120 may, e.g., be configured to generate the one or more audio output signals depending on the audio metadata information only if a difference between the uncompensated length of the diffraction path and the Euclidian distance between the sound source position and the listener position is smaller than a threshold value. Or, the signal generator 120 may, e.g., be configured to generate the one or more audio output signals depending on the audio metadata information only if a quotient between the uncompensated length of the diffraction path and the Euclidian distance between the sound source position and the listener position is smaller than a threshold value. According to an embodiment, the apparatus may, e.g., further comprise an audio decoder for decoding encoded audio signal to determine the audio source signal. The audio metadata determiner 110 may, e.g., be configured to determine the audio metadata information. The signal generator 120 may, e.g., be configured to generate the one or more audio output signals using the audio source signal, which has been decoded from the encoded audio signal, and using the audio metadata information. In an embodiment, the apparatus may, e.g., comprise an audio encoder for encoding an audio source signal as an encoded audio signal. The audio metadata determiner 110 may, e.g., be configured to determine the audio metadata information for a plurality of potential listener positions and for a plurality of potential sound source positions. The apparatus may, e.g., be configured to transmit the encoded audio signal and the audio metadata information to an audio decoder. Moreover, a system is provided. The system comprises said apparatus comprising the audio encoder and an apparatus for generating one or more audio output signals comprising an audio decoder and a signal generator 120. The apparatus comprising the audio encoder is configured to transmit the encoded audio signal and the audio metadata information to the apparatus for generating the one or more audio output signals. The audio decoder is configured to decode the encoded audio signal to obtain the audio source signal. The signal generator 120 is configured to obtain selected metadata information by selecting information from the audio metadata information for a current listener position and for a current sound source position. The signal generator 120 is configured to generate one or more audio output signals using the audio source signal and using the selected metadata information. In the following, particular embodiments of the present invention are described. An issue that shall be addressed is explained with reference to Fig. 2. As already explained, Fig.2 illustrates a sound source position S and a listener position L. The direct line of sight, depicted by line 210, touches a corner of the grey object of Fig.2. If the listener and / or the source now move down (e.g., sound source position S moves from [1, 6]Tto [1, 5]Tand / or, e.g., listener position L moves from [10, 3]Tto [10, 2]T, the grey object of Fig. 2 then shadows the direct line between sound source position S and listener position L. When such a situation occurs, state-of-the-art algorithms switch their distance calculation algorithms for determining the distance between the sound source position S and the listener position L from using Euclidian distance (when a direct line-of-sight exists) to determining the length of the path of the diffracted sound in the employed connected grid as a connected grid path (when no direct line-of-sight exists between S and L). As can be seen however in Fig. 2, the 8-connected grid path 215 is significantly longer than the direct line 210 connecting sound source position S and listener position L. As the distance calculation between S and L may, e.g., be used for determining a contribution of the sound source at sound source position S in the final audio output signal(s), this change of the employed distance calculation method results in undesired audible artifacts. For example, an amplitude of the sound source signal contribution in the audio output signal(s) may, e.g., depend on a determined path length of a sound path from the sound source position S to the listener position L. And / or a delay of the sound source signal contribution in the audio output signal may, e.g., also depends on the determined path length from the sound source position S to the listener position L. Some embodiments of the present invention now compensate the difference that is introduced by the immediate switching from the Euclidian distance calculation to the connected grid path length calculation. In particular embodiments, an offset is calculated that results from the switching of the calculation concepts. In the example of Fig. 2, the difference between the 8-connected grid length 215 between sound source position S and listener position L and the Euclidian distance 210 between sound source position S and listener position L may, e.g., be determined. The calculated offset may, e.g., then be applied (e.g., subtracted) from the diffracted sound path which is determined as an 8- connected grid distance. This results in that when immediately switching from the Euclidian distance calculation to the diffracted sound path calculation, a seamless switch is realized. With reference to Fig. 6 however, it should be noted that what is compensated in embodiments is the length difference between line 615 and line 610, e.g., between a line 615 that approximates the Euclidian distance in the connected grid and the actual Euclidian distance 610. That offset, e.g., that length difference between line 615 and line 610 is then used to compensate the calculated length of the actual diffracted sound path in the connected grid, as depicted by line 620 in Fig.6. In the example of Fig. 6, the sound source positon S and / or the listener position L have been moving further away from the point where the switching between the length calculation algorithm occurred. Thus, as the diffracted sound path has become more complex in Fig.6, compensating the difference between line 615 and the actual Euclidian distance 610 by subtracting this offset from the length of diffracted sound path 620 does no longer result in that the compensated length of the diffracted sound falls back to line 610, but a compensation is conducted with fewer impact. Particular embodiments address the issue that the audible discontinuities happen when a listener moves between unoccluded and occluded areas. These discontinuities can be reduced by compensating the difference of the direct line-of- sight distance and the 8-connected grid distance. The Eulidean distance ^^^௨^^^ௗ^^^between the source grid position S and the listener grid position L is given as follows: S ൌ ^sx, sy^TL ൌ ^lx, ly^TΔx ൌ | lx- sx| Δy ൌ | ly- sy|   The 8-connected grid distance ^^^^^ௗbetween the source grid position S and the listener grid position L can also be computed directly from Δx and Δy: The computation of the 8-connected grid distance is based on the convention that only straight or diagonal steps are allowed to arrive from the source grid position S to the listener grid position L. While a straight step has length of 1 pixel grid, a diagonal step hasa length of√2 pixel grids.The difference can then be used as offset for the computed path distance: where ^^^^^ௗdenotes the grid size. E.g., the grid size corresponds to a length of a pixel grid. If the length is only expressed as a length of pixel grids, then ^^^^^ௗcan be assumed to be1, e.g., ^^^^^ௗ ൌ 1.The above formulae for ^^^^^^^௧ can thus be simplified as: ^^^^^^^௧ ൌ ^^^^^ௗ െ ^^^௨^^^ௗ^^^.In the particular example illustrated by Fig.2, the above formulae result in: S ൌ ^sx, sy^Tൌ ^1, 6^TL ൌ ^lx, ly^Tൌ ^10, 3^TΔx ൌ | lx- sx| ൌ 9 Δy ൌ | ly- sy| ൌ 3 ^^ௗ^^^^^^^ ൌ min^Δx, Δy^ ൌ 3^^^௧^^^^^௧ ൌ max^Δx, Δy^ െ min^Δx, Δy^ ൌ 9 – 3 ൌ 6^^^^^ௗ ൌ ^^^௧^^^^^௧ ^ ^^ௗ^^^^^^^√2 = 6 + 3 √2 = 10.243  ൌ 0.756 ^^^^^ௗ.Assuming that ^^^^^ௗ ൌ 1, ^^^^^^^௧ becomes: ^^^^^^^௧ ൌ 0.756.In an embodiment, compensation is applied for determining a diffraction angle: Without ^^^^^^^௧application: =2 ⋅ ^^^^^^^^^^^^^9.487 / 10.243^ = 2 ⋅ ^^^^^^^^^^^^^0.9262^ = 135.7°The proposed solution reduces the audible discontinuity caused by the overestimation of the path length on the 8-connected grid map. Compared to the computational complexity of the path finding algorithm, the proposed post-processing step adds very little overhead to the overall computational complexity. For a 26-connected grid, the following formulae may, e.g., be applied:       ^^ௗ^^^^^^^_ଷ^ ൌ min^Δx, Δy, Δz^     Δ^x|y|z^ᇱ ൌ Δ^x|y|z^ െ ^^ௗ^^^^^^^_ଷ^     Determine ^^^௧^^^^^௧ and ^^ௗ^^^^^^^as specified for the 2D case for the two largest elements of the set { Δx′, Δy′, Δz′} being used as Δx and Δy.    Fig. 4a shows a distance compensation concept according to an embodiment applied to the previously used example of Fig. 3a to Fig. 3c, where the listener moves behind an occluding object (x-axis: audio frame index; y-axis: distance). The overestimation of the path length is visible by the offset between the Euclidean distance for the direct sound (line 310) and the 8-connected grid distance for the diffracted sound (line 420). At audio frame index 0, the offset is very small as to compared to Fig.3a (which does not apply the inventive processing). Fig. 4b shows the resulting diffraction angle 430 according to an embodiment, which is estimated from the ratio between the direct sound distance and the compensated diffraction propagation path length using equation (1). The discontinuity at audio frame index 0, e.g., the point in time where the listener enters the occluded area, is greatly reduced and the angle is estimated to be approximately 180°. Fig.4c illustrates the distance compensation concept according to the embodiment of Fig. 4a and the resulting diffraction angles 430 of Fig.4b when a listener moves for a plurality of frames. In particular, Fig. 4c illustrates quantized listener positions with compensated Jump Point Search resulting in EQs with quantization steps. According to a particular embodiment, concepts for JPS Path Length Compensation are provided. JPS path comprises only steps parallel to the x / y-axis or diagonal steps. The proposed JPS path length compensation technique may, e.g., be summarized as follows; 1) Compute Euclidian distance between source and listener voxel. 2) Compute “JPS direct sound distance” between source and listener voxel: lengthJPS= stepsstraight+√2 stepsdiagonal. 3) Use difference of these distances as offset for the JPS diffraction path length. The provided concept achieves a smooth transition from an unoccluded area (Eucledian distance) to an occluded area (JPS path length). This is achieved by a simple and lightweight post-processing step for the JPS path-finding algorithm. A path cache may, e.g., be employed, which results in no additional runtime complexity. Fig.5a illustrates quantized listener positions resulting in EQs with quantization steps. According to an embodiment, unquantization of the length of the compensated diffraction path may, e.g., be conducted, such that a length difference of the length of the compensated diffraction path, which is caused by a movement of the sound source position and / or of the listener position, is smoothed. Fig. 5b illustrates unquantized listener positions resulting in EQs without quantization steps according to an embodiment. In the state-of-the-art, a diffraction angle may, e.g., be employed to determine a sound source signal contribution in the audio output signal(s) with, e.g., respect to amplitude and / or with respect to delay. The diffraction angle may, e.g., depend on a relationship of an Euclidian distance between a sound source position S and a listener position L. For example, the diffraction angle ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^may, e.g., be employed according to: wherein ^^^^^^^^^^^^^^denotes the Euclidian distance between S and L, and wherein ^^^^^^^^^^^^^^^^^^^^^^denotes a path length of the diffracted sound path between S and L. In that case, according to embodiments, the compensation of the path length of the diffracted path may, e.g., be conducted within the formula used for calculating the diffraction angle. E.g., the above formula may, e.g., be calculated according to: wherein ^^^^^^^^^^^^ denotes the calculated offset, for example, the length difference between the approximated Euclidian distance as connected grid path in the connected grid (e.g., line 615 in Fig.6) and the proper Euclidian distance (e.g., line 610 in Fig.6). In [1], it is explained that the above diffraction angle may, e.g., be employed in to determine the frequency dependent diffracted source attenuation path effect gain contribution ^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^depending on the diffracted path for determining audio output signals. By applying the ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^modified calculation of the diffraction angle according to the embodiments above, the techniques of [1] can be improved. During the rendering stages, the determined gains may, e.g., be applied on the corresponding audio source signal to obtain the audio output signal(s). See, e.g., the chapters in [1] that relate to distance and equalizer determination.   In particular, in MPEG-I, a 1 / r attenuation (possibly accompanied by a limit) may, e.g., be applied, e,g., together with a delay which depends on the distance and the speed of sound; a frequency dependent filter may, e.g., be applied depending on the diffraction angle, and, in case of earphone reproduction, HRTFs may, e.g., be applied depending on a relative direction. Fig. 7 illustrates a schematic illustration of the data processing on a decoder side. Audio input and a bitstream comprising metadata is received, e.g., by a stream manager. A renderer pipeline may, e.g., process the data, e.g., in a voxel diffraction stage, e.g., in a distance stage, in an equalizer stage, and, e.g., in a spatializer stage. It is, of course, understood by the person skilled in the art that the concepts for adapting the length of the diffraction path and for adapting the diffraction angle are applicable to various other audio processing techniques being different from MPEG-I. Moreover, it is understood by a person skilled in the art that a longer diffraction path from an sound source position to a listener position results in a lower amplitude (and, e.g., a smaller gain value) and in a greater delay compared to a shorter diffraction path. Therefore, the person skilled in the art is aware that a simple way, of obtaining a contribution of an audio source signal asin an audio output signal osis applying a gain g on the audio output signal, e.g., as follows: os= g∙as.Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus. Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software or at least partially in hardware or at least partially in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable. Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed. Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus. The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer. The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0002] References [1] ISO / IEC JTC1 / SC29 / WG6 N0211, “WD5 of ISO_IEC 23090-4, MPEG-I Immersive Audio”, October 2023, Hannover, DE. [2] Dijksta, E. W. (1959). A note on two problems in connexion with graphs. Numerische mathematik, 1(1), 269-271. [3] P. E. Hart, N. J. Nilsson and B. Raphael, "A Formal Basis for the Heuristic Determination of Minimum Cost Paths," in IEEE Transactions on Systems Science and Cybernetics, vol. 4, no. 2, pp. 100-107, July 1968, doi: 10.1109 / TSSC.1968.300136. [4] Harabor, D., & Grastien, A. (2011, August). Online graph pruning for pathfinding on grid maps. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 25, No.1, pp.1114-1119). [5] Daniel, K.; Nash, A.; Koenig, S.; Felner, A. (2010). "Theta*: Any-Angle Path Planning on Grids" (PDF). Journal of Artificial Intelligence Research.39: 533–579. [6] Botea, A., Müller, M., & Schaeffer, J. (2004). Near optimal hierarchical path- finding. J. Game Dev., 1(1), 1-30.

Claims

Claims 1. An apparatus, comprising: an audio metadata determiner (110) configured for determining audio metadata information depending on a first path and depending on a second path, wherein sound waves propagate from a first position to a second position, wherein the first path represents a shortest path from a first position to a second position connecting neighbouring grid pixels of a connected grid, wherein the second path represents a straight line from the first positon to the second position, wherein the first position represents a sound source position where a sound source emits sound waves of the audio source signal and the second position represents a listener position where the sound waves are perceived; or wherein sound waves from the sound source position to the listener position propagate passing the first position and passing the second position.

2. An apparatus according to claim 1, wherein the audio metadata information depends on a difference between the first path and the second path.

3. An apparatus according to claim 1 or 2, wherein the audio metadata determiner (110) is configured to determine audio metadata information indicating a difference between a length of the first path and a length of the second path, the length of the second path representing an Euclidean distance between the first position and the second position.

4. An apparatus according to claim 3, wherein the first position represents the sound source position, wherein the second position represents the listener position, wherein the length of the second path represents the Euclidean distance between the sound source position and the listener position.

5. An apparatus according to claim 4,wherein the audio metadata determiner (110) is configured to determine the length of the first path, wherein the audio metadata determiner (110) is configured to determine the length of the second path, and wherein the audio metadata determiner (110) is configured to determine the difference between the length of the first path and the length of the second path to determine the audio metadata information.

6. An apparatus according to one of the preceding claims, wherein the connected grid is an 8-connected grid, wherein one or more grid pixels of a plurality of grid pixels of the 8-connected grid exhibits exactly 8 neighbouring grid pixels of the plurality of grid pixels of the 8- connected grid, and wherein none of the plurality of grid pixels of the 8-connected grid comprises more than 8 neighbouring grid pixels of the plurality of grid pixels of the 8- connected grid.

7. An apparatus according to claim 6, further depending on claim 5, wherein the audio metadata determiner (110) is configured to determine the length of the second path ^^^^^ௗaccording towherein ^^ௗ^^^^^^^ ൌ min^Δx, Δy^^^^௧^^^^^௧ ൌ max^Δx, Δy^ െ min^Δx, Δy^wherein Δx ൌ | lx- sx|Δy ൌ | ly- sy| wherein S indicates the sound source position, with S ൌ ^sx, sy^T, wherein sxindicates a first coordinate of the sound source position, wherein syindicates a second coordinate of the sound source position, wherein L indicates the listener position, with L ൌ ^lx, ly^T, wherein lxindicates a first coordinate of the listener position, wherein lyindicates a second coordinate of the listener position.

8. An apparatus according to claim 7, wherein the audio metadata determiner (110) is configured to determine the difference between the length of the first path and the length of the second path according toor according towhereinwherein ^^^^^ௗindicates a length of a pixel grid of the 8-connected grid.

9. An apparatus according to one of claims 1 to 5, wherein the connected grid is a 26-connected grid, wherein one or more grid pixels of a plurality of grid pixels of the 26-connected grid exhibits exactly 26 neighbouring grid pixels of the plurality of grid pixels of the 26- connected grid,and wherein none of the plurality of grid pixels of the 26-connected grid comprises more than 26 neighbouring grid pixels of the plurality of grid pixels of the 26- connected grid.

10. An apparatus according to one of claims 3 to 9, wherein the audio metadata determiner (110) is configured to determine a compensated length of a diffraction path from an uncompensated length of the diffraction path depending on the audio metadata information, wherein the diffraction path indicates an estimation of a length of a path along which diffracted sound waves propagate from the sound source position to the listener position along neighbouring grid pixels of the connected grid.

11. An apparatus according to claim 10, wherein the audio metadata determiner (110) is configured to determine the compensated length of the diffraction path from the uncompensated length of the diffraction path by subtracting the difference between a length of the first path and a length of the second path from the uncompensated length of the diffraction path.

12. An apparatus according to claim 10 or 11, wherein the audio metadata determiner (110) is configured to conduct unquantization of the compensated length of the diffraction path, such that a length difference of the compensated length of the diffraction path, which is caused by a movement of the sound source position and / or of the listener position, is smoothed.

13. An apparatus according to one of claims 10 to 12, wherein the audio metadata determiner (110) is configured to determine the diffraction path, which exhibits the uncompensated length, using a shortest search path algorithm for determining the estimation of the path along which diffracted sound waves propagate from the sound source position to the listener position along neighbouring grid pixels of the connected grid.

14. An apparatus according to claim 13,wherein the audio metadata determiner (110) is configured to determine the diffraction path, which exhibits the uncompensated length, using a Dijkstra's algorithm or an A* algorithm, or a jump point search algorithm as the shortest search path algorithm.

15. An apparatus according to one of the preceding claims, wherein the apparatus further comprises a signal generator (120) configured for generating one or more audio output signals depending on an audio source signal and depending on the audio metadata information.

16. An apparatus according to one of claims 10 to 14, further depending on claim 15, wherein the signal generator (120) is configured to generate the one or more audio output signals depending on the audio source signal and depending on the compensated length of the diffraction path.

17. An apparatus according to claim 16, wherein the audio metadata determiner (110) is configured to determine a diffraction angle depending on the compensated length of the diffraction path between the sound source position and the listener position and depending on an Euclidean distance between the sound source position and the listener position, and wherein the signal generator (120) is configured to generate the one or more audio output signals depending on the audio source signal and depending on the compensated length of the diffraction path.

18. An apparatus according to claim 17, wherein the audio metadata determiner (110) is configured to determine the diffraction angle ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^according to: ^^^^^^^^^^^^^^^^^^^^^^^^_^^^^^^^^^^ ൌ 2 ⋅ ^^^^^^^^^^^^^^^^^^^^^^^^^^^ / ^^^^^^^^^^^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^^wherein ^^^^^^^^^^^^^^indicates the Euclidean distance between the sound source position and the listener position,wherein ^^^^^^^^^^^^^^^^^^^^^^indicates the diffraction path between the sound source position and the listener position, which exhibits the uncompensated length, wherein ^^^^^^^^^^^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^ indicates the diffraction path between the soundsource position and the listener position, which exhibits the compensated length, and wherein ^^^^^^^^^^^^ indicates the difference between a length of the first path and a length of the second path, wherein the length of the first path indicates the length of the shortest path from the sound source position to the listener position connecting neighbouring grid pixels of the connected grid, wherein the length of the second path indicates the Euclidean distance ^^^^^^^^^^^^^^between the sound source position and the listener position.

19. An apparatus according to claim 17 or 18, wherein the audio metadata determiner (110) is configured to determine one or more gains depending on the diffraction angle, wherein the signal generator (120) is configured to generate the one or more audio output signals depending on the one or more gains and depending on the audio source signal.

20. An apparatus according to one of claims 17 to 19,   wherein the audio metadata determiner (110) is configured to determine one or more equalizer filters and / or a delay depending on the diffraction angle, wherein the signal generator (120) is configured to generate the one or more audio output signals depending on the audio source signal and depending on the one or more equalizer filters and / or the delay.

21. An apparatus according to one of claims 10 to 19, wherein the signal generator (120) is configured to generate the one or more audio output signals depending on the audio metadata information only if a difference between the uncompensated length of the diffraction path and the Euclidian distance between the sound source position and the listener position is smaller than a threshold value, or wherein the signal generator (120) is configured to generate the one or more audio output signals depending on the audio metadata information only if a quotient between the uncompensated length of the diffraction path and the Euclidian distance between the sound source position and the listener position is smaller than a threshold value.

22. An apparatus according to one of the preceding claims, further depending on claim 15, wherein the apparatus further comprises an audio decoder for decoding encoded audio signal to determine the audio source signal, wherein the audio metadata determiner (110) is configured to determine the audio metadata information, wherein the signal generator (120) is configured to generate the one or more audio output signals using the audio source signal, which has been decoded from the encoded audio signal, and using the audio metadata information.

23. An apparatus according to one of claims 1 to 14, wherein the apparatus comprises an audio encoder for encoding an audio source signal as an encoded audio signal, wherein the audio metadata determiner (110) is configured to determine the audio metadata information for a plurality of potential listener positions and for a plurality of potential sound source positions, and wherein the apparatus is configured to transmit the encoded audio signal and the audio metadata information to an audio decoder.

24. A system, comprising: an apparatus according to claim 23, and an apparatus for generating one or more audio output signals comprising an audio decoder and a signal generator (120), wherein the apparatus according to claim 23 is configured to transmit the encoded audio signal and the audio metadata information to the apparatus for generating the one or more audio output signals, wherein the audio decoder is configured to decode the encoded audio signal to obtain the audio source signal, and wherein the signal generator (120) is configured to obtain selected metadata information by selecting information from the audio metadata information for a current listener position and for a current sound source position, and wherein the signal generator (120) is configured to generate one or more audio output signals using the audio source signal and using the selected metadata information.

25. A method, comprising: determining audio metadata information depending on a first path and depending on a second path, wherein sound waves propagate from a first position to a second position, wherein the first path represents a shortest path from a first position to a second position connecting neighbouring grid pixels of a connected grid, wherein the second path represents a straight line from the first positon to the second position, wherein the first position represents a sound source position where a sound source emits sound waves of the audio source signal and the second position represents a listener position where the sound waves are perceived; or wherein sound waves from the sound source position to the listener position propagate passing the first position and passing the second position.

26. A computer program for implementing the method of claim 25 when being executed on a computer or signal processor.

Citation Information

Patent Citations

  • Methods, apparatus and systems for diffraction modelling based on grid pathfinding

    US20230188920A1

  • Methods, apparatus, and systems for processing audio scenes for audio rendering

    WO2023169934A1