Method, apparatus, and system for initial reflection estimation for voxel-based geometric representation

The method addresses the inefficiency of existing reflection estimation techniques by using a voxel-based approach to calculate initial sound source reflections in 3D environments, achieving accurate and low-complexity sound modeling without needing surface orientation information.

JP2025517640APending Publication Date: 2025-06-10DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024565073
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-26
Filing Date
2023-05-22
Publication Date
2025-06-10

Smart Images

  • Figure 2025517640000001_ABST
    Figure 2025517640000001_ABST
Patent Text Reader

Abstract

A method, apparatus, program, and storage medium for improving the estimation of the initial reflection trajectories of audio sources in a three-dimensional audio scene are described. The method includes obtaining a voxel-based representation of the audio scene, information regarding the listener location in the audio scene, and information regarding the audio source location in the audio scene. For one or more points on the connection line between the audio source location and the listener location, for each of these points, a ray direction pattern is applied to obtain a plurality of sound rays starting from each point. Based on the sound rays and the voxel-based representation of the audio scene, a set of collision voxels is determined. The initial reflection trajectory is determined based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the benefit of priority based on U.S. Provisional Patent Application No. 63 / 344,895, filed on May 23, 2022, and U.S. Provisional Patent Application No. 63 / 387,339, filed on December 14, 2022, all of which are incorporated herein by reference.

[0002] [Technical Field] The present disclosure relates to the modeling of audio sources, and more particularly to a voxel - based method and apparatus for estimating early sound source reflections.

Background Art

[0003] The sound reflection of an acoustic reflecting surface can affect the perceived sound of an audio source. The sound that is reflected and received at a target location (e.g., listener position) immediately after the direct sound is referred to herein as an early reflection (ER), and is of particular interest when modeling a sound source because the perceived sound of an audio source can be accurately modeled by considering only the direct sound and the ER. On the other hand, higher - order acoustic reflections are often less important because they have low energy and are temporally / spatially psychoacoustically masked by the ER and other components.

[0004] The ER causes several perceptual effects such as the apparent source width, perceived distance, timbre, and spaciousness. The ER is relatively sparse in time and typically occurs over a relatively short time period within the first approximately 80 ms of the room impulse response (see FIG. 1). FIG. 1 shows the echoogram of a room, including the echoograms of the direct sound source, early reflections, and late reflections. FIG. 1 also enables visualization of the differences between the direct sound, early reflections, and late reflections.

[0005] The psychoacoustic relevance of the ER depends strongly on several factors such as the direction, level, time delay, and spectral content of the audio signal.

[0006] The direction of the ER affects, in particular, the time delay and the frequency response at the listener's ear. Thus, the direction of the ER plays an important role in the perceived reflected sound. When the direction of arrival changes, this means that there have been changes in the path from the source to the listener's ear due to movement, obstacles, etc. Changes in the path length affect the time delay, and due to the shape of the pinna, different frequency responses are generated depending on the direction of arrival at the ear.

[0007] To estimate the trajectory of the ER, the image source (IS) method aims to find the pure specular reflection path between the audio source and the receiver, i.e., the listener. This process is simplified by assuming that sound propagates only along a straight line, i.e., a ray. The audio image source is spawned at the same distance from the boundary as the original source 101 (see Fig. 2) on a line perpendicular to the boundary. Fig. 2 shows the sound source 101, the listener 102, the boundary, and the image source.

[0008] When the sound is reflected at the boundary surface at the same angle as the angle of incidence, the impression is created that the original source 101 is reflected like a mirror at the boundary surface. The reflection by a single boundary represents the (primary) ER.

[0009] However, the boundaries may be unknown or lack precision. One example is the voxel-based representation of a 3D environment used for sound rendering in VR applications. A voxel is a volumetric space having several acoustic attributes, such as reflectivity. When the orientation of the reflecting surface is not explicitly assigned to its characteristics, a single voxel has no orientation information, so a set of voxels should be considered to find the boundary of the IS method. Therefore, complex trigonometric considerations are required to estimate the boundary. FIG. 3 shows an exemplary scenario. In this figure, the gray voxels represent reflective objects, and the gray voxels adjacent to the white voxels represent the reflective boundaries of the object's surface. Without the orientation information of the reflection, a single voxel is insufficient to determine the reflection trajectory of the sound emitted by source 101.

[0010] Therefore, there is a need for an improved and efficient approach for ER estimation in voxel-based environments, especially when the orientation information of the boundaries reflecting audio is not available in advance. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0011] In view of the above, the present disclosure provides a method, an apparatus, and a program for initial sound source reflection estimation in a voxel-based 3D environment (3D voxel grid), each having the features of the independent claims, as well as a computer-readable storage medium.

[0012] According to one aspect of the present disclosure, a method for estimating an initial reflection is provided. Information regarding a voxel-based representation of a 3D audio scene, the listener location of a listener in the 3D audio scene, and the audio source location of an audio source in the 3D audio scene can be obtained (e.g., received or determined). A sound ray direction pattern is applied to one or more points on a connecting line between the audio source location and the listener location, and for each of the one or more points, a plurality of sound rays starting from the respective point can be obtained. A set of collision voxels can be determined based on the plurality of sound rays and the voxel-based representation of the 3D audio scene. An initial reflection trajectory can be determined based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test. For example, for each collision voxel in the set of collision voxels, a path connecting the listener location and the audio source location through the respective collision voxel can be determined. Then, for each path, if the path is geometrically valid, the path can be determined as the initial reflection trajectory.

[0013] By using the above heuristic method, the initial reflection can be efficiently estimated in a voxel-based environment without requiring information on the orientation of the reflection surface of the voxels. Thereby, the sound source can be modeled with high accuracy and low computational complexity, enabling accurate and efficient sound representation in real-time applications, such as VR games.

[0014] In some embodiments, the method may further include determining a sound ray direction pattern. Determining the sound ray direction pattern may include selecting the sound ray direction pattern from some (set of) predetermined sound ray direction patterns or calculating the sound ray direction pattern. Alternatively, the sound ray direction pattern may be fixed. Further alternatively, an indication of the sound ray direction pattern to be used may be received together with a bitstream.

[0015] In some embodiments, the method may further include determining one or more points based on the number of one or more points (e.g., count, cardinality). That is, the number of one or more points can be obtained or determined (e.g., set to be N points), and the obtained (e.g., N) number (count or cardinality) of one or more points can correspond to the coordinates of the one or more points (in the sense that each of the one or more points has its respective coordinates).

[0016] In some embodiments, the ray direction pattern can be defined as (e.g., can include) a predetermined number of rays and the directions of the predetermined rays from a starting point. The predetermined number of rays may be, for example, 6, 8, or 12. The direction of a ray can be defined by the grid indices of the voxel grid.

[0017] In some embodiments, the direction of a predetermined ray can include one or more of a horizontal direction and a vertical direction to neighboring grid indices and a diagonal direction to neighboring grid indices. Thus, the predetermined direction can define the relative direction from the starting point of the ray, i.e., the grid index (l, m, i) in the voxel grid. The relative direction can be represented as follows.

[0018] (+1, 0, 0), (-1, 0, 0), (0, +1, 0), (0, -1, 0), (0, 0, +1), (-0, 0, -1); (+1, +1, 0), (+1, -1, 0), (-1, +1, 0), (-1, -1, 0), (+1, 0, +1), (+1, 0, -1), (-1, 0, +1), (-1, 0, -1), (0, +1, +1), (0, +1, -1), (0, -1, +1), (0, -1, -1); and (+1, +1, +1), (+1, +1, -1), (+1, -1, +1), (+1, -1, -1), (-1, +1, +1), (-1, +1, -1), (-1, -1, +1), (-1, -1, -1).

[0019] In some embodiments, determining the sound ray direction pattern may be based on the scene type of the 3D audio scene, available computing resources, an encoder preset, or a combination thereof.

[0020] In some embodiments, the coordinates of one or more points on the line connecting the audio source location and the listener location may be determined based on the number of one or more points (e.g., count, density).

[0021] In some embodiments, one or more points may be determined to divide the line connecting the audio source location and the listener location into N - 1 equal line segments, where N is the number of one or more points (e.g., count, density). N may be, for example, 2 or more.

[0022] In some embodiments, the number of one or more points may depend on the scene type of the 3D audio scene, available computing resources, an encoder preset, or a combination thereof.

[0023] In some embodiments, the scene type may include indoor scenes and outdoor scenes.

[0024] In some embodiments, each collision voxel may be an occluder voxel in the voxel - based representation of the 3D audio scene.

[0025] In some embodiments, the occluder voxel may represent an acoustic reflecting surface.

[0026] In some embodiments, the occluder voxel may represent any substance other than air in the voxel - based representation of the 3D audio scene. That is, the occluder voxel may represent a reflecting surface, and the non - occluder voxel may represent a non - reflecting surface (or may not define a surface at all).

[0027] In some embodiments, determining a set of collision voxels based on a plurality of sound rays and a voxel-based representation of a 3D audio scene may include determining one or more intersections (e.g., intersection points) between each sound ray of the plurality of sound rays and a blocker voxel. The method may further include, for each sound ray, determining a blocker voxel that includes the intersection closest to the starting point of the respective sound ray as a collision voxel in the set of collision voxels. That is, the collision voxels may be the blocker voxels that each sound ray first hits.

[0028] In some embodiments, determining an initial reflection trajectory based on a set of collision voxels, a listener location, an audio source location, and a geometric validity test may include, for each collision voxel in the set of collision voxels, determining whether the collision voxel can generate a geometrically valid representation of a primary reflection. If it is determined that a collision voxel can generate a geometrically valid representation of a primary reflection, a path connecting the listener location and the audio source location through each collision voxel may be determined as the initial reflection trajectory.

[0029] In some embodiments, determining whether a collision voxel can generate a geometrically valid representation of a primary reflection may include determining the previous voxels of the collision voxel. The previous voxels may include intersections with respective sound rays and may be voxels that precede the collision voxel in the direction of the respective sound rays. A second path connecting the listener location and the audio source location via each previous voxel may be determined. If the second path does not include an intersection with an occlusion voxel, the collision voxel can generate a geometrically valid representation of the primary reflection. Generally, if neither the path connecting the listener location and the previous voxel nor the path connecting the audio source location and the previous voxel includes an intersection with an occlusion voxel, the collision voxel can generate a geometrically valid representation of the primary reflection. In other words, if both the path connecting the listener location and the previous voxel and the path connecting the audio source location and the previous voxel pass a line-of-sight check (a "visibility check"), the collision voxel can generate a geometrically valid representation of the primary reflection.

[0030] This enables efficient selection of collision voxels that cannot reach a geometrically valid path from the audio source location to the listener position.

[0031] Alternatively or additionally, determining an initial reflection trajectory based on a set of collision voxels, a listener location, an audio source location, and a geometric validity test may include, for each collision voxel in the set of collision voxels, determining a path connecting the listener location and the audio source location via the respective collision voxel. For each path, if the path is geometrically valid, the path may be determined as an initial reflection trajectory. A path is geometrically valid if it passes a line-of-sight check (a "visibility check"), i.e., if both the path connecting the listener location and the collision voxel and the path connecting the collision voxel and the audio source location pass the line-of-sight check.

[0032] In some embodiments, the path may include a straight line connecting the audio source location to a collision voxel within the set of collision voxels, and a straight line connecting the same collision voxel within the set of collision voxels to the listener location.

[0033] In some embodiments, if the path does not include intersections with occluder voxels other than the collision voxels of each path, the path can be determined to be geometrically valid. That is, a path having intersections with two or more occluder voxels can be discarded. In other words, a path can be determined to be geometrically valid if it is not blocked by any occluder voxel other than the collision voxels.

[0034] When both tests are performed for collision voxels that can generate a geometrically valid representation of the primary reflection and for geometrically valid paths, collision voxels that cannot generate a geometrically valid representation of the primary reflection can first be screened by determining whether there is an intersection between the occluder voxels and the paths connecting the audio source location, the previous voxels, and the listener location. For the remaining collision voxels, paths connecting the audio source position, the collision voxels, and the listener position can be determined. Finally, it can be determined whether there is an intersection between these paths and the occluder voxels other than the collision voxels.

[0035] By combining two geometric validity tests, only geometrically valid initial reflection trajectories can be determined regardless of the geometric shape of the 3D audio scene.

[0036] In some embodiments, the method may further include selecting a set of the most acoustically relevant initial reflection trajectories from the initial reflection trajectories.

[0037] In some embodiments, selecting the set of initially reflected trajectories that are most acoustically relevant can be based on the lengths of the initially reflected trajectories and / or the reflection coefficients of the collision voxels of each of the initially reflected trajectories. In particular, the acoustically relevant initially reflected trajectories can have, for example, a short length and / or a large reflection coefficient compared to initially reflected trajectories that are not acoustically relevant.

[0038] In some embodiments, the reflection coefficient can depend on the material modeled (or otherwise indicated) by the collision voxel.

[0039] In some embodiments, selecting the set of initially reflected trajectories that are most acoustically relevant can include discarding initially reflected trajectories having a value that exhibits an interior angle close to 180° at the collision voxel. Here, close to 180° means 180° - ε, where ε is a small angle. In some implementations, for example, initially reflected trajectories having said value with an interior angle greater than 160° can be discarded.

[0040] In some embodiments, the value that exhibits an interior angle close to 180° can be the interior angle or length of the initially reflected trajectory.

[0041] In some embodiments, the method can further include outputting the initially reflected trajectories. That is, the initially reflected trajectories or the initially reflected trajectories that are most acoustically relevant can be output for rendering or for further processing such as, for example, occlusion, diffraction, 3D extent, or reverberation processing prior to rendering.

[0042] In some embodiments, the method can further include rendering a three-dimensional audio scene, for example, by a virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) device.

[0043] In some embodiments, the initial reflection trajectory may represent a primary trajectory. In some embodiments, the primary trajectory may be a reflection trajectory having a single reflection between the audio source location and the listener location.

[0044] According to another aspect of the present disclosure, a method for processing a frame (e.g., a time frame) of a 3D audio scene is provided. The reflection trajectory of the frame can be estimated based on the method according to the previous aspect. The estimated initial reflection trajectory can be stored. (e.g., stored locally or submitted to shared storage or cloud storage). Alternatively, the estimated initial reflection trajectory of the previous frame can be accessed (e.g., from local storage, shared storage, or cloud storage). The estimated initial reflection trajectory of the previous frame can be calculated based on the method according to the previous aspect. The estimated initial reflection trajectory of the previous frame can be accessed only if the voxels including the listener location, the voxels including the audio source location, and the geometry of the voxel-based representation of the 3D audio scene do not change between the frame and the previous frame.

[0045] When the 3D audio scene is static, by using the previous estimate of the initial reflection trajectory, the complexity of processing the audio data for the 3D audio scene can be reduced without affecting the accuracy of the output.

[0046] According to another aspect of the present disclosure, there is provided a method of audio processing for creating a locus for geometrically connected audio sources for efficient implementation on a voxel 3D grid. Information related to a sound ray direction pattern “R” may be received. A first set “P” of points for applying ray casting based on the sound ray direction pattern “R” may be determined. A second set “C” of sound ray-voxel “collision” voxels based on the first set of points and reflection voxels “VOX” may be determined. A third set “S-C-L” of valid reflection loci based on the second set “C” of sound ray-voxel “collision” voxels may be determined. A subset of the most acoustically relevant ones may be selected from the third set of valid reflection loci and output.

[0047] Aspects of the present disclosure may be implemented via an apparatus. The apparatus may include a processor and a memory coupled to the processor. The processor may be adapted to perform the methods according to aspects and embodiments of the present disclosure.

[0048] Aspects of the present disclosure may also be implemented via a program. When the instructions of the program are executed by a processor, the processor may implement aspects and embodiments of the present disclosure. A computer-readable storage medium may store the program. Such computer-readable storage media may include, but are not limited to, memory devices such as those described herein, including random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Thus, some inventive aspects of the subject matter described in the present disclosure may be implemented via one or more computer-readable storage media storing software.

[0049] It will be understood that the features of the apparatus and the steps of the method may be interchanged in many ways. In particular, the details of the disclosed method, as will be understood by those skilled in the art, can be realized by the corresponding apparatus (or system), and vice versa. Further, it will be understood that any of the above descriptions made with respect to the method(s) is equally applicable to the corresponding apparatus (or system), and vice versa.

[0050] Exemplary embodiments of the present invention will be described below with reference to the accompanying drawings.

Brief Description of the Drawings

[0051]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

[0052] The drawings and the following description are directed to preferred embodiments for purposes of illustration only. It should be noted that from the following description, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of the claimed invention.

[0053] Here, some embodiments will be referred to in detail, and examples thereof are shown in the accompanying drawings. It should be noted that whenever possible, similar or like reference numerals may be used in the figures and may indicate similar or like functions. The drawings show embodiments of the disclosed system (or method) for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be used without departing from the principles described herein.

[0054] Moving Picture Experts Group (MPEG) is an alliance of working groups jointly established by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC), which develop standards for media encoding, including audio encoding. MPEG is organized under ISO / IEC SC 29, and the Audio Group is currently identified as Working Group (WG) 6.

[0055] The new MPEG-I standard enables acoustic experiences from different viewpoints and / or perspectives, or listening positions, by supporting various movements such as movements using various degrees of freedom in virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) applications, such as scenes and various movements around such scenes.

[0056] For audio rendering in VR, AR, MR, and XR applications, an object-based approach is widely used by representing complex auditory scenes as a plurality of distinct audio objects, where each audio object is associated with parameters or metadata that define the location / position and trajectory of that object in the scene. Alternatively, audio rendering in such environments may use higher-order ambisonics (HOA).

[0057] Voxels for audio rendering are relevant to media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments. Some concepts related to and defining voxels for audio rendering are described below.

[0058] [What is a voxel in audio rendering?] A voxel is a spatial volume to which acoustic characteristics or audio rendering instructions are assigned.

[0059] [What is the voxel size in audio rendering?] The voxel size is an encoder configuration parameter and can be selected (manually or automatically) according to the level of detail of the scene geometry (e.g., in the range of 10 cm to 1 m).

[0060] [How can voxels be obtained?] Voxels for audio rendering can be obtained as follows, namely · Voxelization (or conversion) of mesh-based scene representations, and / or · From the scene representation used for scene generation (or video rendering) (e.g., by downsampling to smaller-sized voxels) can be obtained.

[0061] [How to represent a voxel-based audio scene?] Any voxel - based representation of an audio scene can include an indication of voxels that are not transparent voxels (e.g., occluder voxels), i.e., voxels through which sound cannot propagate or cannot propagate freely, i.e., an indication of the occlusion geometry. This indication can be related to an indication of the coordinates of each voxel (e.g., center coordinates, corner coordinates). These voxel coordinates can be represented, for example, by a grid index. Further, the voxel - based representation can include an indication of the material properties of voxels that are not transparent voxels, such as absorption coefficients, reflection coefficients, etc. In addition to occluder voxels, the voxel - based representation can also show transparent voxels (e.g., air voxels), i.e., voxels through which sound can propagate in the representation of the sound - propagation medium. Thus, some implementations of the voxel - based representation of an audio scene can include an indication of the respective material properties for each voxel within a predefined section of space (e.g., within the boundaries enclosing the audio scene).

[0062] However, existing conventional approaches for using voxels to provide realistic sound for user experiences (including those with movement) in VR, AR, MR, and XR environments are difficult and computationally complex. Estimating the ER trajectory is of particular interest for these use cases. For VR use cases, the psychoacoustic relevance of the ER direction is higher than the ER audio signal level because VR users can see reflective surfaces better than they can estimate their reflection properties (because, for example, reflection properties give an estimate of the reflected energy or audio signal level).

[0063] In conventional approaches for estimating the ER trajectory (e.g., the IS approach), the boundaries for determining reflections must be known (and must be clearly defined). Since a single voxel does not provide sufficient information about the orientation of the reflecting surface boundary, the method of the present invention relies on a heuristic approach to estimate the ER trajectory for generating a primary ER acoustic effect that is perceived with sufficient accuracy or faithfully enough.

[0064] The heuristic approach according to the present disclosure is based on finding a geometrically reasonable reflection trajectory between the audio source and the listener that is sufficient to generate the perceived primary ER acoustic effect by performing several low-complexity steps based on the location of the audio source and the listener and the voxel-based geometric representation.

[0065] Figure 4 shows the general idea of the heuristic approach. The positions of the audio source 101 and the listener 102 are known. Then, a reasonable reflection trajectory from the audio source 101 to the listener 102 on the collision voxel 104 should be estimated without considering the information of the reflecting surface based on a plurality of voxels, that is, the heuristic approach functions for each voxel and with respect to the grid index representing the voxel position.

[0066] The heuristic approach will be described in detail for the example of a specific audio scene shown in FIGS. 7 to 11. However, the present disclosure should not be construed as being limited by this specific example. Further, although this example relates to the 2D case or shows only a 2D projection, it can be seen that the method according to the present disclosure is generally applicable to 3D audio scenes. FIG. 7 shows an exemplary 2D audio scene having occluding voxels (105, dot pattern), non-occluding voxels (plain), an audio source 101, and a listener 102. The locations of the audio source 101 and the listener 102 are marked with S and L, respectively. This example is presented as 2D for illustrative purposes only. The extension of the algorithm to a 3D environment is straightforward.

[0067] The voxel-based representation of the audio scene, information regarding the listener location, and the location of the audio source 101 are inputs to the ER trajectory estimation method. In other words, the voxel-based representation of the audio scene, information regarding the listener location, and the audio source location can be received. Alternatively, the voxel-based representation of the audio scene, information regarding the listener location, and the audio source location can be determined in the ER trajectory estimation method.

[0068] To find the reflection trajectory between the audio source 101 and the listener 102, the ray direction pattern is applied to a point 103 within the audio scene. In the example of FIG. 7, five equally spaced points are shown on the line connecting the audio source 101 and the listener 102. However, the illustrated example should not be construed as limiting the arrangement and number of points. Different numbers of points and different positions of these points can be used in this method. The ray direction pattern may be predetermined. Further, the ray direction pattern may define a predetermined number of rays and the direction of the predetermined (corresponding) rays from the starting point. Exemplary ray direction patterns are shown in FIGS. 5A - 5D. Thus, the predetermined number of rays may be, for example, 6 (FIG. 5A), 8 (FIG. 5C), or 12 (FIG. 5B), or a combination thereof (FIG. 5D), and the predetermined direction of the rays may include the direction from the ray starting point (l, m, i) in the voxel grid. The starting point of the ray in the audio scene may be the point 103. The direction can be any combination of the horizontal and vertical directions to the neighboring grid indices and the diagonal directions to the neighboring grid indices with respect to the ray starting point. The relative directions can be represented as (+1,0,0), (-1,0,0), (0,+1,0), (0,-1,0), (0,0,+1), (-0,0,-1), (+1,+1,0), (+1,-1,0), (-1,+1,0), (+1,0,+1), (+1,0,-1), (-1,0,+1),(-1,0,-1),(0,+1,+1),(0,+1,-1),(0,-1,+1),(0,-1,-1); and (+1,+1,+1),(+1,+1,-1)(+1,-1,+1),(+1,-1,-1),(-1,+1,+1),(-1,+1,-1),(-1,-1,+1),(-1,-1,-1).

[0069] Determining the ray direction pattern can be understood as selecting one of the predetermined ray patterns. Determining the ray direction pattern can be based on the scene type of the audio scene, available computing resources, encoder preset, or a combination thereof. The scene type can include, for example, an indoor scene and an outdoor scene.

[0070] In the next step, it is necessary to define (e.g., determine or calculate) the coordinates of the points in order to apply the sound ray direction pattern to each point 103. The number of points (e.g., count, density) can be determined. The number of one or more points may depend on the scene type of the audio scene, the available computing resources, the encoder preset, or a combination thereof. The scene type may include, for example, an indoor scene and an outdoor scene. Alternatively, the number of points may be fixed. In some implementations, the number of points may alternatively or additionally depend on the selected sound ray direction pattern.

[0071] Furthermore, it has been found that placing points on the line between the audio source 101 and the listener 102 improves the quality and efficiency of ER trajectory estimation. For this purpose, the points 103 (the location of the points 103) can be determined based on the number of one or more points. Furthermore, the location of the points may be determined based on the line connecting the audio source location and the listener location (e.g., arranged on the said line). More specifically, one or more points 103 can be determined such that, for example, the line connecting the audio source location and the listener location is divided into N - 1 equal line segments. Here, N is the number of one or more points, and in this case, it may be 2 or more.

[0072] In particular, when the number of points is selected as 1, one point may correspond to the audio source location. When the number of points is selected as 2, two points may correspond to the audio source location and the listener location, respectively.

[0073] FIG. 8 shows an example in which a sound ray direction pattern having eight sound rays is applied to the point 103 located at the location of the audio source 101. The sound rays are shown as dashed lines.

[0074] In the next step, a set of collision voxels is determined based on a plurality of sound rays and a voxel-based representation of the audio scene. In particular, the set of collision voxels can be determined by looking for intersections between the sound rays in the audio scene and any occluder voxels 105 (dot pattern). The occluder voxels 105 can represent an acoustic reflecting surface. In other words, the occluder voxels 105 can represent any substance other than air or other representations of the sound propagation medium in the voxel-based representation of the audio scene. Then, for each sound ray, the occluder voxel 105 that contains the intersection closest to the starting point of the respective sound ray can be determined as the collision voxel 104. In other words, the collision voxel 104 can be defined as the first occluding voxel that each sound ray hits. This step ensures that only occluding voxels that can represent a reflecting surface are selected. In FIG. 8, all the collision voxels 104 of the sound rays starting from the audio source location are marked with black circles at the ends of the sound rays. In particular, in this example, the lower right sound ray has no intersection with any occluding voxels and thus is not shown in FIG. 8 and is not further considered in the algorithm.

[0075] In the next step, it can be determined whether the collision voxel 104 can generate a geometrically valid representation of the primary reflection. FIG. 6A shows a scenario where the collision voxel can generate a geometrically valid representation of the primary reflection. For the collision voxel 104, a preceding voxel 107 may be determined. The preceding voxel 107 can be a voxel that precedes the collision voxel 104 in the direction of each sound ray. For the preceding voxel 107, a path connecting the audio source location and the listener location via the preceding voxel 107 can be determined. If the path does not include an intersection with a shielding voxel (i.e., passes the line-of-sight or visibility check), the collision voxel 104 is determined as a collision voxel that can generate a geometrically valid representation of the primary reflection. In the case of FIG. 6A, the path does not include an intersection with a shielding voxel. FIG. 6B shows a scenario of the collision voxel 104 that cannot generate a geometrically valid representation of the primary reflection. In this example, the line connecting the preceding voxel and the listener position intersects with a shielding voxel on the right side of the collision voxel 104. The collision voxel 104 that cannot generate a geometrically valid representation of the primary reflection may be discarded.

[0076] In the next step, for each collision voxel 104 determined for the sound ray and the starting point 103, a path for connecting the audio source 101 and the listener 102 via each collision voxel 104 can be determined. The path can include a straight line connecting the audio source location to the collision voxel 104 and a straight line connecting the same collision voxel 104 to the listener location. Alternatively, the path can include a straight line connecting the audio source location to the preceding voxel 107 and a straight line connecting the preceding voxel 107 to the listener location, or can include a path derived from the two possible paths described above.

[0077] FIG. 9 is a diagram showing the example shown in FIG. 8 together with the determined path between the audio source 101 and the listener 102.

[0078] In an optionally final step, it is determined whether the path from the audio source 101 to the listener 102 is geometrically valid. This step can be combined with the previous selection of collision voxels 104 that can generate a geometrically valid representation of the primary reflection. Then, only the paths associated with the collision voxels that can generate a geometrically valid representation of the primary reflection can be considered in subsequent geometric validity tests. Alternatively, subsequent geometric validity tests can consider the paths associated with all collision voxels 104. Since the paths determined in the previous step can be defined only by the straight lines between the audio source 101, the collision voxels 104, and the listener 102, the lines can cross (e.g., intersect or lightly touch) the shielding voxels other than the collision voxels 104. In practice, such reflection trajectories are impossible in the sense that they do not allow the propagation of sound. Therefore, a path that includes a line crossing a shielding voxel other than each collision voxel 104 can be determined to be geometrically invalid. To find the intersections with the shielding voxels other than each collision voxel 104, a line-grid intersection algorithm can be applied to the lines connecting the audio source 101, the collision voxels 104, and the listener 102. As an example, a fast traversal algorithm for ray tracing (Amanatides, J. and A. Woo, A Fast Voxel Traversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987.87.) may be used.

[0079] FIG. 10 is a diagram showing the determined geometrically valid paths as solid lines. In particular, in this example, only one of the seven previously found paths is determined to be geometrically valid.

[0080] As described above, this process is repeated for each point 103 on the line between the audio source 101 and the listener 102.

[0081] Figure 11 shows the final result of the algorithm for an example of N = 7 points 103 and 8 sound lines per point 103. Of the 7 * 8 = 56 possible paths, only 11 are geometrically valid. And these paths can be considered as ER trajectories 106.

[0082] Figures 12(a) to 12(d) show the results of the algorithm for previous audio scenes with different listener locations and N = 7.

[0083] Optionally, the resulting ER trajectory 106 can be output for further processing such as, for example, rendering the audio scene. Alternatively, the determined ER trajectory 106 may be further analyzed for improved audio scene rendering.

[0084] For example, a set of (one or more) acoustically most relevant ER trajectories from the ER trajectory 106 may be selected. The selection can be based on the length of the ER trajectory 106 and / or the reflection coefficient of the collision voxels 104 of the ER trajectory 106. The reflection coefficient can depend on the material modeled by each collision voxel 104. For example, an ER trajectory having a very long path length (e.g., longer than a certain threshold or longer than a certain fraction or multiple of the length of the connection line between the audio source and the listener) can be discarded. Alternatively or additionally, an ER trajectory having a small reflection coefficient (e.g., smaller than a certain threshold) can be discarded.

[0085] Alternatively or additionally, an ER trajectory 106 having a value indicating a large interior angle in the collision voxel 104 may be discarded. The large angle may be defined as an interior angle close to 180°, for example 180° - ε (where ε is a small angle). Alternatively, a large angle in this context may be an interior angle greater than 160°. The value indicating the interior angle may be the interior angle itself or the length of the ER trajectory 106 (note that a large interior angle means a relatively short path length and a small interior angle means a relatively long path length). FIG. 13 shows an example of a geometrically valid ER trajectory 106 having an interior angle close to 180°. As a result, the path length is very close to the direct path between the audio source 101 and the listener 102. Therefore, the ER is masked by the direct sound received from the audio source 101 (i.e., sound without reflection). Therefore, the ER may be determined to be psychoacoustically invalid / irrelevant. As a result, such an ER may be discarded.

[0086] In other examples, two or more determined ER trajectories may be averaged (e.g., spatially averaged). To do so, respective mirror sources may be determined for two or more ER trajectories, and the determined mirror sources may be spatially averaged to obtain an averaged mirror source. Further, the associated gain of the mirror source may be determined (e.g., based on the reflection coefficient and / or path length), and the gain of the averaged mirror source may be obtained by averaging the individual gains.

[0087] By the disclosed method above, the ER direction can be obtained in an efficient manner while still allowing an accurate (e.g., faithful or at least realistic) representation of the ER effect during sound rendering.

[0088] The disclosed method above may be used for each frame for audio processing of a three-dimensional audio scene. Alternatively, a time subdivision (time unit) other than a frame may be used here. The proposed method is independent of the type of time unit considered.

[0089] For a given frame, the reflection trajectory can be estimated based on a method according to the above-disclosed method. The estimated initial reflection trajectory and the coordinates of the listener location and the audio source location can be stored. The estimated initial reflection trajectory may include the respective gains of the audio mirror sources. Generally, the estimated initial reflection trajectory may be related to an indication of the mirror source location and / or an indication of the mirror source gain. Alternatively, the estimated initial reflection trajectory of the previous frame can be accessed. In particular, the stored coordinates and gains of the audio mirror sources of the previous frame can be accessed. The estimated initial reflection trajectory of the previous frame can also be estimated based on the above-disclosed method. The estimated initial reflection trajectory of the previous frame can be accessed only if the voxels including the listener location, the voxels including the audio source location, and the geometry of the voxel-based representation of the three-dimensional audio scene do not change between the frame and the previous frame. The voxels including the listener location may be represented by a listener head position voxel index. The voxels including the audio source location may be represented by an audio point source position voxel index. The geometry of the voxel-based representation of the three-dimensional audio scene can be represented by a 3D voxel matrix (for example, associated with a reflection coefficient). Along the above lines, as shown in the flowchart of FIG. 14, a method 200 for the estimation of the ER trajectory is provided. This method can be implemented in a decoder or a renderer, or in both a decoder and a renderer, in an AR / VR / MR / XR environment. The decoder and / or renderer may be implemented in a network / cloud, or in a processing device such as a mobile device and an AR / VR / MR / XR Google / Lens, or may be distributed across both the network / cloud and the processing device. In addition to the following method steps, method 200 can optionally include all the variations described above with respect to the aforementioned ER estimation algorithm described in relation to FIGS. 7-13.

[0090] In step S201, a voxel-based representation of the three-dimensional audio scene, information regarding the listener location of the listener 102 in the three-dimensional audio scene, and information regarding the audio source location of the audio source 101 in the three-dimensional audio scene are obtained. The voxel-based representation, the information regarding the listener location, and the information regarding the audio source location can each be received and / or pre-determined (i.e., calculated previously, stored, and then read from memory).

[0091] In step S202, a sound ray direction pattern is applied to one or more points 103 on the connection line between the audio source location and the listener location, and for each of the one or more points 103, a plurality of sound rays starting from the respective point 103 are obtained. The sound ray direction pattern can be received and / or pre-determined. Alternatively, the sound ray direction pattern may be determined in the context of the proposed method. Determining the sound ray direction pattern can be understood, for example, as selecting one from a set of pre-defined sound ray patterns. Determining the sound ray direction pattern can be based on the scene type of the audio scene, available computing resources, encoder presets, or a combination thereof. The scene type can include, for example, an indoor scene and / or an outdoor scene. In some embodiments, the scene type can be completely indoor, completely outdoor, or a combination of both indoor and outdoor.

[0092] In step S203, based on the plurality of sound rays determined in step S202 and the voxel-based representation of the three-dimensional audio scene, a set of collision voxels is determined. The set of collision voxels can be determined according to the method described in connection with FIG. 15.

[0093] In step S204, an initial reflection trajectory is determined based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test. The initial reflection trajectory can be determined according to the method described in connection with FIGS. 16 and / or 17.

[0094] In any step S205, a set of initial reflection trajectories that are acoustically most relevant is selected from the ER trajectories 106. This selection may mean discarding at least one of the ER trajectories 106.

[0095] In any step S206, the ER trajectories 106 are output for rendering a three-dimensional audio scene. The ER trajectories 106 may be the ER trajectories of step S205 or step S206.

[0096] FIG. 15 shows a method 300 for determining a set of collision voxels. The method 300 may perform, for example, step S203.

[0097] In step 301, one or more intersections between each sound ray of a plurality of sound rays and the occlusion voxels 105 are determined.

[0098] In step 302, for each sound ray, the occlusion voxel 105 that includes the intersection closest to the starting point of the respective sound ray is determined as the collision voxel 104. The collision voxels 104 determined in this way form a set of collision voxels.

[0099] FIG. 16 shows a method 400 for determining initial reflection trajectories. The method 400 may perform, for example, step S204.

[0100] In step 401, for each collision voxel in the set of collision voxels, it is determined whether the collision voxel can generate a geometrically valid representation of a primary reflection.

[0101] In step 402, if a collision voxel can generate a geometrically valid representation of a primary reflection, a path connecting the listener location and the audio source location through each collision voxel is determined as an initial reflection trajectory.

[0102] FIG. 17 shows a method 500 for determining an initial reflection trajectory. The method 500 may perform step S204, for example.

[0103] In step 501, for each collision voxel in the set of collision voxels, a path connecting the listener location and the audio source location through each respective collision voxel 104 is determined. The path can include two straight line segments as described above. In other words, the audio source 101 and the listener 102 can be connected in a straight line through the collision voxel 104.

[0104] In step 502, for each path determined in step S501, if the path is geometrically valid, each respective path is determined to be an ER trajectory 106. The path can be determined to be geometrically valid if it is not blocked by shielding voxels other than each respective collision voxel.

[0105] Although a method for estimating the ER trajectory has been described above, the present disclosure similarly relates to corresponding apparatuses and the like. Embodiments providing such an apparatus will next be described with reference to FIG. 18.

[0106] As shown in FIG. 18, the apparatus 400 includes a processor 401 and a memory 402. The memory 402 is configured to store program code. The processor 401 is configured to execute the instructions of the program code, and as a result, the apparatus 400 performs the ER trajectory estimation method in any one of the above-described embodiments and implementation forms. The processor 401 can also receive suitable input data (such as, for example, a voxel grid, voxel data, an audio source, and a listener location, etc.) depending, inter alia, on the use case and / or implementation form. The processor 401 can be adapted to perform the methods / techniques described throughout the present disclosure (such as, for example, methods 200, 300, 400, and 500 shown above with reference to FIGS. 14 to 17 respectively) depending on the use case and / or implementation form, and generate corresponding output data (such as, for example, an ER trajectory, etc.). The apparatus may be part of a virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) apparatus. Further, the apparatus may be associated with a decoder device (decoder-side device) or a rendering device, for example, in the context of a VR / AR / MR / XR environment.

[0107] Aspects of the systems described herein may be implemented in a suitable computer-based sound processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks including any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0108] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein can be implemented using any number of combinations of hardware, firmware, and / or as data and / or instructions implemented in various machine-readable or computer-readable media with respect to their behavior, register transfers, logical components, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions can be embodied include various forms of physical (non-transitory) non-volatile storage media such as, but not limited to, optical storage media, magnetic storage media, or semiconductor storage media.

[0109] Although one or more embodiments have been described by way of example with respect to specific embodiments, it should be understood that one or more embodiments are not limited to the disclosed embodiments. It is a block diagram showing the configuration of an image forming apparatus according to an embodiment of the present invention. Accordingly, the appended claims should be given the broadest interpretation so as to encompass all such modifications and similar configurations.

[0110] Interpretation A computing device implementing the above technology can have the following exemplary architecture. Other architectures including those with more or fewer components are also possible. In some implementations, the exemplary architecture includes one or more processors (e.g., dual-core Intel® Xeon® processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can communicate and exchange data via one or more communication channels (e.g., buses), and the communication channels can utilize various hardware and software to facilitate the transfer of data and control signals between components.

[0111] The term "computer-readable media" refers to media involved in providing instructions to a processor for execution, including but not limited to non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media includes, but is not limited to, coaxial cables, copper wire, and optical fibers.

[0112] The computer-readable medium can further include an operating system (e.g., Linux (registered trademark) operating system), a network communication module, an audio interface manager, an audio processing manager, and a live content distributor. The operating system performs basic tasks including recognizing input from network interfaces and / or devices and providing output thereto, tracking and managing files and directories on a computer-readable medium (e.g., memory or storage device), controlling peripheral devices, and managing traffic on one or more communication channels, but not limited thereto. The network communication module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.).

[0113] The architecture can be implemented in a parallel processing or peer-to-peer infrastructure or on a single device having one or more processors. The software can include multiple software components or can be a single code body.

[0114] The described features can be advantageously implemented in one or more computer programs executable on a programmable system that includes a data storage system, at least one input device, and at least one output device, coupled to receive data and instructions from, and to send data and instructions to, at least one programmable processor. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a particular activity or to cause a particular result. The computer program can be written in any form of programming language (including, for example, Objective-C, Java (registered trademark)) including a compiled or interpreted language, and can be deployed in any form, including as a stand-alone program, or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0115] Processors suitable for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and any one or more processors or cores of any kind of computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. In general, a computer also includes, or is operatively coupled to communicate with, one or more mass storage devices for storing data files, such devices including magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks, including all forms of nonvolatile memory. The processor and memory may be supplemented by, or incorporated in, an ASIC (Application Specific Integrated Circuit).

[0116] For providing interaction with a user, the features can be implemented on a computer having a display device such as a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display) monitor, or a retinal display device for displaying information to the user. The computer can have a touch surface input device (e.g., a touch screen), or a keyboard, and a pointing device such as a mouse or a trackball, by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user.

[0117] The features can be implemented in a computer system that includes backend components such as a data server, or includes middleware components such as an application server or an Internet server, or includes frontend components such as a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, for example, computers and networks that form a LAN, a WAN, and the Internet.

[0118] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between a client and a server arises from computer programs that are executed on respective computers and have a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a client device (for the purpose of, e.g., displaying the data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., as a result of user interaction) can be received at the server from the client device.

[0119] A system of one or more computers can be configured to perform particular actions by having software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes actions to be performed on the system. One or more computer programs can be configured to perform particular actions by including instructions that, when executed by a data processing apparatus, cause the apparatus to perform actions.

[0120] This specification includes many details of specific implementations, which should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. The specific features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, features are described above as acting in a certain combination and may even initially be claimed as such, but one or more features from the claimed combination can, in some cases, be deleted from the combination, and the claimed combination can be directed to a sub-combination, or a variation of a sub-combination.

[0121] Similarly, operations are shown in the drawings in a particular order, which should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the shown operations be performed, in order to achieve a desired result. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

[0122] Unless otherwise specified, as is apparent from the following description, throughout the description of the present invention, descriptions using terms such as "processing", "calculating", "computing", "determining", "analyzing", etc. refer to operations and / or conversions of data represented as physical quantities such as electronic quantities into other data also represented as physical quantities by a computer or computing system, or similar electronic computing devices, or operations and / or processes thereof.

[0123] Throughout the present invention, references to "one embodiment", "some embodiments", or "an embodiment" mean that the particular features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment", "in some embodiments", or "in an embodiment" in various places throughout the present invention are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from the present invention.

[0124] As used herein, unless otherwise specified, the use of adjectives such as "first", "second", "third", etc. to indicate an order for describing common objects merely indicates that different examples of similar objects are being referred to, and is not intended to mean that the objects so described must be in a given order, whether temporally, spatially, in ranking, or in any other way.

[0125] Also, it should be understood that the expressions and terms used in this specification are for purposes of explanation and should not be regarded as limiting. The use of "including", "comprising", or "having", and variations thereof, means including the items listed thereafter and their equivalents, as well as additional items. Unless specifically specified or limited, the terms "attached", "connected", "supported", and "coupled", and variations thereof, are used broadly and include both direct and indirect attachment, connection, support, and coupling.

[0126] In the following claims and the description of this specification, any one of the terms "comprising", "comprised of", or "which comprises" is an open term meaning including at least the elements / features that follow it, but not excluding others. Thus, when used in the claims, the term "comprising" should not be construed as being limited to the means or elements or steps listed thereafter. For example, the scope of the expression "a device comprising A and B" should not be limited to a device consisting only of elements A and B. Any one of the terms "including", "which includes", or "that includes" used in this specification is also an open term meaning including at least the elements / features that follow that term, but not excluding others. Thus, "comprising" is synonymous with "including" and means including.

[0127] In the foregoing description of exemplary embodiments of the present invention, it should be understood that various features of the present invention may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of streamlining the present invention and facilitating understanding of one or more of the various aspects of the invention. However, this method of the present invention should not be construed as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as reflected by the following claims, aspects of the invention lie in less than all of the features of a single foregoing disclosed exemplary embodiment. Accordingly, the claims that follow this specification are hereby expressly incorporated herein, and each claim stands on its own as a separate exemplary embodiment of the present invention.

[0128] Furthermore, some of the exemplary embodiments described herein include some features included in other exemplary embodiments but not others, and combinations of features of different exemplary embodiments are meant to be within the scope of the present invention and, as will be understood by those skilled in the art, form different exemplary embodiments. For example, in the following claims, any of the claimed exemplary embodiments may be used in any combination.

[0129] In the description provided herein, numerous specific details are set forth. However, it should be understood that exemplary embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0130] Accordingly, while the best mode contemplated by the inventors of the present invention has been described, those skilled in the art will recognize that additional changes could be made thereto without departing from the spirit of the invention and all such changes and modifications are intended to be claimed within the scope of the invention. For example, any of the above-described formulas are merely representative examples of procedures that may be used. Functions may be added to or removed from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or removed from the described methods within the scope of the present disclosure.

[0131] Various aspects and implementations of the present disclosure may also be understood from the following enumerated exemplary embodiments (EEEs), which are not the claims of the patent.

[0132] [EEE1] A method for estimating an initial reflection trajectory of an audio source in a three-dimensional audio scene, comprising: obtaining a voxel-based representation of the three-dimensional audio scene, information regarding a listener location of a listener in the three-dimensional audio scene, and information regarding an audio source location of the audio source in the three-dimensional audio scene; applying a sound ray direction pattern to one or more points on a connection line between the audio source location and the listener location to obtain, for each of the one or more points, a plurality of sound rays starting from the respective point; determining a set of collision voxels based on the plurality of sound rays and the voxel-based representation of the three-dimensional audio scene; determining an initial reflection trajectory based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test; and a method including the above steps.

[0133] [EEE1A] A method for estimating an initial reflection trajectory of an audio source in a three-dimensional audio scene, comprising: Obtaining a voxel-based representation of the three-dimensional audio scene, information regarding the listener location of a listener in the three-dimensional audio scene, and information regarding the audio source location of the audio source in the three-dimensional audio scene; Applying a sound ray direction pattern to one or more points on a connection line between the audio source location and the listener location to obtain, for each of the one or more points, a plurality of sound rays starting from the respective point; Determining a set of collision voxels based on the plurality of sound rays and the voxel-based representation of the three-dimensional audio scene; For each collision voxel in the set of collision voxels, determining a path connecting the listener location and the audio source location through the respective collision voxel; For each path, if the path is geometrically valid, determining the path as an initial reflection trajectory; A method comprising the above.

[0134] [EEE2] Determining the sound ray direction pattern The method according to EEE1 or EEE1A, further comprising the above.

[0135] [EEE3] Determining the one or more points based on the obtained density of the one or more points The method according to EEE1, 1A or 2, further comprising the above.

[0136] [EEE4] The sound ray direction pattern defines a predetermined number of sound rays and the direction of a predetermined sound ray from a starting point. The method according to any one of EEE1 to 3 or 1A.

[0137] [EEE5] The predetermined number of sound rays is 6, 8 or 12. The method according to EEE4.

[0138] [EEE6] The voxel positions in the three-dimensional audio grid are defined by grid indices, and the direction of the predetermined sound ray is one or more of the horizontal and vertical directions of the grid index with respect to the neighboring grid indices and the diagonal direction of the grid index with respect to the neighboring grid indices. The method described in EEE5.

[0139] [EEE7] Determining the sound ray direction pattern is based on the scene type of the three-dimensional audio scene, available computing resources, encoder preset, or a combination thereof. The method described in EEE2.

[0140] [EEE8] The coordinates of the one or more points on the line connecting the audio source location and the listener location are determined based on the density of the one or more points. The method described in EEE3.

[0141] [EEE9] The one or more points are determined to divide the line connecting the audio source location and the listener location into N - 1 equal line segments, where N is the density of the one or more points and is greater than or equal to 2. The method described in EEE8.

[0142] [EEE10] The density of the one or more points depends on the scene type of the three-dimensional audio scene, available computing resources, encoder preset, or a combination thereof. The method described in EEE3.

[0143] [EEE11] The scene type includes indoor scenes and outdoor scenes. The method described in EEE7 or 10.

[0144] [EEE12] Each collision voxel in the set of the collision voxels is a shielding voxel in the voxel-based representation of the three-dimensional audio scene, and the method according to any one of EEE1 to 11 or EEE1A.

[0145] [EEE13] The shielding voxel represents an acoustic reflection surface, and the method according to EEE12.

[0146] [EEE14] The shielding voxel represents any substance other than air in the voxel-based representation of the three-dimensional audio scene, and the method according to EEE12.

[0147] [EEE15] Determining the set of the collision voxels based on the plurality of sound rays and the voxel-based representation of the three-dimensional audio scene includes: determining one or more intersections between each sound ray of the plurality of sound rays and the shielding voxel; for each sound ray, determining, as a collision voxel in the set of the collision voxels, a shielding voxel including the intersection closest to the starting point of the respective sound ray; and the method according to any one of EEE12 to 14.

[0148] [EEE16] Determining the initial reflection trajectory based on the set of the collision voxels, the listener location, the audio source location, and the geometric validity test includes: for each collision voxel in the set of the collision voxels, determining whether the collision voxel can generate a geometrically valid representation of a primary reflection; if the collision voxel can generate a geometrically valid representation of a primary reflection, determining, as the initial reflection trajectory, a path connecting the listener location and the audio source location through the respective collision voxel; and The method according to any one of EEE1 to 15 or EEE1A.

[0149] [EEE17] Determining whether the collision voxel can generate a geometrically valid representation of the primary reflection is determining the preceding voxels of the collision voxel, where the preceding voxels include intersections with the respective sound rays and are voxels preceding the collision voxel in the direction of the respective sound rays, determining a second path connecting the listener location and the audio source location through the respective preceding voxels, if the second path does not include an intersection with an occlusion voxel, determining that the collision voxel can generate a geometrically valid representation of the primary reflection, including The method according to EEE16.

[0150] [EEE18] Determining the initial reflection trajectory based on the set of collision voxels, the listener location, the audio source location, and the geometric validity test is for each collision voxel in the set of collision voxels, determining a path connecting the listener location and the audio source location through the respective collision voxel for each path, if the path is geometrically valid, determining the path as an initial reflection trajectory, including The method according to any one of EEE1 to EEE17 or EEE1A.

[0151] [EEE19] The path includes a straight line connecting the audio source location to a collision voxel in the set of collision voxels and a straight line connecting the same collision voxel in the set of collision voxels to the listener location. The method according to EEE16 or EEE18.

[0152] [EEE20] If the path does not include an intersection with a shielding voxel other than the collision voxel of each respective path, the path is determined to be geometrically valid. The method according to EEE18.

[0153] [EEE21] Selecting a set of initial reflection trajectories that are acoustically most relevant from the initial reflection trajectories The method according to any one of EEE1 to 20 or EEE1A, further comprising.

[0154] [EEE22] Selecting a set of initial reflection trajectories that are acoustically most relevant is based on the length of the initial reflection trajectories and / or the reflection coefficient of the collision voxels of the initial reflection trajectories, according to the method described in EEE21.

[0155] [EEE23] The reflection coefficient depends on the substance modeled by the collision voxel, according to the method described in EEE22.

[0156] [EEE24] Selecting a set of initial reflection trajectories that are acoustically most relevant includes discarding initial reflection trajectories having a value indicating an interior angle greater than 160° at the collision voxel, according to any one of EEE21 to 23.

[0157] [EEE25] The value indicating an interior angle greater than 160° is the interior angle or length of the initial reflection trajectory, according to the method described in EEE24.

[0158] [EEE26] Outputting the initial reflection trajectories for rendering of the three-dimensional audio scene The method according to any one of EEE1 to 25 or EEE1A, further comprising.

[0159] [EEE27] The rendering is performed by a virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) device, and is the method described in EEE26.

[0160] [EEE28] The initial reflection trajectory represents a primary trajectory, and is the method described in any one of EEE1 to 27 or EEE1A.

[0161] [EEE29] The primary trajectory is a reflection trajectory having a single reflection between the audio source location and the listener location, and is the method described in EEE28.

[0162] [EEE30] The method is performed by a decoder or a renderer, and is the method described in any one of EEE1 to EEE29 or EEE1A.

[0163] [EEE31] A method for processing a frame of a three-dimensional audio scene, comprising: estimating an initial reflection trajectory of the frame based on the method according to any one of claims 1 to 30, and storing the estimated initial reflection trajectory, or if the voxels including the listener location, the voxels including the audio source location, and the geometry of the voxel-based representation of the three-dimensional audio scene do not change between the frame and the previous frame, accessing the estimated initial reflection trajectory of the previous frame estimated based on the method according to any one of claims 1 to 30, The method includes the above steps.

[0164] [EEE32] An apparatus including a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method described in any one of EEE1 to 31 or EEE1A.

[0165] [EEE33] A program that, when executed by a processor, includes instructions that cause the processor to perform the method according to any one of EEE1 to 31 or EEE1A.

[0166] [EEE34] A computer-readable storage medium storing the program according to EEE33.

[0167] [EEE35] An audio processing method for creating a trajectory for a geometrically connected audio source for efficient implementation on a voxel 3D grid, Receiving information regarding a sound ray direction pattern "R", Determining a first set "P" of points for applying ray casting based on the sound ray direction pattern "R", Determining a second set "C" of sound ray - voxel "collision" voxels based on the first set of points and a reflection voxel "VOX", Determining a third set "S - C - L" of valid reflection trajectories based on the second set "C" of sound ray - voxel "collision" voxels, Selecting and outputting a subset of the most acoustically relevant reflection trajectories from the third set of valid reflection trajectories, A method including the above.

[0168] [EEE36] The method according to EEE35, wherein the second set "C" of sound ray - voxel "collision" voxels is determined based on the sound ray direction pattern "R" applied to the first "P" of points and the reflection voxel "VOX".

[0169] [EEE37] The method according to EEE35, wherein the reflection trajectory "S - C - L" represents a primary trajectory.

[0170] [EEE38] The method according to EEE35, further comprising checking whether a first line connecting a listener and a collision voxel L-C and a second line connecting an audio source and a collision voxel "S-C" cross any of the cut-off / shielding / reflection voxels, and based on a determination that there is no crossing, determining that this is a reasonable approximation of a primary reflection.

[0171] [EEE39] A non-transitory computer program that, when executed by a processor, includes instructions that cause the processor to perform the method according to any one of EEE35 to 38.

[0172] [EEE40] An apparatus configured to perform the method of EEE35 to 38.

Claims

1. A method for estimating an initial reflection trajectory of an audio source in a three-dimensional audio scene, comprising: obtaining a voxel-based representation of the three-dimensional audio scene, information regarding a listener location of a listener in the three-dimensional audio scene, and information regarding an audio source location of the audio source in the three-dimensional audio scene; applying a sound ray direction pattern to one or more points on a connection line between the audio source location and the listener location to obtain, for each of the one or more points, a plurality of sound rays starting from the respective point; determining a set of collision voxels based on the plurality of sound rays and the voxel-based representation of the three-dimensional audio scene; determining an initial reflection trajectory based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test; A method comprising the above.

2. Further comprising determining the sound ray direction pattern The method according to claim 1.

3. Further comprising determining the one or more points based on the obtained density of the one or more points The method according to claim 1 or 2.

4. The sound ray direction pattern defines a predetermined number of sound rays and directions of the predetermined sound rays from a starting point. The method according to any one of claims 1 to 3.

5. The method according to claim 4, wherein the predetermined number of sound rays is 6, 8, or 12.

6. Voxel positions in the three-dimensional audio grid are defined by grid indices, and the directions of the predetermined sound rays include one or more of a horizontal direction and a vertical direction of the grid index with respect to a neighboring grid index and a diagonal direction of the grid index with respect to the neighboring grid index. The method according to claim 5.

7. Determining the sound ray direction pattern is based on a scene type of the three-dimensional audio scene, available computing resources, an encoder preset, or a combination thereof. The method according to claim 2.

8. Coordinates of the one or more points on the line connecting the audio source location and the listener location are determined based on the density of the one or more points. The method according to claim 3.

9. The one or more points are determined to divide the line connecting the audio source location and the listener location into N - 1 equal line segments, where N is the density of the one or more points and is greater than or equal to 2. The method according to claim 8. Claim 10 The density of the one or more points depends on the scene type of the three-dimensional audio scene, available computing resources, encoder preset, or a combination thereof. The method according to claim 3. Claim 11 The scene type includes indoor scenes and outdoor scenes. The method according to claim 7 or 10. Claim 12 Each collision voxel in the set of collision voxels is an occlusion voxel in the voxel-based representation of the three-dimensional audio scene. The method according to any one of claims 1 to 11. Claim 13 The occlusion voxel represents an acoustic reflecting surface. The method according to claim 12. Claim 14 The occlusion voxel represents any substance other than air in the voxel-based representation of the three-dimensional audio scene. The method according to claim 12. Claim 15 Determining the set of collision voxels based on the plurality of sound rays and the voxel-based representation of the three-dimensional audio scene includes: determining one or more intersections between each sound ray of the plurality of sound rays and the occlusion voxel; and for each sound ray, determining the occlusion voxel including the intersection closest to the starting point of the respective sound ray as a collision voxel in the set of collision voxels. including The method according to any one of claims 12 to 14. Claim 16 Determining the initial reflection trajectory based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test includes: for each collision voxel in the set of collision voxels, determining whether the collision voxel can generate a geometrically valid representation of a primary reflection; and if the collision voxel can generate a geometrically valid representation of a primary reflection, determining the path connecting the listener location and the audio source location through the respective collision voxel as the initial reflection trajectory. including The method according to any one of claims 1 to 15. Claim 17 Determining whether the collision voxel can generate a geometrically valid representation of a primary reflection includes: Determining a preceding voxel of the collision voxel, the preceding voxel including an intersection with each of the sound rays and being a voxel preceding the collision voxel in the direction of each of the sound rays, Determining a second path connecting the listener location and the audio source location through each of the preceding voxels, If the second path does not include an intersection with an occlusion voxel, determining that the collision voxel can generate a geometrically valid representation of a primary reflection, including, The method according to claim 16.

18. Determining an initial reflection trajectory based on the set of collision voxels, the listener location, the audio source location, and a geometric validity test includes, For each collision voxel in the set of collision voxels, determining a path connecting the listener location and the audio source location through each of the collision voxels, For each path, if the path is geometrically valid, determining the path as an initial reflection trajectory, including, The method according to any one of claims 1 to 17.

19. The path includes a straight line connecting the audio source location to a collision voxel in the set of collision voxels and a straight line connecting the same collision voxel in the set of collision voxels to the listener location, The method according to claim 16 or 18.

20. If the path does not include an intersection with an occlusion voxel other than the collision voxel of each path, the path is determined to be geometrically valid, The method according to claim 18.

21. Selecting a set of acoustically most relevant initial reflection trajectories from the initial reflection trajectories, further including, the method according to any one of claims 1 to 20.

22. Selecting the set of acoustically most relevant initial reflection trajectories is based on the length of the initial reflection trajectories and / or the reflection coefficient of the collision voxels of the initial reflection trajectories, The method according to claim 21.

23. The reflection coefficient depends on the material modeled by the collision voxel, The method according to claim 22.

24. Selecting the set of acoustically most relevant initial reflection trajectories includes discarding initial reflection trajectories having a value indicating an interior angle close to 180° at the collision voxel, The method according to any one of claims 21 to 23.

25. The value indicating the interior angle close to 180° is the interior angle or the length of the initial reflection trajectory. The method according to claim 24. **Claim 26** Outputting the initial reflection trajectory for rendering the three-dimensional audio scene The method according to any one of claims 1 to 25, further comprising: **Claim 27** The rendering is performed by a virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) device. The method according to claim 26. **Claim 28** The initial reflection trajectory represents a primary trajectory. The method according to any one of claims 1 to 27. **Claim 29** The primary trajectory is a reflection trajectory having a single reflection between the audio source location and the listener location. The method according to claim 28. **Claim 30** The method according to any one of claims 1 to 29, performed by a decoder or a renderer. **Claim 31** A method for processing a frame of a three-dimensional audio scene, comprising: estimating an initial reflection trajectory of the frame based on the method according to any one of claims 1 to 30 and storing the estimated initial reflection trajectory; or if the voxels including the listener location, the voxels including the audio source location, and the geometry of the voxel-based representation of the three-dimensional audio scene do not change between the frame and the previous frame, accessing the estimated initial reflection trajectory of the previous frame estimated based on the method according to any one of claims 1 to 30. The method comprising: **Claim 32** An apparatus including a processor and a memory coupled to the processor, wherein the processor is adapted to perform the method according to any one of claims 1 to 31. **Claim 33** A program including instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 31. **Claim 34** A computer-readable storage medium storing the program according to claim 33.