Methods, apparatus, and systems for environment type modelling of early reflection gain(s)
By dynamically determining an environment-dependent gain factor based on the quantity of valid early reflection trajectories and total rays, the method addresses the limitations of fixed gain approaches in modeling early reflections, enhancing the realism and immersion of audio experiences in changing environments.
Patent Information
- Application Number
- PCT/EP2024/083655
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
Current methods for modeling early reflections in audio scenes use a fixed gain that depends on the environment, which can be suboptimal for scenarios involving transitions between indoor and outdoor environments.
A method for determining an environment-dependent gain factor for early reflections by analyzing the quantity of valid early reflection trajectories and the total quantity of rays in a three-dimensional audio scene, allowing for dynamic adjustment of energy levels based on environmental changes.
This approach improves the immersion in virtual reality and similar applications by ensuring that the energy level of early reflections accurately fits the perceived environment, providing a more realistic audio experience.
Smart Images

Figure EP2024083655_05062025_PF_FP_ABST
Abstract
Description
[0001] METHODS, APPARATUS, AND SYSTEMS FOR ENVIRONMENT TYPE MODELLING OF EARLY REFLECTION GAIN(S)
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims benefit of priority to U.S. Provisional Patent Application No. 63 / 602,974 filed November 27, 2023, U.S. Provisional Patent Application No. 63 / 570,389 filed March 27, 2024 and European Patent Application 24167049.6 filed March 27, 2024, which are incorporated herein by reference.
[0004] TECHNICAL FIELD
[0005] The present disclosure relates to modelling of early sound source reflection and more particularly to methods and devices for estimating environment dependent gains for the modelling of early reflections.
[0006] BACKGROUND
[0007] Sound reflections of an acoustically reflective surface can influence the perceived sound of an audio source. Sounds that are reflected and received shortly after direct sound at a target location (e.g., a listener position), which herein will be referred to as Early Reflection (ER), are of particular interest when modelling a sound source, as the perceived sound of an audio source can be accurately modelled with only considering direct sound and ERs. Higher order acoustic reflections on the other hand are often less important because they are lower in energy and temporally / spatially psychoacoustically masked by ERs and other components.
[0008] Therefore, correctly modeling ERs for Virtual Reality (VR), Extended Reality (XR), Augmented Reality (AR) or similar applications is of high relevancy for the immersion perceived by a user. Currently, ERs are modelled with a fixed gain that depends on the environment, i.e., indoor or outdoor scenarios. Using a fixed gain may provide a sufficiently realistic psychoacoustic impression for some scenarios, but may be suboptimal for certain edge cases, e.g., switching from an indoor to an outdoor environment and vice versa.
[0009] Thus, there is a need for an improved approach of determining an environment dependent gain for ER modeling.
[0010] SUMMARY
[0011] In view of the above, the present disclosure provides methods, apparatus, and programs, as well as computer-readable storage media for an environmental factor determination for modeling early reflections in an audio scene, having the features of the respective independent claims. The audio scene may be received from an external source via a bitstream. The bitstream may have to be decoded to retrieve the audio scene. Alternatively, the audio scene may be retrieved from local memory. The audio scene may be compressed in the local memory and may be decompressed in order to further process the audio scene.
[0012] According to an aspect of the disclosure, a method of determining an environmental gain factor for early reflections is provided. The gain factor may indicate whether a three-dimensional audio scene is indoor, outdoor or a mixture of indoor and outdoor. A representation of the three- dimensional audio scene, information on a listener location of a listener in the three-dimensional audio scene, and information on an audio source location of the audio source in the three- dimensional audio scene may be obtained, e.g., received and / or determined. Further, a predefined quantity of rays for a ray direction pattern and a quantity of points for applying the ray direction pattern may be obtained, e.g., received and / or determined. An indication for a quantity of valid early reflection trajectories for the audio source and the listener may be determined based on the representation of the three-dimensional audio scene, the ray direction pattern, the quantity of points, the listener location and the audio source location. A valid early reflection trajectory may represent a path from the audio source location to the listener location that can produce a geometrically valid representation of a first-order reflection of the audio source, i.e., the path may not be obscured such that the first-order reflection of sound from the audio source may travel this path. The indication for the quantity of valid early reflection trajectories may be the quantity of valid early reflection trajectories or a value indicative of the quantity of valid early reflection trajectories. The gain factor may be determined based on the indication for the quantity of valid early reflection trajectories and a total quantity of rays, wherein the total quantity of rays is determined based on the predefined quantity of rays and the quantity of points.
[0013] By determining the gain factor for early reflections modeling in this way, the energy level for modeling early reflections can be dynamically adjusted when the environment of the audio scene changes, or a location of the audio source or the listener changes. The audio scene modelled for early reflections based on the gain factor may improve the immersion by a user listing and viewing the scene in a VR, XR or similar device, as the observed energy level of the early reflections better fit the perceived environment.
[0014] In some embodiments, the method may further include outputting the gain factor for rendering of the three-dimensional audio scene. Outputting the gain factor may include providing the gain factor to a process for determining energy levels of the early reflections. The energy levels of the early reflections may be mixed with other elements related to the rendering of a VR audio scene.
[0015] In some embodiments, determining the gain factor based on the indication for the quantity of valid early reflection trajectories and the total quantity of rays may include dividing the indication for the quantity of valid early reflection trajectories by the total quantity of rays. In some implementations, the division may be implemented by a multiplication of the indication for the quantity of valid early reflection trajectories with the inverse of the total quantity of rays.
[0016] In some embodiments, the gain factor may be defined as X = a fa+ b fb, fb= 1 — fa, wherein fais a function of the total quantity of rays and the indication for the quantity of valid early reflection trajectories, and wherein a and b are scalars for scaling the gain factor. The values of a and b may be set by a content creator to reflect the artistic intent. Alternatively, a and b may be set to tune the gain factor for different indoor and outdoor scenarios. In some implementations, famay be viewed as a measure of how close the audio scene is to a pure indoor scene, while fbmay be viewed as a measure of how close the audio scene is to a pure outdoor scene. A pure outdoor scene may be a scene with no reflective surfaces, while a pure indoor scene may be a room with 4 walls or a circular room. In a preferred embodiment, a = 2, and b = 1. Then the gain factor may be defined as X = ( — - — ) Nhits+ 1, wherein Nraysis the total quantity of rays and Nhitsis the Ways / yindication for the quantity of valid early reflection trajectories. The total quantity of rays may be defined as the quantity of rays multiplied by the quantity of points. In some embodiments, the gain factor is used to determine an early reflections gain. The early reflections gain may be determined for a sector used for modelling the early reflections. The sector may be determined by dividing a 360-degree angle around the listener position in equiangular portions. For example, there may be 4, 8 or 16 sectors with 90°, 45° or 22,5°, respectively. A quantity of equiangular portions may be indicative of a precision for the modelling of the early reflections.
[0017] In some embodiments, the early reflections gain may be directly proportional to the gain factor multiplied by a mean value of reflection attenuation gains for the audio source in the sector. The mean value of reflection attenuation gains for the audio source in the sector may be defined as Fk= wherein k is the index of the sector, gER^ is the reflection attenuation gain for early reflection trajectory i in sector k, nkis the total quantity of early reflection trajectories in sector k, and Nsis the maximum quantity of reflections per audio source for the reflection modelling.
[0018] In some embodiments, the early reflections gain may be determined for an image audio source in the sector. The image audio source may be a virtual audio source perceived by the listener due to an early reflection. The location of the image audio source may be determined by an image source method. The location of the image audio source may be a location of the audio source mirrored on a reflective surface in the three-dimensional audio scene. The image audio source signal may be determined by downmixing all determined early reflection signals from the respective trajectories in one sector to create a single substitution early reflection signal for the sector. The resulting single image audio source location and gain may be obtained by averaging the location and gain of all image audio sources of the one sector, i.e., the image audio sources corresponding to the determined early reflections in the sector.
[0019] In some embodiments, the representation of the three-dimensional audio scene may a voxel-based representation.
[0020] In some embodiments, determining the indication for the quantity of valid early reflection trajectories for the audio source and the listener may include applying the ray direction pattern to one or more points scattered in a proximity to a region connecting the audio source location and the listener location to obtain, for each of the one or more points, a plurality of rays originating at the respective point. A set of collision voxels may be determined based on the plurality of rays and the voxel-based representation of the three-dimensional audio scene. Collision voxels in the set of collision voxels may be counted to determine the quantity of collision voxels if the indication of the quantity of valid early reflection trajectories is the quantity of collision voxels. Alternatively, if the indication of the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories, early reflection trajectories may be determined based on the set of collision voxels, the listener location, the audio source location and a geometrical validity test. For example, for each collision voxel in the set of collision voxels, a preceding voxel may be determined. The preceding voxel may be a voxel containing an intersection with the respective ray, preceding the collision voxel in the direction of the respective ray. For each preceding voxel, a path connecting the listener location and the audio source via the preceding voxel may be determined. Then, for each path, the path may be determined as an early reflection trajectory if the path is geometrically valid. The early reflection trajectories may be counted to obtain the quantity of valid early reflection trajectories.
[0021] In some embodiments, the proximity to the region connecting the audio source location and the listener location may be a line connecting the audio source location and the listener location
[0022] In some embodiments, the method may further include determining the one or more points based on a number (e.g., count, cardinality) of the one or more points. That is, a number of the one or more points may be obtained or determined (e.g., set to be N points) and the resulting (e.g., N) number (count or cardinality) of the one or more points may correspond to coordinates of the one or more points (e.g., in the sense that for each of the one or more points there are respective coordinates).
[0023] In some embodiments, the ray direction pattern may define the predefined quantity of rays and predefined directions of rays from an origin. The predefined number of rays may be 6, 8, 12, or 26 for example. The directions of rays can be defined by grid indices of the voxel grid.
[0024] In some embodiments, the predefined directions of rays may include one or more of: horizontal and vertical directions to neighboring grid indices; and diagonal directions to neighboring grid indices. Therefore, the predefined directions may define relative directions from an origin of the rays, i.e., a grid index (l,m,i) in the voxel grid. The relative directions can be expressed as:
[0025] (+1,0,0), (-1,0,0), (0,+l,0), (0,-l,0), (0,0, +1), (-0,0,-l); (+l,+l,O), (+1,- 1,0), (-l,+l,0), (-l,-l,0), (+1,0, +1), (+l,0,-l), (-1,0, +1), (-l,0,-l), (0,+l,+l), (0,+l,-l), (0,-1, +1), (0,-1 ,- 1); and
[0026] (+1,+1,+1), (+1,+1,-1), (+1,-1, +1), (+1,— 1,— 1), (— 1,+1,+1), (— 1,+1,— 1), (-1,-1, +1), (— 1,— 1,— 1).
[0027] In some embodiments, coordinates of the one or more points on the line connecting the audio source location and the listener location may be determined based on the number (e.g., count, cardinality) of the one or more points.
[0028] In some embodiments, the one or more points may be determined to split the line connecting the audio source location and the listener location into N-l equal segments, where N is the number (e.g., count, cardinality) of the one or more points. N may be larger than or equal to 2, for example.
[0029] In some embodiments, each collision voxel may be an occluder voxel in the voxel-based representation of the three-dimensional audio scene.
[0030] In some embodiments, the occluder voxel may represent an acoustically reflective surface.
[0031] In some embodiments, the occluder voxel may represent any material in the voxel-based representation of the three-dimensional audio scene other than sound transmission media, e.g., air. That is, the occluder voxel may represent a reflective surface and a non-occluding voxel may represent a non-reflective surface (or not define a surface at all).
[0032] In some embodiments, determining the set of collision voxels based on the plurality of rays and the voxel-based representation of the three-dimensional audio scene may include determining one or more intersections (e.g., intersection points) between each ray of the plurality of rays and the occluder voxels. The method may further include, for each ray, determining an occluder voxel containing an intersection closest to the origin of the respective ray as a collision voxel in the set of collision voxels. That is, the collision voxel may be an occluder voxel first hit by a respective ray.
[0033] In some embodiments, determining whether the path connecting the listener location and the audio source location via the preceding voxel is geometrically valid may include passing a line- of-sight check (“visibility check”). The line-of-sight check may include traversing the path and determining whether the path intersects with an occluder voxel. If the path does not intersect an occluder voxel, the corresponding path may be determined as a valid early reflection trajectory. In other words, a path may be determined as geometrically valid if it is not obstructed by any occluding voxels.
[0034] To achieve some computational saving, the line-of-sight check between the audio source location and the preceding voxel may not need be performed when the corresponding collision voxel is obtained from a respective ray originating from the audio source location. This is due to the fact that the collision voxel is the first occluder to be hit by the respective ray, ensuring the line-of- sight fulfillment. Similarly, the line-of-sight check between the listener location and the preceding voxel may not need be performed when the corresponding collision voxel is obtained from a respective ray originating from the listener location.
[0035] Thereby, paths that cannot lead to a geometrically valid path from the audio source location to the listener position can be efficiently sorted out.
[0036] In some embodiments, the path may include a straight line connecting the audio source location to the preceding voxel and a straight line connecting the same preceding voxel to the listener location.
[0037] Aspects of the present disclosure may be implemented via an apparatus. The apparatus may include a processor and memory coupled to the processor. The processor may be adapted carry out the method according to aspects and embodiments of the present disclosure.
[0038] Aspects of the present disclosure may be implemented via a program. When instructions of the program are executed by a processor, the processor may carry out aspects and embodiments of the present disclosure. A computer-readable storage medium may store the program. Such computer-readable storage media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc.. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more computer-readable storage media having software stored thereon.
[0039] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus (or system), and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) are understood to likewise apply to the corresponding apparatus (or system), and vice versa. BRIEF DESCRIPTION OF DRAWINGS
[0040] Example embodiments of the disclosure are explained below with reference to the accompanying drawings, wherein
[0041] Fig. 1 is a diagram showing an example of an echogram of a room,
[0042] Fig. 2 schematically illustrates an example of a transition between an outdoor and an indoor environment,
[0043] Fig. 3 schematically illustrates an example of determining ERs with the IS method,
[0044] Fig. 4 schematically illustrates a voxel grid, an audio source, a listener, and a single collision voxel considered for reflection trajectory estimation,
[0045] Fig. 5 schematically illustrates an example 2D audio scene with occluding voxels (dotted), nonoccluding voxels (plain), an audio source, and a listener according to embodiments of the disclosure,
[0046] Fig. 6 schematically illustrates the example 2D audio scene of Fig. 5 and a ray direction pattern applied to the audio source location and collision voxels hit by the rays according to embodiments of the disclosure,
[0047] Fig. 7 schematically illustrates the example 2D audio scene of Fig. 6 and lines connecting the audio source to the listener via the collision voxels according to embodiments of the disclosure,
[0048] Fig. 8 schematically illustrates the example 2D audio scene of Fig. 7 and a geometrically valid ER trajectory according to embodiments of the disclosure,
[0049] Fig. 9 schematically illustrates the example 2D audio scene of Fig. 5 and all geometrically valid ER trajectories according to embodiments of the disclosure,
[0050] Figs. 10A and 10B schematically illustrate examples of determining energy levels of ERs in a sector for an outdoor and indoor environment,
[0051] Fig. 11 schematically illustrates an example of a listener transition from one room to another room.
[0052] Fig. 12 schematically illustrates a ray direction pattern according to embodiments of the disclosure, Fig. 13 is a flowchart illustrating an example of a method of determining a gain factor for modeling ERs according to embodiments of the disclosure,
[0053] Fig. 14 is a flowchart illustrating an example of a method of estimating a quantity of ERs in a voxel-based audio scene representation according to embodiments of the disclosure,
[0054] Fig. 15 schematically illustrates an example of an apparatus for determining a gain factor for modelling early reflections of an audio source according to embodiments of the disclosure, and
[0055] Fig. 16 schematically illustrates an example of a mixed indoor-outdoor environment.
[0056] DETAIEED DESCRIPTION
[0057] The Figures (Figs.) and the following description relate to preferred embodiments by way of illustration only. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of what is claimed.
[0058] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying figures. It is noted that wherever practicable similar or like reference numbers may be used in the figures and may indicate similar or like functionality. The figures depict embodiments of the disclosed system (or method) for purposes of illustration only. One skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.
[0059] ERs evoke several perceptual effects such as apparent source width, perceived distance, timbre, and spaciousness. ERs are relatively sparse in time and span a relatively short time usually contained within the first ~80ms of a room impulse response (see Fig. 1). Figure 1 illustrates an echogram of a room, including the echogram for a direct sound source, early reflections, and late reflections. Figure 1 also allows for visualization as to the differences between direct sound, early reflections and late reflections.
[0060] The psychoacoustical relevance of the ER largely depends on several factors such as the direction, level, time delay and spectral content of the audio signal. The energy level of ERs depends on the environment type in which the audio source and the listener are currently positioned. Therefore, accurate modeling of the ERs energy level has a high relevancy for the psychoacoustic effect of an audio scene. The audio scene may be representative of audio experienced in a 3D audio environment, e.g., an indoor or outdoor environment. As an example, modeling the energy level of ERs too low, e.g. lower than the energy level of the direct sound, in an indoor environment could break the immersion of the user, as the sound experience may not fit to the visual experience.
[0061] An example of the influence of different environments on ERs is depicted in Fig. 2. In particular, scenario (A) depicts an outdoor environment, and scenario (B) depicts an indoor environment. In an indoor environment, the listener is typically surrounded by walls from all directions.
[0062] Therefore, the listener may experience a large energy level from the ERs due to the many reflective surfaces. On the other hand, in the outdoor scenario, a few or no reflective surfaces may surround the listener and the audio source. Therefore, the energy level of ERs may be much lower than in the outdoor case. As the energy level of ERs heavily depends on the number of reflective surfaces and consequently the number of ERs, determining the energy level for modelling of ERs may be based on the number or quantity of ERs or another value indicative of the quantity of ERs. Therefore, a value indicative of the quantity of ER trajectories may have to be determined.
[0063] To estimate the trajectories of ERs, the Image-Source (IS) method aims to find the purely specular reflection paths between an audio source and a receiver, i.e., a listener. This process is simplified by assuming that sound propagates only along straight lines, i.e., rays. The audio image source is spawn on a line perpendicular to the boundary and at the same distance from it as the original source 101 (see Fig. 3). Figure 3 illustrates a sound source 101, listener 102, a boundary and an image source.
[0064] As the sound is reflected of the boundary surface with the same angle as the incident angle the impression is created that the original source 101 is mirrored at the boundary surface. A reflection by a single boundary then represents an (1storder) ER.
[0065] To model the energy level for ERs, a single ER gain is determined for a sector k. The single ER gain may correspond to a reflection gain of a single image source modeled for sector k. Sector k may be defined by an angle and a direction from a listener position. For example, sector k may defined by an angle of 90 degrees around a direction. In other words, 360 degrees around the listener position may be divided in equiangular pieces to define the different sectors. For 90- degree sectors, a reflection gain would be determined for 4 sectors. The number and corresponding size of the sectors may correspond to an accuracy and a computational effort of modeling ERs.
[0066] The ER gain for a sector k may be determined according to the following relationship: (1)
[0067] In this case g|Ris the single ER gain for ERs in sector k, X is an environment dependent gain factor and Fkis a mean reflection attenuation gain for all rendered audio sources in sector k.
[0068] Fkmay be calculated as wherein Nsis the maximum number of reflections per audio source for the reflection modelling, gERkis a reflection attenuation gain of ER trajectory i in sector k and nkis the number of ER trajectories in sector k. gERkmay be based on a distance dependent attenuation of the respective ER trajectory i and a material coefficient of the material by which the ER is reflected. The distance dependent attenuation may be linear or quadratic and the material coefficient may be between 0 and 1.
[0069] To summarize, for modeling ER energy levels, a single image audio source may be determined for all ERs corresponding to one audio source in sector k, e.g., based on the IS method, and for the single image audio source a single ER gain may be determined. In other words, all ERs for one audio source in sector k are combined into a single substitution ER path. This implies that the maximum number of substitution ER paths is the number of sectors multiplied by the number of audio sources in the audio scene.
[0070] The following description focuses on the environmental dependent gain factor X, as this factor is currently a fixed value dependent on the environment. In the current MPEG standards, X is set equal to 1 for outdoor scenarios. This value corresponds to the assumption that the reflection energy for outdoor environments is not dominant, i.e., that reflection energy cannot exceed direct energy.
[0071] On the other hand, for indoor scenarios X is set equal to 2. This corresponds to the assumption that reflections energy is dominant for indoor environments, i.e., reflection energy can exceed direct energy. In particular, for example 2 = V4 may be a good representation for indoor environments with on average of 4 ERs from 4 side walls.
[0072] Limiting the gain factor X to two distinct values, which also have to be preset for the audio scene, may lead to problems, especially for scenes with complex environments, e.g., for audio scenes with a mixture of indoor and outdoor environments. Therefore, it would be beneficial if the gain factor X would be a continues variable, and would automatically change when the environment changes, i.e., when the listener position, the audio source position and / or the environment itself changes, e.g., a wall is destroyed.
[0073] In the following, determination of a continues gain factor X will be described in relation to a voxel-based representation of the audio environment. It is however noted that the voxel-based representation only serves as an example for understanding the invention and should not be construed as a limitation. Other types of audio scene representations may also be used for determining the gain factor X as long as a value indicative of the quantity of ER trajectories is determined and the determination is based on ray casting.
[0074] Voxels for audio rendering are relevant for media environments implemented in both hardware and software, such as video game and / or VR, AR, MR and XR environments. The following describe and define some concepts relating to voxels for audio rendering:
[0075] What is voxel for audio rendering?
[0076] A Voxel is a space volume with acoustic properties or audio rendering instructions assigned to it.
[0077] What is voxel size for audio rendering?
[0078] Voxel size is encoder configuration parameter, and it can be (manually or automatically) selected according to a scene geometry level of details (e.g., in the range of 10 cm - 1 m).
[0079] How voxels can be obtained?
[0080] Voxels for audio rendering can be obtained by: • voxelization (or conversion) of a mesh-based scene representation, and / or
[0081] • from scene representation used for scene generation (or even video rendering) (e.g., by down-sampling of voxels of smaller size).
[0082] How to represent voxel-based audio scenes?
[0083] Any voxel-based representation of an audio scene may contain an indication of voxels that are not transmission voxels (e.g., that are occluder voxels), i.e., voxels in which sound cannot propagate or cannot freely propagate - a representation of occluding geometries. This indication may relate to an indication of coordinates (e.g., center coordinates, corner coordinates) of the respective voxels. The coordinates of these voxels may be represented by grid indices, for example.
[0084] Additionally, the voxel-based representation may include indications of material properties of the voxels that are not transmission voxels, such as absorption coefficients, reflection coefficients, etc.. In addition to the occluder voxels, the voxel-based representation may also indicate transmission voxels (e.g., air voxels), i.e., voxels in which sound can propagate a representation of sound propagation media. Accordingly, some implementations of voxel-based representations of audio scenes may include, for each voxel in a predefined section of space (e.g., within boundaries enclosing the audio scene), and indication of a respective material property.
[0085] For a voxel-based representation a heuristic approach for estimating ER trajectories has been disclosed in WO2023 / 227544 Al, which is incorporated herein by reference. The determination of the gain factor X may however not be limited to this specific approach of ER trajectory estimation. Other methods for determining a value indicative of the quantity of ER trajectory in a voxel-based representation may be equally used for the gain factor X, as long as they are based on ray casting. The heuristic approach is understood as a preferred embodiment, as it may enable a very computationally efficient method of determining the gain factor X.
[0086] The heuristic approach is based on finding geometrically valid reflection trajectories between audio source and listener sufficient for generating the perceived 1storder ER sound effect by performing several low complexity steps based on the location of audio source and listener, and the voxel-based geometry representation.
[0087] Fig. 4 depicts the general idea of the heuristic approach. A position of an audio source 101 and a position of a listener 102 are known. Then, valid reflections trajectories from the audio source 101 to the listener 102 over a collision voxel 104 should be estimated without considering information of a reflective surface based on multiples of voxels, i.e., the heuristic approach works on a voxel-by- voxel basis and on grid indices representing the voxel positions.
[0088] The heuristic approach will now be explained in detail for a specific audio scene example depicted in Figs. 5 to 9. The disclosure should however not be construed as to be limited by this specific example. Moreover, while the example relates to a 2D case or shows a 2D projection only, it is understood that the approach according to the present disclosure is generally applicable to 3D audio scenes. Fig. 5 depicts an example 2D audio scene with occluding voxels (105 dotted), non-occluding voxels (plain), audio source 101, and listener 102. The locations of the audio source 101 and listener 102 are marked with S and L, respectively. The example is in 2D for illustration purposes only. The extension of the algorithm to a 3D environment is straightforward.
[0089] The voxel-based representation of the audio scene, information on the listener location, and the audio source 101 location are an input to the ER trajectory estimation method. In other words, the voxel-based representation of the audio scene, information on the listener location, and the audio source location may be received. Therefore, the voxel-based representation of the audio scene, information on the listener location, and the audio source location may be included in a bitstream that is received by an audio rendering device. Alternatively, the voxel-based representation of the audio scene, information on the listener location, and the audio source location may already be included in the memory of or locally provided to the audio rendering device. Yet alternatively, the voxel-based representation of the audio scene, information on the listener location, and the audio source location may be determined in the ER trajectory estimation method.
[0090] To find reflection trajectories between the audio source 101 and the listener 102, a ray direction pattern is applied to points 103 in the audio scene. In the example of Fig. 5, five equally spaced points on the line connecting the audio source 101 and the listener 102 are depicted. The depicted example, however, should not be construed to limit the positioning and number of the points. A different number of points and different positions of these points can be employed for the method. The ray direction pattern may be determined beforehand. Further, the ray direction pattern may define a predefined number of rays and predefined (corresponding) directions of rays from an origin.
[0091] In a next step, coordinates of points need to be defined (e.g., determined or calculated) for application of the ray direction pattern to the respective points 103. The number (e.g., count, cardinality) of points may be determined. The number of points may be fixed. In some implementations, the number of points may alternatively or additionally depend on the chosen ray direction pattern.
[0092] Further, it has been found that locating the points on a line between the audio source 101 and the listener 102 improves the quality and efficiency of ER trajectory estimation. To this end, the points 103 (location of the points 103) may be determined based on the quantity of the one or more points. Additionally, the location of the points may be determined based on the line connecting the audio source location and the listener location (e.g., to be arranged on said line). More particularly, the one or more points 103 may be determined such that the line connecting the audio source location and the listener location is split into N-l equal segments, for example. Here, N is the quantity of the one or more points and may be larger than or equal to 2 in this case.
[0093] Notably, for the quantity of points being chosen as 1, the single point may correspond to the audio source location. For the quantity of points chosen as 2, the two points may correspond to the audio source location and the listener location, respectively.
[0094] Fig. 6 depicts an example where a ray direction pattern with 8 rays is applied to a point 103 located at the location of the audio source 101. The rays are depicted as dashed lines.
[0095] In a next step, a set of collision voxels is determined based on the plurality of rays and the voxelbased representation of the audio scene. In particular, the set of collision voxels may be determined by searching for intersections between the rays and any occluder voxel 105 (dotted) in the audio scene. Occluder voxels 105 may represent an acoustically reflective surface. In other words, occluder voxels 105 may represent any material in the voxel-based representation of the audio scene other than air or other representations of sound propagation media. Then, for each ray, an occluder voxel 105 containing an intersection closest to the origin of the respective ray may be determined as a collision voxel 104. In other words, a collision voxel 104 may be defined as the first occluding voxel hit by the respective ray. This step ensures that only occluding voxels are selected which may represent a reflective surface. In Fig. 6 all collision voxels 104 for the rays originating at the audio source location are marked with a bullet point at the end of the rays. Notably, in this example the lower right ray has no intersection with any occluding voxels and therefore is not depicted in Fig. 6 and not further considered in the algorithm.
[0096] In one embodiment, the quantity of collision voxels 104 may be determined as an indication of the quantity of ER trajectories. In particular, for each point 103 on the line connecting the listener location and the source location, the collision voxels may be determined by applying the respective ray pattern. Then the total number of collision voxels found by this method may be used as an indication for the quantity of ER trajectories. This approach is motivated by the fact that the total number of collision voxels may be a good indication for the quantity of (valid) ER trajectories, at least for some scenarios. For example, in an outdoor scenario with no reflective surfaces, the number of collision voxels will be zero, as will be the number of ER trajectories. Further, in a scenario with a source and a listener in a room with a square base and no obstacles, each collision voxel will lead to a valid early reflection trajectory. In other, more complex environments, the quantity of collision voxels may still provide a good estimation on the number of valid ER trajectories.
[0097] The quantity of collision voxels may then be used to determine the continues gain factor X, as will be explained with reference to Figs. 10 to 14.
[0098] Alternatively, the algorithm may be continued to determine the quantity of valid ER trajectories.
[0099] In a next step, for each collision voxel 104, a preceding voxel 106 may be determined. The preceding voxel 106 may be a voxel containing an intersection with the ray, preceding the respective collision voxel in the direction of the ray. In other words, the preceding voxel 106 may be the last non-occluding voxel before the respective collision voxel 104 in the direction of the ray.
[0100] In a next step, for each preceding voxel 106 determined for collision voxel 104, a path may be determined to connect the audio source 101 and the listener 102 via the respective preceding voxel 106. The path may comprise a straight line connecting the audio source location to the preceding voxel 106 and a straight line connecting the same preceding voxel 106 to the listener location.
[0101] Fig. 7 shows the example depicted in Fig. 6 with the determined paths between audio source 101 and listener 102.
[0102] In a final step, it is determined whether the paths from audio source 101 to listener 102 are geometrically valid. The following geometric validity test may consider the paths relating to all preceding voxels 106. As the paths determined in the previous step may be solely defined by straight lines between audio source 101, the preceding voxel 106, and listener 102, lines may traverse (e.g., intersect or graze) occluding voxels. In reality, such a reflection trajectory would not be possible in the sense that it would not permit propagation of sound. Therefore, paths comprising lines traversing occluding voxels may be determined to be geometrically invalid. To find intersections with occluding voxels, a line-grid intersection algorithm may be applied to the lines connecting audio source 101, preceding voxel 106, and listener 102. As an example, the Fast traversal algorithm for ray tracing (cf. Amanatides, J. and A. Woo, A Fast Voxel Traversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987. 87.) may be used.
[0103] To achieve computational savings, the line-of-sight check between the audio source location and the preceding voxel 106 may not need to be performed when the corresponding collision voxel 104 is obtained from a respective ray originating from the audio source location. This is due to the fact that the collision voxel 104 is the first occluder voxel 105 to be hit by the respective ray, ensuring the line-of-sight fulfillment. Similarly, the line-of-sight check between the listener location and the preceding voxel 106 may not need be performed when the corresponding collision voxel 104 is obtained from a respective ray originating from the listener location.
[0104] Fig. 8 depicts the determined geometrically valid path as a solid line. Notably, only two paths of the previously found 7 paths are determined as geometrically valid in this example.
[0105] As previously stated, the process is repeated for each point 103 on the line between audio source 101 and listener 102.
[0106] Fig. 9 depicts the final result of the algorithm for the example of NP= 7 points 103 and 8 rays per point 103. From the 7 * 8 = 56 possible paths, only 11 are determined as geometrically valid. These paths may then be considered as (valid) ER trajectories 107.
[0107] In the following, either the quantity of collision voxels 104 or the quantity of valid ER trajectories may be considered as an indication for the quantity of valid ER trajectories. The indication for the quantity of valid ER trajectories may then be used to determine gain factor X.
[0108] Returning back to the determination of energy levels of the ERs, Figs. 10A to 10B schematically illustrate a voxel-based representation of an outdoor (Fig. 10A) and an indoor environment (Fig. 10B). In contrast to Fig. 9, only ER trajectories 107 (solid and dashed lines) are shown, which correspond to sector k (roughly a quarter circular segment). As previously stated, for modeling the energy levels of the ERs, for each audio source in sector k, the corresponding image source gains and positions are all combined into a single substitution ER path 108 (dotted line).
[0109] Therefore, the ERs are modeled by one image audio source The resulting levels of early reflections energy for the indoor and outdoor case are the same in the sector k, when the ER modelling method does not account for the environment type. As previously stated, a fixed gain factor X is applied to account for the two different environment types.
[0110] To overcome the shortcomings of a fixed gain factor X, a continues gain factor X is proposed that is determined based on the indication for the quantity of valid ER trajectories, i.e., the quantity of valid ER trajectories itself or the quantity of collision voxels. To determine X, a parameter is defined as wherein Nhitsis either the quantity of valid ER trajectories or the quantity of collision voxels and Nraysis the total quantity of rays that have been casted.
[0111] Parameter famay indicate whether an environment represents an indoor environment, depending on how Nhitsis determined. To illustrate this, Fig. 11 depicts a transition of a listener 102 from a first room (1) to second room (2), wherein the source 101 is positioned in the second room. When Nhiteisthe quantity of collision voxels, Nhitswill be equal to Nraysirrespective of the position of the listener 102 in first room and the second room. In an outdoor scenario with almost no reflective services, Nhitsmay be zero or close to zero. Therefore, fabeing one or close to one may indicate that the listener 102 is in an indoor environment.
[0112] When Nhiteisdefined as the quantity of valid ER trajectories, Nhitswill be close to Nrays, when the listener 102 and the source 101 are located in the same room (2). When the listener is however positioned in the first room (1), only few ERs will go through the doorway connecting the first room and the second room. Therefore, in this case famay be close to zero, despite the listener 102 and the source 101 being indoor. Therefore, famay not indicate whether the listener 102 is located indoor.
[0113] Further, determining Nraysmay depend on the particular method for finding valid ERs. For the method of determining early reflection trajectories as presented in Figs. 5 to 9, Nrays= NRNP, wherein NRis the quantity of rays per point 103, and NPis the quantity of points 103. In the example of Fig. 9, NR= 8 and NP= 7, while the number of valid ER trajectories 106 is 11.
[0114] Therefore, for a total quantity of rays Nrays= 56 cast, only Nhits= 11 valid ER trajectories are estimated. f a, can therefore be determined as
[0115] In a preferred embodiment, NRmay be equal to 26, as 26 may be the maximum number of rays that can be defined by integer arithmetic. The corresponding ray pattern is depicted in Fig. 12.
[0116] The gain factor X is then determined as (4) wherein a and b are positive scalars. The values a and b may be set by the content creators (e.g., extracted from the bitstream, controlled by an application, manually or automatically set at the renderer interface) to reflect the artistic intent. In other words, through a specific setting of a and b a specific artistic effect may be achieved, e.g., in the case of Fig. 11, a large value for a may increase the perceptibility of the source 101 if the listener 102 is in the first room.
[0117] In a preferred embodiment, a = 2 and b = 1. Therefore, the gain factor can be expressed as
[0118] By determining the gain factor in this way, the environment factor for the ER energy level can be adapted on the fly, even for complicated environments, especially for a mix of an indoor and outdoor environment. Experiments have shown that a gain factor calculated in this way leads to a more realistic impression of ERs, i.e., the audio experienced by the user matches the environment as observed by the user. Further, by using values already known when estimating early reflection trajectories, i.e., NR, NP, and Nbits, the method of calculating X does not lead to any relevant increase in computational complexity. In particular, only one multiplication and one addition are needed to calculate gain factor X for one particular audio source and listener position. Gain factor X is then used to determine the ER gain gRRper sector k, as has been previously described. The reflection gain is then applied to the image source audio signal, i.e. to the audio signal of the image source IRR. In a final step of the rendering process, image source audio signals in all sectors may then be mixed with other signals affected or generated by other processing / modeling stages in the audio processing chain, such as distance compensation, occlusion, diffraction and reverberation stages, to output the final rendered audio.
[0119] In line with the above, a method 1300 is provided for determining a gain factor for the ER energy level as depicted in the flowchart of Fig. 13. The method may be implemented in a decoder or renderer or in both decoder and Tenderer in an AR / VR / MR / XR environment. The decoder and / or Tenderer may be implemented in the network / cloud or a processing device such as a mobile device and an AR / VR / MR / XR google / lens or distributed in both the network / cloud and a processing device. In addition to the following method steps, method 1300 may optionally include all variations described above with respect to the aforementioned gain factor determination discussed in connection with Fig. 3 to Fig. 12.
[0120] In step 1302, a representation of the three-dimensional audio scene, information on a listener location of a listener 102 in the three-dimensional audio scene, and information on an audio source location of the audio source 101 in the three-dimensional audio scene are obtained either separately and / or in various combinations. Each of the representation of the three-dimensional audio scene, information on the listener location, and information on the audio source location may be received and / or predetermined (i.e., previously calculated, stored, and then read from memory).
[0121] In step 1304, a predefined quantity of rays for a ray direction pattern and a quantity of points for applying the ray direction pattern are obtained. The predefined quantity of rays may be received and / or predetermined (i.e., previously calculated, stored, and then read from memory).
[0122] In step 1306, an indication for a quantity of valid early reflection trajectories for the audio source and the listener location is determined based on the representation of the three-dimensional audio scene, the ray direction pattern, the quantity of points, the listener location and the audio source location.
[0123] In step 1308, a gain factor is determined based on the indication for the quantity of valid early reflection trajectories and a total quantity of rays, wherein the total quantity of rays is determined based on the predefined quantity of rays and the quantity of points. The total quantity of rays may be determined by multiplying the predefined quantity of rays with the quantity of points.
[0124] In optional step 1310, the gain factor may be output for rendering of the three-dimensional audio scene. Outputting the gain factor may comprise determining of an early reflection energy level based on the gain factor and outputting the early reflection energy level for further processing. Further processing may comprise mixing the early reflections modeled with the early reflection energy with other elements of audio rendering, such as distance compensation, occlusion, diffraction and reverberation stages, to provide a final rendering of the audio scene.
[0125] Further, method 1400 of Fig. 14 is provided for determining the indication for the quantity of valid early reflections. Method 1400 may be one particular example for step 1306, when the representation of the three-dimensional audio scene is voxel-based. However, other methods may be used for determining an indication for the quantity of valid early reflections 10, and other types of audio scene representations, i.e., other than voxel-based, may be used, as long as total quantity of cast rays and an indication for a quantity of valid early reflections determined from the cast rays can be obtained. The indication for the quantity of valid early reflections may be the indication for a quantity of valid early reflections or a suitable value indicating the quantity of valid early reflections. The suitable value may depend on the environment and the method used for determining the early reflections. For example, in the case of a voxel-based environment, a quantity of collision voxels determined when casting rays of a ray direction pattern may provide a good indication for the quantity of valid early reflections, at least for some environments.
[0126] Method 1400 may be an exemplary implementation of method step 1306.
[0127] In method 1400, the indication for the quantity of valid early reflection trajectories may be the quantity of valid early reflection trajectories or a quantity of collision voxels.
[0128] In step S1402, a ray direction pattern is applied to one or more points 103 in a proximity to a region connecting the audio source location and the listener location to obtain, for each of the one or more points 103, a plurality of rays originating at the respective point(s) 103. The ray direction pattern may be predetermined and / or retrieved from memory storage. The proximity to the region connecting the audio source location and the listener location may be a line connecting the audio source location and the listener location.
[0129] In step S1404, a set of collision voxels is determined based on the plurality of rays determined at step 1402 and a voxel-based representation of the three-dimensional audio scene. Determining the set of collision voxels may comprise determining one or more intersections between each ray of the plurality of rays and the occluder voxels 105. Occluder voxels 105 may represent an acoustically reflective surface. Further, for each ray, an occluder voxel 105 containing an intersection closest to the origin of the respective ray may be determined as a collision voxel 104. The collision voxels 104 determined in this manner may form the set of collision voxels. The voxel-based representation may be received and / or predetermined (i.e., previously calculated, stored, and then read from memory).
[0130] In step S1406, collision voxels in the set of collision voxels are counted to determine the quantity of collision voxels if the indication for the quantity of valid early reflection trajectories is the quantity of collision voxels. Therefore, the quantity of the collision voxel may be the output of the method and may be used to determine the gain factor.
[0131] Alternatively, if the indication for the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories, the method skips step S1406 and continues with step S1408.
[0132] In step S1408, early reflection trajectories are determined based on the set of collision voxels, the listener location, the audio source location, and a geometrical validity test. Determining early reflection trajectories may comprise to determine, for each collision voxel in the set of collision voxels, a preceding voxel 106. The preceding voxel 106 may be a voxel containing an intersection with the respective ray, preceding the collision voxel 104 in the direction of the respective ray. Then, a path connecting the listener location and the audio source location via a preceding voxel 106 may be determined as an early reflection trajectory if the path is not obstructed by occluding voxels.
[0133] In step S1410, the early reflection trajectories are counted to obtain the quantity of valid early reflection trajectories. In this case, the quantity of valid early reflection trajectories may be the output of the method and may be used to determine the gain factor.
[0134] In addition, an embodiment of this application further provides an apparatus. As shown in Fig.15, the apparatus 1500 includes a processor 1501 and memory 1502. The memory 1502 is configured to store program code. The processor 1501 is configured to run instructions in the program code, so that the apparatus 1500 performs the determination of a gain factor for modelling early reflections of an audio source in any one of the above implementations. The processor 1501 may also receive, among others, suitable input data (e.g., audio source location, listener location, etc.), depending on various use cases and / or implementations. The processor 1501 may be adapted to carry out the methods / techniques (e.g., methods 1300 and 1400 as illustrated above with reference to Figs. 13and 14, respectively) described throughout the present disclosure and to generate correspondingly output data (e.g., the gain factor etc.), depending on various use cases and / or implementations.
[0135] Aspects of the methods described herein may be implemented in an appropriate computer-based sound processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0136] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer- readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0137] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0138] Exemplary Details
[0139] 1 Motivation
[0140] 1.1 Acoustic environments for early reflection modelling
[0141] The CE on ER modelling for voxels was presented at the 12thWG6 meeting. This CE technology was tested and adopted to the MPEG-I Immersive Audio specification. The CE algorithm calculates the ER gains g|Rfor the outdoor environment type as:
[0142] The following change was introduced in WD5 based on the indoor environment type assumption:
[0143] 1.2 Issue with single and fixed ER gain specification
[0144] The hardcoded gain and / or single tuneable parameter determines the ERs energy level for the whole scene. This level setting is made at the encode side for the initial scene geometry. If the scene contains an indoor, outdoor, and mixed environments simultaneously (or the scene geometry is modifiable), the audio Tenderer cannot account for its type and will apply the same pre-determined ER energy tuning parameter regardless of the user / source position(s).
[0145] Fig. 16 provides an example of a scene with two environment types. If a content creator needs to realize two different levels of ERs for the indoor and outdoor environments, one has to either to: split the scene into several sub- scenes with different ER gains or
[0146] • pre-determine several regions for available environment types, assign the corresponding different ER gains and switch between them by using L2 updates based on the user / source position(s).
[0147] These approaches are not realizable in the RM6 without extra code modifications and / or result in perceptual rendering audio output discontinuity if the listener moves from the region of one environment type to another one (e.g., continuous transition between indoor and outdoor environments), see Figure 16.
[0148] 1.3 Solution introducing environment type dependent ER gain calculation
[0149] 1.3.1 Proposal
[0150] The CE solution proposes to: 1) introduce two tuneable ER energy level modelling parameters for the indoor and outdoor environment types (i.e., gERIndoorJuningand gEROutdoor-tuning) specified by a content creator via the EIF data
[0151] 2) use the existing voxel ER modelling functionality to estimate the environment type; namely to introduce the indoor environment characterization factor findoordescribing “indoorness” of the space between listener and audio source(s). The function findooris determined by the ratio between the number of ER ray-voxel hits Nhitsto the total quantity of rays that have been casted Nraysas:
[0152] 3) apply the linear interpolation function for the indoor and outdoor gain values using the function findOOre[0, 1]; (findoor is equal to 1 for indoor and 0 of outdoor environment type):
[0153] 4) use this environment type dependent function fERfor the ER level adjustments (instead of the hard-coded constant scalar value of “2”):
[0154] 1.3.2 Activation and control
[0155] The activation and control of the CE functionality is performed by the encoder (via the corresponding EIF data) of two ERs tuning gains:
[0156] • EarlyTuninglndoorGainDb - for gERIndoor tuning(indoor environment type)
[0157] • EarlyTuningOutdoorGainDb - for gEROutdoor tuning(outdoor environment type)
[0158] The default value for these content creator ER tuning variables is OdB.
[0159] The gERIndoor tuningand gEROutdoorJuningvalue range is [-119, 12] (dB). If EarlyTuninglndoorGainDb (EarlyTuningOutdoorGainDb) < -119 (dB) => gERIndoor tuning=0 (gERoutdoor_tuning0)- Note: if both values are equal to each other, the interpolation process can be simplified:
[0160] EarlyTuninglndoorGainDb == EarlyTuningOutdoorGainDb => fER=
[0161] SER]nCi00r_tuning
[0162] Note: if both values are equal to +6dB the WG6 output can be reproduced: EarlyTuninglndoorGainDb == EarlyTuningOutdoorGainDb == 6 => fER= 2
[0163] 1.3.3 Advantages
[0164] This CE solution eliminates limitation of the MPEG-I Audio rendering technology and allows to support any environment type (i.e., from ideal indoor to ideal outdoor and any mixed environment in between). During the audio rendering process, the CE algorithm will automatically characterize the current environment type (taking in account the listener and audio source location) and adjusts the corresponding ER energy level according to the content creator specified modification range:
[0165] 1.4 Specification updates 1.4.1 EIF(v8) amendments
[0166] • Replace in “3.11 Acoustic Tuning Parameters” : by
[0167] • Add in “3.11 Acoustic Tuning Parameters” :
[0168] • Replace in “3.11 Acoustic Tuning Parameters” : < / VoxDirectOcclusionControl>
[0169] <VoxDiffractionControl value="0.75" / >
[0170] < V oxEarlyReflectionControl>
[0171] <NumOrigins value="5" / > <NumClusters value="8" / >
[0172] <CombineReflectionsPerCluster / >
[0173] < / V oxEarlyReflectionControl>
[0174] < V oxEarlyReflectionControl> <NumOrigins value="5" / >
[0175] <NumClusters value="8" / >
[0176] <CombineReflectionsPerCluster / >
[0177] <EarlyTuningIndoorGainDb value="3.0" / >
[0178] <EarlyTuningOutdoorGainDb value="-6.0" / > < / VoxEarlyReflectionControl>
[0179] L4.2 WD (v6) amendments
[0180] L4.2.1 Bitstream
[0181] • Add in Table 50 — Syntax of voxReflectionConfigurationP ammeter s( / •
[0182] Table 50 — Syntax of voxReflectionConfigurationParameters()
[0183] Table 50 — Syntax of voxReflectionConfigurationParameters()
[0184] • Add in “6.3.2.4 Voxel pay load data structure”: voxReflectionTuninglndoorOutdoorTypeFlag
[0185] Flag indicating the presence of additional gains (dB) to be used for ER rendering for indoor / outdoor environment types. voxReflectionEarlyTuninglndoorGainDb
[0186] This value is the additional tuning gain (dB) to be used for ER rendering for the indoor environment type.
[0187] It shall range between -119 to +12 (dB). Default value is 0 (dB). voxReflectionEarlyTuningOutdoorGainDb
[0188] This value is the additional tuning gain (dB) to be used for ER rendering for the outdoor environment type.
[0189] It shall range between -119 to +12 (dB). Default value is 0 (dB).
[0190] 1.4.2.2 Processing • Replace in “ 1.3.1.2.3 Processing steps at the listener or audio source updates”:
[0191] • Add at the end of “6.6.11.2 Data elements and variables”: fERER gain adjustment function for indoor / outdoor environment types findoorThe indoor environment type characterization function gERmdoor tuning Indoor environment type ER tuning gain obtained from voxReflectionEarlyTuninglndoorGainDb
[0192] §ERoutdoor tuning Outdoor environment type ER tuning gain obtained from voxReflectionEarlyTuningOutdoorGainDb
[0193] NhitsNumber of early reflection ray hits
[0194] NraysTotal quantity of rays that have been casted
[0195] • Add after “1.3.1.2.3 Processing steps at the listener or audio source updates”:
[0196] 1.3.1.2.4 Early reflection gain adjustment to environment type
[0197] The ER gain adjustment function fERis obtained as: where the indoor environment characterization function findoor isobtained by the ratio between the number of ER ray hits Nhitsand the number of all considered rays Nrays:
[0198] Nrays = NRNp (X)
[0199] Nhits = number{Chit} (X)
[0200] Conclusion
[0201] The proposed CE technology:
[0202] • removes the ER modelling technology constraint on a specific environment type
[0203] • offers better control over ER energy level specification to a content creator
[0204] • provides smooth transition between indoor and outdoor environments for the voxel-based ER audio rendering Interpretation
[0205] A computing device implementing the techniques described above can have the following example architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the example architecture includes one or more processors (e.g., dual-core Intel® Xeon® Processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display) and one or more computer-readable mediums (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can exchange communications and data over one or more communication channels (e.g., buses), which can utilize various hardware and software for facilitating the transfer of data and control signals between components.
[0206] The term “computer-readable medium” refers to a medium that participates in providing instructions to processor for execution, including without limitation, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory) and transmission media. Transmission media includes, without limitation, coaxial cables, copper wire and fiber optics.
[0207] Computer-readable medium can further include operating system (e.g., a Linux® operating system), network communication module, audio interface manager, audio processing manager and live content distributor. Operating system can be multi-user, multiprocessing, multitasking, multithreading, real time, etc. Operating system performs basic tasks, including but not limited to: recognizing input from and providing output to network interfaces and / or devices; keeping track and managing files and directories on computer-readable mediums (e.g., memory or a storage device); controlling peripheral devices; and managing traffic on the one or more communication channels. Network communications module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols, such as TCP / IP, HTTP, etc.).
[0208] Architecture can be implemented in a parallel processing or peer-to-peer infrastructure or on a single device with one or more processors. Software can include multiple software components or can be a single body of code.
[0209] The described features can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, a browser-based web application, or other unit suitable for use in a computing environment.
[0210] Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors or cores, of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer will also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, ASICs (application- specific integrated circuits).
[0211] To provide for interaction with a user, the features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or a retina display device for displaying information to the user. The computer can have a touch surface input device (e.g., a touch screen) or a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user.
[0212] The features can be implemented in a computer system that includes a back-end component, such as a data server, or that includes a middleware component, such as an application server or an Internet server, or that includes a front-end component, such as a client computer having a graphical user interface or an Internet browser, or any combination of them. The components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, e.g., a LAN, a WAN, and the computers and networks forming the Internet.
[0213] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[0214] A system of one or more computers can be configured to perform particular actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0215] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0216] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0217] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the present invention discussions utilizing terms such as “processing”, “computing”, “calculating”, “determining”, “analyzing” or the like, refer to the action and / or processes of a computer or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities.
[0218] Reference throughout this invention to “one example embodiment”, “some example embodiments” or “an example embodiment” means that a particular feature, structure or characteristic described in connection with the example embodiment is included in at least one example embodiment of the present invention. Thus, appearances of the phrases “in one example embodiment”, “in some example embodiments” or “in an example embodiment” in various places throughout this invention are not necessarily all referring to the same example embodiment. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this invention, in one or more example embodiments.
[0219] As used herein, unless otherwise specified the use of the ordinal adjectives “first”, “second”, “third”, etc., to describe a common object, merely indicate that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking, or in any other manner.
[0220] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted”, “connected”, “supported”, and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.
[0221] In the claims below and the description herein, any one of the terms comprising, comprised of or which comprises is an open term that means including at least the elements / features that follow, but not excluding others. Thus, the term comprising, when used in the claims, should not be interpreted as being limitative to the means or elements or steps listed thereafter. For example, the scope of the expression a device comprising A and B should not be limited to devices consisting only of elements A and B. Any one of the terms including or which includes or that includes as used herein is also an open term that also means including at least the elements / features that follow the term, but not excluding others. Thus, including is synonymous with and means comprising.
[0222] It should be appreciated that in the above description of example embodiments of the present invention, various features of the present invention are sometimes grouped together in a single example embodiment, Fig., or description thereof for the purpose of streamlining the present invention and aiding in the understanding of one or more of the various inventive aspects. This method of invention, however, is not to be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed example embodiment. Thus, the claims following the Description are hereby expressly incorporated into this Description, with each claim standing on its own as a separate example embodiment of this invention.
[0223] Furthermore, while some example embodiments described herein include some but not other features included in other example embodiments, combinations of features of different example embodiments are meant to be within the scope of the present invention, and form different example embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed example embodiments can be used in any combination.
[0224] In the description provided herein, numerous specific details are set forth. However, it is understood that example embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0225] Thus, while there has been described what are believed to be the best modes of the present invention, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the present invention, and it is intended to claim all such changes and modifications as fall within the scope of the present invention. For example, any formulas given above are merely representative of procedures that may be used.
[0226] Functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure.
[0227] Enumerated Example Embodiments
[0228] Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0229] EEE 1. A method of determining a gain factor for modelling early reflections of an audio source in a three-dimensional audio scene, wherein the gain factor indicates whether the three- dimensional audio scene is indoor, outdoor or a mixture of indoor and outdoor, the method comprising: obtaining a representation of the three-dimensional audio scene, information on a listener location of a listener in the three-dimensional audio scene, and information on an audio source location of the audio source in the three-dimensional audio scene; obtaining a predefined quantity of rays for a ray direction pattern and a quantity of points for applying the ray direction pattern in a proximity to a region connecting the audio source location and the listener location; determining an indication for a quantity of valid early reflection trajectories for the audio source and the listener based on the representation of the three-dimensional audio scene, the ray direction pattern, the quantity of points, the listener location and the audio source location; and determining the gain factor based on the indication for the quantity of valid early reflection trajectories and a total quantity of rays, wherein the total quantity of rays is determined based on the predefined quantity of rays and the quantity of points.
[0230] EEE 2. The method of EEE 1, further comprising: outputting the gain factor for rendering of the three-dimensional audio scene.
[0231] EEE 3. The method of EEE 1 or 2, wherein determining the gain factor based on the indication for the quantity of valid early reflection trajectories and the total quantity of rays comprises dividing the indication for the quantity of valid early reflection trajectories by the total quantity of rays. EEE 4. The method of EEE 3, wherein the gain factor is defined as wherein fais a function of the total quantity of rays and the indication for the quantity of valid early reflection trajectories, and wherein a and b are scalars for scaling the gain factor for an artistic intent.
[0232] EEE 5. The method of EEE 4, wherein a = 2, and b = 1, and the gain factor is defined as , wherein Nraysis the total quantity of rays and Nhitsis the indication for y the quantity of valid early reflection trajectories.
[0233] EEE 6. The method of any one of the preceding EEEs, wherein the total quantity of rays is defined as the quantity of rays multiplied by the quantity of points.
[0234] EEE 7. The method of any one of the preceding EEEs, wherein the gain factor is used to determine an early reflections gain.
[0235] EEE 8. The method of EEE 7, wherein the early reflections gain is determined for a sector used for modelling the early reflections.
[0236] EEE 9. The method of EEE 8, wherein the sector is determined by dividing a 360-degree angle around the listener position in equiangular portions.
[0237] EEE 10. The method of EEE 9, wherein a quantity of equiangular portions is indicative of a precision for the modelling of the early reflections.
[0238] EEE 11. The method of any one of EEEs 8 to 10, wherein the early reflections gain is directly proportional to the gain factor multiplied by a mean value of reflection attenuation gains for the audio source for the sector.
[0239] EEE 12. The method of EEE 11, wherein the mean value of reflection attenuation gains for the audio source for the sector is defined as wherein k is the index of the sector, gER^ is the reflection attenuation gain for early reflection trajectory i in sector k, nkis the total quantity of early reflection trajectories in sector k, and Nsis the maximum quantity of reflections per audio source for the reflection modelling.
[0240] EEE 13. The method of any one of EEEs 8 to 12, wherein the early reflections gain is determined for an image source in the sector.
[0241] EEE 14. The method of EEE 13, wherein the image source is a virtual audio source perceived by the listener due to the early reflections.
[0242] EEE 15. The method of EEE 13 or 14, wherein a location of the image source is the location of the audio source mirrored on a reflective surface in the three-dimensional audio scene.
[0243] EEE 16. The method of any one of the preceding EEEs, wherein the representation of the three-dimensional audio scene is a voxel-based representation.
[0244] EEE 17. The method of EEE 16, wherein the indication for the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories or a quantity of collision voxels and wherein determining the indication for the quantity of valid early reflection trajectories for the audio source and the listener comprises: applying the ray direction pattern to the points in the proximity to the region connecting the audio source location and the listener location to obtain, for each of the points, a plurality of rays originating at the respective point; determining a set of collision voxels based on the plurality of rays and the voxel-based representation of the three-dimensional audio scene; counting collision voxels in the set of collision voxels to determine the quantity of collision voxels if the indication for the quantity of valid early reflection trajectories is the quantity of collision voxels; or determining early reflection trajectories based on the set of collision voxels, the listener location, the audio source location and a geometrical validity test; and counting the early reflection trajectories to obtain the quantity of valid early reflection trajectories if the indication for the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories.
[0245] EEE 18. The method of EEE 17, further comprising: determining the points based on the quantity of points.
[0246] EEE 19. The method of any one of the preceding EEEs, wherein the ray direction pattern defines the predefined quantity of rays and predefined directions of rays from an origin.
[0247] EEE 20. The method of any one of the preceding EEEs, wherein the predefined quantity of rays is 6, 8, 12 or 26.
[0248] EEE 21. The method of any one of EEEs 16 to 18 or EEEs 19 to 20, when depending on EEE 16, wherein a voxel position in the three-dimensional audio scene is defined by grid indices and the predefined directions of rays comprise one or more of: horizontal and vertical directions of a grid index to neighboring grid indices; and diagonal directions of the grid index to the neighboring grid indices.
[0249] EEE 22. The method of EEE 18, wherein the region connecting the audio source location and the listener location is a line connecting the audio source location and the listener location and points are positioned on the line.
[0250] EEE 23. The method of EEE 22, wherein coordinates of the points on the line connecting the audio source location and the listener location are determined based on the quantity of points.
[0251] EEE 24. The method of EEE 23, wherein the points are determined to split the line connecting the audio source location and the listener location into N-l equal segments where N is the quantity of points and is larger than or equal to 2.
[0252] EEE 25. The method of EEE 18, wherein the quantity of points depends on available computational resources, an encoder preset, or a combination thereof. EEE 26. The method of EEE 17 or any one of EEEs 18 to 25 when depending on EEE 17, wherein each collision voxel in the set of collision voxels is an occluder voxel in the voxelbased representation of the three-dimensional audio scene.
[0253] EEE 27. The method of EEE 26, wherein the occluder voxel represents an acoustically reflective surface.
[0254] EEE 28. The method of EEE 26, wherein the occluder voxel represents any material in the voxel-based representation of the three-dimensional audio scene other than sound transmission media.
[0255] EEE 29. The method of any one of EEEs 26 to 28, wherein determining the set of collision voxels based on the plurality of rays and the voxel-based representation of the three- dimensional audio scene comprises: determining one or more intersections between each ray of the plurality of rays and the occluder voxels; and for each ray, determining an occluder voxel containing an intersection closest to the origin of the respective ray as a collision voxel in the set of collision voxels.
[0256] EEE 30. The method of EEE 17 or any one of EEEs 18 to 29 when depending on EEE 17, wherein determining the early reflection trajectories comprises: for each collision voxel in the set of collision voxels, determining a preceding voxel of the respective collision voxel, wherein the preceding voxel is a voxel containing an intersection with the respective ray, preceding the respective collision voxel in the direction of the respective ray; for each preceding voxel, determining a path connecting the listener location and the audio source location via the preceding voxel; and if the path can produce a geometrically valid representation of a first-order reflection, determining the path as an early reflection trajectory.
[0257] EEE 31. The method of EEE 30, wherein determining whether the path can produce the geometrically valid representation of the first-order reflection comprises: determining that the path can produce a geometrically valid representation of a first-order reflection if the path does not contain an intersection with an occluder voxel.
[0258] EEE 32. The method of EEE 30 or 31, wherein the path comprises a straight line connecting the audio source location to the preceding voxel and a straight line connecting the same preceding voxel to the listener location.
[0259] EEE 33. The method of EEE 2 or any one of EEEs 3 to 32 when depending on EEE 2, wherein the rendering is to be performed by a virtual reality, VR, augmented reality, AR, mixed reality, MR, and / or extended reality, XR device.
[0260] EEE 34. The method of any one of the preceding EEEs, wherein the valid early reflection trajectories represent 1storder trajectories.
[0261] EEE 35. The method of EEE 34, wherein the 1storder trajectories are reflection trajectories with a single reflection between the audio source location and the listener location.
[0262] EEE 36. The method of any one of the preceding EEEs, wherein the method is performed by a decoder or Tenderer.
[0263] EEE 37. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of EEEs 1 to 36.
[0264] EEE 38. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEEs 1 to 36.
[0265] EEE 39. A computer-readable storage medium storing the program according to EEE 38.
[0266] EEE 40. A method of estimating early reflection gains of an audio source in a three- dimensional audio scene, the method comprising: obtaining a voxel-based representation of the three-dimensional audio scene; determining data related to an indoor parameter for the three-dimensional audio scene, wherein the indoor parameter is based on a ratio between the number of the 1st order reflection trajectories Nhitsto the number of all considered rays Nrays; determining early reflection gains for the voxel-based representation of the three- dimensional audio scene, based on the indoor parameter, where the early reflection gains are based on a function X; outputting the early reflection gains adjustment to an indoor-outdoor environment type.
[0267] EEE 41. The method of EEE 40, wherein the early reflection gains are determined based on: where the function X represents an interpolation between the early reflection gains Xindoorand X0utdoorcontrolled by the variables findoorand foutdoor characterizing early reflection environment.
[0268] EEE 42. The method of EEE 41, wherein the variable foutdoor is alinear function of the variable findoordetermined based on: f10utdoor — • 1 > f Undoor
[0269] EEE 43. The method of EEE 41, wherein the function X is determined based on:
[0270] X ^Indoor ^Indoor 4" Xoutcjoorfoutdoor
[0271] EEE 44. The method of EEE 43, wherein the early reflection gains are set to:
[0272] • v^Indoor >_ 9 • y''Outdoor -— 1
[0273] EEE 45. The method of EEE 41, wherein the “indoor” parameter variable fjndoor describing early reflection environment is determined based on: EEE 46. The method of EEE 45, wherein the early reflection gains are determined using one constant integer value addition operation and one constant float value (1 / Nrays) multiplication operation: EEE 47. The method of EEE 46, wherein the number of the 1st order reflection trajectories Nhitsand the number of all considered rays Nraysare determined based on environment characteristics estimated along the line segment connecting the audio source and listener positions. EEE 48. The method of EEE 46, wherein the environment characteristics estimation based on the constant number of the emitted rays from each origin point NRand number origin points NPobtained by equidistantly dividing the line segment connecting the audio source and listener positions: EEE 49. A non-transitory computer program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEEs 40-48.
[0274] EEE 50. An apparatus configured to perform the method of any one of EEEs 40-48.
Claims
CLAIMS1. A method of determining a gain factor for modelling early reflections of an audio source in a three-dimensional audio scene, wherein the gain factor indicates whether the three- dimensional audio scene is indoor, outdoor or a mixture of indoor and outdoor, the method comprising: obtaining a representation of the three-dimensional audio scene, information on a listener location of a listener in the three-dimensional audio scene, and information on an audio source location of the audio source in the three-dimensional audio scene; obtaining a predefined quantity of rays for a ray direction pattern and a quantity of points for applying the ray direction pattern in a proximity to a region connecting the audio source location and the listener location; determining an indication for a quantity of valid early reflection trajectories for the audio source and the listener based on the representation of the three-dimensional audio scene, the ray direction pattern, the quantity of points, the listener location and the audio source location, wherein a valid early reflection trajectory represents a path from the audio source location to the listener location that can produce a geometrically valid representation of a first-order reflection of the audio source and wherein the indication for the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories or a value indicative of the quantity of valid early reflection trajectories; determining the gain factor based on the indication for the quantity of valid early reflection trajectories and a total quantity of rays, wherein the total quantity of rays is determined based on the predefined quantity of rays and the quantity of points; and outputting the gain factor for rendering of the three-dimensional audio scene.
2. The method of claim 1, wherein determining the gain factor based on the indication for the quantity of valid early reflection trajectories and the total quantity of rays comprises dividing the indication for the quantity of valid early reflection trajectories by the total quantity of rays.
3. The method of claim 2, wherein the gain factor is defined asX = a fa+ b fb, fa= 1 - fb,wherein fais a function of the total quantity of rays and the indication for the quantity of valid early reflection trajectories, and wherein a and b are scalars for scaling the gain factor for an artistic intent.
4. The method of claim 3, wherein a = 2, and b = 1, and the gain factor is defined asX = ( — - — ) Nhits+ 1, wherein Nraysis the total quantity of rays and Nhitsis the indication for \Nrays / ythe quantity of valid early reflection trajectories.
5. The method of any one of the preceding claims, wherein the total quantity of rays is defined as the quantity of rays multiplied by the quantity of points.
6. The method of any one of the preceding claims, wherein the gain factor is used to determine an early reflections gain.
7. The method of claim 6, wherein the early reflections gain is determined for a sector used for modelling the early reflections.
8. The method of claim 7, wherein the sector is determined by dividing a 360-degree angle around the listener position in equiangular portions.
9. The method of claim 8, wherein a quantity of equiangular portions is indicative of a precision for the modelling of the early reflections.
10. The method of any one of claims 7 or 9, wherein the early reflections gain is directly proportional to the gain factor multiplied by a mean value of reflection attenuation gains for the audio source for the sector.
11. The method of claim 10, wherein the mean value of reflection attenuation gains for the audio source for the sector is defined as Fkwherein k is the index of thesector, gERkis the reflection attenuation gain for early reflection trajectory i in sector k, nkis thetotal quantity of early reflection trajectories in sector k, and Nsis the maximum quantity of reflections per audio source for the reflection modelling.
12. The method of any one of claims 7 to 11, wherein the early reflections gain is determined for an image source in the sector.
13. The method of claim 12, wherein the image source is a virtual audio source perceived by the listener due to the early reflections.
14. The method of claim 12 or 13, wherein a location of the image source is the location of the audio source mirrored on a reflective surface in the three-dimensional audio scene.
15. The method of any one of the preceding claims, wherein the representation of the three-dimensional audio scene is a voxel-based representation.
16. The method of claim 15, wherein the indication for the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories or a quantity of collision voxels and wherein determining the indication for the quantity of valid early reflection trajectories for the audio source and the listener comprises: applying the ray direction pattern to the points in the proximity to the region connecting the audio source location and the listener location to obtain, for each of the points, a plurality of rays originating at the respective point; determining a set of collision voxels based on the plurality of rays and the voxel-based representation of the three-dimensional audio scene; counting collision voxels in the set of collision voxels to determine the quantity of collision voxels if the indication for the quantity of valid early reflection trajectories is the quantity of collision voxels; or determining early reflection trajectories based on the set of collision voxels, the listener location, the audio source location and a geometrical validity test; and counting the early reflection trajectories to obtain the quantity of valid early reflection trajectories if the indication for the quantity of valid early reflection trajectories is the quantity of valid early reflection trajectories.
17. The method of claim 16, further comprising: determining the points based on the quantity of points.
18. The method of any one of the preceding claims, wherein the ray direction pattern defines the predefined quantity of rays and predefined directions of rays from an origin.
19. The method of any one of the preceding claims, wherein the predefined quantity of rays is 6, 8, 12 or 26.
20. The method of claims 15 to 17 or claims 18 to 19, when depending on claim 15, wherein a voxel position in the three-dimensional audio scene is defined by grid indices and the predefined directions of rays comprise one or more of: horizontal and vertical directions of a grid index to neighboring grid indices; and diagonal directions of the grid index to the neighboring grid indices.
21. The method of claim 17, wherein the region connecting the audio source location and the listener location is a line connecting the audio source location and the listener location and points are positioned on the line.
22. The method of claim 21, wherein coordinates of the points on the line connecting the audio source location and the listener location are determined based on the quantity of points.
23. The method of claim 22, wherein the points are determined to split the line connecting the audio source location and the listener location into N-l equal segments where N is the quantity of points and is larger than or equal to 2.
24. The method of claim 17, wherein the quantity of points depends on available computational resources, an encoder preset, or a combination thereof.
25. The method of claim 16 or any one of claims 17 to 25 when depending on claim 16, wherein each collision voxel in the set of collision voxels is an occluder voxel in the voxel-based representation of the three-dimensional audio scene.
26. The method of claim 25, wherein the occluder voxel represents an acoustically reflective surface.
27. The method of claim 25, wherein the occluder voxel represents any material in the voxel-based representation of the three-dimensional audio scene other than sound transmission media.
28. The method of claims 25 to 27, wherein determining the set of collision voxels based on the plurality of rays and the voxel-based representation of the three-dimensional audio scene comprises: determining one or more intersections between each ray of the plurality of rays and the occluder voxels; and for each ray, determining an occluder voxel containing an intersection closest to the origin of the respective ray as a collision voxel in the set of collision voxels.
29. The method of claim 16 or any one of claims 17 to 28 when depending on claim 16, wherein determining the early reflection trajectories comprises: for each collision voxel in the set of collision voxels, determining a preceding voxel of the respective collision voxel, wherein the preceding voxel is a voxel containing an intersection with the respective ray, preceding the respective collision voxel in the direction of the respective ray; for each preceding voxel, determining a path connecting the listener location and the audio source location via the preceding voxel; and if the path can produce a geometrically valid representation of a first-order reflection, determining the path as an early reflection trajectory.
30. The method of claim 29, wherein determining whether the path can produce the geometrically valid representation of the first-order reflection comprises:determining that the path can produce a geometrically valid representation of a first-order reflection if the path does not contain an intersection with an occluder voxel.
31. The method of claim 29 or 30, wherein the path comprises a straight line connecting the audio source location to the preceding voxel and a straight line connecting the same preceding voxel to the listener location.
32. The method of any previous claim, wherein the rendering is to be performed by a virtual reality, VR, augmented reality, AR, mixed reality, MR, and / or extended reality, XR device.
33. The method of any previous claim, wherein the first-order trajectories are reflection trajectories with a single reflection between the audio source location and the listener location.
34. The method of any one of the preceding claims, wherein the method is performed by a decoder or Tenderer.
35. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 34.
36. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 34.
37. A computer-readable storage medium storing the program according to claim 36.
Citation Information
Patent Citations
Methods, apparatus, and systems for early reflection estimation for voxel-based geometry representation(s)
WO2023227544A1