Method, apparatus, and system for processing an audio scene for audio rendering
The method simplifies voxel-based audio scene updates and diffraction path determination using 2D projection maps, addressing computational complexity and enabling real-time realistic sound rendering in three-dimensional audio scenes.
Patent Information
- Application Number
- JP2025545796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-17
- Filing Date
- 2024-02-15
- Publication Date
- 2026-02-17
AI Technical Summary
Existing methods for diffraction modeling in three-dimensional audio scenes, particularly those using voxels, are computationally complex and difficult to implement in real-time due to frequent recalculation of diffraction paths, leading to high computational burdens and potential adverse user experiences in virtual environments.
A method for updating and decompressing voxel-based audio scenes using a simplified geometric representation, applying 3D smoothing filters, and employing 2D projection maps for diffraction path determination, reducing computational complexity while maintaining realistic sound rendering.
Enables efficient and realistic acoustic diffraction modeling in three-dimensional audio scenes, reducing computational workload and allowing for real-time rendering in complex virtual environments.
Smart Images

Figure 2026505661000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 485,731, filed February 17, 2023, which is incorporated herein by reference in its entirety. Technical Field This disclosure relates to techniques for processing audio scene information for audio rendering. In particular, this disclosure is directed to a voxel-based scene representation for three-dimensional audio scenes and audio rendering that takes into account diffraction effects caused by elements of the three-dimensional audio scene, for example, by applying diffraction modeling for audio occlusion. [Background technology]
[0002] A new audio standard known as Moving Picture Experts Group Immersive Audio (MPEG-I) has been proposed to enable audio experiences from different viewpoints and / or perspectives or listening positions by supporting various movements in and around a scene, such as movements using various degrees of freedom, such as three degrees of freedom (3DoF) or six degrees of freedom (6DoF), in virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) applications. 6DoF interaction extends 3DoF spherical video / audio experiences, which are limited to head rotation (pitch, yaw, and roll), to include translational movements (forward / backward, up / down, and left / right) in addition to head rotation, allowing navigation within a virtual environment (e.g., physically walking around a room).
[0003] For audio rendering in VR, AR, MR, and XR applications, an object-based approach has been widely adopted by representing a complex auditory scene as multiple distinct audio objects, each associated with parameters or metadata that define its location and trajectory in the scene. Alternatively, audio rendering in such environments also uses higher order ambisonics (HOA). However, new uses of "voxels" for rendering audio scenes are currently being explored, including for use in new immersive audio experiences. Voxels for audio rendering are relevant for media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments.
[0004] A voxel is a spatial volume to which acoustic properties or audio rendering instructions are assigned. The voxel size may be an encoder configuration parameter and may be selected (manually or automatically) according to the level of detail of the scene geometry (e.g., in the range of 10 cm to 1 m).
[0005] The voxels for audio rendering can be obtained by: Voxelization (or conversion) of mesh-based scene representations From the scene representation used for scene generation (or even video rendering) (for example, by downsampling to smaller voxels)
[0006] However, traditional approaches to using voxels to provide realistic sound for user experiences (including those involving movement) in virtual (VR, AR, MR, and XR) environments remain challenging and computationally complex.
[0007] In particular, due to the complexity of representing occlusion / diffraction-related object geometry (e.g., walls and holes), the dimensionality of space for audio rendering (e.g., 3D virtual reality), and requirements regarding realism and content creator intent for modeled effects (e.g., audibility range), diffraction modeling, especially modeling of acoustic diffraction effects in virtual environments (e.g., virtual reality or game worlds), is often abandoned entirely or replaced by direct signal propagation approaches. This is because typical techniques for modeling diffraction in three-dimensional audio scenes, such as for computer-mediated reality applications, require recalculation of diffraction paths and other diffraction information whenever either the audio scene, the user position, or the audio source position changes, making it difficult to accurately reproduce realistic acoustic effects in real time in a changing, complex three-dimensional virtual environment.
[0008] For example, the diffraction paths may change as the user and / or the audio source moves through the three-dimensional audio scene. Additionally, the diffraction paths may change when the audio scene itself changes, such as by indicating a door or window opening or closing. Frequent recalculation of the diffraction paths may be computationally expensive, which may require relatively powerful computing devices to implement the computer-mediated reality application and / or may in some cases adversely affect the user experience.
[0009] Thus, there is a need for improved techniques for diffraction modeling in three-dimensional audio scenes, particularly those that utilize voxels. There is a particular need for such techniques that can reduce the computational burden on devices (e.g., decoders, renderers) that implement such techniques, while allowing realistic sound effects to be accurately reproduced in real time in three-dimensional virtual environments. In other words, there is currently a need for improved methods and apparatus for realistic, yet computationally feasible, acoustic diffraction modeling for processing audio content for rendering in (virtual) three-dimensional audio scenes. Summary of the Invention [Problem to be solved by the invention]
[0010] In view of this need, the present disclosure provides a method for processing audio scene information (in particular voxel-based audio scene information) for audio rendering, an apparatus for processing audio scene information for audio rendering, a computer program, and a computer-readable storage medium, having the features of the respective independent claims. [Means for solving the problem]
[0011] An aspect of the present disclosure relates to a method for updating an audio scene for three-dimensional audio rendering. The method may include obtaining a voxelized representation of the audio scene. The voxelized representation may include a set of voxels forming a connected geometric region on a voxel grid of the voxelized representation. In particular, the geometric region may have a rectangular parallelepiped shape, and the set of voxels may include at least a first bounding voxel and a second bounding voxel that define the rectangular parallelepiped shape of the geometric region. Furthermore, the voxels within the geometric region may share common voxel characteristics. Furthermore, the audio scene may be updated, and in response to updating the audio scene, the method may further include determining an updated pair of bounding voxels that define an updated rectangular parallelepiped region for the geometric region. Alternatively or additionally, the method may further include determining updated voxel characteristics for voxels in the rectangular parallelepiped region. The method may further include generating a representation of the audio scene based on the updated pair of boundary voxels and the set of voxels having the updated voxel properties.
[0012] Configured as described above, the proposed method provides a simple approach for compressing voxel-based scene representations (i.e., voxelized representations of audio scenes), which allows for a simple and efficient scene update mechanism during changes in a three-dimensional audio scene, for example, when a user and / or an audio source moves through the three-dimensional audio scene. It can be appreciated that such an update method can be implemented in an encoder to provide an (updated) audio scene, i.e., an updated representation of the audio scene, to a decoder for rendering the updated representation of the audio scene. Thereby, such a proposed voxel-based scene representation technique may allow for creating complex dynamic scenes and encoding them efficiently.
[0013] In some embodiments, the method may further include determining at least a first boundary voxel and a second boundary voxel for a set of voxels from the plurality of voxels of the voxelized representation. In particular, the boundary voxels may describe a scene element that may be part of an audio scene to be updated. Furthermore, existing data associated with voxel characteristics for the scene element may be overwritten depending on the updated boundary voxel pair and / or the updated voxel characteristics.
[0014] In some embodiments, updating the audio scene may be based on update conditions including one or more of time, user input, input from a presentation engine, or input from application logic. Additionally, update attributes indicating the update conditions may be included as metadata in the bitstream along with the compressed representation of the audio scene based on the determined set of voxels. Additionally, updating the audio scene may be further based on triggers received in real time from a user.
[0015] According to another aspect, a method for decompressing a compressed voxel-based audio scene from a bitstream is provided. The method may include receiving a bitstream including a voxelized representation of the audio scene. Specifically, the voxelized representation may include a plurality of voxels arranged on a voxel grid, each voxel having an associated voxel characteristic. Furthermore, the method may include decoding a set of voxels forming a connected geometric region on the voxel grid and an indication that voxels within the geometric region share a first voxel characteristic. The method may also include decoding a subset of the set of voxels associated with a scene subelement within the geometric region and an indication that the subset of the set of voxels is assigned a second voxel characteristic. Furthermore, the method may further include generating an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels. Generating an updated representation of the audio scene may involve overwriting voxels of said subset with second voxel properties.
[0016] When configured as described above, the proposed method provides a simple approach for decompression of voxel-based scene representations (i.e., voxelized representations of audio scenes), which allows a simple and efficient scene update mechanism during changes in the three-dimensional audio scene, for example, when a user and / or an audio source moves through the three-dimensional audio scene. It will be appreciated that such a decompression method can be implemented in a decoder to render the (updated) audio scene, and the updated representation of the audio scene may be output to a renderer to render the updated representation of the audio scene.
[0017] According to another aspect, a method for processing an audio scene for three-dimensional audio rendering is provided. The method may include receiving a voxelized representation of the audio scene. In particular, the voxelized representation may include a set of voxels defining a connected geometric region on a voxel grid of the voxelized representation. The method may further include decoding the received voxelized representation of the audio scene. The method may also include applying at least one three-dimensional (3D) smoothing filter to the decoded voxelized representation of the audio scene.
[0018] In some embodiments, voxels within a geometric region may share common voxel characteristics. Specifically, the common voxel characteristics may include a reference to a material that describes an obscuration characteristic associated with a set of voxels for the geometric region. Moreover, the method may further include applying the at least one 3D smoothing filter to one or more obscuration coefficients that are indicative of the obscuration characteristic.
[0019] In some embodiments, the at least one 3D smoothing filter may be defined based on data transmitted in the bitstream or may be hard-coded in the renderer.
[0020] In some embodiments, the at least one 3D smoothing filter may be applied to a decoded set of voxels of a voxelized representation of an audio scene. Alternatively or additionally, the at least one 3D smoothing filter may be associated with a scene element identifier for a geometric region within the audio scene. In this case, the at least one 3D smoothing filter may be applied to the geometric region having the scene element identifier.
[0021] In some embodiments, the at least one 3D smoothing filter may be applied depending on the user position.
[0022] Configured as described above, by post-processing the decoded voxelized representation of the audio scene (i.e., as a decoded voxel matrix) with a 3D filter, the proposed method enables additional representation capabilities that can transform the regular shapes of geometric regions formed by a collection of voxels (e.g., rectangular, cuboid, etc.) into smoother shapes for objects, thereby achieving a better correspondence to the intended geometry (e.g., from the content creator) without requiring a large amount of metadata overhead for transmission. At the same time, applying a 3D filter is a simple and convenient way to smooth (average) the shapes of objects, as it can be easily implemented on the decoder side. Thus, an improved audio scene representation that approximates a more realistic acoustic environment can be provided to the renderer in a bandwidth- and cost-efficient manner.
[0023] According to another aspect, a method for processing an audio scene for three-dimensional audio rendering is provided. The method may include obtaining a scene configuration representing the audio scene. Specifically, the scene configuration may include a source position of an audio source, a given user position, and a scene description. More specifically, the scene description may include a voxel matrix for a voxelized representation of the audio scene and associated obscurance and diffraction coefficients. The method may also include obtaining a two-dimensional (2D) projection map associated with the voxelized representation of the audio scene. The method may also include determining a diffraction path between the audio source and the given user position based on the 2D projection map. Moreover, the method may further include obtaining auralization data for rendering the audio scene based on a result of the determination.
[0024] As configured above, by considering a voxelized representation of the 3D audio scene, the representation complexity can be significantly reduced. By further determining the diffraction path based on a 2D projection map (i.e., projecting the 3D audio scene onto the 2D projection map), it is possible to employ a reliable and efficient pathfinding algorithm (e.g., a 2D pathfinding algorithm) depending on the specific requirements of the rendering environment, thereby reducing the computational complexity of rendering the audio scene in the renderer.
[0025] Furthermore, it is understood that the determined (diffraction) paths (which can be output by the pathfinding algorithm) contain sufficient information for the generation of virtual sound sources at virtual source positions that realistically simulate the effects of sound diffraction in the original three-dimensional audio scene. Thanks to the reduced complexity achieved by the proposed method, this allows providing a realistic listening experience in a three-dimensional audio scene with a reasonable computational effort. In particular, this enables realistic sound rendering in three-dimensional audio scenes even for real-time applications such as virtual reality applications or computer / console games.
[0026] In some embodiments, the two-dimensional projection map may be obtained by applying a projection operation to a voxelized representation of the audio scene. Further, the diffraction paths may be determined based on one or more two-dimensional (2D) pathfinding algorithms using the 2D projection map for the voxelized representation of the audio scene.
[0027] In some embodiments, the 2D projection map may be either received from a bitstream transmitted by an encoder or calculated at the renderer by applying filtering to the voxel matrix of the scene description based on the source position and / or a given user position.
[0028] In some embodiments, the 2D projection map may be obtained by selecting from among at least one of a horizontal projection associated with the diffraction path and a vertical projection associated with the diffraction path. In particular, the horizontal projection may be related to the voxelized representation by a horizontal projection operation that projects onto a horizontal plane. Alternatively, the vertical projection may be related to the voxelized representation by a vertical projection operation that projects onto a vertical plane.
[0029] In some embodiments, the selection may be based on the route and / or direction of the diffraction path. In some embodiments, to determine the diffraction path, each voxel of the voxel matrix may be assigned a set of diffraction coefficients based on the geometry and / or respective material properties of the occluding object.
[0030] In some embodiments, the method may further include selecting a subset of voxels in the voxel matrix based on the determined diffraction path. Specifically, the subset of voxels may cause one or more changes in the diffracted sound direction. Moreover, the method may further include calculating coefficients of an EQ filter for the diffracted audio signal based on the selected subset of voxels. In particular, the calculation of the coefficients of the EQ filter may depend on at least one of a linear distance between the audio object and the listener, a length of the diffraction path, a diffraction coefficient of the corresponding voxel, and an angle of change in the direction of the diffraction path.
[0031] In some embodiments, the method may further include applying diffraction modeling to the voxelized representation of the audio scene.
[0032] Thus, the proposed method of applying a 2D projection plane (map) of a 3D voxel-based scene representation for diffraction modeling allows for a significant reduction in the computational workload for the rendering tool (i.e., renderer). The proposed method further allows for the flexible selection of one or more appropriate projection planes for diffraction calculations based on occluding structures. In particular, further projection planes may be considered to introduce possible additional paths for diffraction modeling, thereby increasing the accuracy of modeling the perceptual effect of acoustic occlusion / diffraction.
[0033] According to a further aspect, a method for processing an audio scene in a rendering device for three-dimensional audio rendering is provided. The method may include receiving a three-dimensional audio scene and information about a sound source at a source position. The method may also include determining a rendering mode of the rendering device. For a given listener position, the method may further include determining a virtual sound source at a virtual source position based on the source position to simulate the effect of acoustic diffraction by the three-dimensional audio scene on a source signal of the sound source at the source position. Specifically, the determination of the virtual sound source may be based on one or more of pre-computed data along a predefined user path and real-time data calculated using at least one of mesh-based diffraction modeling and voxel-based diffraction modeling, either of which may be selected depending on the given listener position and / or the determined rendering mode of the rendering device.
[0034] As configured above, the proposed method for processing audio scenes in a rendering device for three-dimensional audio rendering utilizes diffraction modeling for voxel-based audio scene representations to reduce computational complexity. Depending on the virtual environment being rendered, such as the user's path and / or position, changes in the audio scene, etc., the proposed method offers the possibility to select an appropriate implementation of diffraction modeling to achieve reduced complexity and / or improved bandwidth efficiency, which allows further optimization of hardware / software usage for rendering audio scenes in complex, evolving three-dimensional virtual environments in real time. It should be further noted that the choice of implementation for diffraction modeling may also depend on the device performing the rendering, where different modes may be configured to meet the respective hardware / software requirements.
[0035] In some embodiments, the method may further include determining a scene state of the three-dimensional audio scene as a known state and, in response, determining a virtual sound source based on pre-computed data along a pre-defined user path. Further, the rendering mode may include a reduced complexity mode and / or a bandwidth efficient mode.
[0036] In some embodiments, when the rendering mode is determined to be a reduced complexity mode, the method may further include determining the virtual sound source based on pre-computed data along a pre-defined user path or applying diffraction modeling to the three-dimensional audio scene using the pre-computed data along the pre-defined user path. Alternatively or additionally, the method may further include determining the virtual sound source based on real-time data calculated using voxel-based diffraction modeling. Whether to use pre-computed data along the pre-defined user path or real-time data calculated using voxel-based diffraction modeling may depend on a given listener position.
[0037] In some embodiments, when the rendering mode is determined to be a bandwidth-efficient mode, the method may further include determining the virtual sound source based on real-time data calculated using mesh-based diffraction modeling and / or voxel-based diffraction modeling, where whether to apply mesh-based diffraction modeling, voxel-based diffraction modeling, or both may depend on a given listener position.
[0038] In some embodiments, the method may also include determining a diffraction modeling order indicative of the complexity of the diffraction path for the given listener position. Additionally, the method may further include determining whether to use mesh-based diffraction modeling or voxel-based diffraction modeling based on the determined diffraction modeling order for the given listener position.
[0039] In some embodiments, when the diffraction modeling order for a given listener position is determined to be higher than a predefined diffraction order (e.g., third order), the method may further include determining a virtual sound source based on real-time data calculated using voxel-based diffraction modeling.
[0040] In some embodiments, when it is determined that for a given listener position, pre-computed data along a pre-defined user path is not available or real-time data calculated using mesh-based diffraction modeling is not available, the method may further include determining a virtual sound source at the given listener position based on real-time data calculated using voxel-based diffraction modeling.
[0041] In light of the above, the present disclosure proposes a computationally efficient method for modeling three-dimensional (audio) scenes, particularly for acoustic diffraction modeling of three-dimensional (audio) scenes. To this end, the present disclosure utilizes a simplified (but sufficiently accurate) geometric representation using a voxelization-based method. Furthermore, the two-dimensional space for diffraction modeling of the relevant geometric representation may be used together with a means for controlling sound effects that approximate acoustic hiding / diffraction phenomena for content creators and encoder operators, thereby allowing perceptually realistic simulation of acoustic hiding / diffraction effects for dynamic and interactive three-dimensional virtual environments. This can therefore improve the overall user experience and facilitate wider deployment of virtual reality (VR) applications.
[0042] According to another aspect, there is provided an apparatus for processing audio scene information for audio rendering, the apparatus may include a processor and a memory coupled to the processor that stores instructions for the processor, the processor may be configured to perform all steps of the methods according to the aforementioned aspects and embodiments thereof.
[0043] According to a further aspect, a computer program is described that may include executable instructions for performing the methods or method steps outlined throughout this disclosure when executed by a computing device (e.g., a processor).
[0044] According to another aspect, a computer-readable storage medium is described that may store a computer program adapted to run on a computing device (e.g., a processor) and, when run on the computing device, to perform the methods or method steps outlined throughout this disclosure.
[0045] It should be noted that the methods and systems, including preferred embodiments thereof, as outlined in this disclosure may be used independently or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this disclosure may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner.
[0046] It will be understood that apparatus features and method steps can be interchanged in many ways. In particular, details of the disclosed methods can be implemented by a corresponding apparatus, and vice versa, as will be understood by those skilled in the art. Furthermore, it will be understood that any statements made above regarding a method (and, e.g., steps thereof) apply equally to a corresponding apparatus (and, e.g., blocks, stages, units thereof), and vice versa. [Brief explanation of the drawings]
[0047] The invention will now be described, by way of example only, with reference to the accompanying drawings, in which: [Figure 1a] 1 shows a schematic of an example of a processing chain for processing audio scene information for audio rendering. [Figure 1b] Schematic showing example diffraction paths for a source and listener position in a voxel-based 3D audio scene. [Figure 1c] 1 is a flowchart illustrating an example method for processing audio scene information for audio rendering, according to an embodiment of the present disclosure. [Figure 2a] 1A and 1B illustrate schematic diagrams of examples of portions of voxel-based audio scenes according to embodiments of the present disclosure. [Figure 2b] 1A and 1B illustrate schematic diagrams of examples of portions of voxel-based audio scenes according to embodiments of the present disclosure. [Figure 2c] 1A and 1B illustrate schematic diagrams of examples of portions of voxel-based audio scenes according to embodiments of the present disclosure. [Figure 3]1 illustrates schematically an example of a voxel-based audio scene to which embodiments of the present disclosure may be applied. [Figure 4a] 10A-10C illustrate schematic examples of audio scene updates illustrating the effect of opening and closing a sliding door, according to embodiments of the present disclosure. [Figure 4b] 10A-10C illustrate schematic examples of audio scene updates illustrating the effect of opening and closing a sliding door, according to embodiments of the present disclosure. [Figure 4c] 10A-10C illustrate schematic examples of audio scene updates illustrating the effect of opening and closing a sliding door, according to embodiments of the present disclosure. [Figure 5a] 1 is a flowchart illustrating an example method for updating an audio scene for three-dimensional audio rendering, according to an embodiment of the present disclosure. [Figure 5b] 1 is a flowchart illustrating an example method for updating an audio scene for three-dimensional audio rendering, according to an embodiment of the present disclosure. [Figure 6] 10A-10C illustrate schematic examples of filtering effects on voxel block elements, according to embodiments of the present disclosure. [Figure 7] 1 is a flowchart illustrating an example method for processing an audio scene using 3D filtering for three-dimensional audio rendering, according to an embodiment of the present disclosure. [Figure 8] 10A-10C schematically illustrate example 2D projection surfaces of 3D voxel-based scene representations for diffraction modeling according to embodiments of the present disclosure. [Figure 9] 1 shows a flowchart illustrating an example method for processing an audio scene for three-dimensional audio rendering, according to an embodiment of the present disclosure. [Figure 10a] 1 illustrates schematically examples of possible use cases for techniques according to embodiments of the present disclosure. [Figure 10b] 10A-10C illustrate complexity metrics as a function of time for different operational modes / implementations of processing audio scene information for audio rendering, according to an embodiment of the present disclosure. [Figure 11] 10A-10C illustrate non-limiting examples of rendering an indoor audio scene using different implementation approaches, according to embodiments of the present disclosure. [Figure 12] 1 is a flowchart illustrating an example method for processing an audio scene in a rendering device for three-dimensional audio rendering, according to an embodiment of the present disclosure. [Figure 13] 1 is a block diagram that schematically illustrates an example of an apparatus for implementing a method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0048] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the accompanying drawings, in which the same elements may be designated by the same reference numerals, and repeated description thereof may be omitted.
[0049] Voxel-based audio scene representation First, we provide an overview of voxel-related concepts for audio scene representation.
[0050] What are voxels for audio rendering? A voxel is understood as a spatial volume that is assigned acoustic properties or audio rendering instructions.
[0051] What is Voxel Size for Audio Rendering? Voxel size can be an encoder configuration parameter: it may be selected (manually or automatically) depending on the level of detail of the scene geometry (e.g., in the range of 10 cm to 1 m).
[0052] How big an audio scene can it handle? A large audio scene does not necessarily result in a large number of voxels and high rendering complexity. For example, a large audio scene can be represented as follows: A collection of independent sub-scenes (and a way to "teleport" between these representations without a "restart" of the renderer) A collection of scene updates (based on user position)
[0053] How can we deal with the discontinuity problem caused by voxel granularity? Strong discontinuities in the sound level (and jumps in the direction of the diffraction signal) can be avoided by applying interpolation (eg, in time and space).
[0054] How to represent a voxel-based audio scene? Any voxel-based representation of an audio scene may include an indication of voxels that are not transparent (e.g., occluder voxels), i.e., voxels through which sound cannot or cannot propagate freely, i.e., a representation of occlusion geometry. This indication may be related to an indication of the coordinates (e.g., center coordinates, corner coordinates, etc.) of each voxel. These voxel coordinates may be represented, for example, by a grid index. Furthermore, the voxel-based representation may include an indication of the material properties of non-transparent voxels, such as absorption coefficients, reflection coefficients, etc. In addition to occluder voxels, the voxel-based representation may also indicate transparent voxels (e.g., air voxels), i.e., voxels through which sound can propagate, i.e., a representation of the sound propagation medium. Thus, some implementations of a voxel-based representation of an audio scene may include an indication of the respective material properties for each voxel in a predefined section of space (e.g., within a boundary surrounding the audio scene).
[0055] Techniques for Processing Audio Scene Information FIG. 1(a) schematically illustrates a processing chain 100 that may be used to process audio scene information for audio rendering. Specifically, the processing chain 100 may be used to convert voxel-related data into parameters and signals necessary for auralization (or audio rendering in general). The processing chain 100 may be implemented in software, hardware, or a combination thereof. For example, the processing chain 100 may be implemented by a renderer / decoder coupled to AR / VR / MR / XR devices, such as AR / VR / MR / XR goggles. Specific implementations may include game consoles, set-top boxes, personal computers, etc.
[0056] The processing chain receives an audio scene description 20 from a bitstream (or storage / memory) 10. The audio scene description 20 may contain a representation of a 3D audio scene and information about source positions of sound sources within the audio scene. The representation of the 3D audio scene may be, for example, voxel-based.
[0057] The processing chain 100 further receives an indication of a user position (listener position) 30 of a user (listener) within the audio scene. The audio scene description 20 and the user position 30 are provided to a diffraction direction calculation block (diffraction calculation block) 40 to determine (e.g., calculate) diffraction information. The diffraction information may relate to acoustic diffraction paths within the audio scene between the source position and the listener position. The diffraction information is then provided to a diffraction modeling tool 50 to apply diffraction modeling and optionally obscuration modeling based on the diffraction information. The obscuration modeling calculates attenuation gain for a line between the listener and the audio source. The diffraction modeling tool 50 may output sonified audio data (3DoF sonification data), for example, including the position, orientation, and frequency-dependent gain of the object to be rendered. The diffraction modeling tool output may be further processed by other rendering stages, such as Doppler, directivity, and distance attenuation. In general, the diffraction modeling tool 50 is said to output diffraction information, as described in more detail below. The sonified audio data may then be used for audio playback, for example.
[0058] In summary, a processing chain such as that shown in Fig. 1(a) can be used to convert voxel-related data into parameters for sonification and parameters for signals. The diffraction direction calculation block 40 and the diffraction modeling tool 50 can be considered as non-limiting examples of rendering tools. In general, rendering tools can generate 3DoF sonifier data.
[0059] As described above, the scene description may include voxel matrices and associated coefficients (e.g., reflection coefficients, obscuration coefficients, absorption coefficients, transmission coefficients, etc.). These coefficients may describe the material or material properties of each voxel. Rendering tools may include, for example, obscuration and diffraction modeling tools. 3DoF sonifier data may include, for example, object position, orientation, and frequency-dependent gain.
[0060] As mentioned above, a voxel-based representation of a 3D audio scene defines the psychoacoustically relevant geometric elements and sound propagation media. In some implementations, the scene description may use the following parameters / interfaces (e.g., the following agreed data formats or agreed points of data exchange) to provide this information to rendering tools:
[0061] Scene Size: In absolute units (e.g., meters) By number of voxels and / or voxel size Scene Anchor: As a coordinate anchor (mapping absolute coordinates to voxel indices) As a scene anchor (mapping a subscene to a subset of voxels) Scene Content Data: Reference to material properties (e.g., transmission, reflection, etc. coefficients) that approximate the acoustic effects caused by occluders (sound obstacles) located within the corresponding volume Reference to sound propagation medium properties (e.g., sound speed, energy absorption, distance attenuation curve, etc.) that approximate the acoustic effects caused by the medium located within the corresponding volume Rendering control parameters that describe the intended occlusion modeling effect For example, a "global" or "local" occluder type that determines the length (and shape) of the occlusion effect shadow behind this voxel. Rendering control parameters that describe the intended sound diffraction modeling effect For example, voxel types that control / cause a change in sound direction (i.e., diffracted sound paths cannot penetrate this volume) Content control parameters describing audio signal relevance and scene authoring For example, audio signal ID and / or signal gain determine which signals are perceptually relevant (rendered) in the corresponding volume. Rendering control parameters that describe the intended reverberation modeling effect For example, the voxel type that controls the reverberation settings (e.g., RT60, DDR, RIR, etc.). Scene content update: Referenced to update the trigger event All data can be audio object dependent (to support content creator intent in flexible audio scene authoring).
[0062] The 3DoF auralizer data can include the following information: Parameters and associated signals for a collection of audio objects (and HOAs) Parameters include the metadata output of the rendering tool (i.e., position, orientation, and gain to simulate the effects of occlusion, diffraction, early reflections, parameters for reverberation coefficients, IR, etc.) The associated signal represents the audio output of the rendering tool (i.e. the downmixed or duplicated audio signal). Scene state identifiers (i.e., metadata that allows mapping scene descriptions and user inputs to 3DoF sonifier data)
[0063] Figure 1(b) shows an example of a possible scene state and the diffraction paths for this scene state, which is understood to be associated with or include a listener position 210 and an audio scene description (including a representation of a three-dimensional audio scene and a source position 220).
[0064] The example in Figure 1(b) relates to a voxel-based representation of a three-dimensional audio scene. This voxel-based representation shows voxels 230, which represent "air" or empty voxels (i.e., voxels through which sound can propagate, or which are transparent voxels), and occluding body voxels 240, which represent voxels through which sound cannot propagate or propagate freely. Thus, occluding body voxels may be understood to refer to voxels filled with materials other than air, which may reflect, block, or otherwise alter sound propagation. For occluding body voxels 240, the representation may further show respective transmission coefficients, reflection coefficients, and potentially absorption coefficients related to the material properties of these voxels. These coefficients may be linked to the ID or index of each voxel in the voxel-based representation. In general, the voxel-based representation may define psychoacoustically relevant geometric elements and sound propagation media in an audio scene.
[0065] The listener position 210 is the parameter L VOX and the source position 220 is given by another parameter S VOX is shown by
[0066] The diffraction path between the source position 220 and the listener position 210 may be determined using a pathfinding algorithm that takes as input the listener position 210, the source position 220, and a representation of the three-dimensional audio scene (or a two-dimensional representation derived therefrom, e.g., a 2D projection or a 2D matrix). For example, an algorithm for determining diffraction information may take as input the listener position 210, the source position 220, and a representation of the three-dimensional audio scene, and VOX and the position of the diffraction corner 250, denoted by in For example, the diffraction information can be [C vox ,r in ]=DiffractionDirectionCalculation(L VOX ,S VOX ,VoxDataDiffractionMap) where DiffractionDirectionCalculation denotes an algorithm for determining diffraction information (a "pathfinding algorithm"), and VoxDataDiffractionMap is a voxel-based representation of a 3D audio scene or a processed version thereof (e.g., a 2D projection or a 2D matrix derived therefrom). VOX are understood to denote the coordinates of the diffraction corners (e.g., the coordinates of the respective voxels that contain the diffraction corners, voxel / grid coordinates, or voxel / grid indices).
[0067] Here, DiffractionDirectionCalculation can involve any promising pathfinding algorithm, such as the Fast Traversal Algorithm for Ray Tracing (see Non-Patent Document 1) and the JPS Algorithm (see Non-Patent Document 2). Furthermore, using a voxel-based scene representation, a 3D pathfinding algorithm can be directly applied to obtain the shortest path between the source position 220 and the listener position 210. Alternatively, a 2D pathfinding algorithm can be applied for this task using an appropriate 2D projection plane of the 3D voxel-based scene representation. For indoor (e.g., multi-room) sound simulation, the corresponding 2D projection plane can be similar to a floor plan describing the "sound propagation path topology." For outdoor sound simulation scenarios, it can be beneficial to consider a second (e.g., perpendicular) 2D projection plane to account for diffraction paths beyond sound obstacles or obscuring structures. While the pathfinding approach remains the same for all projection planes, its application provides additional paths that can be used for diffraction modeling. [Non-Patent Document 1] Amanatides, J. and A. Woo, A Fast Voxel Traversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987. 87 [Non-patent document 2] Harabor, DD and A. Grastien, Online Graph Pruning for Pathfinding On Grid Maps. Proceedings of the Twenty-Fifth AAAI Conference on Artificial Intelligence, 2011
[0068] The pathfinding algorithm is assumed to output a diffraction path consisting of multiple straight path segments (line segments) connecting the source position 220 to the listener position and linked sequentially end-to-end. Each transition from one path segment to another involves a change in the direction of the diffraction path.
[0069] According to the algorithm for determining diffraction information, the diffraction corner C VOX is a set of voxels C that are on or near the diffraction path and represent corner voxels on the diffraction map (shown by the voxel-based representation). set voxels adjacent to a corner voxel (in the VOX is the set of voxels that form the diffraction path (P set ) to the route (P set ) to change direction (from the listener position Lc) to the "visible" corner voxel (C set If there are multiple such corners, the voxels that belong to the diffraction path (P set ) is selected as the corner furthest from the listener position.
[0070] In general, a diffraction path algorithm can be said to determine diffraction information about the acoustic diffraction paths in the audio scene between a source position and a listener position.
[0071] This diffraction information may be sufficient for the renderer to recover / determine the virtual source position of a virtual audio source that encapsulates the effect of the acoustic diffraction effect.VOX coordinates and the diffraction path length r in For example, the virtual source position can be reconstructed by calculating the direction (e.g., azimuth, or azimuth and elevation) of the diffraction corner as seen from the listener position. Using this direction, the path length r of the diffraction path can be calculated. in The virtual source position can be determined by considering ∑ i = ...
[0072] Note that the diffraction information can be represented in a variety of ways. One option is to express the diffraction information in terms of the path length r as described above. in and diffraction corner C VOX Diffraction information containing / storing coordinates (e.g., grid coordinates).
[0073] Based on the above, the following data elements may be defined:
[0074] An example of scene state N1 is N1={L VOX ,S VOX ,VoxDataDiffractionMap} That is, the listener position L VOX , source position S VOX , and may be related to or include a voxel-based representation of the audio scene (e.g., VoxDataDiffractionMap).
[0075] The scene state identifier for scene state N1 is SceneStateIdentifier=HASH(N1) where HASH is a hash function that generates a hash value for the scene state N1, e.g., that maps the scene state to a fixed-size value. In general, a scene state identifier may be said to indicate or identify a certain scene state.
[0076] Further, examples of diffraction information N2 are: N2={CVOX ,r in} where r in is the path length of the diffraction path, and C VOX denotes the position (eg, voxel position) of the diffraction corner, as described above.
[0077] A quantized version of the diffraction information N2 may be denoted by N3, where N3=voxSceneDiffractionPreComputedPathData(N1) N3(~=N2)=DiffractionDirectionCalculation(N1) where voxSceneDiffractionPreComputedPathData() is a bitstream syntax that parses the bitstream and extracts the pre-computed (stored and quantized) diffraction information, and DiffractionDirectionCalculation() represents a function that performs online calculation of the diffraction information, which may be implemented, for example, in the diffraction direction calculation block 40 of FIG. 1(a).
[0078] User position voxel coordinate L vox is fixed, so the diffraction information, e.g., C vox and r in can also be considered as relating to the 3DOF sonifier data.
[0079] A technical benefit and effect of the techniques of the present disclosure is that a scene state identifier or other information derived from a scene state can be used to avoid the application of diffraction modeling or rendering tools (if corresponding processing has already been performed for this scene state and diffraction information or 3DoF sonifier data is available). In this scenario, a renderer can access the diffraction information / 3DoF sonifier data (for a known scene state) without applying rendering tools by: Reusing previously calculated (pre-calculated) data, or Applying said data calculated by another renderer.
[0080] Thus, a technical benefit is that the techniques disclosed herein relate to lossless functionality aimed at low complexity modes (complexity vs. bitrate), allowing flexibility in choosing an implementation mode suitable for efficient real-time rendering of audio scenes depending on the virtual environment being rendered.
[0081] 1(c) is a flowchart illustrating an example method 300 for processing audio scene information for audio rendering, according to an embodiment of the present disclosure. The method 300 may be implemented in software, hardware, or a combination thereof. For example, the processing chain 100 may be implemented by a renderer / decoder coupled to an AR / VR / MR / XR device, such as AR / VR / MR / XR goggles. Specific implementations may include a game console, a set-top box, a personal computer, etc.
[0082] Method 300 includes steps S310-S350, which may be executed, for example, by a decoder / renderer. These steps may be executed, for example, whenever a scene state changes. If the scene state is understood as relating to or including the listener position 210 and the audio scene description (including the representation of the three-dimensional audio scene and the source position 220), for example, as implemented by the above-described scene state N1, a change in the scene state may relate to one or more of a change in the listener position 210, a change in the source position, and a change in (the representation of) the three-dimensional audio scene. Alternatively, steps S310-S350 may be executed for each of multiple processing cycles of the decoder / renderer. However, if the audio scene description is not changed, step S310 may be omitted. It should also be understood that steps S310-S350 do not have to be executed in the order shown in FIG. 1(c).
[0083] In step S310, an audio scene description is received. The audio scene description includes a representation of a three-dimensional audio scene and information about source positions of sound sources within the audio scene. For example, the audio scene description may include, for example, the elements S defined above. vox and VoxDataDiffractionMap.
[0084] In step S320, information on the listener's position within the audio scene is received. The listener position can be determined, for example, by the element L defined above. VOX It can correspond to.
[0085] In step S330, diffraction information about sound diffraction paths in the audio scene between the source position and the listener position is obtained. The obtained diffraction information may indicate a virtual source position of a virtual sound source. For example, the virtual source position may be located at a diffraction corner C as viewed from the listener position. VOX The virtual source distance can be calculated by multiplying the diffraction path length r by the direction (e.g., azimuth, or azimuth and elevation). in Therefore, the diffraction information can be expressed as C vox and r in may include instructions for:
[0086] In step S340, audio rendering is performed for the sound source based on the diffraction information, which may include, for example, diffraction modeling.
[0087] To this end, the virtual source position of a virtual source can be determined based on the diffraction information. The virtual source can be an audio source that encapsulates the effects of acoustic diffraction between the source position and the listener position in the three-dimensional audio scene. For example, the virtual source position can be determined based on the diffraction information of the C VOX and r in Based on this, it can be determined by: Diffraction corner C as seen from the listener's position VOX Determine the direction (e.g., azimuth, or azimuth and elevation) of The determined direction is used as the virtual source direction of the virtual source as seen from the listener's position. Diffraction path length r in of the virtual source distance of the virtual source from the listener position, or Virtual source gain compensation derived from diffracted and direct path lengths Use as.
[0088] Audio rendering may then include, for example, rendering virtual sound sources at the virtual source locations.
[0089] In step S350, a representation of the diffraction information is output. For example, outputting the representation of the diffraction information may include outputting a data element that includes the diffraction information and information about the scene state. The scene state may be represented by an audio scene description (e.g., S vox and VoxDataDiffractionMap) and listener position (e.g., L VOX ).
[0090] The output may be provided to a look-up table (LUT). The LUT contains as its entries different items of diffraction information indexed with information about the respective scene states (e.g., indexed with respective scene state identifiers). The LUT is thus said to contain diffraction information and information about the scene states. The LUT may be stored and / or provided and later retrieved, e.g., by another decoder, from the bitstream or from shared storage (e.g., cloud or server-based), e.g., upon application request. A hash value of the scene state or the scene state identifier may be used to retrieve the actual desired entry from the LUT.
[0091] Additionally, the representation of the diffraction information may be output to a bitstream (e.g., an outgoing bitstream) and / or storage (e.g., memory, cache, file, etc.). The storage may be local or shared (e.g., cloud-based). In general, the representation of the diffraction information may be output to any suitable medium for storing digital or computer-related information. The output may be directed, at least in part, to an external or shared data source or data repository.
[0092] In some implementations, a representation of the diffraction information may be output as part of the voxSceneDiffractionPreComputedPathData() syntax element according to ISO / IEC 23090-4 (Non-Patent Document 3) or any future standard derived therefrom. [Non-patent document 3] Coded representation of immersive media ― Part 4: MPEG-I immersive audio, https: / / www.iso.org / standard / 84711.html
[0093] Efficient Voxel-Based Audio Scene Representation Current representation formats for voxel-based scenes include, for example, *.vox, *.binvox, etc.
[0094] This disclosure provides the following compression approach, for example for the MPEG-I audio standard: A set of voxels with the same acoustic properties (e.g., the same material properties) or the same audio rendering instructions can be identified by two points forming a rectangular region on the voxel grid. All voxels in this rectangular region share the acoustic properties or audio rendering instruction set assigned to the corresponding two points, as follows:
[0095] That is, the scene geometry is determined by a set of such point pairs (as an example of a representation of a geometric domain), e.g. <voxbox id="V_ID" material="P_ID" Point_S="X1 Y1 Z1" Point_E="X2 Y2 Z2" / > where V_ID is the rectangular voxel block element identifier; P_ID is the acoustic property or audio rendering instruction set identifier (e.g., occlusion, reflection, RT60 data); and X1, Y1, Z1 and X2, Y2, Z2 are the grid indices of the corresponding two points that define the rectangular voxel block element. Thus,<VoxBox_id, material, Point_S, Point_E / > may correspond to a scene element as defined above, i.e., a geometric region (e.g., defined by a pair of points) together with its voxel characteristics (e.g., P_ID) and optionally its identifier (e.g., V_ID).
[0096] The voxel size can be determined by the number of voxels as follows: <voxsize size="N" / > where N is the number of voxels along the first longest scene dimension.
[0097] Thus, a voxel-based representation of an audio scene according to embodiments of the present disclosure may include representations or indications of one or more rectangular geometric regions (rectangular spatial regions, rectangular volumes) having identical (i.e., the same, common) acoustic properties (e.g., materials, absorption coefficients, reflection coefficients, etc.) or identical rendering instructions. The acoustic properties or rendering instructions for a given voxel may be non-limiting examples of voxel properties of a given voxel. Representations of indications of rectangular geometric regions may, for example, relate to scene elements. It is understood that the rectangular regions are each non-trivial in the sense that they each contain more than a single voxel and consist of a collection of connected (i.e., contiguous) voxels.
[0098] The shape of each of these geometric regions can be defined by first and second boundary voxels (e.g., the above-mentioned pairs of points). These first and second boundary voxels can be associated with diagonal corners (extreme-corner voxels) of a rectangular prism, such as the extreme-corner voxel with the smallest x, y, and z coordinate values or coordinate index, and the extreme-corner voxel with the largest x, y, and z coordinate values or coordinate index. Other choices of diagonal extreme-corner voxels are possible as well.
[0099] Further, in the voxel-based representation, each geometric region may be represented by an indication of at least first and second boundary voxels and an indication of common voxel characteristics of voxels within the geometric region. Further, the representation of the geometric region may include an identifier (ID) for the geometric region.
[0100] FIG. 2(a) shows an example of a geometric region 1002 within a voxel grid 1007. The geometric region 1002 includes multiple voxels 1003 that form a rectangular parallelepiped. The shape or size of the geometric region can be represented by first and second boundary voxels 1004-1, 1004-2, such as the extreme corner or extreme corner voxels. In this example, the boundary voxels correspond to the lower left front corner and the upper right rear corner, respectively. FIG. 2(b) shows a side view of the geometric region 1002, as viewed from the right front of FIG. 2(a).
[0101] For the proposed compression approach, each time a next point pair (i.e., each time a next geometric region) is reached, the voxel properties in the corresponding rectangular region may be redefined. That is, the voxel properties of the subsequent geometric region may overwrite any previously assigned voxel properties for the voxels of the subsequent geometric region. In some implementations, a smaller geometric region that is entirely contained within a larger geometric region may redefine or overwrite the voxel properties of voxels in the smaller geometric region with the voxel properties of voxels in the smaller geometric region. Here, it is understood that corresponding voxel properties are overwritten while other voxel properties are maintained. For example, if a subsequent geometric region defines acoustic properties for its voxels, these acoustic properties are used to overwrite the acoustic properties defined for the voxels of the previous geometric region, while any rendering instructions for the voxels of the previous geometric region are maintained.
[0102] As noted above, the order between geometric regions may be derived, for example, from whether the geometric regions are completely contained within each other, from the order of representations or indications of the geometric regions in the bitstream, or from a predefined order referencing identifiers (IDs) of the geometric regions.
[0103] FIG. 2(c) shows schematically how voxel properties assigned to a first geometric region defined by the extreme corner voxels 1004-1, 1004-2 can be overwritten by voxel properties of a second geometric region 1002.
[0104] 3 shows a schematic example of a voxel-based description of a three-dimensional audio scene that includes multiple rectangular volumes of common voxel characteristics, potentially with smaller rectangular sub-volumes that redefine the voxel characteristics. The audio scene may include voxels 1101 that are local sound occluders that occlude within a given acoustic environment of the audio scene. The audio scene may further include voxels 1102 that are local sound occluders that separate acoustic environments from each other. The voxel-based scene representation approach proposed in this disclosure allows for the creation of complex dynamic scenes and their efficient encoding.
[0105] An Update Mechanism for Voxel-Based Audio Scene Representation As described above, voxel-based representations can be compressed by defining the shape of a scene geometry (geometric region) by a collection of point pairs (i.e., bounding voxels), which may also be referred to as a scene definition, each of which may be associated with a particular scene state. To provide audio rendering of a three-dimensional environment for a realistic user experience, this disclosure further proposes a scene update mechanism that allows switching between different scene states to effect a scene update, e.g., as a user and / or audio source moves through the three-dimensional audio scene. Such an update mechanism can be implemented in an encoder to provide an updated audio scene, i.e., an updated representation of the audio scene, to a decoder, or can be implemented in a decoder to render an updated representation of the audio scene.
[0106] It can be appreciated that the above compression techniques for voxel-based scene representations may further allow for a simple and efficient scene update mechanism. This may be based on a similar concept of redefining or overwriting voxel properties (i.e., old data) for voxels within a smaller geometric region with voxel properties (i.e., new data) for the smaller geometric region. More specifically, a new set of voxel geometry specifications may be associated with the following attributes that define the conditions for applying the update: Time (e.g., absolute, time offset, etc.) User input (e.g., position, orientation, interaction, etc.) Input from the presentation engine or application logic (e.g., animated hidden objects) A combination of the above (e.g., user proximity conditions).
[0107] In other words, similar to the proposed voxel-based representation for scene definitions that apply boundary voxels, updated regions for scene geometry can also be defined by updated pairs of boundary voxels. Moreover, to represent an audio scene update, it may be sufficient to simply determine updated voxel characteristics for voxels within a geometric region. An audio scene or scene geometry update with an updated region may be associated with a different scene definition than the previous (non-updated) scene definition. Thus, for the representation of an updated audio scene, a scene definition based on the voxel-based representation described above can be used in a similar way, while also considering additional conditions, for example, regarding when each scene definition can occur. Here, a simple example of an audio scene update describing the effect of opening and closing a sliding door is shown in Figure 4.
[0108] Similar to FIG. 2, FIG. 4(a) illustrates a geometric region 1002 within a voxel grid 1007 within an audio environment 1000. The geometric region 1002 includes a plurality of voxels 1003, and the shape or size of the geometric region 1002 may be represented by first and second boundary voxels 1004-1, 1004-2, such as the extreme corners or the extreme corner voxels. In this example, the boundary voxels correspond to the lower left front corner and the upper right rear corner, respectively. The geometric region 1002 may be considered, for example, as a wall within the audio environment 1000.
[0109] Unlike the example shown in FIG. 2, the geometric region 1002 in FIG. 4(a) includes a door portion 1005 that can be opened or closed, e.g., by sliding, based on a user trigger. For example, in the virtual environment, a user can approach the sliding door 1005 and press a button to trigger the sliding door 1005 to open or close. FIGS. 4(b) and 4(c) refer to two different scene states responsive to updates representing a sliding door effect. When the sliding door 1005 opens, the audio scene is switched / updated from FIG. 4(b) to FIG. 4(c), and when the sliding door 1005 closes, the audio scene is switched / updated from FIG. 4(c) to FIG. 4(b). Furthermore, more than two scene states representing the process of moving the sliding door 1005 (opening or closing) may be considered to render the corresponding movement of the sliding door 1005.
[0110] Specifically, a door part 1005 may be formed as a subregion of the geometric region 1002 by a set of voxels from the plurality of voxels 1003. As described above, the set of voxels within the subregion may be redefined or overwritten to create the door part 1005. That is, the set of voxels within the subregion may be associated with the door part 1005. Similar to the geometric region 1002, the shape or size of the door part 1005 may be represented by voxels at the extreme corners or at the extreme corners, as shown, for example, by the boundary voxels at the upper-left front corner 1005-1 and the lower-right back corner (not shown) in FIG. 4( a). It should be noted that the voxel positions shown in FIG. 4( a) are non-limiting examples illustrating possible boundary voxel positions that may be selected to represent a particular geometric region, and that voxels at other locations, such as the lower-left front corner and the upper-right back corner, may also be selected as boundary voxels to represent the same geometric region.
[0111] In response to a user trigger to open the sliding door 1005, the state of the door 1005 switches from 1005b to 1005c, as shown in FIGS. 4(b) and 4(c), respectively. These correspond to the reduced door size caused by the door opening. Meanwhile, the gap portion 1006 between the door portions 1005b and 1005c and the wall 1002 expands from state 1006b to state 1006c, as shown in FIGS. 4(b) and 4(c), respectively. To achieve this, the voxel properties of the voxels within the door portion and / or the gap portion may be overwritten (or redefined) with updated voxel properties to obtain an updated scene definition corresponding to the scene state of FIG. 4(c).
[0112] More specifically, voxel characteristics associated with the sliding door 1005 to be updated / overwritten may include, for example, the door portion's material (wood, metal, etc.) and / or air material. When the door is opened, the door portion's wood / metal material may be overwritten with air material from one voxel to another within the region to be updated (i.e., updated from door to gap), which may be defined by an updated pair of boundary voxels. For example, voxels 1006b-1, 1006b-2 and 1006c-1, 1006c-2 may be determined as boundary voxels of the region being updated to be the gap portion. In the illustrated example, boundary voxels 1006b-1 and 1006b-2 define the corresponding gap portion 1006b for the scene state of FIG. 4(b), and boundary voxels 1006c-1 and 1006c-2 define the corresponding gap portion 1006c for the scene state of FIG. 4(c). When the scene state switches from FIG. 4(b) to FIG. 4(c) in response to a user triggering the door to open, an updated pair of boundary voxels 1006c-1 and 1006c-2 defining an updated gap portion (region) 1006c is applied, in which voxel properties indicating wood / metal material for the door portion may be overwritten with air material to change the associated voxels into a gap portion.
[0113] Similarly, when closing a door, the air material in the associated voxel property may be overwritten with the wood / metal material from one voxel to another within the region to be updated (i.e., updated from gap to door) which may be defined by the updated pair of boundary voxels. That is, when the scene state switches from Figure 4(c) to Figure 4(b) (and then from Figure 4(b) to Figure 4(a)) in response to a user triggering the door to close, the pair of boundary voxels defining the gap portion (region) will be updated, i.e., the gap portion will be updated from 1006c (defined by 1006c-1 and 1006c-2) to 1006b (defined by 1006b-1 and 1006b-2). That is, the updated pair of boundary voxels 1006b-1 and 1006b-2 that define the updated gap portion (region) 1006b is applied to specify voxels (e.g., voxels within gap portion 1006c but outside gap portion 1006b) whose associated voxel properties for the air material are overwritten with the wood / metal material of the door portion.
[0114] It should be noted that the above-described audio scene update representing the effect of opening and closing a sliding door should be considered as a non-limiting example of the scene update mechanism proposed by this disclosure, which may also be applied to any other audio scene that can be captured by a voxel-based representation, taking into account various material / media properties for a given voxel and various shapes / geometry / arrangements of audio scene elements, as described above.
[0115] In addition to different scene definitions associated with different scene states, the proposed scene update mechanism also considers one or more update conditions for each scene definition (state). In other words, different scene states can be defined along with update conditions to specify when or how a particular scene state can occur.
[0116] Continuing with the above example of a sliding door effect, when a user triggers the opening of a sliding door, the following definition may be initiated, overwriting the door section's material with an air material to obtain a gap between the door section and the wall. This scene change may occur immediately or based on a time condition that specifies that the change should occur after a period of time, for example, 0.5 seconds. For example, the gap may become larger in (another) 0.5 seconds, and the door may be fully open in another 0.5 seconds. Similar definitions may be assigned to any possible (update) conditions such as those mentioned above, for example, time, user interaction, etc.
[0117] It can be understood that the previous concept of expressing a scene definition without conditions can be seen as simply initializing / predefining an initial state of the scene, which may then be realized using conditional updates, i.e., different states of the scene are executed based on the updated conditions. Thus, the definition may also include time delays, space, conditions leading to the current state, etc., which may be included in the bitstream. Furthermore, a new definition / state can be obtained after an update, which may also depend on an external trigger by the user (i.e., not included in the bitstream).
[0118] Updates may also occur in real time based on input from a representation engine, a computation engine, or application logic. In the above example of opening and closing a sliding door, if the presentation engine receives a bitstream containing information about the update condition, then in response to a user trigger, the door animation is performed to provide a "gap portion" to the renderer based on the update condition. In some other examples, the presentation engine may also be used without taking a bitstream into account. Thus, for example, a simple interface may be provided to provide the initialization of the scene data, and then the audio scene may be updated or modified according to the operation logic of the representation engine (e.g., based on properties, rules, coordinates, conditions, etc. defined in the logic), thereby simplifying the update process for a three-dimensional audio scene. A similar interface may also be used to take into account a bitstream coming from a remote engine and stored locally.
[0119] In the example above of opening and closing a sliding door for an animated sliding door implementation, the update conditions included in the bitstream may include several different scene states / definitions at different times during the opening / closing process. The update conditions included in the bitstream may further include a time variable indicating when the opening / closing process should be performed, for example, the duration after which the process should start, or whether the process should start immediately in response to a user trigger. Thus, if two states (e.g., open and closed) are considered, the bitstream may provide information that the material properties of the voxels between the two points (voxels) defining the door should be changed to an air material to open the door, as well as time information indicating the duration of time that must be waited to open the door.
[0120] Meanwhile, in response to another user trigger to close the door, the same voxel points may be applied, but the material properties between those points will be replaced with the door material, and the time variable indicates when the door closing effect will occur. For example, if the time variable indicates a duration of 1 second, the material update between these two points will occur after 1 second, i.e., the door closing animation will occur 1 second after the user triggers the door closing process.
[0121] Note that user triggers are generated externally (e.g., from the user) and are not included in the bitstream. However, user triggers may be sent directly to the rendering engine. The bitstream forwarded to the renderer may contain different scene states (e.g., information about the different states a sliding door may have), update conditions, and time information that indicates to the rendering engine when the updates are to be performed. As mentioned above, the engine may be provided with some interface for receiving the update conditions and external triggers from the bitstream. Note further that the renderer may be provided with a bitstream coming from the content creator side and interactions (triggers) coming from the user, and rendering may be triggered by an interface that defines the update conditions in response to an external trigger (e.g., by a short delay) rather than being triggered directly by the bitstream.
[0122] 5(a) shows an example flowchart of a method 500a for updating an audio scene for three-dimensional audio rendering according to an embodiment of the present disclosure. The method 500a may be implemented in software, hardware, or a combination thereof in an encoder to provide an updated representation of the audio scene (i.e., an updated audio scene) to a decoder for rendering the updated representation of the audio scene.
[0123] Method 500a includes steps S510 to S530, which may be performed, for example, by an encoder. In particular, these steps may be performed, for example, when an audio scene update occurs, i.e., when the scene state (or scene definition) of the audio scene changes. As described above, the scene state may be related to or include a listener position and an audio scene description (including a representation of a three-dimensional audio scene and a source position), and a change in the scene state may be related to one or more of a change in listener position, a change in source position, and a change in (the representation of) the three-dimensional audio scene. Furthermore, it should be noted that steps S510 to S530 may be performed for each of multiple processing cycles of the encoder and need not be performed in the order shown in FIG. 5(a).
[0124] In step S510, a voxelized representation of the audio scene is obtained. The voxelized representation includes a set of voxels that form a connected geometric region on a voxel grid of the voxelized representation. For example, the voxelized representation of the audio scene may be similar to the example shown in FIG. 2, where the geometric region has a rectangular parallelepiped shape and the set of voxels includes at least a first boundary voxel and a second boundary voxel that define the rectangular parallelepiped shape of the geometric region. In particular, the voxels within the geometric region may share common voxel properties.
[0125] In step S520, in response to an audio scene update, an updated pair of boundary voxels defining the updated region and / or updated voxel properties for voxels within the updated region are determined. Specifically, the updated pair of boundary voxels defines an updated rectangular parallelepiped region for the geometric region, and updated voxel properties are associated with voxels within the updated rectangular parallelepiped region. Note that in this step, determining either the updated pair of boundary voxels or the updated voxel properties may be sufficient to update the audio scene. However, in response to a scene update, both the updated pair of boundary voxels and the updated voxel properties may be determined to provide a better perceptual effect to the user.
[0126] In step S530, an (updated) audio scene representation is generated based on the updated pairs of boundary voxels and / or the set of voxels with updated voxel properties. The generated representation of the updated audio scene may then be provided to a decoder for rendering the updated representation of the audio scene. It will be further appreciated that the proposed updating method may generate multiple representations of the (updated) audio scene for multiple scene states (definitions) that are delivered to the decoder (e.g., via a bitstream) in order to efficiently render complex dynamic scenes.
[0127] A similar method 500b may be implemented on the decoder side to render an (updated) audio scene to be output to a renderer. Method 500b may provide functionality for decompressing a compressed voxel-based (updated) audio scene from a received bitstream. Method 500b includes steps S540-S570, which may be executed, for example, in response to a user trigger or interaction indicating an audio scene update. Similar to method 500a for the encoder, steps S540-S570 may be executed for each of multiple processing cycles of the decoder and need not be executed in the order shown in FIG. 5(b).
[0128] In step S540, a bitstream including a voxelized representation of an audio scene is received. Similarly, the voxelized representation includes a plurality of voxels arranged in a voxel grid, each voxel having an associated voxel characteristic. Method 550b further includes step S550 of decoding a set of voxels on the voxel grid that form a connected geometric region and an indication that voxels within the geometric region share a first voxel characteristic. Method 550b also includes decoding (S560) a subset of the set of voxels associated with a scene subelement within the geometric region and an indication that the subset of the set of voxels has been assigned a second voxel characteristic. Additionally, method 550b includes generating (S570) an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels, which involves overwriting the voxels of the subset with the second voxel characteristic. The updated representation of the audio scene may be output to a renderer for rendering the updated representation of the audio scene, allowing a simple and efficient scene update mechanism during modification of the 3D audio scene.
[0129] Post-processing of the voxel matrix for the received scene description It is further understood that after decompression, the decompressed voxel matrix may be additionally post-processed to enhance its representation capabilities. For example, one or more three-dimensional (3D) filters (e.g., Gaussian filters) may be applied to the voxel matrix after decoding / decompressing the scene representation to create blocks (i.e., voxel block elements for portions of a scene) with smooth structures. The averaging provided by the filters may smooth out angular objects / elements in the scene, approximating a more realistic acoustic environment (e.g., looking or sounding more natural). Figure 6 schematically illustrates an example of the filtering effect on voxel block elements according to an embodiment of the present disclosure. In this non-limiting example, the block elements may represent tree crowns (601a, 601b) as parts of tree leaves. In other examples, the proposed method may also be applied to block elements of other types of shapes in an audio scene.
[0130] Figure 6(a) shows a representation of a voxel block element after decompression but before applying a 3D filter. The voxel block element may have a rectangular shape (601a) for the tree crown, defined by a cube / block assigned to the transparency characteristics of the tree leaves. For example, two points (voxels) can be used to code a cube of voxels representing the tree crown with reference to material describing the obscuration characteristics of its leaves. Figure 6(b) shows a representation of the voxel block element after applying a 3D filter. Application of a 3D smoothing filter can transform the rectangular shape of the tree crown (601a) into an object with a smoother shape (601b) to achieve a better correspondence to the intended geometry without significant metadata overhead.
[0131] It will be appreciated that applying smoothing filtering before initiating the rendering tool does not require extra bitrate (e.g., no extra overhead bits for coding the intended geometry), providing an improved audio scene representation that approximates a more realistic acoustic environment for rendering in a bandwidth- and cost-efficient manner. It will be further appreciated that for objects / elements with specific shapes, 3D filters may be predefined or transmitted via the bitstream. Specifically, such 3D smoothing filters may include (small-sized) matrices, depending on the resolution and the number of times they are applied. In one example, a 3x3 matrix applied three times may be sufficient for a tree approximately 10 meters away.
[0132] In particular, such smoothing (e.g., averaging) filters may be applied to corresponding concealment coefficients (during direct sound concealment modeling) to make, for example, the periphery of a tree canopy more acoustically transparent than its core. Furthermore, such 3D filters may be based on data transmitted in the bitstream or may be hard-coded in the renderer.
[0133] In some examples, a 3D filter may be associated with a filter identifier to specify the filter to be applied and how many times the filter should be applied. Filtering may also be modified depending on the user position. For example, if a user approaches an object (e.g., a tree) from a far away position, filtering (e.g., a smoother or average in this case) may be activated or modified accordingly to better render the scene. Also, different filtering shapes may be defined and selected to improve user perception in dynamic audio environments.
[0134] After filtering, the filtered scene description may be delivered to the renderer, unchanged. In other words, filtering may be applied to the decompressed voxel matrix before applying the rendering matrix. Thus, the filtering operation may be decoupled from the scene description received in the bitstream to increase application flexibility. For example, based on a user's movement in the audio scene, different filtering may be applied to the same (e.g., original) scene description, which can be repeatedly retrieved from the bitstream.
[0135] 7 shows an example flowchart of a method 700 for processing an audio scene for three-dimensional audio rendering, according to an embodiment of the present disclosure. The method 700, including steps S710 to S730, may be implemented in a decoder to provide a renderer with an improved representation of the audio scene.
[0136] In step S710, a voxelized representation of an audio scene is received. The voxelized representation includes a set of voxels that define a connected geometric region on a voxel grid of the voxelized representation. In step S720, the received voxelized representation of the audio scene is decoded. In step S730, at least one 3D smoothing filter is applied to the decoded voxelized representation of the audio scene.
[0137] Efficient computation of diffraction paths for diffraction modeling As mentioned above, a 3D pathfinding algorithm can be directly applied to obtain the shortest path between the source and listener positions using a voxel-based scene representation, or, to reduce the computational workload, a 2D pathfinding algorithm can be applied for this task using an appropriate 2D projection plane of the 3D voxel-based scene representation. Such an appropriate 2D projection plane can be adapted to the location for the sound simulation, such as using a floor plan (e.g., horizontal) for indoor scenarios and a second (e.g., vertical) 2D projection plane for outdoor scenarios to account for diffraction paths beyond obstacles or obscuring structures. That is, the same pathfinding approach (algorithm) can be adapted to different projection planes to account for additional paths for more accurate diffraction modeling.
[0138] For example, horizontal projection of a three-dimensional audio scene has been applied to calculate the virtual source position of a virtual sound source. Because humans are generally more sensitive to the horizontal dimension (i.e., azimuth) than the vertical dimension in audio spatial localization, when applying the corresponding 2D projection plane for diffraction modeling, the resolution of the vertical axis is lower than that of the horizontal axis. For indoor scenes, it may be sufficient to apply only the vertical or horizontal plane with lower resolution to calculate the diffraction path. However, for outdoor scenes, the present disclosure proposes additionally applying a vertical projection plane (with better resolution) to improve the representation of outdoor scenes.
[0139] For example, as with the application of horizontal projections, the same pathfinding algorithms may be used for vertical analysis, in which case the scene geometry may be divided across the vertical axis to calculate the diffraction path on the vertical axis. Combining the application of horizontal and vertical projections as 2D projection planes for diffraction modeling may result in two or more virtual audio objects that can be found on the horizontal and vertical planes, respectively.
[0140] Figure 8 schematically illustrates two example 2D projection planes of a 3D voxel-based scene representation for diffraction modeling according to an embodiment of the present disclosure. Here, the audio scene may include, for example, a house, as shown in Figure 8(a), which depicts an obscuring structure for an outdoor scenario. Figures 8(b) and 8(c) illustrate corresponding vertical and horizontal planes, respectively, as 2D projection planes for diffraction modeling, where two respective virtual audio sources (VS_v, VS_h) and their positions can be calculated based on information about the original audio source (S) and the listener position (L).
[0141] In this way, flexibility in selecting one or more appropriate projection planes for diffraction calculations based on the occlusion structure is allowed, thereby improving the accuracy of modeling the perceptual effect of occlusion / diffraction.
[0142] Specifically, the 2D projection plane can be calculated on the encoder side and then transmitted in the bitstream, or calculated on the renderer side (e.g., by the "default" projection plane calculation method). For example, in the renderer, filtering (or "matrix slicing cut") can be applied to the 3D voxel matrix, taking into account the position of the user and / or audio source. The resulting diffraction paths can then be used by taking into account the diffraction coefficients. Assuming the use of a virtual source rendering method, the same (or similar) tools and interfaces as for reflection modeling can be applied. Note that while reflection coefficients are used for reflection modeling, modeling of diffraction effects is based on diffraction coefficients.
[0143] The diffraction coefficients may be obtained in an encoder, where a set of diffraction coefficients (similar to those of reflection coefficients) is assigned to each voxel based on the geometry of the occluding object and its material properties. Alternatively, the diffraction coefficients may be obtained in a renderer, where a subset of voxels that cause a directional change in the diffracted sound is selected based on the calculated diffraction path. Furthermore, in the renderer, EQ coefficients for the diffracted audio signal may be further calculated based on the obtained subset of voxels (and the corresponding assigned diffraction coefficients).
[0144] In particular, the EQ filter for a diffracted audio signal can be calculated based on the direct distance between the audio object and the listener. Furthermore, the calculation of the EQ filter for a diffracted audio signal can also depend on the diffracted path length (or an approximation thereof), the diffraction coefficient of the corresponding voxel (or projection plane matrix element), and / or the direction change angle of the diffracted path.
[0145] 9 shows an example flowchart of a method 900 for processing an audio scene for three-dimensional audio rendering, according to an embodiment of the present disclosure. Method 900 may be implemented, for example, in diffraction calculation block 40 and / or diffraction modeling tool 50 in processing chain 100. In particular, method 900 includes steps S910 through S940 for providing sonified audio data (e.g., 3DoF sonifier data) to be further processed by other rendering stages.
[0146] In step S910, a scene configuration is obtained that describes an audio scene, the scene configuration including source positions of audio sources, a given user position, and a scene description including a voxel matrix and associated obscuration and diffraction coefficients for a voxelized representation of the audio scene.
[0147] The method 900 further includes a step S920 of obtaining a two-dimensional projection map associated with the voxelized representation of the audio scene. The method 900 further includes a step S930 of determining a diffraction path between an audio source and a given user position based on the 2D projection map. The method 900 further includes a step S940 of obtaining auralization data for rendering the audio scene based on the results of said determination.
[0148] Application of Diffraction Modeling It will be understood that for the implementation of audio scene rendering, the diffraction information for diffraction modeling (e.g., diffraction path information regarding acoustic diffraction paths in the audio scene between a source position and a listener position) and the results of diffraction modeling (e.g., sonified audio data) may be obtained from pre-computed results, i.e., diffraction information and diffraction modeling results obtained at an earlier time may be reused at a later time. Such pre-computed results may be obtained from the bitstream, and the use of pre-computed diffraction data can avoid complex calculations at the decoder side.
[0149] FIG. 10(a) schematically illustrates an example of a possible use case for reusing precomputed diffraction data according to an embodiment of the present disclosure. Two listeners (users) A and B are shown in different positions within an audio scene (e.g., a house with different areas and levels). Users A and B may be users individually or collaboratively exploring a VR environment containing the audio scene, e.g., as part of a game, virtual tour, or the like. Users exploring a common VR environment may, for example, be running a social VR environment. Having different listener positions within the audio scene, users A and B generate different rendering results and different diffraction information. The present disclosure foresees each user (or their respective device / decoder / renderer) making the calculated diffraction information available to other users. Once user B enters an area of the audio scene where user A was previously present, he or she can benefit from user A's precomputed diffraction information, and vice versa. For example, user A's diffraction information may be made available to user B via a LUT that indexes various items of diffraction information with corresponding scene states or scene state identifiers. By exchanging diffraction information between different devices / decoders / renderers, depending on the user's movement pattern within the audio scene, the computational load for both user devices / decoders / renderers can be reduced.
[0150] Also, because users (listeners) tend to behave similarly, diffraction information (diffraction data) is accumulated, especially for relevant (e.g., frequently occurring) scene states. Achieving this with encoder-side precomputation of diffraction information is very difficult because the encoder does not have access to the actual listener positions and can only assume them. Furthermore, the use of data storage (e.g., physical / shared storage or bitstream bandwidth) would be much less efficient for encoder-side precomputation. This is due to the portion of the precomputed diffraction information pertaining to irrelevant or less relevant scene states in this case.
[0151] For example, the proposed features and techniques can create LUTs that correspond to the user's actual 6DoF behavior (not that assumed on the encoder side), and thus relate to smart user-oriented LUT creation.
[0152] FIG. 10(b) illustrates complexity metrics for various implementations of processing audio scene information or audio rendering as a function of time, assuming a simple maze as the audio scene. We further assume that the user moves randomly through the maze, thus revisiting previously visited locations. Graph 810 relates to the case where no precomputed diffraction information is available (e.g., no diffraction information is provided with the bitstream and memory / caching is disabled). In this case, the computational load on the renderer is substantially constant and relatively high. Graph 820 relates to the case where precomputed diffraction information is available locally (e.g., no diffraction information is provided with the bitstream and local memory / caching is enabled). In this case, the processing load on the renderer decreases over time as more and more diffraction information is accumulated locally. In other words, the scene states encountered are increasingly related to (locally) known scene states. Graph 830, finally, relates to the case where precomputed diffraction information is provided externally (e.g., complete diffraction information is provided with the bitstream). In this case, the computational load on the renderer is always low, since a significant portion of the scene state relates to a known scene state and the diffraction information can be retrieved externally (e.g., from the bitstream or on demand from external / shared storage) without local computation.
[0153] Furthermore, the diffraction modeling methods described above can be implemented on both mesh-based and voxel-based audio scene representations. In particular, diffraction modeling methods that work with voxel-based audio scene representations are computationally advantageous and can be used as an alternative method (e.g., running instead of the mesh-based method) for lower complexity / bitrate rendering modes. Alternatively, they can also be used as a supplementary method (e.g., running in parallel with the mesh-based method) for cases where mesh-based diffraction modeling cannot reliably compute diffraction data (in real time) for complexity-related reasons (e.g., requiring too many computational resources or unsupported higher-order diffraction modeling modes for some audio scene locations) or where mesh-based diffraction modeling cannot provide diffraction data (pre-computed on the encoder side) for encoding- or bitrate-related reasons (e.g., diffraction data is not available for some audio scene regions).
[0154] Therefore, the present disclosure proposes to use the results obtained from both methods to determine the most consistent and appropriate data set for diffraction modeling. Figure 11 shows a non-limiting example of rendering an indoor audio scene using the above-described approaches, i.e., reusing pre-computed diffraction data, mesh-based diffraction modeling, and voxel-based diffraction modeling, according to an embodiment of the present disclosure. Three virtual audio source objects (of diffraction modeling-related data) obtained by the corresponding approaches are indicated by point A for the pre-computed data, point B for the mesh-based data, and point C for the voxel-based data.
[0155] In this example, for an indoor scene consisting of several rooms, original sound sources S1 and S2, and an expected user path (P), diffraction data may be pre-computed (and stored in the bitstream) along the expected user path P between S1 and S2. Alternatively, the diffraction data may be computed (in real-time) using mesh-based diffraction modeling methods (around the sound sources, region S) and / or voxel-based diffraction modeling methods (for the rest of the scene, region R).
[0156] In this hypothetical audio environment as shown in Figure 11, two audio objects S1 and S2, e.g., a TV set located in Living Room 1 and a computer located in Bedroom 2, are shown as audio sources emitting sounds (i.e., audio signals). In concept, a user may move, e.g., from Living Room 1 to Bedroom 2, or vice versa, as defined / designed by a content creator. A scene description of this audio environment for audio rendering can be created by dividing the entire environment into multiple regions depending on the user's movement path. For example, it can be divided into the region of the user path P itself, a region S that the user can easily reach during movement (e.g., a likely area where the user may be present), and a region R that is the rest of the scene.
[0157] For rendering the region of the user path P (which may be, for example, a small predefined path), it may be beneficial to apply precalculated diffraction data, since the scene parameters may already have been generated and subsequently encoded in the bitstream for transmission. Thus, the diffraction data required for rendering do not need to be calculated in real time but can simply be obtained from the bitstream (i.e., previous results such as the diffraction path). In particular, object coordinates and gains may also be obtained from the diffraction path in the bitstream, allowing the rendering tool to be partially or even completely skipped. In other words, the predefined user path P may relate to a known scene state, for which rendering can be performed by providing precalculated diffraction information externally (via the bitstream), as described above, without local calculations by the rendering tool (i.e., reducing the computational load on the renderer).
[0158] Thus, once it is determined that the user is located within the region of the predefined path P (i.e., the scene state is known), the diffraction data is only required to be retrieved externally (e.g., from the bitstream or on demand from external / shared storage) while skipping the calculations in the rendering tool, which provides high-quality rendering with minimal computational effort. On the other hand, reusing precomputed data may require higher bandwidth for transmission.
[0159] For rendering region S, a mesh-based diffraction modeling method may be applied, i.e., rendering a mesh representation of the geometry in the audio scene. Note that this region may be limited by the diffraction modeling order, e.g., up to the third diffraction order. Thus, the calculated path may cross corner obstacles, and beyond this limit, it may become cost-inefficient for real-time calculations. Such a region may be determined by the scene creator and may be based on a trade-off between computational effort and the size of the scene elements. That is, fewer calculations are required for smaller sizes. It is understood that mesh-based diffraction modeling approaches can also provide high-quality rendering, albeit requiring complex local calculations by the rendering tool. On the other hand, such approaches do not require high bandwidth for transmission (i.e., increased bandwidth efficiency).
[0160] For rendering region R (i.e., regions that are not easily reachable by the user and are not part of the predefined user path), voxel-based diffraction modeling techniques may be applied. This is because voxels are not limited by diffraction orders, making it easier to calculate distant sounds with sufficient accuracy (e.g., modeling results based on azimuthal / horizontal planes are at least very representative, although they do not necessarily provide highly detailed diffraction information, including the vertical analysis described above). Therefore, voxel-based diffraction modeling techniques allow for simplified calculations by the rendering tool compared to mesh-based diffraction modeling, and can also increase bandwidth efficiency for transmission compared to reusing precomputed data. This is because, in this case, the diffraction information is not included in the bitstream but is calculated in real time at the decoder / renderer side. However, applying voxel-based techniques can still ensure sufficiently accurate rendering of the audio scene.
[0161] Thus, the above techniques (i.e., precomputed data reuse, and mesh-based and voxel-based diffraction modeling) may be combined with each other to render a complete audio scene. Depending on the scene content and / or the hardware / software resources available in the device / system, any combination of these techniques may be selected / determined to be applied in parallel or alternatively to render the same scene. Thus, the same scene can be rendered on different devices with various hardware / software requirements. For example, voxel-based modeling may be used for devices with low power usage, while mesh-based modeling may be used for devices with higher power usage. When rendering on a device that requires low power consumption but can receive data over an available channel with large transmission bandwidth, it may be beneficial to reuse precomputed data to have improved rendering quality.
[0162] Furthermore, depending on the capabilities of the device, it is also acceptable to use different modes to render the same scene on one device. For example, the device may be provided with modes that reuse pre-computed data, low complexity (power / computation), and / or low bitrate (reduced transmission bandwidth). These different modes may further be assigned to different parts of the scene, as described above. Also, note that for transition regions, interpolation may be applied, or specific regions may be defined to match properties (e.g., obstacles and boundaries) to ensure smooth rendering of the scene. Different implementations may depend on environmental changes.
[0163] 12 shows an example flowchart of a method 1200 for processing an audio scene in a rendering device for three-dimensional audio rendering according to an embodiment of the present disclosure. Method 1200 may be implemented in software, hardware, or a combination thereof, for example, in a renderer / decoder coupled to an AR / VR / MR / XR device as shown by processing chain 100, or more specifically, for example, by diffraction calculation block 40 and / or diffraction modeling tool 50 in processing chain 100. Method 1200 includes steps S1210-S1230 for providing sonified audio data (e.g., 3DoF sonifier data) to be further processed by other rendering stages.
[0164] The method 1200 includes receiving S1210 a three-dimensional audio scene and information about a sound source at a source position. The method includes determining S1220 a rendering mode of a rendering device. Additionally, the method includes determining S1230, for a given listener position, a virtual sound source at the virtual source position based on the source position to simulate the effect of acoustic diffraction by the three-dimensional audio scene on a source signal of the sound source at the source position.
[0165] In particular, said determination of virtual sound sources is based on one or more of the following: Pre-calculated data along pre-defined user paths; and Real-time data calculated using at least one of mesh-based diffraction modeling and voxel-based diffraction modeling depending on the given listener position and / or the determined rendering mode of the rendering device.
[0166] As described above, the rendering modes include a reduced-complexity mode and / or a bandwidth-efficient mode. The rendering modes may also include a high-quality rendering mode. If the rendering mode is determined to be a reduced-complexity mode, the virtual sound sources may be determined based on pre-calculated data along a pre-defined user path, or diffraction modeling may be applied to the three-dimensional audio scene using pre-calculated data along a pre-defined user path. Alternatively or additionally, the virtual sound sources may be determined based on real-time data calculated using voxel-based diffraction modeling depending on a given listener position.
[0167] On the other hand, if the rendering mode is determined to be a bandwidth-efficient mode, the virtual sound sources may be determined based on real-time data calculated using mesh-based diffraction modeling and / or voxel-based diffraction modeling, depending on the given listener position.
[0168] Additionally, in response to determining that the scene state of the three-dimensional audio scene is in a known state (e.g., within a predefined path P as shown in FIG. 11 ), a virtual sound source may be determined based on precalculated data along a predefined user path. As described above, the determination of the virtual sound source using precalculated data or real-time calculation may be based on the area / region to be rendered within the scene (i.e., the area / region where the listener is located). For certain regions with a high diffraction modeling order, e.g., higher than 3, voxel-based diffraction modeling may be used to determine the virtual sound source. Otherwise, mesh-based diffraction modeling may be applied for the real-time calculation of the virtual sound source.
[0169] Therefore, it is possible to select an appropriate implementation method for acquiring diffraction information to be rendered on one or more devices with different capabilities, thereby achieving reduced computational complexity and / or efficient use of transmission bandwidth for rendering audio scenes in complex, evolving 3D virtual environments in real time.
[0170] Mesh Base Voxel coordinate / index representation The following efficient representation of voxel indices may be applicable for transmission or storage of both voxel grid and diffraction map (VoxDataDiffractionMap) entries, for example. It may substitute for any fixed-length representation of voxel indices (voxel coordinates).
[0171] The following steps may be performed in the context of the proposed representation:
[0172] Step 1: Determine the amount of bits (i.e., number, count) needed for the current grid resolution / diffraction map dimension. For a 3D voxel grid and a 2D diffraction map, these numbers NbitsVox and NbitsMap, respectively, can be determined, for example, as follows: NbitsVox=ceil(log2(L*W*H-1)) NbitsMap=ceil(log2(L*W-1)) where L, W, and H (length, width, and height) are the dimensions of the voxel grid and the diffraction map. The values can be different for the voxel grid and the diffraction map.
[0173] Step 1 can be applied on both the encoder and decoder sides.
[0174] Step 2: The voxel index (x,y,z) and the diffraction map index (x,y) are mapped onto a packed representation index (Idx) and encoded using Nbits_vox and Nbits_map bits, respectively. In one embodiment, (x,y,z) are zero-based, and the packed representation index may range from 0 to L*W*H-1 for voxels and from 0 to L*W-1 for diffraction maps. The mapping from index (x,y,z) to packed representation index may be, for example, as follows: Idx(x,y,z)=(x-1)+((y-1)*L)+((z-1)*L*W)
[0175] Step 2 may be performed only on the encoder side.
[0176] In the above, a packed representation index is an index that can uniquely identify a voxel within a voxel grid or a diffraction map. In other words, voxels within a voxel grid may be assigned unique consecutive indices such that each voxel within the voxel grid can be uniquely identified by a single integer. Thus, a packed representation index may be used for any indication of a voxel location within a voxel grid or a two-dimensional map. In particular, a packed representation index may be used to indicate any voxel location referred to throughout this disclosure.
[0177] The assignment of unique indices to voxels may follow a predefined pattern, for example, the voxel grid may be scanned / traversed in the x, y, and z directions in that order to successively assign a unique index to each voxel.
[0178] The mapping from packed representation indices back to voxel and diffraction map indices may be, for example, something like: x=floor(Idx)%L+1 y=floor(Idx / L)%W+1 z=floor(Idx / L / W)%H+1 where % represents the modulo operator.
[0179] In line with the above, the bitstream payload element voxSceneDiffractionPreComputedPathData() given in Table 1 may use the following pseudocode: [Table 1] Here, voxDiffractionMapPosPackedS and voxDiffractionMapPosPackedE denote the packed expression indexes.
[0180] Entropy Coding Entropy coding methods can be applied to sequences of integers representing acoustic features, voxel grid coordinates, voxel grid indices, etc.
[0181] For example, the entropy coding method may be applied to the sequence of integers described above (representing the acoustic characteristics or audio rendering instruction set reference P_ID and the grid indices X1, Y1, Z1 and X2, Y2, Z2), or a packed representation thereof. Additionally, entropy coding may be applied to a sequence of integers derived from the aforementioned representation of diffraction path information.
[0182] Thus, allowing complex dynamic scenes to be created and efficiently encoded is a technical benefit and advantage of the voxel-based scene representation as described herein.
[0183] Device While methods and processing chains have been described above, it should be understood that the present disclosure also relates to apparatus (e.g., computing devices or generally devices having processing capabilities) for implementing these methods and processing chains (or generally techniques).
[0184] An example of such an apparatus 1300 is shown schematically in FIG. 13. The apparatus 1300 comprises a processor 1301 and a memory 1302 coupled to the processor 1301. The memory 1302 may store instructions for execution by the processor 1301. The processor 1301 may be adapted to implement processing chains described throughout this disclosure and / or to perform methods described throughout this disclosure (e.g., methods of processing audio scene information for audio rendering). The apparatus 1300 may receive inputs (e.g., audio scene descriptions, listener positions, etc.) and generate outputs (e.g., representations of diffraction information, etc.).
[0185] interpretation Aspects of the systems described herein may be implemented in a suitable computer-based audio processing network environment (e.g., a server or cloud environment) for processing digital or digitized audio files. Some of these systems may include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0186] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described as data and / or instructions embodied in various machine-readable or computer-readable media using any number of combinations of hardware, firmware, and / or in terms of their behavior, register transfers, logical components, and / or other characteristics. The computer-readable media on which such formatted data and / or instructions may be embodied include various forms of physical (non-transitory) non-volatile storage media, such as, but not limited to, optical, magnetic, or semiconductor storage media.
[0187] Specifically, it should be understood that the embodiments may include hardware, software, and electronic components or modules, which, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, those skilled in the art will recognize, based on reading this detailed description, that in at least one embodiment, electronic-based aspects may be implemented in software (e.g., stored on a non-transitory computer-readable medium) executable by one or more electronic processors, such as microprocessors and / or application-specific integrated circuits ("ASICs"). Thus, it should be noted that a number of hardware- and software-based devices and a number of different structural components may be utilized to implement the embodiments. For example, the computer-implemented neural networks described herein may include one or more electronic processors, one or more computer-readable media modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0188] While one or more implementations have been described in connection with specific embodiments by way of example, it is to be understood that the one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0189] It is also to be understood that the phraseology and terminology used herein is for purposes of description and should not be regarded as limiting. The use of "including," "comprising," or "having," and variations thereof, is intended to encompass the listed items and equivalents thereof as well as additional items. Unless otherwise specified or limited, the terms "mounted," "connected," "supported," and "coupled," and variations thereof, are used broadly and encompass both direct and indirect mounting, connecting, supporting, and coupling.
[0190] Itemized Exemplary Embodiments Various aspects and implementations of the present invention can be understood from the following enumerated example embodiments (EEE), which are not claims.
[0191] [EEE1] 1. A method of updating an audio scene for three-dimensional audio rendering, the method comprising: obtaining a voxelized representation of the audio scene, the voxelized representation including a set of voxels forming a connected geometric region on a voxel grid of the voxelized representation, the geometric region having a rectangular parallelepiped shape, the set of voxels including at least a first boundary voxel and a second boundary voxel that define the rectangular parallelepiped shape of the geometric region, and voxels in the geometric region sharing a common voxel characteristic; determining an updated pair of bounding voxels that defines an updated rectangular parallelepiped region for said geometric region in response to updating said audio scene; and / or determining updated voxel properties for voxels in the rectangular region; generating a representation of the audio scene based on the updated pair of boundary voxels and the set of voxels with the updated voxel properties; A method comprising: [EEE2] The method as recited in EEE1, further comprising determining at least the first boundary voxel and the second boundary voxel for the set of voxels from a plurality of voxels of the voxelized representation. [EEE3] The method of any one of EEE1 and EEE2, wherein the boundary voxels describe scene elements that are part of the audio scene to be updated, and wherein existing data relating to voxel properties for the scene elements is overwritten in dependence on the updated pair of boundary voxels and / or the updated voxel properties. [EEE4] 4. The method of any one of EEE1 to 3, wherein the updating of the audio scene is based on update conditions including one or more of time, user input, input from a presentation engine, or input from application logic. [EEE5] The method according to EEE4, wherein update attributes indicating the update conditions are included as metadata in a bitstream together with a compressed representation of the audio scene based on a determined set of voxels. [EEE6] The method of any one of EEE1 to EEE5, wherein the updating of the audio scene is further based on a trigger received from a user in real time. [EEE7] 10. An encoder having a processor and a memory coupled to the processor, the processor adapted to cause the encoder to perform a method according to any one of EEE1 to EEE6. [EEE8] 1. A method for decompressing a compressed voxel-based audio scene from a bitstream, the method comprising: receiving the bitstream including a voxelized representation of the audio scene, the voxelized representation including a plurality of voxels arranged in a voxel grid, each voxel having associated voxel properties; decoding a set of voxels forming a connected geometric region on the voxel grid and an indication that voxels in the geometric region share a first voxel characteristic; decoding a subset of the set of voxels associated with a scene subelement within the geometric region and an indication that the subset of the set of voxels is assigned a second voxel characteristic; generating an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels, overwriting voxels of the subset with the second voxel properties; A method comprising: [EEE9] 1. A method of processing an audio scene for three-dimensional audio rendering, the method comprising: receiving a voxelized representation of the audio scene, the voxelized representation including a set of voxels defining a connected geometric region on a voxel grid of the voxelized representation; decoding the received voxelized representation of the audio scene; applying at least one 3D smoothing filter to the decoded voxelized representation of the audio scene; A method comprising: [EEE10] The method of claim 8, wherein voxels in the geometric region share common voxel properties, the common voxel properties including a reference to material that describes obscurance properties associated with the set of voxels for the geometric region. [EEE11] The method of EEE10, further comprising applying the at least one 3D smoothing filter to one or more concealment coefficients indicative of the concealment characteristics. [EEE12] 12. The method of any one of EEE9 to 11, wherein the at least one 3D smoothing filter is defined based on data transmitted in a bitstream or is hard-coded in a renderer. [EEE13] 13. The method of any one of EEE9 to 12, wherein the at least one 3D smoothing filter is applied to a decoded set of voxels of the voxelized representation of the audio scene, and / or the at least one 3D smoothing filter is associated with a scene element identifier for a geometric region within the audio scene, and the at least one 3D smoothing filter is applied to the geometric region having that scene element identifier. [EEE14] The method of any one of EEE9 to 13, wherein the at least one 3D smoothing filter is applied depending on a user position. [EEE15] A non-transitory computer program product having instructions which, when executed by a processor, cause the processor to perform a method according to any one of EEE1 to EEE14. [EEE16] 1. A method of processing an audio scene for three-dimensional audio rendering, the method comprising: obtaining a scene configuration representing the audio scene, the scene configuration including a source position of an audio source, a given user position, and a scene description, the scene description including a voxel matrix for a voxelized representation of the audio scene and associated obscuration and diffraction coefficients; obtaining a two-dimensional (2D) projection map associated with the voxelized representation of the audio scene; determining a diffraction path between the audio source and the given user position based on the 2D projection map; obtaining auralization data for rendering the audio scene based on the result of the determination; and A method comprising: [EEE17] The method according to EEE16, wherein the 2D projection map is obtained by applying a projection operation to the voxelized representation of the audio scene, and the diffraction path is determined based on one or more 2D pathfinding algorithms using the 2D projection map for the voxelized representation of the audio scene. [EEE18] 18. The method of claim 16 or 17, wherein the 2D projection map is either received from a bitstream transmitted by an encoder or calculated in a renderer. [EEE19] 19. The method of any one of EEE16 to 18, wherein the 2D projection map is calculated in a renderer by applying filtering to the voxel matrix of the scene description based on the source position and / or the given user position. [EEE20] 20. The method of any one of EEE16 to 19, wherein the 2D projection map is obtained by selecting among at least one of a horizontal projection associated with a diffraction path and a vertical projection associated with the diffraction path, the horizontal projection relating to the voxelized representation by a horizontal projection operation projecting onto a horizontal plane and the vertical projection relating to the voxelized representation by a vertical projection operation projecting onto a vertical plane. [EEE21] The method of claim EEE20, wherein the selection is based on the route and / or direction of the diffraction path. [EEE22] 22. The method of any one of EEE16 to 21, wherein to determine the diffraction paths, each voxel of the voxel matrix is assigned a set of diffraction coefficients based on the geometry and / or respective material properties of an occluding object. [EEE23] The method of any one of EEE16 to 22, further comprising selecting a subset of voxels in the voxel matrix based on the determined diffraction path, the subset of voxels causing one or more changes in diffracted sound direction. [EEE24] The method according to EEE23 when citing EEE22, further comprising calculating coefficients of an EQ filter for the diffracted audio signal based on the selected subset of voxels, wherein the calculation of the coefficients of the EQ filter depends on at least one of the direct distance between the audio object and the listener, the length of the diffraction path, the diffraction coefficients of corresponding voxels, and the angle of change of direction of the diffraction path. [EEE25] 25. The method of any one of EEE16 to 24, further comprising applying diffraction modeling to the voxelized representation of the audio scene. [EEE26] 1. A method of processing an audio scene in a rendering device for three-dimensional audio rendering, the method comprising: receiving a three-dimensional audio scene and information about a sound source at a source location; determining a rendering mode for said rendering device; determining, for a given listener position, a virtual sound source at a virtual source position based on the source position to simulate an effect of sound diffraction by the three-dimensional audio scene on a source signal of the sound source at the source position, the determining of the virtual sound source comprising: Pre-computed data along pre-defined user paths, as well as real-time data calculated using at least one of mesh-based diffraction modeling and voxel-based diffraction modeling depending on the given listener position and / or the determined rendering mode of the rendering device; Based on one or more of the following stages and A method comprising: [EEE27] The method of claim EEE26, further comprising determining that a scene state of the three-dimensional audio scene is a known state, and in response determining the virtual sound source based on the pre-computed data along the pre-defined user path. [EEE28] 28. The method of claim 26 or 27, wherein the rendering modes include a reduced complexity mode and / or a bandwidth efficient mode. [EEE29] The method according to EEE28, when it is determined that the rendering mode is a reduced complexity mode, further comprising: determining the virtual sound source based on the pre-computed data along the pre-defined user path, or applying diffraction modeling to the three-dimensional audio scene using the pre-computed data along the pre-defined user path, and / or determining the virtual sound source based on the real-time data calculated using the voxel-based diffraction modeling in dependence on the given listener position. [EEE30] The method of claim 8, further comprising, when the rendering mode is determined to be a bandwidth efficient mode, determining the virtual sound source based on the real-time data calculated using the mesh-based diffraction modeling and / or the voxel-based diffraction modeling depending on the given listener position. [EEE31] determining a diffraction modeling order indicative of the complexity of the diffraction path for the given listener position; determining whether to use the mesh-based diffraction modeling or the voxel-based diffraction modeling based on the determined diffraction modeling order for the given listener position; 31. The method of any one of EEE26 to 30, further comprising: [EEE32] 8. The method of claim 31, further comprising: when it is determined that the diffraction modeling order for the given listener position is higher than a predefined diffraction order, determining the virtual sound source based on the real-time data calculated using the voxel-based diffraction modeling. [EEE33] 33. The method of any one of EEE26 to 32, further comprising, when it is determined that for the given listener position, the pre-calculated data along the pre-defined user path is not available or real-time data calculated using the mesh-based diffraction modeling is not available, determining the virtual sound source at the given listener position based on the real-time data calculated using the voxel-based diffraction modeling. [EEE34] A non-transitory computer program product having instructions which, when executed by a processor, cause the processor to perform a method according to any one of EEE26 to EEE33.
Claims
1. 1. A method of updating an audio scene for three-dimensional audio rendering, the method comprising: obtaining a voxelized representation of the audio scene, the voxelized representation including a set of voxels forming a connected geometric region on a voxel grid of the voxelized representation, the geometric region having a rectangular parallelepiped shape, the set of voxels including at least a first boundary voxel and a second boundary voxel that define the rectangular parallelepiped shape of the geometric region, and voxels in the geometric region sharing a common voxel characteristic; determining updated pairs of bounding voxels that define an updated rectangular region for the geometric region in response to updating the audio scene; and / or determining updated voxel properties for voxels in the rectangular region; generating a representation of the audio scene based on the updated pairs of boundary voxels and the set of voxels with the updated voxel properties; A method comprising:
2. The method of claim 1 , further comprising determining at least the first and second boundary voxels for the set of voxels from a plurality of voxels of the voxelized representation.
3. 3. The method of claim 1, wherein the boundary voxels describe scene elements that are part of the audio scene to be updated, and existing data relating to voxel properties for the scene elements is overwritten depending on the updated pair of boundary voxels and / or the updated voxel properties.
4. 4. The method of claim 1, wherein the updating of the audio scene is based on update conditions including one or more of time, user input, input from a presentation engine, or input from application logic.
5. 5. The method of claim 4, wherein update attributes indicative of the update conditions are included as metadata in a bitstream together with the compressed representation of the audio scene based on a determined set of voxels.
6. The method of claim 1 , wherein the updating of the audio scene is further based on a trigger received from a user in real time.
7. An encoder comprising a processor and a memory coupled to the processor, the processor adapted to cause the encoder to perform a method according to any one of claims 1 to 6.
8. 1. A method for decompressing a compressed voxel-based audio scene from a bitstream, the method comprising: receiving the bitstream including a voxelized representation of the audio scene, the voxelized representation including a plurality of voxels arranged in a voxel grid, each voxel having associated voxel properties; decoding a set of voxels forming a connected geometric region on the voxel grid and an indication that voxels in the geometric region share a first voxel characteristic; decoding a subset of the set of voxels associated with a scene subelement within the geometric region and an indication that the subset of the set of voxels is assigned a second voxel characteristic; generating an updated representation of the audio scene based on the set of voxels and the subset of the set of voxels, overwriting voxels of the subset with the second voxel properties; A method comprising:
9. 1. A method of processing an audio scene for three-dimensional audio rendering, the method comprising: receiving a voxelized representation of the audio scene, the voxelized representation comprising a set of voxels defining a connected geometric region on a voxel grid of the voxelized representation; decoding the received voxelized representation of the audio scene; applying at least one 3D smoothing filter to the decoded voxelized representation of the audio scene; A method comprising:
10. 10. The method of claim 9, wherein voxels in the geometric region share common voxel properties, the common voxel properties including a reference to a material that describes obscurance properties associated with the set of voxels for the geometric region.
11. The method of claim 10 , further comprising applying the at least one 3D smoothing filter to one or more concealment coefficients indicative of the concealment characteristics.
12. 12. The method of claim 9, wherein the at least one 3D smoothing filter is defined based on data transmitted in a bitstream or is hard-coded in a renderer.
13. 13. The method of claim 9, wherein the at least one 3D smoothing filter is applied to a decoded set of voxels of the voxelized representation of the audio scene, and / or the at least one 3D smoothing filter is associated with a scene element identifier for a geometric region within the audio scene, and the at least one 3D smoothing filter is applied to the geometric region having that scene element identifier.
14. 14. The method according to any one of claims 9 to 13, wherein the at least one 3D smoothing filter is applied depending on the user position.
15. A non-transitory computer program comprising instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 14.
16. 1. A method of processing an audio scene for three-dimensional audio rendering, the method comprising: obtaining a scene configuration representing the audio scene, the scene configuration including source positions of audio sources, a given user position, and a scene description, the scene description including a voxel matrix for a voxelized representation of the audio scene and associated obscuration and diffraction coefficients; obtaining a two-dimensional (2D) projection map associated with the voxelized representation of the audio scene; determining a diffraction path between the audio source and the given user position based on the 2D projection map; obtaining auralization data for rendering the audio scene based on the result of said determination; A method comprising:
17. 17. The method of claim 16, wherein the two-dimensional projection map is obtained by applying a projection operation to the voxelized representation of the audio scene, and the diffraction paths are determined based on one or more 2D path-finding algorithms using the 2D projection map for the voxelized representation of the audio scene.
18. The method of claim 16 or 17, wherein the 2D projection map is either received from a bitstream transmitted by an encoder or calculated in a renderer.
19. 19. The method of claim 16, wherein the 2D projection map is calculated in a renderer by applying filtering to the voxel matrix of the scene description based on the source position and / or the given user position.
20. 20. The method of any one of claims 16 to 19, wherein the 2D projection map is obtained by selecting among at least one of a horizontal projection associated with a diffraction path and a vertical projection associated with the diffraction path, the horizontal projection relating to the voxelized representation by a horizontal projection operation projecting onto a horizontal plane, and the vertical projection relating to the voxelized representation by a vertical projection operation projecting onto a vertical plane.
21. The method of claim 20 , wherein the selection is based on a route and / or a direction of the diffraction path.
22. 22. The method of claim 16, wherein to determine the diffraction path, each voxel of the voxel matrix is assigned a set of diffraction coefficients based on the geometry and / or respective material properties of an occluding object.
23. 23. The method of any one of claims 16 to 22, further comprising selecting a subset of voxels in the voxel matrix based on the determined diffraction path, the subset of voxels causing one or more changes in diffracted sound direction.
24. 24. The method of claim 23 when relying on claim 22, further comprising a step of calculating coefficients of an EQ filter for the diffracted audio signal based on the selected subset of voxels, wherein the calculation of the coefficients of the EQ filter depends on at least one of a direct distance between an audio object and a listener, a length of the diffraction path, the diffraction coefficients of corresponding voxels, and an angle of change of direction of the diffraction path.
25. 25. The method of any one of claims 16 to 24, further comprising applying diffraction modelling to the voxelised representation of the audio scene.
26. 1. A method of processing an audio scene in a rendering device for three-dimensional audio rendering, the method comprising: receiving a three-dimensional audio scene and information about a sound source at a source location; determining a rendering mode of the rendering device; determining, for a given listener position, a virtual sound source at a virtual source position based on the source position to simulate an effect of sound diffraction by the three-dimensional audio scene on a source signal of the sound source at the source position, the determination of the virtual sound source comprising: Pre-computed data along pre-defined user paths, as well as real-time data calculated using at least one of mesh-based diffraction modeling and voxel-based diffraction modeling depending on the given listener position and / or a determined rendering mode of the rendering device; Based on one or more of the following stages and A method comprising:
27. 27. The method of claim 26, further comprising determining that a scene state of the three-dimensional audio scene is a known state, and in response determining the virtual sound source based on the pre-computed data along the pre-defined user path.
28. 28. The method of claim 26 or 27, wherein the rendering modes include a reduced complexity mode and / or a bandwidth efficient mode.
29. 29. The method of claim 28, further comprising, when it is determined that the rendering mode is a reduced complexity mode, determining the virtual sound source based on the pre-calculated data along the pre-defined user path, or applying diffraction modeling to the three-dimensional audio scene using the pre-calculated data along the pre-defined user path, and / or determining the virtual sound source based on the real-time data calculated using the voxel-based diffraction modeling depending on the given listener position.
30. 29. The method of claim 28, when it is determined that the rendering mode is a bandwidth efficient mode, further comprising: determining the virtual sound source based on the real-time data calculated using the mesh-based diffraction modeling and / or the voxel-based diffraction modeling depending on the given listener position.
31. determining a diffraction modeling order indicative of the complexity of the diffraction path for the given listener position; determining whether to use the mesh-based diffraction modeling or the voxel-based diffraction modeling based on the determined diffraction modeling order for the given listener position; 31. The method of any one of claims 26 to 30, further comprising:
32. 32. The method of claim 31 , further comprising: when it is determined that the diffraction modeling order for the given listener position is higher than a predefined diffraction order, determining the virtual sound source based on the real-time data calculated using the voxel-based diffraction modeling.
33. 33. The method of claim 26, further comprising: when it is determined that for the given listener position, the pre-calculated data along the pre-defined user path is not available or real-time data calculated using the mesh-based diffraction modeling is not available, determining the virtual sound source at the given listener position based on the real-time data calculated using the voxel-based diffraction modeling.
34. 34. A non-transitory computer program comprising instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 26 to 33.