Multi-directional audio diffraction modeling for voxel-based audio scene representations

By retrieving pre-stored diffraction information from neighboring voxels at the listener's location, the multi-directional audio rendering direction is determined, solving the problem of voxel rendering discontinuity in VR, AR, MR, and XR environments, and improving the realism of audio rendering and user experience.

CN120858587APending Publication Date: 2025-10-28DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480014825.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-28
Filing Date
2024-02-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing voxel-based audio rendering methods in VR, AR, MR, and XR environments are computationally complex and cause discontinuities, affecting user experience, especially when audio rendering is unrealistic near occluded objects.

Method used

By retrieving pre-stored diffraction information from neighboring voxels at the listener's location, the multi-directional audio rendering direction is determined, avoiding frequent path searches. Pre-computation using DLUT data reduces computational burden.

Benefits of technology

It improves the realism and quality of audio rendering, reduces computational complexity, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120858587A_ABST
    Figure CN120858587A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method of processing audio scene information relating to a voxel-based audio scene represented by a two-dimensional projection map for rendering audio from an audio source at a source location to a listener at a listener location. One such method includes obtaining first diffraction information relating to an acoustic path between the source location and the listener location, the listener location associated with a first voxel of the projection map; determining first direction information indicating a first rendering direction based on the first diffraction information; retrieving, in a neighborhood of the first voxel, pre-stored second diffraction information relating to an acoustic path between the source location and a second voxel of the projection map; determining second direction information indicating a second rendering direction based on the second diffraction information; and outputting the first direction information and the second direction information for rendering. The disclosure further relates to corresponding devices, computer programs and computer-readable storage media.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 487,176, filed February 27, 2023, and European Patent Application No. 23,159,262.7, filed February 28, 2023, the entire contents of each of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to techniques for processing audio scene information for audio rendering. Specifically, this disclosure relates to voxel-based scene representation and audio rendering. Background Technology

[0004] The Moving Picture Experts Group (MPEG) is a working group alliance jointly established by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC) to set standards for media encoding, including audio encoding. MPEG is organized according to ISO / IEC SC 29, and the audio group is currently identified as Working Group (WG) 6. WG 6 is currently working on a new audio standard (also known as MPEG-I Immersive Audio, ISO / IEC 23090-4).

[0005] The new MPEG-I standard enables acoustic experiences from different viewpoints and / or perspectives or listening positions by supporting scenes and various movements around such scenes (e.g., using various degrees of freedom, such as three degrees of freedom (3DOF) or six degrees of freedom (6DoF)) in virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) applications). The 6DoF interactive extension is limited to 3DoF spherical video / audio experiences based on head rotation (tilt, yaw, and roll) to include translational movement (forward / backward, up / down, and left / right) to allow navigation in virtual environments (e.g., physically walking around a room), in addition to head rotation.

[0006] For audio rendering in VR, AR, MR, and XR applications, object-based methods have been widely used to represent complex auditory scenes as multiple individual audio objects, each associated with parameters or metadata defining the object's position / location and trajectory within the scene. Alternatively, audio rendering in such environments also utilizes higher-order high-fidelity stereo image replication (HOA). However, new uses for "voxels" in rendering audio scenes are being explored, such as their application in novel immersive audio experiences. Voxels used for audio rendering are related to media environments implemented in hardware and software, such as video games and / or VR, AR, MR, and XR environments.

[0007] A voxel is a spatial volume that has acoustic properties or is assigned to an audio rendering instruction. The voxel size can be configured as a parameter for the encoder, and it can be selected (manually or automatically) based on the level of detail of the scene geometry (e.g., in the range of 10cm to 1m).

[0008] Voxels used for audio rendering can be obtained in the following ways:

[0009] ● Voxelization (or transformation) of mesh-based scene representation

[0010] ● Scene representation used for scene generation (or even video rendering) (e.g., downsampling from smaller voxels).

[0011] However, conventional methods of using voxels to deliver realistic sound for user experiences (including experiences involving movement) in VR, AR, MR, and XR environments remain challenging and computationally complex.

[0012] Whenever any of the audio scene, user position L, or audio source position O changes, typical techniques for diffraction modeling in a 3D audio scene (e.g., for computer-mediated reality applications) require recalculation of the diffraction path and other diffraction information. For example, the diffraction path may change as the user and / or audio source moves through the 3D audio scene. Furthermore, the diffraction path may change when the audio scene itself changes (e.g., via an open or closed door, window, or similar object). Frequent recalculation of diffraction paths can be computationally expensive, requiring relatively powerful computing devices for implementing computer-mediated reality applications and / or, in some cases, adversely affecting the user experience.

[0013] Furthermore, typical techniques for diffraction modeling in 3D audio scenes employ single-path audio diffraction modeling (for voxel-based audio scene representations) and consider only one diffraction direction along the shortest path from the audio source location O to the listener location L. This can lead to... Figure 1 and Figure 2 The following issues are illustrated in the examples. These diagrams show two-dimensional projection mappings of voxel-based scene representations (e.g., corresponding to or derived from voxel-based scene representations). Occluded voxels 110, 210 are indicated by white squares, and the remaining voxels 120, 220 are empty voxels or (air voxels, typically in which sound can propagate freely). Audio directions 160, 260 toward the unoccluded audio source are indicated by short thin lines, while audio diffraction directions 170, 270 toward the virtual sound source determined by diffraction modeling are indicated by short thick lines. If sound waves can bypass obstacles from the side (or both sides), then there exists a region 130 near the corner of the obstacle or a region 140, 250 behind the obstacle, where the audio diffraction direction can change rapidly (and significantly) from one voxel to the next.

[0014] These large discontinuities in the diffraction orientation can be perceived when a listener moves from one voxel to another. They are particularly noticeable if the user performs a continuous change in listening position through natural head or body movement. In this case, the perceived discontinuities can adversely affect the realism and quality of the rendered audio output and degrade the listening experience.

[0015] The special cases of the problem at hand are represented by a scene with a single voxel occlusion or obstacle element, such as Figure 2 This is illustrated in the example. In this case, a single diffracted sound direction 270 only bypasses the obstacle 210 from one side. This can produce an unrealistic audio rendering effect because a listener positioned behind the obstruction 210 or obstacle expects to perceive sounds from both sides simultaneously.

[0016] Therefore, there is a need for improved techniques for diffraction modeling in two-dimensional or three-dimensional audio scenes (specifically, two-dimensional or three-dimensional audio scenes utilizing voxels). In particular, there is a need for such techniques to improve the realism and quality of rendered audio output for listeners in 3DoF / 6DoF environments. Furthermore, there is a need for techniques that do not increase the computational burden on the devices implementing these techniques (e.g., decoders, renderers). Summary of the Invention

[0017] In view of this need, this disclosure provides a method for processing audio scene information (specifically, voxel-based audio scene information) for audio rendering, an apparatus for processing audio scene information for audio rendering, a computer program, and a computer-readable storage medium having the features of the respective independent claims.

[0018] One aspect of this disclosure relates to a method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene (obtaining information usable for the rendering). The method may include obtaining first diffraction information associated with an acoustic path (e.g., a diffraction path) between the source location and the listener location within the audio scene. The listener location may be associated with (e.g., corresponding to or contained therein) a first voxel of the projection map. The method may further include determining first direction information indicating a first rendering direction (e.g., toward a first virtual sound source, embodied by a first azimuth angle) based on the first diffraction information. The method may further include retrieving second diffraction information associated with an acoustic path (diffraction path) between the source location and a second voxel of the projection map within the audio scene. The second diffraction information may be pre-stored diffraction information. Furthermore, the second voxel may be a voxel in a neighborhood (e.g., a predefined neighborhood) of the first voxel. For example, the (predefined) neighborhood can be defined by a neighborhood matrix centered on the first voxel. The method may further include determining second directional information indicating a second rendering direction (e.g., toward a second virtual sound source, represented by a second azimuth angle) based on the second diffraction information. Retrieving the second diffraction information may not involve applying a pathfinding algorithm. The method may further include outputting the first and second directional information for rendering. It should be understood that the first and second rendering directions may be (sufficiently) different from each other.

[0019] As configured above, the proposed method can provide additional audio rendering directions, which are specifically chosen to avoid abrupt changes in diffraction direction when the listener moves from one voxel to the next near an extended occlusion. Importantly, this is accomplished with very little computational overhead because the method appeals to pre-stored diffraction information of neighboring voxels of the listener's position voxel (i.e., the first voxel).

[0020] In some embodiments, the method may further include selecting the second voxel from the neighborhood of the first voxel according to a predefined selection rule.

[0021] In some embodiments, the neighborhood of the first voxel may be a predefined neighborhood relative to the first voxel, associated with a set of predefined voxels relative to the first voxel. Next, selecting the second voxel may involve sequentially selecting voxels from the set of predefined voxels according to a predefined selection order. Using a predefined selection order for all listener positions can result in a more consistent listener experience.

[0022] In some embodiments, selecting the second voxel may further include, for each selected voxel, determining whether pre-stored diffraction information is available for the selected voxel. Selecting the second voxel may further include, if the pre-stored diffraction information is available, retrieving the pre-stored diffraction information and determining a rendering direction for the selected voxel based on the retrieved diffraction information.

[0023] In some embodiments, selecting the second voxel may further include comparing the determined rendering direction with the first rendering direction. Selecting the second voxel may further include, if the difference between the determined rendering direction and the first rendering direction is greater than a predefined threshold (angle threshold), then the selected voxel is taken as the second voxel, and the determined rendering direction is taken as the second rendering direction.

[0024] In some embodiments, the method may further include determining first and second gains, respectively, associated with the first rendering direction and the second rendering direction, based on the spatial relationship between the first voxel and the second voxel. The method may further include outputting the determined first and second gains for rendering.

[0025] In some embodiments, the first and second gains can be determined based on whether the first and second voxels are laterally adjacent or diagonally adjacent in the projection map.

[0026] In some embodiments, the first and second gains can be determined based on a predefined Gaussian kernel, which may be centered on the first voxel.

[0027] In some embodiments, the method may include determining whether the first voxel is adjacent to an isolated occlusion object of the projection map. The method may further include determining second direction information indicating a second rendering direction based on the first rendering direction and the spatial relationship between the first voxel and the occlusion object (e.g., the azimuth angle of a vector pointing from the first voxel to the occlusion object).

[0028] In some embodiments, obtaining the first diffraction information may involve applying a pathfinding algorithm. Importantly, retrieving the second diffraction information may not involve applying a pathfinding algorithm, but may rely entirely on pre-stored diffraction information (e.g., in the form of a DLUT database as defined elsewhere in this disclosure).

[0029] In some embodiments, the method may further include outputting a representation of the first diffraction information for storage. Therefore, the first diffraction information may be part of pre-stored diffraction information for reuse in a later stage to obtain the first diffraction information and / or the second diffraction information.

[0030] In some embodiments, the diffraction information may include indications of corner voxels (diffraction corners) on the acoustic path (diffraction path) for which the diffraction path changes direction, and for which the straight line from the first voxel to the corner voxel is not obstructed. This may apply to both the first diffraction information and the second diffraction information.

[0031] In some embodiments, the diffraction information may further include an indication of the length of the acoustic path.

[0032] Another aspect of this disclosure relates to a method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene. The method may include determining whether a first voxel of the projection map associated with (e.g., corresponding to or including) the listener location is adjacent to an occluder voxel of the projection map. The method may further include, if the first voxel is adjacent to the occluder voxel, obtaining first diffraction information related to an acoustic path between the source location and the listener location within the audio scene. The method may further include determining first direction information indicating a first rendering direction based on the first diffraction information. The method may further include determining (e.g., calculating) second direction information indicating a second rendering direction based on the first rendering direction and the spatial relationship between the first voxel and the occluder voxel. Additionally, the method may include, prior to determining the second direction information, a check determining whether the first voxel is adjacent to a corner voxel indicated by the first diffraction information. This determination corresponds to checking whether acoustic diffraction is actually caused by the obscuring object.

[0033] As configured above, the proposed method can provide additional audio rendering directions to improve the realism of the rendered audio output when the listener approaches a single occluded object pixel. Importantly, this is accomplished without significantly increasing the computational overhead.

[0034] In some embodiments, the second rendering direction can be determined by rotating the first rendering direction by 90 degrees (i.e., π / 2).

[0035] In some embodiments, the rotation direction (i.e., rotation direction) for rotating the first rendering direction can be determined based on the first rendering direction and the direction from the first voxel to the occluding object.

[0036] In some embodiments, the second rendering direction may be determined such that the direction (e.g., orientation) from the first voxel to the occluding object pixel is within a sector (angular sector) spanned by the first rendering direction and the second rendering direction.

[0037] In some embodiments, determining whether the first voxel is adjacent to the isolation occlusion object voxel may include comparing the first voxel with a set of predefined voxels indicated to be adjacent to the isolation occlusion object voxel.

[0038] Another aspect of this disclosure relates to a method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene. The method may include determining path-search-based diffraction information associated with an acoustic path between the source location and the listener location within the audio scene by applying a path-search algorithm. The listener location may be associated with (e.g., corresponding to or contained therein) a first voxel of the projection map. Furthermore, the path-search-based diffraction information may include indications of corner voxels (diffraction corners) on the acoustic path (diffraction path), for which the diffraction path changes direction, and for which the straight line from the first voxel to the corner voxel is not occluded. The method may further include identifying a set of voxels intersecting a portion of the acoustic path extending between the first voxel and the corner voxel. The method may further include determining the derived diffraction information of each voxel in the set of identified voxels based on the path-search-based diffraction information. The method may further include outputting the path-search-based diffraction information and the derived diffraction information for storage. The diffraction information may be output to bitstreams, local storage, cloud-based storage devices, databases, files, etc.

[0039] By reusing computed diffraction information to derive the diffraction information of voxels near the listener voxel, the proposed method can very efficiently populate the database used to store pre-computed diffraction information without inherently incurring additional computational burden. Furthermore, when the listener moves near an obstruction or occluding object, the proposed method ensures sufficient pre-stored diffraction information for multi-directional diffraction modeling, as described throughout this disclosure.

[0040] In some embodiments, the derived diffraction information of each voxel in the set of identified voxels may indicate the same corner voxel as the path-search-based diffraction information.

[0041] In some embodiments, the path-search-based diffraction information may further include an indication of the length of the acoustic path. Next, determining the derived diffraction information for a given voxel of the identified voxel set may include determining the derived length of the acoustic path based on the length of the acoustic path and the distance between the first voxel and the given voxel.

[0042] According to another aspect, an apparatus is provided for processing audio scene information for audio rendering. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform all the steps of the method according to the foregoing aspects and embodiments thereof.

[0043] According to a further aspect, a computer program is described. The computer program may include executable instructions that, when executed by a computing device (e.g., a processor or a group of processors), are used to perform the methods or method steps outlined throughout this disclosure.

[0044] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted to execute on a computing device (e.g., a processor or a group of processors), and when implemented on the computing device, to perform a method or method steps outlined throughout this disclosure.

[0045] It should be noted that the methods, apparatus, and systems (including their preferred embodiments) outlined in this disclosure can be used alone or in combination with other methods, apparatus, and systems disclosed in this document. Furthermore, all aspects of the methods and systems outlined in this disclosure can be combined arbitrarily. Specifically, the features of the claims can be combined with each other in any manner.

[0046] It will be understood that the device features and method steps can be interchanged in various ways. Specifically, the details of the disclosed methods can be implemented by the corresponding device and vice versa, as will be understood by those skilled in the art. Furthermore, any statements made above regarding the methods (and, for example, their steps) should be understood to apply equally to the corresponding devices (and, for example, their blocks, levels, and units) and vice versa. Attached Figure Description

[0047] The invention is described below by way of example with reference to the accompanying drawings, in which...

[0048] Figure 1An example of a voxel-based audio scene is illustrated, where unidirectional audio diffraction is modeled and a large change in diffraction orientation can occur from one voxel to another.

[0049] Figure 2 An example of a voxel-based audio scene is illustrated, where unidirectional audio diffraction modeling can lead to unrealistic or unnatural audio output.

[0050] Figure 3 Illustrative illustration of embodiments according to the present disclosure Figure 1 An example of multi-directional audio diffraction modeling in an audio scene;

[0051] Figure 4 Illustrative illustration of embodiments according to the present disclosure Figure 2 An example of multi-directional audio diffraction modeling in an audio scene;

[0052] Figure 5 Examples of diffraction paths in a voxel-based audio scene according to embodiments of the present disclosure are illustrated, as well as voxels from which the diffraction paths can be derived.

[0053] Figure 6 This is a flowchart illustrating an example of a method for processing audio scene information according to embodiments of the present disclosure;

[0054] Figure 7 This is a flowchart illustrating another example of a method for processing audio scene information according to embodiments of the present disclosure;

[0055] Figure 8 Describe the neighborhood N of a single occlusion voxel K according to an embodiment of this disclosure. (K) Examples;

[0056] Figure 9A and Figure 9B Examples illustrating the diffraction path of a listener's location near an obscuring object element, to which the technology according to embodiments of this disclosure can be applied;

[0057] Figure 10 An example of the neighborhood N of the listener position voxel L according to an embodiment of the present disclosure is described;

[0058] Figure 11A and Figure 11B Further examples illustrate the diffraction path of a listener's location near an occluded object element, to which the technology according to embodiments of the present disclosure may be applied;

[0059] Figures 12A to 12E Examples of Gaussian kernels for gain determination according to embodiments of this disclosure are described;

[0060] Figure 13A and Figure 13BThis is a flowchart illustrating another example of a method for processing audio scene information according to embodiments of the present disclosure;

[0061] Figure 14 This is a diagram illustrating the complexity of different operating modes / implementations for processing audio scene information for audio rendering as a function of time, according to embodiments of the present disclosure.

[0062] Figure 15 Examples of voxel-based audio scenes to which embodiments of the present disclosure may be applied are illustrated.

[0063] Figures 16A to 16E Examples of multidirectional audio diffraction modeling for different values ​​of the angle threshold h are described according to embodiments of the present disclosure;

[0064] Figures 17A to 17C This describes instances of multipath diffraction directions in different audio scenarios when multidirectional audio diffraction modeling is applied according to embodiments of the present disclosure.

[0065] Figure 18 This describes an example of multipath diffraction directions in a simple maze test scenario when multidirectional audio diffraction modeling is applied according to embodiments of the present disclosure; and

[0066] Figure 19 An apparatus for implementing a method according to an embodiment of the present disclosure is illustrated schematically. Detailed Implementation

[0067] In the following description, exemplary embodiments of the present disclosure will be illustrated with reference to the accompanying drawings. Identical elements in the drawings may be indicated by identical element symbols, and repeated descriptions thereof may be omitted.

[0068] Voxel-based audio scene representation

[0069] First, an overview of voxel-related concepts for audio scene representation will be given.

[0070] What are voxels used for audio rendering?

[0071] A voxel is understood as a spatial volume that has acoustic properties or is assigned audio rendering instructions.

[0072] What is the voxel size used for audio rendering?

[0073] The voxel size is a configurable parameter for the encoder. It can be selected (manually or automatically) based on the level of detail in the scene geometry (e.g., in the range of 10cm to 1m).

[0074] What size audio scene can it handle?

[0075] Large audio scenes do not necessarily lead to a large number of voxels and high rendering complexity. For example, a large audio scene can be represented as follows:

[0076] ● A set of independent sub-scenes (and methods for "transferring" these representations without a renderer "restart").

[0077] ● A set of scene updates (based on user location)

[0078] How to handle the discontinuity problem caused by voxel granularity?

[0079] Any strong discontinuities in sound levels (and jumps in the direction of diffraction signals) can be avoided by applying interpolation (e.g., in time and space).

[0080] How to represent a voxel-based audio scene?

[0081] Any voxel-based representation of an audio scene may contain indications of voxels that are not transmission voxels (e.g., occlusion voxels) (i.e., voxels in which sound cannot propagate or cannot propagate freely) (representation of occlusion geometry). This indication may be associated with an indication of the coordinates of the corresponding voxel (e.g., center coordinates, corner coordinates, etc.). For example, the coordinates of these voxels may be represented by a grid index. Additionally, the voxel-based representation may contain indications of the material properties of voxels that are not transmission voxels, such as absorption coefficients, reflection coefficients, etc. Besides occlusion voxels, the voxel-based representation may also indicate transmission voxels (e.g., air voxels), i.e., voxels in which sound can propagate (representation of the sound propagation medium). Therefore, some embodiments of a voxel-based representation of an audio scene may include, for each voxel in a predefined spatial segment (e.g., within the boundary enclosing the audio scene), an indication of the corresponding material properties.

[0082] For example, according to the MPEG-I standard, a voxel-based audio scene description can be in the form of a voxSceneDiffractionMap() syntax element or a part thereof, as given in Table 1.

[0083] Table 1 - Syntax of voxSceneDiffractionMap()

[0084]

[0085] The syntax of escapedValue() in Table 1 should be as defined in ISO / IEC 23003-3.

[0086] The syntax elements in Table 1 can be defined as follows:

[0087] numberOfVoxDiffractionMapElements This element represents the number of block elements that define the voxel scene diffraction map.

[0088] voxDiffractionMapValue This element represents the voxel type ID of the diffraction map in the voxel block element.

[0089] The element `voxDiffractionMapPosPackedS` represents a compressed form of the variable `voxDiffractionMapPosS`, indicating the voxel index of the first (starting) voxel that defines the diffraction map block elements.

[0090] The element `voxDiffractionMapPosPackedE` represents a compressed form of the variable `voxDiffractionMapPosE`, indicating the voxel index of the second (end) voxel that defines the diffraction map block element.

[0091] Instance processing chain for processing audio scene information

[0092] In some embodiments, the techniques according to this disclosure may be performed relative to a processing chain for processing audio scene information for audio rendering. This processing chain can be used to convert voxel-related data into parameters and signals required for auditory processing (or, generally, audio rendering). The processing chain may be implemented in software, hardware, or a combination thereof. For example, the processing chain may be implemented by a renderer / decoder coupled to an AR / VR / MR / XR device (e.g., AR / VR / MR / XR goggles). Specific implementations may include game control panels, set-top boxes, personal computers, etc.

[0093] The processing chain receives an audio scene description from a bitstream (or storage device / memory). The audio scene description may include a representation of the three-dimensional audio scene (including, for example, its two-dimensional projection mapping) and information about the source locations of sound sources within the audio scene. For example, the representation of the three-dimensional audio scene may be voxel-based.

[0094] The processing chain further receives an indication of the user's (listener's) location (listener's position) in the audio scene. The audio scene description and user location can be provided to a diffraction direction calculation block (diffraction calculation block) for determining (e.g., calculating) diffraction information. The diffraction information can be correlated with the acoustic path (acoustic diffraction path) between the source location and the listener's location within the audio scene. The diffraction information can then be provided to a diffraction modeling tool for applying diffraction modeling and, optionally, occlusion modeling based on the diffraction information. Occlusion modeling calculates the attenuation gain of the straight line between the listener and the audio source. The diffraction modeling tool can output auditory data (3DoF auditory data), which includes, for example, the position, orientation, and / or gain (e.g., frequency-dependent gain) of the object to be rendered. The output of the diffraction modeling tool can be further processed by other rendering stages, such as Doppler, directivity, distance attenuation, etc. Generally, the output of the diffraction modeling tool can be referred to as diffraction information, as detailed below. The auditory data can then be used, for example, for audio playback.

[0095] In summary, the processing chain described above can be used to convert voxel-related data into parameters and signals for auditory processing. The diffraction direction calculation block and diffraction modeling tools can be considered as non-limiting instances of rendering tools. Generally, rendering tools can generate 3DoF auditory data.

[0096] For this processing chain, the method according to embodiments of this disclosure can be performed, for example, in a diffraction direction calculation block (diffraction calculation block) for determining (e.g., calculating) diffraction information. However, this disclosure should not be construed as limiting itself to such processing chains.

[0097] Example scenario description

[0098] Scene descriptions may include voxel matrices and associated coefficients (e.g., reflectance, occlusion, absorption, transmission, etc.). These coefficients may indicate the material or material properties of the corresponding voxels. Rendering tools may include, for example, occlusion and diffraction modeling tools. 3DoF auralizer data may include, for example, object position, orientation, and frequency-dependent gain.

[0099] Furthermore, the voxel-based representation of a 3D audio scene defines the psychoacoustic geometric elements and sound propagation media. In some implementations, the scene description may provide information to the rendering tool using parameters / interfaces (e.g., agreed-upon data formats or agreed-upon data exchange points):

[0100] Scene size:

[0101] -Using absolute units (e.g., meters)

[0102] -In units of voxel number and / or voxel size

[0103] Scene Anchor:

[0104] - Regarding coordinate anchors (mapping absolute coordinates to voxel indices)

[0105] - Regarding scene anchors (mapping sub-scenes to voxel subsets)

[0106] Scene content data:

[0107] -Refer to material properties (e.g., transmission coefficient, reflection coefficient, etc.) that approximate the acoustic effects caused by an obstruction (sound barrier) positioned within the corresponding volume.

[0108] -The properties of sound propagation media are approximated by acoustic effects (e.g., sound speed, energy absorption, distance attenuation curves, etc.) caused by the medium positioned in the corresponding volume.

[0109] - Rendering control parameters that describe the expected occlusion modeling effect

[0110] For example, the "global" or "local" occlusion type determines the length (and shape) of the occlusion effect shadow behind this voxel.

[0111] - Rendering control parameters describing the expected sound diffraction simulation effect

[0112] For example, voxel types that control / cause a change in the direction of sound (i.e., the path of diffracted sound cannot penetrate this volume).

[0113] -Content control parameters describing audio signal correlation and scene creation

[0114] For example, rendering control parameters that determine which signal is perceived in the corresponding volume, related to the (rendered) audio signal ID, and / or signal gain, describing the expected reverberation modeling effect.

[0115] - For example, the voxel type that controls reverberation settings (e.g., RT60, DDR, RIR, etc.).

[0116] Scene content update:

[0117] - Updated triggering events based on reference.

[0118] All data can depend on audio objects (to support the content creator's intent in flexible audio scene creation).

[0119] 3DoF vision data may include the following information:

[0120] - A set of parameters for an audio object and its associated signals (and HOA).

[0121] ○ The parameters include the metadata output of the rendering tool (i.e., the gain of position, orientation and simulated occlusion, diffraction, early reflection effects, reverberation coefficient parameters, IR, etc.).

[0122] ○ The associated signal represents the audio output of the rendering tool (i.e., down-mixing or copying of the audio signal).

[0123] - Scene state identifier (i.e., metadata that allows mapping scene descriptions and user inputs to 3DoF vision data)

[0124] Technology for processing audio scene information

[0125] Broadly speaking, this disclosure provides extensions and improvements to existing audio diffraction modeling methods for voxel-based audio scene representation. These extensions allow for the acquisition of additional diffraction sources to enable multi-directional audio diffraction modeling methods (i.e., supporting multiple audio diffraction paths around acoustic obstacles). This disclosure achieves multi-path diffraction modeling by modifying existing single-path diffraction modeling frameworks. This is accomplished computationally efficiently by querying pre-computed diffraction data from neighboring / nearby voxels at the listener's location. This method avoids performing additional path search processing steps and requires no additional data in the bitstream payload.

[0126] Based on the foregoing, this disclosure extends audio diffraction modeling to enable the generation of additional diffracted audio sources with dissimilar incoming directions. The techniques proposed by this disclosure only affect... Figure 1 The regions 130 and 140 (hereinafter referred to as case (A) regions) shown above exhibit discontinuities in the diffraction azimuth direction, and as shown in the figure. Figure 2 The region 250 shown is closely adjacent to the area of ​​a single voxel occlusion or obstacle (hereinafter referred to as the case (B) region).

[0127] For regions 130 and 140 (case (A)) that have discontinuities in the diffraction azimuth direction, according to the technique of this disclosure, for each affected voxel, additional [something is added]... Figure 3 In the example, the additional sound rendering direction 380 is indicated by a short dashed line. Furthermore, for regions 250 (case (B)) closely adjacent to a single voxel occlusion or obstacle area, according to the techniques of this disclosure, additional sound rendering is added for each affected voxel. Figure 4 In the example, the additional sound rendering direction is indicated by the short dashed line at 480.

[0128] The multi-directional audio diffraction modeling proposed in this disclosure is beneficial to the audio rendering quality and user experience of the aforementioned regions, namely, regions with more than one occlusion element (case (A)) and a single occlusion element (case (B)). For the remainder of the scene space, the default single-path diffraction mode may still be preferred to keep the audio rendering computational complexity unchanged and low. Therefore, the method according to embodiments of this disclosure may include a pre-check to determine whether the listener's position is in one of the regions of case (A) or case (B). Alternatively, as described above, the method according to embodiments of this disclosure may ensure that single-path diffraction modeling is applied to all other cases through appropriate checks within the diffraction modeling loop.

[0129] A feasible method for implementing multipath or multidirectional diffraction modeling schemes may involve performing an additional path search process. The calculation of additional audio diffraction sources may be based on the results of a first (“primary”) diffraction audio source. However, introducing an additional path search step significantly increases the computational complexity of the diffraction modeling algorithm. Therefore, this disclosure proposes an alternative method that does not require an additional path search step.

[0130] Specifically, the method according to this disclosure uses pre-computed (or realized) data stored, for example, in the form of a so-called Diffraction Lookup Table (DLUT). Only one diffraction path is considered for each voxel, and this path is used to extract all the data required for the proposed multipath diffraction modeling. The application of DLUT data (or, in general, pre-stored data) results in improved quality without a significant change in computational complexity, at the cost of additional local memory usage.

[0131] Next, examples of algorithms for processing audio scene information associated with voxel-based audio scenes will be described. The voxel-based audio scene is assumed to be represented by a two-dimensional projection map, which can be obtained from three-dimensional voxel-based scene information using techniques known to those skilled in the art (e.g., slicing or projection techniques). The algorithm is understood to provide the necessary or useful information for rendering audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene.

[0132] Based on the main concepts of this disclosure, the proposed algorithm is based on the "primary" diffraction audio source data D. primary (L) (also known as first diffraction information) is used to determine the "secondary" diffraction audio source data D of the listener's location L. secondary (L) (also known as second diffraction information). The diffraction information is determined based on whether it is case (A) or case (B), the "secondary" diffraction audio source data D. secondary (L) Based on the "primary" diffraction audio source data D primary (L) and the diffraction audio source data D of the neighboring voxel V of the listener's position L. primary(V) and threshold h to determine (case (A)), or based on the "primary" diffraction audio source data D primary (L), the listener's position L and the position K of the obstructing object are used to determine (case (B)).

[0133] Therefore, the "secondary" diffraction audio source data D secondary (L) is determined based on the following

[0134] A) "Primary" diffraction audio source data for listener positions adjacent to or near voxel V. primary (V) DLUT data:

[0135] D secondary (L):D secondary (D primary (L),D primary (V),h) for situation (A)

[0136] or

[0137] B) Listener position L and position of single voxel occlusion element K:

[0138] D secondary (L):D secondary (D primary (L), L, K) for situation (B)

[0139] Notably, pre-stored data (e.g., DLUT data) may contain pre-computed (or cached) information about the “primary” diffraction audio source. An implementation example of the contents of DLUT data is given below. “Primary” diffraction audio source data D primary This is information that can be retrieved from pre-stored data (e.g., DLUT data) (if available). Conversely, "minor" diffraction audio source data D secondary It is not part of the pre-stored data (e.g., DLUT data), but is calculated based on the pre-stored data during the audio rendering process.

[0140] The following definitions and symbols will be used in the remainder of this disclosure:

[0141] D primary (L) Data from the “primary” diffraction source of the listener voxel L (e.g., first diffraction information).

[0142] D primary (V) Data from the “primary” diffraction source of the listener’s neighboring voxel V

[0143] D secondary (L) Data from the “secondary” diffraction source of the receiver voxel L (e.g., second diffraction information).

[0144] asecondary (L) The location of the “secondary” diffractive audio source (e.g., the second rendering direction).

[0145] O Audio source location

[0146] L Listener position voxel (or listener position)

[0147] C L The corner voxel of the receiver position voxel L (diffraction corner).

[0148] a L L and C L The orientation between them (e.g., the first rendering direction), such as from L to C. L azimuth of the vector

[0149] g primary The gain of the "primary" diffracted audio source (e.g., the first gain).

[0150] V Listener location near voxel

[0151] C V The corner voxel (diffraction corner) of the voxel V∈N adjacent to the listener's position.

[0152] a V V and C V The orientation between them, for example, from V to C V azimuth of the vector

[0153] g secondary Gain of the "secondary" diffracted audio source (e.g., second gain).

[0154] N neighborhood (neighborhood matrix) includes all voxels surrounding the listener voxel L.

[0155] h "Primary" and "Secondary" diffraction azimuth difference thresholds

[0156] N (K) The neighborhood matrix (or, in general, the neighborhood) includes all voxels surrounding a single voxel occlusion element K.

[0157] K - Location of single voxel occlusion element

[0158] a K The orientation between L and K, for example, the azimuth angle of the vector pointing from L to K.

[0159] P projection mapping

[0160] r path Length of diffraction path on projection map

[0161] Next, processing steps performed by the method according to this disclosure after the listener's position or source position voxel is updated will be described. Similarly, these steps may be performed after an update to the audio scene description. These steps (i.e., steps 1 to 8 described below) form the algorithm according to this disclosure. However, it should be understood that some embodiments of this disclosure may not require all steps and may involve a subset of steps 1 to 8.

[0162] Step 1 :

[0163] The "primary" audio diffraction data D was obtained in the following manner. primary (L)(First Diffraction Information), including the C of the receiver's position voxel L L

[0164] - Retrieve "primary" audio diffraction data from pre-stored data (e.g., the DLUT dataset) C L ,or

[0165] - Run a pathfinding algorithm (e.g., JPS) and extend the pre-stored data of all voxels (e.g., the DLUT dataset) from the user's location voxel L to the corresponding corner voxel C. L As described below Figure 5 The examples are shown below.

[0166] And the diffraction path r can be obtained for the user's location L voxel. path The length of the data is used as a part of the diffraction data. It can be used, for example, to calculate the corresponding gain of an audio diffraction source.

[0167] Figure 5 This example illustrates an instance of the acoustic path (acoustic diffraction path) 30 between an audio source at location O10 and a listener at location L20 in a voxel-based audio scene represented by a two-dimensional projection map P. The projection map P indicates "air" voxels or empty voxels (i.e., voxels in which sound can propagate, or transmission voxels) 520 and occlusion voxels 510 (i.e., voxels in which sound cannot propagate or cannot propagate freely). Therefore, occlusion voxels 510 can be understood as related to voxels filled with materials other than air, and capable of reflecting, blocking, or otherwise altering sound propagation. For occlusion voxels 510, the voxel-based representation of the audio scene can further indicate the corresponding transmission, reflection, and potential absorption coefficients related to the material properties of these voxels. In the voxel-based representation, these coefficients can be linked to the ID or index of their respective voxels. Generally, a voxel-based representation defines the psychoacoustic geometric elements and sound propagation media in the audio scene.

[0168] As mentioned above, the (first) diffraction information (“primary” audio diffraction data D) primary(L)) Location C containing corner voxels L 40. It may further include acoustic paths r path The length.

[0169] When pre-stored data is not available for the "primary" audio diffraction data D primary (L) When the path search algorithm is used, it can be used to determine the acoustic path (diffraction path) between the source position 10 and the listener position 20. This path search algorithm takes the listener position 20, the source position 10, and a representation of the three-dimensional audio scene (e.g., a two-dimensional projection map or a two-dimensional matrix) as input. For example, an algorithm for determining diffraction information can take the listener position 20, the source position 10, and the representation of the three-dimensional audio scene as input, and output the diffraction corner angle (corner voxel) 40 (by C). L The position of the indicator and the variable r that optionally represents the length of the diffraction path. path For example, diffraction information can be determined based on the following:

[0170] [CL,r path ]=DiffractionDirectionCalculation(L,O,VoxDataDiffractionMap)

[0171] Where DiffractionDirectionCalculation indicates the algorithm used to determine diffraction information (“path-finding algorithm”), and VoxDataDiffractionMap indicates a voxel-based representation of the 3D audio scene or its processed version (e.g., a 2D projection map or 2D matrix derived therefrom). C L It is understood as the coordinates indicating the diffraction corner (e.g., the coordinates of the corresponding voxel containing the diffraction corner, voxel / mesh coordinates, or voxel / mesh index).

[0172] Here, DiffractionDirectionCalculation can refer to any feasible pathfinding algorithm, such as, for example, the fast traversal algorithm for ray tracing (see A Fast Voxel Traversal Algorithm for RayTracing, Amanatides, J. and A. Woo, Proceedings of EuroGraphics, 1987. 87.) and the JPS algorithm (see Online Graph Pruning for Pathfinding On Grid Maps, Harabor, DD and A. Grastien, Proceedings of the Twenty-Fifth AAAI Conference on Artificial Intelligence, 2011.). Furthermore, 2D pathfinding algorithms can be applied to this task using appropriate 2D projection planes (e.g., projection maps) based on 3D voxel-based scene representations.

[0173] The path search algorithm is assumed to output a diffraction path consisting of multiple straight path segments (line segments) that connect source position 10 to listener position 20 in an end-to-end sequential manner. Each transition from one path segment to another is associated with a change in the direction of the diffraction path.

[0174] According to the algorithm used to determine diffraction information, the diffraction corner angle C L A corner voxel can be identified as located on or near the diffraction path and adjacent to the diffraction map (indicated by the voxel-based representation) (in a set of voxels C representing the corner voxels on the diffraction map). set Voxels (in the middle). For example, a set of voxels (P) forming the diffraction path can be used. set Select the diffraction corner angle C L As a corner voxel near the "visible" (from the listener's position Lc) (belonging to C) set The voxels of ) thus lead to the path (P set Change direction. If there is more than one such corner, then choose to follow the diffraction path (P). set The corner closest to the listener's position.

[0175] Generally speaking, the diffraction path algorithm can be called the diffraction information related to the acoustic diffraction path between the source location and the listener location in the audio scene.

[0176] This diffraction information can provide sufficient information for the renderer to recover / determine the virtual source location of the virtual audio source, representing the effect of encapsulated acoustic diffraction. Diffraction angle C L Coordinates and diffraction path length r path This is the case. For example, when viewed from the listener's position, the virtual source position can be recovered by calculating the direction of the diffraction angle (e.g., azimuth, or azimuth and elevation). Using this direction and the path length r of the diffraction path... path The location of the virtual source can be determined by measuring the distance from the virtual source to the listener's location.

[0177] It should be noted that diffraction information can be represented in different ways. As mentioned above, one option is to include / store the path length r. path and diffraction corner angle (corner voxel) C L The diffraction information of the coordinates (e.g., grid coordinates).

[0178] Return to Figure 5 To enhance and accelerate the filling of pre-stored data (e.g., DLUT data), this disclosure proposes to also use diffraction information calculated for a given listener position L for those with the same diffraction corner angle (corner angle voxel C). L Voxels 530 were used to derive the diffraction information of these voxels.

[0179] Figure 6 This is a flowchart illustrating an example of a method 600 for enhancing and accelerating the filling of pre-stored data (e.g., DLUT data). Generally, method 600 is a method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to provide information for rendering audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene (e.g., information available for rendering). Method 600 includes steps S610 to S640, which can be performed whenever diffraction information is no longer available for a given listener location (or generally, a given combination of listener location, source location, and audio scene description (e.g., projection map)).

[0180] exist Step S610,Path-based diffraction information is determined by applying a path-search algorithm to the acoustic path (diffraction path) between the source location and the listener location within the audio scene. This can be accomplished, for example, as described above. The listener location is associated with (e.g., corresponding to or contained therein) a first voxel of the projection map. The determined path-based diffraction information includes indications of corner voxels (diffraction corners) on (or near) the acoustic path, for which the diffraction path changes direction, and for which, from the first voxel to the corner voxel (e.g., voxel C as defined above). L The straight line is not obscured.

[0181] exist Step S620, Identify a set of voxels that intersect with a portion of the acoustic path extending between the first voxel and the corner voxel. By definition, these are voxels that will produce the same corner voxel as the corner voxel determined for the first voxel in step S610 when the path search algorithm is applied to these voxels.

[0182] exist Step S630, Based on the path-search-based diffraction information, derived diffraction information is derived for each voxel in a set of identified voxels. As described above, the derived diffraction information for each voxel in a set of identified voxels will indicate the same corner voxel as the path-search-based diffraction information. However, the length of the corresponding diffraction path will be different (i.e., shorter).

[0183] Therefore, if the path-search-based diffraction information further includes an indication of the acoustic path length, then it is necessary to determine (e.g., calculate) the length of the acoustic path for a set of identified voxels. Thus, determining the derived diffraction information for a given voxel of a set of identified voxels may include determining the derived length of the acoustic path based on the length of the acoustic path (for the first voxel) and the distance between the first voxel and the given voxel. For example, the derived length of the acoustic path can be determined by subtracting the distance between the first voxel and the given voxel from the path length determined for the first voxel.

[0184] Finally, Step S640 The path-search-based diffraction information and the derived diffraction information are output for storage, for example, as part of pre-stored data (e.g., DLUT data). For example, the diffraction information can be output to a bitstream, local memory, cloud-based storage, database, file, etc. Therefore, the diffraction information determined in steps S620 and S630 can be reused for rendering in a later stage.

[0185] Next, further processing steps that may be performed after the listener location or source location voxel is updated according to the method of this disclosure will be described.

[0186] Step 2 :

[0187] For an assessment of whether implementation status (B) is applicable, please refer to [link / reference]. Figure 4 and Figure 9A , Figure 9B This is true if the following conditions are met simultaneously:

[0188] -L∈N (K) The listener's position voxel L belongs to the neighborhood N of a single voxel occlusion element K. (K) ,like Figure 8 China Showcase

[0189] -Voluxol L and C L Adjacent to each other, for example, |L–C L |<1.5 (in voxel size)

[0190] exist Figure 9A and Figure 9B The example illustrates the second of these conditions. The projection mapping again indicates the occluding object pixel 910 and the empty pixel 920. It has a diffraction corner angle C. L The diffraction path 30 of 40 extends between the source position O10 and the listener position L20, which is adjacent to the occlusion element K50 (isolation occlusion object element). Figure 9A and Figure 9B Of the two, the receiver position voxel L is adjacent to the diffraction corner voxel C. L This suggests that the audio diffraction is actually caused by the isolated occlusion object K50, and that case (B) applies. As described below... Figure 11A and Figure 11B The example shown illustrates a case where case (B) is not applicable. In this case, it is still necessary to check whether case (A) applies. Note that case (A) also applies when the listener's position L is adjacent to the occlusion element K, but case (B) does not. This is done through steps 3 through 7 described below.

[0191] If both conditions above are met simultaneously (i.e., case (B) applies), then the "minor" diffraction audio source information D secondary via direction a secondary Determined as:

[0192] a secon dary(L)=a L +sign(a L –a K π / 2

[0193] By corresponding to the azimuth vector a of the "main" diffraction audio source L (L and C) LThe orientation between them (first rendering direction) revolves around the vector a pointing from the listener position L to the position of the single voxel occlusion element K. K The vector a is determined by a π / 2 rotation of (the orientation between L and K). secondary .

[0194] In addition, the energy gain g of the "primary" diffraction audio source primary (First gain) and g of the "secondary" diffracted audio source secondary The second gain is calculated as follows:

[0195] g primary =2 / 3,

[0196] g secondary =1 / 3

[0197] Figure 7 This is a flowchart illustrating an example of a method 700 for processing audio scene information related to a voxel-based audio scene represented by a two-dimensional projection map. The method provides information for rendering audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene (e.g., for rendering purposes), consistent with the foregoing. Method 700 includes steps S710 to S740, which may be performed, for example, after an update to the source location, listener location, or audio scene description.

[0198] exist Step S710 Determining whether a first voxel of a projection map associated with a listener's location (e.g., corresponding to or containing the listener's location) is adjacent to an isolation occlusion object pixel of the projection map. For example, a list of isolation occlusion objects and / or a list of voxels adjacent to isolation occlusion objects may be pre-stored. Therefore, determining whether a first voxel is adjacent to an isolation occlusion object pixel may include comparing the first voxel with a predefined set of voxels indicated as adjacent to isolation occlusion objects.

[0199] Additionally, this step may include an extra check to determine whether the first voxel is adjacent to a corner voxel indicated by the first diffraction information. If so, it can be inferred that the audio diffraction of the audio source at the first voxel is actually caused by the isolated occluding voxel.

[0200] Therefore, step S710 can correspond to checking the two conditions defined in step 2 above.

[0201] exist Step S720 If the first voxel is adjacent to the isolated occlusion voxel (and optionally, if the additional check of step S710 is satisfied), then first diffraction information related to the acoustic path between the source location and the listener location within the audio scene is obtained.

[0202] exist Step S730Based on the first diffraction information, the first direction information indicating the first rendering direction is determined. The first rendering direction may correspond to the orientation vector a of the "primary" diffraction audio source defined above. L (L and C) L (The directions between them).

[0203] exist Step S740 The second rendering direction information, indicating the second rendering direction, is determined based on the first rendering direction and the spatial relationship between the first voxel and the occluding object pixel. The second rendering direction can correspond to the orientation 'a' defined above. secondary As described above, in one instance, the second rendering direction can be determined by rotating the first rendering direction by 90 degrees. The rotation direction (rotation direction) for rotating the first rendering direction can be determined based on the first rendering direction and the direction from the first voxel L to the occluding object voxel K. Specifically, the second rendering direction can be determined via the following...

[0204] a secon dary(L)=a L +sign(a L –a K π / 2

[0205] In any case, a second rendering direction should be determined such that the direction from the first voxel L to the occluding object voxel K (by orientation a) K The defined direction is within the sector spanned by the first rendering direction and the second rendering direction (e.g., see [link]). Figure 9A and Figure 9B ).

[0206] Next, the further processing steps performed after the listener location or source location voxel is updated according to the method of this disclosure will be described.

[0207] If the determination in step 2 produces a positive result (i.e., case (B) applies), then steps 3 through 7, as described below, are skipped, and the algorithm proceeds to step 8.

[0208] However, if the check in step 2 is not met, then the processing of application case (A) is as described in steps 3 to 7 below.

[0209] Figure 11A and Figure 11B This illustrates an example where step 2 produces a negative result and case (A) applies (or is applicable). The projection mapping again indicates the occluded object pixel 1110 and the empty voxel 1120. It has a diffraction corner angle C. L The diffraction path 30 of 40 extends between the source position O10 and the listener position L20, which in these instances (but not required for the application of case (A)) is adjacent to the occlusion element K50 (isolation occlusion object element). Figure 11A and Figure 11B In both cases, the listener position voxel L is not adjacent to the diffraction corner voxel, implying that case (B) is not applicable. In both instances, a neighboring voxel V 60 of the listener position voxel L is selected, and the diffraction information of this neighboring voxel V is obtained from pre-stored data (e.g., the DLUT database). This (second) diffraction information indicates the diffraction path of the neighboring voxel V, which includes the (second) diffraction corner C. V 70.

[0210] Step 3 :

[0211] From the listener's position voxel L neighborhood (by neighborhood N (e.g., Figure 10 The neighborhood matrix N described in the text specifies the selection of a neighboring voxel V∈N.

[0212] If all neighboring voxels have been processed by steps 3 through 6 without resulting in the definition of a "minor" diffraction audio source, then the process exits without the definition of a "minor" diffraction audio source.

[0213] The order in which neighborhood voxels V are selected from the neighborhood matrix N can be given by the following:

[0214] V(L)=V i ∈N(L)

[0215] For example, V = {(L x ,L y+1 ),(L x ,L y-1 ),(L x+1 ,L y ),(L x-1 ,L y (L) x+1 ,L y+1 ),(L x-1 ,L y+1 ),(L x+1 ,L y-1 ),(L x-1 ,L y-1 The following search order for )}:

[0216]

[0217] Based on this search order, the nearest voxels to the right of L will be selected first, then the nearest voxels to the left of L will be selected second, and the voxels above L will be selected third, and so on.

[0218] However, the above is merely an example, and any search order can be used, as long as it is predetermined and consistently used.

[0219] Step 4 :

[0220] Check if the pre-stored data (e.g., the DLUT dataset) contains diffraction data for neighboring voxels V (i.e., the corresponding corner voxels C). V And the path length from voxel V to audio object voxel O.

[0221] This suggests that the following conditions can be checked before searching the DLUT dataset:

[0222] - Neighboring voxel V within the projection mapping dimension P

[0223] - The neighboring voxel V has the sound transmission medium type (i.e., the air voxel) on the projection map P.

[0224] - The adjacent voxel V is not in the direct (i.e., unobstructed) line of sight from the listener to the audio source.

[0225] In fact, in some implementations, these implicit conditions can be checked before proceeding to diffraction modeling according to steps 2 through 7.

[0226] If there is no pre-stored data for a neighboring voxel V (or if the above conditions are not met), then proceed to step 3 to select the next voxel.

[0227] Step 5 :

[0228] Retrieve the corner voxel position C from pre-stored data of neighboring voxels V (e.g., the DLUT dataset). V Information. An approximate value for the diffraction path length r can also be obtained from pre-stored data. path This is used to calculate the final gain in the final rendering stage. Based on this, the gain from V to C can be determined. V Angle (azimuth, bearing) a V .

[0229] Step 6 :

[0230] Check if the difference in diffraction corner angles is greater than the threshold h:

[0231] |a V -a L |≥h

[0232] Where a L It's from L to C L Angle, and a V It's from V to C V The angle.

[0233] If the conditions are not met, proceed to step 3 to select the next element V.

[0234] If the conditions are met, then the "secondary" diffraction audio source D secondary via direction a secondary Determined as:

[0235] a secondary (L)=a V

[0236] Step 7 :

[0237] The energy gain g of the "main" diffraction audio source primary (First gain) and g of the "secondary" diffracted audio source secondary The second gain is calculated as follows:

[0238] g secondary =1 / 3, g primary =2 / 3, if |LV|==1 (in voxel size)

[0239] g secondary =1 / 5, g primary =4 / 5, if |LV|==2 (in voxels)

[0240] Here, |LV| is the so-called Manhattan metric, but other metrics with corresponding adaptations may also be used.

[0241] Generally, primary and secondary gains are determined based on the spatial relationship between the listener's position voxel L and its neighboring voxels V. In one instance, primary and secondary (or first and second) gains may be determined based on a predefined Gaussian kernel centered on the listener's position voxel L. Figure 12A Examples of Gaussian kernels for the 3×3, 5×5, and 7×7 neighborhoods of the listener's position L voxel are shown. Figure 12B and Figure 12C Explain how a 3×3 neighborhood Gaussian kernel can be used to determine the gain when the listener's position L voxel and neighboring voxels V are horizontally adjacent and diagonally adjacent, respectively. Figure 12D This explains why the 3×3 neighborhood Gaussian kernel is used to determine the gain expansion in the case of two neighboring voxels. Finally, Figure 12E This explains that the Gaussian kernel of the 3×3 neighborhood is used to determine the gain extension in the case of three neighboring voxels.

[0242] Figure 13AThis is a flowchart illustrating an example of a method 1300 for processing audio scene information related to a voxel-based audio scene represented by a two-dimensional projection map. The method is used to provide information for rendering audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, consistent with the above. Method 1300 includes steps S1305 to S1330, which may be performed, for example, after an update to the source location, listener location, or audio scene description.

[0243] exist Step S1305 First diffraction information is obtained relating to the acoustic path between the source location and the listener location within the audio scene. The listener location is associated with (e.g., corresponds to or is contained therein) a first voxel of the projection map. This first voxel corresponds to the listener location L voxel as defined above.

[0244] The first diffraction information (and similarly the second diffraction information retrieved in step S1315 below) includes the corner voxel (diffraction corner) C on the acoustic path. L (or C) V (Regarding the indication of the second diffraction information), for the corner voxel, the diffraction path changes direction, and for the corner voxel, the straight line from the first voxel to the corner voxel is not obstructed, and may further include the length r of the acoustic path. path Instructions.

[0245] The first diffraction information can be obtained by retrieving it from pre-stored data or (if unavailable) running a pathfinding algorithm. This can be done, for example, according to step 1 or method 600 described above. If the first diffraction information is determined using a pathfinding algorithm, a representation of the first diffraction information can be output for storage and later reuse, for example, as part of pre-stored data (e.g., a DLUT database). Thus, the pre-stored data is continuously expanded, thereby increasing the likelihood that the diffraction information is available for future rendering operations.

[0246] While path-finding algorithms can be used to determine the first diffraction information, retrieving the second diffraction information, as described below, does not involve applying any path-finding algorithm but relies strictly on pre-stored data.

[0247] exist Step S1310 Based on the first diffraction information, the first direction information indicating the first rendering direction is determined. The first rendering direction may correspond to the (azimuth) angle α defined in step 3 above. L .

[0248] exist Step S1315The process retrieves second diffraction information related to the acoustic path between the source location and the second voxel in the projection mapping within the audio scene. Importantly, the second diffraction information is pre-stored. The second voxel is understood as a voxel in the neighborhood of the first voxel. This neighborhood can be a predefined neighborhood, for example, defined by a neighborhood matrix centered on the first voxel. The second voxel may correspond to a voxel determined by successively selecting and checking neighboring voxels V, as described in steps 3 through 6 above.

[0249] As mentioned above, retrieving the second diffraction information does not involve applying a path search algorithm.

[0250] exist Step S1320 The second direction information, indicating the second rendering direction, is determined based on the second diffraction information. The second rendering direction may correspond to the (azimuth) angle α defined in step 6 above. V .

[0251] Based on the check performed in step 6, it should be understood that the first rendering direction and the second rendering direction are (sufficiently) different from each other.

[0252] exist Step S1325 (This can be optional), the first and second gains associated with the first and second rendering directions, respectively, are determined based on the spatial relationship between the first and second voxels. This can be done, for example, according to step 7 described above. That is, the first and second gains can be determined based on whether the first and second voxels are laterally adjacent or diagonally adjacent in the projection map. Furthermore, the first and second gains can be determined based on a predefined Gaussian kernel centered on the first voxel (having the same size as the predefined neighborhood of the first voxel).

[0253] In the steps S1330 The first and second direction information are output for rendering. Additionally, the first and second gains can be output for rendering at this step.

[0254] Method 1300 may further include the step of selecting a second voxel from the neighborhood of the first voxel according to a predefined selection rule. Figure 13B The details are described in the text. Figure 13B This is a flowchart illustrating an example of method 1350 related to an implementation of steps S1315 and S1320 of method 1300. Method 1350 includes steps S1360 to S1380. The selection of method 1350 may correspond to steps 3 to 6 described above.

[0255] exist Step S1360 Assuming the neighborhood of the first voxel is a predefined neighborhood relative to the first voxel and is associated with a set of predefined voxels relative to the first voxel, then voxels are selected sequentially from the set of predefined voxels according to a predefined selection order. For example, the selection order defined in step 3 above can be used.

[0256] exist Step S1365 For each selected voxel, determine whether the pre-stored diffraction information is available for the selected voxel. This corresponds to the check in step 4 above.

[0257] exist Step S1370 If the pre-stored diffraction information is available, then the pre-stored diffraction information is retrieved, and the rendering direction is determined based on the retrieved diffraction information. This corresponds to step 5 above.

[0258] exist Step S1375 The determined rendering direction is compared with the first rendering direction.

[0259] exist Step S1380 If the difference between the determined rendering direction and the first rendering direction is greater than a predefined threshold, then the selected voxel is used as the second voxel, and the determined rendering direction is used as the second rendering direction. Steps S1375 and S1380 can be performed according to the check in step 6 above.

[0260] If the selected voxel is found to be invalid in steps S1365 to S1380, that is, no valid second rendering direction result is generated (e.g., because the pre-stored diffraction information is unavailable or the check in step S1380 fails), then the next voxel in the neighborhood of the first voxel can be selected according to the predefined selection order, and steps S1365 to S1380 are performed for this next voxel, and so on.

[0261] Returning to the proposed algorithm, the algorithm can end in the final step (i.e., step 8).

[0262] Step 8 :

[0263] The rendering process continues based on the obtained audio diffraction source.

[0264] Processing during scene initialization or scene update

[0265] Next, we will describe the additional processing steps that can be performed during the scene initialization and scene update phases.

[0266] The following 3x3 kernel matrix comparison can be used during scene initialization and update phases to detect subsets of voxels (to address application scenario (B)):

[0267] if

[0268] The neighborhood N of a single voxel occlusion element K can be determined by evaluating the results of applying the filter kernel to the scene projection map. (K) .exist Figure 8 The description of the neighborhood N (K) Examples.

[0269] Control parameters of the secondary diffraction audio source

[0270] Next, the control parameters of the "secondary" diffraction audio source will be described.

[0271] This disclosure allows for the controlled determination of “secondary” diffraction audio sources.

[0272] ● For case (A), the generation of the "secondary" diffraction path is controlled by an adjustable parameter h. The threshold h determines how much the "secondary" diffraction direction must differ from the "primary" diffraction direction to be considered in audio rendering. For example, a suitable fixed setting for the h value is π / 6 (30 degrees).

[0273] ● For case (B), always calculate and consider the "minor" diffraction path. The difference between the "primary" and "minor" diffraction directions is always equal to π / 2 (90 degrees). Case (B) is considered separately from case (A) because in this case, the "minor" diffraction direction is always essential, and its calculation only requires consideration of the "primary" D. primary (L) Knowledge of diffraction audio sources.

[0274] exist Figures 16A to 16E The threshold h is shown in the middle. Figure 1 The effects of voxel-based audio scenes are shown in these figures. These figures illustrate the two-dimensional projection mapping of the voxel-based scene representation. Occluded voxels 1610 are indicated by white squares, along with the remaining voxels 1620 (air voxels). Audio directions 1660 toward unoccluded audio sources are indicated by short thin lines 1660, while audio diffraction directions 1670 toward virtual sound sources as determined by the multi-directional diffraction modeling proposed in this disclosure are indicated by short thick lines 1670 and short dashed lines 1680. Figures 16A to 16E These involve h = 0°, h = 10°, h = 20°, h = 30° and h > 180° respectively.

[0275] Pre-stored diffraction information

[0276] Pre-stored diffraction information may be in the form of DLUT data. Details are described below. However, these details are not limited to DLUT data, but apply to all forms of pre-stored diffraction information used for the purposes of this disclosure.

[0277] DLUT data content

[0278] Pre-computed (or realized) DLUT data contains:

[0279] ● Projection mapping Represents (or voxel scene identifier)

[0280] ● Diffraction path Starting position (Corresponding voxel L index)

[0281] ● Diffraction path Finish line (Corresponding voxel O index)

[0282] ● Diffraction corner (i.e., more distant voxel C) L As can be seen from the starting position L, the diffraction path changes its direction around the obstacle.

[0283] ●Total diffraction path trajectory Path length Approximation (i.e., the path from the start point to the end point around all the occluding elements that cause diffraction effects on the projected map).

[0284] Examples of detailed bitstream syntax are given in Tables 1 and 2.

[0285] DLUT data source

[0286] Pre-computed (or cashed-out) DLUT data is available at:

[0287] ● encoder The content of the DLUT data (transmitted in a bitstream):

[0288] - Projection mapping:

[0289] ■Custom definitions (e.g., using proprietary projection mapping creation tools)

[0290] - Remaining DLUT data:

[0291] ■Custom definitions (e.g., using a proprietary path finder)

[0292] ● renderer The content of the DLUT data (and realized in local storage):

[0293] - Projection mapping:

[0294] ■Default value For example During scene initialization (or scene update phase) use "Slicing and cutting" method

[0295] - Remaining DLUT data:

[0296] The "primary" audio diffraction path calculation performed at the audio renderer.

[0297] ■If the user visits the corresponding voxel (required for the current rendering output)

[0298] ■ If the rendering application triggers diffraction path calculation, and the calculation resource is available (not required for the current render output).

[0299] Encoder-side DLUT data calculation

[0300] A straightforward approach to providing DLUT data to the renderer is to pre-compute audio diffraction paths from all sound source locations to all voxels accessible to the user in the scene. This can be done before audio rendering at the encoder side. However, it is impractical to put every possible diffraction path data into the bitstream because this straightforward approach results in a high computational workload on the encoder and a large amount of data related to diffraction modeling.

[0301] It is more advantageous to pre-compute diffraction paths on the encoder side for only a subset of all possible user locations (and audio object locations). This subset can be semi-automatically determined and controlled by the content creator (with knowledge of the use case content and the anticipated points of interest for the user). That is, the encoder application can attempt to predict the most likely user location and only put the relevant data into the DLUT payload.

[0302] DLUT data calculation on the renderer side

[0303] Pre-compiling DLUT data on the renderer side has the following advantages:

[0304] ● Diffraction modeling tools Small position flow Size (e.g., empty initial DLUT data)

[0305] ● The obtained DLUT data is user-dependent and corresponds to the data rendered by the user during scene rendering. Actual visit The voxels (i.e., the area that the user is interested in).

[0306] ●During scene rendering time, new data Continuous revision DLUT data improves DLUT scene space coverage, while also improving audio rendering quality and reducing computational complexity.

[0307] exist Figure 14 The example illustrates the requirements for scene rendering time. Figure 14This diagram illustrates the complexity of different implementations of processing audio scene information or audio rendering as a function of time, assuming a simple maze as the audio scene. It further assumes that the user moves randomly through the maze, thus revisiting previously visited locations. Figure 1410 addresses the case where pre-computed diffraction information is unavailable in any way (e.g., the bitstream does not provide diffraction information, and memory / caching is disabled). In this case, the computational load on the renderer is substantially constant and relatively high. Figure 1420 addresses the case where pre-computed diffraction information is available locally (e.g., the bitstream does not provide diffraction information, and local memory / caching is enabled). In this case, the processing load on the renderer decreases over time because more and more diffraction information items accumulate locally. In other words, more and more scene states encountered will be associated with (locally) known scene states. Finally, Figure 1430 addresses the case where pre-computed diffraction information is provided externally (e.g., the bitstream provides complete diffraction information). In this scenario, the computational load on the renderer remains low because a significant portion of the scene state is related to a known scene state, and diffraction information can be retrieved externally (e.g., from a bitstream or by requesting from external / shared memory) without requiring local computation.

[0308] Combination methods

[0309] This disclosure proposes an optimized combination of these two DLUT data generation methods (or, in general, pre-stored data generation methods). Specifically, content creators can define pre-designed DLUT data entries in the bitstream (using scene knowledge and user intent), see Figure 11. For example, according to step 1 or method 600 described above, additional DLUT data is continuously calculated and added by the running audio renderer (based on the listener's movement and interaction with the scene) during scene rendering time.

[0310] Figure 15 This describes an instance of an environment with a subspace 1520 associated with pre-computed DLUT data from the encoder and a content creator's expected user action path 1510 representing scene geometry elements as represented by projection maps. For the remaining subspace 1530, no pre-designed LUT data entries are provided in the bitstream, and additional DLUT data is continuously computed by the renderer and added to the DLUT data during scene rendering time.

[0311] (Pre-stored) diffraction information (e.g., DLUT data) may be stored as a portion of the voxSceneDiffractionPreComputedPathData() syntax element according to ISO / IEC 23090-4 (Encoded representation of immersive media – Part 4: MPEG-I immersive audio, https: / / www.iso.org / standard / 84711.html) or any future standard derived therefrom. Table 2 shows the voxSceneDiffractionPreComputedPathData() syntax element according to the MPEG-I standard.

[0312] Table 2 - Syntax of voxSceneDiffractionPreComputedPathData()

[0313]

[0314] `voxSceneDiffractionPreComputedPathData()` is a bitstream syntax for parsing a bitstream and retrieving pre-computed (stored) diffraction information. This voxel payload data structure can have the following elements:

[0315] The numberOfVoxDiffractionPathData element represents the number of pre-computed diffraction path datasets.

[0316] The `voxDiffractionPathStartVoxelPacked` element represents a compressed form of the variable `voxDiffractionPathStartVoxel`, which indicates the voxel index of the path-starting voxel for the pre-computed diffraction path.

[0317] The `voxDiffractionPathEndVoxelPacked` element represents a compressed form of the variable `voxDiffractionPathEndVoxel`, which indicates the voxel index of the path-ending voxel of the pre-computed diffraction path.

[0318] The voxDiffractionPathDataExistFlag element indicates whether the diffraction path exists.

[0319] The `voxDiffractionSourceDirectionPacked` element represents a compressed form of the variable `voxDiffractionSourceDirection`, which indicates the voxel index used to determine the orientation value of the diffraction source.

[0320] The `voxDiffractionPathLength` element represents the length of the diffraction path on the 2D matrix of the diffraction map.

[0321] voxSceneDimensions: The number of voxels in each scene dimension.

[0322] The `escapedValue()` element implements a general method for transmitting integer values ​​using different bit counts. It features a two-level skip mechanism that allows the representable range of the value to be extended by transmitting additional bits consecutively. The syntax of `escapedValue()` should be as defined in ISO / IEC 23003-3.

[0323] Therefore, an additional technical benefit and effect of the technology according to this disclosure is that, if the scene state has been processed accordingly and diffraction information or 3DoF viewer data is available, then the scene state identifier or other information (scene description and source location) derived from this scene state can be used to avoid applying diffraction modeling tools or rendering tools. In this case, the renderer can access diffraction information / 3DoF viewer data (for a known scene state) by reusing data that was pre-computed before scene rendering time (pre-computed) or computed at scene rendering time, without applying rendering tools.

[0324] Equipment for implementing the method according to this disclosure

[0325] Although methods and processing chains have been described above, it should be understood that this disclosure also relates to equipment (e.g., computer equipment or, in general, processing-capable equipment) for implementing these methods and processing chains (or, in general, techniques).

[0326] exist Figure 19 An example of this device 1900 is illustrated schematically. Device 1900 includes a processor 1901 and a memory 1902 coupled to the processor 1901. Memory 1902 may store instructions executed by the processor 1901. The processor 1901 may be adapted to implement a processing chain throughout the description of this disclosure and / or perform methods throughout the description of this disclosure (e.g., methods for processing audio scene information for audio rendering, such as...). Figure 6 Method 600 Figure 7 Method 700 and / or Figure 13A , Figure 13B Methods 1300 and 1350). Device 1900 may receive input 1930 (e.g., audio scene description, listener location, etc.) and generate output 1940 (e.g., direction information, gain, etc.).

[0327] Simulation results

[0328] Figures 17A to 17CExamples of multipath diffraction directions of audio source 1710 in different audio scenarios are shown when multi-directional audio diffraction modeling is used, according to embodiments of the present disclosure.

[0329] Figure 18 An example of the multipath diffraction direction of an audio source 1810 in a simple maze test scenario is shown when multi-directional audio diffraction modeling is used, according to an embodiment of the present disclosure.

[0330] explain

[0331] Aspects of the systems described herein can be implemented in a suitable computer-based sound processing network environment (e.g., a server or cloud environment) for processing digital or digitized audio files. Parts of these systems may include one or more networks comprising any desired number of individual machines, and one or more routers (not shown) for buffering and routing data transmitted between computers. This network may be built on various network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0332] One or more of the components, blocks, processes, or other functional components may be implemented by a computer program executed by a processor-based computing device controlling the system. It should also be noted that the various functions disclosed herein may be described using hardware, firmware, and / or any number of combinations of data and / or instructions embodied in various machine-readable or computer-readable media (in terms of their behavior, register transfers, logical components, and / or other characteristics). Such computer-readable media embodying this formatted data and / or instructions include, but are not limited to, various forms of physical (non-transitory), non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0333] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules, which may be described and illustrated for purposes of discussion as if most components were implemented solely in hardware. However, those skilled in the art, upon reading this detailed description, will recognize that in at least one embodiment, the electronic aspects may be implemented in software (e.g., stored on a non-transitory computer-readable medium) executable by one or more electronic processors (e.g., microprocessors and / or application-specific integrated circuits (“ASICs”)). Therefore, it should be noted that embodiments may be implemented using multiple hardware and software-based devices and multiple different structural components. For example, a computer-implemented neural network described herein may include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., system buses) connecting the various components.

[0334] While one or more embodiments have been described by way of example and with regard to particular embodiments, it should be understood that one or more embodiments are not limited to the disclosed embodiments. Rather, they are intended to cover various modifications and similar arrangements that will be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation to cover all such modifications and similar arrangements.

[0335] Furthermore, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered restrictive. The use of “comprising,” “including,” or “having,” and variations thereof, is intended to cover the items listed thereafter and their equivalents, as well as additional items. Unless otherwise specified or limited, the terms “installation,” “connection,” “support,” and “coupling,” and variations thereof, are used extensively and cover both direct and indirect installation, connection, support, and coupling.

[0336] Examples and Implementations

[0337] Various aspects of the embodiments of the invention can also be understood from the following enumerated examples (EEE), which are not claims.

[0338] EEE1. A method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, the method comprising:

[0339] Obtain first diffraction information related to the acoustic path between the source location and the listener location within the audio scene, wherein the listener location is associated with a first voxel of the projection map;

[0340] Based on the first diffraction information, first direction information indicating the first rendering direction is determined;

[0341] Retrieve second diffraction information related to the acoustic path between the source location within the audio scene and the second voxel of the projection mapping, wherein the second diffraction information is pre-stored diffraction information, and wherein the second voxel is a voxel in the neighborhood of the first voxel;

[0342] Based on the second diffraction information, determine the second direction information indicating the second rendering direction; and

[0343] The first direction information and the second direction information are output for rendering.

[0344] EEE2. The method according to EEE1 further includes selecting the second voxel from the neighborhood of the first voxel according to a predefined selection rule.

[0345] EEE3. According to the method of EEE2, wherein the neighborhood of the first voxel is a predefined neighborhood relative to the first voxel, associated with a set of predefined voxels relative to the first voxel; and

[0346] Selecting the second voxel includes:

[0347] Voxels are selected sequentially from the set of predefined voxels according to a predefined selection order.

[0348] EEE4. The method according to EEE3, wherein selecting the second voxel further comprises:

[0349] For each selected voxel, determine whether the pre-stored diffraction information can be used for the selected voxel;

[0350] If the pre-stored diffraction information is available, then the pre-stored diffraction information is retrieved, and the rendering direction is determined based on the retrieved diffraction information.

[0351] EEE5. The method according to EEE4, wherein selecting the second voxel further includes:

[0352] Compare the determined rendering direction with the first rendering direction; and

[0353] If the difference between the determined rendering direction and the first rendering direction is greater than a predefined threshold, then the selected voxel is used as the second voxel, and the determined rendering direction is used as the second rendering direction.

[0354] EEE6. The method according to any one of EEE1 to EEE5, further comprising determining a first and a second gain respectively associated with the first rendering direction and the second rendering direction based on the spatial relationship between the first voxel and the second voxel.

[0355] EEE7. According to the method of EEE6, the determination of the first and second gains is based on whether the first and second voxels are laterally adjacent or diagonally adjacent in the projection map.

[0356] EEE8. The method according to EEE6 or EEE7, wherein the first and second gains are determined based on a predefined Gaussian kernel.

[0357] EEE9. The method according to any one of EEE1 to EEE8, comprising:

[0358] Determine whether the first voxel is adjacent to the isolated occlusion voxel of the projection map; and

[0359] If the first voxel is adjacent to the isolated occluding object, then second direction information indicating the second rendering direction is determined based on the first rendering direction and the spatial relationship between the first voxel and the occluding object.

[0360] EEE10. The method according to any one of EEE1 to EEE9, wherein obtaining the first diffraction information involves applying a path search algorithm.

[0361] EEE11. The method according to any one of EEE1 to EEE10, further comprising outputting a representation of the first diffraction information for storage.

[0362] EEE12. The method according to any one of EEE1 to EEE11, wherein the diffraction information includes an indication of a corner voxel on the acoustic path, for which the diffraction path changes direction, and for which the corner voxel, the straight line from the first voxel to the corner voxel is not obstructed.

[0363] EEE13. According to the method of EEE12, wherein the diffraction information further includes an indication of the length of the acoustic path.

[0364] EEE14. A method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, the method comprising:

[0365] Determine whether a first voxel of the projection map associated with the listener's location is adjacent to an isolated occlusion voxel of the projection map;

[0366] If the first voxel is adjacent to the isolated occlusion voxel, then first diffraction information related to the acoustic path between the source location and the listener location in the audio scene is obtained;

[0367] Based on the first diffraction information, first direction information indicating the first rendering direction is determined; and

[0368] The second direction information indicating the second rendering direction is determined based on the first rendering direction and the spatial relationship between the first voxel and the occluding voxel.

[0369] EEE15. The method according to EEE14, wherein the second rendering direction is determined by rotating the first rendering direction by 90 degrees.

[0370] EEE16. The method according to EEE15, wherein a rotation direction for rotating the first rendering direction is determined based on the first rendering direction and the direction from the first voxel to the occluding object pixel.

[0371] EEE17. The method according to EEE15 or EEE16, wherein the second rendering direction is determined such that the direction from the first voxel to the occluding object voxel is within the sector spanned by the first rendering direction and the second rendering direction.

[0372] EEE18. The method according to any of EEE14 to EEE17, wherein determining whether the first voxel is adjacent to the isolation occlusion object voxel includes comparing the first voxel with a set of predefined voxels indicated to be adjacent to the isolation occlusion object voxel.

[0373] EEE19. A method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, the method comprising:

[0374] The path search-based diffraction information is determined by applying a path search algorithm to determine the acoustic path between the source location and the listener location within the audio scene, wherein the listener location is associated with a first voxel of the projection map, and wherein the path search-based diffraction information includes an indication of a corner voxel on the acoustic path, for which the diffraction path changes direction, and for which the straight line from the first voxel to the corner voxel is not occluded.

[0375] Identify a set of voxels that intersect with a portion of the acoustic path extending between the first voxel and the corner voxel;

[0376] Based on the path-search-based diffraction information, the derived diffraction information of each voxel in the set of identified voxels is determined; and

[0377] The path-search-based diffraction information and the derived diffraction information are output for storage.

[0378] EEE20. According to the method of EEE19, wherein the derived diffraction information of each voxel in the set of identified voxels indicates the same corner voxel as the path-search-based diffraction information.

[0379] EEE21. The method according to EEE20, wherein the path-search-based diffraction information further includes an indication of the length of the acoustic path; and

[0380] Determining the derived diffraction information of a given voxel of the identified voxels includes determining the derived length of the acoustic path based on the length of the acoustic path and the distance between the first voxel and the given voxel.

[0381] EEE22. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to perform the method according to any one of EEE1 to EEE21.

[0382] EEE23. A program comprising, when executed by a processor, instructions that cause the processor to perform a method according to any one of EEE1 to EEE21.

[0383] EEE24. A computer-readable storage medium storing the program described in EEE23.

Claims

1. A method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, the method comprising: Obtain first diffraction information related to the acoustic path between the source location and the listener location within the audio scene, wherein the listener location is associated with a first voxel of the projection map; Based on the first diffraction information, first direction information indicating the first rendering direction is determined; Retrieve second diffraction information related to the acoustic path between the source location within the audio scene and the second voxel of the projection mapping; The second diffraction information is pre-stored diffraction information, and the second voxel is a voxel in the neighborhood of the first voxel; Based on the second diffraction information, second direction information indicating the second rendering direction is determined; and The first direction information and the second direction information are output for rendering.

2. The method of claim 1, further comprising selecting the second voxel from the neighborhood of the first voxel according to a predefined selection rule.

3. The method of claim 2, wherein the neighborhood of the first voxel is a predefined neighborhood relative to the first voxel, associated with a set of predefined voxels relative to the first voxel; and Selecting the second voxel includes: Voxels are selected sequentially from the set of predefined voxels according to a predefined selection order.

4. The method of claim 3, wherein selecting the second voxel further comprises: For each selected voxel, determine whether the pre-stored diffraction information can be used for the selected voxel; If the pre-stored diffraction information is available, then the pre-stored diffraction information is retrieved, and the rendering direction is determined based on the retrieved diffraction information.

5. The method of claim 4, wherein selecting the second voxel further comprises: Compare the determined rendering direction with the first rendering direction; and If the difference between the determined rendering direction and the first rendering direction is greater than a predefined threshold, then the selected voxel is used as the second voxel, and the determined rendering direction is used as the second rendering direction.

6. The method according to any one of claims 1 to 5, further comprising determining a first gain and a second gain respectively associated with the first rendering direction and the second rendering direction based on the spatial relationship between the first voxel and the second voxel.

7. The method of claim 6, wherein determining the first and second gains is based on whether the first and second voxels are laterally adjacent or diagonally adjacent in the projection map.

8. The method of claim 6 or 7, wherein the first and second gains are determined based on a predefined Gaussian kernel.

9. The method according to any one of claims 1 to 8, comprising: Determine whether the first voxel is adjacent to the isolated occlusion voxel of the projection map; and If the first voxel is adjacent to the isolated occluding object, then second direction information indicating the second rendering direction is determined based on the first rendering direction and the spatial relationship between the first voxel and the occluding object.

10. The method according to any one of claims 1 to 9, wherein obtaining the first diffraction information involves applying a path search algorithm.

11. The method according to any one of claims 1 to 10, further comprising outputting a representation of the first diffraction information for storage.

12. The method according to any one of claims 1 to 10, wherein the diffraction information includes an indication of a corner voxel on the acoustic path, the diffraction path changing direction for the corner voxel, and the straight line from the first voxel to the corner voxel not being obstructed for the corner voxel.

13. The method of claim 12, wherein the diffraction information further includes an indication of the length of the acoustic path.

14. A method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, the method comprising: Determine whether a first voxel of the projection map associated with the listener's location is adjacent to an isolated occlusion voxel of the projection map; If the first voxel is adjacent to the isolated occlusion voxel, then first diffraction information related to the acoustic path between the source location and the listener location in the audio scene is obtained; Based on the first diffraction information, first direction information indicating the first rendering direction is determined; and The second direction information indicating the second rendering direction is determined based on the first rendering direction and the spatial relationship between the first voxel and the occluding voxel.

15. The method of claim 14, wherein the second rendering direction is determined by rotating the first rendering direction by 90 degrees.

16. The method of claim 15, wherein a rotation direction for rotating the first rendering direction is determined based on the first rendering direction and the direction from the first voxel to the occluding object pixel.

17. The method of claim 15 or 16, wherein the second rendering direction is determined such that the direction from the first voxel to the occluding object pixel is within a sector spanned by the first rendering direction and the second rendering direction.

18. The method according to any one of claims 14 to 17, wherein determining whether the first voxel is adjacent to the isolation occlusion voxel comprises comparing the first voxel with a set of predefined voxels indicated to be adjacent to the isolation occlusion voxel.

19. A method for processing audio scene information associated with a voxel-based audio scene represented by a two-dimensional projection map, the method being used to render audio from an audio source at a source location in the audio scene to a listener at a listener location in the audio scene, the method comprising: The path search-based diffraction information is determined by applying a path search algorithm to determine the acoustic path between the source location and the listener location within the audio scene, wherein the listener location is associated with a first voxel of the projection map, and wherein the path search-based diffraction information includes an indication of a corner voxel on the acoustic path, for which the diffraction path changes direction, and for which the straight line from the first voxel to the corner voxel is not occluded. Identify a set of voxels that intersect with a portion of the acoustic path extending between the first voxel and the corner voxel; The derived diffraction information of each voxel in the set of identified voxels is determined based on the path-search-based diffraction information. and The path-search-based diffraction information and the derived diffraction information are output for storage.

20. The method of claim 19, wherein the derived diffraction information of each voxel in the set of identified voxels indicates the same corner voxel as the path-search-based diffraction information.

21. The method of claim 20, wherein the path-search-based diffraction information further includes an indication of the length of the acoustic path; and Determining the derived diffraction information of a given voxel of the identified voxels includes determining the derived length of the acoustic path based on the length of the acoustic path and the distance between the first voxel and the given voxel.

22. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to perform the method according to any one of claims 1 to 21.

23. A program comprising, when executed by a processor, instructions that cause the processor to perform the method according to any one of claims 1 to 21.

24. A computer-readable storage medium storing the program according to claim 23.