Method, device and system for processing audio scene information

By encoding path information of audio scenes in voxel meshes and utilizing differential encoding of corner voxels and path lengths, the high computational complexity and storage requirements of voxel rendering of audio scenes in VR, AR, MR, and XR environments are solved, achieving efficient audio scene rendering and lossless restoration.

CN121532825APending Publication Date: 2026-02-13DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202480047238.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-15
Filing Date
2024-06-05
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

When using voxel rendering to render audio scenes in VR, AR, MR, and XR environments, existing technologies suffer from high computational complexity and large storage and bandwidth requirements, especially the high computational cost and storage requirements caused by frequent recalculation of diffraction paths when users and audio sources move.

Method used

By encoding the path information of the audio scene, utilizing the corner voxels and path length in the voxel grid, differential coding and mode selection are employed to reduce the bit stream and storage requirements of the encoded path information, while ensuring lossless recovery of the original path information.

Benefits of technology

It effectively reduces the amount of data for encoding path information, lowers computational complexity and storage requirements, while ensuring efficient rendering of audio scenes and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532825A_ABST
    Figure CN121532825A_ABST
Patent Text Reader

Abstract

The invention relates to a method of processing audio scene information. One such method includes obtaining a voxel-based audio scene representation of an audio scene; sequentially encoding items of path information for a first location and a second location in the two-dimensional voxel grid, each item of path information specifying the first location, the second location, a path length of the acoustic path, and corner voxels on the acoustic path; and for the current path information item, generating an encoded path information item based on the path information item. The encoded path information item includes indications of respective first and second locations. If the corner voxel specified by the current path information item is different from the corner voxel specified by the previous path information item, the encoded path information item includes an indication of the corner voxel; if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, the encoded path information item includes an indication of "the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item", rather than an indication of the corner voxel. The disclosure also relates to a corresponding apparatus, a computer program and a computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 508367, filed June 15, 2023, which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to techniques for processing audio scene information, such as for storage, transmission, and / or audio rendering. In particular, this disclosure relates to voxel-based scene representation and the encoding of path information (e.g., acoustic path information, such as diffraction path information) for voxel-based scene representation. Background Technology

[0004] The Moving Picture Experts Group (MPEG) is a working group alliance jointly established by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC) to develop standards for media coding, including audio coding. MPEG is part of ISO / IEC SC29, and the audio group is currently designated as Working Group 6 (WG 6). WG 6 is currently developing a new audio standard (also known as MPEG-I Immersive Audio, ISO / IEC 23090-4).

[0005] The new MPEG-I standard enables acoustic experiences from different viewpoints and / or perspectives or listening positions by supporting scenes and various movements around those scenes, such as the use of various degrees of freedom (e.g., three-DOF or six-DOF) movements in virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) applications. 6DoF interaction extends the 3DoF spherical video / audio experience, which was previously limited to head rotation (pitch, yaw, and roll), to include translational movement (forward / backward, up / down, and left / right), thus allowing navigation within the virtual environment (e.g., actually walking inside a room) in addition to head rotation.

[0006] For audio rendering in VR, AR, MR, and XR applications, object-based approaches have been widely adopted by representing complex auditory scenes as multiple individual audio objects, each associated with parameters or metadata defining that object's location / position and trajectory within the scene. Alternatively, higher-order high-fidelity stereo (HOA) is also used for audio rendering in such environments. However, new uses for rendering audio scenes using "voxels" are now being explored, such as for novel immersive audio experiences. Voxels used for audio rendering are relevant to media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments.

[0007] A voxel is a spatial volume assigned acoustic properties or audio rendering instructions. Voxel size can be an encoder configuration parameter, and it can be selected (manually or automatically) based on the level of scene geometry detail (e.g., in the range of 10cm–1m).

[0008] Voxels used for audio rendering can be obtained in the following ways:

[0009] • Voxelization (or transformation) of mesh-based scene representation

[0010] • From the scene representation used for scene generation (or even video rendering) (e.g., by downsampling smaller voxels).

[0011] However, conventional methods of using voxels to deliver realistic sound for user experiences (including those involving movement) in VR, AR, MR, and XR environments remain challenging and computationally complex.

[0012] Typical techniques for diffraction modeling in 3D audio scenes, such as those used in computer-mediated real-world applications, require recalculating diffraction paths and other diffraction information whenever the audio scene, user location, or audio source location changes. For example, diffraction paths may change as the user and / or audio source moves within the 3D audio scene. Furthermore, diffraction paths may change when the audio scene itself changes (e.g., by indicating an open or closed door or window). Frequent recalculation of diffraction paths can be computationally expensive, requiring relatively powerful computing devices for computer-mediated real-world applications, and / or in some cases, may negatively impact the user experience. On the other hand, storing pre-calculated acoustic paths (e.g., acoustic diffraction paths) can require significant storage and / or bandwidth.

[0013] For example, US11606662B2, US6313841B1, US10275937B2, and US10293259B2 each involve (pre-)computation of paths in a given scenario or environment. However, doing so can generate a relatively large amount of data (especially for a large number of paths), particularly when the focus is on lossless transmission or storage of path information.

[0014] Therefore, improved techniques are needed to encode path information (e.g., acoustic path information, such as diffraction path information) in voxel-based audio scenes. In particular, techniques capable of reducing the bandwidth or storage requirements for processing encoded path information are needed. Summary of the Invention

[0015] In view of this need, this disclosure provides a method for processing audio scene information (particularly voxel-based audio scene information), an apparatus for processing audio scene information, a computer program, and a computer-readable storage medium having the features of the respective independent claims.

[0016] One aspect of this disclosure relates to a method for processing audio scene information, such as encoding audio scene information, particularly acoustic path information. The method may include obtaining a voxel-based audio scene representation of the audio scene. The method may further include: sequentially encoding path information items for corresponding first and second locations in a two-dimensional voxel grid associated with the voxel-based audio scene representation, for one or more (e.g., multiple) first locations and one or more (e.g., multiple) second locations. The voxel grid may relate to a two-dimensional projection map generated from the voxel-based audio scene representation. Each path information item may specify the first location, the second location, the path length of the acoustic path between the first and second locations in the voxel grid, and corner voxels where the acoustic path changes direction. A corner voxel may be an unoccluded voxel pointing to a second voxel. The method may further include: for a current path information item, generating an encoded path information item based on that path information item. The encoded path information item may include an indication of the corresponding first location and an indication of the corresponding second location. If the corner voxel specified by the current path information item is different from the corner voxel specified by the previous path information item, then the encoded path information item may include an indication of the corner voxel (e.g., the location of the corner voxel). This indication of the corner voxel may be, for example, an absolute or non-differential indication of the corner voxel using voxel coordinates or voxel indexes. On the other hand, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item may include an indication that "the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item," rather than an indication of the corner voxel itself.

[0017] By encoding path information items sequentially and reusing information from the previous path information item, the proposed method reduces the bitstream and storage requirements for processing encoded path information. Nevertheless, the original path information can still be fully recovered; that is, the proposed method provides lossless encoding of path information in voxel-based scenarios.

[0018] In some embodiments, the encoded path information item may further include an indication of path length. The indication of path length may be, for example, an absolute or non-differential indication of path length in units of voxels or voxel side lengths. In some embodiments, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item may further include an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item.

[0019] This allows for a further reduction in the amount of data required to encode path information, while still allowing for lossless recovery of the original path information.

[0020] In some embodiments, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, the encoded path information item may further include an indication of whether the encoded path information item includes an indication of the path length or an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item. In this case, the encoded path information item may further include an indication of the path length or an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item.

[0021] This allows for a further reduction in the overall amount of data required to encode path information, while still allowing for lossless recovery of the original path information. In some embodiments, the method may further include: for each first location, traversing a voxel grid according to a predetermined pattern to determine a sequence of second locations. The method may further include: for a corresponding first location, sequentially encoding path information items for the determined sequence of second locations.

[0022] In some embodiments, the predetermined pattern can traverse the voxel grid along rows and columns in a raster scan manner.

[0023] In some embodiments, the difference may be encoded by 2 bits. Alternatively, the difference may take one of four predetermined values ​​(e.g., the potential values ​​of the difference may be limited to a set of four distinct values).

[0024] Using the above pattern of traversing the voxel grid, it can be ensured that once a corner voxel does not change from one path information item to the next, the difference between the corresponding path lengths can only take one of four values, which can be efficiently encoded using only two bits.

[0025] In some embodiments, each encoded path information item may include an indication of whether a first mode or a second mode is used. Each encoded path information item may also include an indication of a first location. Each encoded path information item may also include an indication of a second location. Each encoded path information item may also include an indication of whether an acoustic path exists for the first and second locations. If the first mode is used, each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item. Furthermore, in the first mode, if the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, each encoded path information item may also include an indication of path length. Furthermore, in the first mode, each encoded path information item may otherwise include an indication of corner voxel and an indication of path length. If the second mode is used, each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item. Furthermore, in the second mode, if the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then each encoded path information item may also include an indication of the difference between the previous path length and the current path length. The previous path length can be the path length specified by the previous path information item, and the current path length can be the path length specified by the path information item corresponding to the encoded path information item. Additionally, in the second mode, each encoded path information item may otherwise include an indication of the corner voxel and an indication of the path length.

[0026] In some embodiments, each encoded path information item may include an indication of whether a first mode or a second mode is used. Each encoded path information item may also include an indication of a first location. Each encoded path information item may also include an indication of a second location. Each encoded path information item may also include an indication of whether an acoustic path exists for the first and second locations. If the first mode is used, then each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item. Then, if the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item may also include an indication of whether the encoded path information item includes an indication of path length or an indication of the difference between the previous path length and the current path length, together with the indication of path length or the indication of the difference between the previous path length and the current path length. Here, the previous path length may be the path length specified by the previous path information item. The current path length may be the path length specified by the path information item corresponding to the encoded path information item. Furthermore, in the first mode, the encoded path information item may otherwise include an indication of a corner voxel and an indication of the path length. If the second mode is used, each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item. Furthermore, in the second mode, if the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then each encoded path information item may also include an indication of the difference between the previous path length and the current path length. Here, the previous path length may be the path length specified by the previous path information item, and the current path length may be the path length specified by the path information item corresponding to the encoded path information item. Furthermore, in the second mode, each encoded path information item may otherwise include an indication of a corner voxel and an indication of the path length.

[0027] In some embodiments, the method may further include outputting encoded path information items to a bitstream.

[0028] Another aspect of this disclosure relates to a method for processing audio scene information, such as decoding audio scene information, particularly acoustic path information. The method may include receiving a bitstream comprising a sequence of encoded path information items for one or more first locations and one or more second locations in a two-dimensional voxel grid associated with a voxel-based audio scene representation. Each encoded path information item may correspond to a corresponding path information item specifying the first location, the second location, the path length of an acoustic path between the first and second locations in the voxel grid, and corner voxels along the acoustic path where the acoustic path changes direction. A corner voxel may be an unoccluded voxel pointing to a second voxel. The method may further include sequentially decoding the encoded path information items to generate corresponding path information items. Generating a corresponding path information item for a currently encoded path information item may include determining whether the currently encoded path information item includes an indication that "the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item." Then, if the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, the method may include setting the corner voxel specified by the path information item corresponding to the previously encoded path information item as the corner voxel of the path information item corresponding to the currently encoded path information item. On the other hand, if the corner voxel specified by the corresponding path information item is different from the corner voxel specified by the path information item corresponding to the previously encoded path information item, the method may include an indication to extract the corner voxel from the currently encoded path information item.

[0029] In some embodiments, generating the corresponding path information item may further include an indication of the path length extracted from the currently encoded path information item.

[0030] In some embodiments, if the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then generating the corresponding path information item may further include an indication of extracting the difference between the previous path length and the current path length. The previous path length may be the path length specified by the path information item corresponding to the previously encoded path information item. The current path length may be the path length specified by the path information item corresponding to the currently encoded path information item.

[0031] In some embodiments, if the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then generating the corresponding path information item may further include extracting an indication of whether the encoded path information item includes an indication of the path length or an indication of the difference between the previous path length and the current path length. The previous path length may be the path length specified by the path information item corresponding to the previously encoded path information item. The current path length may be the path length specified by the path information item corresponding to the currently encoded path information item. Generating the corresponding path information item may further include extracting an indication of the path length or an indication of the difference between the previous path length and the current path length.

[0032] In some embodiments, one or more second locations may refer to locations obtained by traversing a voxel grid according to a predetermined pattern for each first location to define a sequence of second locations. Furthermore, for a corresponding first location, encoded path information items can be decoded sequentially according to the sequence of second locations.

[0033] In some embodiments, the predetermined pattern can traverse the voxel grid along rows and columns in a raster scan manner.

[0034] In some embodiments, the difference may be encoded using 2 bits. Alternatively, the difference may take one of four predetermined values.

[0035] In some embodiments, each encoded path information item may include an indication of whether a first mode or a second mode is used. Each encoded path information item may also include an indication of a first location. Each encoded path information item may also include an indication of a second location. Each encoded path information item may also include an indication of whether an acoustic path exists at the first and second locations. If the first mode is used, then each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item. If the first mode is used, and if the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then each encoded path information item may also include an indication of path length. Otherwise, each encoded path information item may also include an indication of corner voxel and an indication of path length. If the second mode is used, each encoded path information item may further include an indication of whether the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item. If the second mode is used, and if the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then each encoded path information item may further include an indication of the difference between the previous path length and the current path length. Here, the previous path length may be the path length specified by the path information item corresponding to the previous encoded path information item, and the current path length may be the path length specified by the path information item corresponding to this encoded path information item. Otherwise, each encoded path information item may further include an indication of the corner voxel and an indication of the path length.

[0036] In some embodiments, each encoded path information item may include an indication of whether a first mode or a second mode is used. Each encoded path information item may also include an indication of a first location. Each encoded path information item may also include an indication of a second location. Each encoded path information item may also include an indication of whether an acoustic path exists at the first and second locations. If the first mode is used, then each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item. If the first mode is used, and if the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then each encoded path information item may also include an indication of whether the encoded path information item includes an indication of path length or an indication of the difference between the previous path length and the current path length, together with the indication of path length or the indication of the difference between the previous path length and the current path length. The previous path length may be the path length specified by the path information item corresponding to the previous encoded path information item. The current path length can be the path length specified by the path information item corresponding to the encoded path information item. Otherwise, each encoded path information item may also include an indication of a corner voxel and an indication of the path length. If the second mode is used, each encoded path information item may also include an indication of whether the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item. If the second mode is used, and if the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then each encoded path information item may also include an indication of the difference between the previous path length and the current path length. Here, the previous path length can be the path length specified by the path information item corresponding to the previous encoded path information item, and the current path length can be the path length specified by the path information item corresponding to the encoded path information item. Otherwise, each encoded path information item may also include an indication of a corner voxel and an indication of the path length.

[0037] According to another aspect, an apparatus for processing audio scene information is provided. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform all steps of the method according to the foregoing aspects and embodiments thereof.

[0038] According to another aspect, a computer program is described. This computer program may include executable instructions for performing, when executed by a computing device (e.g., a processor), the methods or method steps outlined throughout this disclosure.

[0039] According to another aspect, a computer-readable storage medium is described. This storage medium can store a computer program adapted to be executed on a computing device (e.g., a processor) and, when executed on that computing device, to perform the methods or method steps outlined herein.

[0040] It should be noted that the methods and systems outlined in this disclosure, and their preferred embodiments, can be used alone or in combination with other methods and systems disclosed in this document. Furthermore, all aspects of the methods and systems outlined in this disclosure can be combined arbitrarily. In particular, the features of the claims can be combined with each other in any manner.

[0041] It will be appreciated that the device features and method steps can be interchanged in a variety of ways. In particular, the details of the disclosed method(s) can be implemented by the corresponding device and vice versa, as will be appreciated by those skilled in the art. Moreover, any foregoing statements made with respect to the method(s) (and their steps, for example) should be understood to apply equally to the corresponding device (and its blocks, stages, units, for example). Attached Figure Description

[0042] The invention will now be explained by way of example with reference to the accompanying drawings, in which:

[0043] Figure 1 An example of a processing chain used to process audio scene information for audio rendering is illustrated schematically.

[0044] Figure 2 An example of diffraction paths at the source and listener locations in a voxel-based 3D audio scene is illustrated schematically.

[0045] Figure 3 This is a flowchart illustrating an example of a method for processing audio scene information to perform audio rendering according to an embodiment of the present disclosure;

[0046] Figure 4 The illustration is based on an embodiment of the present disclosure. Figure 3 A flowchart illustrating the implementation details of the method;

[0047] Figures 5 to 7 An example of a processing chain for processing audio scene information for audio rendering according to an embodiment of the present disclosure is illustrated schematically.

[0048] Figure 8This is a diagram illustrating the complexity as a function of time for different operation modes / implementations of processing audio scene information for audio rendering according to embodiments of the present disclosure.

[0049] Figure 9 Examples of possible use cases of the technology according to embodiments of this disclosure are illustrated schematically;

[0050] Figures 10A-10C An example of a voxel-based audio scene according to an embodiment of the present disclosure is illustrated schematically;

[0051] Figure 11 An example of a voxel-based audio scene to which embodiments of the present disclosure may be applied is illustrated schematically;

[0052] Figure 12 This is a flowchart illustrating an example of a method for processing audio scene information to encode acoustic path information according to an embodiment of the present disclosure;

[0053] Figure 13 This is a flowchart illustrating an example of a method for processing audio scene information to decode acoustic path information according to an embodiment of the present disclosure;

[0054] Figure 14 This is a diagram illustrating the complexity as a function of time for different operation modes / implementations of processing audio scene information for audio rendering according to embodiments of this disclosure; and

[0055] Figure 15 This is a block diagram schematically illustrating an example of an apparatus for implementing a method according to an embodiment of the present disclosure. Detailed Implementation

[0056] In the following description, exemplary embodiments of the present disclosure will be illustrated with reference to the accompanying drawings. The same elements in the drawings may be indicated by the same reference numerals, and repeated descriptions may be omitted.

[0057] Voxel-based audio scene representation

[0058] First, an overview of voxel-related concepts used to represent audio scenes will be given.

[0059] What are voxels used for audio rendering?

[0060] A voxel is understood as a spatial volume to which acoustic properties or audio rendering instructions are assigned.

[0061] What is the voxel size used for audio rendering?

[0062] Voxel size can be an encoder configuration parameter. It can be selected (manually or automatically) based on the level of scene geometry detail (e.g., in the range of 10 cm–1 m).

[0063] How large of an audio scene can it handle?

[0064] Large audio scenes do not necessarily lead to a large number of voxels and high rendering complexity. For example, a large audio scene can be represented as:

[0065] • A set of independent sub-scenes (and methods for "transferring" between these representations without requiring a renderer "restart");

[0066] • A set of scene updates (based on user location).

[0067] How can we address the discontinuity issues caused by voxel granularity?

[0068] Any strong discontinuities in the sound level (and jumps in the direction of the diffraction signal) can be avoided by applying interpolation (e.g., in time and space).

[0069] How to represent a voxel-based audio scene?

[0070] Any voxel-based representation of an audio scene can include indications of voxels that are not transmissive voxels (e.g., as occlusion voxels)—representations of occlusion geometry. This indication can involve the coordinates of the corresponding voxels (e.g., center coordinates, corner coordinates, etc.). These voxel coordinates can be represented, for example, by a grid index. Furthermore, voxel-based representations can include indications of the material properties (such as absorption coefficients, reflection coefficients, etc.) of voxels that are not transmissive voxels. In addition to occlusion voxels, voxel-based representations can also indicate transmissive voxels (e.g., air voxels), i.e., voxels in which sound can propagate—representations of the sound propagation medium. Accordingly, some implementations of voxel-based audio scene representations can include indications of the corresponding material properties for each voxel within a predefined spatial portion (e.g., within the boundary surrounding the audio scene).

[0071] Technology for processing audio scene information

[0072] Figure 1 A processing chain 100 is schematically illustrated that can be used to process audio scene information for audio rendering. Specifically, the processing chain 100 can be used to convert voxel-related data into parameters and signals required for audible rendering (or general audio rendering). The processing chain 100 can be implemented in software, hardware, or a combination thereof. For example, the processing chain 100 can be implemented by a renderer / decoder coupled to an AR / VR / MR / XR device (such as AR / VR / MR / XR glasses). Specific implementations may include game consoles, set-top boxes, personal computers, etc.

[0073] The processing chain receives an audio scene description 20 from the bitstream (or storage device / memory) 10. The audio scene description 20 may include a representation of a three-dimensional audio scene and information about the source locations of sound sources within the audio scene. The representation of the three-dimensional audio scene may be, for example, voxel-based.

[0074] Processing chain 100 also receives an indication of the user's (listener's) location 30 within the audio scene. The audio scene description 20 and the user's location 30 are provided to a diffraction direction calculation block (diffraction calculation block) 40 for determining (e.g., calculating) diffraction information. The diffraction information may relate to the acoustic diffraction path between the source location and the listener's location within the audio scene. The diffraction information is then provided to a diffraction modeling tool 50 for applying diffraction modeling and optional occlusion modeling based on the diffraction information. Occlusion modeling calculates the attenuation gain of the straight line between the listener and the audio source. The diffraction modeling tool 50 can output audible audio data (3DoF audibler data) including, for example, the location, orientation, and frequency-dependent gain of the object to be rendered. The output of the diffraction modeling tool can be further processed by other rendering stages, such as Doppler, directivity, distance attenuation, etc. In general, the diffraction modeling tool 50 can be described as outputting diffraction information, as detailed below. The audible audio data can then be used, for example, for audio playback.

[0075] In short, such as Figure 1 The processing chain shown can be used to convert voxel-related data into parameters and signals for audibility. The diffraction direction calculation block 40 and diffraction modeling tool 50 can be considered as non-limiting examples of rendering tools. Generally, rendering tools can generate 3DoF audibler data.

[0076] As described above, the scene description may include a voxel matrix and associated coefficients (e.g., reflection coefficients, occlusion coefficients, absorption coefficients, transmission coefficients, etc.). These coefficients may indicate the material or material properties of the corresponding voxels. Rendering tools may include, for example, occlusion and diffraction modeling tools. 3DoF audible data may include, for example, object position, orientation, and frequency-dependent gain.

[0077] As mentioned above, the voxel-based representation of a 3D audio scene defines the psychoacoustic geometric elements and sound propagation medium. In some implementations, the scene description can provide information to the rendering tool using parameters / interfaces such as the following agreed-upon data formats or agreed-upon data exchange points:

[0078] Scene size:

[0079] - In absolute units (e.g., meters);

[0080] - By voxel count and / or voxel size.

[0081] Scene anchor point:

[0082] - In terms of coordinate anchor points (mapping absolute coordinates to voxel indices);

[0083] - In terms of scene anchors (mapping sub-scenes to voxel subsets).

[0084] Scene content data:

[0085] - References to material properties, which are approximated by the acoustic effects (e.g., transmission, reflection, etc.) caused by obstructions (sound barriers) located in the corresponding volume;

[0086] - References to the properties of the sound propagation medium, which approximate the acoustic effects caused by the medium located in the corresponding volume (e.g., sound speed, energy absorption, distance attenuation curves, etc.).

[0087] - Rendering control parameters that describe the expected occlusion modeling effect, such as the "global" or "local" occlusion type, which determines the length (and shape) of the occlusion effect shadow behind this voxel.

[0088] - Rendering control parameters that describe the expected sound diffraction modeling effect; for example, the type of voxel that controls / causes a change in sound direction (i.e., the path of the diffracted sound cannot penetrate the volume).

[0089] - Content control parameters that describe audio signal relevance and scene creation, such as determining which signal has perceptual relevance (to be rendered) in the corresponding volume, the audio signal ID, and / or signal gain.

[0090] - Rendering control parameters that describe the desired reverb modeling effect, such as the voxel type that controls reverb settings (e.g., RT60, DDR, RIR, etc.).

[0091] Scene content update:

[0092] - Referenced to update the triggering event.

[0093] All data can rely on audio objects (to support content creators' intentions in creating flexible audio scenarios).

[0094] 3DoF audible amplifier data may include the following information:

[0095] - The parameters and associated signals of this group of audio objects (and HOAs);

[0096] ○ Parameters include the metadata output of the rendering tool (i.e., position, orientation, and gains for simulated occlusion, diffraction, early reflection, reverberation coefficient parameters, IR, etc.).

[0097] ○ Associated signals represent the audio output of the rendering tool (i.e., the audio signal that has been mixed or copied).

[0098] - Scene state identifier (i.e., metadata that allows scene descriptions and user inputs to be mapped to 3DoF audible data).

[0099] Figure 2 Examples of possible scene states and their diffraction paths are illustrated. It should be understood that a scene state involves or includes listener location 210 and an audio scene description (including a representation of the three-dimensional audio scene and source location 220).

[0100] Figure 2 The example relates to a voxel-based representation of a three-dimensional audio scene. This voxel-based representation indicates "air" voxels or empty voxels (i.e., voxels in which sound can propagate, or transmissive voxels) 230 and occlusion voxels 240 (i.e., voxels in which sound cannot propagate or cannot propagate freely). Accordingly, occlusion voxels can be understood as voxels filled with material other than air and capable of reflecting, blocking, or otherwise altering sound propagation. For occlusion voxels 240, the representation may also indicate corresponding transmission, reflection, and potential absorption coefficients related to the material properties of these voxels. These coefficients may be linked to the ID or index of their corresponding voxels in the voxel-based representation. In general, a voxel-based representation can define psychoacoustic-related geometric elements and sound propagation media in an audio scene.

[0101] Listener location 210 is determined by parameter L VOX Indicated, and source location 220 is determined by another parameter S VOX instruct.

[0102] A pathfinding algorithm can be used to determine the diffraction path (or, in general, the acoustic path) between source location 220 and listener location 210. This algorithm takes as input a representation of listener location 210, source location 220, and a three-dimensional audio scene (or a two-dimensional representation derived therefrom, such as a 2D projection or a 2D matrix, for example, in the form of a voxel grid). For example, an algorithm for determining diffraction information can take the representation of listener location 210, source location 220, and three-dimensional audio scene as input and can output a path defined by C. VOX The location of the indicated corner voxel (e.g., diffraction angle) 250° and the variable r representing the length of the diffraction path. in For example, diffraction information (or acoustic path information in general) can be determined based on the following:

[0103]

[0104] Where DiffractionDirectionCalculation indicates the algorithm used to determine diffraction information (“pathfinding algorithm”), and VoxDataDiffractionMap indicates a voxel-based representation of the 3D audio scene or a processed version thereof (e.g., a 2D projection or a 2D matrix derived therefrom). C VOX This is understood as the coordinates indicating the diffraction angle (e.g., the coordinates of the corresponding voxel of the diffraction angle, voxel / mesh coordinates, or voxel / mesh index). The diffraction angle can also be referred to as the corner voxel.

[0105] Here, DiffractionDirectionCalculation can involve any feasible pathfinding algorithm, such as, for example, the fast traversal algorithm for ray tracing (see Amanatides, J. and A. Woo, A Fast VoxelTraversal Algorithm for Ray Tracing. Proceedings of EuroGraphics, 1987. 87.) and the JPS algorithm (see Harabor, DD and A. Grastien, Online Graph Pruning for PathfindingOn Grid Maps. Proceedings of the Twenty-Fifth AAAI Conference on Artificial Intelligence, 2011.). Alternatively, a 3D pathfinding algorithm can be directly applied to obtain the shortest path between source location 220 and listener location 210 using a voxel-based scene representation. Alternatively, a 2D pathfinding algorithm can be applied to this task using an appropriate 2D projection plane of the 3D voxel-based scene representation. For indoor (e.g., multi-room) sound simulation, the corresponding 2D projection plane can be analogous to a plan view describing the "sound propagation path topology". For outdoor sound simulation scenarios, it may be of interest to consider a second (e.g., vertical) 2D projection plane to account for diffraction paths over one or more sound obstacles or occlusion structures. The pathfinding method remains the same for all projection planes, but its application provides additional paths that can be used for diffraction modeling.

[0106] Suppose the pathfinding algorithm outputs an acoustic path (e.g., a diffraction path) connecting the source location 220 to the listener location and consisting of multiple voxel indicators (e.g., voxel indices or voxel coordinates). The acoustic path can be said to change direction at these voxels. For visualization, the acoustic path can in some cases be viewed as involving multiple consecutive straight path segments (line segments). Each transition from one path segment to another involves a change in the direction of the diffraction path.

[0107] According to the algorithm used to determine diffraction information, the diffraction angle (or corner voxel) C VOX It can be identified as a voxel located on or near the diffraction path and obstructing the diffraction pattern (indicated by a voxel-based representation) (a set of voxels C10 ... set (in the middle) adjacent voxels. For example, diffraction angle C VOX From a set of voxels (P) that form the diffraction path, set Choose from the options, as the path that leads to the nearest (P) set The obstructing object whose direction changes (from the listener's position Lc) is "visible" (belonging to C) set The voxel is ). If there is more than one such angle, then choose along the diffraction path (P). set The one furthest from the listener's location.

[0108] Generally speaking, diffraction path algorithms can be described as determining diffraction information related to the acoustic path (e.g., acoustic diffraction path) between the source location and the listener location within an audio scene.

[0109] This diffraction information (or general acoustic path information) can be sufficient for the renderer to recover / determine the location of the virtual source of the virtual audio source encapsulating the acoustic diffraction effect. For the diffraction angle C... VOX Coordinates and diffraction path length r in That's the situation. For example, the virtual source location can be reconstructed by calculating the direction of the diffraction angle as seen from the listener's location (e.g., azimuth, or azimuth and elevation). This direction is then used, and the path length r of the diffraction path is taken. in The virtual source location can be determined by the distance of the virtual source to the listener's location.

[0110] Note that diffraction information can be represented in different ways. As mentioned above, one option is to include / store the path length r. in and diffraction angle C VOX Diffraction information of coordinates (e.g., grid coordinates).

[0111] Based on the above, the following data elements can be defined:

[0112] An example of scene state N1 can be represented as

[0113]

[0114] That is, it may involve or include the listener's location L. VOX Source location S VOX And voxel-based representations of audio scenes (e.g., VoxDataDiffractionMap).

[0115] The scene state identifier of scene state N1 can be defined as

[0116]

[0117] HASH is a hash function that generates the hash value of scene state N1; for example, it maps scene states to values ​​of a fixed size. Generally speaking, a scene state identifier can be described as indicating or identifying a specific scene state.

[0118] Furthermore, an example of diffraction information N2 can be represented as

[0119]

[0120] Where r in It is the path length of the diffraction path, and C vox Locations indicating diffraction angles (e.g., voxel locations), as described above.

[0121] The quantized version of the diffraction information N2 can be indicated by N3, where

[0122]

[0123] Where `voxSceneDiffractionPreComputedPathData()` parses the bitstream and retrieves pre-computed (stored and quantized) diffraction information (e.g., from...). Figure 6 The bitstream syntax (generated by the processing chain 500) and DiffractionDirectionCalculation() represent the function that performs online calculations of diffraction information, which can be used, for example, in... Figure 5 and Figure 7 The diffraction direction calculation is implemented in block 40.

[0124] Diffraction information (e.g., C) vox and r in This can also be considered as involving 3DOF audible data, because the user position voxel coordinates L vox It is fixed.

[0125] Examples of the voxSceneDiffractionPreComputedPathData() syntax element according to the MPEG-I standard are given in Table 1. This voxel payload data structure can have the following elements:

[0126] numberOfVoxDiffractionPathData:

[0127] This element represents the number of pre-computed diffraction path datasets.

[0128] voxDiffractionPathStartVoxelPacked:

[0129] This element represents the packaged form of the variable `voxDiffractionPathStartVoxel`, indicating the path-starting voxel of the pre-computed diffraction path (e.g., ...). Figure 2 S in vox or L vox The voxel index of the diffraction path. voxDiffractionPathStartVoxel can be, for example, a 2D position on a diffraction map indicating the starting position of the diffraction path.

[0130] voxDiffractionPathEndVoxelPacked:

[0131] This element represents the packaged form of the variable `voxDiffractionPathEndVoxel`, indicating the path-ending voxel of the pre-computed diffraction path (e.g., ...). Figure 2 L in vox or S vox The voxel index of the path is: voxDiffractionPathEndVoxel. This can be, for example, a 2D position on a diffraction map indicating the end of the diffraction path.

[0132] voxDiffractionPathDataExistFlag:

[0133] This element indicates whether a diffraction path exists.

[0134] voxDiffractionSourceDirectionPacked:

[0135] This element represents the packaged form of the variable `voxDiffractionSourceDirection`, indicating the voxel index used to determine the voxel orientation value of the diffraction source. This could correspond to, for example, a corner voxel C. vox .

[0136] voxDiffractionPathLength:

[0137] This element represents diffraction. Figure 2 The diffraction path length on the D matrix. This can correspond to, for example, the path length r. in .

[0138] voxSceneDimensions:

[0139] The number of voxels in each scene dimension.

[0140] escapedValue():

[0141] This element implements a common method for transmitting integer values ​​using a variable number of bits. It features a two-stage escape mechanism, allowing the range of representable values ​​to be extended by continuously transmitting additional bits. The syntax of `escapedValue()` should conform to the definition in ISO / IEC 23003-3.

[0142] Table 1 — Syntax of voxSceneDiffractionPreComputedPathData()

[0143]

[0144] Variables retrieved from voxSceneDiffractionPreComputedPathData() (e.g., ... Figure 6 (As shown in the diagram) is further processed to output diffraction information N3.

[0145] The technical benefits and effects of the technology disclosed herein are that, if the corresponding processing for this scene state has been completed and diffraction information or 3DoF audible data is available, then the scene state identifier or other information derived from the scene state can be used to avoid applying diffraction modeling tools or rendering tools. In such a scenario, the renderer can access the diffraction information / 3DoF audible data (for a known scene state) without applying rendering tools:

[0146] - Reuse previously calculated (pre-calculated) data, or

[0147] - Apply data calculated by another renderer.

[0148] Therefore, the technical benefit and effect is that the technology according to this disclosure relates to lossless functionality for low-complexity modes (complexity and bit rate).

[0149] To fully implement this approach, this disclosure proposes providing an interface for a processing chain (e.g., in a decoder / renderer) that processes audio scene information for audio rendering, for providing / outputting diffraction information for later use or by different decoders / renderers. This interface is understood as a data interface for outputting data in a predefined format, thereby allowing consistent reuse, particularly by other decoders / renderers. This interface can be implemented and / or utilized in any combination of software and hardware. Specifically, this may involve providing / outputting data elements that include diffraction information and information about the scene state (such as, for example, a scene state identifier). The data elements may have a predefined format, such as predefined data fields. Using this interface, the processing chain can provide the calculated diffraction information or 3DoF audible data, along with the scene state identifier, to other decoders / renderers and / or store it for later reuse.

[0150] In some implementations, the above may involve providing / outputting (or receiving on the receiving side) pre-computed acoustic path information for one or more (e.g., all feasible) combinations of source locations and listener locations in a voxel-based scene representation (e.g., a two-dimensional voxel grid).

[0151] Example 1: If the decoder / renderer has already obtained the acoustic path information / diffraction information (e.g., diffraction path) for a given user location (listener location), then the decoder / renderer can reuse it until the user leaves the corresponding voxel volume (or the scene description is updated).

[0152] Example 2: If the calculated acoustic path information / diffraction information corresponds to a scene state unknown to other decoders, then they can reuse the acoustic path information / diffraction information and avoid running their own diffraction modeling tools or rendering tools.

[0153] The exchange and sharing of acoustic path / diffraction information between different decoders can be accomplished using a database that can be included in the bitstream (e.g., accessed via application requests).

[0154] Figure 3 This is a flowchart illustrating an example of a method 300 for processing audio scene information to perform audio rendering according to an embodiment of the present disclosure. Method 300 can be implemented in software, hardware, or a combination thereof. For example, processing chain 100 can be implemented by a renderer / decoder coupled to an AR / VR / MR / XR device (such as AR / VR / MR / XR glasses). Specific implementations may include game consoles, set-top boxes, personal computers, etc.

[0155] Method 300 includes steps S310 to S350, which can be performed, for example, by a decoder / renderer. These steps can be performed whenever the scene state changes. Since the scene state is understood to involve or include listener location 210 and audio scene description (including a representation of the 3D audio scene and source location 220), implemented, for example, by the scene state N1 described above, a change in the scene state may involve one or more of the following: a change in listener location 210, a change in source location, and a change in the 3D audio scene (representation). Alternatively, steps S310 to S350 can be performed for each of multiple processing cycles of the decoder / renderer. If the audio scene description does not change, then step S310 can be omitted. It should also be understood that steps S310 to S350 do not need to be performed according to... Figure 3 The execution sequence is shown.

[0156] At step S310, an audio scene description is received. The audio scene description includes a representation of the three-dimensional audio scene and information about the source locations of sound sources within the audio scene. For example, the audio scene description may include, for example, the element S defined above. vox And VoxDataDiffractionMap.

[0157] In step S320, information about the listener's location within the audio scene is received. The listener's location may correspond to, for example, the element L defined above. VOX .

[0158] In step S330, diffraction information related to the acoustic diffraction path between the source location and the listener location within the audio scene is obtained. The obtained diffraction information can indicate the virtual source location of the virtual sound source. For example, when viewed from the listener's location, the virtual source location can have a diffraction angle C... VOX The same direction (e.g., azimuth, or azimuth and elevation). The virtual source distance can correspond to the length r of the diffraction path. in Accordingly, the diffraction information can include the C defined above. vox and r in Instructions.

[0159] At step S340, audio rendering is performed on the sound source based on diffraction information. This may include, for example, diffraction modeling.

[0160] Therefore, the location of a virtual source can be determined based on diffraction information. A virtual source can be an audio source that encapsulates the acoustic diffraction effects between a source location and a listener location in a 3D audio scene. For example, it can be based on C... VOX and r in The virtual source location is determined in the following way:

[0161] • Determine the diffraction angle C when viewed from the listener's location.VOX The direction (e.g., azimuth, or azimuth and elevation).

[0162] • Use a defined direction as the virtual source direction when viewed from the listener's location; and

[0163] • Use the diffraction path length r in As

[0164] ○ The distance of the virtual source from the listener's location, or

[0165] ○ Virtual source gain compensation derived from diffraction and direct path length.

[0166] Then, audio rendering can include, for example, rendering a virtual sound source at a virtual source location.

[0167] At step S350, a representation of the diffraction information is output. For example, the representation of the diffraction information may include data elements that output diffraction information and information about the scene state. The scene state may include an audio scene description (e.g., S...). vox (and VoxDataDiffractionMap) and listener locations (e.g., L) VOX ).

[0168] The output can be provided to a lookup table (LUT). The LUT comprises different diffraction information items indexed by information about the corresponding scene state (e.g., indexed by the corresponding scene state identifier) ​​as its entries. Therefore, the LUT can be described as including diffraction information and information about the scene state. The LUT can be stored and / or provided for later retrieval from the bitstream or from shared storage (e.g., cloud-based or server-based) by other decoders, for example, through application requests. The actual desired entry can be retrieved from the LUT using either the hash value of the scene state or the scene state identifier.

[0169] Furthermore, the representation of diffraction information can be output to a bitstream (e.g., outgoing bitstream) and / or a storage device (e.g., memory, cache, file, etc.). The storage device can be local or it can be shared (e.g., cloud-based). Generally, the representation of diffraction information can be output to a suitable medium for storing digital or computer-related information. The output can be at least partially directed to an external or shared data source or database.

[0170] In some implementations, the representation of diffraction information can be output as part of the voxSceneDiffractionPreComputedPathData() syntax element according to ISO / IEC 23090-4 (Coded representation of immersive media — Part 4: MPEG-I immersive audio, https: / / www.iso.org / standard / 84711.html) or any future standard derived from it.

[0171] For example, the syntax elements of voxSceneDiffractionMap() can be given in Table 2.

[0172] Table 2 — Syntax of voxSceneDiffractionMap()

[0173]

[0174] voxSceneDiffractionMap() provides a compact representation of the 2D diffraction map (VoxDataDiffractionMap). This 2D representation is similar to the 3D representation used for voxel-based 3D audio scenes.

[0175] A MapElement is defined by two points (x, y indices) on the diffraction map and their corresponding values. These two points span a rectangle, and all covered grid cells are assigned the value voxDiffractionMapValue.

[0176] The bitstream element numberOfVoxDiffractionMapElements represents the number of MapElements.

[0177] The bitstream element `voxDiffractionMapValue` represents a binary value that controls the pathfinding algorithm. It is useful because this value indicates whether a path can pass through a grid cell. This value is defined for all entries on the diffraction map.

[0178] The bitstream element `voxDiffractionMapPosPackedS` represents a packed representation of the starting grid cell of the `MapElement` using two indices. It can be an array illustrating the set of all starting grid cells.

[0179] The bitstream element `voxDiffractionMapPosPackedE` represents a packed representation of the two indices of the ending grid cell of the `MapElement`. It can be an array illustrating the set of all ending grid cells.

[0180] Both voxDiffractionMapPosPackedS and voxDiffractionMapPosPackedE are useful because they allow for a compact representation of the data, where a single voxDiffractionMapValue is used for all grid cells between the two variables.

[0181] Figure 4 This is a flowchart illustrating an example of a method 400 including steps that can be performed to implement the steps of method 300. Method 400 includes steps S410 to S460. Steps S410 to S450 can implement step 330 of method 300. Furthermore, step S460 can correspond to step S350.

[0182] In step S410, the current scene state is determined based on the audio scene description and the listener's location.

[0183] At step S420, it is determined whether the current scene state corresponds to a known scene state for which pre-computed diffraction information is available (e.g., can be retrieved). The pre-computed diffraction information can be retrieved from a bitstream (incoming bitstream) or a storage device (particularly, including external or shared storage devices). Determining whether the current scene state corresponds to a known scene state may include determining a hash value based on the current scene state. It may also include comparing the hash value of the current scene state with the hash value of a known (e.g., previously encountered) scene state.

[0184] If it is determined that the current scene state corresponds to the known scene state (as in step S430), then the method proceeds to step S440.

[0185] At step S440, diffraction information is determined by extracting pre-computed diffraction information of the known scene state from the bit stream or storage device. The storage device may involve local storage devices (e.g., memory, cache, file, etc.) or shared storage devices (e.g., cloud storage devices, server storage devices).

[0186] Extracting pre-computed diffraction information from a known scene state may include receiving a lookup table or lookup table entries from a bitstream (incoming bitstream) or a storage device. The lookup table can be considered a representation of the diffraction information. It may include multiple pre-computed diffraction information items, each associated with a corresponding known scene state. The pre-computed diffraction information and associated known scene states may correspond to the aforementioned data elements. The known scene state may include or indicate a known audio scene description and a known listener location.

[0187] Selecting the relevant entries from the received lookup table, or selecting the relevant entries to receive (if not all lookup tables, but only their entries), may involve using hash values, as described above.

[0188] On the other hand, if it is determined that the current scene state does not correspond to the known scene state (NO at step S430), then the method proceeds to step S450.

[0189] In step S450, based on the source location, listener location, and representation of the three-dimensional audio scene, a pathfinding algorithm is used to determine diffraction information. This can be done based on a reference... Figure 2 The process described is completed.

[0190] At step S460, the diffraction information obtained via step S440 or step S450 is output. This step can correspond to step S350 described above.

[0191] In summary, the proposed methods may (in particular) include the following:

[0192] • Check if the pre-computed diffraction path information (i.e., the pre-computed diffraction information) can be retrieved from the cache and reused for the current scene state and listener position.

[0193] • Check if the pre-computed diffraction path information (i.e., the pre-computed diffraction information) of the current scene state can be obtained from the memory cache or bitstream and reused.

[0194] The scene state is determined by the input parameter L of the function DiffractionDirectionCalculation(). vox S vox The VoxDataDiffractionMap definition includes functions such as pathfinding algorithms and voxel C. vox The selection and diffraction path length estimation steps. Diffraction path information (e.g., diffraction information) is transmitted via the output parameter C. vox r inDefinition. This diffraction path information (if available) can be obtained directly from the bitstream syntax voxSceneDiffractionPreComputedPathData() of the corresponding scene state, thus avoiding a call to the function DiffractionDirectionCalculation().

[0195] When considering the current scenario state L vox S vox VoxDataDiffractionMap obtains diffraction path information C vox r in This information can be cached in memory (and made available outside the renderer) for later reuse by the renderer (or other renderer instances).

[0196] In other words, the "diffraction pathfinding" according to this disclosure (e.g., implemented by method 300 and / or method 400) may involve the following processing:

[0197] - Check if the pre-computed diffraction path information of the current scene state can be obtained from the memory cache or bitstream and reused.

[0198] The scene state is determined by the input parameter L of the function DiffractionDirectionCalculation(). vox S vox The VoxDataDiffractionMap definition includes functions such as pathfinding algorithms and voxel C. vox The steps involve selecting and estimating the diffraction path length.

[0199]

[0200] Diffraction path information is transmitted via output parameter C vox r in Definition. This diffraction path information (if available) can be obtained directly from the bitstream syntax voxSceneDiffractionPreComputedPathData() of the corresponding scene state, thus avoiding a call to the function DiffractionDirectionCalculation().

[0201] - When considering the current scenario state L vox S vox VoxDataDiffractionMap obtains diffraction path information C vox r in This information can be cached in memory (and made available outside the renderer) for later reuse by the renderer (or other renderer instances).

[0202] In the above, the bitstream syntax definition can be written in the function() style of MPEG standard documents. It defines how to read / parse data (bitstream elements) from the bitstream. In this case, it is used to obtain the variables / information needed to recover the diffraction path information.

[0203] Figure 5 , Figure 6 and Figure 7 An example of a processing chain 500 based on the above is shown, which can be used to process audio scene information for audio rendering. Specifically, processing chain 500 can be used to convert voxel-related data into parameters and signals required for audibility.

[0204] Figure 5 This involves situations where the current scene state is unknown. Figure 1 Unlike the processing chain 100, the bitstream / memory 510 additionally includes data elements containing diffraction information and associated scene states, for example, in the form of a lookup table as described above.

[0205] Similar to processing chain 100, processing chain 500 receives audio scene description 20 from bitstream (or storage device / memory) 510. Processing chain 500 also receives indication of user location (listener location) 30 of the user (listener) within the audio scene.

[0206] The diffraction direction calculation block (diffraction calculation block) 40 for determining (e.g., calculating) diffraction information and the diffraction modeling tool 50 for applying diffraction modeling and optional occlusion modeling based on the diffraction information can be the same as those in the processing chain 100.

[0207] However, with Figure 1Unlike processing chain 100, audio scene description 20 and listener location 30 are used to determine scene state 515 or scene state identifier. This scene state 515 (e.g., scene state N1 as defined above) or scene state identifier (e.g., HASH(N1)) is provided / input to scene state analysis block 520, which determines whether the current scene state 515 corresponds to a known scene state 530 (yes at block 535) or does not (no at block 535). In this example, the current scene state 515 does not correspond to a known scene state (i.e., the current scene state 515 is an unknown scene state). Therefore, audio scene description 20 and listener location 30 are input to diffraction direction calculation block 40 in the same manner as used for processing chain 100 to generate diffraction information 550. The diffraction information 550 is then used for rendering / diffraction modeling, as in the case of processing chain 100. However, additionally, the diffraction information 550 (e.g., diffraction information N2 as defined above or its quantized version N3) is output via the interface for later reuse by the renderer or other (external) rendering instances. Specifically, the diffraction information 550, along with the corresponding scene state, can be output to the bitstream (or memory / storage device) 510 via the output interface 555.

[0208] Figure 6 The current scene state 515 refers to a known scene state.

[0209] Similarly, the current scene state 515 is provided / input to the scene state analysis block 520 to determine whether the current scene state 515 corresponds to the known scene state 530. In this example, the current scene state 515 corresponds to the known scene state. Therefore, instead of inputting the audio scene description 20 and the listener location 30 to the diffraction direction calculation block 40 to calculate / generate diffraction information, diffraction information is extracted / received from the bitstream (or storage device / memory) 510, as described above (e.g., via step S450 of method 400). Nevertheless, even if the diffraction information is not calculated locally, it can be output to the bitstream (or memory / memory) 510, as... Figure 5 That is the case. The reason is that if the diffraction information is obtained from one source (e.g., in a bitstream or storage device), then it can be provided to the corresponding (one or more) other sources in this way.

[0210] Figure 7 The complete processing chain 500 is shown, including data paths for known and unknown scenario states 515.

[0211] Figure 8This is a graph illustrating the complexity metric of different implementations of processing audio scene information or audio rendering over time, assuming a simple maze as the audio scene. Further, it is assumed that the user randomly navigates the maze, thus revisiting previously visited locations. Graph 810 represents the case where no pre-computed diffraction information is available (e.g., the bitstream does not provide diffraction information, and memory / caching is disabled). In this case, the computational load on the renderer is essentially constant and relatively high. Graph 820 represents the case where pre-computed diffraction information is locally available (e.g., the bitstream does not provide diffraction information, and local memory / caching is enabled). In this case, the processing load on the renderer decreases over time as more and more diffraction information items accumulate locally. In other words, the scene states encountered increasingly involve (locally) known scene states. Graph 830 finally represents the case where pre-computed diffraction information is provided externally (e.g., the bitstream provides complete diffraction information). In this scenario, the computational load on the renderer remains low because a large portion of the scene state involves known scene states, and diffraction information can be retrieved from external sources (e.g., from a bitstream or by requesting external / shared storage) without performing local computation.

[0212] Figure 9 Examples of possible use cases for the technology according to embodiments of this disclosure are illustrated schematically. Two listeners (users) A and B are shown in different locations within an audio scene (e.g., a house with different areas and floors). Users A and B may be users exploring a VR environment including the audio scene individually or together, for example, as part of a game, virtual tour, etc. Users exploring a public VR environment may be running, for example, social VR. With different listener locations within the audio scene, users A and B will produce different rendering results and different diffraction information. This disclosure anticipates that each user (or its corresponding device / decoder / renderer) makes their calculated diffraction information available to other users. Once user B enters the audio scene area previously occupied by user A, they may benefit from user A's pre-calculated diffraction information, and vice versa. For example, user A's diffraction information can be made available to user B via a LUT indexed with a corresponding scene state or scene state identifier. By exchanging diffraction information between different devices / decoders / renderers, depending on the user's movement patterns within the audio scene, the computational load on both user devices / decoders / renderers can be reduced.

[0213] Furthermore, since users (listeners) tend to behave similarly, diffraction information (diffraction data) is particularly focused on accumulating relevant (e.g., frequently occurring) scene states. This is very difficult to achieve for encoder-side pre-computation of diffraction information, as the encoder cannot access the actual listener locations and can only assume them. Moreover, using data storage devices (e.g., physical / shared storage or bitstream bandwidth) for encoder-side pre-computation is much less efficient, because in this case, some of the pre-computed diffraction information involves irrelevant or less relevant scene states.

[0214] For example, the proposed features and techniques can create LUTs that correspond to the user’s actual 6DoF behavior (rather than assumed behavior on the encoder side), and thus can be said to involve intelligent, user-oriented LUT creation.

[0215] Representation and encoding of acoustic path information

[0216] In the foregoing, methods for providing, outputting, storing, exchanging, and / or reusing pre-computed scene state information have been described. In conjunction with or beyond this, it may be of interest to provide an interface for providing, outputting, storing, exchanging, and / or reusing acoustic path information between encoders and decoders / renderers (e.g., encoder-decoder, decoder1–decoder2, or decoder1–decoder1). To do so, independent of the computational and exchange details of the acoustic path information, this disclosure provides a scheme for efficiently representing (e.g., encoding) acoustic path information, for example, for storage or transmission. Thus, this scheme can be used as a standalone scheme or in combination with techniques described elsewhere in this disclosure.

[0217] Here, acoustic path information can refer to initially pre-computed acoustic path information, such as for all feasible start and end point pairs (start and end voxels) in the voxel mesh (note that for some pairs, there may not be a valid path). In this case, the acoustic path information can be provided, for example, by the encoder. Alternatively, acoustic path information can refer to acoustic path information computed by the decoder / renderer at runtime, for example, for later reuse by the same decoder or for use by different decoders. This disclosure provides different modes for representing (e.g., encoding) acoustic path information, depending on the corresponding use case, such as depending on the amount and / or nature of the pre-computed acoustic path information.

[0218] In the reference model (RM) used to represent or encode acoustic path information, each acoustic diffraction path is encoded with "from-to" voxel coordinate pairs (e.g., user and object voxel coordinates).

[0219] • If an acoustic diffraction path exists for a given pair of voxel coordinates:

[0220] Then, the diffraction angle coordinates (determining the location of the diffraction source) and the diffraction path length (determining the level / gain of the diffraction source) are explicitly encoded and transmitted.

[0221] • Otherwise:

[0222] No additional information is encoded or transmitted.

[0223] However, the RM method used to represent the pre-computed acoustic path data (acoustic diffraction path data) has high redundancy because:

[0224] • Repetition of voxel coordinates for static scene states (in terms of user or object position);

[0225] as well as

[0226] • Redundant path lengths indicate accuracy (encoded as IEEE float32).

[0227] This disclosure addresses these two issues in the following ways:

[0228] • Avoid repetition of voxel coordinates (e.g., through the nested order of voxel pair encoding); and

[0229] • Apply differential diffraction path length encoding (e.g., by sequentially traversing the voxel grid and representing the path length by the number of voxels or the difference between the voxel count and the previous value).

[0230] Generally, acoustic path information may include one or more path information items (acoustic path information items), each relating to a corresponding acoustic path. Each path information item specifies the path length of the acoustic path between a first location, a second location, and the first and second locations in the voxel grid, as well as corner voxels. The first location (e.g., the first voxel) may relate to the starting location of the acoustic path (e.g., the starting voxel), such as the source location S defined above. VOX The second location (e.g., the second voxel) can refer to the end location of the acoustic path (e.g., the end voxel), such as the listener location L as defined above. VOXFor example, the start location can be given by the syntax element `pcpdStartVoxelPacked` as defined below, and the end location can be given by the syntax element `pcpdEndVoxelPacked` as defined below. In any case, depending on the use case and requirements, the assignment of the first and second locations to the start and end locations may be reversed in some implementations. A corner voxel can be a voxel on an acoustic path where the acoustic path changes direction. In some implementations, it may be additionally required that the corner voxel be visible from the second location (e.g., the end location), because there must be a line of sight between the second location and the corner voxel in the voxel mesh. If more than one voxel conforms to this definition, then the voxel that is farthest from the second location (or closest to the first location) among the conforming voxels can be designated as the corner voxel.

[0231] This disclosure seeks to efficiently encode sequences of path information items. Generally, when encoding a given path information item, the techniques of this disclosure seek to reuse information associated with a previous path information item. For example, even if one or both of the first and second locations differ from one path information item to the next, the corner voxels and / or path lengths may be the same or similar.

[0232] To increase the probability of reusing information among path information items, encoding and decoding are performed in a nested manner according to the techniques of this disclosure. That is, path information items in the path information sequence are grouped in the sequence according to their corresponding first locations. Furthermore, for each first location, the path information items are preferably grouped such that the corresponding second locations of adjacent path information items in the sequence are close to each other, for example, adjacent in a voxel grid.

[0233] An example implementation of the technology disclosed herein seeks to at least reuse information about corner voxels. Reference will now be made to... Figure 12 and Figure 13 Examples of methods for describing the corresponding encoded acoustic path information.

[0234] Figure 12 This is a flowchart illustrating a method 1200 for processing audio scene information (in particular, encoding acoustic path information including multiple path information items for a given voxel-based audio scene representation).

[0235] At step S1210, a voxel-based audio scene representation of the audio scene is obtained (e.g., received, extracted from a bitstream, read from a storage device, etc.).

[0236] At step S1220, for one or more (e.g., multiple) first locations and one or more (e.g., multiple) second locations in a two-dimensional voxel grid associated with the voxel-based audio scene representation, path information items are sequentially encoded for the corresponding first and second locations. That is, the path information items can be arranged in a given sequence (e.g., a predefined sequence) and can be encoded one after another according to that sequence. This sequence can be such that path information items specifying the same first location are next to each other, i.e., for subsequences within the sequence. In this sense, the above sequence can be described as involving a nested traversal of associated first and second locations, where path information items specifying the same first location are grouped. Preferably, for each such group, the second locations specified by the path information items in that group are traversed according to a predefined pattern, as described in more detail below.

[0237] As described above, each path information item can specify the first location, the second location, the path length of the acoustic path between the first and second locations in the voxel mesh, and the corner voxels where the acoustic path changes direction. It should be understood that the voxel mesh can involve a two-dimensional projection map generated from a voxel-based audio scene representation.

[0238] At step S1230, for the current path information item, an encoded path information item is generated based on the (current) path information item. The generated encoded path information item includes at least an indication of the corresponding first location and an indication of the corresponding second location. It may include additional encoded information, as described in detail below.

[0239] The further encoding process for the current path information item depends on whether the corner voxel specified by the current path information item is different from the corner voxel specified by the previous path information item (i.e., the previous one in the sequence) (and therefore, the further content of the encoded path information item is also different). It should be understood that method 1200 may include a step of determining whether this is the case (not shown in the figure).

[0240] Step S1240 involves the case where the corner voxel specified by the current path information item is different from the corner voxel specified by the previous path information item. The encoded path information item then includes an indication of the corner voxel's location. In other words, the indication of the corner voxel is included (or added to) the encoded path information item. This indication of the corner voxel can be, for example, an absolute, explicit, and / or non-differential indication of the corner voxel using its voxel coordinates or voxel index. This indication can involve, for example, the syntax element `pcpdSourceDirectionPacked` as defined below. Furthermore, the encoded path information item can include an indication of a corner voxel that is different from the corner voxel in the previous path information item, for example, in the form of a single bit flag. This bit flag can involve the flag `pcpdUsePrevSourceDirection == "false"` as defined below (in the first mode, such as the selective mode defined below), or the flag `pcpdUsePrevData == "false"` as defined below (in the second mode, such as the full mode defined below).

[0241] Step S1250 involves the case where the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item (i.e., the corner voxel specified by the current path information item is exactly the same as the corner voxel specified by the previous path information item). The encoded path information item then includes an indication that "the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item," rather than an indication of the corner voxel itself. In other words, the indication that "the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item" is included (or added to) the encoded path information item. This indication may involve, for example, a single bit flag. This bit flag may involve the flag pcpdUsePrevSourceDirection == "true" (in the first mode, e.g., selective mode), or the flag pcpdUsePrevData == "true" (in the second mode, e.g., full mode), as defined below.

[0242] Method 1200 may also include the step of outputting encoded path information items to a bit stream (not shown in the figure), for example for transmission or storage.

[0243] Starting with the techniques described above, this disclosure provides two different modes (encoding modes) that can be selected for encoding acoustic path information: a first mode (e.g., a selective mode), which can be used, for example, when only a relatively small number of path information items need to be encoded; and a second mode (e.g., a full mode), which can be used, for example, when a large number of path information items need to be encoded (e.g., for all feasible first and second location pairs in an acoustic scene). Whether the first mode or the second mode is used for a path information item (or a sequence of path information items) can be signaled by a flag in the bitstream (valid for the entire sequence) or included in the encoded path information item. This flag can be, for example, the flag pcpdFullMode defined below.

[0244] In the first mode (e.g., pcpdFullMode == "false"), the generated encoded path information item also includes an indication of the path length (e.g., pcpdPathLength as defined below). This path length indication can be, for example, an absolute, explicit, and / or non-differential indication of the path length in units of voxels or voxel side lengths. Therefore, the first mode can reuse previous corner voxel indications, but can include the path length indication regardless of whether the corner voxels change from one path information item to the next.

[0245] As an alternative implementation to the first mode, if the corner voxel specified by the current path information entry is the same as the corner voxel specified by the previous path information entry, then the generated encoded path information entry may include an indication (e.g., a 1-bit flag) of whether the encoded path information entry includes an indication of the path length or an indication of the difference between the path length specified by the current path information entry and the path length specified by the previous path information entry. Depending on this indication, the encoded path information entry may then include either an indication of the path length or an indication of the difference between the path length specified by the current path information entry and the path length specified by the previous path information entry. The latter may involve, for example, a 2-bit value.

[0246] In the second mode (e.g., pcpdFullMode == "true"), the generated encoded path information entry does not necessarily include an indication of the path length. That is, in the second mode, if the corner voxel specified by the current path information entry is the same as the corner voxel specified by the previous path information entry, then the encoded path information entry also includes an indication of the difference between the path length specified by the current path information entry and the path length specified by the previous path information entry. This indication can be in the form of, for example, the syntax element pcpdPathLengthDelta defined below.

[0247] For path information items in a specific order, the differences between the aforementioned path lengths can be encoded very efficiently. Therefore, for encoding in the second mode, for each first location, the voxel grid can be traversed according to a predetermined pattern to determine the sequence of second locations, and for the corresponding first location, path information items are encoded sequentially for the determined sequence of second locations. That is, the sequence of path information items is determined by traversing the voxel grid, and encoding is performed according to this sequence.

[0248] An example of this predefined pattern is to traverse the voxel grid along rows and columns using a raster scan.

[0249] When using this mode, the difference in path lengths described above can be encoded very efficiently, using only 2 bits. In other words, the difference can (only) take one of four predetermined values, as these four predetermined values ​​are sufficient to encode the difference (assuming the corner voxels between the current and previous path information items are the same). Examples of these four predetermined values ​​are given in Table 3 below, where it is assumed that the voxel side length is 1.

[0250] Consistent with the above, the bitstream may include a sequence of encoded path information items, indicating whether a first mode or a second mode is used (e.g., the bit flag pcpdFullMode defined below). Furthermore, each encoded path information item may have the following:

[0251] • Indication of the first location (e.g., pcpdStartVoxelPacked as defined below)

[0252] • Instructions for a second location (e.g., pcpdEndVoxelPacked as defined below) and

[0253] • An indication of whether an acoustic path exists at the first and second locations (e.g., the bit flag pcpdPathExists defined below).

[0254] If the first mode is used (e.g., selective mode) (e.g., pcpdFullMode == "false"), then the encoded path information items also include:

[0255] • An indication of whether the corner voxel specified by the path information entry corresponding to the encoded path information entry is the same as the corner voxel specified by the previous path information entry (e.g., a 1-bit flag, such as the bit flag pcpdUsePrevSourceDirection defined below).

[0256] • If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item (e.g., pcpdUsePrevSourceDirection == "true"), then include an indication of the path length (e.g., pcpdPathLength as defined below).

[0257] • Otherwise (e.g., pcpdUsePrevSourceDirection == "false"), including indications of corner voxels (e.g., pcpdSourceDirectionPacked as defined below) and indications of path length (e.g., pcpdPathLength as defined below).

[0258] In an alternative embodiment of the first mode (e.g., selective mode), the encoded path information item may further include (instead of the above):

[0259] • An indication of whether the corner voxel specified by the path information entry corresponding to the encoded path information entry is the same as the corner voxel specified by the previous path information entry (e.g., a 1-bit flag, such as the bit flag pcpdUsePrevSourceDirection defined below).

[0260] • If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item (e.g., pcpdUsePrevSourceDirection == "true"), then the path length is an indication of whether it is non-differentially encoded (e.g., in an absolute item or explicitly) or differentially encoded (e.g., in a relative item, as an incremental value) (e.g., a 1-bit flag).

[0261] ○ If the path length is non-differential encoded, then the path length is indicated (non-differential, absolute, or explicit) (e.g., pcpdPathLength as defined below).

[0262] ○ If the path length is differentially encoded, then it is an indication (e.g., a 2-bit value) of the difference between the previous path length and the current path length, where the previous path length is the path length specified by the previous path information entry, and the current path length is the path length specified by the path information entry corresponding to the encoded path information entry;

[0263] • Otherwise, if the corner voxel specified by the path information item corresponding to the encoded path information item is not the same as the corner voxel specified by the previous path information item (e.g., pcpdUsePrevSourceDirection == "false"), then the corner voxel indicator (e.g., pcpdSourceDirectionPacked as defined below) and the path length indicator (e.g., pcpdPathLength as defined below) are used.

[0264] If the second mode (e.g., full mode) is used (e.g., pcpdFullMode == "true"), then the encoded path information items also include:

[0265] • An indication of whether the corner voxel specified by the path information entry corresponding to the encoded path information entry is the same as the corner voxel specified by the previous path information entry (e.g., the bit flag pcpdUsePrevData defined below).

[0266] • If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item (e.g., pcpdUsePrevData == "true"), then an indication of the difference between the previous path length and the current path length (e.g., a 2-bit value, such as pcpdPathLengthDelta as defined below), where the previous path length is the path length specified by the previous path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item;

[0267] • Otherwise (e.g., pcpdUsePrevData == "false"), then the corner voxel is indicated (e.g., pcpdSourceDirectionPacked as defined below) and the path length is indicated (e.g., pcpdPathLength as defined below).

[0268] Figure 13 This is a flowchart illustrating a method 1300 for processing audio scene information (specifically, decoding acoustic path information including multiple (encoded) path information items) for a given voxel-based audio scene representation. Decoding method 1300 may include steps mirroring those of the corresponding encoding method 1200. Therefore, it should be understood that the data elements mentioned below (e.g., indicators, flags, etc.) may correspond to those data elements defined above in the context of method 1200. In other words, the bitstream received by method 1300 may be the bitstream output by method 1200.

[0269] At step S1310, a bit stream is received. The bit stream includes a sequence of path information items for encoding one or more first locations and one or more second locations in a two-dimensional voxel grid associated with a voxel-based audio scene representation.

[0270] Each encoded path information item corresponds to a specific path information item, which specifies the path length of the acoustic path between the first location, the second location, the first location and the second location in the voxel grid, and the corner voxels where the acoustic path changes direction. It should be understood that the bitstream received in this step can be a bitstream generated or output by method 1200 as described above.

[0271] In step S1320, the encoded path information items are decoded sequentially to generate (e.g., recover) the corresponding path information items.

[0272] As described above, the encoded path information items can be arranged in a given sequence (e.g., a predefined sequence) in the bitstream, and they can be decoded one after another according to this sequence. This sequence can be such that encoded path information items specifying the same first location are grouped together, i.e., for subsequences within a sequence. In this sense, the above sequence can be described as involving a nested traversal of related first and second locations, where encoded path information items specifying the same first location are grouped. Preferably, for each such group, the second location specified by the encoded path information items in that group is traversed according to a predefined pattern, as described above.

[0273] The further decoding process for the current path information item depends on whether the currently encoded path information item includes an indication that "the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item".

[0274] Therefore, in step S1330, for the currently encoded path information item, it is determined whether the currently encoded path information item includes an indication that "the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item". This indication may be, for example, a 1-bit flag. This bit flag may relate to the flag pcpdUsePrevSourceDirection (in the first mode, such as selective mode) or to the flag pcpdUsePrevData (in the second mode, such as full mode).

[0275] Step S1340 involves the case where the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item (e.g., pcpdUsePrevSourceDirection == "true" in the first mode or pcpdUsePrevData == "true" in the second mode). In this case, the corner voxel specified by the path information item corresponding to the previously encoded path information item is set as the corner voxel of the path information item corresponding to the currently encoded path information item. That is, the corner voxel of the previously decoded path information item is reused as the corner voxel of the currently decoded path information item.

[0276] Step S1350 involves cases where the corner voxel specified by the corresponding path information item differs from the corner voxel specified by the path information item corresponding to the previously encoded path information item (e.g., pcpdUsePrevSourceDirection == "false" in the first mode or pcpdUsePrevData == "false" in the second mode). In this case, the (absolute, explicit, and / or non-differential) indication of the corner voxel (e.g., pcpdSourceDirectionPacked) is extracted from the currently encoded path information item.

[0277] As described above, this disclosure provides two different modes that can be selected for encoding and decoding acoustic path information: a first mode (e.g., a selective mode), which can be used, for example, when only a relatively small number of path information items need to be encoded; and a second mode (e.g., a full mode), which can be used, for example, when a large number of path information items need to be encoded (e.g., for all feasible first and second location pairs in the acoustic scene).

[0278] In the first mode, the encoded path information item obtained from the bitstream also includes an indication of the path length (e.g., pcpdPathLength), as described above. This path length indication can be, for example, an absolute, explicit, and / or non-differential indication of the path length in units of voxels or voxel side lengths. Therefore, generating a path information item corresponding to the currently encoded path information item also includes extracting the path length indication from the currently encoded path information item.

[0279] In the second mode, the encoded path information item does not necessarily include an indication of the path length, but if the corner voxel remains unchanged, it may instead include an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item (e.g., pcpdPathLengthDelta).

[0280] Accordingly, if the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the decoding in the second mode (e.g., in step S1340) may further include an indication to extract the difference between the previous path length and the current path length. Here, the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the currently encoded path information item.

[0281] As described above, one or more second locations may refer to locations obtained by traversing a voxel grid according to a predetermined pattern for each first location to define a sequence of second locations, and for each corresponding first location. It should then be understood that the encoded path information items are decoded sequentially according to the sequence of second locations.

[0282] For example, the predetermined pattern can be traversed along the rows and columns of the voxel grid using a raster scan. In this case, the aforementioned difference can be encoded with 2 bits, or in other words, it can (only) take one of four predetermined values.

[0283] Next, an example implementation of the above scheme will be described.

[0284] Consistent with the above, this implementation supports encoding pre-computed acoustic path information (e.g., acoustic diffraction path data) in the following modes:

[0285] • “Full” scene data mode, corresponding to the second mode above (e.g., full mode), in which diffraction data is pre-computed for each voxel of the scene (i.e., for low-complexity rendering scenes).

[0286] as well as

[0287] • “Selective” scene data mode, corresponding to the first mode mentioned above (e.g., selective mode), in which diffraction data is computed for a subset of voxels of the scene.

[0288] These two modes provide the ability to support applications in different scenarios coded according to this disclosure, such as:

[0289] • “Full” mode: “Encoder to renderer” (e.g., all data is pre-computed and then used by the renderer);

[0290] as well as

[0291] • “Selective” mode: “Encoder / renderer to renderer” (e.g., some data is pre-computed, but new data can be added or swapped during the rendering process).

[0292] Correspondingly, the "selective" mode offers more real-time-dependent capabilities (e.g., one renderer can act as an encoder for another renderer), but the "full" mode typically provides smaller encoded data size and / or lower rendering complexity.

[0293] In a specific example, the bitstream syntax (data interface) can be defined as follows:

[0294] if(pcpdDataPresent) {1 bit

[0295] pcpdNumStartPostions = escapedValue(8,16,32)

[0296] if(!pcpdFullMode) { / / Selective mode 1 bit

[0297] for(int i = 0; i < pcpdNumStartPostions; ++i) {

[0298] pcpdNumEndPositions = escapedValue(8,16,32)

[0299] pcpdStartVoxelPacked;NBitsMap

[0300] for(int j = 0; j < pcpdNumEndPositions; ++j) {

[0301] pcpdEndVoxelPacked;NBitsMap

[0302] if(pcpdPathExists) {1 bit

[0303] / / Differential coding

[0304] if(!pcpdUsePrevSourceDirection) {1 bit

[0305] pcpdSourceDirectionPacked;NBitsMap

[0306] }

[0307] pcpdPathLength; 32-bit float

[0308] }

[0309] }

[0310] }

[0311] } else { / / Full mode

[0312] for(int i = 0; i < pcpdNumStartPostions; ++i) {

[0313] if(pcpdDataPresentForPos) {1 bit

[0314] pcpdStartVoxelPacked;NBitsMap

[0315] for(int x = 1; x <= voxSceneDimensions[0]; ++x) {

[0316] for(int y = 1; y <= voxSceneDimensions[1]; ++y) {

[0317] / / pcpdEndVoxelPacked is defined by x and y

[0318] if(pcpdPathExists) {1 bit / / Differential encoding}

[0319] if (!pcpdUsePrevData) {1 bit

[0320] pcpdSourceDirectionPacked;NBitsMap

[0321] pcpdPathLength; 32-bit float

[0322] } else {

[0323] pcpdPathLengthDelta; 2 bits

[0324] }

[0325] }

[0326] }

[0327] }

[0328] }

[0329] }

[0330] }

[0331] }

[0332] Note: NbitsMap = ceil(log2(voxSceneDimensions[0]*voxSceneDimensions[1]-1)

[0333] The bitstream syntax described above can be considered a replacement for the bitstream syntax given in Table 1 above.

[0334] Table 3 — Values ​​of pcpdPathLengthDelta

[0335]

[0336] For the example bitstream syntax above, the encoding and decoding processes can be performed as follows.

[0337] The encoder (or renderer) can select a pre-computed acoustic diffraction path data encoding mode (signaled by pcpdFullMode or other suitable flags) based on the compression performance produced or desired for the current application or scene. If an acoustic diffraction path exists in a "selective" encoding mode (e.g., pcpdFullMode == "false"), and previous data is unavailable (e.g., pcpdUsePrevSourceDirection == "false"), then the coordinates (e.g., packed coordinates) and path length data of the corner voxels are explicitly read from the bitstream (e.g., in RM); on the other hand, if previous data is available (e.g., pcpdUsePrevSourceDirection == "true"), then the corner coordinate data of the last transmitted voxels is used in the calculation. It is worth noting that in the "selective" encoding mode, the path length must be transmitted for each path (e.g., if pcpdPathExists == "true").

[0338]

[0339] Alternatively, in the “selective” encoding mode, if previous data is available, the bit stream may include an indication (e.g., a 1-bit flag) of whether the path length is explicitly transmitted or can be differentially encoded (e.g., using a 2-bit value) with reference to previous data.

[0340] If an acoustic diffraction path exists in “full” encoding mode (e.g., pcpdFullMode == “true”) and previous data is unavailable (e.g., pcpdUsePrevData == “false”), then the coordinates (e.g., packed coordinates) and path length data of the corner voxels are explicitly read from the bitstream (as in RM); if previous data is available (e.g., pcpdUsePrevData == “true”), then the last transmitted voxel corner and path length values ​​are used in the calculation.

[0341]

[0342] The path length difference (Delta) is encoded with 2 bits (e.g., via pcpdPathLengthDelta) because the path length difference with the previous path length can be + / -1 or + / -d, where d = sqrt(2) - 1; this is because the path length is calculated as the sum of the horizontal and diagonal steps on a uniform voxel grid.

[0343] Next, the data compression performance estimation results according to the technology disclosed herein will be described.

[0344] Table 4 and Figure 14 The image shows a bitrate comparison for all MPEG-I CfP Test1 scenes (VoxData payload only). The decoder output maintains bit accuracy relative to the RM. Overall, Test1 bitrate savings are approximately 38% (full mode) and approximately 16% (selective mode).

[0345] condition:

[0346] - RM: Current Reference Model (v25 bitstream)

[0347] - PCPD (Full): Pre-computed path data encoding (full mode)

[0348] - PDPD (Selective): Pre-computed path data encoding (selective mode)

[0349] Table 4 — Data Compression Performance Results

[0350]

[0351] Representation of voxel coordinates / index

[0352] The following efficient representation of voxel indexes can be used for the transmission or storage of, for example, voxel grids and diffraction map entries. It can replace any fixed-length representation of voxel indexes (voxel coordinates).

[0353] The following steps can be performed in the context of the proposed representation:

[0354] Step 1: Determine the number of bits (i.e., quantity, count) required for the current grid resolution / diffraction pattern dimension. For 3D voxel grids and 2D diffraction patterns, these numbers NbitsVox and NbitsMap can be determined, for example, as follows:

[0355]

[0356] Where L, W, and H (length, width, and height) are the dimensions of the voxel grid and diffraction pattern. The values ​​for the voxel grid and diffraction pattern may differ. Step 1 can be applied to both the encoder and decoder sides.

[0357] Step 2: Map the voxel indices (x, y, z) and diffraction pattern indices (x, y) to the packed representation indices (Idx), and encode them using Nbits_vox and Nbits_map bits respectively. In one embodiment, (x, y, z) is zero-based, and the packed representation indices can be in the range of 0 to L*W*H-1 for voxels and 0 to L*W-1 for diffraction patterns. The mapping from indices (x, y, z) to packed representation indices can be as follows:

[0358]

[0359] Step 2 can be performed only on the encoder side.

[0360] In the foregoing, a packing representation index is an index that can uniquely identify a voxel in a diffraction pattern or voxel grid. In other words, voxels in a voxel grid can be assigned unique, consecutive indices such that each voxel in the voxel grid can be uniquely identified by a single integer. Accordingly, the packing representation index can be used for any indication of a voxel location in a voxel grid or a two-dimensional plot. In particular, the packing representation index can be used to indicate any voxel location mentioned throughout this disclosure.

[0361] Unique indices can be assigned to voxels according to a predefined pattern. For example, the voxel grid can be scanned / traversed sequentially along the x, y, and z directions to continuously assign unique indices to the corresponding voxels.

[0362] Mapping from the packed representation index back to the voxel and diffraction map indexes can be done, for example, as follows:

[0363]

[0364] The % operator represents the modulo operator.

[0365] Device

[0366] While methods and processing chains have been described above, it should be understood that this disclosure also relates to apparatus (e.g., computer apparatus or apparatus generally having processing capabilities) for implementing these methods and processing chains (or techniques in general).

[0367] Figure 15 An example of such a device 1500 is schematically illustrated. Device 1500 includes a processor 1501 and a memory 1502 coupled to the processor 1501. Memory 1502 may store instructions for execution by the processor 1501. The processor 1501 may be adapted to implement a processing chain throughout the present disclosure and / or perform methods throughout the present disclosure (e.g., a method of processing audio scene information for audio rendering). Device 1500 may receive input (e.g., an audio scene description, listener location, etc.) and generate output (e.g., a representation of diffraction information, acoustic path information, etc.).

[0368] explain

[0369] The aspects of the systems described herein can be implemented in a suitable computer-based sound processing network environment (e.g., a server or cloud environment) for processing digital or digitized audio files. Parts of these systems may include one or more networks comprising any desired number of independent machines, including one or more routers (not shown) for buffering and routing data between computers. Such networks can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0370] One or more of the components, modules, processes, or other functional components may be implemented by a computer program executed by a processor-based computing device controlling the system. It should also be noted that, in terms of their behavior, register transfers, logical components, and / or other characteristics, the various functions disclosed herein can be described using any number of hardware, firmware combinations, and / or as data and / or instructions contained in various machine-readable or computer-readable media. Computer-readable media that may contain such formatted data and / or instructions include, but are not limited to, various forms of physical (non-transitory), non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0371] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules, which, for ease of discussion, may be described and illustrated as if most components were implemented solely in hardware. However, those skilled in the art, upon reading this detailed description, will recognize that in at least one embodiment, the electronic aspects may be implemented by software (e.g., stored on a non-transitory computer-readable medium) executable by one or more electronic processors (such as microprocessors and / or application-specific integrated circuits (“ASICs”). Thus, it should be noted that embodiments may be implemented using multiple hardware and software-based devices and multiple different structural components. For example, the computer-implemented neural network described herein may include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., system buses) connecting the various components.

[0372] While one or more embodiments have been described by way of example and specific examples, it should be understood that one or more embodiments are not limited to the disclosed embodiments. Rather, as those skilled in the art will appreciate, they are intended to cover various modifications and similar arrangements. Therefore, the scope of the appended claims should be given the broadest interpretation to cover all such modifications and similar arrangements.

[0373] Furthermore, it should be understood that the wording and terminology used herein are for descriptive purposes only and should not be considered limiting. The use of “including,” “comprising,” or “having,” and variations thereof is intended to cover the items listed thereafter and their equivalents, as well as additional items. Unless otherwise specified or limited, the terms “installation,” “connection,” “support,” and “coupling,” and variations thereof, are used extensively and cover both direct and indirect installation, connection, support, and coupling.

[0374] Example implementation of enumeration

[0375] Various aspects and implementations of the invention can also be appreciated from the following enumerated exemplary embodiments (EEE), which are not claims.

[0376] EEE1. A method for processing audio scene information, the method comprising:

[0377] Obtain a voxel-based audio scene representation;

[0378] For one or more first locations and one or more second locations in a 2D voxel grid associated with a voxel-based audio scene representation, path information items for the corresponding first and second locations are encoded sequentially, wherein each path information item specifies the path length of the acoustic path between the first location, the second location, the first location and the second location in the voxel grid, and the corner voxels where the acoustic path changes direction along the acoustic path; and

[0379] For the current path information item, generate an encoded path information item based on that path information item.

[0380] The coded path information includes indications for the corresponding first location and indications for the corresponding second location.

[0381] If the corner voxel specified by the current path information item is different from the corner voxel specified by the previous path information item, then the encoded path information item includes an indication of the corner voxel.

[0382] Specifically, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item includes an indication that "the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item", rather than an indication of the corner voxel.

[0383] EEE2. According to the method of EEE1, the encoded path information item further includes an indication of the path length.

[0384] EEE3. According to the method of EEE1, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item also includes an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item.

[0385] EEE4. According to the method of EEE1, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item further includes:

[0386] An indication of whether the encoded path information item includes an indication of the path length or an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item; and

[0387] An indication of the path length or the difference between the path length specified by the current path information item and the path length specified by the previous path information item.

[0388] EEE5. The method according to any of the preceding EEEs further includes:

[0389] For each first location, the voxel grid is traversed according to a predetermined pattern to determine the sequence of second locations, and for the corresponding first location, the path information items of the determined sequence of second locations are encoded sequentially.

[0390] EEE6. According to the method of EEE5, wherein the predetermined pattern traverses the voxel grid along the rows and columns of the voxel grid in a raster scan manner.

[0391] EEE7. According to the method described in EEE5 or EEE6, when subordinate to EEE3.

[0392] The difference is encoded by 2 bits; or

[0393] The difference is one of four predetermined values.

[0394] EEE8. The method according to any of the preceding EEE methods, wherein each encoded path information item includes:

[0395] Instructions on whether to use the first mode or the second mode;

[0396] Instructions for the first location;

[0397] Instructions for the second location; and

[0398] Regarding the indication of whether there is an acoustic path at the first and second locations,

[0399] If the first mode is used, then each encoded path information item also includes:

[0400] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item;

[0401] If the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the previous path information item, then the path length is indicated; and

[0402] Otherwise, the corner voxel indicator and the path length indicator; and

[0403] If the second mode is used, then each encoded path information item also includes:

[0404] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item;

[0405] If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then the indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the previous path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item; and

[0406] Otherwise, the corner voxel indication and the path length indication.

[0407] EEE9. The method according to any one of EEE1 to EEE7, wherein each encoded path information item includes:

[0408] Instructions on whether to use the first mode or the second mode;

[0409] Instructions for the first location;

[0410] Instructions for the second location; and

[0411] Regarding the indication of whether there is an acoustic path at the first and second locations,

[0412] If the first mode is used, then each encoded path information item also includes:

[0413] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item;

[0414] If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then the indication of "whether the encoded path information item includes an indication of the path length or an indication of the difference between the previous path length and the current path length" is provided, where the previous path length is the path length specified by the previous path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item, together with the indication of the path length or the indication of the difference between the previous path length and the current path length; and

[0415] Otherwise, the corner voxel indicator and the path length indicator; and

[0416] If the second mode is used, then each encoded path information item also includes:

[0417] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item;

[0418] If the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the previous path information item, then the indication of the difference between the previous path length and the current path length; and

[0419] Otherwise, the corner voxel indication and the path length indication.

[0420] EEE10. The method according to any one of EEE1 to EEE9 further includes:

[0421] Output the encoded path information items to the bit stream.

[0422] EEE11. A method for processing audio scene information, the method comprising:

[0423] Receive a bitstream comprising a sequence of encoded path information items for one or more first locations and one or more second locations in a two-dimensional voxel grid associated with a voxel-based audio scene representation, each encoded path information item corresponding to a corresponding path information item, the path information item specifying the first location, the second location, the path length of the acoustic path between the first location and the second location in the voxel grid, and the corner voxels along the acoustic path where the acoustic path changes direction; and

[0424] The encoded path information items are decoded sequentially to generate the corresponding path information items;

[0425] For the path information item currently encoded, the corresponding path information items generated include:

[0426] Determine whether the currently encoded path information item includes an indication that "the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item";

[0427] If the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the corner voxel specified by the path information item corresponding to the previously encoded path information item is set as the corner voxel of the path information item corresponding to the currently encoded path information item; and

[0428] If the corner voxel specified by the corresponding path information item is different from the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the indication of the corner voxel is extracted from the currently encoded path information item.

[0429] EEE12. According to the method described in EEE11, generating the corresponding path information item further includes:

[0430] Extract the path length indication from the currently encoded path information item.

[0431] EEE13. According to the method of EEE11, wherein if the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then generating the corresponding path information item further includes:

[0432] An indication to extract the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the currently encoded path information item.

[0433] EEE14. According to the method of EEE11, wherein if the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then generating the corresponding path information item further includes:

[0434] Extract the indication of whether the encoded path information item includes an indication of the path length or an indication of the difference between the previous path length and the current path length, where the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the currently encoded path information item; and

[0435] An indication of the path length or the difference between the previous path length and the current path length.

[0436] EEE15. The method according to any one of EEE11 to EEE14, wherein the one or more second locations relate to locations obtained by traversing a voxel grid according to a predetermined pattern for each first location to define a sequence of second locations, and for a corresponding first location, an encoded path information item is decoded sequentially according to the sequence of second locations.

[0437] EEE16. According to the method of EEE15, wherein the predetermined pattern traverses the voxel grid along the rows and columns of the voxel grid in a raster scan manner.

[0438] EEE17. According to the method described in EEE15 or EEE16, when subordinate to EEE13.

[0439] The difference is encoded by 2 bits; or

[0440] The difference is one of four predetermined values.

[0441] EEE18. The method according to any one of EEE11 to EEE17, wherein each encoded path information item includes:

[0442] Instructions on whether to use the first mode or the second mode;

[0443] Instructions for the first location;

[0444] Instructions for the second location; and

[0445] Indication regarding the existence of acoustic paths at the first and second locations;

[0446] If the first mode is used, then each encoded path information item also includes:

[0447] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item;

[0448] If the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then the path length is indicated; and

[0449] Otherwise, the corner voxel indicator and the path length indicator; and

[0450] If the second mode is used, then each encoded path information item also includes:

[0451] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item;

[0452] If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item; and

[0453] Otherwise, the corner voxel indication and the path length indication.

[0454] EEE19. The method according to any one of EEE11 to EEE17, wherein each encoded path information item includes:

[0455] Instructions on whether to use the first mode or the second mode;

[0456] Instructions for the first location;

[0457] Instructions for the second location; and

[0458] Indication regarding the existence of acoustic paths at the first and second locations;

[0459] If the first mode is used, then each encoded path information item also includes:

[0460] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item;

[0461] If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then the indication of "whether the encoded path information item includes an indication of the path length or an indication of the difference between the previous path length and the current path length" is provided, where the previous path length is the path length specified by the path information item corresponding to the previous encoded path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item, together with the indication of the path length or the indication of the difference between the previous path length and the current path length; and

[0462] Otherwise, the corner voxel indicator and the path length indicator; and

[0463] If the second mode is used, then each encoded path information item also includes:

[0464] An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item;

[0465] If the corner voxel specified by the path information entry corresponding to the encoded path information entry is the same as the corner voxel specified by the path information entry corresponding to the previously encoded path information entry, then the indication of the difference between the previous path length and the current path length; and

[0466] Otherwise, the corner voxel indication and the path length indication.

[0467] EEE20. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to perform the method according to any one of EEE1 to EEE19.

[0468] EEE21. A program comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of EEE1 to EEE19.

[0469] EEE22. A computer-readable storage medium for storing a program of EEE21.

Claims

1. A method for processing audio scene information, the method comprising: Obtain a voxel-based audio scene representation; For one or more first locations and one or more second locations in a two-dimensional voxel grid associated with a voxel-based audio scene representation, path information items for the corresponding first and second locations are encoded in sequence, wherein each path information item specifies the path length of the acoustic path between the first location, the second location, the first location and the second location in the voxel grid, and the corner voxel on the acoustic path where the direction of the acoustic path changes. as well as For the current path information item, generate an encoded path information item based on that path information item. The coded path information includes indications for the corresponding first location and indications for the corresponding second location. Wherein, if the corner voxel specified by the current path information item is different from the corner voxel specified by the previous path information item, the encoded path information item includes an indication of the corner voxel. Wherein, if the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, the encoded path information item includes an indication that the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, rather than an indication of the corner voxel.

2. The method of claim 1, wherein the encoded path information item further includes an indication of path length.

3. The method according to claim 1, wherein, If the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, the encoded path information item also includes an indication of the difference between the path length specified by the current path information item and the path length specified by the previous path information item.

4. The method according to claim 1, wherein, If the corner voxel specified by the current path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item also includes: The encoded path information entry includes either an indication of the path length or an indication of the difference between the path length specified by the current path information entry and the path length specified by the previous path information entry; and An indication of the path length or the difference between the path length specified by the current path information item and the path length specified by the previous path information item.

5. The method according to any one of the preceding claims, further comprising: For each first location, the voxel grid is traversed according to a predetermined pattern to determine the sequence of second locations, and for the corresponding first location, the path information items of the determined sequence of second locations are encoded sequentially.

6. The method of claim 5, wherein the predetermined pattern traverses the voxel grid along rows and columns in a raster scan manner.

7. The method according to claim 5 or 6 when dependent on claim 3, The difference is encoded by 2 bits; or The difference is one of four predetermined values.

8. The method according to any one of the preceding claims, wherein each encoded path information item comprises: Instructions on whether to use the first mode or the second mode; Instructions for the first location; Instructions for the second location; as well as Regarding the indication of whether there is an acoustic path at the first and second locations, If the first mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item; If the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the previous path information item, then the path length is indicated; and Otherwise, the corner voxel indication and the path length indication; and If the second mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item; If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then the indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the previous path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item; and Otherwise, the corner voxel indication and the path length indication.

9. The method according to any one of claims 1 to 7, wherein each encoded path information item comprises: Instructions on whether to use the first mode or the second mode; Instructions for the first location; Instructions for the second location; as well as Regarding the indication of whether there is an acoustic path at the first and second locations, If the first mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item; If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item, then the encoded path information item includes either an indication of the path length or an indication of the difference between the previous path length and the current path length, together with the indication of the path length or the indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the previous path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item; and Otherwise, the corner voxel indication and the path length indication; and If the second mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the previous path information item; If the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the previous path information item, then the indication of the difference between the previous path length and the current path length; and Otherwise, the corner voxel indication and the path length indication.

10. The method according to any one of claims 1 to 9, further comprising: Output the encoded path information items to the bit stream.

11. A method for processing audio scene information, the method comprising: Receive a bitstream comprising a sequence of encoded path information items for one or more first locations and one or more second locations in a two-dimensional voxel grid associated with a voxel-based audio scene representation, each encoded path information item corresponding to a corresponding path information item specifying the path length of the acoustic path between the first location, the second location, the first location and the second location in the voxel grid, and the corner voxels on the acoustic path where the direction of the acoustic path changes. as well as The encoded path information items are decoded sequentially to generate the corresponding path information items; For the path information item currently encoded, the corresponding path information items generated include: Determine whether the currently encoded path information item includes an indication that is the same as the corner voxel specified by the corresponding path information item and the corner voxel specified by the path information item corresponding to the previous encoded path information item; If the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the corner voxel specified by the path information item corresponding to the previously encoded path information item is set as the corner voxel of the path information item corresponding to the currently encoded path information item; and If the corner voxel specified by the corresponding path information item is different from the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the indication of the corner voxel is extracted from the currently encoded path information item.

12. The method according to claim 11, wherein generating the corresponding path information item further includes: Extract the path length indication from the currently encoded path information item.

13. The method according to claim 11, wherein, If the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then generating the corresponding path information item also includes: An indication to extract the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the currently encoded path information item.

14. The method of claim 11, wherein, If the corner voxel specified by the corresponding path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then generating the corresponding path information item also includes: The extracted path information item is an indication of whether it includes an indication of the path length or an indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the currently encoded path information item; and An indication of the path length or the difference between the previous path length and the current path length.

15. The method according to any one of claims 11 to 14, wherein the one or more second locations relate to locations obtained by traversing a voxel grid according to a predetermined pattern for each first location to define a sequence of second locations, and for a corresponding first location, an encoded path information item is decoded sequentially according to the sequence of second locations.

16. The method of claim 15, wherein the predetermined pattern traverses the voxel grid along rows and columns in a raster scan manner.

17. The method according to claim 15 or 16 when dependent on claim 13, The difference is encoded by 2 bits; or The difference is one of four predetermined values.

18. The method according to any one of claims 11 to 17, wherein each encoded path information item comprises: Instructions on whether to use the first mode or the second mode; Instructions for the first location; Instructions for the second location; as well as Indication regarding the existence of acoustic paths at the first and second locations; If the first mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item; If the corner voxel specified by the path information item corresponding to this encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then the path length is indicated; and Otherwise, the corner voxel indication and the path length indication; and If the second mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item; If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previously encoded path information item, then the indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the path information item corresponding to the previously encoded path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item; and Otherwise, the corner voxel indication and the path length indication.

19. The method according to any one of claims 11 to 17, wherein each encoded path information item comprises: Instructions on whether to use the first mode or the second mode; Instructions for the first location; Instructions for the second location; as well as Indication regarding the existence of acoustic paths at the first and second locations; If the first mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item; If the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item, then the encoded path information item includes either an indication of the path length or an indication of the difference between the previous path length and the current path length, together with the indication of the path length or the indication of the difference between the previous path length and the current path length, wherein the previous path length is the path length specified by the path information item corresponding to the previous encoded path information item, and the current path length is the path length specified by the path information item corresponding to the encoded path information item; and Otherwise, the corner voxel indication and the path length indication; and If the second mode is used, then each encoded path information item also includes: An indication of whether the corner voxel specified by the path information item corresponding to the encoded path information item is the same as the corner voxel specified by the path information item corresponding to the previous encoded path information item; If the corner voxel specified by the path information entry corresponding to the encoded path information entry is the same as the corner voxel specified by the path information entry corresponding to the previously encoded path information entry, then the indication of the difference between the previous path length and the current path length; and Otherwise, the corner voxel indication and the path length indication.

20. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to perform the method according to any one of claims 1 to 19.

21. A program comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 19.

22. A computer-readable storage medium storing the program as described in claim 21.

Citation Information

Patent Citations

  • Indirect illumination method and 3D graphics processing device

    US10275937B2

  • Control of audio effects using volumetric data

    US10293259B2

  • Modeling acoustic effects of scenes with dynamic portals

    US11606662B2

  • Parallel volume rendering system with a resampling module for parallel and perspective projections

    US6313841B1