Method, system, and apparatus for acoustic 3D spread modeling for voxel-based geometric representation
The method addresses the challenge of rendering 3D acoustic effects in voxel-based environments by assigning audio sources to positions within the scene using line segments, reducing computational load and improving audio realism.
Patent Information
- Application Number
- JP2024573273
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-25
- Filing Date
- 2023-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Existing audio rendering technologies in VR, AR, MR, and XR environments struggle with efficiently representing and rendering the acoustic effects of 3D extents using voxel-based geometry, leading to high computational loads and a need for simplified processes.
A method and apparatus for rendering audio in an audio scene by receiving a voxel-based audio scene representation, determining line segments through intersections, and assigning audio sources to positions within the scene based on these segments, utilizing occlusion and diffraction modeling to enhance audio rendering.
This approach reduces computational load and enhances the realism of audio rendering by accurately positioning audio sources within voxel-based 3D environments, allowing for more natural sound perception without explicit signaling of source coordinates.
Smart Images

Figure 2025522412000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 352,360, filed Jun. 15, 2022, and U.S. Provisional Application No. 63 / 441,120, filed Jan. 25, 2023, the entire disclosures of which are hereby incorporated by reference. Technical Field Broadly speaking, the present disclosure relates to a method of rendering audio within an audio scene, particularly based on a voxel - based audio scene representation of the audio scene. The present disclosure further relates to respective apparatuses and computer program products.
[0002] Some embodiments will be described herein with particular reference to the disclosure, but it will be understood that the present disclosure is not limited to such fields of use and is applicable in a broader context.
Background Art
[0003] Any discussion of background art throughout the present disclosure should in no way be construed as an admission that such technology is widely known or forms part of the common general knowledge in the art.
[0004] The Moving Picture Experts Group (MPEG) is an alliance of working groups jointly established by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC), and sets standards for media coding, including audio coding. MPEG is organized under ISO / IEC SC29, and the Audio Group is currently identified as Working Group (WG) 6. WG6 is currently working on the MPEG - I audio standard.
[0005] The new MPEG-I standard enables acoustic experiences from different viewpoints and / or perspectives or listening positions by supporting scenes and various movements around such scenes, such as movements using various degrees of freedom like 3 degrees of freedom (3DOF) or 6 degrees of freedom (6DoF) in virtual reality (VR), augmented reality (AR), mixed reality (MR) and / or extended reality (XR) applications. 6DoF interaction extends the 3DOF spherical video / audio experience limited to head rotation (pitch, yaw, roll) to include translational movements (forward / backward, up / down, left / right) in addition to head rotation, allowing navigation within the virtual environment (e.g., physically walking indoors).
[0006] For audio rendering in VR, AR, MR and XR applications, the object-based approach is widely used by representing complex auditory scenes as a plurality of distinct audio objects, each object being associated with parameters or metadata that define the position / location and trajectory of that object within the scene. Alternatively, audio rendering in such environments also uses higher-order ambisonics (HOA).
[0007] Audio objects are typically represented as point sources (without spread). As used herein, an audio source having an "extent" is an audio source waveform associated with a spatial region (the region being larger than a point). For example, a piano can be represented as an audio source having a cubic extent (e.g., stereo or monaural L / R) instead of a mere point source. The use of extent allows for an improvement in the user's audio experience, for example when the user is around a virtual piano object within a VR, AR, MR, or XR environment. In this example, the extent representing the piano for audio rendering need not have exact physical details like an actual piano.
[0008] To reflect the acoustic effects of an audio object having an extent, such an audio object may be represented by a voxel-based geometry. Voxels for audio rendering are important for media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments.
[0009] However, there still exists an existing need for improved rendering of the acoustic effects of 3D extents represented by voxel-based geometry, and in particular, it may be desirable to simplify the process and reduce the computational load.
Summary of the Invention
Means for Solving the Problems
[0010] In view of the above, the present disclosure provides a method, an apparatus, and a program for rendering audio in an audio scene, and a computer-readable storage medium, each having the features of the independent claims.
[0011] According to a first aspect of the present disclosure, a method for rendering audio in an audio scene is provided. The method can include receiving a voxel-based audio scene representation of the audio scene. The audio scene representation may include an indication of extent voxels representing a 3D extent, together with a plurality of audio source signals for audio sources related to the 3D extent. The method may further include obtaining (e.g., determining, calculating) the coordinates of intersections within the 3D extent. The method may further include determining one or more line segments extending along each coordinate direction of the audio scene representation through the intersections. The end points of each line segment may be determined based on the coordinates of one or more of the extent voxels. The method may also include allocating an audio source among the plurality of audio sources to an audio source position in the audio scene based on the one or more line segments.
[0012] In some embodiments, the intersection point may be one of the geometric center of the 3D extent and the center of gravity of the 3D extent.
[0013] In some embodiments, the endpoints of each line segment may be determined based on the extreme coordinate values of the 3D extent along their respective coordinate directions, and thus the length of the line segment corresponds to the maximum dimension of the projection of the 3D extent in their respective coordinate directions.
[0014] In some embodiments, the audio scene representation may further show occluded voxels. Assigning an audio source may include assigning the audio source to coordinates within voxels other than the occluded voxels.
[0015] In some embodiments, the audio scene representation may further show non-filled voxels (e.g., air voxels). Assigning an audio source may include assigning the audio source to coordinates on each line segment that are closest to the endpoints of the respective line segment and are within the spreading voxels or non-filled voxels.
[0016] In some embodiments, assigning an audio source may further include determining one or more possible target positions for assigning the audio source based on the line segment.
[0017] In some embodiments, the audio scene representation may further show non-filled voxels (e.g., air voxels). Determining the one or more possible target positions may include selecting coordinates for the one or more possible target positions that are closest to the endpoints of the respective line segment and are within the spreading voxels or non-filled voxels.
[0018] In some embodiments, determining the one or more possible target positions may include selecting the coordinates for the one or more possible target positions that are closest to the endpoints of each line segment and that are within the spread voxel.
[0019] In some embodiments, the method may further include selecting the audio source positions from the possible target positions based on a predefined minimum distance between the audio sources. Also, the method may include assigning an audio source among the plurality of audio sources to the selected audio source position.
[0020] In some embodiments, the method may further include obtaining a mapping indicating the assignment of audio source signals to the audio source positions.
[0021] In some embodiments, the method may further include assigning a gain to the audio source positions based at least in part on the mapping.
[0022] In some embodiments, the method may further include obtaining the coordinates of the listener position. Also, the method may include rendering the audio source signals of the assigned audio sources based on a reference distance between the listener position and the 3D spread.
[0023] In some embodiments, rendering may further include rendering the audio source signals based on occlusion and diffraction modeling.
[0024] According to a second aspect of the present disclosure, an apparatus for rendering audio in a voxel-based audio scene representation is provided. The apparatus may include one or more processors configured to execute a method that may include receiving a voxel-based audio scene representation of an audio scene, the audio scene representation including an indication of extent voxels representing a 3D extent, together with a plurality of audio source signals for audio sources associated with the 3D extent. The method may further include obtaining coordinates of intersections within the 3D extent. The method may further include determining one or more line segments extending along respective coordinate directions of the audio scene representation through the intersections. Endpoints of each line segment may be determined based on coordinates of one or more of the extent voxels. The method may also include, based on the one or more line segments, assigning audio sources of the plurality of audio sources to audio source positions within the audio scene.
[0025] Aspects of the present disclosure may be implemented via an apparatus. The apparatus may include a processor and a memory coupled to the processor. The processor may be adapted to implement the methods according to aspects and embodiments of the present disclosure.
[0026] Aspects of the present disclosure may be implemented via a program. When instructions of the program are executed by a processor, the processor may implement aspects and embodiments of the present disclosure. A computer-readable storage medium may store the program. Such a computer-readable storage medium may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Thus, some innovative aspects of the subject matter described in the present disclosure may be implemented via one or more computer-readable storage media storing software.
[0027] It will be understood that the features of the apparatus and the steps of the method can be interchanged in many ways. In particular, the details of the disclosed method can be implemented by the corresponding apparatus (or system), and vice versa. This will be understood by those skilled in the art. Furthermore, any of the above descriptions made with respect to the method is understood to be equally applicable to the corresponding apparatus (or system), and vice versa.
Brief Description of the Drawings
[0028] Here, with reference to the accompanying drawings, exemplary embodiments of the present disclosure will be described merely by way of example.
[0029]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0030] In the drawings, when connection elements such as solid lines, dashed lines, and arrows are used to illustrate a connection, relationship, or association between two or more other schematic elements, the absence of such a connection element is not intended to imply that the connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the present disclosure. Further, for ease of explanation, a single connection element is used to represent multiple connections, relationships, or associations between elements. For example, when a connection element represents the communication of a signal, data, or instruction, it should be understood by those skilled in the art that such an element may represent one or more signal paths that affect the communication, as necessary.
Mode for Carrying Out the Invention
[0031] An audio source having extent is an audio source waveform associated with a spatial region (larger than a point). The spatial region can be modeled by a geometry (2D or 3D). A voxel is a 3D volume representation and thus can model such a geometry. The use of voxels for audio rendering is relevant to diverse media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments. A voxel is a spatial volume having assigned acoustic properties or audio rendering instructions. The voxel size is an encoder configuration parameter and can be selected (manually or automatically) according to the level of detailed scene geometry (e.g., in the range of 10 cm to 1 m). Voxels for audio rendering can be obtained in the following ways. · Voxelization (or conversion) of a mesh - based scene representation · From a scene representation used for scene generation (or even video rendering), for example, by downsampling to voxels of a smaller size.
[0032] The methods and apparatus described herein relate to how to render the acoustic effects of a 3D extent when the 3D extent is represented by voxel - based geometry. More specifically, the methods and apparatus described herein relate to how to obtain the coordinates of "joint" (point) audio sources.
[0033] Typically, in order to approximate a spreading spatial region, multiple audio sources are required to model an audio source having a spread. These (target) audio sources can be derived from a given audio source(s) associated with the spread, specified by a scene creator, for example using a scene description. The term "congruent" as used here can be said to imply that these target sources are related to each other since they represent the spatial region of the spread in one dimension. Since there are three dimensions, at least one pair of audio sources is required for each dimension. As an example, the scene description specifies stereo channels with a cubic spread to represent a virtual piano object. Then, processing may be performed in a renderer to derive three pairs of "congruent" "target" audio sources placed at six different positions within the vicinity of the spread.
[0034] That is, the methods and apparatus described herein find a respective number of (point) audio sources, e.g., N = [1,…,6] and their coordinates (positions) P 1,…,N and aim to map the audio signals S 1,…,M to the respective positions P 1,…,N and gains, which is based on a given scene description including, for example, listener position coordinates L, 3D spread material ID (representing an audio object 3D spread geometry approximation), 3D grid index VOX (representing a set of 3D spreads), and a set of M audio signals (mono, stereo, etc.), as well as, for example, the minimum distance Δ min between two "congruent" (point) audio sources, a mapping matrix F for assigning audio signals to the resulting point source positions (and gains), and a reference distance.
[0035] The methods and apparatus described herein allow for modeling an audio object having an extent represented by a voxel - based geometry without explicitly signaling audio source coordinates (e.g., without explicitly transmitting and receiving this information within a bitstream). That is, it can be said that the methods and apparatus described herein emphasize the manner in which "congruent" audio source coordinates (positions) are determined within the vicinity of the extent, assuming that the extent is represented by a voxel - based geometry. The resulting positions / coordinates are voxel coordinates. Since these are calculated on the renderer side, there is no need to know them in advance and no explicit signaling / transmission is required.
[0036] Advantageously, this allows for automatically obtaining signal audio source coordinates for a complex voxel - based 3D extent geometry on the decoder side, especially when the decoder operates in a manner compliant with an audio standard such as the MPEG - set standard. Another advantage is that this allows for support of 3D extent geometry modification at the decoder (without the need to re - encode the modified scene).
[0037] Encoding of the 3D extent geometry is performed at the encoder and transmitted to the decoder to convey information about the extent geometry to the decoder / renderer. Similar to many other objects in the scene, the extent can be modified either on the encoder side or on the decoder / renderer side. Modification on the encoder side requires a "re - encoding" of the extent to be transmitted to the decoder. This does not apply to modifications on the decoder / renderer side. The methods described herein are implemented on the decoder / renderer side, i.e., any modification to the extent is made on the decoder / renderer side, so no "re - encoding" is required.
[0038] How to represent a voxel-based audio scene Any voxel - based representation of an audio scene may include an indication of voxels that are not transparent voxels (e.g., occluding voxels), i.e., voxels through which sound cannot propagate or cannot propagate freely, i.e., an indication of the occlusion geometry. This indication may be related to an indication of the coordinates of each voxel (e.g., center coordinates, corner coordinates, etc.). These voxel coordinates may be represented, for example, by a grid index. Further, the voxel - based representation may include an indication of the material properties of voxels that are not transparent voxels, such as absorption coefficients, reflection coefficients, etc. In addition to occluding voxels, the voxel - based representation may also indicate transparent voxels or non - filled voxels (e.g., air voxels), i.e., voxels through which sound can propagate, i.e., a representation of the sound - propagation medium. Thus, some implementations of the voxel - based representation of an audio scene may include an indication of the respective material properties for each voxel within a pre - defined section of space (e.g., within the boundaries enclosing the audio scene).
[0039] Method for rendering audio in an audio scene Referring to FIG. 1, an example of a method for rendering audio within an audio scene is shown. This method is executed on the decoder / renderer side and may be implemented by each decoder / renderer. For example, all method steps may be executed in real - time in a single device which may be a VR / AR / MR / XR device.
[0040] Step S101 Then, a voxel - based audio scene representation of the audio scene is received. The audio scene representation includes an indication of spread voxels representing a 3D spread, along with a plurality of audio source signals for audio sources related to the 3D spread. In other words, it can be said that the 3D spread corresponds to an audio object having a spread with a geometric shape represented by the spread voxels.
[0041] An example of a voxel-based audio scene representation of an audio scene is schematically shown in FIG. 2. The example of FIG. 2 is a 2D cut through a voxel-based 3D audio scene representation including a 3D extent. FIG. 2 shows a grid pattern representing the voxelization of the audio scene representation. In the example of FIG. 2, according to one embodiment, extent voxels 205 and non-filled voxels (e.g., air voxels) 206 are shown. That is, in addition to the extent voxels representing the 3D extent, the audio scene representation may also show voxels representing a part of the acoustic environment of the 3D extent. The non-filled voxels can be said to represent the sound transmission medium. The sound transmission medium may be, for example, air and / or water.
[0042] Referring again to the example of FIG. 1, Step S102 wherein the coordinates of the intersection within the 3D extent are obtained (e.g., determined, calculated). In one embodiment, the intersection may be one of the geometric center and the center of gravity of the 3D extent. In the example of FIG. 2, the geometric center 201 and the center of mass (center of gravity) 202 of the 3D extent, which can be alternatively used, are schematically shown.
[0043] Without intending to be limiting, the intersection can be the origin O of the Cartesian coordinate system. In the context of an example of a Cartesian coordinate system, the intersection (center of the 3D extent) C of the voxel-based 3D extent representation VOX x,y,z can be determined using the "min / max" approach as follows. x,y,z
Equation
Number
[0044] Referring again to the example of FIG. 1, Step S103Then, one or more line segments are determined, each passing through an intersection point and extending along each coordinate direction of the audio scene representation (e.g., along the x, y, and z coordinate axes). The endpoints of each line segment are determined based on the coordinates of one or more of the said spreading voxels. For example, as detailed below, the endpoints of each line segment may be determined based on the extreme coordinate values of the 3D spread along each coordinate direction. That is, for example, for a line segment extending along the x-axis, the endpoints can be determined based on the extreme coordinates of the 3D spread along the x-axis.
[0045] In the example of FIG. 2, in a 2D cut, two such line segments 203, 204 are shown, each passing through the geometric center 201 of the 3D spread and having respective endpoints 203a, 203b, 204a, 204b. In the case of a Cartesian coordinate system, the lines are the X, Y, Z axes (with the intersection point being the origin of the coordinate system), and the line segments may be segments of the X, Y, Z axes. In the 2D cut of FIG. 2, line 203 may be the Y-axis and line 204 may be the X-axis.
[0046] Referring again to the example of FIG. 1, Step S104 then, an audio source among a plurality of audio sources is assigned to an audio source position within the audio scene based on one or more line segments. As used herein, "assigned" can be said to mean that the target audio source is generated (e.g., based on a given / specified audio source of a certain spread) and linked / mapped to the calculated coordinate position. That is, in step S104, a set of target audio sources placed at the calculated positions in the vicinity of the spread can be output. These target sources (instead of a given / specified audio source with a spread) can be used to replace the task of rendering an "audio source with a spread" by rendering a set of point sources. The one or more line segments are constructed to assist in determining the target audio source positions. Note that in S103, these line segments are output.
[0047] Referring now to FIGS. 3 and 4, two examples of assigning audio sources to audio source positions within an audio scene are schematically shown. That is, FIGS. 3 and 4 represent two possible implementations of method step S104. These implementations differ in the way the target source positions indicated by 308a, 308b, 309a, and 309b are determined. In particular, FIGS. 3 and 4 also represent respective 2D cuts.
[0048] FIGS. 3 and 4 show examples of indications of the spread voxels 305, non-filled voxels 306, and occluded voxels 307. The occluded voxels may represent, for example, acoustic occlusions existing between a 3D spread and the listener.
[0049] In some embodiments, the endpoints of each line segment may be determined in step S103 based on the extreme coordinate values of the 3D spread along their respective coordinate directions, and thus the length of the line segment may correspond to the maximum dimension of the projection of the 3D spread in their respective coordinate directions.
[0050] As described above, the intersection point inside the 3D spread may be used as the origin O of the rectangular coordinate system, and each line segment may represent a line segment of the X, Y, and Z axes. Thus, in an example of the rectangular coordinate system, the maximum dimensions (characteristic dimension extreme points) D of the projection of the 3D spread in their respective coordinate directions max x,y,z and D min x,y,z can be determined (extracted) as follows.
Equation
[0051] The examples of FIGS. 3 and 4 show respective maximum dimensions 303a, 303b and 304a, 304b.
[0052] In some embodiments, such as those shown in FIG. 3, in step S104, allocating an audio source (to the audio source position within the audio scene) may include allocating the audio source to coordinates within a voxel other than the occlusion voxel 307 (e.g., 308a). The occlusion voxels 307 can each affect the sound perceived by the listener at their respective listener positions, and by allocating the audio source to coordinates within a voxel other than the occlusion voxel 307, the respective audio source signals of the allocated audio sources can be rendered such that the sound perceived by the listener seems realistic.
[0053] Further, in some embodiments, such as those shown in FIG. 4, allocating an audio source (to the audio source position within the audio scene) may further include allocating the audio source to coordinates on each line segment that is closest to the endpoints of the respective line segment and that is within a spread voxel or an unfilled voxel (e.g., 308a, 308b, 309a, 309b in FIG. 4).
[0054] Instead of or in addition to the above embodiments, allocating an audio source may further include determining one or more possible target positions for allocating the audio source based on the line segments. Determining the one or more possible target positions may include selecting coordinates for the one or more possible target positions that are closest to the endpoints of the respective line segments and that are within a spread voxel or an unfilled voxel. In a further embodiment, determining the one or more possible target positions may include selecting coordinates for the one or more possible target positions that are closest to the endpoints of the respective line segments and that are within a spread voxel.
[0055] Depending on the usage scenario, by assigning the audio source to the coordinates within the spread voxel or by selecting the coordinates for each possible target position that is closest to the endpoints of each line segment and within the spread voxel, it becomes possible to render each audio source signal such that the sound perceived by the listener seems more natural compared to the coordinates within the non-filled voxels.
[0056] In an example of a Cartesian coordinate system, the possible target positions (isotropic point source coordinates in 3D spread) P max x,y,z , 308a, 309a and P min x,y,z , 308b, 309b may be selected as follows. P max x = [p max x , C y , C z is the voxel on the line [D max x , D min x that is not the "hidden" voxel for this audio object and is closest to D max x on it. P min x = [p min x , C y , C z is the voxel on the line (D max x , D min x that is not the "hidden" voxel for this audio object and is closest to D min x on it.
[0057] Applying the same procedure, P max y,z and P min y,z can be obtained. The "non-blocking" voxels are
Number
[0058] In some embodiments, the method may further include selecting an audio source position from possible target positions based on a predefined minimum distance between audio sources. The method may include assigning an audio source among a plurality of audio sources to the selected audio source position. For example, the number N = [1,…,6] of (point) audio sources can be calculated by considering three variables.
Number
[0059] In an embodiment, the method may further include obtaining a mapping indicating the (e.g., intended or desired) assignment of audio source signals to audio source positions (or possible target positions). For example, the following mapping of the position coordinates P 1,…,M of M audio signals S 1,…,6 may be read from the bitstream payload.
Table 1
[0060] In some embodiments, the method may further include assigning a gain to an audio source location. This assignment may be at least partially based on the mapping described above. An appropriate signal gain may further be assigned based on the number of selected audio sources in order to ensure energy conservation.
[0061] Referring now to the examples of FIGS. 5 - 9, a use case of an example of a method for rendering audio within an audio scene as described herein is shown. In this exemplary use case, the 3D extent to be rendered / modeled is, by way of example, based on the tram 500. FIGS. 7 - 9 show the respective "visible" line segments 501, 502, 503, and the respective assigned coordinates / target position coordinates 504a, 504b, 505a, 505b, 506a, 506b, which have been determined according to the method described herein.
[0062] In this exemplary use case, the renderer is tasked with appropriately rendering the sound of a virtual tram within a VR / AR / XR / MR scene. It may be seen that the tram is running on a busy street. Since the audio emitted from the tram comes from several parts distributed along the length of the tram, the tram cannot be modeled as a single point source. First, a "tram" object (Figure 5) within the VR / AR / XR / MR scene and an accompanying "spatial audio source(s)" to represent the sound of the "tram" may be specified by the scene creator as part of the "scene description" of the VR / AR / XR / MR scene. A specified "spatial" model representing the tram for audio rendering is shown in Figure 6 as an example. Figures 7, 8, and 9 show possible embodiments when applying the method shown in Figure 1 to determine a set of target audio source positions 504a, 504b, 505a, 505b, 506a, 506b in the vicinity of the space. Then, actual target audio sources corresponding to those positions are generated. In this context, Figures 8 and 9 are taken from Figure 7 by cutting the tram space representation to show the target audio source positions. Thus, the sound of the tram within the VR / AR / XR / MR scene results from the rendering of those target audio sources.
[0063] Referring now to the example of Figure 10, in one embodiment, the method may further include obtaining the coordinates of the listener position 510 and rendering the audio source signals of the assigned audio sources based on a reference distance 511 between the listener position 510 and the 3D space 500. For example, the distance from the listener to the object (L,P) may be subtracted by the distance from the listener L to the closest point R of the space VOX. Here, |L - R| == min(|L - VOX|).
[0064] In some embodiments, rendering may further include rendering (point) audio source signals based on (voxel-based) occlusion and diffraction modeling. That is, for example, a selected subset {P max x,y,z , P min x,y,z} of point audio sources may be rendered by applying voxel-based occlusion and diffraction modeling. The example of FIG. 10 shows a 3D extent representation of a streetcar 500 hidden by an obstacle 512. Coordinates 504a, 504b, 506a, 518 are thus occluded, and as a result, a set 516, 516a-d of virtual coordinates is generated by diffraction modeling.
[0065] The 3D extent modeling method described herein assumes the application of diffraction modeling. However, this method can also be used without applying diffraction modeling. In this case, the method is applied to the 3D extent subset visible to the listener 510. To obtain a subset that is "visible" to the listener (visible implies that there is no acoustic shield between the listener and the corresponding point), the following methods can be used: a method based on ray tracing, or a method by checking for occlusion along the line between the listener and the subset of the 3D extent representation. This subset can be determined by Monte Carlo or any other subsampling method.
[0066] Exemplary algorithm In other words, a method for rendering audio in an audio scene can be described as follows. The following represents an exemplary implementation of the method shown in FIG. 1. It is assumed that the decoder has already received "scene description" information including the aspects shown in step S101.
[0067] Given Scene description : The position coordinates L of the listener 3D extent material ID (representing the 3D extent geometry approximation of the audio object) Set of 3D grid indices VOX (representing a 3D extent set) Set of M audio signals (mono, stereo, etc.) Modeling settings : Minimum distance Δ between two "identical" point audio sources min Mapping matrix F for assigning audio signals to the obtained point source positions (and gains) Reference distance To find Number of point sound sources N = [1, …, 6] and their coordinates P 1,…,N Audio signal S 1,…,M At position P 1,…,N And map to the gain
[0068] Solution For example, use the "minimum / maximum (geometric center)" approach to determine the 3D extent center representation coordinates C of the voxel-based 3D extent representation VOX x,y,z 3D extent center representation coordinates C of x,y,z (Step S102):
Number
[0069] A center representation is required to extract three characteristic dimensions from the three-dimensional 3D extent representation.
[0070] Next extreme points D max x,y,z And D min x,y,z Determine the characteristic dimension representation of the 3D extent by (step S103):
Number
[0071] 3D extent identical point source coordinates Pmax x,y,z and P min x,y,z are determined (step S104). Here, P max x = [p max x , C y , C z is not a "hidden" voxel for this audio object, and is the voxel on line [D max x , D min x ) closest to D max x in it. Here, p min x = [p min x , C y , C z is not a "hidden" voxel for this audio object, and is the voxel on line (D max x , D min x closest to D min x in it.
[0072] Applying the same procedure, P max y,z and P min y,z can be obtained. A "non-shielding" voxel can be defined as any one of
Number
[0073] Considering the three variables
Number
[0074] All W x,y,z ≤ Δ min If so, the 3D spread is not modeled and the audio object is represented by a single audio point source located at voxel C x,y,z
[0075] Based on the bit - stream payload, obtain a mapping of the M audio signals S 1,…,M to the position coordinates P 1,…,6 of, for example, as in the exemplary mapping given in Table 1 above. In some implementations, the mapping M may be read or extracted from the bit - stream.
[0076] Assign appropriate signal gains based on the number of selected audio sources (to ensure energy conservation).
[0077] Apply voxel - based occlusion and diffraction modeling to render a selected subset {P max x,y,z , P min x,y,z} of the point audio sources (see Figure 10 as an example).
[0078] Apply reference - distance processing. That is, subtract the distance from the listener L to the closest point R of the spread VOX from the distance from the listener L to the object (L, P). Here, |L - R| == min(|L - VOX|)
[0079] Apply voxel - based occlusion and diffraction modeling to render a selected subset {P max x,y,z , P min x,y,z Render {}.
[0080] The rendering may be performed by a renderer that can simulate acoustic hiding and diffraction modeling.
[0081] FIG. 10 shows, as an example of how diffraction processing can be applied to the various methods described herein, an object blocked by an obstacle, i.e., a 3D extent representation of a streetcar. The central points are hidden, and as a result, a set of virtual points is generated by diffraction modeling. 513 and 514 are lines of sight not blocked by the hiding object 512. 515 is the direction (azimuth) in which the object 518 is perceived. These objects are then perceived (modeled) as 516. As a result, the coordinates 504a, 504b, 506a, 506b (missing in the figure) belonging to 518 are modeled by 516a, 516d, 516b, 516c belonging to 516.
[0082] The 3D extent modeling method described here assumes the application of diffraction modeling.
[0083] However, this method can also be used without applying diffraction modeling. In this case, the method is applied to a 3D extent subset that is visible to the listener. To obtain a subset that is "visible" to the listener (visible implies that there is no acoustic shielding between the listener and the corresponding point), the following methods can be used: a method based on ray tracing, or a method by checking for shielding on the line between the listener and the subset of the 3D extent representation. This subset can be determined by Monte Carlo or any other sub-sampling method.
[0084] Referring to the example of FIG. 11, an apparatus 1100 including one or more processors 1101, 1102 according to an embodiment of the present disclosure is shown. The one or more processors 1101, 1102 may be configured to execute the methods described herein.
[0085] A computing device implementing the above-described techniques can have the following exemplary architecture. Other architectures are possible, including architectures having more or fewer components. In some implementations, the exemplary architecture includes one or more processors (e.g., a dual-core Intel® processor), one or more output devices (e.g., an LCD), one or more network interfaces, one or more input devices (e.g., a mouse, keyboard, touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can communicate and exchange data through one or more communication channels (e.g., a bus), and various hardware and software can be utilized to facilitate the transfer of data and control signals between components.
[0086] The term "computer-readable medium" refers to a medium that participates in providing instructions to a processor for execution and includes, but is not limited to, non-volatile media (e.g., optical disk or magnetic disk), volatile media (e.g., memory), and transmission media. Transmission media includes, but is not limited to, coaxial cable, copper wire, and fiber optic.
[0087] The computer-readable medium can further include an operating system (e.g., Linux (registered trademark) operating system), a network communication module, an audio interface manager, an audio processing manager, and a live content distributor. The operating system can be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system performs basic tasks including, but not limited to, recognizing input from a network interface and / or device and providing output thereto, tracking and managing files and directories on a computer-readable medium (e.g., memory or storage device), controlling peripheral devices, and managing traffic on one or more communication channels. The network communication module includes various components for establishing and maintaining a network connection (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.).
[0088] The architecture can be implemented in a parallel processing or peer-to-peer infrastructure, or in a single device with one or more processors. The software can include multiple software components or can be a single code body.
[0089] The described features can be advantageously implemented in one or more computer programs executable on a programmable system that includes a programmable processor coupled to receive data and instructions from, and to send data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, within a computer to perform an activity or cause a result. A computer program can be written in any form of programming language (including, for example, Objective-C, Java) that includes compiled languages or interpreter-type languages, and can be deployed in any form, which includes being deployed as a stand-alone program, or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.
[0090] Suitable processors for executing the command program include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one of a single processor or multiple processors or cores of any type of computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer includes or is operatively coupled to communicate with one or more mass storage devices for storing data files. Such devices include magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Storage devices suitable for embodying computer program instructions and data physically include any form of non-volatile memory, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by or incorporated into an ASIC (application specific integrated circuit).
[0091] To provide interaction with a user, these features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or retinal display device for displaying information to the user. The computer can have a touch surface input device (e.g., a touch screen) or a keyboard, and a pointing device such as a mouse or trackball by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user.
[0092] These features can be implemented in a computer system including back-end components such as a data server, a computer system including middleware components such as an application server or an Internet server, or a computer system including front-end components such as a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, for example, LANs, WANs, and the computers and networks that form the Internet.
[0093] A computing system can include clients and servers. Clients and servers are generally separated from each other and typically interact through a communication network. The relationship between a client and a server results from computer programs that are executed on respective computers and have a client-server relationship with each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (for the purpose of, e.g., displaying the data to a user who interacts with the client device and receiving user input from the user). Data generated at the client device (e.g., as a result of user interaction) can be received at the server from the client device.
[0094] One or more computer systems can be configured to perform specific actions by installing in the system software, firmware, hardware, or any combination thereof that causes the system to perform actions during operation. One or more computer programs can be configured to perform specific actions by including instructions that, when executed by a data processing device, cause the device to perform actions.
[0095] This specification includes many details of individual implementations, which should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. In this specification, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately, or in any suitable sub-combination, in multiple embodiments. Furthermore, features may be described as acting in certain combinations and may even initially be claimed as such, but one or more features from the claimed combination may, in some cases, be deleted from such combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0096] Similarly, operations are depicted in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, in order to achieve desirable results, or that all of the operations shown be performed. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0097] Unless otherwise specified, as will be apparent from the following discussion, throughout this disclosure, discussions using terms such as "processing", "computing", "calculating", "determining", "analyzing", etc. refer to actions and / or processes of a computer or computing system, or similar electronic computing device, which manipulate and / or transform data represented as physical quantities, such as electronic quantities, into other data also represented as physical quantities.
[0098] References throughout this disclosure to "one exemplary embodiment", "some exemplary embodiments", or "an exemplary embodiment" mean that a particular feature, structure, or characteristic described in connection with the exemplary embodiment is included in at least one exemplary embodiment of the disclosure. Thus, the appearances of the phrases "in one exemplary embodiment", "in some exemplary embodiments", or "in an exemplary embodiment" in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiment. Further, the particular features, structures, or characteristics may be combined in any suitable manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from this disclosure.
[0099] As used herein, unless otherwise specified, the use of ordinal adjectives such as "first", "second", "third", etc. to describe a common object merely indicates that different instances of similar objects are being referred to, and is not intended to mean that the objects so described must be in a given sequence in any way, whether temporal, spatial, ranking, or otherwise.
[0100] Also, it should be understood that the terminology and expressions used herein are for the purpose of description and should not be regarded as limiting. The use of "including", "comprising", or "having" and their variations is meant to include the recited items and their equivalents, as well as additional items. Unless specifically stated otherwise or limited, the terms "attached", "connected", "supported", and "coupled", and their variations, are used in a broad sense and include both direct and indirect attachment, connection, support, and coupling.
[0101] In the following claims and the description of this specification, any of the terms "comprising", "consisting of", or "having" are open terms meaning including at least the recited element / feature but not excluding others. Thus, when used in the claims, the term "having" / "including" should not be construed as a limitation to the recited means, element, or step. For example, the expression "an apparatus having A and B" should not be limited to an apparatus consisting only of elements A and B. As used herein, any of the terms "including", "containing", or "comprising" are also open terms meaning including at least the element / feature recited by that term but not excluding others. Thus, "including" is synonymous with "having" and means having.
[0102] In the above description of exemplary embodiments of the present disclosure, it should be understood that various features of the present disclosure may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of facilitating the flow of the disclosure and aiding in the understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed exemplary embodiment. Thus, the claims following the specification are hereby expressly incorporated into this specification, with each claim standing on its own as a separate exemplary embodiment of the present disclosure.
[0103] Furthermore, some exemplary embodiments described herein include some features included in other exemplary embodiments but not other features, but combinations of features of different exemplary embodiments are intended to be within the scope of the present disclosure and form different exemplary embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any of the claimed exemplary embodiments can be used in any combination.
[0104] In the description provided herein, numerous specific details are set forth. However, it is understood that the exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure the understanding of this document.
[0105] Accordingly, while what is considered to be the best aspects of the present disclosure is described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of the present disclosure. It is intended to claim all such changes and modifications that fall within the scope of the present disclosure. For example, the formulas given above are merely representative of the procedures that can be used. Functions may be added to or removed from the block diagrams, and operations may be exchanged between functional blocks. Steps can be added or removed from the methods described within the scope of the present disclosure.
[0106] Various aspects and implementations of the present disclosure can also be understood from the following enumerated example embodiments (EEE), which are not the claims.
[0107] 〔EEE1〕 A method for modeling an extended audio object for audio rendering in a virtual or augmented reality environment, the method comprising: determining a 3D spread center representation of a voxel-based 3D spread representation; determining a 3D spread characteristic dimension representation based on the 3D spread representation; determining 3D spread congruent point source coordinates based on the 3D spread characteristic dimension representation or the 3D spread center representation, a method. 〔EEE2〕 receiving a mapping of M audio signals to position coordinates; further comprising assigning signal gains of the M audio signals to the point source, the method according to EEE1. 〔EEE3〕 the method according to EEE1, further comprising rendering the point source based on voxel-based occlusion and diffraction modeling. 〔EEE4〕 the method according to EEE1, wherein the center representation can be determined based on the geometric center of the voxel or by a centroid approach. 〔EEE5〕 The method according to EEE1, wherein the dimensional representation can be based on endpoints or by calculating corresponding offsets from a center. 〔EEE6〕 The method according to any one of EEE1 to 5, wherein the method is applied to a subset of the voxel-based 3D extent representation. 〔EEE7〕 The method according to EEE6, wherein the subset of the voxel-based 3D extent representation corresponds to voxels that are not acoustically shielded (visible). 〔EEE8〕 A non-transitory computer program comprising instructions that, when executed by a processor, cause the processor to execute the method according to any one of EEE1 to 7. 〔EEE9〕 An apparatus configured to implement the method according to any one of EEE1 to 7.
Claims
1. A method for rendering audio in an audio scene, the method comprising: receiving (S101) a voxel-based audio scene representation of the audio scene, the audio scene representation including indications of spread voxels (205; 305) representing a 3D extent, together with a plurality of audio source signals for audio sources associated with the 3D extent; obtaining (S102) coordinates of intersections (201, 202; 301) within the 3D extent; determining (S103) one or more line segments (203, 204; 303, 304) extending along respective coordinate directions of the audio scene representation through the intersections (201, 202; 301), wherein endpoints (203a, 203b, 204a, 204b; 303a, 303b, 304a, 304b) of each line segment (203, 204; 303, 304) are determined based on coordinates of one or more spread voxels (205; 305); assigning (S104) audio sources of the plurality of audio sources to audio source positions (308a, 308b, 309a, 309b) in the audio scene based on the one or more line segments (203, 204; 303, 304); A method as described above.
2. The method according to claim 1, wherein the intersection is one of a geometric center and a center of gravity of the 3D extent.
3. The method according to claim 1, wherein endpoints of each line segment are determined based on extreme coordinate values of the 3D extent along respective coordinate directions, and a length of the line segment corresponds to a maximum dimension of a projection of the 3D extent in each respective coordinate direction.
4. The audio scene representation further indicates occluded voxels, and assigning the audio sources includes assigning the audio sources to coordinates within voxels other than the occluded voxels. The method according to claim 1.
5. The audio scene representation further indicates non-filled voxels, and assigning the audio sources includes assigning the audio sources to coordinates on each line segment that are closest to endpoints of the respective line segment and are within a spread voxel or a non-filled voxel. The method according to claim 4.
6. The method according to claim 1, wherein allocating the audio source further includes determining one or more possible target positions for allocating the audio source based on the line segment.
7. The audio scene representation further shows non-filled voxels, Determining the one or more possible target positions includes selecting coordinates for the one or more possible target positions that are closest to the endpoints of each line segment and are within the spreading voxels or non-filled voxels. The method according to claim 6.
8. The method according to claim 6, wherein determining the one or more possible target positions includes selecting coordinates for the one or more possible target positions that are closest to the endpoints of each line segment and are within the spreading voxels.
9. The method further includes: selecting the audio source position from the possible target positions based on a predefined minimum distance between audio sources; and allocating the audio source from among the plurality of audio sources to the selected audio source position. The method according to claim 6.
10. The method according to claim 1, further including obtaining a mapping indicating the allocation of the audio source signal to the audio source position.
11. The method according to claim 10, further including allocating a gain to the audio source position based at least in part on the mapping.
12. The method further includes: obtaining coordinates of the listener position; and rendering the sound source signal of the allocated audio source based on a reference distance between the listener position and the 3D spread. The method according to claim 1.
13. The method according to claim 12, wherein rendering further includes rendering the audio source signal based on occlusion and diffraction modeling.
14. An apparatus (1100) for rendering audio in a voxel-based audio scene, the apparatus including one or more processors (1101, 1102) configured to execute a method, the method comprising: Receiving a voxel-based audio scene representation of the audio scene (S101), wherein the audio scene representation includes an indication of spread voxels (205; 305) representing a 3D spread, together with a plurality of audio source signals for audio sources related to the 3D spread; Obtaining coordinates of intersections (201, 202; 301) inside the 3D spread (S102); Determining one or more line segments (203, 204; 303, 304) extending along respective coordinate directions of the audio scene representation and passing through the intersections (201, 202; 301), wherein endpoints (203a, 203b, 204a, 204b; 303a, 303b, 304a, 304b) of each line segment (203, 204; 303, 304) are determined based on coordinates of one or more spread voxels (205; 305); Assigning audio sources among the plurality of audio sources to audio source positions (308a, 308b, 309a, 309b) in the audio scene based on the one or more line segments (203, 204; 303, 304) (S104); An apparatus comprising the above.
15. A program including instructions that, when executed by a processor, cause the processor to execute the method according to any one of Claims 1 to 13.
16. A computer-readable storage medium storing the program according to Claim 15.
Citation Information
Patent Citations
Diffraction modelling based on grid pathfinding
WO2021198152A1