Method, system, and apparatus for acoustic 3D spread modeling for voxel-based geometric representation

The method simplifies the rendering of 3D sound effects in voxel-based geometry by assigning audio sources to positions based on line segments, reducing computational load and enhancing audio experience in VR, AR, and XR environments.

JP7855735B2Active Publication Date: 2026-05-08DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2023-06-13
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for rendering 3D extended sound effects using voxel-based geometry in VR, AR, MR, and XR applications face challenges in computational load and complexity, necessitating improved rendering techniques.

Method used

A method and apparatus for rendering audio in a voxel-based audio scene representation by determining line segments through intersections, assigning audio sources to positions based on these segments, and applying gain adjustments, while considering voxel types and listener position, to model audio objects with spread without explicit signaling of source coordinates.

Benefits of technology

This approach simplifies the rendering process, reduces computational load, and enhances the audio experience by accurately modeling 3D sound propagation, allowing for natural sound perception without the need for re-encoding modified scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007855735000010
    Figure 0007855735000010
  • Figure 0007855735000011
    Figure 0007855735000011
  • Figure 0007855735000012
    Figure 0007855735000012
Patent Text Reader

Abstract

The present application describes a method for rendering audio within an audio scene. The method includes receiving a voxel-based audio scene representation of the audio scene, where the audio scene representation includes indications of extent voxels representing a 3D extent, together with a plurality of audio source signals for audio sources associated with the 3D extent; obtaining coordinates of intersections within the 3D extent; determining one or more line segments extending along respective coordinate directions of the audio scene representation through the intersections, where endpoints of each line segment are determined based on coordinates of one or more extent voxels; and assigning, based on the one or more line segments, audio sources among the plurality of audio sources to audio source positions within the audio scene. Further, respective apparatuses and computer program products are described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Application No. 63 / 352,360, filed on 15 June 2022, and U.S. Provisional Application No. 63 / 441,120, filed on 25 January 2023, all of which are incorporated herein by reference in their entirety. Technical field This disclosure broadly relates to a method for rendering audio within an audio scene, particularly based on a voxel-based audio scene representation of the audio scene. This disclosure further relates to the respective apparatus and computer program products.

[0002] While this specification describes several embodiments with particular reference to the disclosure, it will be understood that this disclosure is not limited to such articulations but is applicable in a broader context. [Background technology]

[0003] Any discussion of background technology throughout this disclosure should not be considered in any way as an acknowledgment that such technology is widely known or forms part of common technical knowledge in the art.

[0004] The Motion Picture Encoding (MPEG) is a union of working groups jointly established by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC) to set standards for media encoding, including audio encoding. MPEG is organized under ISO / IEC SC29, and the audio group is currently identified as Working Group (WG) 6. WG6 is currently working on the MPEG-I audio standard.

[0005] The new MPEG-I standard enables different perspectives and / or viewpoints or listening positions for audio experiences by supporting various movements in and around a scene, such as motion with various degrees of freedom, like 3 degrees of freedom (3DOF) or 6 degrees of freedom (6DoF), in virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) applications. 6DoF interaction extends the 3DoF spherical video / audio experience, which is limited to head rotation (pitch, yaw, roll), to include translational motion (forward / backward, up / down, left / right) in addition to head rotation, allowing for navigation within the virtual environment (e.g., physically walking around a room).

[0006] For audio rendering in VR, AR, MR, and XR applications, object-based approaches are widely used by representing complex auditory scenes as multiple distinct audio objects, each associated with parameters or metadata that define its position / location and trajectory within the scene. Alternatively, audio rendering in such environments also utilizes higher-order ambisonics (HOA).

[0007] Audio objects are typically represented as point sources (without extent). As used here, an audio source with "extent" is an audio source waveform associated with a spatial domain (a domain larger than a point). For example, a piano can be represented as an audio source with a cubic extent (e.g., stereo or mono L / R) instead of simply as a point source. The use of extent allows for improved audio experience for the user, for example, when the user is around a virtual piano object in a VR, AR, MR, or XR environment. In this example, the extent representing the piano for audio rendering does not need to have the strict physical details of a real piano.

[0008] To reflect the sonic effects of expansive audio objects, such audio objects may be represented by voxel-based geometry. Voxels for audio rendering are important for media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments.

[0009] However, there is still an existing need for improved rendering of 3D extended sound effects represented by voxel-based geometry, and it is particularly desirable to simplify the process and reduce the computational load. [Overview of the project] [Means for solving the problem]

[0010] In view of the foregoing, the present disclosure provides a method, apparatus, and program for rendering audio in an audio scene, as well as a computer-readable storage medium, each having the characteristics of an independent claim.

[0011] According to a first aspect of this disclosure, a method for rendering audio in an audio scene is provided. This method may include receiving a voxel-based audio scene representation of the audio scene. The audio scene representation may include indications of spread voxels representing a 3D spread, along with a plurality of audio source signals for audio sources associated with the 3D spread. This method may further include obtaining (e.g., determining, calculating) the coordinates of intersections in the 3D spread. This method may further include determining one or more line segments that pass through the intersections and extend along the respective coordinate directions of the audio scene representation. The endpoint of each line segment may be determined based on the coordinates of one or more spread voxels. This method may also include assigning an audio source from the plurality of audio sources to an audio source position in the audio scene based on the one or more line segments.

[0012] In some embodiments, the intersection may be one of the geometric center of the 3D spread and the centroid of the 3D spread.

[0013] In some embodiments, the endpoints of each line segment may be determined based on the extreme coordinate values ​​of the 3D spread along each coordinate direction, so that the length of the line segment corresponds to the maximum dimension of the projection of the 3D spread in each coordinate direction.

[0014] In some embodiments, the audio scene representation may further indicate a hidden voxel. Assigning an audio source may include assigning the audio source to coordinates within a voxel other than the hidden voxel.

[0015] In some embodiments, the audio scene representation may further represent unfilled voxels (e.g., air voxels). Assigning an audio source may involve assigning the audio source to the coordinates on each line segment that are closest to the endpoint of each line segment and are located within an expanding voxel or an unfilled voxel.

[0016] In some embodiments, assigning an audio source may further include determining one or more possible target locations for assigning the audio source based on a line segment.

[0017] In some embodiments, the audio scene representation may further represent unfilled voxels (e.g., air voxels). Determining the one or more possible target locations may involve selecting coordinates for the one or more possible target locations that are closest to the endpoint of each line segment and that lie within an expanding voxel or an unfilled voxel.

[0018] In some embodiments, determining the one or more possible target locations may involve selecting coordinates for the one or more possible target locations that are closest to the endpoints of each line segment and lie within the spread voxel.

[0019] In some embodiments, the method may further include selecting an audio source location from the possible target locations based on a predefined minimum distance between audio sources. The method may also include assigning an audio source from among the plurality of audio sources to the selected audio source location.

[0020] In some embodiments, the method may further include obtaining a mapping that indicates the assignment of audio source signals to audio source locations.

[0021] In some embodiments, the method may further include assigning gain to the audio source location based at least partially on the mapping.

[0022] In some embodiments, the method may further include obtaining the coordinates of the listener's position. The method may also include rendering the audio source signals of the assigned audio sources based on a reference distance between the listener's position and the 3D spread.

[0023] In some embodiments, rendering may further include rendering an audio source signal based on concealment and diffraction modeling.

[0024] A second aspect of this disclosure provides an apparatus for rendering audio in a voxel-based audio scene representation. The apparatus may include one or more processors configured to perform a method which may include receiving a voxel-based audio scene representation of an audio scene, the audio scene representation including indications of spread voxels representing a 3D spread, along with a plurality of audio source signals for audio sources associated with the 3D spread. This method may further include obtaining the coordinates of intersections in the 3D spread. This method may further include determining one or more line segments that pass through the intersections and extend along the respective coordinate directions of the audio scene representation. The endpoints of each line segment may be determined based on the coordinates of one or more of the spread voxels. This method may also include assigning audio sources from the plurality of audio sources to audio source locations in the audio scene based on the one or more line segments.

[0025] Aspects of this disclosure may be implemented via a device. The device may include a processor and memory coupled to the processor. The processor may be adapted to carry out the aspects of this disclosure and the methods according to the embodiments.

[0026] Aspects of this disclosure may be implemented through a program. When program instructions are executed by a processor, the processor can implement aspects and embodiments of this disclosure. A computer-readable storage medium can store a program. Such a computer-readable storage medium may include, but is not limited to, memory devices such as those described herein, including random access memory (RAM) devices and read-only memory (ROM) devices. Thus, some innovative aspects of the subject matter described herein can be implemented through one or more computer-readable storage media on which the software is stored.

[0027] It will be understood that the features of the apparatus and the steps of the method can be replaced in many ways. In particular, the details of the disclosed method can be implemented by the corresponding apparatus (or system), and vice versa. This will be understood by those skilled in the art. Furthermore, it will be understood that any of the above descriptions made relating to the method are similarly applicable to the corresponding apparatus (or system), and vice versa. [Brief explanation of the drawing]

[0028] Herein, exemplary embodiments of the present disclosure will be described simply by reference to the accompanying drawings.

[0029] [Figure 1] An example of a method for rendering audio within an audio scene according to embodiments of this disclosure is shown. [Figure 2] An example of a voxel-based audio scene representation of an audio scene according to an embodiment of the present disclosure is shown. [Figure 3] This embodiment of the disclosure shows an example of assigning an audio source to an audio source location within an audio scene. [Figure 4] Another example of assigning an audio source to an audio source location within an audio scene, according to embodiments of this disclosure, is shown. [Figure 5] This disclosure provides an exemplary use case of an example of a method for rendering audio within an audio scene according to embodiments of this disclosure. [Figure 6] This disclosure provides an exemplary use case of an example of a method for rendering audio within an audio scene according to embodiments of this disclosure. [Figure 7] This disclosure provides an exemplary use case of an example of a method for rendering audio within an audio scene according to embodiments of this disclosure. [Figure 8] This disclosure provides an exemplary use case of an example of a method for rendering audio within an audio scene according to embodiments of this disclosure. [Figure 9]This disclosure provides an exemplary use case of an example of a method for rendering audio within an audio scene according to embodiments of this disclosure. [Figure 10] Examples of reference distances between listener position and 3D spread, as well as examples of concealment and diffraction modeling, are shown according to embodiments of this disclosure. [Figure 11] Examples of devices including one or more processors according to embodiments of this disclosure are shown.

[0030] Where connecting elements such as solid lines, dashed lines, and arrows are used in drawings to illustrate connections, relationships, or associations between two or more other schematic elements, the absence of such connecting elements is not intended to imply that such connections, relationships, or associations cannot exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure this disclosure. Furthermore, for the sake of clarity, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, where a connecting element represents the communication of signals, data, or instructions, it should be understood by those skilled in the art that such an element may, as necessary, represent one or more signal paths affecting the communication. [Modes for carrying out the invention]

[0031] An audio source with spread is an audio source waveform associated with a spatial domain (larger than a point). A spatial domain can be modeled by geometry (2D or 3D). A voxel is a 3D volumetric representation and therefore can model such geometry. The use of voxels for audio rendering is relevant to a variety of media environments implemented in both hardware and software, such as video games and / or VR, AR, MR, and XR environments. A voxel is a spatial volume with assigned acoustic properties or audio rendering instructions. Voxel size is an encoder configuration parameter and can be selected (manually or automatically) according to the level of detail of the scene geometry (e.g., in the range of 10cm to 1m). Voxels for audio rendering can be obtained in the following ways: • Voxelization (or conversion) of mesh-based scene representations - From a scene representation used for scene generation (or even video rendering), for example, by downsampling smaller-sized voxels.

[0032] The methods and apparatus described herein relate to how to render the acoustic effects of a 3D spread when the 3D spread is represented by voxel-based geometry. More specifically, the methods and apparatus described herein relate to how to obtain the coordinates of a "joint" (point) audio source.

[0033] Typically, to approximate the spatial domain of the extent, multiple audio sources are required to model an audio source with extent. These (target) audio sources may be derived from a given audio source(s) related to the extent, specified, for example, by the scene creator using a scene description. The term “congruent” as used here implies that these target sources are related to each other, since they represent the spatial domain of the extent in one dimension. Since there are three dimensions, at least one pair of audio sources is required for each dimension. As an example, a scene description specifies a stereo channel with a cubic extent to represent a virtual piano object. The renderer may then process to derive three pairs of “congruent” “target” audio sources positioned at six different locations within the extent neighborhood.

[0034] In other words, the methods and apparatus described herein use a number of (point) audio sources, for example, N=[1,…,6] and their coordinates (positions) P. 1,…,N Find (for example, select, decide) the audio signal S 1,…,M at each position P 1,…,N The aim is to map to gain, for example, a given scene description including listener position coordinates L, 3D spread material ID (representing an approximation of the audio object's 3D spread geometry), 3D grid index VOX (representing a set of 3D spreads), and a set of M audio signals (mono, stereo, etc.), as well as, for example, the minimum distance Δ between two "congruent" (point) audio sources. min This is based on a modeling setup that includes a mapping matrix F for assigning audio signals to the obtained point source positions (and gains), and a reference distance.

[0035] The methods and apparatus described herein allow for the modeling of audio objects having spread represented by voxel-based geometry without explicitly signaling audio source coordinates (for example, without explicitly sending or receiving this information in a bitstream). In other words, the methods and apparatus described herein emphasize how "congruent" audio source coordinates (positions) are determined within the neighborhood of a spread, assuming that the spread is represented by voxel-based geometry. The resulting positions / coordinates are voxel coordinates. These do not need to be known in advance as they are calculated on the renderer side, and therefore no explicit signaling / transmission is required.

[0036] Advantageously, this allows the decoder to automatically obtain the signal audio source coordinates for complex voxel-based 3D spread geometry, especially if the decoder operates in a manner compliant with audio standards such as those set by MPEG. Another advantage is that this allows for support for 3D spread geometry modifications in the decoder (without the need to re-encode the modified scene).

[0037] 3D spread geometry encoding is performed by the encoder and sent to the decoder, which then delivers information about the spread geometry to the decoder / renderer. Like many other objects in a scene, spread can be modified either on the encoder side or the decoder / renderer side. Modifications made in the encoder require "re-encoding" of the spread, which should then be sent to the decoder. This is not the case for modifications made on the decoder / renderer side. The methods described here are performed on the decoder / renderer side; that is, any modifications to spread are made on the decoder / renderer side and therefore "re-encoding" is not required.

[0038] How to represent the voxel-based audio scene Any voxel-based representation of an audio scene may include indications for voxels that are not transparent voxels (e.g., occluded voxels), i.e., voxels on which sound cannot propagate or freely, i.e., representations of occlusion geometry. This indication may relate to the indication of the coordinates of each voxel (e.g., center coordinates, corner coordinates, etc.). These voxel coordinates may be represented, for example, by grid indices. Furthermore, voxel-based representations may include indications of material properties of non-transparent voxels, such as absorption coefficients and reflection coefficients. In addition to occluded voxels, voxel-based representations may also indicate representations of transparent or unfilled voxels (e.g., air voxels), i.e., voxels on which sound can propagate, i.e., representations of sound propagation media. Thus, some implementations of voxel-based representations of audio scenes may include indications of the respective material properties for each voxel within a predefined section of space (e.g., within the boundary surrounding the audio scene).

[0039] How to render audio in an audio scene Referring to Figure 1, an example of how to render audio within an audio scene is shown. This method is performed on the decoder / renderer side and may be implemented by each decoder / renderer. For example, all method steps may be performed in real time on a single device, which may be a VR / AR / MR / XR device.

[0040] Step S101 Next, a voxel-based audio scene representation of the audio scene is received. The audio scene representation includes indications for spread voxels representing 3D spread, along with multiple audio source signals for audio sources associated with the 3D spread. In other words, 3D spread can be said to correspond to audio objects having a spread that has a geometric shape represented by spread voxels.

[0041] An example of a voxel-based audio scene representation of an audio scene is schematically shown in Figure 2. The example in Figure 2 is a 2D cut through a voxel-based 3D audio scene representation that includes 3D spread. Figure 2 shows a grid pattern representing the voxelization of the audio scene representation. In the example in Figure 2, according to one embodiment, spread voxels 205 and unfilled voxels (e.g., air voxels) 206 are shown. That is, in addition to spread voxels representing 3D spread, the audio scene representation may also show voxels representing parts of the acoustic environment of the 3D spread. Unfilled voxels can be said to represent sound transmission media. Sound transmission media may be, for example, air and / or water.

[0042] Referring again to the example in Figure 1, Step S102 Then, the coordinates of the intersection within the 3D spread are obtained (e.g., determined, calculated). In one embodiment, the intersection may be one of the geometric center and the centroid of the 3D spread. In the example in Figure 2, the geometric center 201 and the center of mass (centroid) 202 of the 3D spread, which can be used alternatively, are schematically shown.

[0043] While not intended to be restrictive, the intersection can be considered the origin O of a Cartesian coordinate system. In the context of the Cartesian coordinate system example, voxel-based 3D spread representation VOX x,y,z Intersection (center of 3D spread) C x,y,z This can be determined using the "minimum / maximum" approach as follows:

number

[0044] Referring again to the example in Figure 1, Step S103Then, one or more line segments are determined, each passing through the intersection and extending along the respective coordinate directions of the audio scene representation (for example, along the x, y, and z coordinate axes). The endpoints of each line segment are determined based on the coordinates of one or more of the spread voxels. For example, as detailed below, the endpoints of each line segment may be determined based on the extreme coordinate values ​​of the 3D spread along the respective coordinate direction. That is, for example, for a line segment extending along the x coordinate axis, the endpoints can be determined based on the extreme coordinates of the 3D spread along the x coordinate axis.

[0045] In the example in Figure 2, two line segments 203 and 204 are shown in the 2D cut, passing through the geometric center 201 of the 3D extension and having endpoints 203a, 203b, 204a, and 204b, respectively. In a Cartesian coordinate system, lines are the X, Y, and Z axes (with the intersection as the origin of the coordinate system), and line segments may be segments of the X, Y, and Z axes. In the 2D cut in Figure 2, line 203 may be the Y axis, and line 204 may be the X axis.

[0046] Referring again to the example in Figure 1, Step S104 Then, one of the multiple audio sources is assigned to an audio source position in the audio scene based on one or more line segments. Here, "assigned" means that the target audio source is generated (for example, based on a given / specified audio source of a certain spread) and linked / mapped to a calculated coordinate position. That is, in step S104, a set of target audio sources placed at a calculated position in the neighborhood of the spread may be output. These target sources (instead of a given / specified audio source with a spread) can be used to replace the task of rendering an "audio source with a spread" by rendering a set of point sources. The one or more line segments are constructed to assist in determining the target audio source positions. Note that in S103, these line segments are output.

[0047] Referring now to FIGS. 3 and 4, two examples of assigning an audio source to an audio source position within an audio scene are schematically shown. That is, FIGS. 3 and 4 represent two possible implementations of method step S104. These implementations differ in the way the target source positions indicated by 308a, 308b, 309a, 309b are determined. In particular, FIGS. 3 and 4 also represent respective 2D cuts.

[0048] FIGS. 3 and 4 show examples of indications of spread voxels 305, non-filled voxels 306, and occluded voxels 307. Occluded voxels may represent, for example, acoustic occlusions existing between a 3D spread and a listener.

[0049] In some embodiments, the endpoints of each line segment may be determined in step S103 based on the extreme coordinate values of the 3D spread along their respective coordinate directions, and thus the length of the line segment may correspond to the maximum dimension of the projection of the 3D spread in their respective coordinate directions.

[0050] As described above, an intersection point inside the 3D spread may be used as the origin O of the Cartesian coordinate system, and each line segment may represent a line segment of the X, Y, and Z axes. Thus, in an example of the Cartesian coordinate system, the maximum dimension (characteristic dimension extreme points) D of the projection of the 3D spread in each coordinate direction max x,y,z and D min x,y,z can be determined (extracted) as follows.

Equation

[0051] The examples of FIGS. 3 and 4 show respective maximum dimensions 303a, 303b and 304a, 304b.

[0052] In some embodiments, as shown in Figure 3, assigning an audio source (to an audio source location in the audio scene) in step S104 may include assigning an audio source to coordinates in a voxel other than the hidden voxel 307 (for example, 308a). The hidden voxel 307 can influence the sound perceived by the listener at each listener position, and by assigning an audio source to coordinates in a voxel other than the hidden voxel 307, each audio source signal of the assigned audio source can be rendered so that the sound perceived by the listener appears realistic.

[0053] Furthermore, in some embodiments, such as those shown in Figure 4, assigning an audio source (to an audio source location within an audio scene) may further include assigning an audio source to the coordinates on each line segment that are closest to the endpoint of each line segment and are located within an expanded voxel or an unfilled voxel (for example, 308a, 308b, 309a, 309b in Figure 4).

[0054] In place of, or in addition to, the above embodiments, assigning an audio source may further include determining one or more possible target locations for assigning an audio source based on a line segment. Determining the one or more possible target locations may include selecting coordinates for the one or more possible target locations that are closest to the endpoint of each line segment and are located within a spreading voxel or an unfilled voxel. In further embodiments, determining the one or more possible target locations may include selecting coordinates for the one or more possible target locations that are closest to the endpoint of each line segment and are located within a spreading voxel.

[0055] Depending on the use case, it becomes possible to render each audio source signal in such a way that the sound perceived by the listener sounds more natural compared to coordinates in an unfilled voxel, by assigning the audio source to coordinates within a spread voxel, or by selecting coordinates for each of the one or more possible target positions that are closest to the endpoint of each line segment and are within the spread voxel.

[0056] In the example of a Cartesian coordinate system, the possible target position (congruent point source coordinates of 3D extent) P max x,y,z , 308a, 309a and P min x,y,z 308b and 309b may be selected as follows: P max x =[p max x ,C y ,C z ] is not a "hidden" voxel for this audio object, line [D max x ,D min x ) D above max x It is the closest voxel. P min x =[p min x ,C y ,C z ] is not a "hidden" voxel for this audio object, but a line (D max x ,D min x ] Above D min x It is the closest voxel.

[0057] Apply the same procedure, P max y,z and P min y,z This can be achieved. A "non-obstructive" voxel is

number

[0058] In some embodiments, the method may further include selecting an audio source location from possible target locations based on a predefined minimum distance between audio sources. The method may also include assigning an audio source from among multiple audio sources to the selected audio source location. For example, the number of (point) audio sources N = [1, ..., 6] can be calculated by considering three variables.

number

[0059] In one embodiment, the method may further include obtaining a mapping that indicates the (e.g., intended or desired) assignment of audio source signals to audio source locations (or possible target locations). For example, M audio signals S 1,…,M Position coordinates P 1,…,6 The following mappings to may be read from the bitstream payload. [Table 1]

[0060] In some embodiments, the method may further include assigning gain to the audio source locations. This assignment may be at least in part based on the mapping described above. The appropriate signal gain may further be assigned based on the number of selected audio sources to ensure energy conservation.

[0061] Referring here to the examples in Figures 5–9, an example use case of the method of rendering audio in an audio scene as described herein is shown. In this exemplary use case, the 3D extent to be rendered / modeled is, exemplary, based on a tram 500. Figures 7–9 show the respective “visible” line segments 501, 502, 503 and their respective assigned coordinates / target position coordinates 504a, 504b, 505a, 505b, 506a, 506b, which are determined according to the method described herein.

[0062] In this exemplary use case, the renderer is tasked with appropriately rendering the sound of a virtual tram in a VR / AR / XR / MR scene. The tram may be seen running along a busy street. The audio emitted from the tram comes from several parts distributed along the length of the tram, so the tram cannot be modeled by a single point source. First, a “tram” object in the VR / AR / XR / MR scene (Figure 5) and accompanying “spread audio sources (one or more)” to represent the sound of the “tram” may be specified by the scene creator as part of the “scene description” of the VR / AR / XR / MR scene. A specified “spread” model representing the tram for audio rendering is shown as an example in Figure 6. Figures 7, 8, and 9 show possible embodiments of determining a set of target audio source locations 504a, 504b, 505a, 505b, 506a, 506b in the vicinity of the spread by applying the method shown in Figure 1. Subsequently, actual target audio sources corresponding to those locations are generated. In this context, Figures 8 and 9 are taken from Figure 7 by cutting the tram spread representation to indicate the target audio source locations. Thus, the sound of the tram in the VR / AR / XR / MR scene arises from the rendering of those target audio sources.

[0063] Referring here to the example in Figure 10, in one embodiment, the method may further include obtaining the coordinates of the listener position 510 and rendering the audio source signal of the assigned audio source based on a reference distance 511 between the listener position 510 and the 3D spread 500. For example, the distance from the listener L to the nearest point R of the spread VOX may be subtracted from the distance (L,P) from the listener to the object, where |LR| == min(|L - VOX|).

[0064] In some embodiments, rendering may further include rendering (point) audio source signals based on (voxel-based) occlusion and diffraction modeling. That is, for example, a selected subset of point audio sources {P max x,y,z ,P min x,y,z The} may be rendered by applying voxel-based occlusion and diffraction modeling. An example in Figure 10 shows a 3D spread representation of a tram 500 occluded by an obstacle 512. Coordinates 504a, 504b, 506a, and 518 are thus occluded, and as a result, diffraction modeling generates a set of virtual coordinates 516, 516a-d.

[0065] The 3D spread modeling method described herein assumes the application of diffraction modeling. However, this method can also be used without applying diffraction modeling. In this case, the method is applied to the 3D spread subset visible to listener 510. To obtain a “visible” subset for the listener (visibility implies that there are no acoustic occlusions between the listener and the corresponding point), the following methods may be used: a ray tracing-based method, or a method by checking for linear occlusions between the listener and the subset of the 3D spread representation. This subset can be determined by Monte Carlo or any other subsampling method.

[0066] Exemplary algorithm In other words, the method for rendering audio in an audio scene can be described as follows. The following represents an exemplary implementation of the method shown in Figure 1. This assumes that the decoder has already received “scene description” information, including the aspects shown in step S101.

[0067] things that are given Scene description : Listener's position coordinate L 3D Spread Material ID (Represents the approximation of the 3D spread geometry of the audio object) 3D grid index set VOX (represents a set of 3D spreads) A set of M audio signals (mono, stereo, etc.) Modeling settings : Minimum distance Δ between two "congruent" point audio sources min Mapping matrix F for assigning audio signals to the obtained point source positions (and gains) Reference distance What to discover The number of point sound sources N = [1, ..., 6] and their coordinates P 1,…,N Audio signal S 1,…,M Position P 1,…,N and mapping to gain

[0068] solution For example, using the "minimum / maximum (geometric center)" approach, voxel-based 3D spread representation VOX x,y,z 3D spread center representation coordinate C x,y,z Determine (Step S102):

number

[0069] A central representation is necessary to extract three characteristic dimensions from a 3D spread representation.

[0070] Next extreme point D max x,y,z and D min x,y,z This determines the characteristic dimensional representation of the 3D extent (step S103):

number

[0071] 3D spread congruent point source coordinates Pmax x,y,z and P min x,y,z are determined (step S104). Here, P max x = [p max x , C y , C z is not a "hidden" voxel for this audio object, and is the voxel on the line [D max x , D min x ) closest to D max x . Here, p min x = [p min x , C y , C z is not a "hidden" voxel for this audio object, and is the voxel on the line (D max x , D min x closest to D min x . <0OO00438> Applying the same procedure, P max y,z and P min y,z can be obtained. The "non-shielding" voxels can be defined as any of

Number

[0073] Considering the three variables<OO000449>

Number

[0074] W x,y,z > Δ min If so, the audio source Pmax x,y,z and P min x,y,z The "x-", "y-", and "z-" pairs are taken into consideration for 3D spread modeling. This is done to prevent phasing audio artifacts caused by two correlated audio signals being rendered too close together.

[0074] All W x,y,z ≤Δ min In this case, the 3D spread is not modeled, and the audio object is a voxel C x,y,z It is represented by a single audio point source located at [location].

[0075] Based on the bitstream payload, M audio signals S, as shown in the exemplary mapping in Table 1 above. 1,…,M Position coordinates P 1,…,6 A mapping is obtained to M. For example, in some implementations, the mapping M may be read from or extracted from a bitstream.

[0076] The appropriate signal gain is assigned based on the number of selected audio sources (to ensure energy conservation).

[0077] Voxel-based concealment and diffraction modeling is applied to a selected subset of point audio sources {P max x,y,z ,P min x,y,z Render the} (see Figure 10 as an example).

[0078] Apply a baseline distance calculation. That is, subtract the distance from the listener L to the nearest point R in the spread VOX from the distance (L,P) from the listener to the object. Here, |LR| == min(|L - VOX|).

[0079] Applying voxel-based hiding and diffraction modeling to a selected subset of point audio sources {P maxx,y,z ,P min x,y,z Render}.

[0080] Rendering may be performed by a renderer capable of simulating acoustic obscuration and diffraction modeling.

[0081] Figure 10 shows a 3D spread representation of an object obscured by an obstacle, namely a streetcar, as an example of how diffraction processing can be applied to the methods described herein. The central points are occluded, and as a result, a set of virtual points is generated by diffraction modeling. 513 and 514 are lines of sight not obstructed by the occluding object 512. 515 is the direction (azimuth angle) in which object 518 is perceived. These objects are then perceived (modeled) as 516. As a result, coordinates 504a, 504b, 506a, and 506b (missing in the figure) belonging to 518 are modeled by 516a, 516d, 516b, and 516c belonging to 516.

[0082] The 3D spread modeling method described here assumes the application of diffraction modeling.

[0083] However, this method can also be used without applying diffraction modeling. In this case, the method is applied to a 3D spread subset that is visible to the listener. To obtain a subset that is "visible" to the listener (visibility implies that there are no acoustic occlusions between the listener and the corresponding point), the following methods can be used: a method based on ray tracing, or a method by checking for linear occlusions between the listener and the subset of the 3D spread representation. This subset can be determined by Monte Carlo or any other subsampling method.

[0084] Referring to the example in Figure 11, an apparatus 1100 comprising one or more processors 1101, 1102 according to an embodiment of the present disclosure is shown. The one or more processors 1101, 1102 may be configured to perform the methods described herein.

[0085] A computing device implementing the techniques described above may have the following exemplary architecture. Other architectures are also possible, including architectures with more or fewer components. In some implementations, the exemplary architecture includes one or more processors (e.g., dual-core Intel® processors), one or more output devices (e.g., LCDs), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, hard disk, optical disc, flash memory, etc.). These components can communicate and exchange data through one or more communication channels (e.g., buses), and various hardware and software can be utilized to facilitate the transfer of data and control signals between components.

[0086] The term "computer-readable medium" refers to, but is not limited to, any medium involved in providing instructions to a processor for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media include, but are not limited to, coaxial cables, copper wires, and optical fibers.

[0087] Computer-readable media may further include an operating system (e.g., the Linux® operating system), a network communication module, an audio interface manager, an audio processing manager, and a live content distributor. The operating system may be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system performs basic tasks including, but not limited to, recognizing input from network interfaces and / or devices and providing output to them, tracking and managing files and directories on computer-readable media (e.g., memory or storage devices), controlling peripheral devices, and managing traffic on one or more communication channels. The network communication module includes various components for establishing and maintaining network connectivity (e.g., software for implementing communication protocols such as TCP / IP and HTTP).

[0088] The architecture can be implemented using parallel processing or peer-to-peer infrastructure, or on a single device with one or more processors. The software can consist of multiple software components or be a single code body.

[0089] The described features can be advantageously implemented in one or more computer programs executable on a programmable system comprising a data storage system, at least one input device, and at least one output device coupled to receive data and instructions from and transmit data and instructions to them. A computer program is a set of instructions that can be used directly or indirectly within a computer to perform an activity or produce a result. A computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0090] A suitable processor for executing a program of instructions includes, for example, both general-purpose and dedicated microprocessors, as well as one of a single processor or multiple processors or cores in any type of computer. Generally, a processor receives instructions and data from read-only memory or random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer includes or is operationally coupled to one or more mass storage devices for storing data files. Such devices include magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Storage devices suitable for materially embodying computer program instructions and data include, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and all forms of non-volatile memory, including CD-ROMs and DVD-ROM disks. Processors and memory can be complemented by or incorporated into ASICs (Application-Specific Integrated Circuits).

[0091] To provide user interaction, these features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or retinal display device for displaying information to the user. The computer may have a touch surface input device (e.g., a touchscreen) or keyboard and a pointing device such as a mouse or trackball that allows the user to provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0092] These features can be implemented in computer systems including backend components such as data servers, computer systems including middleware components such as application servers or internet servers, or computer systems including frontend components such as client computers with graphical user interfaces or internet browsers, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication, such as communication networks. Examples of communication networks include, for example, computers and networks that form LANs, WANs, and the Internet.

[0093] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact through a communication network. The client-server relationship is established by computer programs running on each computer that have a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to the client device (for example, to display data to a user interacting with the client device and to receive user input from the user). Data generated by the client device (e.g., the results of user interaction) can be received by the server from the client device.

[0094] One or more computer systems can be configured to perform a specific action by installing software, firmware, hardware, or a combination thereof on the system that causes the system to perform an action while it is running. One or more computer programs can be configured to perform a specific action by containing instructions that cause the data processing device to perform an action when executed by the device.

[0095] This specification includes many specific implementation details, which should not be construed as limitations on the scope of any invention or claim, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments. Furthermore, features may be described above as acting in some combination, and may even be initially claimed as such, but one or more features from a claimed combination may, in some cases, be removed from such combination, and the claimed combination may be directed towards a subcombination or a variation of a subcombination.

[0096] Similarly, although the operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in a specific or sequential order shown, or that all illustrated operations be performed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0097] Unless otherwise explicitly stated, as will be evident from the following discussion, any discussion using terms such as “processing,” “computing,” “calculating,” “determining,” and “analyzing” throughout this disclosure is understood to refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data, which is expressed as a physical quantity such as an electronic quantity, into other data, which is similarly expressed as a physical quantity.

[0098] Any reference throughout this disclosure to “one exemplary embodiment,” “several exemplary embodiments,” or “a certain exemplary embodiment” means that any particular feature, structure, or characteristic described in relation to that exemplary embodiment is included in at least one exemplary embodiment of this disclosure. Therefore, the phrases “one exemplary embodiment,” “several exemplary embodiments,” or “a certain exemplary embodiment” appearing in various places throughout this disclosure do not necessarily all refer to the same exemplary embodiment. Furthermore, any particular feature, structure, or characteristic can be combined in any suitable manner as will be apparent to those skilled in the art from this disclosure in one or more exemplary embodiments.

[0099] Where used herein, unless otherwise explicitly stated, the use of ordinal adjectives such as “first,” “second,” and “third” to describe a common object simply indicates that different instances of a similar object are being referred to, and is not intended to imply that the objects described in this way must be in a given sequence in any temporal, spatial, ranking, or any other manner.

[0100] Furthermore, it should be understood that the expressions and terminology used herein are for illustrative purposes only and should not be considered limiting. The use of “include,” “contain,” or “have,” and their variations, means to include the enumerated items and their equivalents, as well as any additional items. Unless specifically stated or limited otherwise, the terms “attached,” “connected,” “supported,” and “joined,” and their variations, are used broadly and include both direct and indirect attachment, connection, support, and joining.

[0101] In the following claims and description herein, any term “equipped with,” “consisting of,” or “having” is an open term meaning that it includes at least the elements / features listed, but does not exclude others. Therefore, when used in the claims, the terms “having / including” should not be interpreted as limiting to the means, elements, or steps listed. For example, the expression “apparatus having A and B” should not be limited to an apparatus consisting only of elements A and B. Any term “containing,” “encompassing,” or “inclusive” as used herein is also an open term meaning that it includes at least the elements / features listed in that term, but does not exclude others. Therefore, “containing” is synonymous with “having” and means “having.”

[0102] In the above description of the exemplary embodiments of the Disclosure, it should be understood that various features of the Disclosure may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of improving the flow of the Disclosure and aiding in the understanding of one or more of the various inventive aspects. However, this method of disclosure should not be construed as reflecting an intention that the claims require more features than expressly described in each claim. Rather, as reflected in the following claims, the inventive aspects are fewer than all the features of a single, aforementioned exemplary embodiment. Thus, the claims following the specification are hereby expressly incorporated herein, and each claim stands alone as a separate exemplary embodiment of the Disclosure.

[0103] Furthermore, while some exemplary embodiments described herein include some features found in other exemplary embodiments, combinations of features from different exemplary embodiments are intended to be within the scope of this disclosure and form different exemplary embodiments, as will be understood by those skilled in the art. For example, any of the claimed exemplary embodiments may be used in any combination within the following claims.

[0104] Numerous specific details are described in the description provided herein. However, it should be understood that exemplary embodiments of this disclosure may be carried out without these specific details. On the other hand, well-known methods, structures, and techniques are not described in detail so as not to obscure the understanding of this paper.

[0105] Therefore, while the best aspects of this disclosure are described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of this disclosure. All such changes and modifications that fall within the scope of this disclosure are intended to be requested. For example, the formulas given above are merely representative of the procedures that may be used. Functions may be added to or removed from block diagrams, and operations may be swapped between function blocks. Steps may be added to or removed from methods described within the scope of this disclosure.

[0106] Several aspects are described below. [Aspect 1] A method for rendering audio in an audio scene, the method being: Step (S101) of receiving a voxel-based audio scene representation of the audio scene, wherein the audio scene representation includes indications for expansion voxels (205;305) representing 3D expansion, along with a plurality of audio source signals for audio sources associated with the 3D expansion; The step (S102) involves obtaining the coordinates of the intersection (201, 202; 301) within the aforementioned 3D spread; Step (S103) is to determine one or more line segments (203, 204; 303, 304) that pass through the aforementioned intersection (201, 202; 301) and extend along the respective coordinate directions of the audio scene representation, wherein the endpoints (203a, 203b, 204a, 204b; 303a, 303b, 304a, 304b) of each line segment (203, 204; 303, 304) are determined based on the coordinates of one or more spreading voxels (205; 305); Step (S104) is to assign an audio source from among the plurality of audio sources to an audio source position (308a, 308b, 309a, 309b) in the audio scene based on one or more line segments (203, 204; 303, 304) and Methods that include... [Aspect 2] The method according to embodiment 1, wherein the intersection is one of the geometric center of the 3D spread and the centroid of the 3D spread. [Aspect 3] The method according to embodiment 1, wherein the endpoints of each line segment are determined based on the extreme coordinate values ​​of the 3D spread along their respective coordinate directions, and the length of each line segment corresponds to the maximum dimension of the projection of the 3D spread in each coordinate direction. [Aspect 4] The aforementioned audio scene representation further shows concealed voxels, Assigning the audio source includes assigning the audio source to coordinates within a voxel other than the hidden voxel. The method described in Embodiment 1. [Aspect 5] The aforementioned audio scene representation further shows unfilled voxels, Assigning the audio source involves assigning the audio source to the coordinates on each line segment that are closest to the endpoint of each line segment and are located within an extended voxel or an unfilled voxel. The method according to aspect 4. [Aspect 6] The method according to embodiment 1, wherein assigning the audio source further comprises determining one or more possible target positions for assigning the audio source based on the line segment. [Aspect 7] The aforementioned audio scene representation further shows unfilled voxels, Determining the one or more possible target locations involves selecting coordinates for the one or more possible target locations that are closest to the endpoints of each line segment and are located within an extended voxel or an unfilled voxel. The method according to embodiment 6. [Aspect 8] The method according to embodiment 6, wherein determining the one or more possible target locations includes selecting coordinates for the one or more possible target locations that are closest to the endpoint of each line segment and located within the spread voxel. [Aspect 9] This method further: A step of selecting the audio source position from the possible target positions based on a predefined minimum distance between audio sources; The step includes assigning one of the plurality of audio sources to the selected audio source position. The method according to embodiment 6. [Aspect 10] The method according to embodiment 1, further comprising the step of obtaining a mapping indicating the assignment of the audio source signal to the audio source location. [Aspect 11] The method according to embodiment 10, further comprising assigning gain to the audio source location based at least partially on the mapping. [Aspect 12] This method further: The step of obtaining the coordinates of the listener's location; The steps include rendering the sound source signal of the assigned audio source based on a reference distance between the listener position and the 3D spread, The method described in Embodiment 1. [Aspect 13] The method according to embodiment 12, further comprising rendering the audio source signal based on concealment and diffraction modeling. [Aspect 14] A device (1100) for rendering audio in a voxel-based audio scene, the device comprising one or more processors (1101, 1102) configured to perform a method, the method being: Step (S101) of receiving a voxel-based audio scene representation of the audio scene, wherein the audio scene representation includes indications for expansion voxels (205;305) representing 3D expansion, along with a plurality of audio source signals for audio sources associated with the 3D expansion; The step (S102) involves obtaining the coordinates of the intersection (201, 202; 301) within the aforementioned 3D spread; Step (S103) is to determine one or more line segments (203, 204; 303, 304) that pass through the aforementioned intersection (201, 202; 301) and extend along the respective coordinate directions of the audio scene representation, wherein the endpoints (203a, 203b, 204a, 204b; 303a, 303b, 304a, 304b) of each line segment (203, 204; 303, 304) are determined based on the coordinates of one or more spreading voxels (205; 305); Step (S104) is to assign an audio source from among the plurality of audio sources to an audio source position (308a, 308b, 309a, 309b) in the audio scene based on one or more line segments (203, 204; 303, 304) and A device including a device. [Aspect 15] A program that, when executed by a processor, includes instructions that cause the processor to perform the method described in any one of the embodiments 1 to 13. [Aspect 16] A computer-readable storage medium storing the program described in embodiment 15. Various aspects and implementations of this disclosure can also be understood from the following enumerated example embodiments (EEEs) which are not part of the claims.

[0107] [EEE1] A method for modeling augmented audio objects for audio rendering in a virtual or augmented reality environment, the method being: Determine the central representation of 3D spread in voxel-based 3D spread representation; Based on the 3D spread representation, the 3D spread characteristic dimension representation is determined; This includes determining the coordinates of a 3D spread congruent point source based on a 3D spread feature dimension representation or a 3D spread center representation. method. [EEE2] Receive a mapping of M audio signals to their positional coordinates; The further includes assigning the signal gains of the M audio signals to the point source. Methods described in EEE1. [EEE3] The method according to EEE1, further comprising rendering the point source based on voxel-based concealment and diffraction modeling. [EEE4] The method according to EEE1, wherein the central representation can be determined based on the geometric center of the voxel or by a centroid approach. [EEE5] A method according to EEE1, wherein the dimension representation may be based on an endpoint or by calculating a corresponding offset from the center. [EEE6] The method according to any one of EEE1 to 5, wherein the method is applied to a subset of the voxel-based 3D spread representation. [EEE7] The method according to EEE6, wherein the subset of the voxel-based 3D spread representation corresponds to acoustically unobstructed (visible) voxels. [EEE8] A non-temporary computer program that, when executed by a processor, includes instructions causing the processor to perform any of the methods described in any one of EEE1 through 7. [EEE9] An apparatus configured to carry out the method described in any one of EEE1 to EEE7.

Claims

1. A method for rendering audio in an audio scene, performed by one or more processors, the method being: Step (S101) of receiving a voxel-based audio scene representation of the audio scene, wherein the audio scene representation includes indications for expansion voxels (205; 305) representing 3D expansion, along with a plurality of audio source signals for audio sources associated with the 3D expansion; The step (S102) is to obtain the coordinates of the intersections (201, 202; 301) within the aforementioned 3D spread; Step (S103) is to determine one or more line segments (203, 204; 303, 304) that pass through the aforementioned intersection (201, 202; 301) and extend along the respective coordinate directions of the audio scene representation, wherein the endpoints (203a, 203b, 204a, 204b; 303a, 303b, 304a, 304b) of each line segment (203, 204; 303, 304) are determined based on the coordinates of one or more spreading voxels (205; 305); Step (S104) of assigning an audio source signal from among the plurality of audio source signals to an audio source position (308a, 308b, 309a, 309b) in the audio scene based on one or more line segments (203, 204; 303, 304) and Methods that include...

2. The method according to claim 1, wherein the intersection is one of the geometric center of the 3D spread and the centroid of the 3D spread.

3. The method according to claim 1, wherein the endpoints of each line segment are determined based on the extreme coordinate values ​​of the 3D spread along their respective coordinate directions, and the length of each line segment corresponds to the maximum dimension of the projection of the 3D spread in each coordinate direction.

4. The aforementioned audio scene representation further shows the concealed voxels, Assigning the aforementioned audio source signal includes assigning the aforementioned audio source signal to coordinates within a voxel other than the hidden voxel. The method according to claim 1.

5. The aforementioned audio scene representation further shows unfilled voxels, Assigning the aforementioned audio source signal includes assigning the audio source signal to the coordinates on each line segment that are closest to the endpoint of each line segment and located within a spread voxel or an unfilled voxel. The method according to claim 4.

6. The method according to claim 1, wherein assigning the audio source signal further comprises determining one or more possible target positions for assigning the audio source signal based on the line segment.

7. The aforementioned audio scene representation further shows unfilled voxels, Determining the one or more possible target locations involves selecting coordinates for the one or more possible target locations that are closest to the endpoints of each line segment and are located within an extended voxel or an unfilled voxel. The method according to claim 6.

8. The method according to claim 6, wherein determining the one or more possible target locations includes selecting coordinates for the one or more possible target locations that are closest to the endpoint of each line segment and located within the spread voxel.

9. This method further: A step of selecting the audio source position from the possible target positions based on a predefined minimum distance between audio sources; The step includes assigning an audio source signal from among the plurality of audio source signals to the selected audio source position. The method according to claim 6.

10. The method according to claim 1, further comprising the step of obtaining a mapping indicating the assignment of the audio source signal to the audio source location.

11. The method according to claim 10, further comprising assigning gain to the audio source location based at least partially on the mapping.

12. This method further: The step of obtaining the coordinates of the listener's location; The steps include rendering the assigned audio source signal based on a reference distance between the listener's position and the 3D spread, The method according to claim 1.

13. The method according to claim 12, further comprising rendering the audio source signal based on concealment and diffraction modeling.

14. A device (1100) for rendering audio in a voxel-based audio scene, the device comprising one or more processors (1101, 1102) configured to perform a method, the method being: Step (S101) of receiving a voxel-based audio scene representation of the audio scene, wherein the audio scene representation includes indications for expansion voxels (205; 305) representing 3D expansion, along with a plurality of audio source signals for audio sources associated with the 3D expansion; The step (S102) is to obtain the coordinates of the intersections (201, 202; 301) within the aforementioned 3D spread; Step (S103) is to determine one or more line segments (203, 204; 303, 304) that pass through the aforementioned intersection (201, 202; 301) and extend along the respective coordinate directions of the audio scene representation, wherein the endpoints (203a, 203b, 204a, 204b; 303a, 303b, 304a, 304b) of each line segment (203, 204; 303, 304) are determined based on the coordinates of one or more spreading voxels (205; 305); Step (S104) of assigning an audio source signal from among the plurality of audio source signals to an audio source position (308a, 308b, 309a, 309b) in the audio scene based on one or more line segments (203, 204; 303, 304) and A device including a device.

15. A program that, when executed by a processor, includes instructions causing the processor to perform the method described in any one of claims 1 to 13.

16. A computer-readable storage medium storing the program described in claim 15.

Citation Information

Patent Citations

  • Diffraction modelling based on grid pathfinding

    WO2021198152A1