Rendering of Occluded Audio Elements
The two-stage occlusion detection method for audio elements with extents optimizes raycast usage, enabling accurate and efficient occlusion rendering in XR environments by distinguishing complete and partial occlusions, reducing computational burden.
Patent Information
- Application Number
- JP2024569170
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-13
- Filing Date
- 2023-07-03
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-07-03
AI Technical Summary
Existing occlusion rendering techniques for audio elements with an extent face significant processing complexity due to the need for a large number of rays to detect partial occlusion, which is not efficiently handled by current methods.
A two-stage occlusion detection method is employed, where the first stage identifies complete occlusions by determining the edges of the audio element's extent, and the second stage calculates occlusion filters for sub-areas within the modified extent using a sparse grid of rays, optimizing the number of raycasts required.
This approach allows for accurate detection of occlusion with reduced computational complexity, ensuring smooth and responsive occlusion rendering without perceptible steps, particularly in extended reality (XR) environments.
Smart Images

Figure 2025521414000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments related to the rendering of occluded audio elements are disclosed.
Background Art
[0002] Spatial audio rendering is a process used to present audio within an extended reality (XR) scene (e.g., virtual reality (VR), augmented reality (AR), mixed reality (MR) scene), where sound arrives from physical sound sources within the scene at specific positions and gives the listener the impression of having a specific size and shape (i.e., extent). The presentation can be done via headphones or other speakers. When the presentation is done via headphones, the processing used is called binaural rendering and uses the spatial cues of human spatial hearing that enable identifying from which direction the sound is arriving. The cues involve interaural time delay (ITD), interaural level difference (ILD), and / or spectral differences.
[0003] The most common form of spatial audio rendering is based on the concept of point sources, where each sound source is defined as emitting sound from one specific point. Since each sound source is defined as emitting sound from one specific point, the sound source has no size or shape. Different methods have been developed for rendering sound sources with an extent (size and shape).
[0004] One such known method is to create multiple copies of a monaural audio element at positions around the audio element. This configuration creates the perception of a spatially homogeneous object having a particular size. This concept is used, for example, in the "object spread" and "object divergence" features of the MPEG-H 3D Audio standard (see References [1] and [2]), and in the "object divergence" feature of the EBU Audio Definition Model (ADM) standard (see Reference [4]). This idea of using a monaural audio source has been further developed as described in Reference [7], where the area-volume geometry of the sound object is projected onto a sphere around the listener and the sound is rendered to the listener using a pair of head-related (HR) filters evaluated as the integral of all HR filters covering the geometric projection of the object on the sphere. For a spherical volume sound source, this integral has an analytical solution. However, for any area-volume sound source geometry, the integral is evaluated by sampling the projected sound source surface on the sphere using so-called Monte Carlo ray sampling.
[0005] Another rendering method is to render a spatially diffuse component in addition to the monaural audio signal, which creates the perception of a somewhat diffuse object that does not have a distinct pinpoint location, as opposed to the original monaural audio element. This concept is used, for example, in the "object diffuseness" feature of the MPEG-H 3D Audio standard (see Reference [3]) and in the "object diffuseness" feature of the EBU ADM (see Reference [5]).
[0006] Combinations of the two methods described above are also known. For example, the "object extent" feature of the EBU ADM combines the creation of multiple copies of a monaural audio element with the addition of a diffuse component (see Reference [6]).
[0007] In many cases, the actual shape of the audio element can be adequately described by a basic shape (e.g., a sphere or a box). However, the actual shape may be more complex and may need to be described in a more detailed format (e.g., a mesh structure or a parametric description format).
[0008] In the case of heterogeneous audio elements, as described in reference [8], the audio element includes at least two audio channels (i.e., audio signals) to describe the spatial variation over its range.
[0009] In some XR scenes, there may be objects that block at least a portion of the audio elements within the XR scene. In such scenarios, the audio elements are said to be at least partially occluded.
[0010] That is, occlusion occurs when, from the perspective of a listener (e.g., a human listener) at a given listening position, the audio object is completely or partially blocked behind some object such that the direct sound from the occluded portion of the object does not reach the listener or hardly reaches the listener. Depending on the material of the occluding object, the occlusion effect can be either complete occlusion (e.g., when the occluding object is a thick wall) or incomplete occlusion (e.g., when the occluding object is made of a thin fabric such as a curtain) (also known as "soft" occlusion). Soft occlusion can often be well described by a filter having a specific frequency response that matches the acoustic characteristics of the material of the occluding object. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0011] Currently, there are several problems. For example, available occlusion rendering techniques can handle point sound sources that can easily detect the occurrence of occlusion using raytracing between the listener position and the point sound source position. However, for audio elements with an extent, the situation is more complex because the occluding object may only occlude a part of the extent of the audio element. To obtain sufficient resolution of the occlusion effect, a large number of rays need to be projected towards its extent, which will add a significant amount of processing complexity to the rendering of audio objects.
Means for Solving the Problem
[0012] Therefore, on one side, an improved method for rendering an audio element associated with a first extent is provided. In one embodiment, the method includes determining a first point in the first extent that is not completely occluded. The method also includes determining a second extent for the audio element, where determining includes determining a first edge of the second extent using the first point. The method also includes, after determining the second extent, dividing the second extent into a set of one or more sub-areas including a first sub-area. The method also includes determining a first gain value for a first sample point of the first sub-area. The method further includes rendering the audio element using the first gain value (e.g., generating an output signal using the first gain value).
[0013] In another aspect, there is provided a computer program comprising instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to execute the methods disclosed herein. In one embodiment, there is provided a carrier containing the computer program, the carrier being one of an electrical signal, an optical signal, a wireless signal, and a computer-readable storage medium. In another aspect, there is provided a rendering device configured to execute the methods disclosed herein. The rendering device may include a memory and a processing circuit coupled to the memory.
[0014] The advantage of the embodiments disclosed herein is that it enables occlusion of an extent to be detected with good accuracy using a limited number of raycasts.
Brief Description of the Drawings
[0015] The accompanying drawings, which are incorporated herein and form a part of this specification, illustrate various embodiments.
[0016]
Figure 1
[0017]
Figure 2
[0018]
Figure 3
[0019]
Figure 4
[0020]
Figure 5
[0021]
Figure 6A
[0022]
Figure 6B
[0023]
Figure 7A
[0024]
Figure 7B
[0025]
Figure 8
[0026]
Figure 9A
[0027]
Figure 9B
[0028]
Figure 10
[0029]
Figure 11
[0030]
Figure 12
[0031] The occurrence of occlusion can be detected using a raytracing method in which a direct sound path (or simply "path") between the listener position and the position of the audio element is searched for any object that occludes the audio element. FIG. 1 shows an example of two point sound sources (S1 and S2), where one (i.e., S2) is occluded by an object (O) (referred to as an "occluding object") from the listener's perspective, and the other (i.e., S1) is not occluded from the listener's perspective. In this case, the occluded audio element S2 should be muted in a way that corresponds to the acoustic characteristics of the data of the occluding object. If the occluding object is a thick wall, the rendering of the direct sound from the occluded audio element should be muted more or less completely.
[0032] For a given frequency range, any given portion of an audio element may be completely occluded, partially occluded, or not occluded at all. The frequency range may be the entire frequency range perceivable by humans or a subset of that frequency range. In one embodiment, a portion of an audio element is completely occluded in a given frequency range when an occlusion gain factor (or abbreviated as "gain") associated with a portion of the audio element meets a predetermined condition. For example, when the occlusion gain (which may or may not be frequency-dependent) associated with a portion of the audio element is below a threshold gain value (T), the portion of the audio element is completely occluded in the given frequency range, where the value T is a selected value (e.g., T = 0 is one possibility). That is, for example, any one or more occluding objects that pass less than a certain amount of sound are considered complete occlusion. In another embodiment, there is a frequency-dependent determination in which the amount of occlusion in different frequency bands is compared to a predetermined table of thresholds for these frequency bands. In yet another embodiment, the current signal power of an audio signal representing an audio source is used to estimate the actual sound power passed to a listener, and then the sound power is compared to an auditory threshold. In short, a completely occluded audio element (or a part thereof) can be defined as a sound path that is not perceptually important because the sound is highly suppressed. This includes cases where it is completely blocked by occlusion, i.e., no sound passes through, or where only a very small amount of the original sound energy passes through the occluding object and does not contribute sufficiently to have a perceptual impact on the overall rendering of the audio source.
[0033] For example, when there is a "hard" occluding object on the sound path, i.e., a virtual straight line from the listening position to a portion of the audio element, a portion of the audio element is completely occluded. An example of a hard occluding object is a thick brick wall. On the other hand, for example, when there is a "soft" occluding object on the sound path, a portion of the audio element may be partially occluded. An example of a soft occluding object is a thin curtain.
[0034] If one or several soft occluding objects are on the sound path, the occlusion effect can be calculated as a filter corresponding to the audio transmission characteristics of the data. This filter can be specified as a list of frequency ranges, and for each of the listed frequency ranges, a corresponding gain can be specified. If two or more soft occluding objects are in the path, the filters of the data of those objects can be multiplied together to form one composite filter corresponding to the audio transmission characteristics of that path.
[0035] Ray tracing can be started by specifying a start point and an end point, or, in addition to the horizontal and vertical angles, optionally by specifying the start point and direction of the ray in a polar format that also implies a length. Occlusion detection is repeated periodically in time, or whenever there is an update to the scene, so that the renderer has the latest occlusion information.
[0036] In the case of the audio element 202 having the extent 204, as shown in FIG. 2, the extent of the audio element can be occluded only partially by the occluding object 206. This means that it is necessary to change the rendering of the audio element 202 so as to reflect which part of the extent is occluded and which part is not occluded. The extent 204 can be the actual extent of the audio element 202 as seen from the listener position, or a projection of the audio element 202 as seen from the listener position. Here, the projection can be, for example, a projection of the extent of the audio element onto a sphere around the listener, or a projection of the extent of the audio element onto a plane between the audio element and the listener.
[0037] 1.1. Modes of Occlusion Detection and Rendering
[0038] The process of detecting the occlusion of the extent involves checking the paths between the listener's perspective, typically the listener's position ("listening position"), and each of a number of points on the extent for occluding the object. Some processing is required in both the geometric calculations involved in ray tracing and the calculations of the audio transmission filters. This means that the number of these paths (i.e., the points on the extent) to be checked should be minimized.
[0039] When rendering the occlusion effect of an audio object with an extent, there are perceptually important aspects. The human auditory system is very good at creating the angle of an audio object in the horizontal plane, often called the azimuth, because it can utilize the timing difference between the sounds reaching the right and left ears respectively. Under ideal circumstances, a human can distinguish a difference in horizontal angle of just 1°. For the vertical angle, often called the elevation, there is no timing difference that can assist the auditory system. Instead, the only clue that spatial hearing uses to distinguish different vertical angles is the difference in frequency response resulting from filtering from the ears, which is different for different vertical angles. Therefore, the accuracy of the vertical angle is less perceptually important than the horizontal angle.
[0040] Even if the vertical positions of the upper and lower ends of an extent are not perceptually important, the detection of these ends can potentially affect the overall energy of the audio object. The required resolution due to changes in energy can be higher than that due to the perceived changes in spatial position.
[0041] In the case of an audio object with an extent, the outer edges are the most prominent features, but the energy distribution across that extent also needs to be reasonably well reflected. It is important that changes in position, energy, or filtering are smooth and do not change in discrete steps unless there are sudden movements of the audio source, listener, or some occluder (or any combination of them).
[0042] The easiest way to avoid step - wise changes in occlusion detection is to use a large number of ray casts so that smooth changes in occlusion can be tracked at a high resolution. This can make the steps small enough to be imperceptible. However, a large number of ray casts significantly increases the complexity of the algorithm.
[0043] Another solution is to add temporal smoothing of occlusion detection and / or occlusion rendering. This makes abrupt steps uniform and the occlusion effect operate more smoothly. The drawback to this is that the response in occlusion detection / rendering is slowed and does not directly react to fast movement. Generally, there is a trade-off between resolution and temporal smoothing in order to make detection and rendering as fast as possible without generating audible steps.
[0044] 1.2 Two-stage Occlusion Detection
[0045] To achieve high resolution of the most important aspects of occlusion detection of certain audio sources while minimizing the number of raycasts required, the process can be carried out in two stages as follows.
[0046] First, so-called cropping occlusion is detected. This is an occlusion that completely obscures at least one edge of the extent. In the case of cropping occlusion (i.e., when the entire edge of the extent is completely occluded), a modified extent is calculated where the completely occluded portion is discarded.
[0047] Second, using the modified extent with the completely occluded portion discarded, a set of raycasts (e.g., an evenly distributed set) is sent to measure the amount of occlusion and an occlusion filter representing different sub-areas of the modified extent is calculated.
[0048] The first stage focuses on determining whether the edges of the extent are completely occluded, and if so, determining the corresponding edges of the modified extent (i.e., the occluded edges). Here, an iterative search algorithm can be used to find the occluded edges. Since the first stage only detects complete occlusion, it is not necessary to calculate an occlusion filter for each raycast. This stage will be further described in Section 1.3.
[0049] The second stage operates on the modified extent with the completely occluded parts discarded. However, the "modified" extent may actually be the same as the original extent rather than a modified version of it (this will be further explained below with respect to Figure 5). In any case, the focus of this stage is to identify the occlusions that occur within the so-called modified extent and calculate the occlusion filters corresponding to the occlusions in different sub-areas of the modified extent. This stage will be further described in Section 1.4.
[0050] 1.3 Optimized Detection of Cropping Occlusions
[0051] The detection of cropping occlusions examines the complete occlusion of the edges of the extent of the audio object. Since the auditory system is not sufficiently capable of identifying the exact shape of the audio object, this can be simplified to identifying the width and height of the parts of the extent that are not completely occluded. This can be done by using an iterative search algorithm such as binary search to find the points on the extent that represent the points with the highest and lowest horizontal angles and the highest and lowest vertical angles that are not completely occluded. By using an iterative search algorithm, it is possible to identify the occlusion edges with high precision with the fewest possible raycasts.
[0052] A total gain factor is calculated that describes the overall gain of the modified extent compared to the original extent, along with the modified extent. If part of the original extent is occluded, it should be reflected in the total gain of the rendered audio element. The total gain factor g OV is expressed by the following formula.
Equation
Equation
[0053] The detection of cropping occlusion starts by casting a sparse grid of rays towards the extent to obtain a first rough estimate of the edges. Points representing the horizontal and vertical, highest and lowest angles that are not completely occluded are stored as starting points for an iterative search to find the exact edges.
[0054] Figure 3 shows an example where the extent 304 of the audio element 302 (a rectangular extent in this example) is occluded by the occluders 310 and 311. In one embodiment, occlusion detection is performed in two stages. In the first stage, a cropping occlusion is detected, and in this example, a modified extent 404 (see Figure 4) representing a part of the extent 304 (e.g., the rectangular part of the extent 304) is determined. That is, in this example, since the entire edge 340 (i.e., the left edge) of the extent 304 is completely occluded, the modified extent 404 is smaller than the extent 304. More specifically, in this example, the modified extent 404 has a left edge different from that of the extent 304, but the right edge, the upper edge, and the lower edge are all the same because they are not completely occluded.
[0055] The ray tracing positions are visualized as black dots in Figure 3. In this example, as shown in Figure 3, the left edge 340 of the extent 304 is completely occluded by the object 310. The ray tracing point P1 is the point representing the leftmost point of the unoccluded extent. A binary search between the points P1 and P2 can be used to find the occlusion edge 350. This edge 350 is then used as the left edge of the modified extent (the "cropped extent") (e.g., see Figure 4) used in the next stage. The occluder 311 does not occlude any edges and has no effect on the modified extent.
[0056] In one embodiment, after casting a grid of light rays towards the extent, the non - occlusion point P1 representing the lowest horizontal angle is stored as min_azimuth_point. To find the exact edge of the occlusion, a binary search can be used. In the binary search, the lower end and the upper limit are used. In this case, the lower limit can be initialized to min_azimuth_point or P1. The upper limit is initialized to a point with a lower azimuth angle (left side in this example) that is known to be occluded or known to be on the edge of the extent. In this case, this can be P2. Next, the search is started by evaluating the occlusion at a point between the lower limit and the upper limit. If this mid - point is occluded, the upper limit is set to this mid - point. If this mid - point is not occluded, the lower limit is set to this mid - point. Then, this process is repeated until the distance between the lower limit and the upper limit is below a certain threshold or can be repeated N times. N is a predefined set value. Then, the mid - point between the upper limit and the lower limit is used to describe the azimuth angle of the left edge of the modified extent.
[0057] FIG. 5 shows an example where none of the edges of the extent 304 are completely occluded. Thus, in this example, the determined modified extent is the same as the extent 304.
[0058] Cropping occlusion detection does not detect the exact shape of the occlusion. Instead, as shown in Figure 4, it only detects the rectangular portion of the extent that is not fully occluded. However, this covers many typical cases, such as when the extent is partially covered by a wall or when viewing / hearing an audio object through a window. The cropping occlusion stage can be thought of as a way to define a frame around the portion of the extent that is not fully occluded. Within this frame, partial or soft occlusion may also occur. Outside the cropped extent, no further checks for occlusion are necessary.
[0059] In the case of occlusions where the shape of the occluding object is more complex or there is soft occlusion, the second stage is used to describe the effect of the occlusion within the modified extent.
[0060] The density of the sparse grid of rays used as the starting point for edge search does not directly affect the accuracy of edge detection. However, the grid of rays needs to be dense enough to detect at least one point of the extent that is not occluded, which can then be used as the starting point for iterative search for the edges of the modified extent. If most of the extent is occluded and only a small portion is not occluded and the sparse grid fails to identify the non-occluded portion of the extent, there may be situations where iterative search cannot be performed appropriately. Sections 1.5 - 1.7 show some examples of ways to optimize the sampling grid so that small non-occluded portions are also detected without making the sample grid very dense.
[0061] 1.4 Optimized Detection of Occlusions within the Cropped Extent
[0062] The second stage of occlusion detection checks for occlusions within the modified (``cropped'') extent. An example of this is shown in Figure 4.
[0063] This is done, as shown in Figure 4, by dividing the modified extent 404 into one or more sub-areas and calculating an occlusion filter for each sub-area of the modified extent. The occlusion filter for a sub-area describes the amount of occlusion in different frequency bands of the sub-area.
[0064] Thus, in one embodiment, the modified extent (i.e., extent 304 or 404) is divided into several sub-areas. The number of sub-areas may vary, for example, depending on the size of the extent and may be adaptive. Usually, the number of sub-areas required is related to how the extent will be rendered later. If the rendering is based on virtual speakers and the number of virtual speakers is small, a virtual speaker setup with limited spatial resolution is used for rendering, so there is little need for many sub-areas. If the extent is very small, no division is required and only one sub-area is defined, which is equal to the whole of the modified extent.
[0065] Examples of typical sub-area divisions with different numbers of sub-areas are shown below.
[0066] One sub-area: No division,
[0067] Two sub-areas: Left, right,
[0068] Three sub-areas: Left, center, right,
[0069] Four sub-areas: Upper left, upper right, lower left, lower right,
[0070] Five sub-areas: Upper left, upper right, center, lower left, lower right,
[0071] Six sub - areas: upper - left, upper - center, upper - right, lower - left, lower - center, lower - right.
[0072] The sub - areas do not necessarily have to be of the same size. However, if they are of the same size, the rendering is simplified because the energy contribution of each sub - area is the same.
[0073] For each sub - area, a set of rays is cast to obtain an estimate of how much occlusion exists for this particular part of the modified extent. For each ray - cast, the occlusion filter is formed from the acoustic transmission parameters of any material through which the ray passes. The filter can be represented as a list of gain factors for different frequency bands. If a ray passes through two or more occluders, the occlusion filter is calculated by multiplying the gains of the different materials in each frequency band. If a ray is completely occluded, the occlusion filter can be set to 0.0 for all frequencies. If a ray does not pass through any occluding object, the occlusion filter can be counted as having a gain of 1.0 for all frequencies. If a ray does not hit the extent of an audio object, it can be treated as a completely occluded ray or simply discarded.
[0074] For each sub - area, the occlusion filters of each ray - cast are accumulated to form one occlusion filter representing the occlusion within that sub - area. Next, the cumulative gain per frequency band for that sub - area can be calculated, for example, using the following.
Equation
Number
[0075] In a specific example, assume that two light rays are cast towards the sub - area of the extent, the first light ray passes through a thin occluding object made of a first material (e.g., cotton), and the second light ray passes through a thick occluding object made of a second material (e.g., brick). That is, the points within the extent through which the first light ray passes are occluded by the thin occluding object, and the points within the extent through which the second light ray passes are occluded by the thick occluding object. Also, assume that each material is associated with a different filter (i.e., a set of frequency range and gain coefficient for each frequency range) as shown in the following table.
Table 1
[0076] Here, G SA,F1 = sqrt((g11 + g21) / 2); G SA,F2 = sqrt((g12 + g22) / 2); G SA,F3 = sqrt((g13 + g23) / 2). That is, the sub - area is associated with three different cumulative gain values (G SA,F1 , G SA,F2 , G SA,F3 ) one for each frequency (or frequency range).
[0077] The distribution pattern of light rays across each sub - area should preferably be uniform. The simplest form of a uniform distribution pattern would be a regular grid. However, a regular grid pattern means that many sample points are created at the same horizontal angle and many sample points are created at the same vertical angle. Since many occlusion situations involve occluders with straight - vertical or horizontal edges such as walls, doorways, windows, etc., this can increase the problem of step - wise behavior. This problem is shown in FIGS. 6A and 6B.
[0078] FIGS. 6A and 6B show an example of occlusion detection using 24 uniform grid light rays. The extent 604 is shown as viewed from the listening position, and the occluder 610 is moving from left to right to cover more of the extent. In FIG. 6A, the occluder 610 is blocking 12 light rays (the light rays are shown as black dots). In FIG. 6B, the occluder has moved further to the right, where it is blocking 15 light rays. As the occluder moves further, the amount of occlusion changes in discrete steps, and an instantaneous change in the audio level can be heard.
[0079] Instead of using a regular grid as shown in FIGS. 6A and 6B, some form of random sampling distribution such as completely random sampling, clustered random sampling, or regular sampling with random offsets may be used. Generally, a good distribution pattern is one in which the sample points do not repeat the same vertical or horizontal angle. Such a pattern can be constructed from a regular grid where increasing offsets are added to the vertical positions of the samples within each horizontal row and increasing offsets are added to the horizontal positions of the samples within each vertical column. Such a distorted grid pattern distributes the sample points so that the horizontal and vertical positions of all sample points are distributed as evenly as possible. FIGS. 7A and 7B show an example of a grid with increasing offsets in the horizontal position of the sample points. As can be seen from the figures, when the occluder moves, only one extra ray is occluded. This means that, as shown in FIGS. 6A and 6B, the detection resolution is improved by a factor of three compared to using a uniform grid with the same number of sample points.
[0080] 1.5 Time-Varying Raycast Sampling Grid
[0081] The number of raycasts used for the two stages of detection can be adaptive. For example, the number of raycasts can be adapted so that the resolution is kept constant regardless of the size of the extent, so that fewer rays are used when there are a large number of other active processes in the renderer, or so that the number of rays depends on the current renderer load.
[0082] Another way to vary the number of rays is to utilize past detection results, such that the resolution is increased for a period after a certain occlusion has been detected. In this way, a sparser set of rays can be used to detect whether an occlusion exists, and whenever an occlusion is detected, the resolution for the next update of the occlusion state can be increased. Then, as long as some occlusion is still detected, the increased resolution can then be maintained over an additional period.
[0083] Yet another way to vary the raycast sampling grid over time is to use a series of grids that complement each other such that the spatial resolution can be increased by using the cumulative results of two or more consecutive grids. This means that the results are averaged over a longer time frame, and thus the response for occlusion detection is slower, similar to when temporal smoothing is applied. One way to overcome this is to use only the sequential grid when no occlusion has been detected previously for a period, and when an occlusion is detected, turn off the sequential grid and instead use one sampling grid with high resolution. Such consecutive grids may be pre-computed or may be generated on-the-fly by adding an offset to one predefined grid.
[0084] 1.6 Reuse of Raytracing Information from Stage 1 in Stage 2
[0085] It is possible to reuse raytracing information from stage 1 in stage 2 if the occlusion filter for each ray cast in stage 1 is evaluated, stored, and can be included in the calculation of the cumulative occlusion filter for each sub-area.
[0086] 1.7 Reuse of Occlusion Information from Past Occlusion Detection Updates
[0087] Scene updates are often smooth, so occlusion changes are usually gradual. In many cases, information from the previous occlusion detection can be used as a good starting point for the next update. One way to utilize the previous detection is to add points from within the corrected extent of the past update when performing the first-stage detection of the cropping occlusion. For example, the center point of the corrected extent of the previous occlusion detection update can be added as an additional sample point in the first stage. For example, combining a sparse continuous grid of sample points with additional sample points from the previous corrected extent can provide a very efficient way to detect the cropping occlusion.
[0088] FIG. 8 is a flowchart showing a process 800 for rendering an audio element associated with a first extent according to an embodiment. The first extent may be the actual extent of the audio element as seen from the listener position, or a projection of the audio element as seen from the listener position, where the projection may be, for example, a projection of the extent of the audio element onto a sphere around the listener, or a projection of the extent of the audio element onto a plane between the audio element and the listener. International Patent Application Publication No. WO2021180820 describes a technique for projecting an audio object in a complex shape. For example, this publication describes a method for representing an audio object with respect to a listener's listening position in an augmented reality scene, the method including obtaining first metadata describing a first three-dimensional (3D) shape associated with the audio object, and transforming the obtained first metadata to generate transformed metadata describing a two-dimensional (2D) plane or a one-dimensional (1D) line, the 2D plane or 1D line representing at least a portion of the audio object, transforming the first metadata obtained to generate the transformed metadata, determining a set of description points, and using the description points to determine the 2D plane or 1D line, the 2D plane or 1D line passing through an anchor point. The anchor point may be i) a point on the surface of the 3D shape closest to the listener's listening position in the augmented reality scene, ii) a spatial average of points on or within the 3D shape, or iii) the centroid of the portion of the shape visible to the listener, and the set of description points further includes a first point on the first 3D shape representing a first edge of the first 3D shape with respect to the listener's listening position, and a second point on the first 3D shape representing a second edge of the first 3D shape with respect to the listener's listening position.
[0089] Process 800 starts at step s802. Step s802 includes determining a first point in the first extent that is not fully occluded. This step corresponds to the step of the first stage of the two-stage process described above, and the first point may correspond to point P1 in FIG. 3.
[0090] Step s804 includes determining a second extent (referred to above as the modified extent) for the audio element, where determining includes determining a first edge of the second extent using the first point. This step is also a step of the first stage described above. The first edge of the second extent is edge 350 when edge 340 is fully occluded as shown in FIG. 3, and may be edge 340 when edge 340 is not fully occluded as shown in FIG. 5.
[0091] After determining the second extent in step s806, step s806 includes dividing the second extent into a set of one or more sub-areas including the first sub-area.
[0092] Step s808 includes determining a first gain value for the first sample point of the first sub-area (e.g., for the first frequency).
[0093] Step s810 includes rendering the audio element (e.g., generating an output signal using the first gain value) using the first gain value.
[0094] Usage example
[0095] FIG. 9A shows an XR system 900 to which an embodiment can be applied. The XR system 900 includes speakers 904 and 905 (which can be speakers of headphones worn by a listener) and a display device 910 configured to be worn by the listener. As shown in FIG. 9B, the XR system 910 may include an orientation detection unit 901, a position detection unit 902, and a processing device 903 (directly or indirectly) coupled to an audio renderer 951 for generating output audio signals (e.g., a left audio signal 981 for the left speaker and a right audio signal 982 for the right speaker as shown). The audio renderer 951 generates an output signal based on an input audio signal, metadata regarding the XR scene experienced by the listener, and information regarding the position and orientation of the listener. Metadata for the XR scene includes metadata for each object and audio element included in the XR scene, and the metadata for the object may include information about the size of the object and the occlusion gain for the object (e.g., the metadata may be applicable to different frequencies or frequency ranges for each occlusion gain, and may specify a set of occlusion gains (or a set of occlusion factors from which the occlusion gain can be derived)). The audio renderer 951 may be a component of the display device 910 or may be located remotely from the listener (e.g., the renderer 951 may be implemented in the "cloud").
[0096] The orientation detection unit 901 is configured to detect a change in the direction of the listener and provide information regarding the detected change to the processing device 903. In some embodiments, the processing device 903 determines the absolute orientation (with respect to some coordinate systems) taking into account the detected change in the orientation detected by the orientation detection unit 901. Also, for example, different systems for determining orientation and position may exist, such as a system using a lighthouse tracker (rider). In one embodiment, the orientation detection unit 901 can determine the absolute orientation (associated with a certain coordinate system) when a given change in the detected orientation is provided. In this case, the processing device 903 can simply multiplex the absolute orientation data from the orientation detection unit 901 and the position data from the position detection unit 902. In some embodiments, the orientation detection unit 901 may comprise one or more accelerometers and / or one or more gyroscopes.
[0097] FIG. 10 shows an exemplary implementation of an audio renderer 951 for generating sound for an XR scene. The audio renderer 951 includes a controller 1001 and a signal modifier 1002 for modifying an audio signal 961 (e.g., an audio signal of a multi-channel audio element) based on control information 1010 from the controller 1001. The controller 1001 may be configured to receive one or more parameters and trigger the signal modifier 1002 to perform a modification on the audio signal 961 based on the received parameters (e.g., increasing or decreasing the volume level). The received parameters include information 963 regarding the position and / or orientation of the listener (e.g., the direction and distance to the audio element), metadata 962 regarding an audio element (e.g., audio element 602) within the XR scene, and metadata regarding an object blocking the audio element (in some embodiments, the controller 1001 itself generates the metadata 962). Using the metadata and the position / orientation information, the controller 1001 may calculate one or more gain factors (e.g., a total gain factor and a cumulative gain for each sub-area) for the audio elements within the XR scene that are at least partially occluded as described above.
[0098] FIG. 11 shows an exemplary implementation of the signal changer 1002 according to an embodiment. The signal modifier 1002 includes a direction mixer 1104, a filter 1106, and a speaker signal generator 1108.
[0099] The direction mixer 1104 receives an audio input 961, which in this example includes a pair of audio signals 1101 and 1102 associated with the audio element, and generates a set of k virtual speaker signals (VS1, VS2, …, VSk) based on the audio input and control information 1171. In one embodiment, the signal for each virtual speaker can be derived, for example, by appropriate mixing of the signals including the audio input 961. For example, VS1 = α×L + β×R. Here, L is the input audio signal 1101, R is the input audio signal 1102, and α and β are coefficients that depend, for example, on the position of the listener relative to the audio element and the position of the virtual speaker to which VS1 corresponds.
[0100] In an example where the audio element is associated with three virtual speakers (SpL, SpC, R), k is equal to 3 for the audio element, VS1 can correspond to SpL, VS2 can correspond to SpC, and VS3 can correspond to SpR. The control information 1171 used by the direction mixer to generate the virtual speaker signals can include the position of each virtual speaker relative to the audio element. In some embodiments, when the audio element is occluded, the controller 1001 adjusts the position of one or more of the virtual speakers associated with the audio element, provides the position information to the direction mixer 1104, and then the direction mixer generates signals for the virtual speakers (i.e., VS1, VS2, …, VSk) using the updated position information.
[0101] Filter 1106 can adjust one or more gains of the virtual speaker signals based on control information 1172, and the control information may include the above-described cumulative gain factor and total gain factor calculated by controller 1001. That is, for example, when an audio element is at least partially occluded, controller 1001 can control filter 1106 to adjust one or more gains of the virtual speaker signals by providing one or more gain factors to filter 1106. For example, if the entire left portion of the audio element is occluded, controller 1001 provides control information 1172 to filter 1106, and filter 1106 reduces the gain of VS1 by 100% (i.e., sets the gain factor = 0 so that VS1’ = 0). As another example, if only 50% of the left portion of the audio element is occluded and 0% of the central portion is occluded, controller 1001 can provide filter 1106 with control information 1172 that reduces the gain of VS1 by 50% (i.e., VS1’ = 50% VS1) and does not reduce the gain of VS2 at all (i.e., sets the gain factor = 1 so that VS2' = VS2). As another example, if VS1 is a signal related to a specific sub - area, and the cumulative gain of this sub - area is g SA1 and the total gain is g OV assuming it is, in one embodiment, VS1' = VS1×g SA1 ×g OV is.
[0102] Using the virtual speaker signals VS1’, VS2’, …, VSk’, speaker signal generator 1108 generates output signals (e.g., output signal 981 and output signal 982) for driving a speaker (e.g., a headphone speaker or other speaker). In one embodiment where the speaker is a headphone speaker, speaker signal generator 1108 can perform conventional binaural rendering to generate the output signal. In embodiments where the speaker is not a headphone speaker, speaker signal generator 1108 can perform conventional speaker panning to generate the output signal.
[0103] Figure 12 is a block diagram of an audio rendering device 1200 according to some embodiments for executing the methods disclosed herein (for example, the audio renderer 951 may be implemented using the audio rendering device 1200). As shown in Figure 12, the audio rendering device 1200 includes a processing circuit (PC) 1202. The processing circuit (PC) 1202 may include one or more processors (P) 1255 (for example, one or more other processors such as a general-purpose microprocessor and / or an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.). The plurality of processors may be co-located within a single housing or within a single data center, or may be geographically distributed (i.e., the device 1200 may be a distributed computing device). The audio rendering device 1200 may include at least one network interface 1248. The at least one network interface 1248 may include a transmitter (Tx) 1245 and a receiver (Rx) 1247 to enable the device 1200 to transmit and receive data with other nodes connected to a network 110 (for example, an Internet Protocol (IP) network) to which the network interface 1248 is (directly or indirectly) connected. In that case, the network interface 1248 may be wirelessly connected to the network 110. The network interface 1248 is connected to an antenna device and a storage unit (also referred to as a "data storage system") 1208 that may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where the PC 1202 includes a programmable processor, a computer program product (CPP) 1241 may be provided. The CPP 1241 includes a computer-readable medium (CRM) 1242 that stores a computer program (CP) 1243 comprising computer-readable instructions (CRI) 1244. The CRM 1242 may be a non-transitory computer-readable medium such as a magnetic medium (for example, a hard disk), an optical medium, a memory device (for example, a random access memory, a flash memory), etc.In some embodiments, when the CRI 1244 of the computer program 1243 is executed by the PC 1202, the CRI is configured to cause the audio rendering device 1200 to execute the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, the audio rendering device 1200 may be configured to execute the steps described herein without requiring code. That is, for example, the PC 1202 may be composed of only one or more ASICs. Therefore, the features of the embodiments described herein may be implemented by at least either hardware or software.
[0104] Overview of Various Embodiments
[0105] A1. A method 800 for rendering an audio element 302 associated with a first extent 304, the method comprising: determining s802 a first point (e.g., P1) in the first extent that is not fully occluded; determining s804 a second extent 304, 404 for the audio element, wherein determining the second extent includes determining a first edge of the second extent using the first point; after determining the second extent, dividing s806 the second extent into a set of one or more sub-areas including a first sub-area; determining s808 a first gain value for a first sample point of the first sub-area; and rendering s810 the audio element using the first gain value.
[0106] A2. The method according to embodiment A1, wherein the step of determining the first edge of the second extent using the first point includes determining whether the first point is on the first edge of the first extent or within a threshold distance of the first edge of the first extent.
[0107] A3. The step of determining the first edge of the second extent further includes setting the first edge of the second extent equal to the first edge of the first extent as a result of determining that the first point is on or within a threshold distance of the first edge of the first extent, the method according to embodiment A2.
[0108] A4. The step of determining the first edge of the second extent further includes determining a third point between the first point and a second point that is completely occluded, and determining whether the third point is completely occluded, or defining the first edge of the second extent using the third point, the method according to embodiment A2.
[0109] A5. The step of determining the first edge of the second extent further includes determining whether the third point is completely occluded, and if the third point is determined to be completely occluded, determining a fourth point between the first point and the third point, or if the third point is determined not to be completely occluded, determining a fourth point between the second point and the third point, the method according to embodiment A4.
[0110] A6. The step of determining the first edge of the second extent further includes defining the first edge of the second extent using the fourth point, the method according to embodiment A5.
[0111] A7. The step of determining the second extent further includes determining a fifth point in the first extent that is not completely occluded, and determining a second edge of the second extent using the fifth point, the method according to any one of embodiments A1 to A6.
[0112] A8. The step of determining the second edge of the second extent using the fifth point includes determining whether the fifth point is on the second edge of the first extent or within a threshold distance of the second edge of the first extent, according to the method described in Embodiment A7.
[0113] A9. The step of determining the second edge of the second extent further includes setting the second edge of the second extent equal to the second edge of the first extent as a result of determining that the fifth point is on the second edge of the first extent or within the threshold distance of the second edge of the first extent, according to the method described in Embodiment A8.
[0114] A10. The step of determining the first gain value for the first sample point of the first subarea includes determining whether the virtual line extending from the listening position to the first sample point passes through one or more objects, according to the method described in any one of Embodiments A1 - A9.
[0115] A11. The virtual line passes through at least a first object, and the step of determining the first gain value for the first sample point of the first subarea further includes obtaining first metadata associated with the first object and determining the first gain value using the first metadata, according to the method described in Embodiment A10.
[0116] A12. The virtual line further passes through a second object, and the step of determining the first gain value for the first sample point of the first subarea further includes obtaining second metadata associated with the second object and further determining the first gain value using the second metadata, according to the method described in Embodiment A11.
[0117] A13. The step of rendering the audio element using the first gain value includes calculating a first cumulative gain value for the first sub - area using the first gain value, and rendering the audio element using the first cumulative gain value. The method according to any one of Embodiments A1 to A12.
[0118] A14. The step of rendering the audio element using the first cumulative gain value includes modifying an audio signal related to the first sub - area based on the first cumulative gain value to generate a modified audio signal, and rendering the audio element using the modified audio signal. The method according to Embodiment A13.
[0119] A15. The step of determining the first gain value for the first sample point of the first sub - area includes casting a ray of a diagonal grid towards the first sub - area, and one of the rays intersects the sub - area at the first sample point. The method according to any one of Embodiments A1 to A14.
[0120] A16. Further includes the step of calculating a total gain coefficient g OV and the step of rendering the audio element using the first gain value includes rendering the audio element using the first gain value and g OV . The method according to any one of Embodiments A1 to A15.
[0121] A17. The first extent has a first area A1, the second extent has a second area A2, A2 < A1, and the step of calculating g OV includes the step of calculating A2 / A1. The system according to Embodiment 16.
[0122] A18. The step of calculating g OV further includes the step of obtaining the square root of A2 / A1. The system according to Embodiment 17.
[0123] B1. A computer program including instructions that, when executed by a processing circuit of an audio rendering device, cause the audio rendering device to execute the method according to any one of Embodiments A1 to A18.
[0124] B2. A carrier including the computer program of Embodiment B1, wherein the carrier is one of an electrical signal, an optical signal, a wireless signal, and a computer-readable storage medium.
[0125] C1. An audio rendering device configured to execute the method according to any one of Embodiments A1 to A18.
[0126] C2. The audio rendering device according to Embodiment C1, comprising a memory and a processing circuit coupled to the memory.
[0127] Although various embodiments are described herein, it should be understood that they are presented by way of example only and not by way of limitation. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments. Further, unless otherwise indicated herein or clearly contradicted by context, any combination of the above objectives in all possible variations thereof is encompassed by the present disclosure.
[0128] In addition, the processes described above and illustrated in the drawings are shown as a series of steps, but this is done for illustration only. Accordingly, some steps may be added, some steps may be omitted, the order of steps may be rearranged, or some steps may be executed in parallel.
[0129] References
[0130] [1] MPEG-H 3D Audio, Clause 8.4.4.7: “Spreading”
[0131] [2] MPEG-H 3D Audio, Clause 18.1: “Element Metadata Preprocessing”
[0132] [3] MPEG-H 3D Audio, Clause 18.11: “Diffuseness Rendering”
[0133] [4] EBU ADM Renderer Tech 3388, Clause 7.3.6: “Divergence”
[0134] [5] EBU ADM Renderer Tech 3388, Clause 7.4: “Decorrelation Filters”
[0135] [6] EBU ADM Renderer Tech 3388, Clause 7.3.7: “Extent Panner”
[0136] [7] Schissler, C., et. al., “Efficient HRTF-based Spatial Audio for Area and Volumetric Sources,” IEEE Transactions on Visualization and Computer Graphics, Vol. 22, No. 4, pp. 1356-1366, April 2016.
[0137] [8] Patent Publication WO2020144062, “Efficient spatially-heterogeneous audio elements for Virtual Reality.”
[0138] [9] Patent Publication WO2022218986, “RENDERING OF OCCLUDED AUDIO ELEMENTS,” (Application No. PCT / EP2022 / 059762)
[0139]
[10] Patent Publication WO2021180820, “RENDERING OF AUDIO OBJECTS WITH A COMPLEX SHAPE”
Claims
1. A method (800) for rendering an audio element (302) associated with a first extent (304), comprising: determining (s802) a first point (P1) that is not fully occluded in the first extent; determining (s804) a second extent (304, 404) for the audio element, wherein determining the second extent includes determining a first edge (350) of the second extent using the first point; after determining the second extent, dividing (s806) the second extent into a set of one or more sub-areas including a first sub-area; determining (s808) a first gain value for a first sample point of the first sub-area; and rendering (s810) the audio element using the first gain value. A method as described above.
2. The method according to claim 1, wherein determining the first edge (350) of the second extent using the first point includes determining whether the first point is on the first edge of the first extent or within a threshold distance of the first edge of the first extent.
3. The method according to claim 2, wherein determining the first edge of the second extent further includes setting the first edge of the second extent equal to the first edge of the first extent as a result of determining that the first point is on the first edge of the first extent or within the threshold distance of the first edge of the first extent.
4. The method according to claim 2, wherein determining the first edge of the second extent further includes: determining a third point between the first point and a second point (P2) that is fully occluded; and determining whether the third point is fully occluded or defining the first edge of the second extent using the third point. The method according to claim 2, further including the above steps.
5. The method according to claim 2, wherein determining the first edge of the second extent further includes: determining whether the third point is fully occluded; When it is determined that the third point is completely occluded, determining a fourth point between the first point and the third point, or, When it is determined that the third point is not completely occluded, determining a fourth point between the second point and the third point, and The method according to claim 4, further comprising.
6. The step of determining the first edge of the second extent is The method according to claim 5, further comprising defining the first edge of the second extent using the fourth point.
7. The step of determining the second extent is Determining a fifth point in the first extent that is not completely occluded, and Determining a second edge of the second extent using the fifth point, and The method according to any one of claims 1 to 6, further comprising.
8. The step of determining the second edge of the second extent using the fifth point includes determining whether the fifth point is on the second edge of the first extent or within a threshold distance of the second edge of the first extent. The method according to claim 7.
9. The step of determining the second edge of the second extent further includes setting the second edge of the second extent to be equal to the second edge of the first extent as a result of determining that the fifth point is on the second edge of the first extent or within the threshold distance of the second edge of the first extent. The method according to claim 8.
10. The step of determining the first gain value for the first sample point of the first sub - area is The method according to any one of claims 1 to 9, including determining whether the virtual line extending from the listening position to the first sample point passes through one or more objects.
11. The virtual line passes through at least a first object, and the step of determining the first gain value for the first sample point of the first sub - area is Obtaining first metadata associated with the first object, and a step of determining the first gain value using the first metadata; The method according to claim 10, further comprising: **Claim 12** The virtual line further passes through a second object, and the step of determining the first gain value for the first sample point of the first sub-area includes: a step of obtaining second metadata associated with the second object; a step of further determining the first gain value using the second metadata; The method according to claim 11, further comprising: **Claim 13** The step of rendering the audio element using the first gain value includes calculating a first cumulative gain value for the first sub-area using the first gain value, and rendering the audio element using the first cumulative gain value. The method according to any one of claims 1 to 12. **Claim 14** The step of rendering the audio element using the first cumulative gain value includes modifying an audio signal related to the first sub-area based on the first cumulative gain value to generate a modified audio signal, and rendering the audio element using the modified audio signal. The method according to claim 13. **Claim 15** The step of determining the first gain value for the first sample point of the first sub-area includes: casting a ray of an oblique grid towards the first sub-area, one of the rays intersecting the sub-area at the first sample point. The method according to any one of claims 1 to 14. **Claim 16** Total gain factor g OV further comprising a step of calculating, and the step of rendering the audio element using the first gain value comprises the step of rendering the audio element using the first gain value and g OV The method according to any one of claims 1 to 15, comprising a step of rendering the audio element using **Claim 17** The first extent has a first area A1, The second extent has a second area A2, and A2 < A1. g OV The step of calculating includes the step of calculating A2 / A1, the method according to claim 16. **Claim 18** g OV The step of calculating OV further includes the step of obtaining the square root of A2 / A1, the method according to claim 17. **Claim 19** A computer program (1243) including instructions (1244) for causing an audio rendering apparatus (1200) to execute the method according to any one of claims 1 to 18 when executed by a processing circuit (1202) of the audio rendering apparatus. **Claim 20** A carrier including the computer program according to claim 19, wherein the carrier is one of an electrical signal, an optical signal, a wireless signal, and a computer-readable storage medium (1242). **Claim 21** An audio rendering device (1200) for rendering an audio element (302) associated with a first extent (304), the audio rendering device (1200) comprising: Determining (s802) a first point (P1) in the first extent that is not completely occluded; Determining (s804) a second extent (304, 404) for the audio element, wherein determining the second extent includes determining a first edge (350) of the second extent using the first point; After determining the second extent, dividing (s806) the second extent into a set of one or more sub-areas including a first sub-area; Determining (s808) a first gain value for a first sample point of the first sub-area; Rendering (s810) the audio element using the first gain value; An audio rendering device configured to execute a method comprising the steps. **Claim 22** The audio rendering device according to claim 21, further configured to execute the method according to any one of claims 2 to 18.
Citation Information
Patent Citations
Audio device and processing method thereof
JP2022525902A
Rendering occluded audio elements
JP2024514170A
3D audio rendering using volumetric audio rendering and scripted audio level-of-detail
US20200296533A1