Rendering of occluded audio element

The method uses virtual loudspeakers to adjust gain and position based on occlusion calculations, addressing the challenge of rendering partially occluded audio elements in extended reality, ensuring high-quality spatial audio rendering.

JP2025157249APending Publication Date: 2025-10-15TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025106042
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-04-14
Filing Date
2025-06-24
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Existing spatial audio rendering techniques struggle to effectively handle occluded audio elements in extended reality scenes, particularly for audio elements with a range or mixed extended audio elements, leading to incomplete or distorted sound due to partial occlusion and step-wise changes in occlusion as objects and listeners move.

Method used

A method using a set of virtual loudspeakers to modify signals based on occlusion calculations, adjusting gain factors and positions to accurately render partially occluded audio elements, preserving spatial information and avoiding step-wise distortions.

Benefits of technology

The method ensures high-quality spatial audio rendering by dynamically adjusting signal gain and speaker positions to account for partial occlusions, maintaining audio integrity and reducing abrupt changes in sound output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025157249000001_ABST
    Figure 2025157249000001_ABST
Patent Text Reader

Abstract

To provide a method for rendering an at least partially occluded audio element.SOLUTION: In a method for rendering an audio element 602, an audio element is represented using a set of two or more virtual loudspeakers SpL, SpC, SpR. To reflect the occlusion of the audio element 602 by an object 604, the virtual loudspeaker SpR is moved to the edge where the occlusion is happening, and the virtual loudspeaker SpC is moved to the center of the part that is not occluded. To reflect the occlusion of the audio element 602 by an object 614, the virtual loudspeaker SpR is moved upward and the loudspeaker SpC is also moved upward.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] SUMMARY OF THE INVENTION Embodiments related to rendering occluded audio elements are disclosed. [Background technology]

[0002] Spatial audio rendering is a process used to present audio in an extended reality (XR) scene (e.g., a virtual reality (VR), augmented reality (AR), or mixed reality (MR) scene) to give the listener the impression that the sound is coming from a physical source within the scene at a certain location and having a certain size and shape (i.e., extent). Presentation can be through headphone speakers or other speakers. When presentation is through headphone speakers, the process used is called binaural rendering and uses spatial cues of human spatial hearing that make it possible to determine which direction a sound is coming from. The cues involve inter-aural time delay (ITD), inter-aural level difference (ILD), and / or spectral differences.

[0003] The most common form of spatial audio rendering is based on the point source concept, where each sound source is defined to emanate from one specific point. Because each sound source is defined to emanate from one specific point, the sound sources do not have a size or shape. Different methods have been developed to render sound sources that have a range (size and shape).

[0004] One such known method is to create multiple copies of a mono audio element at positions around the audio element. This configuration results in the perception of a spatially uniform object with a certain size. This concept is used, for example, in the "object spread" and "object divergence" features of the MPEG-H 3D audio standard (see References [1] and [2]) and in the "object divergence" feature of the EBU Audio Specification Model (ADM) standard (see Reference [4]). This idea using a mono audio source is further developed as described in Reference [7], where the area-volume geometry of the sound object is projected onto a sphere around the listener, and the sound is rendered to the listener using a pair of head-related (HR) filters that is evaluated as the integral of all HR filters covering the geometric projection of the object on the sphere. For a spherical volume source, this integral has an analytical solution. However, for arbitrary area-volume source geometries, the integral is evaluated by sampling the projected source surface on the sphere using so-called Monte Carlo ray sampling.

[0005] Another rendering method renders a spatially diffuse component in addition to the mono audio signal, which results in the perception of somewhat diffuse objects that, in contrast to the original mono audio element, do not have a distinct pinpoint location. This concept is used, for example, in the "object diffuseness" feature of the MPEG-H 3D audio standard (see reference [3]) and the "object diffuseness" feature of the EBU ADM (see reference [5]).

[0006] Combinations of the two methods above are also known, for example the "object extent" feature of the EBU ADM, which combines the creation of multiple copies of a mono audio element with the addition of a diffuse component (see reference [6]).

[0007] In many cases, the actual shape of an audio element can be described sufficiently well using primitive shapes (e.g., a sphere or a box). However, sometimes the actual shape is more complex and needs to be described in a more detailed form (e.g., a mesh structure or a parametric description format).

[0008] In the case of a mixed audio element, as described in reference [8], the audio element includes at least two audio channels (ie, audio signals) to describe the spatial variation over the range of the audio element.

[0009] In some XR scenes, there may be objects that occlude at least a portion of an audio element in the XR scene, in such a scenario the audio element is said to be at least partially occluded.

[0010] That is, occlusion occurs when an audio element is completely or partially hidden behind some object such that, from the listener's perspective at a given listening position, no or little direct sound from the occluded portion of the audio element reaches the listener. Depending on the material of the occluding object, the occlusion effect can be either complete occlusion (e.g., when the occluding object is a thick wall) or soft occlusion, where part of the audio energy from the audio element passes through the occluding object (e.g., when the occluding object is made from a thin fabric such as a curtain). Summary of the Invention

[0011] Currently, several challenges exist. For example, while available occlusion rendering techniques address point sources, where the occurrence of occlusion can be easily detected using ray tracing between the listener position and the point source's position, for audio elements with a range, the situation is more complicated because an occluding object may occlude only a portion of the extended audio element. Therefore, more sophisticated occlusion detection techniques (e.g., occlusion detection techniques that determine which portions of an extended audio element are occluded) are needed. For mixed extended audio elements (i.e., audio elements with a range that have non-uniform spatial audio information distributed across the range of the audio element (e.g., an extended audio element represented by a stereo signal)), the situation is even more complicated, as rendering of this type of partially occluded object should take into account what the expected consequences of partial occlusion would be for the spatial audio information reaching the listener. A special version of the latter problem appears when mixed extended audio elements are rendered by a discrete number of virtual loudspeakers. If traditional occlusion were used and operated on individual virtual loudspeakers, and one or more of the virtual loudspeakers were occluded, this would mean, for example, that if two virtual loudspeakers (e.g., a left (L) speaker and a right (R) speaker) were used, essentially all spatial information would be lost whenever either the L virtual loudspeaker or the R virtual loudspeaker was occluded.More generally, for extended objects (and therefore also including non-mixed audio elements, e.g., uniform or diffuse extended audio elements) that are rendered using a discrete number of virtual loudspeakers, there is an issue with the amount of occlusion changing in a step-wise manner as the audio elements, the occluding objects, and / or the listener move relative to each other.

[0012] Thus, in one aspect, a method for rendering an at least partially occluded audio element is provided, where the audio element is represented using a set of two or more virtual loudspeakers, the set including a first virtual loudspeaker. In one embodiment, the method includes modifying a first virtual loudspeaker signal for the first virtual loudspeaker, thereby producing a first modified virtual loudspeaker signal. The method also includes using the first modified virtual loudspeaker signal to render the audio element (e.g., generating an output signal using the first modified virtual loudspeaker signal). In another embodiment, the method includes moving the first virtual loudspeaker from an initial position to a new position. The method also includes generating a first virtual loudspeaker signal for the first virtual loudspeaker based on the new position of the first virtual loudspeaker. The method also includes using the first virtual loudspeaker signal to render the audio element.

[0013] In another aspect, a computer program is provided that includes instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform any of the methods described above. In one embodiment, a carrier is provided that includes the computer program, the carrier being one of an electronic signal, an optical signal, a wireless signal, and a computer-readable storage medium. In another aspect, a rendering device is provided that is configured to perform any of the methods described above. The rendering device may include a memory and a processing circuit coupled to the memory.

[0014] An advantage of the embodiments disclosed herein is that the rendering of at least partially occluded audio elements is done in a manner that preserves the quality of the spatial information of the audio elements.

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various embodiments. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 shows two point sources (S1 and S2) and an occluding object (O). [Figure 2] FIG. 2 shows an audio element having a region that is partially occluded by an occluding object (O). [Figure 3] FIG. 1 illustrates the use of many point sources to represent an audio element. [Figure 4A] 1 is a flowchart illustrating a process according to one embodiment. [Figure 4B] 1 is a flowchart illustrating a process according to one embodiment. [Figure 5] 1 is a flowchart illustrating a process according to one embodiment. [Figure 6] 1A-C are diagrams illustrating various exemplary embodiments. [Figure 7]1A-C are diagrams illustrating various exemplary embodiments. [Figure 8] FIG. 1 illustrates an exemplary embodiment. [Figure 9] 1A-B are diagrams illustrating various exemplary embodiments. [Figure 10] FIG. 1 illustrates an exemplary embodiment. [Figure 11] FIG. 1 illustrates an exemplary embodiment. [Figure 12] 1A-B illustrate a system according to some embodiments. [Figure 13] FIG. 1 illustrates a system, according to some embodiments. [Figure 14] FIG. 2 illustrates a signal modifier according to one embodiment. [Figure 15] FIG. 1 is a block diagram of an apparatus, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0017] The occurrence of occlusion can be detected using ray tracing methods, in which a direct path between the listener position and the audio element's position is searched for any occluding object. Figure 1 shows an example of two point sources (S1 and S2), one of which is occluded by an object (O) (called the "occluding object") and the other is not. In this case, the occluded audio element should be muted in a manner that corresponds to the acoustic properties of the occluding object's material. If the occluding object is a thick wall, the rendering of the direct sound from the occluded audio element should be almost completely muted. For an audio element (E) with a range, as shown in Figure 2, the audio element (E) may be only partially occluded. This means that the rendering of the audio element needs to be modified in a way that reflects which parts of the range are occluded and which parts are not.

[0018] One strategy for solving the occlusion problem for an audio element with a range (see audio element 302 in FIG. 3 ) is to represent the audio element 302 with a number of point sources spread across that range (as shown in FIG. 3 ) and calculate the occlusion effect for each point source individually using one of the known methods for point sources. However, this strategy is highly inefficient due to the large number of point sources that must be used to obtain a sufficiently good resolution of the occlusion effect. Also, even if many point sources are used so that the resolution for the static case is sufficiently good, there will still be stepwise behavior in dynamic scenes where the occlusion effect changes in discrete steps when individual point sources are either occluded or unoccluded. Another drawback of using many point sources to represent a heterogeneous (multi-channel) audio element is that it is non-trivial how to upmix from a few audio channels to a large number of point sources without introducing spatial and / or spectral distortions in the resulting listener signal (due to the fact that adjacent point sources will be highly correlated).

[0019] Accordingly, the present disclosure describes additional embodiments that do not suffer from these drawbacks described in the previous paragraph. In one aspect, a method according to one embodiment includes the following steps.

[0020] 1. Detecting when an audio element seen from a listener position is occluded (e.g., fully occluded or partially occluded) by an occluding object.

[0021] 2. Calculating the amount of occlusion in a set of sub-areas (a.k.a. portions) of the projection of the audio element as seen from the listener position, where the projection can be, for example, a projection of the extent of the audio element onto a sphere around the listener, or a projection of the extent of the audio element onto a plane between the audio element and the listener. International Patent Application Publication No. WO2021180820 describes techniques for projecting audio objects with complex shapes. For example, this publication describes a method for representing an audio object relative to a listening position of a listener in an extended reality scene, the method including obtaining first metadata describing a first three-dimensional (3D) shape associated with the audio object, and transforming the obtained first metadata to produce transformed metadata describing a two-dimensional (2D) plane or a one-dimensional (1D) line, the 2D plane or 1D line representing at least a portion of the audio object, and transforming the obtained first metadata to produce the transformed metadata includes determining a set of description points including anchor points, and determining a 2D plane or 1D line using the description points, where the 2D plane or 1D line passes through the anchor points. The anchor point may be i) a point on the surface of the 3D shape closest to the listener's listening position in the extended reality scene, ii) a spatial average of points on or within the 3D shape, or iii) a centroid of a portion of the shape that is visible to the listener, and the set of description points further includes a first point on the first 3D shape that represents a first edge of the first 3D shape relative to the listener's listening position, and a second point on the first 3D shape that represents a second edge of the first 3D shape relative to the listener's listening position.

[0022] 3. Calculating a gain factor for each virtual loudspeaker signal used in rendering the audio element based on the amount of occlusion in different parts of the area (e.g., the gain factor for the virtual loudspeaker signal for the part of the audio element not affected by the occluding object is set to 1, while the gain factor for the other virtual loudspeaker signal for the part affected by the occluding object is set to a value less than 1); and

[0023] 4. Modifying the positions of zero or more of the virtual loudspeakers to represent the unoccluded portions of the range.

[0024] A. Calculating the amount of occlusion in each sub-area.

[0025] Given knowledge of what sub-areas of an audio element (or more precisely, the projection of the audio element) are at least partially occluded, and given knowledge about the occluding object (e.g., a parameter indicating the amount of audio energy from the audio element that passes through the occluding object), the amount of occlusion can be calculated for each said sub-area. In scenarios where the parameter indicates that no energy from the audio element passes through the occluding object, the amount of occlusion can be calculated as the percentage of the sub-area that is occluded from the listening position.

[0026] The subareas of the projection of an audio element may be defined in many different ways. In one embodiment, there are as many subareas as there are virtual loudspeakers used for rendering, with each subarea corresponding to one virtual loudspeaker. In another embodiment, the subareas are defined independently of the number and / or position of the virtual loudspeakers used for rendering. The subareas may be equal in size. The subareas may be directly adjacent to each other. The subareas together may completely fill the surface area of ​​the projected range of the audio element, i.e., the total size of the projected range is equal to the sum of the surface areas of all the subareas.

[0027] B. Calculating the gain factor:

[0028] For each sub-area, a gain factor may be calculated depending on the amount of occlusion for that area. For example, in some scenarios where the occluding object is a thick brick wall or the like, a sub-area that is fully occluded by the occluding brick wall (the amount is 100%) may be completely muted, and therefore the gain factor should be set to 0.0. For sub-areas where the occlusion amount is 0, the gain factor should be set to 1.0. For other amounts of occlusion, the gain factor should be somewhere between 0.0 and 1.0, although the exact behavior may depend on the spatial nature of the audio element. In one embodiment, the gain factor is It is calculated as g=(1.0-0.01*O), where O is the amount of occlusion in percent.

[0029] In one embodiment, O for a given subarea is a function of a frequency-dependent occlusion factor (OF) and a value P, where P is the proportion of the subarea covered by occluding objects (i.e., the proportion of the subarea that cannot be seen by the listener due to the fact that occluding objects are located between the listener and the subarea). For example, O=OF*P, where OF=Of1 for frequencies below f1, OF=Of2 for frequencies between f1 and f2, and OF=Of3 for frequencies above f2. That is, for a given frequency, different types of occluding objects may have different occlusion factors. For example, for a first frequency, a brick wall may have an occlusion factor of 1, while a sheer cotton curtain may have an occlusion factor of 0.2, and for a second frequency, a brick wall may have an occlusion factor of 0.8, while a sheer cotton curtain may have an occlusion factor of 0.1.

[0030] In another embodiment, the gain factor is calculated using the assumption that the audio element is mostly diffuse in spatial information, and that a 50% amount of occlusion would give a -3 dB reduction in audio energy from that sub-area. g=cos(0.01*O*π / 2) or g=sqrt(1-0.01*O) It can be calculated as:

[0031] Embodiments are not limited to the above examples, as other gain functions for calculating the gain of a sub-area are possible. As illustrated by the two embodiments described above, the effect of occlusion can be gradual when an audio element is partially occluded, and therefore the signal from a virtual loudspeaker is not necessarily completely muted whenever the virtual loudspeaker is occluded for a listener. This prevents, for example, in the case of stereo rendering with two virtual loudspeakers, receiving no sound at all from the left half of an audio element whenever the left virtual loudspeaker is occluded. Furthermore, this prevents undesirable "stepwise" occlusion effects when the occluding object, audio element, and / or listener are moving relative to each other.

[0032] C. Modifying the position of the virtual loudspeakers representing the audio elements

[0033] When a portion of an audio element is occluded, the position of the virtual loudspeaker representing the audio element may be moved so that the virtual loudspeaker better represents the unoccluded portion. If one of the edges of the extent of an audio element is occluded, the virtual loudspeaker(s) representing this edge should be moved to the edge where the occlusion occurs, as shown in Figures 8 and 9B.

[0034] If the occluding object covers the middle of an audio element, the speaker positions remain the same and the effect of the occlusion is represented only by the gain factor of the signal going to each virtual loudspeaker, as shown in Figure 10.

[0035] If an audio element is represented only by a virtual loudspeaker in the horizontal plane, an occlusion covering either the lower or upper part can be rendered by changing the vertical position of the virtual loudspeaker so that the vertical position of the virtual loudspeaker corresponds to the middle of the unoccluded part of the range.

[0036] In another embodiment, the vertical position of each virtual loudspeaker is controlled by the ratio of the amount of occlusion in the upper and lower sub-areas. An example of how this position can be calculated is: P y =O U / O L *P YT +(1-O U / O L )*P YB where P Y is the vertical coordinate of the loudspeaker, and O U and O L are the occlusion amounts in the upper and lower parts of the range. P YT and P YB are the vertical coordinates of the upper and lower edges of the range.

[0037] 4A is a flowchart illustrating a process 400 for rendering an at least partially occluded audio element represented using a set of two or more virtual loudspeakers, the set including a first virtual loudspeaker, according to one embodiment. Process 400 may begin at step s402. Step s402 includes modifying a first virtual loudspeaker signal for the first virtual loudspeaker, thereby producing a first modified virtual loudspeaker signal. Step s404 includes using the first modified virtual loudspeaker signal to render the audio element (e.g., generate an output signal using the first modified virtual loudspeaker signal).

[0038] In some embodiments, the process further includes obtaining information indicating that the audio element is at least partially occluded, and the modifying is performed as a result of obtaining the information.

[0039] In some embodiments, the process further comprises detecting that the audio element is at least partially occluded, and the modifying is performed as a result of the detection.

[0040] In some embodiments, modifying the first virtual loudspeaker signal includes adjusting a gain of the first virtual loudspeaker signal.

[0041] In some embodiments, the process further includes moving the first virtual loudspeaker from an initial position (e.g., a default position) to a new position, and then generating the first virtual loudspeaker signal using information indicative of the new position.

[0042] In some embodiments, the process further includes determining an occlusion amount (O) associated with the first virtual loudspeaker, and modifying the first virtual loudspeaker signal for the first virtual loudspeaker includes modifying the first virtual loudspeaker signal based on O. In some embodiments, modifying the first virtual loudspeaker signal based on O includes modifying the first virtual loudspeaker signal VS1 such that the modified loudspeaker signal is equal to (g*VS1), where g is a gain factor calculated using O and VS1 is the first virtual loudspeaker signal. In one embodiment, g=1−0.01*O or g=sqrt(1−0.01*O). In one embodiment, determining O includes obtaining a specific occlusion factor (Of) for the occluding object and determining a percentage of a sub-area of ​​the projection of the audio element covered by the occluding object, the first virtual loudspeaker associated with the sub-area.

[0043] 4B is a flowchart illustrating a process 450 for rendering an at least partially occluded audio element represented using a set of two or more virtual loudspeakers, the set including a first virtual loudspeaker, according to one embodiment. Process 450 may begin at step s452. Step s452 includes moving the first virtual loudspeaker from an initial position to a new position. Step s454 includes generating a first virtual loudspeaker signal for the first virtual loudspeaker based on the new position of the first virtual loudspeaker. Step s456 includes using the first virtual loudspeaker signal to render the audio element. In some embodiments, the process further includes obtaining information indicating that the audio element is at least partially occluded, the moving being performed as a result of obtaining the information. In some embodiments, the process further includes detecting that the audio element is at least partially occluded, the moving being performed as a result of the detection.

[0044] FIG. 5 is a flowchart illustrating a process 500 for rendering occluded audio elements according to one embodiment. Process 500 may begin at step s502. Step s502 includes obtaining metadata for the audio element and metadata for objects occluding the audio element (the metadata for the occluding objects may include information specifying occlusion factors for the objects at different frequencies). Step s504 includes determining an amount of occlusion for each subarea of ​​the audio element. Step s506 includes calculating a gain factor for each virtual loudspeaker signal based on the amount of occlusion. Step s508 includes determining, for each virtual loudspeaker, whether the virtual loudspeaker should be placed in a new location and placing the virtual loudspeaker in the new location. Step s510 includes generating virtual loudspeaker signals based on the locations of the virtual speakers. Step s512 includes adjusting the gain of one or more of the virtual loudspeaker signals based on the gain factor.

[0045] 6A is an example in which audio element 602 (or, more precisely, the projection of audio element 602 as seen from the listener position) is logically divided into six parts (also known as six sub-areas), with parts 1 and 4 representing the left area of ​​audio element 602, parts 3 and 6 representing the right area, and parts 2 and 5 representing the center. Also, parts 1, 2, and 3 together represent the upper area of ​​the audio element, and parts 4, 5, and 6 represent the lower area of ​​the audio element.

[0046] FIG. 6B illustrates an exemplary scenario in which an audio element 602 seen by a listener is partially occluded by an occluding object 604, which, in this and other examples, has an occlusion factor of 1. By calculating how much of each portion of the audio element 602 is covered by the occluding object 604, the relative gain balance of the left, center, and right portions can be calculated. Similarly, the relative gain balance of the upper area compared to the lower area can be calculated. In the example shown in FIG. 6B, the right area of ​​the audio element should be completely muted since it is completely covered by the object 604, the center area should have a slightly lower gain, and the left area is unaffected. There is no difference in occlusion of the upper area compared to the lower area.

[0047] 6C shows an example scenario in which an audio element 602 is partially occluded by an occluding object 614. In this example, the center and right areas should be partially muted. The lower part should be muted more than the upper part.

[0048] Figure 7A shows an example in which an audio element 602 is represented by three virtual loudspeakers SpL, SpC, and SpR. Figure 7B shows how the positions of the virtual loudspeakers are modified to reflect the occlusion of the audio element 602 by an object 604. Speaker SpR, which represents the right edge of the range, is moved to the edge where the occlusion occurs. Speaker SpC is moved to the center of the unoccluded part. Figure 7C shows how the positions of the virtual loudspeakers are modified to reflect the occlusion of the audio element 602 by an object 614. Speaker SpR, which represents the right edge of the range, is moved up to a new position, and speaker SpC is also moved up.

[0049] 8 shows an example where the right sub-area of ​​audio element 602 is partially occluded. In this case, the virtual loudspeaker representing the right edge is moved so that it is aligned with the edge where the occlusion occurs. The center speaker can be moved to a position that represents the center of the unoccluded portion of the audio element.

[0050] 9 shows an example of an audio element 902 represented by six virtual loudspeakers, where the lower portion of the audio element is occluded. In this case, the virtual loudspeaker representing the bottom edge is moved so that it aligns with the edge where the occlusion occurs.

[0051] 10 shows an example where the middle of audio element 602 is occluded. In this case, the loudspeaker positions are kept the same, since neither the left nor right edges are occluded and do not need to be represented. The occlusion in this case only affects the gain of the signal to each speaker. In this case, the middle speaker is completely muted (i.e., gain factor = 0), and the gain for the left and right speakers has been slightly reduced to reflect that subareas 1, 4, 3, and 6 are also partially occluded.

[0052] 11 shows an example where the center and right areas of audio element 602 are partially occluded. The position of the virtual loudspeaker is modified in elevation to reflect the greater amount of occlusion in these lower parts. Also, the gain of the signal should be reduced to reflect that the center and right areas are partially occluded.

[0053] Exemplary Use Cases

[0054] 12A shows an XR system 1200 to which embodiments may be applied. The XR system 1200 includes speakers 1204 and 1205 (which may be speakers of headphones worn by the listener) and a display device 1210 configured to be worn by the listener. As shown in FIG. 12B , the XR system 1210 may include an orientation sensing unit 1201, a position sensing unit 1202, and a processing unit 1203 coupled (directly or indirectly) to an audio renderer 1251 for producing output audio signals (e.g., a left audio signal 1281 for the left speaker and a right audio signal 1282 for the right speaker, as shown). The audio renderer 1251 produces output signals based on the input audio signals, metadata about the XR scene the listener is experiencing, and information about the listener's location and orientation. The metadata for the XR scene may include metadata for each object and audio element contained in the XR scene, and the metadata for an object may include information about the object's dimensions and occlusion factors for the object (e.g., the metadata may specify a set of occlusion factors, each occlusion factor being applicable for a different frequency or frequency range). The audio renderer 1251 may be a component of the display device 1210, or the audio renderer 1251 may be remote from the listener (e.g., the renderer 1251 may be implemented in the "cloud").

[0055] The orientation sensing unit 1201 is configured to detect changes in the listener's orientation and provide information about the detected changes to the processing unit 1203. In some embodiments, the processing unit 1203 determines an absolute orientation (with respect to some coordinate system) given the detected change in orientation detected by the orientation sensing unit 1201. There may also be different systems for determining orientation and position, for example, systems that use lighthouse trackers (lidar). In one embodiment, the orientation sensing unit 1201 may determine an absolute orientation (with respect to some coordinate system) given the detected change in orientation. In this case, the processing unit 1203 may simply multiplex the absolute orientation data from the orientation sensing unit 1201 and the position data from the position sensing unit 1202. In some embodiments, the orientation sensing unit 1201 may comprise one or more accelerometers and / or one or more gyroscopes.

[0056] 13 shows one example implementation of an audio renderer 1251 for creating sound for an XR scene. The audio renderer 1251 includes a controller 1301 and a signal modifier 1302 for modifying(s) audio signal(s) 1261 (e.g., audio signals of multi-channel audio elements) based on control information 1310 from the controller 1301. The controller 1301 may be configured to receive one or more parameters and trigger the modifier 1302 to implement modifications to the audio signal 1261 (e.g., increase or decrease the volume level) based on the received parameters. The received parameters include information 1263 regarding the listener's position and / or orientation (e.g., direction and distance to the audio element), metadata 1262 regarding the audio element (e.g., audio element 602) in the XR scene, and metadata regarding objects occluding the audio element (e.g., object 154) (in some embodiments, the controller 1301 itself generates the metadata 1262). Using the metadata and position / orientation information, the controller 1301 may calculate another gain factor (g) for audio elements in the XR scene that are at least partially occluded as described above.

[0057] 14 shows an exemplary implementation of the signal modifier 1302, according to one embodiment. The signal modifier 1302 includes a directional mixer 1404, a gain adjuster 1406, and a speaker signal producer 1408.

[0058] Directional mixer 1404 receives audio input 1261, which in this example includes pairs of audio signal 1401 and audio signal 1402 associated with audio elements (e.g., audio element 602), and produces a set of k virtual loudspeaker signals (VS1, VS2, ..., VSk) based on the audio input and control information 1471. In one embodiment, the signal for each virtual loudspeaker may be derived, for example, by appropriate mixing of the signals including audio input 1261. For example, VS1 = α × L + β × R, where L is input audio signal 1401 and R is input audio signal 1402, and α and β are coefficients that depend, for example, on the position of the listener relative to the audio elements and the position of the virtual loudspeaker to which VS1 corresponds.

[0059] In an example where audio element 602 is associated with three virtual loudspeakers (SpL, SpC, and SpR), then k would be equal to 3 for that audio element, and VS1 may correspond to SpL, VS2 may correspond to SpC, and VS3 may correspond to SpR. The control information 1471 used by the directional mixer to create the virtual loudspeaker signals may include the position of each virtual loudspeaker relative to the audio element. In some embodiments, controller 1301 is configured such that when an audio element is occluded, controller 1301 may adjust the position of one or more of the virtual loudspeakers associated with the audio element and provide the position information to directional mixer 1404, which then uses the updated position information to create signals for the virtual loudspeakers (i.e., VS1, VS2, ..., VSk).

[0060] Gain adjuster 1406 may adjust the gain of any one or more of the virtual loudspeaker signals based on control information 1472, which may include the gain factors described above, calculated by controller 1301. That is, for example, when an audio element is at least partially occluded, controller 1301 may control gain adjuster 1406 to adjust the gain of one or more of the virtual loudspeaker signals by providing one or more gain factors to gain adjuster 1406. For example, if the entire left portion of an audio element is occluded, controller 1301 may provide control information 1472 to gain adjuster 1406, causing gain adjuster 1406 to reduce the gain of VS1 by 100% (i.e., gain factor=0, and therefore VS1′=0). As another example, if only 50% of the left portion of an audio element is occluded and 0% of the center portion is occluded, the controller 1301 may provide control information 1472 to the gain adjuster 1406, causing the gain adjuster 1406 to reduce the gain of VS1 by 50% (i.e., VS1′=50% VS1) and not reduce the gain of VS2 at all (i.e., gain factor=1, and therefore VS2′=VS2).

[0061] Using the virtual loudspeaker signals VS1', VS2', ..., VSk', speaker signal producer 1408 produces output signals (e.g., output signal 1281 and output signal 1282) for driving speakers (e.g., headphone speakers or other speakers). In an embodiment where the speakers are headphone speakers, speaker signal producer 1408 may perform conventional binaural rendering to produce the output signals. In an embodiment where the speakers are not headphone speakers, speaker signal producer 1408 may perform conventional speaking panning to produce the output signals.

[0062] 15 is a block diagram of an audio rendering device 1500 according to some embodiments for performing methods disclosed herein (e.g., audio renderer 1251 may be implemented using audio rendering device 1500). As shown in FIG. 15, audio rendering device 1500 includes a processing circuit (PC) 1502 that may include one or more processors (P) 1555 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.), which may be co-sited in a single housing or in a single data center or may be geographically distributed (i.e., device 1500 may be a distributed computing device), and at least one network interface 1548, so that device 1500 can communicate with the network interface. The PC 1502 may comprise at least one network interface 1548, a transmitter (Tx) 1545, and a receiver (Rx) 1547 for enabling the network interface 1548 to transmit data to and receive data from other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 1548 is connected (directly or indirectly) (e.g., the network interface 1548 may be wirelessly connected to the network 110, in which case the network interface 1548 is connected in an antenna configuration), and a storage unit (a.k.a., a "data storage system") 1508, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 1502 includes a programmable processor, a computer program product (CPP) 1541 may be provided. The CPP 1541 includes a computer-readable medium (CRM) 1542, which stores a computer program (CP) 1543 including computer-readable instructions (CRI) 1544. The CRM 1542 may be a non-transitory computer-readable medium, such as a magnetic medium (eg, a hard disk), an optical medium, or a memory device (eg, a random access memory, a flash memory).In some embodiments, the CRI 1544 of the computer program 1543, when executed by the PC 1502, is configured such that the CRI causes the audio-rendering device 1500 to perform the steps described herein (e.g., steps described herein with reference to flowcharts). In other embodiments, the audio-rendering device 1500 may be configured to perform the steps described herein without the need for code. That is, for example, the PC 1502 may simply consist of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.

[0063] Overview of Various Embodiments

[0064] A1. A method for rendering an at least partially occluded audio element (602, 902) represented using a set of two or more virtual loudspeakers (e.g., SpL and SpR), the set including a first virtual loudspeaker (e.g., any one of SpL, SpC, SpR), the method including modifying a first virtual loudspeaker signal (e.g., VS1, VS2, or ...) for the first virtual loudspeaker, thereby producing a first modified virtual loudspeaker signal, and using the first modified virtual loudspeaker signal to render the audio element (e.g., generating an output signal using the first modified virtual loudspeaker signal).

[0065] A2. The method of embodiment A1, further comprising obtaining information indicating that the audio element is at least partially occluded, and wherein the modifying is performed as a result of obtaining the information.

[0066] A3. The method of embodiment A1 or A2, further comprising detecting that the audio element is at least partially occluded, and wherein the modifying is performed as a result of the detection.

[0067] A4. The method of any one of embodiments A1-A3, wherein modifying the first virtual loudspeaker signal includes adjusting a gain of the first virtual loudspeaker signal.

[0068] A5. The method of any one of embodiments A1 to A4, further comprising moving the first virtual loudspeaker from an initial position (e.g., a default position) to a new position, and then generating the first virtual loudspeaker signal using information indicating the new position.

[0069] A6. The method of any one of embodiments A1 to A5, further comprising determining a first occlusion amount (OA1), and wherein the step of modifying the first virtual loudspeaker signal for the first virtual loudspeaker comprises modifying the first virtual loudspeaker signal based on OA1.

[0070] A7. The method of embodiment A6, wherein modifying the first virtual loudspeaker signal based on OA1 includes modifying the first virtual loudspeaker signal so that the modified loudspeaker signal is equal to g1*VS1, where g1 is a gain coefficient calculated using OA1 and VS1 is the first virtual loudspeaker signal.

[0071] A8. The method of embodiment A7, wherein g1 is a function of OA1 (eg, g1=(1-(0.01*OA1)), or g1=sqrt(1-0.01*OA1)).

[0072] A9. A method according to any one of embodiments A6 to A8, wherein the audio element is at least partially occluded by an occluding object, and determining OA1 includes obtaining an occlusion coefficient for the occluding object and determining a proportion of a first sub-area of ​​the projection of the audio element that is covered by the occluding object, and wherein a first virtual loudspeaker is associated with the first sub-area.

[0073] The method of embodiment A9, wherein obtaining an occlusion factor includes selecting an occlusion factor from a set of occlusion factor, the selection being based on a frequency associated with an audio element. For example, each occlusion factor (OF) included in the set of occlusion factor is associated with a different frequency range, and the selection is based on the frequency associated with the audio element, such that the selected OF is associated with a frequency range that encompasses the frequency associated with the audio element.

[0074] A11. The method of embodiment A9 or A10, wherein determining OA1 includes calculating OA1=Of1*P, where Of1 is an occlusion factor and P is a proportion.

[0075] A12. The method of any one of embodiments A1 to A11, further comprising: modifying a second virtual loudspeaker signal for a second virtual loudspeaker, thereby producing a second modified virtual loudspeaker signal; and using the first modified virtual loudspeaker signal and the second modified virtual loudspeaker signal to render an audio element.

[0076] A13. The method of embodiment A12, further comprising determining a second occlusion amount (OA2) associated with the second virtual loudspeaker, and wherein the step of modifying the second virtual loudspeaker signal comprises modifying the second virtual loudspeaker signal based on OA2.

[0077] A14. The method of embodiment A13, wherein modifying the second virtual loudspeaker signal based on OA2 includes modifying the second virtual loudspeaker signal so that the second modified loudspeaker signal is equal to g2*VS2, where g2 is a gain coefficient calculated using OA2 and VS2 is the second virtual loudspeaker signal.

[0078] A15. The method of embodiment A13 or A14, wherein determining OA2 includes determining a proportion of a second sub-area of ​​the projection of the audio element that is covered by the occluding object, and a second virtual loudspeaker is associated with the second sub-area.

[0079] B1. A method for rendering an at least partially occluded audio element (602, 902) represented using a set of two or more virtual loudspeakers, the set including a first virtual loudspeaker and a second virtual loudspeaker, the method including: moving the first virtual loudspeaker from an initial position to a new position; generating a first virtual loudspeaker signal for the first virtual loudspeaker based on the new position of the first virtual loudspeaker; and using the first virtual loudspeaker signal to render the audio element.

[0080] B2. The method of embodiment B1, further comprising obtaining information indicating that the audio element is at least partially occluded, and wherein the moving is performed as a result of obtaining the information.

[0081] B3. The method of embodiment B1 or B2, further comprising detecting that the audio element is at least partially occluded, and wherein the moving is performed as a result of the detection.

[0082] C1. A computer program comprising instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform a method according to any one of the preceding embodiments.

[0083] C2. A carrier containing the computer program as described above, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.

[0084] D1. An audio rendering device configured to perform the method according to any one of the preceding embodiments.

[0085] D2. The audio rendering device of embodiment D1, wherein the audio rendering device comprises a memory and a processing circuit coupled to the memory.

[0086] While various embodiments have been described herein, it should be understood that these embodiments have been presented by way of example only, and not limitation. Thus, the breadth and scope of the present disclosure should not be limited by any of the exemplary embodiments described above. Moreover, unless otherwise indicated herein or clearly contradicted by context, any combination of the above-described objects in all possible variations thereof is encompassed by the present disclosure.

[0087] Additionally, while the processes described above and illustrated in the figures have been shown as a sequence of steps, this has been done for purposes of illustration only, and it is therefore contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.

[0088] References [1] MPEG-H 3D Audio, Clause 8.4.4.7: “Spreading” [2] MPEG-H 3D Audio, Clause 18.1: “Element Metadata Preprocessing” [3] MPEG-H 3D Audio, Clause 18.11: “Diffuseness Rendering” [4] EBU ADM Renderer Tech 3388, Clause 7.3.6: “Divergence” [5] EBU ADM Renderer Tech 3388, Clause 7.4: “Decorrelation Filters” [6] EBU ADM Renderer Tech 3388, Clause 7.3.7: “Extent Panner” [7] Efficient HRTF-based Spatial Audio for Area and Volumetric Sources“, IEEE Transactions on Visualization and Computer Graphics 22(4):1-1·January 2016 [8] Patent Publication WO2020144062, “Efficient spatially-heterogeneous audio elements for Virtual Reality.”

Claims

1. A method (400) for rendering an at least partially occluded audio element (602, 902) represented using a set of two or more virtual loudspeakers (SpL, SpC, SpR), said set including a first virtual loudspeaker, said method comprising: modifying (s402) a first virtual loudspeaker signal for the first virtual loudspeaker, thereby producing a first modified virtual loudspeaker signal; using the first modified virtual loudspeaker signal to render the audio element (s404); and The method (400) includes:

2. 2. The method of claim 1, further comprising obtaining information indicating that the audio element is at least partially occluded, and wherein the modifying (s402) is performed as a result of obtaining the information.

3. The method of claim 1 , further comprising detecting that the audio element is at least partially occluded, and wherein the modifying (s402) is performed as a result of the detection.

4. The method of claim 1 , wherein modifying the first virtual loudspeaker signal comprises adjusting a gain of the first virtual loudspeaker signal.

5. 5. The method of claim 1, further comprising: moving the first virtual loudspeaker from an initial position to a new position; and then generating the first virtual loudspeaker signal using information indicative of the new position.

6. 6. The method of claim 1, further comprising determining a first occlusion amount O1, and wherein the step of modifying the first virtual loudspeaker signal for the first virtual loudspeaker comprises modifying the first virtual loudspeaker signal based on O1.

7. 7. The method of claim 6, wherein modifying the first virtual loudspeaker signal based on O1 comprises modifying the first virtual loudspeaker signal so that the modified loudspeaker signal is equal to g1*VS1, where g1 is a gain factor calculated using O1 and VS1 is the first virtual loudspeaker signal.

8. g1=(1-0.01*O1), or g1=sqrt(1-0.01*O1), The method of claim 7.

9. the audio element is at least partially occluded by an occluding object (604, 614); determining O1 comprises obtaining an occlusion coefficient for the occluding object and determining a percentage of a first sub-area of ​​the projection of the audio element that is covered by the occluding object, wherein the first virtual loudspeaker is associated with the first sub-area; 9. The method according to any one of claims 6 to 8.

10. 10. The method of claim 9, wherein obtaining the occlusion factor comprises selecting an occlusion factor OF from a set of occlusion factor, each OF included in the set of occlusion factor being associated with a different frequency range, the selection being based on frequencies associated with the audio element, such that the selected OF is associated with a frequency range that encompasses the frequencies associated with the audio element.

11. 11. The method of claim 9 or 10, wherein determining O1 comprises calculating O1=Of1*P, where Of1 is the occlusion factor and P is the proportion.

12. modifying the second virtual loudspeaker signal for a second virtual loudspeaker, thereby producing a second modified virtual loudspeaker signal; using the first modified virtual loudspeaker signal and the second modified virtual loudspeaker signal to render the audio element; 12. The method of claim 1, further comprising:

13. 13. The method of claim 12, further comprising determining a second occlusion amount O2 associated with the second virtual loudspeaker, and wherein modifying the second virtual loudspeaker signal comprises modifying the second virtual loudspeaker signal based on O2.

14. 14. The method of claim 13, wherein modifying the second virtual loudspeaker signal based on O2 comprises modifying the second virtual loudspeaker signal so that the second modified loudspeaker signal is equal to g2*VS2, where g2 is a gain factor calculated using O2 and VS2 is the second virtual loudspeaker signal.

15. 15. The method of claim 13 or 14, wherein determining O2 comprises determining a proportion of a second sub-area of ​​the projection of the audio element that is covered by an occluding object, and wherein the second virtual loudspeaker is associated with the second sub-area.

16. 1. A method (450) for rendering an at least partially occluded audio element (602, 902) represented using a set of two or more virtual loudspeakers (SpL, SpC, SpR), said set including a first virtual loudspeaker, said method comprising: moving (s452) the first virtual loudspeaker from an initial position to a new position; generating (s454) a first virtual loudspeaker signal for the first virtual loudspeaker based on the new position of the first virtual loudspeaker; using the first virtual loudspeaker signal to render the audio element (s456); and The method (450) includes:

17. 17. The method of claim 16, further comprising obtaining information indicating that the audio element is at least partially occluded, and wherein the moving (s452) is performed as a result of obtaining the information.

18. The method of claim 16 , further comprising detecting that the audio element is at least partially occluded, and wherein the moving (s452) is performed as a result of the detection.

19. A computer program (1543) comprising instructions (1544) that, when executed by a processing circuit (1502) of an audio renderer device (1500), cause the audio renderer device to perform a method according to any one of claims 1 to 18.

20. 20. A carrier containing the computer program of claim 19, said carrier being one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1542).

21. 1. An audio rendering device (1500) for rendering an at least partially occluded audio element (602, 902) represented using a set of two or more virtual loudspeakers (SpL, SpC, SpR), said set including a first virtual loudspeaker, said audio rendering device comprising: modifying (s402) a first virtual loudspeaker signal for the first virtual loudspeaker, thereby producing a first modified virtual loudspeaker signal; using the first modified virtual loudspeaker signal to render the audio element (s404); and An audio rendering device (1500) configured to:

22. 22. The audio rendering apparatus (1500) of claim 21, further configured to perform the step of obtaining information indicating that the audio element is at least partially occluded, and wherein the modifying is performed as a result of obtaining the information.

23. 22. The audio rendering apparatus (1500) of claim 21, further configured to perform the step of detecting that the audio element is at least partially occluded, and wherein the modifying is performed as a result of the detection.

24. 24. The audio rendering apparatus (1500) of any one of claims 21 to 23, wherein modifying the first virtual loudspeaker signal comprises adjusting a gain of the first virtual loudspeaker signal.

25. 25. The audio rendering apparatus (1500) of claim 21, further configured to perform the steps of moving the first virtual loudspeaker from an initial position to a new position and then generating the first virtual loudspeaker signal using information indicative of the new position.

26. 26. An audio rendering apparatus (1500) according to any one of claims 21 to 25, further configured to perform a step of determining a first occlusion amount O1, and wherein the step of modifying the first virtual loudspeaker signal for the first virtual loudspeaker comprises modifying the first virtual loudspeaker signal based on O1.

27. 27. The audio rendering apparatus (1500) of claim 26, wherein modifying the first virtual loudspeaker signal based on O1 includes modifying the first virtual loudspeaker signal so that the modified loudspeaker signal is equal to g1*VS1, where g1 is a gain factor calculated using O1 and VS1 is the first virtual loudspeaker signal.

28. g1=(1-0.01*O1), or g1=sqrt(1-0.01*O1), 28. An audio rendering device (1500) according to claim 27.

29. the audio element is at least partially occluded by an occluding object (604, 614); determining O1 comprises obtaining an occlusion coefficient for the occluding object and determining a percentage of a first sub-area of ​​the projection of the audio element that is covered by the occluding object, wherein the first virtual loudspeaker is associated with the first sub-area; An audio rendering device (1500) according to any one of claims 26 to 28.

30. 30. The audio rendering device (1500) of claim 29, wherein obtaining the occlusion coefficient includes selecting an occlusion coefficient OF from a set of occlusion coefficients, each OF included in the set of occlusion coefficients being associated with a different frequency range, the selection being based on frequencies associated with the audio element, such that the selected OF is associated with a frequency range that encompasses the frequencies associated with the audio element.

31. 31. An audio rendering apparatus (1500) according to claim 29 or 30, wherein determining O1 comprises calculating O1 = Of1 * P, where Of1 is the occlusion factor and P is the proportion.

32. modifying a second virtual loudspeaker signal for a second virtual loudspeaker, thereby producing a second modified virtual loudspeaker signal; using the first modified virtual loudspeaker signal and the second modified virtual loudspeaker signal to render the audio element; 32. An audio rendering device (1500) according to any one of claims 21 to 31, further configured to perform:

33. 33. The audio rendering apparatus (1500) of claim 32, further configured to perform a step of determining a second occlusion amount O2 associated with the second virtual loudspeaker, and wherein the step of modifying the second virtual loudspeaker signal comprises modifying the second virtual loudspeaker signal based on O2.

34. 34. The audio rendering apparatus (1500) of claim 33, wherein modifying the second virtual loudspeaker signal based on O2 comprises modifying the second virtual loudspeaker signal so that the second modified loudspeaker signal is equal to g2*VS2, where g2 is a gain factor calculated using O2 and VS2 is the second virtual loudspeaker signal.

35. 35. The audio rendering apparatus (1500) of claim 33 or 34, wherein determining O2 comprises determining a proportion of a second sub-area of ​​the projection of the audio element that is covered by the occluding object, and wherein the second virtual loudspeaker is associated with the second sub-area.

36. 1. An audio rendering device (1500) for rendering an at least partially occluded audio element (602, 902) represented using a set of two or more virtual loudspeakers (SpL, SpC, SpR), said set including a first virtual loudspeaker, said audio rendering device comprising: moving (s452) the first virtual loudspeaker from an initial position to a new position; generating (s454) a first virtual loudspeaker signal for the first virtual loudspeaker based on the new position of the first virtual loudspeaker; using the first virtual loudspeaker signal to render the audio element (s456); and An audio rendering device (1500) configured to:

37. 37. The audio rendering apparatus (1500) of claim 36, further configured to perform the step of obtaining information indicating that the audio element is at least partially occluded, and wherein the moving is performed as a result of obtaining the information.

38. 37. The audio rendering apparatus (1500) of claim 36, further configured to perform the step of detecting that the audio element is at least partially occluded, and wherein said moving is performed as a result of said detection.

39. 37. An audio rendering device according to any one of claims 21 to 36, comprising a memory (1542) and a processing circuit (1502) coupled to the memory.