Rendering audio elements

By dividing the speaker system into hemispheres and applying gain adjustments based on listener position, the method addresses the challenge of transitioning between internal and external audio element representations, ensuring a smooth and natural-sounding spatial audio experience.

JP2026012711APending Publication Date: 2026-01-27TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025167739
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-01
Filing Date
2025-10-03
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing spatial audio rendering methods struggle to smoothly transition between internal and external representations of audio elements when a listener moves across their spatial boundaries, particularly failing to handle positions above or below the boundary and causing orientation instability.

Method used

A method is introduced to divide the speaker system into hemispheres and apply gain adjustments to the upper, lower, and rear hemispheres based on the listener's vertical position relative to the audio element's boundaries, ensuring smooth transitions by attenuating sound from inappropriate directions.

Benefits of technology

This approach effectively handles listener positions above or below the audio element's boundary, providing a seamless and natural-sounding audio experience by adjusting speaker gains to maintain consistent orientation and spatial harmony.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012711000001_ABST
    Figure 2026012711000001_ABST
Patent Text Reader

Abstract

To provide a method for successfully handling a situation where a listening point is above or below the range of an audio element.SOLUTION: The method comprises a step s602 of determining top gain values (G _ top) for the top part of the internal representation of the audio element based on L, the perpendicular distance between the reference plane and the listening point, and T, the perpendicular distance between the reference plane and the top point of the extent of the audio element, and a step s604 of determining bottom gain values (G _ bottom) for the bottom part of the internal representation of the audio element based on L, and B, the perpendicular distance between the reference plane and the bottom point of the extent of the audio element.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] SUMMARY OF THE INVENTION Embodiments related to rendering audio elements are disclosed. [Background technology]

[0002] Spatial audio rendering is a process used to present audio in an extended reality (XR) scene (e.g., a virtual reality (VR), augmented reality (AR), or mixed reality (MR) scene) to give the listener the impression that the sound is coming from a physical source within the scene at a certain location and having a certain size and shape (i.e., extent). Presentation can be through headphone speakers or other speakers. When presentation is through headphone speakers, the process used is called binaural rendering and uses spatial cues of human spatial hearing that make it possible to determine which direction a sound is coming from. The cues involve inter-aural time delay (ITD), inter-aural level difference (ILD), and / or spectral differences.

[0003] The most common form of spatial audio rendering is based on the point source concept, where each sound source is defined to emanate from one specific point. Because each sound source is defined to emanate from one specific point, the sound sources do not have a size or shape. Different methods have been developed to render sound sources that have a range (size and shape).

[0004] One such known method is to create multiple copies of a mono audio element at positions around the audio element. This configuration results in the perception of a spatially uniform object with a certain size. This concept is used, for example, in the "object spread" and "object divergence" features of the MPEG-H 3D audio standard (see References [1] and [2]) and in the "object divergence" feature of the EBU Audio Specification Model (ADM) standard (see Reference [4]). This idea using a mono audio source is further developed as described in Reference [7], where the area-volume geometry of the sound object is projected onto a sphere around the listener, and the sound is rendered to the listener using a pair of head-related (HR) filters that is evaluated as the integral of all HR filters covering the geometric projection of the object on the sphere. For a spherical volume source, this integral has an analytical solution. However, for arbitrary area-volume source geometries, the integral is evaluated by sampling the projected source surface on the sphere using so-called Monte Carlo ray sampling.

[0005] Another rendering method renders a spatially diffuse component in addition to the mono audio signal, which results in the perception of somewhat diffuse objects that, in contrast to the original mono audio element, do not have a distinct pinpoint location. This concept is used, for example, in the "object diffuseness" feature of the MPEG-H 3D audio standard (see reference [3]) and the "object diffuseness" feature of the EBU ADM (see reference [5]).

[0006] Combinations of the two methods above are also known, for example the "object extent" feature of the EBU ADM, which combines the creation of multiple copies of a mono audio element with the addition of a diffuse component (see reference [6]).

[0007] In many cases, the actual shape of an audio element can be described sufficiently well using primitive shapes (e.g., a sphere or a box). However, sometimes the actual shape is more complex and needs to be described in a more detailed form (e.g., a mesh structure or a parametric description format).

[0008] Some audio elements are of a nature such that a listener can move within the range for the audio element (i.e., the spatial boundary of the audio element) and expect to hear a plausible audio representation of the audio element. For these audio elements, the range serves as a spatial boundary that defines the edge between the interior and exterior of the audio element. Examples of such audio elements include a forest (birds, wind in the trees), a crowd of people (sounds of people clapping or cheering), and the background sounds of a plaza (sounds of traffic, birds, people walking).

[0009] When the listener moves within the spatial boundaries of the audio element, the audio presentation should be immersive and surround the listener. When the listener moves out of the spatial boundaries, the presentation should now appear to come from the range of the audio element.

[0010] These audio elements can be represented as a number of individual point sources, but it is more efficient to represent them using a single audio signal. For internal audio representation, a listener-centric format is preferred, in which the sound field around the listener is described. Listener-centric formats include channel-based formats such as 5.1, 7.1, and scene-based formats such as Ambisonic. Listener-centric formats are generally rendered using several virtual speakers (or "speakers" for short) positioned around the listener. Summary of the Invention

[0011] Currently, several challenges exist. For example, there is no clearly defined way to directly render a listener-centric audio signal when the listener position is outside the spatial boundary. When the listener is located outside the spatial boundary, a source-centric representation is more suitable because the sound source no longer surrounds the listener, but should instead be rendered as coming from a distance in a certain direction. One solution is to use a listener-centric audio signal for the internal representation and derive a source-centric audio signal from it, which can then be rendered using source-centric techniques. This technique is described in Reference [8]. Furthermore, a technique for rendering an external representation of such audio elements, whose extents can be of arbitrary shape, is described in Reference [9]. However, one challenge with these solutions is making the transition between the internal and external representations smooth and natural-sounding. Reference

[10] describes a method for rendering a smooth transition between the external and internal representations. Reference

[10] describes a method for attenuating the rear hemisphere of a speaker setup used for interior rendering when the listener is close to the surface of the range. This makes the transition more natural, as when the listener is located near the surface of the range, the audio appears to come from within the range rather than surrounding the listener. As the listener moves further inside the range, the attenuation is gradually reduced, so that the listener becomes increasingly completely surrounded by audio from all sides.

[0012] The method described in

[10] for modifying the internal representation when the listener is close to a range surface is based on the alignment of the speaker system of the internal representation to the surface of the range of the audio source. This alignment makes it possible to determine the speakers that represent the outside of the range. Two variants of this method are described: one in which the alignment is performed only in the horizontal plane, and one in which the alignment is based on an observation vector, a vector from the listener position to a target point on the range.

[0013] The problem with the first variant of this method is that, because the matching is done only in the horizontal dimension, there is no way to properly handle listening points that are above or below the range, and therefore there is no way to modify the internal representation rendering so that sound from the audio source appears to come from above or below.

[0014] A second variant of this method uses alignment in both the horizontal and vertical dimensions based on the observation vector, which allows for handling cases where the listener is above or below the range. However, using alignment in both the horizontal and vertical dimensions can cause problems with the stability of the rendering speaker system's orientation, which can change quickly as the listener approaches the range surface. Often, when the listener is within the range, the closest point of the range will be directly below the listener (e.g., on the "floor" of the range). As the listener moves closer to the range surface, at some point the closest point will suddenly be on the nearest "wall" of the range. This will result in a sudden and large rotation of the speaker system as the listener approaches the range surface, which will create audible artifacts. Also, this method will not properly handle cases where both the rear and upper parts of the speaker system, or both the rear and lower parts, should be attenuated.

[0015] Thus, in one aspect, a method for rendering an audio element is provided, the method including at least one of the following steps: 1) determining a top gain value (G_top) for an upper part of an internal representation of the audio element based on L and T, where L is the vertical distance between a reference plane and a listening point and T is the vertical distance between the reference plane and a top point of a range for the audio element, or 2) determining a bottom gain value (G_bottom) for a lower part of the internal representation of the audio element based on L and B, where B is the vertical distance between the reference plane and a bottom point of a range for the audio element.

[0016] In another aspect, a computer program is provided comprising instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform the method described above. In one embodiment, a carrier is provided that includes the computer program, the carrier being one of an electronic signal, an optical signal, a wireless signal, and a computer-readable storage medium. In another aspect, a rendering apparatus is provided that is configured to perform the method described above. The rendering apparatus may include a memory and a processing circuit coupled to the memory.

[0017] An advantage of the embodiments disclosed herein is that they handle well situations where the listening point is above or below the range of the audio element.

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various embodiments. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 illustrates an exemplary speaker system. [Figure 2] FIG. 1 illustrates an example of dividing a speaker system into hemispheres. [Figure 3A] FIG. 1 illustrates a listening point above an audio element. [Figure 3B] FIG. 1 illustrates a listening point below an audio element. [Figure 4] FIG. 1 illustrates horizontal contours for audio elements. [Figure 5] FIG. 1 is a diagram showing various listening points. [Figure 6] 1 is a flowchart illustrating a process according to some embodiments. [Figure 7A] FIG. 1 illustrates a system, according to some embodiments. [Figure 7B] FIG. 1 illustrates a system, according to some embodiments. [Figure 8] FIG. 1 illustrates a system, according to some embodiments. [Figure 9] FIG. 2 illustrates a signal modifier according to one embodiment. [Figure 10] FIG. 1 is a block diagram of an apparatus, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0020] Typically, the internal representation of the audio elements is rendered using a speaker system comprising a set of virtual speakers arranged in a spherical configuration around a listening point. This is illustrated in FIG. 1, which shows an exemplary speaker system 100 comprising a set of virtual speakers S1-S18 arranged in a spherical configuration around a listening position 101 (also referred to as the "listener" 101 or "listening point" 101). The number of speakers and their positions can vary, but they are generally positioned at equal distances from the listener position. Vector F represents the front vector of speaker system 100. The front vector defines the orientation of the speaker system and is independent of the listener's head rotation.

[0021] In one embodiment, the set of speakers S1-S18 is divided into four hemispheres, as shown in Figure 2: front 201, rear 202, top 203, and bottom 204. Regardless of the exact configuration of the speaker sets, the speaker system as a whole has a rotation defined by a front vector. The front vector represents the direction the front hemisphere is pointed. By attenuating the gain of a signal going to a speaker in one of the hemispheres, sound energy from the corresponding direction can be reduced.

[0022] The gain of the rear hemisphere, the gain of the upper hemisphere, and the gain of the lower hemisphere can be attenuated independently to create the effect that sound is coming only from the direction of the audio element. For example, if audio element 302 (see FIG. 3A) is directly in front of listening point 101, the gain of the rear hemisphere should be attenuated. If the listening point is above the audio element, as shown in FIG. 3A, the gain of the upper hemisphere should be attenuated. Similarly, if the listening point is below the audio element, as shown in FIG. 3B, the gain of the lower hemisphere should be attenuated. If the listening point is above the range and near its edge, as shown in FIG. 3A, the gain of both the rear and upper hemispheres can be attenuated.

[0023] Arrow 304 shown in Figures 3A and 3B indicates the front vector of the horizontal alignment of the internal representation speaker setup. In the example shown in Figure 3A, the listening point 101 is above and near the edge of the audio element 302, which in these examples has a simplified rectangular extent. In this situation, the internal representation should be modified so that audio does not sound from above or behind. This can be achieved by attenuating the top 203 and rear 202 hemispheres of the speaker system. In the example shown in Figure 3B, the listening point is below and near the edge of the extent; in this case, the internal representation should be modified so that audio does not sound from below or behind.

[0024] The attenuation can either go all the way to 0 so that the hemispheres can be completely muted, or the attenuation can be limited so that the hemispheres are only attenuated to a certain extent to achieve a softer spatial effect.

[0025] Separate gain coefficients (also known as gain values) are calculated for the rear, upper, and lower hemispheres. These gain coefficients are then applied to the corresponding virtual speaker signals corresponding to each hemisphere. Some speakers in the system may belong to two (or more) hemispheres, and the signals for these speakers should be affected by the gain coefficients for each hemisphere to which the speaker belongs.

[0026] For example, speaker S4 resides in both the upper and rear hemispheres. The corrected signal y4' for that speaker is then y4'=y4*G_rear*G_top where y4 is the original signal before spatial modification, and G_rear and G_top are gain factors for the rear and top hemispheres.

[0027] As an example, for the entire speaker system shown in Figures 1 and 2, the calculation is: It could be something like TIFF2026012711000002.tif102170.

[0028] Calculation of posterior hemisphere gain: In one embodiment, horizontal alignment of the speaker system 100 is used. This alignment rotates the speaker system so that its front vector points horizontally in the direction of the range. This alignment does not take into account the relative height of the range and the listener; the alignment is only used to control the attenuation in the rear hemisphere. Because height information is discarded when performing this alignment, the alignment can be performed on the contour of the range's projection onto a horizontal plane, as shown in FIG. 4. FIG. 4 shows a horizontal contour 400 found by projecting the spherical range 410 of an audio element onto a horizontal plane and finding the contour of that projection. The front vector of the speaker system 100 should point in the negative direction of the normal of the closest point of the horizontal contour to the listening point projected on the horizontal plane.

[0029] The alignment should ensure that the front vector of the speaker system 100 points inward toward the range, and that the left and right of the speaker system are aligned with the horizontal contour of the range. As the listener 101 moves around, the rear hemisphere should always point away from the range. In other words, the front vector of the speaker system should be aligned with the normal of the nearest point on the horizontal contour of the range.

[0030] When the listener is at a distance from the range and not above or below the range, the rear hemisphere always represents the side pointing away from the range and therefore can always be attenuated unless the listener is inside the range.

[0031] When the listener is inside, above, or below the range, the projected listening point will be inside the horizontal contour. In this case, the rear hemisphere should not be attenuated. To have a smooth transition, an inner fade region can be used, where the attenuation is gradually reduced, as described in Reference

[10] . The fade region can also be an outer region, where the attenuation is gradually reduced until the listener crosses the horizontal contour of the range.

[0032] Calculation of upper and lower hemisphere gain: To control the attenuation in the upper hemisphere when the listener is above the range, or the attenuation in the lower hemisphere when the listener is below the range, the height of the listener position should be compared to the height of the range (i.e., the vertical distance between the reference plane and the listening point is compared to the vertical distance between the reference plane and the top point of the range).

[0033] In one embodiment, the upper hemisphere is attenuated if the listener position is higher than the top point of the range (or the top point of a selected portion of the range). For example, in one embodiment, the upper gain factor is a function of the difference between L and T, where L is the vertical distance between the listening point and a reference plane, and T is the vertical distance between the top point of the audio element (or a simplified range representing the audio element) and the reference plane. Similarly, the lower hemisphere is attenuated if the listener position is lower than the bottom point of the range (or the bottom point of a selected portion of the range). For example, in one embodiment, the lower gain factor is a function of the difference between L and B, where B is the vertical distance between the bottom point of the audio element (or a simplified range representing the audio element) and the reference plane.

[0034] This is shown in FIG. 5 . That is, the top point 501 and bottom point 502 of the range 410 of the audio element are used to define where the gain attenuation for the upper and lower hemispheres should begin and end. Optionally, there can be a fade region where the attenuation can be gradually introduced. As shown in FIG. 5 , listening point A1 is above the top point 501 of the range 410, and therefore the upper hemisphere should be attenuated (i.e., the vertical distance 580 between position A1 and the reference plane 590 is greater than the vertical distance 581 between the top point 501 and the reference plane). Listening point A2 is inside the fade region where the attenuation for the upper hemisphere is gradually reduced. Listening point A3 is halfway between the top and bottom of the range, where no attenuation is applied to the upper or lower hemispheres. Listening point A4 is inside the fade region where the attenuation for the lower hemisphere is gradually introduced. Listening point A5 is below the lowest point 502 of the range, where the lower hemisphere should be attenuated (i.e., the vertical distance 582 between position A5 and reference plane 590 is less than the vertical distance 583 between the lowest point 502 and reference plane 590). However, using the top and bottom points of range 410 as the basis for adaptation may not work ideally for very large ranges with more complex shapes, and the height of the range may vary in different parts of the range.

[0035] To handle large, complex ranges, a method can be used that considers only the part of the range that is relevant for a given listening point. This can mean that only the part of the range that is within a certain distance from the listener is taken into account, or that only the part of the range that is seen as the perceptually relevant part of the range using some perceptual model. Thus, in this embodiment, the upper hemisphere is attenuated if the listener position is higher than the top point of the relevant part of the range, and the lower hemisphere is attenuated if the listener position is lower than the bottom point of the relevant part of the range.

[0036] If there is an external representation available, as in Reference

[10] , this may represent a perceptually relevant part of the range, in which case only the points that define the external representation need to be evaluated. However, the external representation is not valid when the listener is inside the range, so in this method it may be beneficial to have a fade region outside the range, such that the attenuation is gradually reduced as one gets closer to the range, and there is no attenuation when the listener is inside the range.

[0037] Modification of the internal representation in the spatial harmonics domain: In some cases, the rendering of the internal representation is not done using virtual speakers, but instead with direct rendering from the internal representation, for example, an Ambisonic signal can be rendered directly to a binaural signal in the spherical harmonics domain. In this case, the attenuation of different hemispheres by applying gain factors to the individual loudspeaker signals cannot be done; instead, a spatial correction needs to be applied in the spatial harmonics domain before rendering takes place. Several methods are known how this spatial correction should be done; for example, so-called spatial capping can be used to perform directional loudness correction on Ambisonic signals, as described in Reference

[11] .

[0038] The same principles can be used to derive the required gain for the upper, lower and rear hemispheres described previously, but then the application of gain modifications is done using one spatial cap function for each hemisphere that should be attenuated.

[0039] 6 is a flow diagram illustrating a process 600 for rendering an audio element, according to one embodiment. Process 600 may begin at step s602 or step s604.

[0040] Step s602 includes determining a top gain value (G_top) for the top part of the internal representation of the audio element based on L and T, where L is the vertical distance between the reference plane and the listening point and T is the vertical distance between the reference plane and the top point of the range of the audio element (e.g., point 501). For example, in one embodiment, when L is greater than T, G_top is inversely proportional to the difference between L and T (e.g., G_top≈α×1 / (LT), where α is a predetermined correction factor). This would mean that G_top is faded out in the area above the top point.

[0041] In another embodiment, G_top is TIFF2026012711000003.tif19170, where β describes the size of the fade region below the top point. In another embodiment, the fade region is above the top point, in which case G_top is calculated as It can be calculated as TIFF2026012711000004.tif19170.

[0042] Step s604 involves determining a bottom gain value (G_bottom) for the bottom part of the internal representation of the audio element based on L and B, where B is the vertical distance between the reference plane and the bottom point of the range of the audio element (e.g., point 502). For example, in one embodiment, when L is smaller than B, G_bottom is inversely proportional to the difference between B and L (e.g., G_bottom ≈ α × 1 / (BL)). This would mean that G_bottom is faded out in the area below the bottom point.

[0043] In another embodiment, G_bottom is TIFF2026012711000005.tif19170, where β describes the size of the fade region above the bottom point 502. In another embodiment, the fade region is below the bottom point, in which case G_bottom is calculated as It can be calculated as TIFF2026012711000006.tif19170.

[0044] In some embodiments, T is the vertical distance between the top point 501 of the selected portion of the range for the audio element and the reference plane, and B is the vertical distance between the bottom point 502 of the selected portion of the range for the audio element and the reference plane.

[0045] In some embodiments, an audio element has an original range, and said range of an audio element is a simplified range for the audio element that represents the original range from a certain listening point.

[0046] In some embodiments, the audio elements are represented using a set of virtual speakers (e.g., speakers S1-S18) comprising a set of one or more upper virtual speakers located above listening point 101 and / or a set of one or more lower virtual speakers located below the listening point.

[0047] In some embodiments, the set of upper virtual speakers comprises a first upper virtual speaker, and the method further includes producing a first upper virtual speaker signal y1 for the first upper virtual speaker, producing a gain-adjusted first upper virtual speaker signal y1′, where y1′=g*y1, where g is a function of at least the upper gain value, and using y1′ to render the audio element.

[0048] In some embodiments, the set of virtual speakers comprises two or more sets of rear virtual speakers, the set of rear virtual speakers comprising a first top virtual speaker, and the method further includes determining a rear gain value (G_rear) for the set of rear virtual speakers, where g is a function of at least the top gain value (G_top) and the rear gain value (G_rear) (i.e., g=f(G_top,G_rear), for example, g=G_top*G_rear).

[0049] In some embodiments, the set of lower virtual speakers comprises a first lower virtual speaker, and the method further includes producing a first lower virtual speaker signal y2 for the first lower virtual speaker, producing a gain-adjusted first lower virtual speaker signal y2′, where y2′=g*y2, where g is a function of at least the lower gain value, and using y2′ to render the audio element.

[0050] In some embodiments, the set of virtual speakers comprises two or more sets of rear virtual speakers, the set of rear virtual speakers comprising a first bottom virtual speaker, and the method further includes determining a rear gain value (G_rear) for the set of rear virtual speakers, where g is a function of at least the bottom gain value (G_bottom) and the rear gain value (G_rear) (i.e., g=f(G_bottom, G_rear), for example, g=G_bottom*G_rear).

[0051] Exemplary Use Cases 7A shows an XR system 700 to which embodiments disclosed herein may be applied. The XR system 700 includes speakers 704 and 705 (which may be speakers of headphones worn by a listener) and an XR device 710 that may include a display for displaying images to a user and, in some embodiments, is configured to be worn by a listener. In the illustrated XR system 700, the XR device 710 has a display and is designed to be worn on the user's head, and is commonly referred to as a head-mounted display (HMD).

[0052] As shown in FIG. 7B, the XR device 710 may include an orientation sensing unit 701, a position sensing unit 702, and a processing unit 703 coupled (directly or indirectly) to an audio renderer 751 to produce output audio signals (e.g., a left audio signal 781 for the left speaker and a right audio signal 782 for the right speaker, as shown).

[0053] The orientation sensing unit 701 is configured to detect changes in the listener's orientation and provide information about the detected changes to the processing unit 703. In some embodiments, the processing unit 703 determines an absolute orientation (with respect to some coordinate system) given the detected change in orientation detected by the orientation sensing unit 701. There may also be different systems for determining orientation and position, for example, systems that use lighthouse trackers (LIDAR). In one embodiment, the orientation sensing unit 701 may determine an absolute orientation (with respect to some coordinate system) given the detected change in orientation. In this case, the processing unit 703 may simply multiplex the absolute orientation data from the orientation sensing unit 701 and the position data from the position sensing unit 702. In some embodiments, the orientation sensing unit 701 may comprise one or more accelerometers and / or one or more gyroscopes.

[0054] The audio renderer 751 produces an audio output signal based on an input audio signal 761, metadata 762 about the XR scene the listener is experiencing, and information 763 about the listener's location and orientation. The metadata 762 for the XR scene may include metadata for each object and audio element contained in the XR scene, and the metadata for an object or audio element may include information about the range of the object or audio element. The metadata 762 may also include control information, such as reverberation time values, reverberation level values, and / or absorption parameters. The audio renderer 751 may be a component of the XR device 710, or the audio renderer 751 may be remote from the XR device 710 (e.g., the audio renderer 751, or components thereof, may be implemented in the so-called "cloud").

[0055] 8 shows one example implementation of an audio renderer 751 for creating sound for an XR scene. The audio renderer 751 includes a controller 801 and a signal modifier 802 for modifying(s) audio signal(s) 761 (e.g., audio signals of multi-channel audio elements) based on control information 810 from the controller 801. The controller 801 may be configured to receive one or more parameters and trigger the modifier 802 to implement a modification to the audio signal 761 (e.g., increase or decrease the volume level) based on the received parameters. The received parameters include information 763 regarding the listener's position and / or orientation (e.g., direction and distance to the audio elements) and metadata 762 regarding the audio elements in the XR scene (in some embodiments, the controller 801 itself generates the metadata 762). Using the metadata and position / orientation information, the controller 801 may calculate another gain factor (also known as an attenuation factor) for the audio elements in the XR scene as described herein.

[0056] 9 shows an exemplary implementation of a signal modifier 802, according to one embodiment. The signal modifier 802 includes a directional mixer 904, a gain adjuster 906, and a speaker signal producer 908.

[0057] The directional mixer 904 receives audio input 761, which in this example includes pairs of audio signal 901 and audio signal 902 associated with audio elements, and produces a set of k virtual speaker signals (y1, y2, ..., yk) based on the audio input and control information 991. In one embodiment, the signal for each virtual speaker may be derived, for example, by appropriate mixing of signals including audio input 761. For example, y1 = f1 × L + f2 × R, where L is input audio signal 901, R is input audio signal 902, and f1 and f2 are coefficients that depend, for example, on the position of the listener relative to the audio elements and the position of the virtual loudspeaker to which y1 corresponds.

[0058] Gain adjuster 906 may adjust the gain of any one or more of the virtual speaker signals based on control information 992, which may include the gain coefficients described above, calculated by controller 901. That is, for example, controller 901 may create specific gain coefficients for the upper, lower, and rear hemispheres and provide these gain coefficients to gain adjuster 906 along with information indicating the signal to which each gain coefficient should be applied.

[0059] Using the virtual speaker signals y1', y2', ..., yk', speaker signal producer 908 produces output signals (e.g., output signals 781 and 782) for driving speakers (e.g., headphone speakers or other speakers). In an embodiment in which the speakers are headphone speakers, speaker signal producer 908 may perform conventional binaural rendering to produce the output signals. In an embodiment in which the speakers are not headphone speakers, speaker signal producer may perform conventional speaking panning to produce the output signals.

[0060] 10 is a block diagram of an audio rendering device 1000 according to some embodiments for performing the methods disclosed herein (e.g., audio renderer 751 may be implemented using audio rendering device 1000). As shown in FIG. 10, audio rendering device 1000 includes a processing circuit (PC) 1002 that may include one or more processors (P) 1055 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.), which may be co-sited in a single housing or in a single data center or may be geographically distributed (i.e., device 1000 may be a distributed computing device), and at least one network interface 1048, so that device 1000 can communicate with the network interface. The PC 1002 may comprise at least one network interface 1048, a transmitter (Tx) 1045, and a receiver (Rx) 1047 for enabling the network interface 1048 to transmit data to and receive data from other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 1048 is connected (directly or indirectly) (e.g., the network interface 1048 may be wirelessly connected to the network 110, in which case the network interface 1048 is connected in an antenna configuration), and a storage unit (a.k.a. "data storage system") 1008, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 1002 includes a programmable processor, a computer-readable storage medium (CRSM) 1042 may be provided. The CRSM 1042 stores a computer program (CP) 1043 comprising computer-readable instructions (CRI) 1044. The CRSM 1042 may be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, or a memory device (e.g., a random access memory, a flash memory).In some embodiments, the CRI 1044 of the computer program 1043, when executed by the PC 1002, is configured such that the CRI causes the audio-rendering device 1000 to perform the steps described herein (e.g., steps described herein with reference to flowcharts). In other embodiments, the audio-rendering device 1000 may be configured to perform the steps described herein without the need for code. That is, for example, the PC 1002 may simply consist of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.

[0061] Overview of Various Embodiments

[0062] A1. A method for rendering an audio element, the method including: determining a top gain value (G_top) for an upper part of an internal representation of the audio element based on L and T, where L is the vertical distance between a reference plane and a listening point and T is the vertical distance between the reference plane and a top point of the range of the audio element (e.g., point 501); and / or determining a bottom gain value (G_bottom) for a lower part of the internal representation of the audio element based on L and B, where B is the vertical distance between the reference plane and a bottom point of the range of the audio element (e.g., point 502).

[0063] A2. The method of embodiment A1, wherein T is the vertical distance between the top point of the selected portion of the range for the audio element and the reference plane, and B is the vertical distance between the bottom point of the selected portion of the range for the audio element and the reference plane.

[0064] A3. The method of embodiment A1 or A2, wherein an audio element has an original range, and the range of an audio element is a simplified range for the audio element that represents the original range from a listening point.

[0065] A4. The method of embodiment A1 or A2, wherein the audio elements are represented using a set of virtual speakers comprising a set of one or more upper virtual speakers located above the listening point and / or a set of one or more lower virtual speakers located below the listening point.

[0066] A5. The method of embodiment A4, wherein the set of upper virtual speakers includes a first upper virtual speaker, and the method further includes: creating a first upper virtual speaker signal y1 for the first upper virtual speaker; creating a gain-adjusted first upper virtual speaker signal y1', where y1' = g * y1, where g is a function of at least the upper gain value; and using y1' to render the audio element.

[0067] A6. The method of embodiment A5, wherein the set of virtual speakers comprises two or more sets of rear virtual speakers, the set of rear virtual speakers comprising a first top virtual speaker, and the method further includes determining a rear gain value (G_rear) for the set of rear virtual speakers, wherein g is a function of at least the top gain value (G_top) and the rear gain value (G_rear) (i.e., g = f(G_top, G_rear), for example, g = G_top * G_rear).

[0068] A7. The method of embodiment A4, wherein the set of lower virtual speakers includes a first lower virtual speaker, and the method further includes: creating a first lower virtual speaker signal y2 for the first lower virtual speaker; creating a gain-adjusted first lower virtual speaker signal y2', where y2' = g * y2, where g is a function of at least the lower gain value; and using y2' to render the audio element.

[0069] A8. The method of embodiment A7, wherein the set of virtual speakers comprises two or more sets of rear virtual speakers, the set of rear virtual speakers comprising a first bottom virtual speaker, and the method further comprises determining a rear gain value (G_rear) for the set of rear virtual speakers, wherein g is a function of at least the bottom gain value (G_bottom) and the rear gain value (G_rear) (i.e., g = f(G_bottom, G_rear), for example, g = G_bottom * G_rear).

[0070] B1. A computer program comprising instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform a method according to any one of the preceding embodiments.

[0071] B2. A carrier containing the computer program as described above, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.

[0072] D1. An audio rendering device configured to perform the method according to any one of the preceding embodiments.

[0073] D2. The audio rendering device of embodiment D1, wherein the audio rendering device comprises a memory and a processing circuit coupled to the memory.

[0074] While various embodiments have been described herein, it should be understood that these embodiments have been presented by way of example only, and not limitation. Thus, the breadth and scope of the present disclosure should not be limited by any of the exemplary embodiments described above. Moreover, unless otherwise indicated herein or clearly contradicted by context, any combination of the above-described objects in all possible variations thereof is encompassed by the present disclosure.

[0075] Additionally, while the processes described above and illustrated in the figures have been shown as a sequence of steps, this has been done for purposes of illustration only, and it is therefore contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.

[0076] "References" [1] MPEG-H 3D Audio, Clause 8.4.4.7: "Spreading" [2] MPEG-H 3D Audio, Clause 18.1: "Element Metadata Preprocessing" [3] MPEG-H 3D Audio, Clause 18.11: "Diffuseness Rendering" [4] EBU ADM Renderer Tech 3388, Article 7.3.6: “Divergence” [5] EBU ADM Renderer Tech 3388, Article 7.4: "Decorrelation Filters" [6] EBU ADM Renderer Tech 3388, Article 7.3.7: "Extent Panner" [7]Efficient HRTF-based Spatial Audio for Area and Volumetric Sources“, IEEE Transactions on Visualization and Computer Graphics 22(4):1-1 January 2016 [8] Patent Publication WO2020144062, “Efficient spatially-heterogeneous audio elements for Virtual Reality.” [9] Patent Publication WO2021180820, "Rendering of Audio Objects with a Complex Shape."

[10] International Patent Application No. PCT / EP2021 / 068833, “Seamless Rendering of Audio Elements with Both Interior and Exterior Representations,” filed July 7, 2021.

[11] M. Kronlachner, F. Zotter, “Spatial transformations for the enhancement of Ambisonic recordings”, ICSA2014

Claims

1. A method (600) for rendering an audio element (302), the method comprising: determining (s602) a top gain value G_top for the top part of the internal representation of the audio element based on L and T, where L is the vertical distance between a reference plane (590) and listening points (A1, A2, A3, A4, A5) and T is the vertical distance between the reference plane and the top point (501) of the range (410) of the audio element; and / or determining (s604) a bottom gain value G_bottom for a bottom part of the internal representation of the audio element based on L and B, where B is the vertical distance between the reference plane and a bottom point of the range of the audio element (502); The method (600) includes:

2. T is the vertical distance between the top point of the selected portion of the range for the audio element and the reference plane; B is the vertical distance between the bottom point of the selected portion of the range for the audio element and the reference plane; The method of claim 1.

3. 3. The method of claim 1, wherein the audio elements have an original range, and the range of the audio elements is a simplified range for the audio elements that represents the original range from a listening position.

4. 3. The method of claim 1, wherein the audio elements are represented using a set of virtual speakers, comprising a set of one or more upper virtual speakers located above a listening point and / or a set of one or more lower virtual speakers located below the listening point.

5. the set of upper virtual speakers comprises a first upper virtual speaker, and the method further comprises: producing a first upper virtual speaker signal y1 for the first upper virtual speaker; producing a gain adjusted first upper virtual speaker signal y1′, where y1′=g*y1, where g is a function of at least the upper gain value; using y1' to render said audio element; The method of claim 4 further comprising:

6. the set of virtual speakers comprises two or more sets of rear virtual speakers; the set of rear virtual speakers comprises the first upper virtual speaker; the method further comprising determining a rear gain value G_rear for the set of rear virtual speakers; g is a function of at least the top gain value G_top and the rear gain value G_rear; The method of claim 5.

7. g = G_top * G_rear, The method of claim 6.

8. wherein the set of lower virtual speakers comprises a first lower virtual speaker, and the method further comprises: producing a first lower virtual speaker signal y2 for the first lower virtual speaker; producing a gain adjusted first lower virtual speaker signal y2′, where y2′=g*y2, where g is a function of at least the lower gain value; using y2' to render said audio element; The method of claim 4 or 5, further comprising:

9. the set of virtual speakers comprises two or more sets of rear virtual speakers; the set of rear virtual speakers comprises the first lower virtual speaker; the method further comprising determining a rear gain value G_rear for the set of rear virtual speakers; g is a function of at least the bottom gain value G_bottom and the rear gain value G_rear; The method of claim 8.

10. g = G_bottom * G_rear, 10. The method of claim 9.

11. the method comprising determining G_top; Determining G_top is If L≧T, set G_top to 0; If (T-β is ≦L and L<T), then set G_top to (T-L) / β, where β describes the size of the fade region below the top point (501); or If L<(T-β), set G_top to 1 11. The method of claim 1, comprising:

12. the method comprising determining G_top; Determining G_top is If L≧(T+β), then set G_top to 0; if (T≦L and L<T+β), then set G_top to 1+((T−L) / β), where β describes the size of the fade region above the top point (501); or If L<T, set G_top to 1 11. The method of claim 1, comprising:

13. the method including determining G_bottom; Determining G_bottom can be done by If L≧B+β, set G_bottom to 1; If (B≦L and L<B+β), then set G_bottom to (L−B) / β, where β describes the size of the fade region above the bottom point (502); or If L<B, set G_bottom to 0 13. The method of any one of claims 1 to 12, comprising:

14. the method including determining G_bottom; Determining G_bottom can be done by If L≧B, set G_bottom to 1; if (B-β≦L and L<B), then set G_bottom to 1+((L-B) / β), where β describes the size of the fade area below the bottom point (502); set G_bottom to 1+((L-B) / β), else If L<(B-β), set G_bottom to 0 13. The method of any one of claims 1 to 12, comprising:

15. A computer program (1043) comprising instructions (1044) that, when executed by a processing circuit (1002) of an audio renderer (751, 1000), causes the audio renderer to perform the method of any one of claims 1 to 14.

16. 16. A carrier containing the computer program of claim 15, said carrier being one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1042).

17. An audio rendering device (751, 1000), comprising: determining a top gain value G_top for the top part of the internal representation of the audio element based on L and T, where L is the vertical distance between a reference plane (590) and the listening points (A1, A2, A3, A4, A5) and T is the vertical distance between said reference plane and the top point (501) of the range (410) of said audio element; and / or determining (s604) a bottom gain value G_bottom for a bottom part of the internal representation of the audio element based on L and B, where B is the vertical distance between the reference plane and a bottom point of the range of the audio element (502); An audio rendering device (751, 1000) configured to:

18. T is the vertical distance between the top point of the selected portion of the range for the audio element and the reference plane; B is the vertical distance between the bottom point of the selected portion of the range for the audio element and the reference plane; 18. An audio rendering device according to claim 17.

19. 19. The audio rendering apparatus of claim 17 or 18, wherein the audio elements have an original range, and the range of the audio elements is a simplified range for the audio elements that represents the original range from a listening position.

20. 19. The audio rendering apparatus of claim 17 or 18, wherein the audio elements are represented using a set of virtual speakers, comprising a set of one or more upper virtual speakers located above a listening point and / or a set of one or more lower virtual speakers located below the listening point.

21. the set of upper virtual speakers comprises a first upper virtual speaker, and the audio rendering device: producing a first upper virtual speaker signal y1 for the first upper virtual speaker; producing a gain adjusted first upper virtual speaker signal y1′, where y1′=g*y1, where g is a function of at least the upper gain value; using y1' to render said audio element; 21. The audio rendering device of claim 20, further configured to:

22. the set of virtual speakers comprises two or more sets of rear virtual speakers; the set of rear virtual speakers comprises the first upper virtual speaker; the audio rendering device is further configured to determine a rear gain value G_rear for the set of rear virtual speakers; g is a function of at least the top gain value G_top and the rear gain value G_rear; 22. An audio rendering device according to claim 21.

23. g = G_top * G_rear, 23. An audio rendering device according to claim 22.

24. wherein the set of lower virtual speakers comprises a first lower virtual speaker, and the method further comprises: producing a first lower virtual speaker signal y2 for the first lower virtual speaker; producing a gain adjusted first lower virtual speaker signal y2′, where y2′=g*y2, where g is a function of at least the lower gain value; using y2' to render said audio element; 22. An audio rendering device according to claim 20 or 21, further comprising:

25. the set of virtual speakers comprises two or more sets of rear virtual speakers; the set of rear virtual speakers comprises the first lower virtual speaker; the audio rendering device is further configured to determine a rear gain value G_rear for the set of rear virtual speakers; g is a function of at least the bottom gain value G_bottom and the rear gain value G_rear; 25. An audio rendering device according to claim 24.

26. g = G_bottom * G_rear, 26. An audio rendering device according to claim 25.

27. The audio rendering device comprises the following steps: if L≧T, setting G_top to 0; if (T-β≦L and L<T), then set G_top to (T-L) / β, where β describes the size of the fade area below said top point (501); or If L<(T-β), set G_top to 1. and determining G_top by executing a process including one of:

27. An audio rendering device according to any one of claims 17 to 26.

28. The audio rendering device comprises the following steps: if L≧(T+β), then setting G_top to 0; if (T≦L and L<T+β), then setting G_top to 1+((T−L) / β), where β describes the size of the fade area above said top point (501); or If L<T, set G_top to 1. and determining G_top by executing a process including one of:

27. An audio rendering device according to any one of claims 17 to 26.

29. The audio rendering device comprises the following steps: if L≧B+β, then setting G_bottom to 1; if (B≦L and L<B+β), then set G_bottom to (L−B) / β, where β describes the size of the fade area above the bottom point (502); or If L<B, set G_bottom to 0. [0047] The method was configured to determine G_bottom by performing a process including one of:

29. An audio rendering device according to any one of claims 17 to 28.

30. The audio rendering device comprises the following steps: if L≧B, then setting G_bottom to 1; if (B-β≦L and L<B), then set G_bottom to 1+((L-B) / β), where β describes the size of the fade area below the bottom point (502); set G_bottom to 1+((L-B) / β), else If L<(B-β), set G_bottom to 0. [0047] The method was configured to determine G_bottom by performing a process including one of:

29. An audio rendering device according to any one of claims 17 to 28.

31. 31. An audio rendering device according to any one of claims 17 to 30, wherein the audio rendering device comprises a memory and a processing circuit coupled to the memory.