Virtual audio mixing for 6DOF virtual environments

By adjusting audio levels and directions based on listener pose using superseding metadata, the method addresses the challenge of maintaining immersion and optimal audio experience in 6DoF environments, ensuring the audio experience aligns with creative intent.

WO2026050020A1PCT designated stage Publication Date: 2026-03-05DOLBY LABORATORIES LICENSING CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/042172
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-27
Filing Date
2025-08-15
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing audio rendering techniques for virtual environments fail to maintain immersion and optimal audio experience when listeners move within a 6DoF environment, as they either maintain a constant sound that mismatches visual cues or apply physics-based rules that deteriorate the listening experience.

Method used

A method that adjusts audio levels, directions, and propagation based on listener pose using superseding pose-dependent modification metadata, allowing creative intent to prevail over physics-based models, ensuring immersive and aesthetically pleasing audio experiences.

Benefits of technology

The method ensures that audio mixes adapt to listener movements in 6DoF environments, maintaining immersion and providing auditory cues while ensuring the audio experience remains optimal and aligned with the creator's intent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025042172_05032026_PF_FP_ABST
    Figure US2025042172_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a method of rendering audio for a virtual environment. The method comprises obtaining input of one or more audio objects and corresponding positional metadata for the one or more audio objects, the one or more audio objects and corresponding positional metadata characterizing an audio mix; obtaining superseding pose-dependent modification metadata describing how the audio mix is intended to vary based on listener pose of a listener in the virtual environment; obtaining listener pose information indicative of the listener pose in the virtual environment; and adjusting levels of the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information. The disclosure further relates to corresponding apparatus, computer programs, and computer-readable storage media.
Need to check novelty before this filing date? Find Prior Art

Description

D23013W001VIRTUAL AUDIO MIXING FOR 6DOF VIRTUAL ENVIRONMENTSCross-Reference to Related Applications

[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 687,568, filed on August 27, 2024, which is incorporated by reference herein in its entirety.Technical Field

[0002] The present disclosure relates to techniques for rendering audio for virtual environments, in particular 6 degrees of freedom (6DoF) virtual environments or 3DoF virtual environments. The present disclosure particularly relates to such techniques that provide for additional flexibility in handling audio mixes in relation to pose changes of a user.Background

[0003] Audio experiences today are typically mixed by a content creator for a specific static listening position. In this context, “mixing” involves taking various audio signals (typically channels) as individual inputs and adjusting the relative intensities (and possibly spatial positions) so that they meet the creator’s creative intent. The resulting adjusted combination is an “intended mix” from one specific listening position.

[0004] When the listener moves in a virtual 6DoF environment to a different viewing or listening position, two scenarios are feasible:

[0005] The sound remains constant. The listener experiences the high aesthetic quality audio mix (e.g., as intended by the content creator) but does not have any audio cue that they have moved. If there are visual cues that the user has moved (e.g., they are viewing a scene from a different viewing position) then this introduces a mismatch in perceptual cues, where the view has changed but the audio has not. This reduces the level of immersion.

[0006] The sound changes according to some (e.g., predetermined) rules, for example physics rules, such as default rules (e.g., physics rules) stored or otherwise implemented by an audio decoder or Tenderer. For example, if a listener moves away from an audio object, then it could become quieter, following some distance-dependent attenuation rule. If they turn their head, the audio object could pan around. This approach can simulate reality, and henceD23013W001 provide a level of immersion, but may result in less-than-optimal audio experience, since the “mix” is no longer ideal for the listening position. For example, for a virtual music concert, a listener may want to change their position to get a better view of a certain element of the virtual environment, such as a certain musician on stage, but applying realistic physics rules may result in deteriorated listening experience for the music content, which may not be intended nor desirable for the listener.

[0007] There is thus need for improved techniques for rendering immersive audio for a virtual environment. There is a particular need for such techniques that can maintain immersion while ensuring optimal audio experience.Summary

[0008] In view of this need, the present disclosure provides methods of rendering audio for a virtual environment, apparatus for rendering audio for a virtual environment, computer programs, and computer-readable storage media.

[0009] One aspect of the present disclosure relates to a method of rendering audio for a virtual environment (e.g., 6DoF virtual environment). The method may include obtaining input of one or more audio objects and corresponding positional metadata (or audio object metadata in general) for the one or more audio objects. The one or more audio objects and corresponding positional metadata may relate to (e.g., characterize) an audio mix (e.g., first audio mix). The method may further include obtaining superseding pose-dependent modification metadata describing how the audio mix (e.g., first audio mix) is intended to vary based on listener pose of a listener in the virtual environment. The audio mix and the superseding pose-dependent modification metadata may be extracted from a bitstream, for example. The superseding pose-dependent modification metadata may have been authored by a content creator, possibly specifically tailored to the audio mix. The method may further include obtaining listener pose information indicative of the listener pose in the virtual environment. The listener pose may include one or more of listener position and listener viewing direction in the virtual environment. The method may yet further include adjusting levels (audio levels) of the one or more audio objects of the audio mix (e.g., first audio mix) based on the superseding pose-dependent modification metadata and the listener pose information. The adjustments may be, for example, relative to the obtained audio mix or relative to an audio mix that is modified based on a default tenderer model (e.g., physics-D23013W001 based model). Adjusting the levels (audio levels) of the one or more audio objects of the first audio mix may result in a second audio mix different from the first audio mix.

[0010] Thereby, the proposed method ensures that creative intent on how an audio mix should sound is still matched for the resulting audio mix even when the listener can freely move in a 6D0F virtual environment, and in particular can move away from a sweet spot or reference listener position for which the audio mix had been intended. By enabling a modification of a default model (e.g., physics-based model) for attenuation, directivity, and / or propagation via the superseding pose-dependent modification metadata provided by the encoder side (in particular, by a content creator), which may take precedence over any default decoder-side or renderer-side model, the audio mix can be modified when the user moves in the virtual environment to have a plausible feel and to provide cues indicative of the listener’s movement in the virtual environment, thereby maintaining full immersion, but to still sound aesthetically pleasing. All this can be achieved for rendering that is fully left to decoder-side devices, enabled by metadata as defined above. Accordingly, the proposed method is applicable, to name few examples, to virtual concerts and other virtual music events where multiple audio objects at different positions, for example on a virtual stage, are provided, or to virtual cinemas and virtual home theaters.

[0011] In some embodiments, the method may further include deriving relative positions of the one or more audio objects and the listener in the virtual environment based on the positional metadata and the listener pose information. Then, the levels of the one or more audio objects may be adjusted based on the superseding pose-dependent modification metadata and the relative positions of the one or more audio objects and the listener in the virtual environment.

[0012] In some embodiments, the method may further include deriving a listener viewing direction in relation to the one or more audio objects based on the positional metadata and the listener pose information. Then, the levels of the one or more audio objects may be adjusted based on the listener viewing direction in relation to the one or more audio objects.

[0013] With this, the proposed method allows to modify attenuation rules that are physics-based or otherwise predefined at the decoder to, for example, ensure that artistic intent is met even when the listener moves relative to the audio objects of the audio mix. At the same time, auditory cues corresponding to the listener’ s movement can be provided, thereby improving immersion.D23013W001

[0014] In some embodiments, the method may further include adjusting directions of arrival for the one or more audio objects of the audio mix based on the superseding posedependent modification metadata and the listener pose information. For example, the directions of arrival may be adjusted for respective audio objects based on the listener position in the virtual environment, such as based on relative distances between the one or more audio objects and the listener position. Depending on circumstances, audio objects may appear at larger or smaller angular distances from each other than would be dictated by a physics-based model (e.g., a geometric or trigonometric model).

[0015] With this, the proposed method allows to modify directional rules that are physics-based or otherwise predefined at the decoder to, for example, ensure that artistic intent is met even when the listener moves relative to the audio objects of the audio mix. At the same time, auditory cues corresponding to the listener’ s movement can be provided, thereby improving immersion.

[0016] In some embodiments, the method may further include adjusting times of arrival or propagation delays for the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information. For example, the times of arrival or propagation delays may be adjusted for respective audio objects based on the listener position in the virtual environment, such as based on relative distances between the one or more audio objects and the listener position.

[0017] With this, the proposed method allows to modify sound propagation rules that are physics-based or otherwise predefined at the decoder to, for example, ensure that artistic intent is met even when the listener moves relative to the audio objects of the audio mix. For instance, the time that sound would normally need to travel to the listener’ s position in the virtual environment may be artificially reduced to better align visual and acoustic cues.

[0018] In some embodiments, the superseding pose-dependent modification metadata may include parameters indicative of: one or more of modifications to gains for the one or more audio objects based on the listener pose, modifications to apparent directions for the one or more audio objects based on the listener pose, and / or modifications to times of arrival or propagation delays for the one or more audio objects based on the listener pose. For example, all of the above may be modified based on listener position. In addition, at least the gains may be modified based on the listener viewing direction, resulting in attention-directed gains.

[0019] In some embodiments, the superseding pose-dependent modification metadata may include parameters indicative of one or more of: modifications to gains for the one orD23013W001 more audio objects based on positions of audio objects other than the one or more audio objects and modifications to apparent directions for the one or more audio objects based on the positions of the audio objects other than the one or more audio objects.

[0020] In some embodiments, the superseding pose-dependent modification metadata may include parameters indicative of one or more of: modifications to gains for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment; modifications to apparent directions for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment; and modifications to times of arrival or propagation delays for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment.

[0021] In some embodiments, the superseding pose-dependent modification metadata may further include a parameter indicative of a reverberation model to be used for the audio mix.

[0022] In some embodiments, the method may further include applying limits to gains for the one or more audio objects based on relative distances between the one or more audio objects and a listener position indicated by the listener pose information. Additionally, or alternatively, the method may further include applying limits to changes of gains for the one of more audio objects relative to a reference audio mix for the audio mix.

[0023] Accordingly, when starting out from a reference listening position or when applying custom rules for attenuation, directivity, and / or propagation, techniques according to the present disclosure can still ensure that gains are kept within certain limits, or that the resulting audio mix is kept within certain limits relative to the reference audio mix (or intended audio mix).

[0024] In some embodiments, the method may further include reading a flag indicative of whether the levels of the one or more audio objects shall be adjusted based on the superseding pose-dependent modification metadata or a default model.

[0025] In some embodiments, adjusting the level of a given one among the one or more audio objects of the audio mix may be further based on a type or class of the audio object. The same may apply to adjustments of directions of arrival and / or times of arrival (propagation delays). Thereby, a distinction may be made for example between audio objects of different importance, such as singers / guitars in a virtual concert on the one hand and crowd cheering on the other hand.D23013W001

[0026] In some embodiments, the superseding pose-dependent modification metadata may include different sets of parameters for different zones (spatial zones) of the virtual environment. Additionally, the method may further include obtaining input of different audio mixes for different zones of the virtual environment.

[0027] Thereby, multi-zone audio content, such as multi-stage concert events can be rendered in a way that meets artistic intent even when the user moves around in the 6D0F virtual environment and immersion is to be maintained.

[0028] In some embodiments, the method may further include obtaining audio mix metadata for the audio mix and for each of one or more reference listener poses in the virtual environment. The method may yet further include adjusting rendering characteristics of the one or more audio objects based on the audio mix metadata and the listener pose information. The reference listener poses may relate to reference listener positions and / or reference listener viewing directions, for example.

[0029] In some embodiments, the method may further include interpolating between reference listener poses for determining interpolated audio mix metadata for the listener pose in the virtual environment. For example, the interpolation may include linear interpolation between reference listener positions closest to the listener position.

[0030] According to another aspect, an apparatus for rendering audio for a virtual environment is provided. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform the methods or method steps outlined throughout the present disclosure.

[0031] According to a further aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., processor).

[0032] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a computing device (e.g., processor) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.

[0033] It should be noted that the methods and apparatus including its preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus outlined in the present disclosure may be arbitrarily combined. InD23013W001 particular, the features of the claims may be combined with one another in an arbitrary manner.

[0034] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa.Brief Description of the Drawings

[0035] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein:

[0036] Fig. 1A and Fig. IB show examples of audio tracks of different audio sources (audio objects),

[0037] Fig. 2 shows an example of basic spatial information for audio in a virtual environment,

[0038] Fig. 3A to Fig. 3C illustrate examples of different reference listener positions in the virtual environment according to embodiments of the disclosure,

[0039] Fig. 4 is a flowchart illustrating an example of a method of rendering audio for a virtual environment according to embodiments of the disclosure,

[0040] Fig. 5 is a flowchart illustrating an example implementation of steps of the method of Fig. 4, according to embodiments of the disclosure,

[0041] Fig. 6 is a flowchart illustrating an example of a method including additional optional steps for the method of Fig. 4, according to embodiments of the disclosure,

[0042] Fig. 7 is a flowchart illustrating another example of a method including additional optional steps for the method of Fig. 4, according to embodiments of the disclosure, and

[0043] Fig. 8 is a block diagram schematically illustrating an example of an apparatus suitable for implementing methods according to embodiments of the disclosure.D23013W001Detailed Description

[0044] In the following, example embodiments of the disclosure will be described with reference to the appended figures. Identical elements in the figures may be indicated by identical reference numbers, and repeated description thereof may be omitted.Introduction

[0045] Broadly speaking, the present disclosure relates to methods of enabling the immersive experience of the audio mix changing in accordance with listener pose changes, e.g., the listener changing position or orientation, while enabling multiple creative mixes or appropriately adjusted mixes to ensure that the audio experience is still aesthetically pleasing from at least a variety of listening positions.

[0046] Accordingly, embodiments and implementations of the present disclosure relate to audio which varies as the user moves within a virtual environment, such as a 6DoF environment. Without intended limitation, use cases may include:

[0047] Rendering audio for viewing movies in a “virtual cinema” or “virtual home theatre”, where a user wears a VR headset with headphones which simulates the experience of viewing a movie on a virtual screen. This invention allows the sound to vary as the user moves around, just as it would in real life, while still allowing for creative mixing to ensure a good audio experience.

[0048] Rendering audio for viewing virtual concerts or other performances, either in a VR headset or on a conventional flat screen. As the user navigates around the virtual environment, the sound is adapted to simulate how it would sound in a real environment, while still allowing for creative mixing to ensure a good audio experience.

[0049] Embodiments and implementations of the present disclosure may include providing (1) audio channels according to, for example, audio objects, as well as (2) metadata describing audio objects as well as desired audio mix(es), for example in the form of instructions on how audio mixes should change in reaction to pose changes of the listener. Multiple sets of metadata can be included, describing a desired audio mix corresponding to, for example, as many listening positions (e.g., reference listening positions) or as many schemes for varying audio mixes in dependence on listener pose change as desired. It is understood that the provided metadata may have precedence over any decoder-side or renderer-side default models (e.g., physics models or physics-inspired models) for how audio mixes should change in reaction to pose changes of the listener.D23013W001

[0050] While the concept of audio objects is well known in the art, the present disclosure leverages means for manipulating audio objects (e.g., adjusting level, directivity, delay, and / or apparent width, etc.) for providing audio mix(es) corresponding to, for example, multiple reference listening positions and / or multiple schemes for varying audio mixes in dependence on listener pose changes.Nomenclature and Definitions

[0051] In the context of the present disclosure, the following nomenclature and definitions will be used for audio channels and audio objects.Audio channels:

[0052] Audio channels correspond to recordings from individual sound sources (e.g., musical instruments, etc.) or audio objects, but may likewise apply to synthesized sounds and effects. In any case, each audio channel may correspond to a given sound source, including real sound sources and synthetic sound sources. Examples of audio channels (audio tracks) corresponding to different sound sources or audio objects are illustrated in Fig. 1A and Fig. IB, where Fig. 1A shows an example of an audio track of a guitar and Fig. IB shows an example of an audio track of a singer, for example as audio objects in a virtual convert or other virtual music event.Audio object metadata:

[0053] Each audio object is understood to be associated with (e.g., correspond to, point to, etc.) a respective audio channel. The audio object metadata describes the spatial location where the audio is emitted from. At a minimum, an audio object can be described by three spatial world coordinates, for example (X, Y, Z) coordinates. By default, such as object would be considered to emit sound isotopically (equally in all directions) from a single point.

[0054] Audio objects can also have spatial size. For example, a piano or guitar emits sound from much more than a single point. Also, a background sound like a river or waterfall may emit from a very large area. The spatial size may be approximated by defining a three- dimensional oval (e.g., ellipsoid) with Rx, Ry, Rz parameters indicating respective radii of the oval. More complex shapes may be defined as well in the metadata, such as donut-shaped objects suitable for an audience or crowd.

[0055] Audio objects can also emit sound anisotropically (e.g., not equally in all directions). This anisotropy may be indicated by one or more parameters or flags in the audio object metadata.D23013W001

[0056] In some implementations, spherical harmonics may be suitable for describing the 3D shape and size of audio objects and their emitted sound field, with corresponding parameters in the audio object metadata.

[0057] Fig. 2 shows an example of basic spatial information for audio in a virtual environment 10. In this virtual environment 10, which may relate to a virtual concert or other virtual music event, audio objects corresponding to a singer 20A and a guitarist (guitar) 20B that are placed on a virtual stage are provided. Spatial information (or in general, audio object metadata) for either audio object, as shown in the figure, may include spatial coordinates (e.g., X, Y, Z coordinates) of the audio object, and optionally one or more of a channel number and an audio object ID (e.g., Guitar, Singer, etc.).Audio mix metadata:

[0058] Audio mix metadata according to embodiments of the disclosure may include any of the optional metadata categories described below.

[0059] Global volume mix metadata may be included to describe the base physics for how sound intensity propagates in the virtual environment. This may include a function describing how the sound intensity of objects decreases as a function of distance D (e.g., in m) to the audio source. For example, the sound intensity may change as a function of V = c0w'lere co>c c2, andc3arecarried as metadata. The default values may correspond to isotropic sound propagation in air, with c0= 0, c2= 1, and c3= 0. Additionally, or alternatively, a limiting function may be defined that imposes limits on how much the level of each object may vary from different listening positions. The limiting function may be, for example, a sigmoidal function with limits on both decreasing intensity and increasing intensity.

[0060] Global timing mix metadata may be included to describe the base physics for how quickly audio propagates in the virtual environment. This amounts to a representation of the speed with which sound travels in the virtual environment. The delay T that an audio event is heard at distance D (e.g., in m) may be T = D / V , where V may be given by the metadata. If this metadata is not explicitly provided, a default value may be assumed for the speed of sound, such as 343.2 m / s (speed of sound at 20°C in dry air), for example.

[0061] Reference listening position metadata may be included to define a number of reference listening positions. For each reference listening position, the location corresponding to the audio mix metadata may be described in suitable coordinates (e.g., X, Y, Z world coordinates). Further, for each reference listening position, an assumption may be made thatD23013W001 the direction is facing directly at the spatial position of the audio object. However, additional metadata may be optionally included for a specific direction.

[0062] Further, multiplicative offsets for the physics-based volumes for each of the audio channels may be included in the reference listening position metadata. This is intended as an offset to the intensity of each object that would have been calculated according only to the base global physics model. For example, if the value is one, the volume of the channel will be determined only by the physics. If the value is greater than one, the volume of the channel will be louder, and below one it may be quieter. These offsets can be creatively adjusted to form a pleasing mix even when moving away from the reference listening position.

[0063] Optionally, the reference listening position metadata may include directional offsets from the physics-based directions for each audio channel. For example, if the listening position is very far away from several nearby audio sources, according to the base physics the listener would not be able to distinguish the spatial directions of the different audio sources — they would all appear to come from nearly the same direction. The directional offsets can allow a creator to specify to enhance the directional component of the audio object beyond what would be predicted by physics. For example, a value of 1 would reproduce the direction predicted by physics, while a value greater than 1 would enhance the spatial direction of the audio object. In this example, if physics predicted the audio object to reach the listener at 1 degree from the direction they are facing, a directional offset higher than 1, (e.g., 10) could increase this to a larger angle from the viewing direction (e.g., 10 degrees). The resulting value could be modified by a roll off curve to limit the direction to between 0 degrees and 90 degrees.

[0064] Further optionally, the reference listening position metadata may include timing offsets from the physics-based timing for each audio channel. For example, if the listening position is very far away from the audio object, then physics may dictate that the sound should take a noticeable time to propagate from the source of the audio object to the listener position. The timing offsets allow a multiplicative scale to the physics-based time offset, so that the delay can be reduced or increased as desired to achieve the desired aesthetic mix. For example, a value of 1.0 would reproduce the delay determined by physics, while a value of 0 would indicate zero delay. A value of >1.0 would further increase the delay.

[0065] Further optionally, the reference listening position metadata may include multiplicative offsets for gains and / or directional offsets based on positions of objects in the virtual environment. These objects may relate to (other) audio objects and / or to any ofD23013W001 occluding, partially occluding, or reflecting elements in the virtual environment. For instance, some audio object could be positioned between the listening position and the audio object position, which would occlude (or shadow) the sound source from the listener, and could thus reduce the level. Another example is the presence of a virtual wall, floor, or ceiling, in the virtual environment, which can cause reflections of the sound source. In this context, the aforementioned multiplicative offsets for gains and / or directional offsets may relate to metadata that allows for a blending of “physics-based” reflections and occlusions with a “purely aesthetic mix”, where such interactions of sound with the virtual environment may be ignored completely (or alternatively, may even be exaggerated) in order to achieve a preferred creative intent. The default model for treating objects in the virtual environment in relation to sound sources may be for example based on “acoustic ray tracing” models which derive the physics-based audio mix based on reflections and occlusions, as the skilled person will appreciate. Metadata as proposed here may effectively control how much or how strongly these physics-based interactions affect the final audio mix.

[0066] Yet further optionally, the reference position metadata may include control parameters for controlling whether or not, or to which degree, reverberation adheres to a physics-based model (e.g., default model) or to an aesthetics-based (e.g., creative intentbased) model. For example, as a listener moves from outdoors to an indoor brick room, a physics-based acoustic renderer may produce different sound due to reverberation. On the other hand, a content creator may want to have the identical sound in both an outdoor venue or the small brick room. This intent of the audio experience can be signaled by metadata along with other physics rules. An example for such metadata could be a “room reverberation” value or other reverberation-related parameter or flag, which scales between a full room reverberation model (e.g., a physics-based model) and a no room reverberation model.

[0067] Examples of reference listening positions are illustrated in Fig. 3A, Fig. 3B, and Fig. 3C.

[0068] Fig. 3A shows an example reference listening position (Listening Position 0) of a listener 30 that stands directly in front of a stage with a singer 20A and a guitarist (guitar) 20B. The singer’s 20 A and guitar’s 20B positions may be given in suitable coordinates, such as XYZ coordinates, for example relative to the reference listening position. Further, volume offsets may be provided for this reference listener position.Example VolumeOffsets = [1, 1.1] would make the singer (channel 1, 2ndnumber of the VolumeOffsets) slightly louder. Yet further, directional offsets may be provided for thisD23013W001 reference position as well. For the example of Listening Position 0 however, unit directional offsets, DirectionalOffsets = [1, 1] may be appropriate, i.e., no change in direction of arrival compared to what is predicted by the global physics model.

[0069] Fig. 3B shows an example reference listening position (Listening Position 1) of the listener 30 that stands on the stage, behind the singer 20A and the guitarist (guitar) 20B. Again, the singer’s 20A and guitar’s 20B positions may be given in suitable coordinates, such as XYZ coordinates, for example relative to the reference listening position. For this reference listening position, unit volume offsets and directional offsets may be appropriate, VolumeOffsets = [1, 1], DirectionalOffsets = [1 ,1].

[0070] Fig. 3C shows an example reference listening position (Listening Position 2) of the listener 30 that stands far away from the stage, slightly towards the left. Again, the singer’s 20A and guitar’s 20B positions may be given in suitable coordinates, such as XYZ coordinates, for example relative to the reference listening position. For this reference listening position, volume offsets may be appropriate that make both channels louder (not necessarily in equal measure) than what pure physics would dictate for this distance, for example VolumeOffsets = [1.4, 1.5]. Moreover, directional offsets may be appropriate that increase the separation of the audio objects (not necessarily in equal measure) compared to what pure physics would dictate for this distance, for example DirectionalOffsets = [1.4, 1.4]. Authoring Metadata

[0071] To author metadata (e.g., superseding pose-dependent modification metadata, audio mix metadata, reference listening position metadata) according to the present disclosure, a user interface may be provided to the content creator allowing them to freely navigate around in the virtual environment. At any point they may be enabled to select a “reference listening position” and adjust the mix to their taste. Alternatively, or additionally, they may be enabled to specify a model for how the audio mix (e.g., levels, directions of arrival, times of arrival, etc.) shall change when the user moves in the virtual environment.

[0072] A corresponding encoding process, for example by an encoder included in or coupled to a content creation system may generate audio content that comprises both audio objects and corresponding metadata (e.g., superseding pose-dependent modification metadata, audio mix metadata, reference listening position metadata). The encoder may be used by the content creator to generate the audio content, including authoring the metadata.

[0073] The audio mix and the corresponding metadata (e.g., superseding posedependent modification metadata) may be provided in the form of a dedicated format (e.g.,D23013W001 content creation format) that includes the audio mix and the metadata. This format may be an output format of content creation / authoring tools used for authoring the audio content. It may include, in addition to the audio mix and the superseding pose-dependent modification metadata, any metadata that is to be delivered to end users via a bitstream. A data file according to the dedicated format may be used by the encoder to generate a bitstream for distribution. By definition of such dedicated format, the authored audio mixes and their superseding pose-dependent modification metadata can be easily exchanged, for example among content creators and between content creators and end users.Distribution of Metadata

[0074] Metadata (e.g., superseding pose-dependent modification metadata, audio mix metadata, reference listening position metadata) according to the present disclosure will be packaged into a format suitable for distribution, and transmitted, for example as a side channel, with corresponding audio (and possibly video) information. As such, the metadata may be part of a corresponding bitstream, for example. It may be provided together with the corresponding audio objects (e.g., audio mixes). As described above, the bitstream may be generated by an encoder that encodes a data file of the aforementioned dedicated format. Playback

[0075] At playback, having available the metadata (e.g., superseding pose-dependent modification metadata, audio mix metadata, reference listening position metadata) as described above opens up the possibility for providing high quality audio mixes, possibly in close alignment with artistic intent, even in a 6D0F environment in which the listener can freely move around. For example, it may be the case that for an audio mix that has been created for a particular listener pose (e.g., location and / or orientation), modification that would be dictated by physics (or an approximation thereof) or any default model when the listener moves around in the 6D0F environment, such as changed levels, directions of arrival, times of arrival, etc. would lead to a modified audio mix that is (despite being in accordance with physics) less acoustically pleasing. Techniques according to the present disclosure allow to modify the rules according to which an audio mix is changed, for example relative to a reference audio mix, when the user moves in the 6D0F environment, to provide an acoustically pleasing audio mix while still providing plausible auditory cues of the user’s movement.

[0076] While the present disclosure may make frequent reference to the listener moving in a 6DoF virtual environment, it is understood that techniques according to embodiments of the disclosure may likewise apply to 3DoF use cases.D23013W001

[0077] Fig. 4 illustrates one example of a method 400 of rendering audio for a virtual environment (e.g., 6DoF environment, or 3DoF environment) for playback that leverages concepts according to the present disclosure. Method 400 comprises steps S410 through S470, of which any or all of steps S440, S460, and S470 may be optional steps. It is understood that the steps described herein are to be seen as processes that are ordered only to the extent that processing results or other outputs from one process may be required for performing another process. Apart from that, the steps disclosed herein may be performed in any order and / or in parallel.

[0078] At step S410, input of one or more audio objects and corresponding positional metadata (e.g., audio object metadata and / or basic spatial information as described above) for the one or more audio objects is obtained. These one or more audio objects and corresponding positional metadata relate to (e.g., characterize) an audio mix. The audio mix may for example be a reference audio mix (intended audio mix) as generated by a content creator for a particular listener pose (e.g., reference listener pose). The (reference) audio mix may include audio levels for the audio objects as intended by the content creator.

[0079] The audio mix may be extracted from a bitstream, for example.

[0080] At step S420, superseding pose-dependent modification metadata is obtained (e.g., received as input). The superseding pose-dependent modification metadata may relate to the audio mix metadata as described above, including for example global volume mix metadata and / or global timing mix metadata. Additionally, or alternatively, the superseding pose-dependent modification metadata may relate to or include one or more of multiplicative offsets, directional offsets, and / or timing offsets, possibly on a per-object basis, as described above. In any case, the superseding pose-dependent modification metadata describes how the audio mix is intended to vary based on listener pose of a listener in the virtual environment.

[0081] The superseding pose-dependent modification metadata may be extracted from a bitstream, for example. It may be provided by a content creator specifically for (and in conjunction with) the audio mix obtained at step S410. Importantly, the superseding posedependent modification metadata may be different from a default model (e.g., physics model or physics-inspired model) that is in place at the decoder or tenderer for describing how the audio mix varies based on listener pose. The superseding pose-dependent modification metadata may be specifically provided together with the audio mix, thereby ensuring full creative control over how the audio mix changes when the listener moves in the virtual environment.

[0082] The superseding pose-dependent modification metadata may includeD23013W001 parameters indicative of one or more of modifications to gains for the one or more audio objects based on the listener pose (e.g., in the form of volume offsets (multiplicative offsets) as described above, and / or in the form of a modified propagation model or modified physics model (e.g., global volume mix metadata), for example via parameters c0, c1?c2, and c3, or via other suitable parameters), modifications to apparent directions for the one or more audio objects based on the listener pose (e.g., in the form of directional offsets as described above), and / or modifications to times of arrival or propagation delays for the one or more audio objects based on the listener pose (e.g., in the form of timing offsets or global timing mix metadata, as described above). For example, all of the above may be modifications that are based on listener position. In addition, at least the gains may be modified based on the listener viewing direction, resulting in attention-directed gains as described elsewhere in the disclosure.

[0083] Further, the superseding pose-dependent modification metadata may include parameters indicative of modifications to gains for the one or more audio objects based on positions of audio objects other than the one or more audio objects and / or indicative of modifications to apparent directions for the one or more audio objects based on the positions of the audio objects other than the one or more audio objects.

[0084] Yet further, the superseding pose-dependent modification metadata may include parameters indicative of one or more of modifications to gains for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment, modifications to apparent directions for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment, and modifications to times of arrival or propagation delays for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment.

[0085] These optional elements of the superseding pose-dependent modification metadata allow a content creator to specify how objects (e.g., sound sources or occluding / partially occluding / reflecting elements) in the virtual environment affect the sound mix. For example, for some object between the listener position and the position of a given audio object, the object could occlude the audio object in relation to the listener, and could thus reduce the audio level for the given audio object. As another example, a virtual wall, floor, or ceiling in the virtual environment can cause reflections of sound from the audio object. The superseding pose-dependent modification metadata not necessarily provides a specific model for determining how any intervening objects would adjust the physics-basedD23013W001 audio mix, but may instead include information (e.g., parameters, flags, etc.) for blending a physics-based model (e.g., physics-based reflections, occlusions, and / or shadowing, for example as per a default model) with an intended mix (e.g., a purely aesthetic mix) where such interactions of the sound from the audio object with the virtual environment are ignored completely or at least partially, in order to achieve a preferred creative intent.

[0086] At step S430, listener pose information indicative of the listener pose in the virtual environment is obtained (e.g., received as input). The listener pose may include (or relate to) one or more of listener position and listener viewing direction in the virtual environment.

[0087] At step S440, which is an optional step, a flag indicative of whether the levels (audio levels, e.g., intensities, gains) of the one or more audio objects shall be adjusted based on the superseding pose-dependent modification metadata or based on a default model is read. The flag may be read for example from the bitstream, allowing a content creator to specify whether a physics-based mix or a creative mix (or an in-between version thereof) shall be used. Alternatively, the flag may be user- selected, allowing the user to select if they want to use a physics-based or a creative mix (or an in-between version thereof).

[0088] At step S450, levels of the one or more audio objects of the audio mix are adjusted based on the superseding pose-dependent modification metadata and the listener pose information. Example details of this step will be described below with reference to Fig. 5. For example, the levels may be adjusted for respective audio objects based on the listener position in the virtual environment, such as based on relative distances between the one or more audio objects and the listener position, for example.

[0089] In any case, adjusting the level of a given one among the one or more audio objects of the audio mix may be further based on a type or class of the audio object.

[0090] It is also understood that the level adjustment of step S450 may be performed at any point of the Tenderer chain. For instance, levels of the initial audio mix may be adjusted, levels may be adjusted during the actual rendering process, or as the case may be, if accessible, levels may also be adjusted after the rendering. The same holds true for the remaining adjustment operations described throughout the application, where applicable.

[0091] At step S460, which is an optional step, directions of arrival for the one or more audio objects of the audio mix are adjusted based on the superseding pose-dependent modification metadata and the listener pose information. Example details of this step will be described below with reference to Fig. 5. For example, the directions of arrival may be adjusted for respective audio objects based on the listener position in the virtual environment,D23013W001 such as based on relative distances between the one or more audio objects and the listener position, for example.

[0092] Also here, adjusting the direction of arrival of a given one among the one or more audio objects of the audio mix may be further based on the type or class of the audio object.

[0093] At step S470, which is an optional step, times of arrival or propagation delays for the one or more audio objects of the audio mix are adjusted based on the superseding pose-dependent modification metadata and the listener pose information. Example details of this step will be described below with reference to Fig. 5. For example, the times of arrival or propagation delays may be adjusted for respective audio objects based on the listener position in the virtual environment, such as based on relative distances between the one or more audio objects and the listener position, for example.

[0094] Also here, adjusting the times of arrival of a given one among the one or more audio objects of the audio mix may be further based on the type or class of the audio object.

[0095] As a further optional step (not shown in Fig. 5), gains for the one or more audio objects may be modified based on (parameters of) the superseding pose-dependent modification metadata and positions of audio objects other than the one or more audio objects, and / or apparent directions for the one or more audio objects may be modified based on (parameters of) the superseding pose-dependent modification metadata and positions of the audio objects other than the one or more audio objects. This may be done in accordance with the considerations presented above in relation to the elements of the superseding posedependent modification metadata.

[0096] As yet another optional step (not shown in Fig. 5), one or more of gains for the one or more audio objects, apparent directions for the one or more audio objects, and times of arrival (or propagation delays) for the one or more audio objects may be modified based on (parameters of) the superseding pose-dependent modification metadata and positions of occluding, partially occluding, or reflecting elements in the virtual environment. This may be done in accordance with the considerations presented above in relation to the elements of the superseding pose-dependent modification metadata.

[0097] Depending on circumstances, the above method may be adapted to account for different listening zones in the virtual environment. These different listening zones may relate to different stages in a virtual music festival with multiple stages or to an entry area thereof, for example.D23013W001

[0098] To account for the multiple listening zones, the superseding pose-dependent modification metadata (e.g., obtained at step S420 of method 400) may include different sets of parameters for different zones (listening zones) of the virtual environment. Further, different audio mixes may be obtained (e.g., at step S410) for different zones of the virtual environment. Then, depending on which zone the listener is currently located in, a respective audio mix may be rendered, wherein characteristics of the audio mix (e.g., gains, directions, and / or times of arrival, as described above) are adjusted depending on the respective set of parameters of the superseding pose-dependent modification metadata.

[0099] Fig. 5 illustrates an example method 500 that may be performed in the context of step S450, step S460, and / or step S470 of method 400. Method 500 comprises steps S510 and S520, of which either or both may be optional.[000100] At step S510. relative positions of the one or more audio objects and the listener in the virtual environment are derived based on the positional metadata (including, for example, coordinates of the audio objects) and the listener pose information (including, for example, the listener location). The relative positions may include for example relative distances from the listener to the audio objects. Deriving the relative positions may be done by conventional trigonometric considerations, applying a suitable distance metric. Having derived the relative positions of the one or more audio objects and the listener in the virtual environment, the levels of the one or more audio objects can be adjusted based on the superseding pose-dependent modification metadata and these relative positions. Additionally, one or more of directions of arrival and / or times of arrival of sound from the audio objects may be adjusted based on the superseding pose-dependent modification metadata and the relative positions.[000101] For example, it may be preferred that for a listener in a virtual music event or concert, the audio level associated with audio elements located on a stage (e.g., singer, guitar) falls off slower when moving away from the stage than would be dictated by a physics-based model.[000102] To this end, attenuations (attenuation gains) may be determined for the audio objects based on the relative positions (e.g., relative distances) using a predefined physicsbased model, followed by a modification of these attenuations (or gains) using multiplicative offsets for respective audio objects as included in the superseding pose-dependent modification metadata. Alternatively, a modified global model for attenuation, as for example defined by global volume mix metadata as included in the superseding pose-dependent modification metadata may be used for determining attenuations (attenuation gains) for theD23013W001 audio objects based on the relative positions (e.g., relative distances). Depending on circumstances, also a combination of both modification mechanisms may be used.[000103] Depending on circumstances, the listener position in relation to the one or more audio objects may also be used as basis for adjusting directions of arrival for the audio objects and / or adjusting times of arrival for the audio objects. This may be done in analogy to the above, for example using directional offsets and / or timing offsets as described above, or global timing mix metadata as described above.[000104] In a subsequent step, the direction that the listener is facing may be taken into account and the volumes for the left and right ears may be raised or lowered according to the listener’s viewing direction, in accordance with conventional techniques.[000105] Independently thereof, gains for the audio objects as such may be adjusted based on the listener’s viewing direction. For example, audio objects may be attenuated more than would be dictated by the physics-based model when the listener faces away from or not directly towards the audio objects. This technique may be used accordingly to implement attention-directed gains.[000106] In line with the above, at step S520, a listener viewing direction in relation to the one or more audio objects (e.g., angles between (absolute) listener viewing direction and viewing directions towards respective audio objects) is derived based on the positional metadata (including, for example, coordinates of the audio objects) and the listener pose information (including, for example, the listener location and the listener viewing direction). Again, this may be done by conventional trigonometric considerations, applying suitable distance and angular metrics. Having derived the listener viewing direction in relation to the one or more audio objects, the levels of the one or more audio objects can be adjusted based on the superseding pose-dependent modification metadata and these relative positions.[000107] For example, the listener may be located in a central area of a virtual multistage music event with multiple audio elements associated with respective stages. Then, gains for audio objects in a (possibly narrow) sector around the listener’s current viewing direction may have unit gains or even increased gains, with noticeable roll off at the sector’s edges. Further, audio objects from which the listener is facing away may have reduced (e.g., very small) gains. This would allow the listener to selectively listen to different stages by directing their view towards respective stages in the virtual environment, for example for choosing a stage area to enter from the central area.D23013W001[000108] Depending on circumstances, the listener viewing direction in relation to the one or more audio objects may also be used as basis for adjusting directions of arrival for the audio objects and / or adjusting times of arrival for the audio objects.[000109] Fig. 6 illustrates an example method 600 that includes additional steps for method 400 of Fig. 4. Method 600 comprises steps S610 and S620, of which either or both may be optional.[000110] At step S610, limits are applied to gains for the one or more audio objects based on relative distances between the one or more audio objects and a listener position indicated by the listener pose information.[000111] At step S620, limits are applied to changes of gains for the one of more audio objects relative to a reference audio mix for the audio mix.[000112] Accordingly, when starting out from a reference listening position or when applying custom rules for attenuation, directivity, and / or propagation, techniques according to the present disclosure can still ensure that gains are kept within certain limits, or that the resulting audio mix is kept within certain limits relative to the reference audio mix.[000113] As an optional step (not shown in Fig. 6), a reverberation model to be used for the audio mix (e.g., a physics-based or default model as opposed to a customized or aesthetically-driven model) may be determined (e.g., selected) based on information (e.g., a parameter or flag) included in the superseding pose-dependent modification metadata. As an example, there may be a blend between different reverberation models based on said information. This determination or blend may be done in accordance with the considerations presented above in relation to the elements of the superseding pose-dependent modification metadata.[000114] Techniques described above relate to modifying characteristics or behavior (e.g., levels / gains, directions of arrival, and / or times of arrival, and possibly reverberation) of audio objects in an audio mix compared to or beyond what would be predicted by a physicsbased model. In combination with this, or alternatively as a stand-alone method, different audio mixes (including audio objects and audio object metadata) for different reference listener positions in the virtual environment may be defined, optionally using interpolation between the reference listener positions. Thereby, it can be ensured that at least for the specified reference listener positions, the audio mix that is eventually rendered exactly matches artistic intent. On the other hand, by allowing the listener to move away from or between reference listener positions, a feel of immersion may be created.D23013W001[000115] At playback, the spatial mix metadata may be interpolated to create a single mix depending on the current listening position. This allows the listener to smoothly move around the virtual environment and have the mix continually updated according to their position. For example, the current spatial position may be first interpolated from the nearest reference listening positions using tetrahedral interpolation, which is well known in the art. The relative weights to the nearest spatial positions may then be used to interpolate the volume and direction components of the audio channels, wherein the relative weights may be based on distances to respective reference listening positions.[000116] Fig. 7 illustrates an example method 700 in line with the above that includes additional steps for method 400 of Fig. 4. Method 700 comprises steps S710 and S720, of which either or both may be optional.[000117] At step S710. audio mix metadata for the audio mix and for each of one or more reference listener poses in the virtual environment is obtained (e.g., received as input). The reference listener poses may relate to reference listener positions, for example.[000118] At step S720, rendering characteristics of the one or more audio objects (e.g., gains / levels, directions of arrival, and / or times of arrival) are adjusted based on the audio mix metadata and the listener pose information.[000119] To this end, method 700 may further comprise interpolating between reference listener poses (e.g., reference listener positions) to determine interpolated audio mix metadata for the listener pose in the virtual environment. For example, the interpolation may include linear interpolation between (two or more) reference listener positions closest to the listener position.Example Implementations[000120] Non-limiting example implementations of the present disclosure in line with the above may be as follows.Example 1:[000121] A method of rendering audio for a virtual environment, comprising: (a) reading one or more audio objects and corresponding positional metadata, (b) obtaining the position of the listener, (c) reading global metadata describing how the mix varies based on listener position, and (d) adjusting the levels of each audio object depending on the relative positions of the audio object and the listener and the listener position metadata. Operation (c) may include: (i) imposing a limit on gains depending on distance (e.g., soft sigmoid or clamp), (ii) imposing a limit on changes of gains relative to original mix, (hi) using aD23013W001 parameterized model, (iv) reading a flag indicating whether to modify based on physics or reference mix, and / or (v) vary processing by the type or class of object.Example 2:[000122] The method of Example 1, further comprising: (a) reading one or more audio mix metadata corresponding to one or more reference listening positions, (b) interpolating audio mix metadata from the one or more reference listening positions based on the listening position of the listener to obtain an interpolated audio mix metadata, and (c) adjusting the audio characteristics of each audio object depending on the relative positions of the audio object and the listener and the interpolated audio mix metadata.Example 3:[000123] A method of rendering audio for a virtual environment, comprising: (a) reading one or more audio objects and corresponding directional metadata, (b) obtaining the direction (e.g., looking / viewing direction) of the listener, (c) reading global metadata describing how the mix varies based on listener direction, and (d) adjusting the levels of each audio object depending on the relative orientations of the audio object and the listener and the listener direction metadata. Operation (c) may include: (i) imposing a limit on gains depending on direction (e.g., soft sigmoid or clamp), (ii) imposing a limit on changes of gains relative to original mix, (iii) using a parameterized model, (iv) reading a flag indicating whether to modify based on physics or reference mix, and / or (v) vary processing by the type or class of object.Example 4:[000124] The method of Example 3, further comprising: (a) reading one or more audio mix metadata corresponding to one or more reference listening positions, (b) interpolating audio mix metadata from the one or more reference listening positions based on the listening position of the listener to obtain an interpolated audio mix metadata, and (c) adjusting the audio characteristics of each audio object depending on the relative orientations of the audio object and the listener and the interpolated audio mix metadata.Example 5:[000125] The method of any of the above examples, including attention-directed gains based on looking direction of the listener. For example, from some location “ZoneA” which is outside of a central performance area of a virtual music event, the behavior could be different from a second location “ZoneB”, which is a central performance area. In ZoneA there could be strong attention-directed level gains, which could fall off when approachingD23013W001ZoneB. The global metadata from Example 1 to Example 4 could vary for different position zones.Example 6:[000126] This example relates to a virtual music event or virtual music concert where multiple performances are divided up into zones. Each zone may have different mixes (e.g., Dolby Atmos® mixes) and corresponding metadata / physical rules. There may be additional zones outside of the key performance zones with yet a different set of rules.Example 7:[000127] In this example, modifications may be made based on listener position, in accordance with techniques described above. This may relate to adjusting levels of arrival, adjusting directions of arrival, and / or adjusting times of arrival. Adjustments may be made relative to what is predicted or dictated by a global physics model (e.g., decoder-implemented physics model).Example 8:[000128] In this example, modifications may be made based on listener look direction and position. This may relate to adjusting levels (and optionally, direction and / or time of arrival). Adjustments may be made relative to what is predicted or dictated by a global physics model (e.g., decoder-implemented physics model).Apparatus[000129] Finally, the present disclosure likewise relates to an apparatus (e.g., computer- implemented apparatus or apparatus having processing capability in general) for performing or implementing methods and techniques described throughout the present disclosure. For example, this apparatus may relate to a device including an audio encoder or a device including an audio decoder. The audio decoder may be part of or coupled to a virtual reality (VR), augmented reality (AR), or computer-mediated reality device (such as VR / AR goggles), for example. The encoder may be part of or coupled to a content creation system used for creating audio content including audio objects and superseding pose -dependent modification metadata, for eventual decoding and rendering by, for example, a decoder included in or coupled to a VR, AR, or computer-mediated reality device.[000130] Fig. 8 shows an example of such apparatus 800. Apparatus 800 comprises a processor 810 and a memory 820 coupled to the processor 810. The memory 820 may store instructions for the processor 810. The processor 810 may also receive, among others, suitable input data 830 (e.g., audio objects, positional metadata, superseding pose-dependentD23013W001 modification metadata, listener pose information, flags, content creator input, device settings, user profiles, etc.), depending on use cases and / or implementations. To this end, apparatus 800 may comprise one or more interfaces 835 for receiving respective inputs. The processor 810 may be adapted to carry out or implement the methods / techniques described throughout the present disclosure (e.g., method 400 of Fig. 4, method 500 of Fig. 5, method 600 of Fig. 6, and / or method 700 of Fig. 7, as well as methods of authoring environmental metadata as described throughout the disclosure) and to generate corresponding output data 840 (e.g., adjusted audio mixes, rendered audio data, audio objects, superseding pose-dependent modification metadata, etc.), depending on use cases and / or implementations. To this end, apparatus 800 may comprise one or more interfaces 845 for outputting respective outputs.[000131] The present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products.Technical Advantages[000132] Important aspects of the present disclosure relate to how the “mix metadata” (e.g., superseding pose-dependent modification metadata) carries information to enable implementing the desired audio adjustments, for example for a variety of listening positions and / or deviating from what would be dictated by a physics-based or otherwise predetermined model. Without any mix metadata, the sound would be dictated only by how the physics of sound propagation is simulated by the playback device, which may not be desirable. Due to their ability to efficiently represent multiple mixes, corresponding to different listening positions, or to carry metadata indicating how characteristics of the audio mix should change based on listener pose, techniques according to the present disclosure are particularly applicable to scenarios where a listener is free to move around in a 6DoF virtual environment. This is contrary to scenarios with only a single listening position, where it may be more useful to pre-bake the audio mix into the audio signal. Using the proposed techniques, auditory cues reflecting the listener’ s movement can be provided to maintain or improve immersion, but it can also be ensured that the resulting audio mix is aesthetically pleasing and / or meets artistic intent.Interpretation[000133] Aspects of the techniques described herein may be implemented in an appropriate computer-based sound processing network environment (e.g., server or cloud environment)D23013W001 for processing digital or digitized audio files. Portions of these systems may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.[000134] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processorbased computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer- readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may he embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.[000135] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, apparatus and devices described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.[000136] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadestD23013W001 interpretation so as to encompass all such modifications and similar arrangements.[000137] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.Enumerated Example Embodiments[000138] Various Aspects and implementations of the invention may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.EEE1. A method of rendering audio for a virtual environment, the method comprising: obtaining input of one or more audio objects and corresponding positional metadata for the one or more audio objects, the one or more audio objects and corresponding positional metadata relating to (e.g., characterizing) an audio mix; obtaining superseding pose-dependent modification metadata describing how the audio mix is intended to vary based on listener pose of a listener in the virtual environment; obtaining listener pose information indicative of the listener pose in the virtual environment; and adjusting levels of the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information.EEE2. The method according to EEE1, further comprising: deriving relative positions of the one or more audio objects and the listener in the virtual environment based on the positional metadata and the listener pose information, wherein the levels of the one or more audio objects are adjusted based on the superseding pose-dependent modification metadata and the relative positions of the one or more audio objects and the listener in the virtual environment.EEE3. The method according to EEE1 or EEE2, further comprising: deriving a listener viewing direction in relation to the one or more audio objects based on the positional metadata and the listener pose information,D23013W001 wherein the levels of the one or more audio objects are adjusted based on the listener viewing direction in relation to the one or more audio objects.EEE4. The method according to any one of the preceding EEEs, further comprising: adjusting directions of arrival for the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information.EEE5. The method according to any one of the preceding EEEs, further comprising: adjusting times of arrival or propagation delays for the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information.EEE6. The method according to any one of the preceding EEEs, wherein the superseding pose-dependent modification metadata includes parameters indicative of one or more of: modifications to gains for the one or more audio objects based on the listener pose; modifications to apparent directions for the one or more audio objects based on the listener pose; and modifications to times of arrival or propagation delays for the one or more audio objects based on the listener pose.EEE7. The method according to any one of the preceding EEE, wherein the superseding pose-dependent modification metadata includes parameters indicative of one or more of: modifications to gains for the one or more audio objects based on positions of audio objects other than the one or more audio objects; and modifications to apparent directions for the one or more audio objects based on the positions of the audio objects other than the one or more audio objects.EEE8. The method according to any one of the preceding EEEs, wherein the superseding pose-dependent modification metadata includes parameters indicative of one or more of: modifications to gains for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment; modifications to apparent directions for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment; and modifications to times of arrival or propagation delays for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment.D23013W001EEE9. The method according to any one of the preceding EEEs, wherein the superseding pose-dependent modification metadata further includes a parameter indicative of a reverberation model to be used for the audio mix.EEE10. The method according to any one of the preceding EEEs, further comprising: applying limits to gains for the one or more audio objects based on relative distances between the one or more audio objects and a listener position indicated by the listener pose information.EEE11. The method according to any one of the preceding EEEs, further comprising: applying limits to changes of gains for the one of more audio objects relative to a reference audio mix for the audio mix.EEE12. The method according to any one of the preceding EEEs, further comprising: reading a flag indicative of whether the levels of the one or more audio objects shall be adjusted based on the superseding pose-dependent modification metadata or a default model.EEE13. The method according to any one of the preceding EEEs, wherein adjusting the level of a given one among the one or more audio objects of the audio mix is further based on a type or class of the audio object.EEE14. The method according to any one of the preceding EEEs, wherein the superseding pose-dependent modification metadata includes different sets of parameters for different zones of the virtual environment.EEE15. The method according to EEE14, comprising: obtaining input of different audio mixes for different zones of the virtual environment.EEE16. The method according to any one of the preceding EEEs, further comprising: obtaining audio mix metadata for the audio mix and for each of one or more reference listener poses in the virtual environment; and adjusting rendering characteristics of the one or more audio objects based on the audio mix metadata and the listener pose information.EEE17. The method according to EEE16, further comprising interpolating between reference listener poses for determining interpolated audio mix metadata for the listener pose in the virtual environment.D23013W001EEE18. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE 1 to EEE 17.EEE19. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE17.EEE20. A computer-readable storage medium storing the program of EEE19.

Claims

D23013W001Claims1. A method of rendering audio for a virtual environment, the method comprising: obtaining input of one or more audio objects and corresponding positional metadata for the one or more audio objects, the one or more audio objects and corresponding positional metadata relating to an audio mix; obtaining superseding pose-dependent modification metadata describing how the audio mix is intended to vary based on listener pose of a listener in the virtual environment; obtaining listener pose information indicative of the listener pose in the virtual environment; and adjusting levels of the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information.

2. The method according to claim 1, further comprising: deriving relative positions of the one or more audio objects and the listener in the virtual environment based on the positional metadata and the listener pose information, wherein the levels of the one or more audio objects are adjusted based on the superseding pose-dependent modification metadata and the relative positions of the one or more audio objects and the listener in the virtual environment.

3. The method according to claim 1 or 2, further comprising: deriving a listener viewing direction in relation to the one or more audio objects based on the positional metadata and the listener pose information, wherein the levels of the one or more audio objects are adjusted based on the listener viewing direction in relation to the one or more audio objects.

4. The method according to any one of the preceding claims, further comprising: adjusting directions of arrival for the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information.

5. The method according to any one of the preceding claims, further comprising:D23013W001 adjusting times of arrival or propagation delays for the one or more audio objects of the audio mix based on the superseding pose-dependent modification metadata and the listener pose information.

6. The method according to any one of the preceding claims, wherein the superseding pose-dependent modification metadata includes parameters indicative of one or more of: modifications to gains for the one or more audio objects based on the listener pose; modifications to apparent directions for the one or more audio objects based on the listener pose; and modifications to times of arrival or propagation delays for the one or more audio objects based on the listener pose.

7. The method according to any one of the preceding claims, wherein the superseding pose-dependent modification metadata includes parameters indicative of one or more of: modifications to gains for the one or more audio objects based on positions of audio objects other than the one or more audio objects; and modifications to apparent directions for the one or more audio objects based on the positions of the audio objects other than the one or more audio objects.

8. The method according to any one of the preceding claims, wherein the superseding pose-dependent modification metadata includes parameters indicative of one or more of: modifications to gains for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment; modifications to apparent directions for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment; and modifications to times of arrival or propagation delays for the one or more audio objects based on positions of occluding, partially occluding, or reflecting elements in the virtual environment.

9. The method according to any one of the preceding claims, wherein the superseding pose-dependent modification metadata further includes a parameter indicative of a reverberation model to be used for the audio mix.D23013W00110. The method according to any one of the preceding claims, further comprising: applying limits to gains for the one or more audio objects based on relative distances between the one or more audio objects and a listener position indicated by the listener pose information.

11. The method according to any one of the preceding claims, further comprising: applying limits to changes of gains for the one of more audio objects relative to a reference audio mix for the audio mix.

12. The method according to any one of the preceding claims, further comprising: reading a flag indicative of whether the levels of the one or more audio objects shall be adjusted based on the superseding pose-dependent modification metadata or a default model.

13. The method according to any one of the preceding claims, wherein adjusting the level of a given one among the one or more audio objects of the audio mix is further based on a type or class of the audio object.

14. The method according to any one of the preceding claims, wherein the superseding pose-dependent modification metadata includes different sets of parameters for different zones of the virtual environment.

15. The method according to claim 14, comprising: obtaining input of different audio mixes for different zones of the virtual environment.

16. The method according to any one of the preceding claims, further comprising: obtaining audio mix metadata for the audio mix and for each of one or more reference listener poses in the virtual environment; and adjusting rendering characteristics of the one or more audio objects based on the audio mix metadata and the listener pose information.

17. The method according to claim 16, further comprising interpolating between reference listener poses for determining interpolated audio mix metadata for the listener pose in the virtual environment.D23013W00118. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 17.

19. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 17.

20. A computer-readable storage medium storing the program of claim 19.

Citation Information

Patent Citations

  • Spatial audio rendering

    WO2024078809A1