Method and system for immersive 3DOF / 6DOF audio rendering

JP2025510923A5Pending Publication Date: 2026-04-07DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently rendering immersive audio on devices with limited power, such as AR glasses, due to high computational complexity and motion/inter-sound latency issues.

Method used

A method for rendering audio that involves a chain of renderers, where the first renderer processes audio data and metadata to generate digested rendering parameters, which are then used by subsequent renderers to reduce computational complexity and minimize latency.

Benefits of technology

This approach allows for efficient division of computational loads and minimizes motion/inter-sound latency, enabling high-quality immersive audio rendering on devices with limited power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Described herein is a method of rendering audio, the method including: receiving, at a first renderer, first audio data and first metadata for the first audio data, the first metadata including one or more canonical rendering parameters; processing, at the first renderer, the first metadata and optionally the first audio data to generate second metadata and optionally the second audio data; and providing, by the first renderer, the second metadata and optionally the second audio data for further processing by the second renderer, the second metadata including the one or more first digested rendering parameters and optionally a first portion of the one or more canonical rendering parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 326,063, filed March 31, 2022 (Reference: D22023USP1), and U.S. Provisional Application No. 63 / 490,197, filed March 14, 2023 (Reference: D22023USP2), all of which are incorporated herein by reference in their entireties.

[0002] The present disclosure relates generally to methods for rendering audio. In particular, the present disclosure relates to rendering audio by two or more renderers (rendering chains). The present disclosure further relates to respective systems and computer program products.

[0003] Although certain embodiments will be described herein with particular reference to that disclosure, it will be understood that the disclosure is not limited to such fields of use but is applicable in a broader context. [Background technology]

[0004] Throughout this disclosure, any discussion of background art should in no way be taken as an admission that such art is well known or forms part of the common general knowledge within the field.

[0005] Extended reality XR (e.g. Augmented Reality (AR) / Mixed Reality (MR) / Virtual Reality (VR)) may increasingly rely on end devices with very limited power. AR glasses are a prominent example. To make them as light as possible, they cannot be equipped with heavy batteries. As a result, only highly complexity-constrained numerical operations are possible on the processors they contain to allow reasonable operating times. Immersive audio, on the other hand, is an essential media component of XR services. The services may support adjusting the presented immersive audio / visual scenes in response to typically 3DoF or 6DoF user (head) movements. Making the corresponding immersive audio renditions in high quality typically demands high numerical complexity.

[0006] Thus, there is an existing need for improved rendering of immersive audio, in particular one that allows for efficient division of the computational load. Summary of the Invention [Means for solving the problem]

[0007] According to a first aspect of the present disclosure, a method of rendering audio is provided. The method may include receiving, at a first renderer, first audio data and first metadata for the first audio data, the first metadata including one or more canonical rendering parameters. The method may further include processing, at the first renderer, the first metadata and optionally the first audio data to generate second metadata and optionally the second audio data, the processing including generating one or more first digested rendering parameters based on the one or more canonical rendering parameters. The method may also include providing, by the first renderer, the second metadata and optionally the second audio data for further processing by the second renderer, the second metadata including the one or more first digested rendering parameters and optionally a first portion of the one or more canonical rendering parameters.

[0008] In some embodiments, some or all of the one or more first digested rendering parameters may be derived from a combination of at least two canonical rendering parameters.

[0009] In some embodiments, generating one or more first digested rendering parameters in the first renderer may further involve calculating the one or more first digested rendering parameters based on (e.g., to represent) an approximated (e.g., first-order) (digest) renderer model for the one or more canonical rendering parameters.

[0010] In some embodiments, the calculating step may involve calculating a first or higher order Taylor expansion of the renderer model based on one or more canonical rendering parameters.

[0011] In some embodiments, the method may further include receiving, at the first renderer, one or more external parameters, and the processing at the first renderer may further be based on the one or more external parameters.

[0012] In some embodiments, the one or more external parameters may include 3DOF / 6DOF tracking parameters, and the processing in the first renderer may be further based on the tracking parameters.

[0013] In some embodiments, the method may further include receiving, at the first renderer, timing information indicating a delay between the first and second renderers, and the processing at the first renderer may further be based on the timing information.

[0014] In some embodiments, the method may further include receiving, at the first renderer, audio captured from the second renderer, and the processing at the first renderer may further be based on the captured audio.

[0015] In some embodiments, the step of further processing by the second renderer may include rendering the output audio in the second renderer based on the second metadata and, optionally, the second audio data.

[0016] In some embodiments, rendering the output audio in the second renderer may further be based on one or more local parameters available in the second renderer.

[0017] In some embodiments, the secondary audio data may be primary pre-rendered audio data.

[0018] In some embodiments, the primary pre-rendered audio data may include one or more of mono audio, binaural audio, multi-channel audio, First Order Ambisonics (FOA) audio or Higher Order Ambisonics (HOA) audio, or combinations thereof.

[0019] In some embodiments, the first renderer may be implemented on one or more servers and the second renderer may be implemented on one or more end devices.

[0020] In some embodiments, one or more end devices may be a wearable device.

[0021] In some embodiments, the further processing by the second renderer may include processing the second metadata and optionally the second audio data at the second renderer to generate third metadata and optionally third audio data, where the processing includes generating one or more second digested rendering parameters based on the rendering parameters included in the second metadata. The further processing may also include providing the third metadata and optionally the third audio data by the second renderer for further processing by the third renderer of the third metadata including the one or more second digested rendering parameters and optionally a second portion of the one or more canonical rendering parameters.

[0022] In some embodiments, the step of further processing by a third renderer may include rendering the output audio in the third renderer based on the third metadata and, optionally, the third audio data.

[0023] In some embodiments, rendering the output audio in the third renderer may further be based on one or more local parameters available in the third renderer.

[0024] In some embodiments, the method may further include receiving, at the first renderer and / or at the second renderer, one or more external parameters, and the processing at the first renderer and / or at the second renderer may further be based on the one or more external parameters.

[0025] In some embodiments, the one or more external parameters may include 3DOF / 6DOF tracking parameters, and the processing in the first renderer and / or in the second renderer may be further based on the tracking parameters.

[0026] In some embodiments, the method may further include receiving, at the second renderer, timing information indicating a delay between the second and third renderers, and the processing at the second renderer may further be based on the timing information.

[0027] In some embodiments, the method may further include receiving, at the first renderer, audio captured from the third renderer, and the processing at the first renderer may further be based on the captured audio.

[0028] In some embodiments, generating the one or more second digested rendering parameters may be based on a first portion of the one or more canonical rendering parameters.

[0029] In some embodiments, the step of generating the one or more second digested rendering parameters may be further based on the one or more first digested rendering parameters.

[0030] In some embodiments, the second portion of the one or more canonical rendering parameters may be smaller than the first portion of the one or more canonical rendering parameters.

[0031] In some embodiments, the tertiary audio data may be secondary pre-rendered audio data.

[0032] In some embodiments, the secondary pre-rendered audio data may include one or more of mono audio, binaural audio, multi-channel audio, First Order Ambisonics (FOA) audio or Higher Order Ambisonics (HOA) audio, or combinations thereof.

[0033] In some embodiments, the first and second renderers may be implemented on one or more servers, and the third renderer may be implemented on one or more end devices.

[0034] In some embodiments, one or more end devices may be a wearable device.

[0035] In some embodiments, the canonical rendering parameters may be rendering parameters that are related to independent audio features.

[0036] In some embodiments, generating the one or more digested rendering parameters may include performing scene simplification.

[0037] In some embodiments, the first, second, and / or third metadata may further include one or more local canonical rendering parameters.

[0038] In some embodiments, the first, second, and / or third metadata may further include one or more local digested rendering parameters.

[0039] In some embodiments, the one or more local canonical rendering parameters or the one or more local digested rendering parameters may be based on one or more device or user parameters, including at least one of a device orientation parameter, a user orientation parameter, a device position parameter, a user position parameter, user personalization information, or user environmental information.

[0040] In some embodiments, the first, second, or third audio data may further include locally captured or locally generated audio data.

[0041] According to a second aspect of the disclosure, a method of rendering audio is provided. The method may include receiving pre-processed metadata and, optionally, pre-rendered audio data at an intermediate renderer. The pre-processed metadata may include one or more of digested rendering parameters and / or canonical rendering parameters. The method may further include processing the pre-processed metadata and, optionally, the pre-rendered audio data at the intermediate renderer to generate secondary pre-processed metadata and, optionally, secondary pre-rendered audio data. The processing may include generating one or more secondary digested rendering parameters based on the rendering parameters included in the pre-processed metadata. The method may also include providing the secondary pre-processed metadata and, optionally, the secondary pre-rendered audio data by the intermediate renderer for further processing by a subsequent renderer. The secondary pre-processed metadata may include one or more secondary digested rendering parameters and, optionally, one or more of the canonical rendering parameters.

[0042] According to a third aspect of the present disclosure, a method of rendering audio is provided. The method may include receiving, at a first renderer, initial first audio data having one or more canonical properties. The method may further include generating, at the first renderer, first digested audio data from the initial first audio data based on the one or more canonical properties and one or more first digested rendering parameters associated with the first digested audio data. The first digested audio data may have fewer canonical properties than the initial first audio data. The method may also include providing, by the first renderer, the first digested audio data and the one or more first digested rendering parameters for further processing by a second renderer.

[0043] In some embodiments, the method may further include receiving, at the first renderer, one or more external parameters, and the generating at the first renderer may further be based on the one or more external parameters.

[0044] In some embodiments, the one or more external parameters may include 3DOF / 6DOF tracking parameters, and the generating in the first renderer may be further based on the tracking parameters.

[0045] In some embodiments, the method may further include receiving, at the first renderer, timing information indicating a delay between the first and second renderers, and the generating at the first renderer may further be based on the timing information.

[0046] In some embodiments, the delay may be calculated in the second renderer.

[0047] In some embodiments, the method may further include adjusting the tracking parameter based on the timing information. Optionally, the adjusting step may include predicting the tracking parameter based on the timing information.

[0048] In some embodiments, the adjusting step may be performed in the second renderer.

[0049] In some embodiments, the step of further processing by the second renderer may include rendering the output audio in the second renderer based at least in part on the first digested audio data and the one or more first digested rendering parameters.

[0050] In some embodiments, rendering the output audio in the second renderer may further be based on one or more local parameters available in the second renderer.

[0051] In some embodiments, the further processing by the second renderer may include processing the first digested audio data and, optionally, the one or more first digested rendering parameters in the second renderer to generate second digested audio data and one or more second digested rendering parameters. The second digested audio data may have fewer canonical properties than the first digested audio data. The further processing by the second renderer may also include providing the second digested audio data and the one or more second digested rendering parameters by the second renderer for further processing by a third renderer.

[0052] In some embodiments, the method may further include receiving, at the first renderer and / or at the second renderer, one or more external parameters, and the generating at the first renderer and / or the processing at the second renderer may further be based on the one or more external parameters.

[0053] In some embodiments, the one or more external parameters may include 3DOF / 6DOF tracking parameters, and the generating in the first renderer and / or processing in the second renderer may be further based on the tracking parameters.

[0054] In some embodiments, the method may further include receiving, at the second renderer, timing information indicating a delay between the second and third renderers, and the processing at the second renderer may further be based on the timing information.

[0055] In some embodiments, the delay may be calculated in a third renderer.

[0056] In some embodiments, the method may further include adjusting the tracking parameter based on the timing information. Optionally, the adjusting step may include predicting the tracking parameter based on the timing information.

[0057] In some embodiments, the adjusting step may be performed in a third renderer.

[0058] In some embodiments, the further processing by the third renderer may include rendering the output audio in the third renderer based at least in part on the second digested audio data and the one or more second digested rendering parameters.

[0059] In some embodiments, rendering the output audio in the third renderer may further be based on one or more local parameters available in the third renderer.

[0060] In some embodiments, the canonical properties may include one or more of extrinsic and / or intrinsic canonical properties. The extrinsic canonical properties may be associated with one or more canonical rendering parameters. The intrinsic canonical properties may be associated with properties of the audio data to preserve its potential to be fully rendered in response to the extrinsic renderer parameters.

[0061] In some embodiments, one or more of the canonical rendering parameters may be tracking parameters.

[0062] In some embodiments, the tracking parameters may be 3DOF / 6DOF tracking parameters.

[0063] In some embodiments, the method may further include receiving, at the first renderer, timing information indicating a delay between the first and second renderers, and the processing at the first renderer may further be based on the timing information.

[0064] In some embodiments, the method may further include adjusting the tracking parameter based on the timing information, and optionally, the adjusting step may include predicting the tracking parameter based on the timing information.

[0065] In some embodiments, some or all of the one or more digested rendering parameters may be derived from a combination of at least two canonical properties.

[0066] In some embodiments, some or all of the one or more digested rendering parameters may be derived from at least one canonical property and the individual initial or digested audio data.

[0067] In some embodiments, generating one or more digested rendering parameters in the individual renderers may further involve computing the one or more digested rendering parameters to represent the approximated renderer model with respect to one or more canonical properties.

[0068] In some embodiments, the calculating step may involve calculating a first or higher order Taylor expansion of the renderer model based on one or more canonical properties.

[0069] In some embodiments, the calculation of one or more digested rendering parameters may involve multiple renderings.

[0070] In some embodiments, computing the one or more digested rendering parameters may involve analysing signal properties of the initial first audio data and identifying parameters associated with an acoustic reception model.

[0071] In some embodiments, the first renderer may be implemented on one or more servers.

[0072] In some embodiments, the second renderer or the third renderer may be implemented on one or more end devices.

[0073] In some embodiments, one or more end devices may be a wearable device.

[0074] According to a fourth aspect of the present disclosure, a method of rendering audio is provided. The method may include receiving digested audio data at an intermediate renderer, the digested audio data having one or more canonical properties and one or more digested rendering parameters. The method may further include processing the digested audio data and, optionally, the one or more digested rendering parameters at the intermediate renderer to generate secondary digested audio data and the one or more secondary digested rendering parameters. The secondary digested audio data may have fewer canonical properties than the digested audio data. The method may also include providing the secondary digested audio data and the one or more secondary digested rendering parameters by the intermediate renderer for further processing by a subsequent renderer.

[0075] According to a fifth aspect of the present disclosure, there is provided a system including one or more processors configured to perform operations as described herein.

[0076] According to a sixth aspect of the present disclosure, there is provided a program comprising instructions that, when executed by a processor, cause the processor to perform a method as described herein. The program may be stored on a computer-readable storage medium.

[0077] It should be understood that system (apparatus) features and method steps can be interchanged in many ways. In particular, details of the disclosed methods can be implemented by corresponding systems (apparatus), and vice versa, as would be understood by a person skilled in the art. It is also understood that any of the above statements made with respect to the method apply equally to the corresponding systems (apparatus), and vice versa. [Brief description of the drawings]

[0078] Exemplary embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0079] [Figure 1] FIG. 1 illustrates an example of a method for rendering audio according to an embodiment of the present disclosure.

[0080] [Diagram 2] FIG. 2 illustrates an example of a system for rendering audio by first and second renderers, according to an embodiment of this disclosure.

[0081] [Diagram 3] FIG. 3 illustrates a further example of a system for rendering audio by first and second renderers, according to an embodiment of the present disclosure.

[0082] [Figure 4] FIG. 4 illustrates an example of a system for rendering audio by first, second, and third renderers, according to an embodiment of this disclosure.

[0083] [Diagram 5] FIG. 5 illustrates a further example of a system for rendering audio by first, second, and third renderers, according to an embodiment of the present disclosure.

[0084] [Figure 6] FIG. 6 illustrates a further example of a system for rendering audio by first, second, and third renderers, according to an embodiment of the present disclosure.

[0085] [Figure 7]FIG. 7 illustrates a further example of a system for rendering audio by first, second, and third renderers, according to an embodiment of the present disclosure.

[0086] [Figure 8] FIG. 8 illustrates an example of a system for rendering audio by first, second, and third renderers in the context of 3GPP IVAS and MPEG-I audio, according to an embodiment of this disclosure.

[0087] [Figure 9] FIG. 9 illustrates an example of a system for rendering audio by first, second, and third renderers including local parameters and local audio, according to an embodiment of the present disclosure.

[0088] [Figure 10] FIG. 10 illustrates another example of a method for rendering audio according to an embodiment of the present disclosure.

[0089] [Figure 11] FIG. 11 illustrates an example of a system for rendering audio by first and second renderers according to an embodiment of the present disclosure.

[0090] [Figure 12] FIG. 12 illustrates a further example of a system for rendering audio by first, second, and third renderers, according to an embodiment of the present disclosure.

[0091] [Figure 13] FIG. 13 diagrammatically illustrates an example of an apparatus for implementing a method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0092] Description of exemplary embodiments Overview One potential solution to address the problem of high computational complexity of immersive audio renditions is to perform the rendering on some entity of the mobile / wireless network to which the end device is connected, or on a powerful mobile UE to which the end device is tethered, rather than on the device itself. In that case, the end device would receive only audio that is already binaurally rendered, for example. 3DoF / 6DoF pose information (head tracking metadata) would need to be transmitted to the rendering entity (network entity / UE). However, the latency for transmission between the end device and the network entity / UE can be very long, around 100 ms. Performing the rendering on the network entity / UE would therefore result in relying on outdated head tracking metadata and in that the binauralized audio performed by the end rendering device would not match the actual pose of the head / end device. This latency is referred to as motion / inter-sound latency. If it is too long, the end user will perceive this as quality degradation and will eventually suffer from motion sickness.

[0093] With respect to the video component of immersive media rendering, this problem has been addressed by a split rendering approach, where approximate parts of the video scene are rendered by the network entity / UE and the final video scene adjustments are made on the end device. However, with respect to audio, this area has been little, if at all, explored to date.

[0094] The MPEG-I Audio Renderer is one example of a 6DoF audio renderer that may be installed in the network entity / UE, and a stripped-down or low-power version of it may be installed in the end device. The low-power version may have certain constraints, such as a limited number of channels and objects, and a lower order high-order Ambisonics HOA (e.g., first order, i.e., FOA). Such a renderer may take as input the channel, object, and HOA signals, as well as 3DoF / 6DoF metadata, and output a binauralized or loudspeaker signal for AR / VR applications. In the specific case of HOA signal rendering, other dedicated 3DoF or 6DoF HOA renderers may be used, such as the MASA renderer and the MPEG-H Audio HOA renderer.

[0095] In many cases, for HOA content, it may be more preferable to enable the transmission of the HOA signal itself to the end device. This is prompted by the fact that scene adjustments (e.g., rotation, magnification) are better performed in this domain, and that the computationally less demanding HOA binauralization can be performed in the end device. This method may also be preferred to avoid the motion / inter-sound latency problem in a split rendering context. In bitrate-constrained scenarios that may arise between the rendering entity and the end device, a lower order HOA, e.g., FOA, is preferred. Appropriate measures on the original HOA signal shall be implemented other than simply shortening the HOA signal itself. At present, there are no concrete interfaces and solutions to this problem in a split rendering context.

[0096] Performing binauralization from the HOA / FOA representation in the end device preserves all inherent controllability of the Ambisonics audio representation to the end device, such as the possibility to perform scene rotation in response to head tracker (pose) metadata. In that sense, Ambisonics can be considered as a "canonical" audio representation. In contrast, after binauralization, a two-channel audio signal is obtained with less controllability. In particular, such a binaural audio signal is no longer head-trackable. However, according to a preferred embodiment, metadata may be associated with the binaural audio signal, which may represent information on how to adjust the binaural audio signal (e.g. in terms of loudness or spectral properties) to make it head-trackable again. The process for binauralizing the canonical audio representation and the generation of such metadata can thus be considered as converting the canonical audio representation into a digested representation, where the digested metadata may support the end device to perform output signal adjustments in response to metadata locally available in the end device. One advantage of this concept is that the operation in the end device may not be significantly more complicated, another advantage is that end devices that are not able to interpret the digest metadata may still be able to output a binaural audio signal as a fallback.

[0097] It should be noted that, although in the following, reference is made to a first, second, and third renderer, this sequence is for illustrative purposes only and is not intended to be limiting: any renderer followed by another renderer may be a front-end renderer, i.e., a first renderer, and any renderer that receives pre-rendered data may be an end renderer, i.e., a second renderer.

[0098] Furthermore, in the following, local and external parameters / data are referred to, but some of the external parameters may also be available locally, in other words, local parameters / data may also be said to be external parameters / data, and vice versa. Method and system for rendering audio - Patents.com

[0099] As a solution to the presented problem, the present disclosure describes a method and device for rendering audio that allows to effectively split the computational load and at the same time minimize the motion / inter-sound latency. One example of a method for (split) rendering (immersive) audio is illustrated in FIG.

[0100] In step S101, at a first renderer, first audio data and first metadata for the first audio data are received, the first metadata including one or more canonical rendering parameters.

[0101] In step S102, in a first renderer, the first metadata and, optionally, the first audio data are processed to generate second metadata and, optionally, second audio data, the processing step including a step of generating one or more first digested rendering parameters based on the one or more canonical rendering parameters.

[0102] In general, one aspect to be considered with the split rendering approach for immersive audio as described herein is that there may be two types of metadata or rendering parameters: canonical (or initial) and digested.

[0103] In some embodiments, canonical rendering parameters, which may also be referred to as initial rendering parameters, may be rendering parameters that relate to independent audio features. Parameters such as position, direction, directivity, range, transitionDistance, inside / outside, and authoring parameters (noDoppler, noDistance, ...) are typically canonical, meaning that they allow to control certain feature types independent of others. Although this is convenient, it may not necessarily lead to the least complex renderer solution. Therefore, it may not be very attractive or possible to apply them on very output limited devices in the final rendering stage on the end device. Canonical rendering parameters may also be said to be related / associated with extrinsic canonical properties.

[0104] The digested parameters refer to basic audio features such as gain, (spectral) shape, or time delay, and are less computationally intensive when applied to the audio signal during a rendering operation. The digested parameters may be obtained by "digesting" a set of canonical parameters associated with the audio (e.g., object metadata) and device parameters such as 3DOF or 6DOF orientation and position. In an embodiment, some or all of the (first) digested rendering parameters may be derived from a combination of at least two of the canonical rendering parameters. The term "digest" as used herein may therefore refer to extracting relevant features from the individual canonical rendering parameters and combining them with digested rendering parameters that may be applied to the audio signal with reduced complexity. The rendering process using the digested parameters may therefore be less computationally intensive and therefore can also be applied on very power-limited devices in the final rendering stage. In an embodiment, some or all of the (second) digested rendering parameters may be derived from a combination of one or more of the canonical rendering parameters and one or more of the previously generated (first) digested parameters, as further described below. Some or all of the digested rendering parameters may further be derived from one canonical rendering parameter, as follows: As a basic example, an object with a (one-dimensional) room coordinate x as a single parameter may be assumed. Direct rendering may require, first, to calculate the distance between the object and the listener (x coordinate of the head tracker) and, second, to apply a distance model and attenuate the audio signal depending on the distance. This may be more or less complicated.The digested rendering parameters may then be simple coefficients by which the end renderer multiplies the x-coordinate of the listener to obtain a scaling factor for the audio signal.

[0105] Referring again to the example of Figure 1, in step S103, second metadata and optionally second audio data are provided by the first renderer for further processing by the second renderer. The second metadata includes one or more first digested rendering parameters and optionally a first portion of the one or more canonical rendering parameters.

[0106] A combination of canonical and digested parameters may be used to control the computational complexity of the 3DoF / 6DoF audio renderer to closely match the needs of the underlying hardware platform. This approach may be used to build a chain of two or more renderers, distributed, for example, across various components of a network, all contributing to the final experience.

[0107] 2 and 3 illustrate an example of a system for rendering audio by a chain of first and second renderers to implement the described methods.

[0108] 2 and 3, the system includes a first renderer 207, which in an embodiment may be implemented on one or more servers, e.g., in a network or on an EDGE server, and a second renderer 209, which in an embodiment may be implemented on one or more end devices of a user. In an embodiment, the one or more end devices may be wearable devices.

[0109] 2 and 3, the first renderer 207 receives first metadata including a number of N canonical rendering parameters 201-205. Notably, in an embodiment, some or all of the digested rendering parameters may be derived from a combination of two or more of the canonical rendering parameters. The metadata may include one or more canonical rendering parameters.

[0110] The first metadata is processed by the first renderer 207 to generate second metadata 208, 203-205. The second metadata includes one or more first digested rendering parameters and, optionally, a first portion of one or more canonical rendering parameters. In the example of Figs. 2 and 3, generating the second metadata by the first renderer 207 includes digesting / processing two canonical rendering parameters 201, 202. As a result of the combined processing of the canonical rendering parameters 201, 202, first digested rendering parameters 208 are generated. It should be noted that the example of Figs. 2 and 3 is non-limiting in that the number of digested rendering parameters generated is also not limited and will depend on the individual use case.

[0111] As can be derived from the examples of Figures 2 and 3, the second metadata includes the first digested rendering parameters 208 received by the first renderer as well as a portion of the canonical rendering parameters 203-205. It should be noted that the inclusion of a portion of the canonical rendering parameters in the second metadata is optional and may depend on the use case. A portion of the canonical rendering parameters may be used in the final rendering stage, but may also be used for further intermediate rendering steps if the chain of renderers includes more than two renderers, as illustrated in the examples of Figures 4-7.

[0112] In the example of Figures 2 and 3, the first renderer 207 also receives the first audio data 206. Depending on the use case, the first audio data 206 may be processed by the first renderer 207 to generate second audio data 211, as illustrated in the example of Figure 3. In an embodiment, the second audio data 211 may be primary pre-rendered audio data. The primary pre-rendered audio data may include one or more of mono audio, binaural audio, multi-channel audio, first order Ambisonics (FOA) audio or higher order Ambisonics (HOA audio), or a combination thereof.

[0113] In the example of Figs. 2 and 3, the second renderer 209 may be said to be the final renderer, which performs the final rendering step. That is, in an embodiment, the output audio is rendered by the second renderer 209 based on the second metadata and, optionally, the second audio data 211. The step of rendering the output audio 210 by the second renderer 209 may also be based on one or more local parameters 212 available in the second renderer 209. The local parameters 212 may be, for example, head tracker data. The one or more local parameters 212 may also be transmitted to the front-end renderer as extrinsic parameters 213. The processing step by the first (front-end) renderer 207 may then also be based on these extrinsic parameters 213. In some embodiments, the one or more extrinsic parameters may include 3DOF / 6DOF tracking parameters, and the processing step in the first renderer may also be based on the tracking parameters.

[0114] Referring now to the example of FIG. 4-7, the chain of renderers may also include more than two renderers. In the example of FIG. 4-7, the chain of renderers includes three renderers. In an embodiment, the first renderer 407 and the second renderer 409 may be implemented on one or more servers, for example in the network and in an EDGE server. The third renderer 411 may be implemented on one or more end devices of the user. The one or more end devices may be wearable devices.

[0115] It should also be noted that in the examples of Figs. 4-7, the first renderer 407 receives first audio data, which may be optionally processed by the second renderer 409 and / or the third renderer 411, depending on the use case. In contrast to the examples of Figs. 2 and 3, in these examples the second renderer 409 represents an intermediate renderer, while the third renderer 411 represents a final renderer, which performs the final rendering step. If the first audio data is processed by the first renderer 407, the generated second audio data may therefore comprise pre-rendered, in particular pre-binauralized, audio. Similar to the second audio data 413, which may be primary pre-rendered audio data in one embodiment, the third audio data 414 may be secondary pre-rendered audio data. In an embodiment, the secondary pre-rendered audio data may include one or more of mono audio, binaural audio, multi-channel audio, object audio, first-order Ambisonics (FOA) audio or higher-order Ambisonics (HOA) audio, or a combination thereof. As can be derived from the example of Fig. 4-7, the second renderer 409 provides third metadata 410, 404-405 and, optionally, third audio data 414 for further processing by the third renderer 411. The third renderer 411 may also represent an intermediate renderer, but in the example of Fig. 4-7, the third renderer 411 renders the output audio 412 based on the third metadata 410, 404-405 and, optionally, the third audio data 414. The step of rendering the audio 412 output by the third renderer 411 may further also be based on one or more local parameters 415 available in the third renderer 411. The local parameters may be, for example, head tracker data. The one or more local parameters 415 may also be transmitted to the front-end renderers as external parameters 416, 417.The processing by the first and / or second (front-end) renderers 407, 409 may then also be based on these extrinsic parameters 416, 417. In some embodiments, the one or more extrinsic parameters may include 3DOF / 6DOF tracking parameters, and the processing in the first renderer and / or in the second renderer may further be based on the tracking parameters.

[0116] Up to the second rendering stage in Figures 4-7 the processing steps are identical to the embodiment of Figures 2 and 3 described above. In contrast to the embodiment of Figures 2 and 3, in the second renderer 409 the second metadata 408, 403-405 and optionally the second audio data 413 are now processed to generate third metadata 410, 404-405 and optionally third audio data 414. The processing step in the second renderer 409 comprises generating one or more second digested rendering parameters 410 based on the rendering parameters 408, 403-405 contained in the second metadata. In this case, since the second metadata may include digested rendering parameters and canonical rendering parameters, the second digested rendering parameters 410 may be derived from a combination of the first digested rendering parameters 408 and the canonical rendering parameters 403 from among the first portion of the canonical rendering parameters, as illustrated in the example of Figures 4-7. Alternatively or additionally, the second digested rendering parameters may also be derived from a combination of two canonical rendering parameters from among the first portion of the canonical rendering parameters. It should also be noted that the number of second digested rendering parameters generated is not limited and may depend on the use case.

[0117] Thus, the generated third metadata includes one or more second digested rendering parameters and, optionally, a second portion of the one or more canonical rendering parameters. In the example of Figures 4-7, the third metadata is illustrated to include one of the second digested rendering parameters 410 and the second portion of the canonical rendering parameters 404-405. As illustrated in the example of Figures 4-7, in an embodiment, the second portion of the one or more canonical rendering parameters 404-405 may be smaller than the first portion of the one or more canonical rendering parameters 403-405.

[0118] Additionally, an MPEG-I 6DoF audio renderer in combination with a 3GPP® IVAS codec and renderer as shown in the example of FIG. 8 may be an example where the concept of canonical and digested parameters, and therefore the “split rendering” approach, is applied. In this example, a “social VR audio bitstream” 801 may be coded using 3GPP® IVAS, containing compressed audio and metadata (metadata A 802). Metadata A 802 may be a set of related canonical or digested parameters or a mixture thereof. Renderer A 803 may take metadata A 802 as input and “transform” it into “low-delay audio” 804 and metadata B 805. “Low-delay audio” 804 may be an intermediate audio format such as pre-binauralized audio. Metadata B 805 may also be a set of related canonical or digested parameters or a mixture thereof. The renderer B806 may take the metadata B805 as input and "transform" it into another audio representation 807 and metadata C808. This another audio representation 807 may be an intermediate audio format such as pre-binauralized audio. The metadata C808 may also be a set of related canonical or digested parameters or a mixture thereof. The renderer C809 may take the metadata C808 as input and "transform" it into a final audio representation 810 such as binauralized audio or loudspeaker feed. The rendering of the final audio output representation 810 by the renderer C809 may also be based on one or more local parameters / data 811 available by the renderer C809. The local parameters may be, for example, head tracker data. The one or more local parameters 811 may also be transmitted to the front-end renderer as external parameters 812.The processing by Renderer A 803 and / or Renderer B 806 may then also be based on one or more of these external parameters 812.

[0119] For example, for an XR use case, the real listening environment representation may contain local parameters and signals (e.g., local audio, RT60, critical distance, mesh, RIR data, pose, position information, properties of output devices (headphones, car speakers), etc.). These local data may be available on the end device side, but applying it directly there is computationally expensive. These parameters may be sent to a pre-rendering entity and processed as external data together with the rest of the XR audio scene data. The resulting pre-rendered, listener-environment adjusted XR audio scene content (together with the associated "digested" parameters) is returned to the end device side in a simpler "digested" representation form (suitable for low-complexity rendering).

[0120] Some pre-rendering processing steps (e.g., listener environment independent) can be performed once for many rendering end devices, which can provide additional computational advantages for multi-user / social XR scenarios.

[0121] Furthermore, the pre-rendering process can take into account the computational / bitrate capabilities and latency requirements associated with the rendering end device. To meet the corresponding requirements, a scene simplification step can be performed during the conversion of the "canonical" to "digested" parameters. For example, this step can include the reduction of: ·Update rate ·Frequency resolution The number of corresponding elements, for example by combining two or more audio objects into one

[0122] Another example for handling different computation or bitrate capabilities of end rendering devices is to associate priority metadata with different sound effects as produced by the renditions or corresponding metadata controlling such effects. End devices that suffer from resource shortages (permanence of transient computation limitations, power limitations due to battery drain) may then use the priority information and downscale with respect to sound effects in a controlled manner to best maintain the overall constrained user experience. Priority metadata may be associated with the received audio or depend on the end user's user preferences / interactions or the end user's situational context (ambient sounds, focus of attention, visual scene).

[0123] To address bitrate-constrained scenarios for HOA signal end device rendering, the following approach can be taken, which assumes that the end device is capable of handling low-power processing such as FOA / binaural rendering and / or simple panning functions. -One or more sector and / or perimeter FOA signals, possibly with additional sector / perimeter based metadata (e.g. direction, sector area), can be extracted and used for final rendering. For example, MPEG-H decoded HOA signals can be processed to generate the above formats and transmit them to a low-power MPEG-I renderer in the end device. -In the case of sources with range, a reduced order HOA signal such as FOA can also be used to control the range width at the end device by additionally including control parameters such as blurring, blending, and filter coefficients as accompanying metadata. - One or more primary signals with accompanying metadata (e.g. directional information) can be extracted and transmitted to the end device. Ambient signals can additionally be transmitted in FOA format or as transport signals with accompanying metadata such as intensity, diffuseness, and spatial energy. In very low bit rate conditions, a parametric representation of the HOA signal plus zero or more transport channels may be extracted or (if the HOA signal is already parametrically encoded, as is typically done in low bit rate HOA processing scenarios) simply automatically forwarded to the end device.

[0124] Note that the transport or primary signals may be rendered by a simple panning function or a channel / object renderer. In such cases, these signals (e.g., transport signal, primary signal, sector and surrounding FOA / HOA) may be appropriately addressed in the MPEG-I audio context, for example, by additionally declaring the signals with accompanying metadata information. Existing MPEG-I audio signal properties (metadata information, interfaces) are declared below for reference.

[0125] A 3DoF / 6DoF audio renderer (e.g., an MPEG-I audio renderer) may provide an input interface for canonical parameters, as further detailed in Tables 1, 2, and 3. In addition to those parameters, the 3DoF / 6DoF audio renderer may provide an input interface for digested parameters, which are combinations of the canonical parameters, for example, as listed above and in Tables 1, 2, and 3. Additionally, the 3DoF / 6DoF audio renderer may provide an input interface for a combination of the canonical and digested parameters. In one embodiment, the digested parameters may be a 3DoF representation, derived from the 6DoF parameters. [Table 1] [Table 2] [Table 3]

[0126] For a split-render approach as described herein, it may be assumed that there is a low-power end-render device (e.g., AR glasses) and at least one pre-rendering entity (EDGE, high-performance UE). In general, to ensure the shortest possible motion / inter-sound latency, ideally all rendering would be performed in the end-render device. However, since the end-render device may not be capable of handling the processing, some of the processing may be performed in a more powerful front-end renderer. Still, not all of the processing can be performed by the front-end renderer due to transmission latency between the last front-end renderer in the renderer chain and the end-render device.

[0127] In the concept applied here, it is assumed that applying the digested parameters in some approximation form can be done with very low complexity by the end renderer stage. For example, if the audio signal is a binaurally pre-rendered signal, this would simply mean gain adjusting, filtering, or time-shifting the two channels. The point is to get the exact parameters and the exact signal to be filtered with these parameters. It is further assumed that the exact parameter digest and the exact application of the digested parameters can be very complex operations that cannot be done by the end renderer, but only by the front renderer.

[0128] As described herein, it is proposed to split the rendering into at least two parts as follows.

[0129] The first part of the rendering may be performed by one or more front-end renderers, which receive the audio signal to be rendered as well as its metadata. Optionally, the (delayed) tracking parameters (3DOF / 6DOF) and the captured audio from the end renderer device may be received. The front-end renderer may render the received audio in response to all the parameters and render the captured audio into a pre-rendered audio signal. This signal is typically binauralized audio, which would be the best possible output signal, except for the delayed tracking data. In addition, the front-end renderer may calculate parameters of the (first or higher order) digest renderer model (first digested parameters), which are essentially first or higher order Taylor expansions of functions of the digested parameters with respect to the tracking parameters, e.g., gain, spectral shape, time delay. That is, in some embodiments, generating one or more first digested rendering parameters in the first renderer as described herein may further involve calculating the one or more first digested rendering parameters based on (e.g., as representative of) an approximated (e.g., first-order) (digested) renderer model for one or more canonical rendering parameters. In some embodiments, the calculating step may involve calculating a first-order or higher-order Taylor expansion of the renderer model based on the one or more canonical rendering parameters to obtain digested rendering parameters. As described above, a first-order or higher-order Taylor expansion of a function of the one or more canonical rendering parameters may also be performed with respect to the tracking parameters if received at the first renderer. Of note, in embodiments relating to a chain of renderers with more than two renderers, second or further digested rendering parameters may be calculated in a similar manner.

[0130] As will be appreciated by those skilled in the art, the computation of a Taylor expansion of order n requires the availability of the nth derivative of the function to be approximated. Numerically, the nth derivative is obtained by evaluating at least n+1 function values. Thus, a pre-renderer must be applied for the n+1 "search" attitudes (or positions) to calculate such n+1 function values ​​of, for example, gain, spectral shape, time delay. Using the function values, the first derivative is approximated by calculating the difference quotient between the function value difference and the difference value of the sought attitude (or position) parameter, such as the attitude angle (or Cartesian position coordinate value). Higher order derivatives are calculated according to similar known techniques.

[0131] Based on these digest renderer model parameters, the end renderer may adjust the received pre-binauralized audio signal in response to the tracking parameters. For example, the left or right binaural audio channel will be gain adjusted by an amount proportional to the first order gain factor times the difference amount of the given tracking parameter (assuming that a zeroth order factor (constant) has been applied in the front renderer). The difference amount is the amount by which the tracking parameter has changed between the value assumed by the front renderer and the actual amount as perceived by the end renderer. It should be noted that the front renderer may perform extrapolations of the tracking parameter evolution to increase the accuracy of the pre-rendered audio. The front renderer may also obtain these extrapolations using additional information, e.g., describing the recalled (or most likely) user position and orientation trajectory (e.g., derived from data describing the scene elements that attract the user's attention), the user listening environment (e.g., real room dimensions), etc.

[0132] In some embodiments, the method described herein may further include receiving, at the first renderer, timing information indicative of a delay between the first (front-end) and second (end) renderers, and the processing at the first renderer may further be based on the timing information. Of note, in embodiments relating to a chain of renderers with more than two renderers, the timing information may be received at the second to last, i.e., last front-end renderer in the chain prior to the end rendering. That is, in some embodiments, the method may further include receiving, at the second renderer, timing information indicative of a delay between the second and third renderers, and the processing at the second renderer may further be based on the timing information.

[0133] The timing information may, for example, indicate the actual round trip delay between the front-end and end renderers. Based on this approach, the parameters to be transmitted to the interface between a front-end renderer (e.g., Renderer B) and an end renderer (e.g., Renderer C) may be the following: External parameters and signals to the front-end renderer: -Tracking parameters (including user / scene interactions, e.g., "audio augmentation", i.e., audio objects of interest, i.e., "cocktail party effect", head tracking parameters (e.g., pose and / or position)) -Timestamp parameters to find the round trip delay -Audio captured from end-render devices -Additional information available only to the end renderer, namely the following description: - Sound playback settings (e.g. loudspeaker settings, headphone type, and HpTF compensation filter) -Listener-related and personalization data (e.g. personalized HRTF, listener EQ settings) - Listener environment (e.g. AR: real room reverberation, VR: playback area dimensions, background noise levels) -Listener posture and / or position (e.g. sitting / standing / VR-treadmill / inside a moving vehicle) External parameters and signals to the end renderer: -Audio content / signal (channels, objects, FOA, binaural audio, mono audio) -Accompanying audio signal metadata (e.g. sectors, mix coefficients, etc. see FOA metadata above) - Primary or higher order digest renderer coefficients -Timestamp parameters to find the round trip delay

[0134] Yet another example of metadata that may be transmitted between renderer instances over the interface relates to reverberation effects. For example, a front-end renderer may add reverberation according to a certain room model. Subsequent renderers (e.g., end renderers) may also have the capability to add reverberation. To avoid both renderers adding reverberation in a parallel fashion, it may be advantageous to inform the individual other renderers that reverberation is being added by one of the instances, so that the individual other renderers do not add this effect or at least take into account the reverberation effect of the other renderer. It may also be possible to split the reverberation effect between renderer instances, such that the less output-constrained renderer generates the computationally demanding parts of the effect (e.g., involving high-resolution filters with many taps), while the less capable end renderer makes only some adjustments to match the actual situational context at the end user / end rendering device.

[0135] In addition to the digested / canonical parameters and associated audio data, locally generated or captured audio 915 may be used as input to the various rendering stages, as illustrated in the example of FIG. 9. For example, renderer B in FIG. 8 may be running on a device (e.g., a smartphone) that also has a microphone for capturing local audio. The "local audio" block may generate associated metadata 913, 914 with the locally captured audio 915, which is input as digested or canonical parameters to the rendering stages and processed as described above. Such metadata may include the capture position in space, either in absolute or relative coordinates. Relative coordinates may be relative to the location of the microphone, the location of the smartphone, the location of the device running another rendering stage (e.g., renderer C running on AR glasses). In addition, the "local audio" block may generate audio data (e.g., earcon data) and associated metadata. Such accompanying metadata associated with the locally generated audio 915 may include a position of the locally generated audio 915 relative to a reference point within the virtual or augmented audio scene. That is, in an embodiment, the first, second, and / or third metadata may further include one or more local canonical rendering parameters. In a further embodiment, the first, second, and / or third metadata may further include one or more local digested rendering parameters. In a further embodiment, the one or more local canonical rendering parameters or the one or more local digested rendering parameters may be based on one or more device or user parameters including at least one of a device orientation parameter, a user orientation parameter, a device position parameter, a user position parameter, user personalization information, or user environment information.In further embodiments, the first, second, or third audio data may further include locally captured or locally generated audio data. The locally captured or locally generated audio data, the local canonical rendering parameters, and the local digested rendering parameters may be said to be associated with / derived from the local data, as described herein.

[0136] In addition to the above, the present disclosure describes a further exemplary method of rendering audio. The method may include receiving pre-processed metadata and, optionally, pre-rendered audio data at an intermediate renderer. The pre-processed metadata may include one or more of digested and / or canonical rendering parameters. The method may further include processing the pre-processed metadata and, optionally, the pre-rendered audio data at the intermediate renderer to generate secondary pre-processed metadata and, optionally, secondary pre-rendered audio data. The processing may include generating one or more secondary digested rendering parameters based on the rendering parameters included in the pre-processed metadata. The method may also include providing the secondary pre-processed metadata and, optionally, the secondary pre-rendered audio data by the intermediate renderer for further processing by a subsequent renderer. The secondary pre-processed metadata may include one or more secondary digested rendering parameters and, optionally, one or more of the canonical rendering parameters.

[0137] Advantageously, the exemplary methods described above may be implemented within an already existing renderer chain, based on an already existing renderer / system, or implemented to create a separate renderer chain. ALTERNATIVE METHOD AND SYSTEM FOR RENDERING AUDIO - Patent application

[0138] As a further solution to the presented problem, the present disclosure describes an alternative method and system for rendering audio that allows for effectively splitting the computational load and at the same time minimizing motion / inter-sound latency. One example of such an alternative method for (split) rendering (immersive) audio 1000 is illustrated in FIG.

[0139] Referring to the embodiment of FIG. 10, in step S1001, initial first audio data having one or more canonical properties is received at a first renderer.

[0140] In step S1002, in a first renderer, first digested audio data and one or more first digested rendering parameters associated with the first digested audio data are generated from the initial first audio data based on one or more canonical properties, where the first digested audio data has fewer canonical properties than the initial first audio data.

[0141] Also in step S1003, the first renderer provides the first digested audio data and the one or more first digested rendering parameters for further processing by the second renderer.

[0142] In general, audio signals such as Ambisonics may be considered as "canonical" audio representations, in the sense that individual audio data may have one or more canonical properties.

[0143] In an embodiment, the canonical properties may include one or more of the extrinsic and / or intrinsic canonical properties, which may be associated with one or more canonical rendering parameters, as already described above.

[0144] Intrinsic canonical properties may be associated with properties of the audio data to preserve the potential for being fully rendered in response to external renderer parameters. For example, a property such as scene rotatability of Ambisonics audio is intrinsically canonical, meaning that it allows for controlling certain feature types, such as scene orientation, independently from other features. Intrinsic canonical properties are also associated with properties of the audio signal that preserve the potential for being fully rendered in response to external renderer parameters, such as pose. As in the case with extrinsic canonical parameters, rendering from a canonical representation is expedient, but this may not necessarily lead to the least complex rendering solution. As an example, binaural rendering of Ambisonics audio may still be too complex for very output-limited end devices. Thus, rendering an audio signal with intrinsic canonical properties on very output-limited end devices may not be very attractive or possible.

[0145] In an embodiment, the method may further include receiving one or more external parameters in the first renderer as described, and the generating in the first renderer may further be based on the one or more external parameters. The one or more external parameters may include 3DOF / 6DOF tracking parameters. The generating in the first renderer may then further be based on the tracking parameters.

[0146] In an embodiment, the method may further include receiving, at the first renderer, timing information indicative of a delay between the first and second renderers. The generating / processing at the first renderer may then be further based on the timing information. The delay may be calculated at the second renderer.

[0147] In an embodiment, the method may still further include adjusting the tracking parameters based on the timing information, and optionally the adjusting step may include predicting the tracking parameters based on the timing information, and the adjusting (predicting) step may be performed in a second renderer.

[0148] It should be noted that although this is described referring to a first and second renderer, the delay derivation and prediction may occur at each front-end renderer (node), not just the first one, i.e., the case involving two renderers may be an example only, where round trip delay measurements may be performed between any two renderers, and adjustments may also occur anywhere in the renderer chain, and the same may apply when more than two renderers are involved.

[0149] Referring to the example of Figure 11, an example of a system for rendering audio by a chain of first and second renderers for implementing the described method is illustrated. In the example of Figure 11, the system includes a first renderer 1102, which in an embodiment may be implemented on one or more servers, e.g., in a network or on an EDGE server, and a second renderer 1105, which may be implemented on one or more end devices of a user. In an embodiment, the one or more end devices may be wearable devices.

[0150] In the embodiment of Fig. 11, a first renderer 1102 receives initial first audio data 1101. The initial first audio data 1101 may correspond to a canonical audio representation. In that sense, the initial first audio data 1101 has one or more canonical properties. As already detailed above, the one or more canonical properties may include one or more of extrinsic and / or intrinsic canonical properties.

[0151] In a first renderer 1102, from initial first audio data 1101, first digested audio data 1103 and one or more first digested rendering parameters 1104 associated with the first digested audio data 1103 are generated based on one or more canonical properties, where the first digested audio data 1103 has fewer canonical properties than the initial first audio data 1101.

[0152] Some or all of the one or more first digested rendering parameters 1104 may be derived from a combination of at least two of the canonical properties of the initial first audio data 1101. Alternatively or in addition, some or all of the one or more first digested rendering parameters 1104 may be derived from a step of combining at least one of the canonical properties of the initial first audio data 1101 and the individual initial first audio data 1101. The digested rendering parameters may also be obtained from a combination of the canonical properties of the audio data and external parameters / data 1108, e.g. local end device parameters / data 1107, e.g. pose and position.

[0153] Generating one or more first digested rendering parameters 1104 in a respective first renderer 1102 may further involve calculating one or more first digested rendering parameters 1004 to represent an approximated renderer model with respect to one or more canonical properties. The calculating step may involve calculating a first or higher order Taylor expansion of the renderer model based on the one or more canonical properties. The calculation of one or more first digested rendering parameters 1104 may involve multiple renderings in some embodiments. That is, multiple renderings may be performed at one renderer (node) of the rendering chain. For example, a front-end renderer, e.g., the first renderer, may render with respect to two hypothetical "exploration" poses. Alternatively, or in addition, the calculation of the one or more first digested rendering parameters 1104 may analyze signal properties of the initial first audio data 1101 and identify parameters related to the audio reception model, which may be applied to all renderers in the chain except the last renderer.

[0154] For example, the digested model parameters may be obtained in a front-end renderer (e.g., a first renderer) by analyzing the (canonical) audio signal in terms of certain signal properties such as direction or distance of the dominant sound source, applying a sound reception model (e.g., a human head model, a distance model), and calculating how the (binaurally) rendered output signal and the associated digested model parameters would change when a change is applied in pose (or position).

[0155] Although the steps of deriving / calculating / generating / obtaining individual digested rendering parameters may be described in terms of primary or secondary throughout this disclosure, individual method steps may also be applied to derive / calculate / generate / obtain individual higher order digested rendering parameters.

[0156] In the example of Fig. 11, the second renderer 1105 may be said to be the final (end) renderer performing the final rendering step. That is, in an embodiment, the output audio 1106 may be rendered by the second renderer 1105 based on the first digested audio data 1103 and at least in part based on one or more first digested rendering parameters 1104. The step of rendering the output audio by the second renderer may furthermore also be based on one or more local parameters 1107 available in the second renderer. The local parameters may be, for example, head tracker data. As already explained above, in an embodiment, the method may further comprise a step of receiving, in the first renderer 1102, one or more external parameters 1108. The step of generating in the first renderer 1102 may then further be based on the one or more external parameters 1108. The one or more external parameters may include 3DOF / 6DOF tracking parameters. The generating in the first renderer 1102 may then be further based on the tracking parameters.

[0157] Referring now to the example of FIG. 12, the chain of renderers may also include more than two renderers. In the example of FIG. 12, the chain of renderers includes three renderers. In an embodiment, the first renderer 1202 and the second renderer 1205 may be implemented on one or more servers, for example in the network and on an EDGE server. The third renderer 1208 may be implemented on one or more end devices of the user. The one or more end devices may be wearable devices.

[0158] In contrast to the embodiment of FIG. 11, in this embodiment the second renderer 1205 represents an intermediate renderer, while the third renderer 1208 represents a final renderer that performs the final rendering step.

[0159] That is, in the embodiment of Fig. 12, in a second renderer 1205, the first digested audio data 1203 and, optionally, one or more first digested rendering parameters 1204 may be processed to generate second digested audio data 1206 and one or more second digested rendering parameters 1207. The second digested audio data may have fewer canonical properties than the first digested audio data. This may be said to result from a successive complexity reduction during the rendering stage.

[0160] In an embodiment, the method may further include receiving in the first renderer 1202 and / or in the second renderer 1205 one or more extrinsic parameters 1212, 1211. The generating in the first renderer and / or the processing in the second renderer 1205 may then be further based on the one or more extrinsic parameters 1212, 1211. The one or more extrinsic parameters 1212, 1211 may include 3DOF / 6DOF tracking parameters. The generating in the first renderer 1202 and / or the processing in the second renderer 1205 may then be further based on the tracking parameters.

[0161] In an embodiment, the method may further include receiving, at the second renderer, timing information indicative of a delay between the second and third renderers. The generating / processing at the second renderer may then be further based on the timing information. The delay may be calculated at the third renderer.

[0162] In an embodiment, the method may still further include adjusting the tracking parameters based on the timing information, and optionally, the adjusting step may include predicting the tracking parameters based on the timing information, and the adjusting (predicting) step may be performed in a third renderer.

[0163] In the example of FIG. 12, the second renderer 1205 provides second digested audio data 1206 and one or more second digested rendering parameters 1207 for further processing by the third renderer 1208.

[0164] In this embodiment, since the third renderer 1208 is the final renderer, further processing by the third renderer 1208 includes rendering output audio 1209 based on the second digested audio data 1206 and at least in part based on one or more second digested rendering parameters 1207. Rendering the output audio by the third renderer may also be based on one or more local parameters 1210 available at the third renderer 1208. The local parameters may be, for example, head tracker data.

[0165] In addition to the above, the present disclosure describes a further exemplary method of rendering audio. The method may include receiving digested audio data at an intermediate renderer, the digested audio data having one or more canonical properties and one or more digested rendering parameters. The method may further include processing the digested audio data and, optionally, the one or more digested rendering parameters at the intermediate renderer to generate secondary digested audio data and one or more secondary digested rendering parameters. The secondary digested audio data may have fewer canonical properties than the digested audio data. The method may also include providing the secondary digested audio data and the one or more secondary digested rendering parameters by the intermediate renderer for further processing by a subsequent renderer. Apparatus for implementing methods according to the present disclosure

[0166] Finally, the present disclosure also relates to apparatus (e.g., computer-implemented apparatus) for implementing the methods and techniques described throughout the present disclosure. FIG. 13 illustrates one embodiment of such an apparatus 1300. In particular, the apparatus 1300 comprises a processor 1310 and a memory 1320 coupled to the processor 1310. The memory 1320 may store instructions for the processor 1310. The processor 1310 may also receive suitable input data 1330, among other things, depending on the use case and / or implementation. The processor 1310 may be adapted to perform the methods / techniques described throughout the present disclosure and generate corresponding output data 1340, depending on the use case and / or implementation. interpretation

[0167] Aspects of the system described herein may be implemented in a suitable computer-based sound processing network environment to process digital or digitized audio files. Part of an adaptive audio system may include one or more networks with any desired number of individual machines, including one or more routers (not shown) that act as buffers and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0168] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described in terms of their behavior, register transfers, logical components, and / or other characteristics, using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media. The computer-readable media in which such formatted data and / or instructions may be embodied include physical (non-transitory) non-volatile storage media in various forms, such as, but not limited to, optical, magnetic, or semiconductor storage media.

[0169] Although one or more implementations have been described by way of example and in terms of specific embodiments, it should be understood that the one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to one skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.

[0170] Various aspects and implementations of the present disclosure can also be understood from the non-claimed exemplary embodiments (EEE) set forth below.

[0171] EEE1. A method for processing audio, comprising: receiving, at a first renderer, one or more canonical rendering parameters; generating, in a first renderer, one or more digested rendering parameters based on the one or more canonical rendering parameters; providing, by the first renderer, one or more digested rendering parameters and, optionally, a portion of the one or more canonical rendering parameters to a second renderer; rendering the audio by a second renderer based on the one or more digested rendering parameters and, optionally, a portion of the one or more canonical rendering parameters; A method comprising:

[0172] EEE2. The method of EEE1, wherein the first renderer is implemented on one or more servers and the second renderer is implemented on one or more wearable devices.

[0173] EEE3. The method of EEE1, wherein the one or more canonical rendering parameters each comprise a parameter that controls a characteristic of the audio that is independent of other characteristics of the audio.

[0174] EEE4. The method of EEE1, wherein the one or more digested rendering parameters include one or more of parameters derived from one or more canonical parameters or one or more device or user parameters.

[0175] EEE5. The method of EEE4, wherein the one or more device or user parameters include at least one of a device orientation parameter, a user orientation parameter, a device position parameter, a user position parameter, user personalization information, or user environmental information.

[0176] EEE6. Rendering the audio with a second renderer includes: providing, by the second renderer, the one or more digested rendering parameters, one or more additional digested rendering parameters derived by the second renderer from a portion of the canonical rendering parameters, and a smaller portion of the one or more canonical rendering parameters to the third renderer; or providing, by a second renderer, a representation of the audio output to be played by the one or more transducers; The method according to claim 1, further comprising at least one of:

[0177] EEE7. The method of EEE1, comprising the step of generating pre-rendered audio by a first renderer based on one or more canonical rendering parameters, wherein the step of rendering the audio by a second renderer is based on the pre-rendered audio.

[0178] EEE8. The method of EEE7, wherein the pre-rendered audio includes at least one of mono audio, binaural audio, multi-channel audio, FOA audio, or HOA audio.

[0179] EEE9. The method of EEE7 comprising generating, by a second renderer, secondary pre-rendered audio based on the one or more digested rendering parameters and, optionally, a portion of the one or more canonical rendering parameters.

[0180] EEE10. The method of EEE8, wherein the secondary pre-rendered audio includes at least one of mono audio, binaural audio, multi-channel audio, FOA audio, or HOA audio.

[0181] EEE11. A system comprising one or more processors configured to perform the operations described in any one of EEE1-10.

[0182] EEE12. A computer program product configured to cause one or more processors to perform the operations recited in any one of EEE1-10.

Claims

1. A method for rendering audio, wherein the method is The first renderer receives first audio data and first metadata for the first audio data, wherein the first metadata includes one or more canonical rendering parameters, each of which includes a parameter that controls a feature of the first audio data that is independent of other features of the first audio data. The first renderer processes the first metadata and optionally the first audio data in order to generate second metadata and optionally the second audio data, wherein the processing includes generating one or more first digested rendering parameters based on one or more canonical rendering parameters, wherein some or all of the first digested rendering parameters are derived from at least two combinations of the canonical rendering parameters. The first renderer provides the second metadata and the first audio data, optionally the second metadata and the second audio data, for further processing by the second renderer, wherein the second metadata includes one or more first digested rendering parameters and optionally a first portion of one or more canonical rendering parameters. A method that includes rendering the output audio in the second renderer, and further depends on one or more local parameters available in the second renderer.

2. The method according to any one of claims 1 to 8, wherein further processing by the second renderer includes rendering output audio in the second renderer based on the second metadata and optionally the second audio data.

3. The method according to claim 11, wherein the primary pre-rendered audio data includes one or more of the following: mono audio, binaural audio, multi-channel audio, primary ambisonic audio, or higher-order ambisonic audio, or a combination thereof.

4. The method according to any one of claims 1-29, wherein generating one or more of the first or second digested rendering parameters includes performing scene simplification.

5. A system comprising one or more processors configured to perform the operations described in any one of claims 1-4.

6. A program, the program including instructions, the instructions, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, wherein the computer-readable storage medium stores the program described in claim 6.