Method and device for efficiently delivering edge-based rendering of 6DOF MPEG-i immersive audio

JP2025000851A5Pending Publication Date: 2025-10-21NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024170433
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-15
Filing Date
2024-09-30
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently delivering and rendering computationally complex 6-degree-of-freedom (6DoF) MPEG-I immersive audio due to resource constraints on end-user devices, particularly in maintaining immersion and responsiveness to user orientation and position changes.

Method used

The solution involves edge-based rendering, where computational-intensive tasks are offloaded to edge computing resources, and the rendered audio is efficiently delivered to user equipment using low-latency communication codecs, preserving spatial audio characteristics and responsiveness to user movements.

Benefits of technology

This approach enables resource-constrained devices to consume complex 6DoF audio scenes with low latency and high fidelity, maintaining immersion and responsiveness to user movements, expanding the market for immersive audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and a device for efficiently delivering edge-base rendering of 6DOF MPEG-I immersive audio.SOLUTION: A device for generating spatialized audio output based on a user position includes means for performing: a step for acquiring a user position value; a step for acquiring an input audio signal for achieving rendering of an input audio signal and related metadata; a step for generating an intermediate format immersive audio signal based on the input audio signal, metadata and the user position value; a step for processing the intermediate format immersive audio signal and acquiring a space parameter and the audio signal; and a step for encoding the space parameter and the audio signal. The space parameter and the audio signal are configured to partially generate spatialized audio output.SELECTED DRAWING: Figure 1a
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application relates to methods and apparatus for efficient delivery of edge-based rendering of six degrees of freedom MPEG-I immersive audio, including, but not limited to, methods and apparatus for efficient delivery of edge-based rendering of six degrees of freedom MPEG-I immersive audio to user equipment based renderers. [Background technology]

[0002] Modern cellular and wireless communication networks such as 5G have brought computational resources for a variety of uses and services closer to the edge, providing the impetus for mission-critical network-enabled applications and immersive multimedia rendering.

[0003] Additionally, these networks significantly reduce latency and bandwidth constraints between the edge computing layer and end-user media consumption devices such as mobile devices, HMDs (configured for augmented reality / virtual reality / mixed reality - AR / VR / XR applications), and tablets.

[0004] Ultra-low latency edge computing resources can be used by end-user devices with end-to-end latency of less than 10 ms (e.g., latencies as low as 4 ms have been reported). Hardware-accelerated SoCs (Systems on Chips) are increasingly being deployed on media edge computing platforms to exploit rich and diverse multimedia processing applications for volumetric and immersive media (e.g., six degrees of freedom - 6DoF audio). These trends make edge computing-based media encoding as well as edge computing-based rendering attractive propositions. Advanced media experiences can be delivered to many devices that do not have the capability to perform highly complex volumetric and immersive media rendering.

[0005] The MPEG-I6DoF audio format standardized by MPEG Audio WG06 is often computationally very complex for immersive audio scenes. The process of encoding, decoding and rendering a scene from the MPEG-I6DoF audio format is, in other words, computationally very complex or expensive. For example, within a moderately complex scene rendering: Audio for second or third order effects (e.g. modeling reflections from an audio source) can result in a large number of effective image sources, which makes not only the rendering (implemented in a listener position dependent manner) but also the encoding (which can happen offline) a very complex proposition.

[0006] Furthermore, the Immersive Audio and Audio Services (IVAS) codec is an extension of the 3GPP EVS codec and is intended for new immersive audio and audio services over communication networks as described above. Such immersive services include, for example, immersive audio and audio for virtual reality (VR). This versatile audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. It is expected to support a variety of input formats, including channel-based and scene-based input. It is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions. Standardization of the IVAS codec is expected to be completed by the end of 2022.

[0007] Metadata-assisted spatial audio (MASA) is one input format proposed for IVAS. It uses audio signals together with corresponding spatial metadata (including, for example, direction in frequency bands and direct-to-total energy ratio). The MASA stream can be obtained, for example, by capturing spatial audio with a microphone of a mobile device, and the spatial metadata settings are estimated based on the microphone signal. The MASA stream can also be obtained from other sources, for example, specific spatial audio microphones (e.g., Ambisonics), studio mixes (e.g., 5.1 mixes), or other content, by means of appropriate format conversion. One such conversion method is disclosed in Tdoc S4-191167 (Nokia Corporation: Description of IVAS MASA C Reference Software; 3GPP TSG-SA4#106 meeting; 21-25 October, 2019, Busan, Republic of Korea). Summary of the Invention

[0008] According to a first aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus comprising means configured to obtain a user position value, obtain at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal, generate an intermediate format immersive audio signal based on the at least one input audio signal, the metadata and the user position value, process the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal, and encode the at least one spatial parameter and the at least one audio signal, where the at least one spatial parameter and the at least one audio signal are configured to at least partially generate the spatialized audio output.

[0009] The means may be further configured to transmit the encoded at least one spatial parameter and at least one audio signal to a further device, the device being further configured to output a binaural or multi-channel audio signal based on processing of the at least one audio signal, the processing being based on the user rotation value and the at least one spatial audio rendering parameter.

[0010] The further device may be operated by a user and the means configured to obtain user position values ​​may be configured to receive the user position values ​​from the further device.

[0011] The means configured to obtain user position values ​​may be configured to receive the user position values ​​from a head mounted device operated by the user.

[0012] The means may be further configured to transmit the user location.

[0013] The means configured to process the intermediate format immersive audio signal to obtain at least one spatial parameter may be configured to generate a metadata assisted spatial audio bitstream.

[0014] The means configured to encode at least one spatial parameter may be configured to generate an immersive voice and audio services bitstream.

[0015] The means configured to encode the at least one spatial parameter and the at least one audio signal may be configured for low latency encoding of the at least one spatial parameter and the at least one audio signal.

[0016] The means configured to process the intermediate format immersive audio signal to obtain at least one spatial parameter and to obtain at least one audio signal comprises: The apparatus may be configured to determine an audio frame length difference between the intermediate format immersive audio signal and the at least one audio signal, and to control buffering of the intermediate format immersive audio signal based on the determination of the audio frame length difference.

[0017] The means may be configured to obtain a user rotation value, and wherein the means further configured to generate the intermediate format immersive audio signal may be configured to generate the intermediate format immersive audio signal further based on the user rotation value.

[0018] The means configured to generate the intermediate format immersive audio signal may be configured to generate the intermediate format immersive audio signal further based on a predetermined or agreed upon user rotation value, and the further device may be configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal, the processing being based on the predetermined or agreed upon user rotation value and the obtained user rotation value for the at least one spatial audio rendering parameter.

[0019] According to a second aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus comprising means configured to: obtain a user position value and a rotation value; obtain an encoded at least one audio signal and at least one spatial parameter, where the encoded at least one audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position value; and generate an output audio signal based on processing the encoded at least one audio signal, the at least one spatial parameter and the user rotation value in six degrees of freedom.

[0020] The apparatus may be operated by a user and the means configured to obtain a user position value may be configured to generate a user position value.

[0021] The means configured to obtain user position values ​​may be configured to receive the user position values ​​from a head mounted device operated by the user.

[0022] The means configured to obtain the encoded at least one audio signal and the at least one spatial parameter may be configured to receive the encoded at least one audio signal and the at least one spatial parameter from a further device.

[0023] The means may be further configured to receive a user position value and / or a user orientation value from the further device.

[0024] The means may be configured to transmit the user position values ​​and / or the user orientation values ​​to a further device, which may be configured to generate an intermediate format immersive audio signal based on the at least one input audio signal, the determined metadata, and the user position values, and to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal.

[0025] The encoded at least one audio signal may be low latency encoded at least one audio signal.

[0026] The intermediate format immersive audio signal may have a format selected based on encoding compressibility of the intermediate format immersive audio signal.

[0027] According to a third aspect, there is provided a method for an apparatus for generating a spatialized audio output based on a user position, the method comprising: obtaining a user position value; obtaining at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal; generating an intermediate format immersive audio signal based on the at least one input audio signal, the metadata and the user position value; processing the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal; and encoding the at least one spatial parameter and the at least one audio signal, wherein the at least one spatial parameter and the at least one audio signal are configured to at least partially generate the spatialized audio output.

[0028] The method may further comprise transmitting the encoded at least one spatial parameter and the at least one audio signal to a further device, which may be configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal, the processing being based on the user rotation value and the at least one spatial audio rendering parameter.

[0029] The further device may be operated by a user, and obtaining the user position value may include receiving the user position value from the further device.

[0030] Obtaining the user position values ​​may include receiving the user position values ​​from a head-mounted device operated by the user.

[0031] The method may further include transmitting the user position value.

[0032] Processing the intermediate format immersive audio signal to obtain at least one spatial parameter can comprise generating a metadata-assisted spatial audio bitstream.

[0033] Encoding the at least one spatial parameter and the at least one audio signal may comprise generating an immersive audio and audio services bitstream.

[0034] Encoding the at least one spatial parameter and the at least one audio signal may comprise low latency encoding the at least one spatial parameter and the at least one audio signal.

[0035] Processing the intermediate format immersive audio signal to obtain at least one spatial parameter and to obtain the at least one audio signal may comprise determining an audio frame length difference between the intermediate format immersive audio signal and the at least one audio signal, and controlling buffering of the intermediate format immersive audio signal based on determining the audio frame length difference.

[0036] The method may further include obtaining a user rotation value, and generating the intermediate format immersive audio signal may comprise generating the intermediate format immersive audio signal further based on the user rotation value.

[0037] Generating the intermediate format immersive audio signal may comprise generating the intermediate format immersive audio signal further based on a predetermined or agreed upon user rotation value, and the further device may be configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal, the processing being based on the predetermined or agreed upon user rotation value and the obtained user rotation value for the at least one spatial audio rendering parameter.

[0038] According to a fourth aspect, there is provided a method for an apparatus for generating a spatialized audio output based on a user position, the method comprising: obtaining a user position value and a rotation value; obtaining at least one encoded audio signal and at least one spatial parameter, where the encoded at least one audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position value; and generating an output audio signal based on processing the encoded at least one audio signal, the at least one spatial parameter, and the user rotation value in six degrees of freedom.

[0039] The device may be operated by a user, and obtaining the user position value may include generating the user position value.

[0040] Obtaining the user position values ​​may include receiving the user position values ​​from a head-mounted device operated by the user.

[0041] Obtaining the encoded at least one audio signal and the at least one spatial parameter may comprise receiving the encoded at least one audio signal and the at least one spatial parameter from a further device.

[0042] The method may further comprise receiving a user position value and / or a user orientation value from the further device.

[0043] The method may comprise transmitting the user position value and / or the user orientation value to a further device, which may be configured to generate an intermediate format immersive audio signal based on the at least one input audio signal, the determined metadata, and the user position value, and to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal.

[0044] The encoded at least one audio signal may be low latency encoded at least one audio signal.

[0045] The intermediate format immersive audio signal may have a format selected based on encoding compressibility of the intermediate format immersive audio signal.

[0046] According to a fifth aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured, with the at least one processor, to at least cause the apparatus to obtain a user position value, obtain at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal, generate an intermediate format immersive audio signal based on the at least one input audio signal, the metadata, and the user position value, process the intermediate format immersive audio signal to obtain at least one spatial parameter and the at least one audio signal, and encode the at least one spatial parameter and the at least one audio signal, where the at least one spatial parameter and the at least one audio signal at least partially generate the spatialized audio output.

[0047] The apparatus may be further caused to transmit the encoded at least one spatial parameter and the at least one audio signal to a further device, the further device being configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal, the processing being based on the user rotation value and the at least one spatial audio rendering parameter.

[0048] The further device may be operated by a user and the device having the user position values ​​obtained can receive the user position values ​​from the further device.

[0049] The apparatus for obtaining user position values ​​may receive the user position values ​​from a head-mounted device operated by a user.

[0050] The device may further cause a user position value to be transmitted.

[0051] An apparatus for processing an intermediate format immersive audio signal to obtain at least one spatial parameter may be capable of generating a metadata-assisted spatial audio bitstream.

[0052] The device for encoding at least one spatial parameter and at least one audio signal is capable of generating an immersive audio and audio services bitstream.

[0053] The device for causing at least one spatial parameter and at least one audio signal to be coded may cause the at least one spatial parameter and the at least one audio signal to be coded with low latency.

[0054] An apparatus for processing an intermediate format immersive audio signal to obtain at least one spatial parameter and to obtain at least one audio signal, comprising: An audio frame length difference between the intermediate format immersive audio signal and the at least one audio signal can be determined, and buffering of the intermediate format immersive audio signal can be controlled based on the determination of the audio frame length difference.

[0055] The apparatus may further be capable of obtaining a user rotation value, and the apparatus for generating an intermediate format immersive audio signal may generate the intermediate format immersive audio signal further based on the user rotation value.

[0056] The device for generating the intermediate format immersive audio signal may generate the intermediate format immersive audio signal further based on a predefined or agreed upon user rotation value, and the further device may be configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal, the processing being based on the predefined or agreed upon user rotation value and the obtained user rotation value for the at least one spatial audio rendering parameter.

[0057] According to a sixth aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured, using the at least one processor, to cause the apparatus to obtain user position values ​​and rotation values, to obtain at least one encoded audio signal and at least one spatial parameter based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position values, and to generate an output audio signal based on processing the at least one encoded audio signal, the at least one spatial parameter and the user rotation values ​​in six degrees of freedom.

[0058] The device may be operated by a user and the device causing the user position value to be obtained may be configured to generate the user position value.

[0059] The apparatus for obtaining the user position values ​​may receive the user position values ​​from a head mounted device operated by a user.

[0060] The device for obtaining the at least one encoded audio signal and the at least one spatial parameter may cause receiving the at least one encoded audio signal and the at least one spatial parameter from a further device.

[0061] The device may be further adapted to receive user position values ​​and / or user orientation values ​​from a further device.

[0062] The device may cause the user position values ​​and / or the user orientation values ​​to be transmitted to a further device, and the further device may be configured to generate an intermediate format immersive audio signal based on the at least one input audio signal, the determined metadata and the user position values, and to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal.

[0063] The encoded at least one audio signal may be low latency encoded at least one audio signal.

[0064] The intermediate format immersive audio signal may have a format selected based on encoding compressibility of the intermediate format immersive audio signal.

[0065] According to a seventh aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus comprising: means for obtaining a user position value; means for obtaining at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal; means for generating an intermediate format immersive audio signal based on the at least one input audio signal, the metadata and the user position value; means for processing the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal; and means for encoding the at least one spatial parameter and the at least one audio signal, where the at least one spatial parameter and the at least one audio signal are configured to at least partially generate the spatialized audio output.

[0066] According to an eighth aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus including: means for obtaining user position values ​​and rotation values; means for obtaining at least one encoded audio signal and at least one spatial parameter, where the encoded at least one audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position values; and means for generating an output audio signal based on processing the encoded at least one audio signal, the at least one spatial parameter and the user rotation value in six degrees of freedom.

[0067] According to a ninth aspect, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] to cause an apparatus for generating a spatialized audio output based on a user position to at least obtain a user position value, obtain at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal, generate an intermediate format immersive audio signal based on the at least one input audio signal, the metadata and the user position value, process the intermediate format immersive audio signal to obtain at least one spatial parameter and the at least one audio signal, and encode the at least one spatial parameter and the at least one audio signal, wherein the at least one spatial parameter and the at least one audio signal are configured to at least partially generate the spatialized audio output.

[0068] According to a tenth aspect of the present application, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] to cause an apparatus to at least perform the following: obtain user position values ​​and rotation values; obtain at least one encoded audio signal and at least one spatial parameter, where the at least one encoded audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position values; and generate an output audio signal based on processing the at least one encoded audio signal, the at least one spatial parameter and the user rotation values ​​in six degrees of freedom.

[0069] According to an eleventh aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing at least an apparatus to execute: obtaining a user position value; obtaining at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal; generating an intermediate format immersive audio signal based on the at least one input audio signal, the metadata, and the user position value; processing the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal; and encoding the at least one spatial parameter and the at least one audio signal, wherein the at least one spatial parameter and the at least one audio signal are configured to at least partially generate a spatialized audio output.

[0070] According to a twelfth aspect, there is provided an apparatus for generating a spatialized audio output based on a user position, the apparatus comprising means configured to: obtain a user position value and a rotation value; obtain an encoded at least one audio signal and at least one spatial parameter, where the encoded at least one audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position value; and generate an output audio signal based on processing the encoded at least one audio signal, the at least one spatial parameter and the user rotation value in six degrees of freedom. A non-transitory computer-readable medium is provided comprising program instructions for causing the apparatus to execute the apparatus.

[0071] According to a thirteenth aspect, there is provided an apparatus comprising: an acquisition circuit configured to acquire a user position value, an acquisition circuit configured to acquire at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal, a generation circuit configured to generate an intermediate format immersive audio signal based on the at least one input audio signal, the metadata and the user position value, and a processing circuit configured to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal, and to encode the at least one spatial parameter and the at least one audio signal, wherein the at least one spatial parameter and the at least one audio signal are configured to at least partially generate the spatialized audio output.

[0072] According to a fourteenth aspect, there is provided an apparatus comprising: an acquisition circuit configured to acquire user position values ​​and rotation values; an acquisition circuit configured to acquire at least one encoded audio signal and at least one spatial parameter, where the encoded at least one audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on user position values; and a generation circuit configured to generate an output audio signal based on processing the encoded at least one audio signal, the at least one spatial parameter and the user rotation values ​​in six degrees of freedom.

[0073] According to a fifteenth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least perform the following: obtain a user position value; obtain at least one input audio signal and associated metadata enabling rendering of the at least one input audio signal; generate an intermediate format immersive audio signal based on the at least one input audio signal, the metadata, and the user position value; process the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal; and encode the at least one spatial parameter and the at least one audio signal, wherein the at least one spatial parameter and the at least one audio signal are configured to at least partially generate a spatialized audio output.

[0074] According to a sixteenth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least perform the following: obtain user position values ​​and rotation values; obtain at least one encoded audio signal and at least one spatial parameter; where the at least one encoded audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position values; and generate an output audio signal based on processing the at least one encoded audio signal, the at least one spatial parameter and the user rotation values ​​in six degrees of freedom.

[0075] The present invention includes means for performing the operations described above.

[0076] The present apparatus is configured to perform the method operations as described above.

[0077] A computer program comprising program instructions for causing a computer to carry out the method described above.

[0078] A computer program product stored on the medium may cause an apparatus to perform the methods described herein.An electronic device may comprise an apparatus as described herein.

[0079] A chipset may comprise the apparatus described herein.

[0080] SUMMARY OF THE PRESENT APPLICATION The embodiments of the present application aim to address problems associated with the state of the art. [Brief description of the drawings]

[0081] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1a] 1a and 1b show diagrammatically a suitable system of apparatus in which some embodiments may be implemented. [Figure 1b]1a and 1b show diagrammatically a suitable system of apparatus in which some embodiments may be implemented. [Diagram 2] FIG. 2 illustrates a schematic of an exemplary conversion between MPEG-I and IVAS frame rates. [Diagram 3] FIG. 3 illustrates a schematic of an edge layer and user equipment device suitable for implementing some embodiments. [Figure 4] FIG. 4 illustrates a flow diagram of an example operation of the edge layer and user equipment devices as illustrated in FIG. 3 according to some embodiments. [Diagram 5] FIG. 5 illustrates a flow diagram of an example operation of a system such as that shown in FIG. 2, according to some embodiments. [Figure 6] FIG. 6 illustrates a schematic diagram of the low latency render output shown in FIG. 2 in further detail, according to some embodiments. [Figure 7] FIG. 7 illustrates diagrammatically an exemplary device suitable for implementing the depicted apparatus. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0082] In the following, suitable apparatus and possible mechanisms for efficient delivery of edge-based rendering of 6DoF MPEG-I immersive audio are described in further detail.

[0083] Acoustic modeling, including early reflection modeling, diffraction, occlusion, and extensive source rendering as described above, can be computationally very demanding for even moderately complex scenes. For example, rendering a large scene (i.e., a scene with many reflective elements) Attempting to perform early reflection modeling for second or higher order reflections to generate outputs with high plausibility is very resource intensive. Therefore, there are significant benefits for content creators in defining flexibility for rendering complex scenes in 6DoF scenes. This is even more important when the rendering device does not have the resources to provide high quality rendering.

[0084] The solution to this rendering device resource problem is to provide rendering of complex 6DoF audio scenes at the edge. In other words, an edge computing layer is used to provide physical or perceptual acoustics based rendering for complex scenes with a high degree of plausibility.

[0085] In the following, a complex 6DoF scene refers to a 6DoF scene with multiple sources (which may be static or dynamic, moving sources) as well as a scene that includes complex scene geometry with multiple surfaces or geometric elements with reflective, occlusion and diffraction properties.

[0086] The application of edge-based rendering assists by enabling the offloading of computational resources required to render a scene, thus enabling consumption of highly complex scenes even on devices with limited computing resources, resulting in a wider target addressable market for MPEG-I content.

[0087] The challenge with such edge rendering approaches is to deliver the rendered audio to the listener efficiently, while reliably responsive to changes in the listener's orientation and position, and maintaining a sense of believability and immersion as the listener changes direction and position.

[0088] Edge rendering of the MPEG-I 6DOF Audio format requires different configuration or control than traditional rendering on the consuming device (such as on an HMD, mobile phone, or AR glasses). When 6DOF content is consumed on the local device via headphones, it is set to headphone output mode. However, headphone output is not optimal for conversion to a spatial audio format such as MASA (Metadata Assisted Spatial Audio), one of the proposed input formats for IVAS. Therefore, even though the end point consumption is headphones, the MPEG-I renderer output cannot be configured to output to "headphones" when consumed via IVAS-assisted edge rendering.

[0089] Therefore, there is a need to have a bitrate-efficient solution that allows edge-based rendering while preserving the spatial audio characteristics of the rendered audio and maintaining a perceived motion responsive to ear latency. This is important because edge-based rendering needs to be feasible on real-world networks (which have bandwidth constraints) and the user should not experience a delayed response to changes in the user's listening position and orientation. This is required to maintain the immersion and plausibility of 6DoF. The ear-to-ear motion latency in the following disclosure is the time required to produce the effect of a perceived audio scene change based on a change in head motion.

[0090] Current approaches attempt to provide fully pre-rendered audio with audio frame delay compensation based on head orientation changes that may be implemented at the edge layer rather than the cloud layer. Thus, the concept described in the embodiments herein is one configured to provide distributed free viewpoint audio rendering or providing apparatus and methods, where pre-rendering is performed at a first instance (network edge layer) based on the user's 6DoF position to provide efficient transmission and render a perceptually motivated transformation to 3DoF spatial audio that is binaurally rendered to the user based on the user's 3DoF orientation at a second instance (UE).

[0091] Thus, the embodiments described herein extend the capabilities of the UE (e.g., adding support for 6DoF audio rendering) and allocate the highest complexity decoding and rendering to the network, enabling efficient high quality transmission and rendering of spatial audio, while achieving lower motion-to-sound latency and thus more accurate and natural spatial rendering according to the user orientation than is achievable by rendering only at the network edge.

[0092] In the following disclosure, edge rendering refers to performing rendering (at least in part) on network edge-layer based computing resources connected to appropriate consuming devices via low-latency (or ultra-low-latency) data links, e.g., computing resources in close proximity to gNodeBs in a 5G cellular network.

[0093] The embodiments herein further propose, in some embodiments, for distributed six degree of freedom (i.e., the listener can move within the scene and the listener position is tracked) audio rendering and delivery, a method is proposed to encode and transmit the 3DoF immersive audio signal generated by the 6DoF audio renderer using a low latency communication codec to achieve bitrate-efficient low latency delivery of partially rendered audio while preserving spatial audio cues and maintaining responsive motion to ear latency.

[0094] In some embodiments, this is achieved by obtaining a user position (e.g., by receiving a position from the UE), obtaining at least one audio signal and metadata enabling 6DoF rendering of the at least one audio signal (MPEG-I EIF), and rendering an immersive audio signal (rendering the MPEG-I rendering to the HOA or LS) using the at least one audio signal, the metadata, and the user head position.

[0095] The embodiments further describe processing the immersive audio signal to obtain at least one spatial parameter and at least one audio carrier signal (e.g., as part of Metadata-Assisted Spatial Audio - MASA format), and encoding and transmitting the at least one spatial parameter and the at least one carrier signal to another device using an audio codec having low latency (such as an IVAS-based codec).

[0096] In some embodiments on other devices, the user head orientation (UE head tracking) can be obtained and the user head orientation and at least one spatial parameter can be used to render at least one transport signal into a binaural output (e.g., rendering IVAS format audio into an appropriate binaural audio signal output).

[0097] In some embodiments, the audio frame length for rendering the immersive audio signal is determined to be the same as the low latency audio codec.

[0098] Additionally, in some embodiments, if the audio frame length for immersive audio rendering cannot be the same as the low latency audio codec, a FIFO buffer is instantiated to accommodate the excess samples. For example, EVS / IVAS expects 960 samples per 20ms frame, which corresponds to 240 samples. If MPEG-I runs at a frame length of 240 samples, it is in lockstep and there is no additional intermediate buffering. If MPEG-I runs at 256, additional buffering is required.

[0099] In some further embodiments, changes in user position are delivered to the renderer at a frequency above a threshold frequency and below a threshold latency to obtain an acceptable user translation to sound latency.

[0100] In such a way, the embodiments enable resource-constrained devices to consume highly complex 6-DOF immersive audio scenes that would otherwise only be feasible on consuming devices with significant computational resources, and the proposed method enables consumption on devices with lower power requirements.

[0101] Furthermore, these embodiments allow for a more detailed rendering of the audio scene compared to rendering with limited computational resources.

[0102] Head movements in these embodiments are reflected instantly in the "on-device" rendering of low-latency spatial audio, and listener translations are reflected as soon as the listener position is relayed to the edge renderer.

[0103] Furthermore, even complex virtual environments that can be dynamically modified can be simulated with high quality virtual acoustics on resource-constrained consuming devices such as AR glasses, VR glasses, and mobile devices.

[0104] With respect to Figures 1a and 1b, an exemplary system in which an embodiment may be implemented is shown. This high-level overview of the end-to-end system for edge-based rendering can be divided into three main parts with respect to Figure 1a of the system. These three parts are cloud tier / server 101, edge tier / server 111, and UE 121 (or user equipment). Furthermore, the high-level overview of the end-to-end system for edge-based rendering can be divided into four parts with respect to Figure 1b of the system where cloud tier / server 101 is divided into content creation server 160 and cloud tier / server / CDN 161. The example of Figure 1b shows that content creation and encoding can be done in separate servers, and the generated bitstream is stored or hosted in an appropriate cloud server or CDN for access by an edge renderer depending on the UE location. The cloud tier / server 101 for the end-to-end system for edge-based rendering may be close to the UE location in the network or may be independent of it. Cloud tier / server 101 is a location or entity where 6DoF audio content is generated or stored. In this example, cloud tier / server 101 is configured to generate / store audio content in MPEG-I immersive audio for 6DoF audio format.

[0105] Thus, in some embodiments, the cloud tier / server 101 shown in Fig. 1a comprises an MPEG-I encoder 103, and in Fig. 1b the content creation server 160 comprises an MPEG-I encoder 103. The MPEG-I encoder 103 is configured to generate MPEG-I 6DoF audio content with the help of a content creator scene description or an Encoder Input Format (EIF) file, associated audio data (raw audio files and MPEG-H encoded audio files).

[0106] Additionally, cloud tier / server 101 includes an MPEG-I content bitstream output 105 as shown in Figure 1a, and cloud tier / server / CDN 161 includes an MPEG-I content bitstream output 105 as shown in Figure 1b. MPEG-I content bitstream output 105 is configured to output or stream the MPEG-I encoder output as an MPEG-I content bitstream over any available or suitable Internet Protocol (IP) network or any other suitable communication network.

[0107] The edge layer / server 111 is the second entity in the end-to-end system. The edge-based computing layer / server or node is selected based on the UE location in the network. This allows for the provisioning of minimum data link latency between the edge computing layer and the end user consumption device (UE 121). In some scenarios, the edge layer / server 111 may be co-located with the base station (e.g., gNodeB) to which the UE 121 is connected, which may result in minimal end-to-end latency.

[0108] In some embodiments as shown in FIG. 1b, where the cloud tier / server / CDN 161 includes an MPEG-I content bitstream output 105, the edge server 111 includes an MPEG-I content buffer 163 for storing the MPEG-I content bitstream (i.e., 6DoF audio scene bitstream) that it retrieves from the cloud or CDN 161. In some embodiments, the edge layer / server 111 includes an MPEG-I edge renderer 113. The MPEG-I edge renderer 113 is configured to retrieve the MPEG-I content bitstream from the cloud tier / server output 105 (or the cloud tier / server in general) or the MPEG-I content buffer 163, and is further configured to retrieve information about the user location (or more generally the consuming device or UE location) from the low latency render adapter 115. The MPEG-I edge renderer 113 is configured to render the MPEG-I content bitstream depending on the user location (x,y,z) information.

[0109] The edge layer / server 111 further comprises a low latency render adapter 115. The low latency render adapter 115 receives the output of the MPEG-I edge renderer 113; It is configured to convert the MPEG-I rendered output into a format suitable for efficient representation for low latency delivery and then output it to the consuming device or UE 121. Thus, the low latency render adapter 115 is configured to convert the MPEG-I output format into the IVAS input format 116.

[0110] In some embodiments, the low latency render adapter 115 may be another stage in the 6DoF audio rendering pipeline. Such an additional rendering stage may produce output that is natively optimized as input for the low latency delivery module.

[0111] In some embodiments, the post-edge / server 111 comprises an edge render controller 117 configured to perform necessary configuration and control of the MPEG-I edge renderer 113 according to renderer setting information received from a player application in the UE 121.

[0112] In these embodiments, the UE 121 is a consuming device used by a listener of the 6DoF audio scene. The UE 121 may be any suitable device. For example, the UE 121 may be a mobile device, a head mounted device (HMD), augmented reality (AR) glasses, or headphones with head tracking. The UE 121 is configured to obtain a user position / orientation. For example, in some embodiments, the UE 121 includes head tracking and position tracking to determine the user's position when the user is consuming the 6DoF content. The user's position 126 in the 6DoF scene is delivered from the UE to an MPEG-I edge renderer (via a low latency render adapter 115) located in the edge layer / server 111 to affect position transformation or modification for 6DoF rendering.

[0113] In some embodiments, a low latency spatial render receiver 123 is configured to receive the output of the low latency render adapter 115 and pass it to a head tracking spatial audio renderer 125.

[0114] The UE 121 may further include a head tracking spatial audio renderer 125. The head tracking spatial audio renderer 125 is configured to receive the output of the low latency spatial render receiver 123 and the user head rotation information and generate an appropriate output audio rendering based thereon. The head tracking spatial audio renderer 125 is configured to implement 3DOF rotational degrees of freedom to which a listener is generally more sensitive.

[0115] In some embodiments, the UE 121 includes a renderer control 127. The renderer control 127 is configured to initiate configuration and control of the edge renderer 113.

[0116] With respect to the Edge Renderer 113 and the Low Latency Render Adapter 115, several requirements follow for implementing edge-based rendering of MPEG-6 DoF audio content and delivering it to end users who may be connected over low-latency high-bandwidth links such as 5G ULLRC (Ultra-Low Latency Reliable Communications) links.

[0117] In some embodiments, the time frame length of MPEG-I 6DoF audio rendering is aligned with the low latency delivery frame length to minimize any intermediate buffering delay. For example, if the low latency transport format frame length is 240 samples (sampling rate 48KHz), in some circumstances, the MPEG-I renderer 113 is configured to operate with an audio frame length of 240 samples. An example of this is shown in FIG. 2 by the top half of the figure, where the MPEG-I output 201 is 240 samples per frame, the IVAS input 203 is also 240 samples per frame, and there is no frame length conversion or buffering.

[0118] 2, for example, the MPEG-I output 211 is 128, 256 samples per frame and the IVAS input 213 is 240 samples per frame. In these embodiments, a FIFO buffer 212 may be inserted whose input is from the MPEG-I output and whose output is to the IVAS input 213, thus performing frame length conversion or buffering.

[0119] In some embodiments, MPEG-I 6DoF audio should be rendered to an intermediate format instead of the default listening mode specified format, if necessary. The need for rendering to an intermediate format is to preserve important spatial properties of the renderer output. This allows faithful playback of the rendered audio with the necessary spatial audio cues when converting to a format more suitable for efficient, low latency delivery. This can therefore maintain the listener's sense of relevance and immersion in some embodiments.

[0120] In some embodiments, the configuration information from the renderer control 127 to the edge renderer control 117 is:

number

[0121] However, when the rendering_mode value is 1, the intermediate format involves further encoding and decoding using a low-latency codec, which is necessary to allow faithful reproduction of the spatial audio characteristics while transporting the audio over a low-latency efficient delivery mechanism. On the other hand, when the rendering_mode value is 2, the renderer output is generated according to the listening_mode value without further compression.

[0122] Thus, a direct format value of 2 is useful for networks where sufficient bandwidth and ultra-low latency network delivery pipes exist (e.g., for dedicated network slices with transmission delays of 1-4 ms). The method utilized in the indirect format with a rendering_mode value of 1 is suitable for networks with greater bandwidth constraints with low latency delivery. [Table 2] 6dof_audio_frame_length is the working audio buffer frame length for ingest and delivered as output. This can be expressed in terms of number of samples. Typical values ​​are 128, 240, 256, 512, 1024, etc.

[0123] In some embodiments, the sampling_rate variable (or parameter) indicates the value of the audio sampling rate per second. Some example values ​​can be 48000, 44100, 96000, etc. In this example, a common sampling rate is used for the MPEG-I renderer as well as the low latency transport. In some embodiments, each can have a different sampling rate.

[0124] In some embodiments, the low_latency_transfer_format variable (or parameter) indicates the low latency delivery codec, which may be any efficient representation codec for spatial audio suitable for low latency delivery.

[0125] In some embodiments, a low_latency_transfer_frame_length variable (or parameter) indicates the low latency delivery codec frame length in terms of number of samples. The low latency transfer formats and frame length possible values ​​are shown below. [Table 3]

[0126] In some embodiments, the intermediate_format_type variable (or parameter) indicates the type of audio output configured for the MPEG-I renderer when, for some reason, the rendering_mode needs to be converted to another format. One such motivation may be to have a format suitable for subsequent compression without reducing spatial properties. For example, an efficient representation for low latency delivery, i.e., rendering_mode value as 1. In some embodiments, there may be other motivations for conversion, which are described in more detail below.

[0127] In some embodiments, the end-user listening mode affects the audio rendering pipeline configuration and the component rendering stages. For example, if the desired audio output is to headphones, the final audio output can be directly synthesized as binaural audio. In contrast, for loudspeaker output, the audio rendering stages are configured to generate loudspeaker output without a binaural rendering stage.

[0128] In some embodiments, depending on the type of rendering_mode, the output of the 6DoF audio renderer (or MPEG-I immersive audio renderer) may differ from listening_mode. In such embodiments, the renderer may be configured to render the audio signal based on the intermediate_format_type variable (or parameter) to facilitate preserving salient audio characteristics of the MPEG-I renderer output when it is distributed via a low-latency efficient distribution format, for example, the following options may be employed: [Table 4]

[0129] Exemplary intermediate_format_type variable (or parameter) options for enabling edge-based rendering of 6DoF audio may be, for example, as follows: [Table 5] 3 and 4 present an example apparatus and flow diagram for edge-based rendering of MPEG-I 6DoF audio content with head-tracked audio rendering.

[0130] In some embodiments, the EDGE layer / server MPEG-I edge renderer 113 comprises an MPEG-I encoder input 301 configured to receive an MPEG-I encoded audio signal as input and pass it to an MPEG-I virtual scene encoder 303.

[0131] Further, in some embodiments, the EDGE layer / server MPEG-I edge renderer 113 comprises an MPEG-I virtual scene encoder 303 configured to receive the MPEG-I encoded audio signal and extract the virtual scene modeling parameters.

[0132] The EDGE layer / server MPEG-I edge renderer 113 further comprises an MPEG-I renderer to VLS / HOA 305. The MPEG-I renderer to VLS / HOA is configured to take the virtual scene parameters and the MPEG-I audio signal, and further the signal user transformations from the user position tracker 304, and generate an MPEG-I rendering in the VLS / HOA format (even in the case of headphone listening by the listener). The MPEG-I rendering is performed to the initial listener position.

[0133] The low latency render adapter 115 further comprises a MASA format converter 307. The MASA format converter 307 is configured to convert the rendered MPEG-I audio signal into a suitable MASA format, which can then be provided to the IVAS encoder 309.

[0134] The low latency render adapter 115 further comprises an IVAS encoder 309. The IVAS encoder 309 is configured to generate an encoded IVAS bitstream.

[0135] In some embodiments, the encoded IVAS bitstream is provided to the IVAS decoder 311 over an IP link to the UE.

[0136] In some embodiments, the UE 121 comprises a low latency spatial render receiver 123, which in turn comprises an IVAS decoder 311 configured to decode the IVAS bitstream and output it as a MASA format.

[0137] In some embodiments, the UE 121 comprises a head tracked spatial audio renderer 125, which in turn comprises a MASA format input 313. The MASA format input receives the output of the IVAS decoder and passes it to a MASA external renderer 315.

[0138] Further, the head tracking spatial audio renderer 125 comprises a MASA external renderer 315 configured in some embodiments to obtain head motion information from the user position tracker 304 and render an appropriate output format (e.g., binaural audio signals for headphones). The MASA external renderer 315 is configured to support 3DoF rotational degrees of freedom with minimal perceptible latency due to local rendering and head tracking. The user translated information as position information is sent back to the edge-based MPEG-I renderer. The position information and optionally rotation of the listener in the 6DoF audio scene in some embodiments is delivered as an RTCP feedback message. In some embodiments, the edge-based rendering delivers the rendered audio information to the UE. This allows the receiver to re-adjust its orientation before switching to a new translational position.

[0139] With reference to FIG. 4, an exemplary operation of the apparatus of FIG. 3 is shown.

[0140] First, in step 401, an MPEG-I encoder output as shown in FIG. 4 is obtained.

[0141] Next, via step 403, the virtual scene parameters are determined, as shown in FIG.

[0142] The user's position / orientation is obtained by step 404 as shown in FIG.

[0143] The MPEG-I audio is then rendered as VLS / HOA format by step 405 based on the virtual scene parameters and the user position, as shown in FIG.

[0144] The VLS / HOA format rendering is converted to MASA format by step 407 as shown in FIG.

[0145] The MASA format is IVAS encoded as shown in FIG.

[0146] The IVAS encoded (partially rendered) audio is then decoded, per step 411, as shown in FIG.

[0147] The decoded IVAS audio is then passed to head (rotation) associated rendering by step 413, as shown in FIG.

[0148] Next, via step 415, as shown in FIG. 4, the decoded IVAS audio is rendered head (rotation) relevant based on the user / head rotation information.

[0149] With reference to FIG. 5, a flow diagram of the operation of the apparatus of FIGS. 3 and 1 is shown in further detail.

[0150] 5 by step 501, an end user selects 6DoF audio content to be consumed. This can be represented by a URL pointer to a 6DOF content bitstream and an associated manifest (e.g., an MPD or Media Presentation Description).

[0151] Then, in some embodiments, as shown in FIG. 5 by step 503 (UE renderer controller), selects render_mode, which can be, for example, 0 or 1 or 2. If render_mode value 0 is selected, the MPD can be used to take the MPEG-I 6DoF content bitstream and render it in a renderer on the UE. If render_mode value 1 or 2 is selected, an edge-based renderer needs to be configured. Necessary information such as render_mode, low_latency_transfer_format, and associated low_latency_transfer_frame_length are signaled to the edge renderer controller. In addition, the end user consumption method, i.e., listener_mode, can also be signaled.

[0152] The UE may be configured to deliver configuration information represented by the following structure:

number

[0153] Step 505 (Edge Renderer Controller) determines a suitable intermediate format as the output format of the MPEG-I renderer in order to convert it to a low latency transport format while minimizing the loss of spatial audio characteristics, as shown in Figure 5. Various possible interim formats are listed above and a suitable interim format is selected.

[0154] Further, as shown in FIG. 5 by steps 507 and 509 in some embodiments, the method (edge ​​renderer controller) obtains supported time frame length information from the MPEG-I renderer (6dof_audio_frame_length) and the low latency transfer format (low_latency_transfer_frame_length).

[0155] Thereafter, as shown in FIG. 5 by step 511, an appropriate queuing mechanism (e.g., a FIFO queue) is implemented to handle the disparity in the audio frame length (determined to occur). Note that in order to successfully deliver 17 audio frames for every 16 MPEG-I renderer output frames, the low latency transport needs to have stricter operational constraints compared to the MPEG-I renderer. For example, the MPEG-I renderer has a frame length of 256 samples, but the IVAS frame length is only 240 samples. In a period, the MPEG-I renderer outputs 16 frames, i.e., 4096, and to avoid delay accumulation, the low latency transport is needed to deliver 17 frames of 240 sample size. This determines the conversion processing, coding, and transmission constraints for the low latency render adapter operation.

[0156] The rendering of the MPEG-I to the determined intermediate format (eg, LS, HOA) is illustrated in FIG.

[0157] The method may then include adding the rendered audio spatial parameters (eg, position and orientation of the rendered intermediate format), as illustrated in FIG. 5 by step 514.

[0158] In some embodiments, as illustrated in FIG. 5 by step 515, the MPEG-I renderer output is converted in an intermediate format to a low latency transport input format (eg, MASA).

[0159] The MPEG-I output rendered in MASA format is then encoded into an IVAS encoded bitstream, as shown in FIG. 5, by step 517 .

[0160] The IVAS encoded bitstream is delivered to the UE via the appropriate network bit pipe, as shown in FIG. 5, by step 519 .

[0161] As shown in FIG. 5, step 521 (UE) in some embodiments decodes a received IVAS encoded bitstream.

[0162] Additionally, in some embodiments as illustrated in FIG. 5 by step 523, the UE performs head-tracked rendering of the decoded output using three rotational degrees of freedom.

[0163] Finally, the UE sends the user translated information as a position feedback message as an RTCP feedback message, as shown in Figure 5 by step 525. The renderer continues to render the scene starting from step 513 at the new position obtained at 525. In some embodiments, if there is a mismatch in the user position and / or rotation information signals due to network jitter, appropriate smoothing in Doppler processing is performed.

[0164] With respect to Figure 6, an example deployment of edge-based rendering of MPEG-I 6DoF audio utilizing 5G network slices enabling ULLRC is shown. The top part of Figure 6 shows a conventional application such as when the MPEG-I renderer in the UE determines that it is unable to render an audio signal but is configured to determine a listening mode.

[0165] 6 illustrates an edge-based rendering apparatus and further illustrates in more detail an example of a low latency render adapter 115. The example low latency render adapter 115 is shown with an intermediate output 601 configured to receive the output of the MPEG-I edge renderer 113 instead of, for example, listening mode.

[0166] Additionally, the low latency render adapter 115 comprises a temporal frame length matcher 605 configured to determine whether there is a frame length difference between the MPEG-I output frames and the IVAS input frames and to implement appropriate frame length compensation as described above.

[0167] Additionally, low latency render adapter 115 is shown to include a low latency converter 607 configured to convert, for example, MPEG-I formatted signals into MASA formatted signals.

[0168] Further, the low latency render adapter 115 comprises a low latency (IVAS) encoder 609 configured to receive MASA or suitable low latency format audio signals and encode them before outputting them as a low latency encoded bitstream.

[0169] The UE as described above may be equipped with a suitable low latency spatial rendering receiver comprising a low latency spatial (IVAS) decoder 123 which performs head tracked rendering and outputs a signal to a head tracked renderer 125 which is further configured to output the listener position to an MPEG-I edge renderer 113.

[0170] In some embodiments, the listener / user (e.g., user equipment) is configured to pass the listener orientation value to the edge layer / server. However, in some embodiments, the edge layer / server is configured to implement low latency rendering for a default or predefined orientation. In such embodiments, rotation delivery may be skipped by assuming a constant orientation for the session, e.g., (0,0,0) for yaw-pitch roll.

[0171] A default or pre-defined orientation of the audio data is then provided to the listening device, with "local" rendering performing panning to the desired orientation based on the listener's head orientation.

[0172] In other words, the exemplary deployment leverages traditional network connectivity, represented by the black arrows in traditional MPEG-I rendering, as well as 5G-enabled ULLRC for edge-based rendering, and delivery of user position and / or orientation feedback to the MPEG-I renderer via a low-latency feedback link. Note that while the example shows 5G ULLRC, any other suitable network or communication link may be used. Since the feedback link is carrying feedback signals to generate the rendering, it is important that it does not require high bandwidth and prioritizes low end-to-end latency. While this example shows IVAS as the low-latency transport codec, other low-latency codecs may also be used. In one embodiment of the implementation, the MPEG-I renderer pipeline is extended to incorporate low-latency transport delivery as an embedded rendering stage of the MPEG-I renderer (e.g., block 601 may be part of the MPEG-I renderer).

[0173] In the above described embodiment, there is an apparatus for generating an immersive audio scene using tracking, which may also be known as an apparatus for generating spatialized audio output based on a listener or user position.

[0174] As detailed in the above embodiments, the objective of these embodiments is to perform high quality immersive audio scene rendering without resource constraints and make the rendered audio available to resource constrained playback devices. This can be done by leveraging edge computing nodes connected via low latency connections between the edge computing nodes and playback consuming devices. Low latency response is required to maintain responsiveness to user movements. Despite the low latency network connection, the embodiments aim to have low latency efficient encoding of the immersive audio scene rendering output in an intermediate format. The low latency encoding helps ensure that additional latency penalties are minimized due to efficient data transfer from the edge to the playback consuming devices. The low latency encoding is a relative value compared to the overall tolerable latency (including the immersive audio scene rendering at the edge node, the encoding of the intermediate format output from the edge rendering, the decoding of the encoded audio, and the transmission latency). For example, a speech codec can have encoding and decoding latency of up to 32 ms. On the other hand, there exist low latency encoding techniques that can be up to 1 ms. The criteria employed in codec selection in some embodiments is to have the transfer of the intermediate rendering output format to be delivered to the playback consuming device with minimal bandwidth requirements and minimal end-to-end coding latency.

[0175] With reference to FIG. 7, an exemplary electronic device is shown that may represent any of the devices shown above. The device may be any suitable electronic device or equipment operated by an end user. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. However, the exemplary electronic device may represent, at least in part, an edge layer / server 111 or a cloud layer / server 101 in a formation of a distributed computing resource.

[0176] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program code, such as the methods described herein.

[0177] In some embodiments, the device 1400 comprises a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 may be any suitable storage means. In some embodiments, the memory 1411 comprises program code sections for storing program code executable on the processor 1407. Furthermore, in some embodiments, the memory 1411 may further comprise a stored data section for storing data, e.g., data that has been processed or is to be processed according to the embodiments described herein. The executed program code stored in the program code sections and the data stored in the stored data sections may be retrieved by the processor 1407 via the memory-processor coupling as needed.

[0178] In some embodiments, the device 1400 comprises a user interface 1405. The user interface 1405 may be coupled to a processor 1407 in some embodiments. In some embodiments, the processor 1407 may control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 may allow a user to input commands to the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 may allow a user to obtain information from the device 1400. For example, the user interface 1405 may comprise a display configured to display information from the device 1400 to a user. The user interface 1405 may comprise a touch screen or touch interface in some embodiments that can both allow information to be input into the device 1400 and further display information to a user of the device 1400. In some embodiments, the user interface 1405 may be a user interface for communicating with a position determiner, as described herein.

[0179] In some embodiments, the device 1400 comprises an input / output port 1409. In some embodiments, the input / output port 1409 comprises a transceiver. The transceiver in such embodiments may be coupled to the processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may in some embodiments be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.

[0180] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an Infrared Data Path (IRDA).

[0181] The transceiver input / output port 1409 may be configured to receive signals and, in some embodiments, determine parameters as described herein by using a processor 1407 executing appropriate code.

[0182] Also, while exemplary embodiments have been described hereinabove, it should be noted that there are several variations and modifications that can be made to the disclosed solutions without departing from the technical scope of the present invention.

[0183] In general, various aspects may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present disclosure may be implemented in hardware, and other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present disclosure is not limited thereto. Although various aspects of the present disclosure may be illustrated and illustrated as block diagrams, flow charts, or using any other diagrammatic representation, it is fully understood that these blocks, devices, systems, techniques, or methods contemplated herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controller, or other computing device, or any combination thereof, as non-limiting examples.

[0184] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) Hardware-only circuit implementations (e.g., analog and / or digital-only circuit implementations); and (b) (where applicable) a combination of hardware circuitry and software, such as (i) a combination of analog and / or digital hardware circuitry and software / firmware; and (ii) Any portion of software (including digital signal processors), software, and hardware processors having memory that work together to cause a device, such as a mobile phone or server, to perform various functions. (c) Includes hardware circuitry and / or a processor, such as a microprocessor or portion of a microprocessor, that requires software (e.g., firmware) for operation, but the software may not be present when not needed for operation.

[0185] This definition of circuitry applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuitry encompasses merely a hardware circuit or processor (or processors), or a portion of a hardware circuit or processor, and its (or their) accompanying software and / or firmware implementations.

[0186] The term circuitry also encompasses, for example, baseband or processor integrated circuits for mobile devices or similar integrated circuits in servers, cellular network devices, or other computing or network devices, where applicable to particular claim elements.

[0187] The embodiments of the present disclosure may be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs, also called program products, including software routines, applets, and / or macros, may be stored in any device-readable data storage medium, which comprise program instructions for performing specific tasks. A computer program product may comprise one or more computer-executable components, which are configured to perform the embodiments when the program is executed. The one or more computer-executable components may be at least one software code or a portion thereof.

[0188] Further, in this regard, it should be noted that any block of logic flow as illustrated may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software may be stored on physical media such as memory chips, or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants CDs. Physical media are non-transitory media.

[0189] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based storage devices, magnetic storage devices and systems, optical memories and systems, fixed and removable memories, etc. The data processor may be of any type suitable for the local technology environment and may comprise, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), an FPGA, a gate level circuit, and a processor based on a multi-core processor architecture.

[0190] The embodiments of the present disclosure may be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is a large-scale, highly automated process. Complex and powerful software tools are available to convert logic level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0191] The scope of protection sought for the various embodiments of the present disclosure is indicated by the independent claims. The embodiments and features described herein that are not included in the scope of the independent claims should be interpreted as examples useful for understanding the various embodiments of the present disclosure.

[0192] The above description has provided a complete and informative description of the exemplary embodiments of the present disclosure, by way of non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art in view of the above description upon perusal of the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the present disclosure will still fall within the scope of the present invention as defined in the appended claims. Indeed, there are further embodiments that include combinations of one or more of the embodiments with any of the other embodiments discussed above.

Claims

1. 1. An apparatus for generating a spatialized audio output based on a user position, the apparatus comprising: obtaining a user position value; obtaining at least one input audio signal and associated metadata enabling rendering of said at least one input audio signal; generating an intermediate format immersive audio signal based on the at least one input audio signal, the metadata, and the user position value; processing the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal; encoding the at least one spatial parameter and the at least one audio signal, the at least one spatial parameter and the at least one audio signal being configured to at least partially generate the spatialized audio output; 10. An apparatus comprising means configured to perform the steps of:

2. the means is further configured to perform the step of transmitting the encoded at least one spatial parameter and the at least one audio signal to a further device; the further device is configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal; the processing is based on a user rotation value and the at least one spatial audio rendering parameter.

10. The apparatus of claim 1.

3. the further device is operated by a user; the means configured to obtain a user position value is configured to receive the user position value from the further device.

3. The apparatus of claim 2.

4. 3. The apparatus of claim 1 or 2, wherein the means configured to obtain the user position values ​​is configured to receive the user position values ​​from a head-mounted device operated by a user.

5. 4. The apparatus of claim 2 or 3, wherein said means is further configured to transmit said user position value.

6. 4. The apparatus of claim 1, wherein the means configured to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal are configured to generate a metadata-assisted spatial audio bitstream.

7. 4. The apparatus according to claim 1, wherein the means configured to encode the at least one spatial parameter and the at least one audio signal are configured to generate an immersive voice and audio service bitstream.

8. 4. The apparatus according to claim 1, wherein the means configured for encoding the at least one spatial parameter and the at least one audio signal are configured for low-latency encoding of the at least one spatial parameter and the at least one audio signal.

9. 4. The apparatus of claim 1, wherein the means configured to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal are configured to determine an audio frame length difference between the intermediate format immersive audio signal and the at least one audio signal, and to control buffering of the intermediate format immersive audio signal based on the determination of the audio frame length difference.

10. the means being further configured to obtain a user rotation value; the means configured to generate the intermediate format immersive audio signal is configured to generate the intermediate format immersive audio signal further based on the user rotation value.

4. An apparatus according to any one of claims 1 to 3.

11. the means configured to generate the intermediate format immersive audio signal is configured to generate the intermediate format immersive audio signal further based on a predetermined or agreed-upon user rotation value; the further device is configured to output a binaural or multi-channel audio signal based on processing the at least one audio signal; the processing is based on the predetermined or agreed upon user rotation value and the obtained user rotation value for the at least one spatial audio rendering parameter.

4. The device according to claim 2 or 3.

12. 1. An apparatus for generating spatialized audio output based on a user position, comprising: The apparatus obtains user position and rotation values, obtains an encoded at least one audio signal and at least one spatial parameter, the encoded at least one audio signal is based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position value; means configured to generate an output audio signal based on processing the encoded at least one audio signal and the at least one spatial parameter and a user rotation value in six degrees of freedom; Device.

13. the device is operated by a user; the means configured to obtain a user position value is configured to generate the user position value; 13. The apparatus of claim 12.

14. The apparatus described in claim 12 or 13, wherein the means configured to acquire a user position value is configured to receive the user position value from a head-mounted device operated by the user.

15. 14. The device according to claim 12 or 13, wherein the means configured to obtain the encoded at least one audio signal and the at least one spatial parameter are configured to receive the encoded at least one audio signal and the at least one spatial parameter from a further device.

16. 16. The device according to claim 15, wherein said means is further configured to receive said user position value and / or user orientation value from said further device.

17. said means being adapted to transmit said user position value and / or user direction value to said further device; the further device is configured to generate an intermediate format immersive audio signal based on at least one input audio signal, the determined metadata and the user position value, and to process the intermediate format immersive audio signal to obtain the at least one spatial parameter and the at least one audio signal.

16. The apparatus of claim 15.

18. 4. The apparatus of claim 1, wherein the intermediate format immersive audio signal has a format selected based on a coding compressibility of the intermediate format immersive audio signal.

19. 1. A method for an apparatus for generating spatialized audio output based on a user position, the method comprising: obtaining a user position value; obtaining at least one input audio signal and associated metadata enabling rendering of said at least one input audio signal; generating an intermediate format immersive audio signal based on the at least one input audio signal, the metadata, and the user position value; processing the intermediate format immersive audio signal to obtain at least one spatial parameter and at least one audio signal; encoding the at least one spatial parameter and the at least one audio signal, the at least one spatial parameter and the at least one audio signal being configured to at least partially generate the spatialized audio output; A method comprising:

20. 1. A method for an apparatus for generating spatialized audio output based on user position, comprising: obtaining user position and rotation values; obtaining at least one encoded audio signal and at least one spatial parameter, wherein the at least one encoded audio signal comprises: based on an intermediate format immersive audio signal generated by processing an input audio signal based on the user position value; generating an output audio signal based on processing the encoded at least one audio signal, the at least one spatial parameter, and a user rotation value with six degrees of freedom; A method comprising: