An apparatus and method for mapping immersive audio with coded audio

By determining and encoding audio bitstream identifiers within immersive audio bitstreams, the proposed solution addresses the issue of incorrect source mapping in existing codecs, ensuring accurate and immersive audio rendering.

WO2026098917A1PCT designated stage Publication Date: 2026-05-15NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2025-10-15
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing immersive audio codecs, such as MPEG-I Immersive audio and 3GPP IVAS, lack the ability to map audio sources at a granular level, leading to incorrect associations between immersive audio scenes and coded audio streams, which can result in erroneous rendering and a compromised immersive experience.

Method used

The proposed solution involves generating correspondence information between immersive audio bitstreams and coded audio bitstreams by determining audio bitstream identifiers based on audio sources within the scene, and including these identifiers in the bitstream configuration packets, such as PACTYP_MPEGI_CFG, to enable accurate mapping and rendering.

Benefits of technology

This approach ensures that audio sources in immersive audio scenes are correctly associated with their corresponding coded audio streams, resulting in a faithful rendering of the intended immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025079711_15052026_PF_FP_ABST
    Figure EP2025079711_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A method for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the method comprising: obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source; obtaining a coded audio bitstream for the immersive audio scene; determining at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AN APPARATUS AND METHOD FOR MAPPING IMMERSIVE AUDIO WITH CODED AUDIO

[0002] Field

[0003] The present application relates to apparatus and methods for mapping immersive audio with coded audio, and not exclusively for mapping immersive audio rendering with coded audio in augmented reality and / or virtual reality apparatus.

[0004] Background

[0005] Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is MPEG-I Immersive audio (ISO / IEC 23090-4) which is being currently standardized in ISO / IEC JTC1 SC29 WG6 (MPEG Audio coding WG). This codec is capable of delivery of immersive audio bitstream for 6DoF (six degrees of freedom) audio rendering where the audio coding technology can be, for example, MPEG-H 3DA (ISO / IEC 23008-3 codec audio signals such as audio objects, audio channels and Higher Order Ambisonics or HOA). Another example of such a codec is the future 3GPP immersive voice and audio services (IVAS) codec with 3DoF (three degrees of rotational freedom: yaw, pitch and roll) which is being designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing.

[0006] Summary

[0007] There is provided according to a first aspect a method for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the method comprising: obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source; obtaining a coded audio bitstream for the immersive audio scene; determining at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0008] Obtaining a coded audio bitstream for the immersive audio scene may comprise: obtaining at least one scene relevant audio signal; and generating the coded audio bitstream based on the at least one scene relevant audio signal. Determining audio bitstream identifiers with correspondence to audio source identifiers may comprise: generating correspondence information between the at least one scene relevant audio signal and the at least one audio source within the obtained audio scene information; and encoding the correspondence information within a bitstream further comprising the audio scene information and the at least one scene relevant audio signal.

[0009] The method may further comprise at least one of: storing the bitstream for retrieval by a renderer; transmitting the bitstream to a renderer.

[0010] Obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source may comprise obtaining or generating the immersive audio scene information within an encoder input format.

[0011] Obtaining at least one scene relevant audio signal may comprise: determining a coding method for the at least one scene relevant audio signal; encoding the at least one scene relevant audio signal based on the determined coding method.

[0012] Determining at least one audio bitstream identifier with correspondence to at least one audio source identifier may comprise determining at least one audio bitstream identifier based on the coding method for the at least one scene relevant audio signal.

[0013] Obtaining a coded audio bitstream for the immersive audio scene may comprise: determining a delivery method for the at least one scene relevant audio signal; encoding an audio bitstream based on the determined delivery method.

[0014] Determining the at least one audio bitstream identifier with correspondence to at least one audio source identifier may comprise determining the at least one audio bitstream identifier based on the delivery method for the at least one scene relevant audio signal.

[0015] The at least one audio bitstream identifier based on the coding method or delivery method for the at least one scene relevant audio signal may be a MHASPacketLabel of a MPEG-H 3D stream for the delivery method for the at least one scene relevant audio signal.

[0016] The at least one audio source may comprise at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

[0017] Including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier may comprise including the correspondence information within PACTYP_MPEGI_CFG.

[0018] According to a second aspect there is provided a method for assisting immersive audio rendering, the method comprising: obtaining a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtaining the immersive audio scene information from the bitstream; obtaining the coded audio bitstream from the bitstream; obtaining the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; mapping the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configuring a renderer based on the mapping; and rendering an output audio signal based on the configured renderer.

[0019] The method may further comprise at least one of: retrieving the bitstream; and receiving the bitstream.

[0020] The audio scene information may be an encoder input format.

[0021] The at least one audio bitstream identifier may be a MHASPacketLabel of a MPEG-H 3D stream.

[0022] The at least one audio source may comprise at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

[0023] The at least one audio bitstream identifier may be within PACTYP_MPEGI_CFG.

[0024] Configuring a renderer based on the mapping may comprise configuring at least one audio data path based on the mapping.

[0025] According to a third aspect there is provided an apparatus for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain immersive audio scene information, for defining an audio scene comprising at least one audio source; obtain a coded audio bitstream for the immersive audio scene; determine at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0026] The apparatus caused to obtain a coded audio bitstream for the immersive audio scene may be further caused to: obtain at least one scene relevant audio signal; and generate the coded audio bitstream based on the at least one scene relevant audio signal.

[0027] The apparatus caused to determine audio bitstream identifiers with correspondence to audio source identifiers may be caused to: generate correspondence information between the at least one scene relevant audio signal and the at least one audio source within the obtained audio scene information; and encode the correspondence information within a bitstream further comprising the audio scene information and the at least one scene relevant audio signal. The apparatus may further be caused to at least one of: store the bitstream for retrieval by a renderer; transmit the bitstream to a renderer.

[0028] The apparatus caused to obtain immersive audio scene information, for defining an audio scene comprising at least one audio source may be caused to obtain or generate the immersive audio scene information within an encoder input format.

[0029] The apparatus caused to obtain at least one scene relevant audio signal may be caused to: determine a coding method for the at least one scene relevant audio signal; encode the at least one scene relevant audio signal based on the determined coding method.

[0030] The apparatus caused to determine at least one audio bitstream identifier with correspondence to at least one audio source identifier may be caused to determine at least one audio bitstream identifier based on the coding method for the at least one scene relevant audio signal.

[0031] The apparatus caused to obtaining a coded audio bitstream for the immersive audio scene may be caused to: determine a delivery method for the at least one scene relevant audio signal; encoding an audio bitstream based on the determined delivery method.

[0032] The apparatus caused to determine the at least one audio bitstream identifier with correspondence to at least one audio source identifier may be caused to determine the at least one audio bitstream identifier based on the delivery method for the at least one scene relevant audio signal.

[0033] The at least one audio bitstream identifier based on the coding method or delivery method for the at least one scene relevant audio signal may be a MHASPacketLabel of a MPEG-H 3D stream for the delivery method for the at least one scene relevant audio signal.

[0034] The at least one audio source may comprise at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

[0035] The apparatus caused to include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier may be caused to include the correspondence information within PACTYP-MPEGLCFG.

[0036] According to a fourth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtain the immersive audio scene information from the bitstream; obtain the coded audio bitstream from the bitstream; obtain the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; map the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configure a renderer based on the mapping; and render an output audio signal based on the configured renderer.

[0037] The apparatus may further be caused to at least one of: retrieve the bitstream; and receive the bitstream.

[0038] The audio scene information may be an encoder input format.

[0039] The at least one audio bitstream identifier may be a MHASPacketLabel of a MPEG-H 3D stream.

[0040] The at least one audio source may comprise at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

[0041] The at least one audio bitstream identifier may be within PACTYP_MPEGI_CFG.

[0042] The apparatus caused to configure the renderer based on the mapping may be caused to configure at least one audio data path based on the mapping.

[0043] According to a fifth aspect there is provided an apparatus for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus comprising means configured to: obtain immersive audio scene information, for defining an audio scene comprising at least one audio source; obtain a coded audio bitstream for the immersive audio scene; determine at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0044] The means configured to obtain a coded audio bitstream for the immersive audio scene may be further configured to: obtain at least one scene relevant audio signal; and generate the coded audio bitstream based on the at least one scene relevant audio signal.

[0045] The means configured to determine audio bitstream identifiers with correspondence to audio source identifiers may be configured to: generate correspondence information between the at least one scene relevant audio signal and the at least one audio source within the obtained audio scene information; and encode the correspondence information within a bitstream further comprising the audio scene information and the at least one scene relevant audio signal.

[0046] The means may further be configured to at least one of: store the bitstream for retrieval by a renderer; transmit the bitstream to a renderer. The means configured to obtain immersive audio scene information, for defining an audio scene comprising at least one audio source may be configured to obtain or generate the immersive audio scene information within an encoder input format.

[0047] The means configured to obtain at least one scene relevant audio signal may be configured to: determine a coding method for the at least one scene relevant audio signal; encode the at least one scene relevant audio signal based on the determined coding method.

[0048] The means configured to determine at least one audio bitstream identifier with correspondence to at least one audio source identifier may be configured to determine at least one audio bitstream identifier based on the coding method for the at least one scene relevant audio signal.

[0049] The means configured to obtaining a coded audio bitstream for the immersive audio scene may be configured to: determine a delivery method for the at least one scene relevant audio signal; encoding an audio bitstream based on the determined delivery method.

[0050] The means configured to determine the at least one audio bitstream identifier with correspondence to at least one audio source identifier may be configured to determine the at least one audio bitstream identifier based on the delivery method for the at least one scene relevant audio signal.

[0051] The at least one audio bitstream identifier based on the coding method or delivery method for the at least one scene relevant audio signal may be a MHASPacketLabel of a MPEG-H 3D stream for the delivery method for the at least one scene relevant audio signal.

[0052] The at least one audio source may comprise at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

[0053] The means configured to include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier may be configured to include the correspondence information within PACTYP-MPEGLCFG.

[0054] According to a sixth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising means configured to: obtain a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtain the immersive audio scene information from the bitstream; obtain the coded audio bitstream from the bitstream; obtain the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; map the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configure a renderer based on the mapping; and render an output audio signal based on the configured renderer.

[0055] The means may further be configured to at least one of: retrieve the bitstream; and receive the bitstream.

[0056] The audio scene information may be an encoder input format.

[0057] The at least one audio bitstream identifier may be a MHASPacketLabel of a MPEG-H 3D stream.

[0058] The at least one audio source may comprise at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

[0059] The at least one audio bitstream identifier may be within PACTYP_MPEGI_CFG.

[0060] The means configured to configure the renderer based on the mapping may be configured to configure at least one audio data path based on the mapping.

[0061] According to a seventh aspect there is provided an apparatus generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus comprising means configured to: obtain immersive audio scene information, for defining an audio scene comprising at least one audio source; obtain a coded audio bitstream for the immersive audio scene; determine at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0062] According to an eighth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising means configured to: obtain a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtain the immersive audio scene information from the bitstream; obtain the coded audio bitstream from the bitstream; obtain the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; map the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configure a renderer based on the mapping; and render an output audio signal based on the configured renderer.

[0063] According to a ninth aspect there is provided an apparatus for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus comprising: obtaining circuitry configured to obtain immersive audio scene information, for defining an audio scene comprising at least one audio source; obtaining circuitry configured to obtain a coded audio bitstream for the immersive audio scene; determining circuitry configured to determine at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; including circuitry configured to include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0064] According to a tenth aspect there is provided an apparatus for assisting immersive audio rendering, the apparatus comprising: obtaining circuitry configured to obtain a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtaining circuitry configured to obtain the immersive audio scene information from the bitstream; obtaining circuitry configured to obtain the coded audio bitstream from the bitstream; obtaining circuitry configured to obtain the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; mapping circuitry configured to map the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configuring circuitry configured to configure a renderer based on the mapping; and rendering circuitry configured to render an output audio signal based on the configured renderer.

[0065] According to an eleventh aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source; obtaining a coded audio bitstream for the immersive audio scene; determining at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0066] According to a twelfth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtaining the immersive audio scene information from the bitstream; obtaining the coded audio bitstream from the bitstream; obtaining the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; mapping the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configuring a renderer based on the mapping; and rendering an output audio signal based on the configured renderer.

[0067] According to a thirteenth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus, for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source; obtaining a coded audio bitstream for the immersive audio scene; determining at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0068] According to a fourteenth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtaining the immersive audio scene information from the bitstream; obtaining the coded audio bitstream from the bitstream; obtaining the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; mapping the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configuring a renderer based on the mapping; and rendering an output audio signal based on the configured renderer. According to a fifteenth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining a coded audio bitstream for the immersive audio scene; determining at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

[0069] According to a sixteenth aspect there is provided a computer readable medium comprising instructions for causing an apparatus for assisting immersive audio rendering, the apparatus caused to perform at least the following: obtaining a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtaining the immersive audio scene information from the bitstream; obtaining the coded audio bitstream from the bitstream; obtaining the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; mapping the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configuring a renderer based on the mapping; and rendering an output audio signal based on the configured renderer.

[0070] An apparatus comprising means for performing the actions of the method as described above.

[0071] An apparatus configured to perform the actions of the method as described above.

[0072] A computer program comprising program instructions for causing a computer to perform the method as described above.

[0073] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.

[0074] An electronic device may comprise apparatus as described herein.

[0075] A chipset may comprise apparatus as described herein.

[0076] Embodiments of the present application aim to address problems associated with the state of the art.

[0077] Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:

[0078] Fig.1 shows schematically an example interaction between the entities such as immersive audio bitstream, coded audio data, renderer and player;

[0079] Figs.2a to 2d shows example MHAS packet structures (reference: ISO / IEC 23008-3);

[0080] Figs.3a to 3d shows further example MHAS packet structures (reference: ISO / IEC 23090-4);

[0081] Fig.4 shows an example mapping using a single common label;;

[0082] Fig.5 shows a flow diagram of the example mapping operation according to some embodiments;

[0083] Fig.6 shows a flow diagram of the example encoding of the correspondence information operation as shown in Fig.5;

[0084] Fig.7 shows an example system within which some embodiments can be implemented; and

[0085] Fig.8 shows an example device suitable for implementing the apparatus shown in previous figures.

[0086] Embodiments of the Application

[0087] The following describes in further detail suitable apparatus and possible mechanisms for describing the correspondence between the audio sources defined in an immersive audio coded bitstream and the audio streams related to the audio sources based on the audio stream identifiers. In such embodiments this correspondence enables a player or renderer entity or a presentation engine to render the one or more audio sources in the audio scene with the correct decoded audio signal from the one or more audio streams (for e.g., multiple audio sources may be encoded as part of the same audio stream as multiple audio elements or multiple audio streams with one or more audio elements per each stream).

[0088] With respect to Fig.1 is shown a schematic view of interactions between different entities such as the bitstream, renderer and player.

[0089] In this example the bitstreams 100 are obtained, and can comprise MPEG-H 3D audio coded audio data 102 and also a MPEG-I immersive audio bitstream 104. These can be passed to the player or player engine (PE) 101. In other words as shown in Fig.1 there can be employed different bitstreams corresponding to 6DoF rendering versus a coded audio data bitstream. The bitstreams can be stored or hosted in a server or any suitable system of apparatus.

[0090] The MPEG-I immersive audio bitstream 104 and MPEG-H 3D audio bitstream 102 can be obtained or received as MHAS (MPEG-H Audio Stream) packets.

[0091] The player or PE 101 having received the MPEG-H 3D audio bitstream 102 and a listener position 114 from a listener position determiner 103 is configured to generate MPEG-H 3D audio coded audio bitstream 106 which is passed to an MPEG-H 3DA decoder 105, which is configured to output a MPEG-H decoded audio bitstream 108 and pass this to an MPEG-I audio decoder and renderer 105. The player or PE 101 having received the MPEG-I immersive audio bitstream 104 is further configured to output the MPEG-I immersive audio bitstream 104 which is then passed to the MPEG-I audio decoder and renderer 105.

[0092] The MPEG-I audio decoder and renderer 105 is then configured to render a suitable audio output 112 based on the obtained MPEG-I immersive audio bitstream 104, the MPEG-H decoded audio bitstream 108 and the listener position 114 from the listener position determiner 103.

[0093] MPEG-H 3DA MHAS packets are described in ISO / IEC 23090-4. Some example packet types are the Audio Codec Configuration (CFG), Audio Scene Information (ASI), 3D Audio frame data (3DA).

[0094] The MPEG-H 3DA can be employed for encoding and decoding audio data when using MPEG-I immersive audio. The MPEG-H 3DA codec however is not employed for rendering the audio data when used together with MPEG-I immersive audio. The MPEG-H 3DA bitstream does not carry explicit timestamps because the effective timestamp in these can be derived based on the number of audio samples and the sampling frequency

[0095] Figs 2a to 2d for example show example structures of MHAS packets.

[0096] For example Fig.2a shows an example MHAS packet stream comprising a time instance i-1 MHAS packet 200, time instance i MHAS packet 202, and time instance i+1 MHAS packet 204. Each of the packet instances furthermore comprises a MHAS header 206 and MHAS payload 208.

[0097] Additionally as shown by Fig.2b is an example MHAS header 206 (which may be between 4-15 bytes in length) in further detail. In this example the MHAS header 206 comprises a type 210 field indicating the type of the MHAS packet, a label field 212 for identifying the packet, and a length field 214 identifying the length of the packet.

[0098] Fig.2c shows a code definition for the example MHAS packet stream and Fig.2d the code definition for a MHAS packet example.

[0099] The MPEG-H 3D audio MHAS packets are delivered in a specific sequence as shown in the examples to ensure the flexibility of change in decoder configuration in each sample (or audio frame to be rendered). The configuration packet, if present, shall precede the audio scene information MHAS packets and the audio frame data MHAS packets. The Audio scene information MHAS packet if present shall precede the audio frame data in order to ensure appropriate rendering by the renderer.

[0100] Furthermore as shown with respect to Figs 3a to 3d the MPEG-I Audio (6DoF rendering) can be specified as a set of packets to carry information such as configuration of renderer, scene information rendering metadata, scene changes. As shown in Fig.3a is shown an example MHAS packet structure for MPEG-I bitstream can for example comprise a bitstream of MHAS packets preceded by a MHAS header 300. The MHAS header comprises a bitstream identifier 310 and bitstream version 312. The packets 302, 304, 306 and 308, furthermore can comprise a MHAS packer header 320, MHAS packet payload 322 and Byte_Align* pointer 324.

[0101] Fig.3b shows the example MHAS header 320 (which may be 9 bytes in length) in further detail. In this example the MHAS header 320 comprises a type 330 field indicating the type of the MHAS packet, a label field 332 for identifying the packet, and a length field 334 identifying the length of the packet.

[0102] Fig.3c shows a code definition for the example MHAS packet for MPEG-I stream and Fig.3d the code definition for a MHAS packet for MPEG-I example.

[0103] Example new MHAS packet types for MPEG-I implementations include: Scene Configuration (PACTYP-MPEGLCFG), Scene Update (PACTYP_MPEGI_UPD), Rendering payload (PACTYP-MPEGLPLD).

[0104] Furthermore, MPEG-I packets also carry timestamps (e.g., to specify when a particular update is applied). The MPEG-I immersive audio scene can be a relatively static scene with some dynamic aspects (e.g., animation which is predefined or known only during runtime). For example, a room where most of the content storyline occurs. Thus, MPEG-I an immersive audio bitstream (even if large) is required during startup (for example an entire scene is needed during startup). There can be a continuous set of scene changes (which can include rendering metadata changes) over a period of time, and these can be delivered after the start of the playback of the scene. For example, streaming of a live performance (data is needed at startup but also over the entire duration of the scene).

[0105] MPEG-H 3DA (audio coding) codec can be employed for encoding audio data in conjunction with MPEG-I immersive audio.

[0106] As per the current specification, a parameter MHASPacketlabel present in both the MPEG-H 3DA MHAS packet header and the MPEG-I immersive audio MHAS packet header can be employed to map the MPEG-I Immersive audio (MIA) streams to the corresponding MPEG-H 3D Audio (MAH) streams.

[0107] This mechanism of having a common MHASPacketLabel between MPEG-I and MPEG-H transport streams may be sufficient in order to be able to associate audio bitstreams carry coded audio data and MPEG-I immersive audio bitstream. For example, MPEG-I immersive audio bitstream corresponding to a particular scene can be associated with MPEG-H 3DA bitstreams corresponding to the particular scene. However, this can be problematic in situations where a mapping is expected between the MPEG-I scene elements (e.g., audio sources in the immersive audio scene) and the MPEG-H 3DA audio signals (Objects, Channels and HOA signals) corresponding to one or more audio sources in the scene. In other words, a scene level correspondence is possible with this common label between MPEG-H 3DA streams and MPEG- I immersive audio streams but greater granularity (at the level of audio sources) than that with the state-of- the-art approach in the ISO / IEC 23090-4 is not possible. This can be shown for example in Fig.4 where employing a single common label to create correspondence between MPEG-H 3DA audio stream and MPEG-I objects is not sufficient. Thus for example there is the SYN or sync 400 packet which is a MPEG- H 3DA MHAS packet 430. This is followed by MPEG-I audio MHAS packets Scene Configuration (CFG) 402, Rendering payload (PLD) 404 and Scene Update (UPD) 406 packets. Then follows further MPEG-H 3DA MHAS packets configuration (CFG) 408 and 3DA payload packets 410 and 412

[0108] Consequently, there can be examples where bitstream identifiers of the MPEG-H 3D audio are not associated with any of the audio sources in the MPEG-I audio scene. Hence, the following examples show additional mechanisms for mapping the object sources, channel sources and HOA sources in the MPEG-I coded audio scene bitstream with relevant MPEG-H audio elements or MAH streams with relevant identifiers.

[0109] Otherwise the MPEG-I immersive audio and MPEG-H 3D audio cannot be interoperable with each other. Furthermore, the embodiments permit or enable delivery of coded audio data corresponding to the audio sources in the immersive audio scene.

[0110] The concept as discussed in further by the following embodiments relates to a representation of 6DoF audio scene and proposes suitable methods, apparatus or computer programs (operating on the apparatus) configured to describe or assist defining a correspondence between the audio sources defined in the immersive audio coded bitstream and the audio streams related to the audio sources based on the audio stream identifiers which enable a player or renderer entity or a presentation engine to render the one or more audio sources in the audio scene with the correct decoded audio signal from the one or more audio streams.

[0111] For example in some embodiments multiple audio sources may be encoded as part of the same audio stream as multiple audio elements or multiple audio streams with one or more audio elements per each stream. The proposed signaling of correspondence information in the immersive audio bitstream enables efficiency and flexibility of encoding the one or more audio sources into one or more encoded audio streams.

[0112] This can, for example, be achieved in some embodiments by extending the immersive audio scene configuration packet to carry at least one coded audio bitstream identifier corresponding to each of the audio source identifiers in the audio scene.

[0113] In some embodiments, the audio sources in the immersive audio coded bitstream can correspond to at least one of the audio object sources, channel sources, HOA sources in described in the immersive audio scene where the coded audio bitstream identifier can be, for example, the MHASPacketLabel of the MPEG- H 3DA (ISO / IEC 23008-3) streams.

[0114] In some further embodiments, a correspondence description between the audio sources in the immersive audio scene and the coded audio streams is achieved using the MHASPacketLabel of the MHAS packet header for the audio stream such that the MHASPacketLabel is constrained to be the audio source identifier. In such a manner an association, mapping or correspondence between the audio source in the audio scene with the audio stream can be defined. In some other embodiments, a single coded audio bitstream can correspond to a single audio source in the immersive audio scene. Furthermore, a single coded audio stream can correspond to a single or multiple audio source(s) in the immersive audio scene. This is feasible, for example, with MPEG-H 3DA bitstreams, where an MHAS stream can carry one or more audio elements.

[0115] The correspondence information between the coded audio data and the audio sources in the immersive audio scene, in some embodiments, can be carried as part of the scene configuration packet, the scene payload MHAS packet or the scene update MHAS packet.

[0116] In such a manner an example scenario can be considered, without loss of generality, to describe the technical effect of the invention.

[0117] In this example scenario a content creator can author an immersive audio scene with 3 audio sources (speech for a person as the first source, music for a TV as the second course, fan sound for a fan in the VR scene). The coded audio data for the three audio sources can be delivered as three audio objects as part of a single MPEG-H 3D audio track.

[0118] In a conventional application (in other words without employing the following embodiments) and using the current solution in the MPEG-I immersive audio standard, all the audio objects carried as coded audio object data in MPEG-H 3DA stream can be successfully associated with the immersive audio scene scene. However, due to the absence of any further information, there is no possibility to correctly associate and auralize the first audio source with the speech audio object, the second audio source with the music audio object and the fan audio source with the fan audio object. A mismatch could therefore result in a rendered audio scene which could be totally erroneous and break the immersive experience.

[0119] In employing the following embodiments, the audio sources are able to be associated with the correct audio objects and thus the rendered output audio signals enable the listener to experience the immersive scene as specified by the scene creator.

[0120] With respect to Fig.5 is shown a flow diagram of example operations according to some embodiments.

[0121] For example, as shown in Fig.5 by 501 , is the operation of receiving audio scene information and the scene relevant audio data.

[0122] The audio scene information which may be in any suitable format such as Encoder Input Format as described in MPEG-I immersive audio Encoder Input Format, version 9 , N00248, the associated document management number is MDS23718. or any other suitable format is obtained together with the relevant audio data. The audio data is related to the audio sources in the immersive audio scene. These audio sources can be audio objects, channels, HOA sources, or any other spatial audio format.

[0123] Following this, as shown in Fig.5 by 503, is the operation of determining a coding method for audio data. The audio data needs to be coded in a suitable format to facilitate efficient streaming and storage. The selected coding method directly influences the representation and delivery of coded audio data. For example, in embodiments where a multichannel codec such as MPEG-H 3DA is employed, then 5.1 channel or HOA sources can be coded into a single bitstream. In some examples where a mono codec is employed, then each of the channels of 5.1 or HOA sources can be individually encoded. Thus, the selection of the coding method directly affects the correspondence information that is determined and also signaled.

[0124] After the selection of the coding method is the operation, as shown in Fig.5 by 505, of preparing the coded version of the scene relevant audio data. Preparation of the coded audio data provides the possibility of knowing or determining or even selecting the audio bitstream identifiers for the encoder. This information is then subsequently employed to generate the correspondence information.

[0125] In some embodiments in the scene specific EIF file, there is determined or generated an association between the uncompressed audio file and the audio sources with the help of the audio file name. This is converted to bitstream information during streaming of the immersive audio scene bitstream and the coded audio bitstreams, which do not carry the physical file names.

[0126] Furthermore as shown in Fig.5 by step 507 is the operation of determining audio source and coded audio correspondence. In this operation coded audio bitstream identifiers are obtained, determined or otherwise after the encoding of the uncompressed audio data. The correspondence information can then be subsequently included in the immersive audio scene configuration data structure. This therefore enables or facilitates an initialization of the renderer and the related audio data paths in the playback device.

[0127] Having generated or determined the audio source and coded audio correspondence information then is the operation of encoding the correspondence information as shown in Fig.5 by 509. In some embodiments the generated correspondence information is encoded and included as part of the immersive audio scene specific bitstream. In some embodiments where MHAS transmission is employed, the correspondence information can be included in a suitable configuration message such as a PACTYP_MPEGI_CFG packet.

[0128] Additionally is shown in Fig.5 by 511 the operation of configuring audio data paths for the renderer for each audio source based on the correspondence information in the bitstream. This operation is implemented as the correspondence information is known at the level of the bitstream identifier. Consequently, the player configures the audio data paths between the audio stream decoder and the correct corresponding audio sources. This is ensured by provided coded audio data bitstreams to the correct decoder instance, based on the decoder audio data path connection to the audio source(s).

[0129] Finally, as shown in Fig.5 by step 513 is the operation of rendering the audio output signal. In this operation the audio is rendered by the immersive audio renderer based on the immersive audio scene specific bitstream and the received decoded audio data.

[0130] With respect to Fig.6 is shown in further detail the operation of encoding the correspondence information in the bitstream as shown in Fig.5 by 509 according to some embodiments. For example, initially is shown the operation in Fig.6 by 601 of obtaining the immersive audio scene description.

[0131] Furthermore is shown, in Fig.6 by 603, the operation of obtaining coded audio bitstream for the immersive audio scene.

[0132] Following this is the operation, as shown in Fig.6 by 605, of including coded audio bitstream identifiers with correspondence to audio source identifiers to be included in the immersive audio scene bitstream with correspondence to audio source identifiers. This is implemented based on the audio sources in the immersive audio scene.

[0133] Then is the operation, as shown in Fig.6 by 607, the operation of encoding the correspondence information in the configuration packet in the scene specific bitstream.

[0134] Fig.7 shows schematically an example system where the embodiments are implemented in an encoder device 703 on a suitable sever computer 701 which performs part of the functionality; writes data into a bitstream and which can be transmitted via a network 711 transmits that for a playback device 721 , which decodes the bitstream, performs reverberator processing according to the embodiments and outputs audio for headphone 791 listening.

[0135] The encoder side 703 of Fig.7 can be performed on content creator computers and / or network server computers 701. The output of the encoder is the immersive audio and coded audio bitstream 710 which is made available for downloading or streaming, for example via the network 711.

[0136] The decoder / renderer 741 functionality runs on an end-user-device or playback device 721 , which can be a mobile device, personal computer, sound bar, tablet computer, car media system, home HiFi or theatre system, head mounted display for AR or VR, smart watch, or any suitable system for audio consumption.

[0137] The encoder 703 is configured to receive the encoder input format or any other scene description information (for example, Scene.XML) 702 and the audio signals 704. Generally, the scene description contains an acoustically relevant description of the contents of the scene, and contains, for example, the scene geometry as a mesh or as voxels, acoustic materials, acoustic environments with reverberation parameters, positions of sound sources, and other audio element related parameters such as whether reverberation is to be rendered for an audio element or not.

[0138] The encoder 703 in some embodiments comprises an MPEG-H 3D audio encoder 707 configured to receive the audio signals 704 and is configured to determine the coding and delivery method for the scene relevant audio data (such as shown in Fig.5 by 503) and prepare coded version of the scene relevant audio data (as shown in Fig.5 by 505) and generate the encoded audio bitstream 708.

[0139] The encoder 703 in some embodiments comprises an Immersive audio encoder (e.g., MPEG-I Immersive audio) 705 which is configured to receive the scene description (for example in a format such as El F) 702 and the encoded audio / or audio signals and determine any correspondence information between the coded audio data and the audio sources in the received audio scene information and subsequently, generate representation for correspondence information (as shown in Fig.5 by 507). Additionally this correspondence information is encoded with the other scene description information (as shown in Fig.5 by 509). The Immersive audio encoder (e.g., MPEG-I Immersive audio) 705 thus outputs the MPEG-I immersive audio bitstream with mapping (or correspondence) information 706.

[0140] The encoded audio bitstream 708 and the MPEG-I immersive audio bitstream with mapping information 706 can then be combined to form the Immersive audio and coded audio bitstream 710.

[0141] In other words there can be employed two encoders. The first is the audio data compression encoder (e.g., MPEG-H 3D audio codec) and the second is the immersive audio rendering (e.g., MPEG-I immersive audio codec) related encoder to generate the scene specific immersive audio bitstream. In the embodiments as described herein, the encoder pipeline is structured such that the coded audio bitstream is available prior to immersive audio bitstream encoding. This is done to facilitate the inclusion of the mapping information between the coded audio data and the corresponding audio sources in the immersive audio scene.

[0142] The decoder 1941 in some embodiments comprises a bitstream decoder 1951 configured to decode the bitstream.

[0143] The playback device 721 can comprise a presentation engine 731 or PE which is configured to receive the coded audio data bitstream (MPEG-H 3DA) comprising the scene information with mapping information and pass this to a decoder for coded audio data bitstream to generate the decoded audio (MPEG- H 3DA) and pass this to a decoded audio buffer (MPEG-H 3DA) 755 and decode the mapping or correspondence information and pass this to a MPEG-I and MPEG-H mapping determiner / MPEG-l decoder 743.

[0144] The PE 731 is further configured to receive or otherwise obtain the immersive audio bitstream (MPEG-I) 751 and also pass this to the MPEG-I and MPEG-H mapping determiner / MPEG-l decoder 743.

[0145] The PE 731 further comprises a renderer 741 , which in turn comprises the MPEG-I and MPEG-H mapping determiner / MPEG-l decoder 743 which is then configured to determine the MPEG-I and MPEG-H mapping based on the mapping or correspondence information. This mapping determination can then be furthermore passed back to the decoded audio buffers (MPEG-H 3DA) 755 for enabling a configuration of an audio data path for the renderer for each audio source (as shown in Fig.5 by 511). The renderer pipeline 745 of the renderer 757 is configured to receive the decoded MPEG-I data and the MPEG-H 3DA data as well as head pose (position and orientation) generated by the head pose generator 757 and generate a renderer output audio delivered with aligned audio-visual entities and pass this to the output buffers 759 prior to output to the headset, headphones, or suitable transducer system (shown as head mounted device HMD 791) for outputting the audio scene to the listener. Thus with respect to the playback side, there is a module or suitable means, for example the decoder for Coded audio data bitstream (MPEG-H 3DA) 753, configured to inspect the immersive audio scene specific bitstream to extract the mapping information and use that information (for example the MPEG- I and MPEG-H mapping determiner 743) to configure the audio data decoding pipeline such that the right audio bitstream is fed to auralize the corresponding audio sources in the audio scene.

[0146] In some embodiments a suitable Bitstream based signaling to describe the mapping or correspondence between the audio sources in the immersive audio scene and the audio data for the audio sources is described herein. As an example the following table shows an example syntax of a MPEG-I immersive audio scene configuration function which could be labelled mpegiSceneConfig () (or any suitable function label)

[0147] In the following table is shown an example data structure for a scene configuration packet for MPEG-I immersive audio is the PACTYP_MPEGI_CFG. The data structure mpegiSceneConfig is one based on the example shown in table 3 of ISO / IEC 23090-4 DIS text with added source entity type information as shown below by underlining.

[0148] This packet is delivered in the beginning of the bitstream. This packet prepares the decoder to be configured according to the authored scene specific bitstream. Consequently, this provides a suitable location in the bitstream to include information regarding the audio data that is related to the audio sources in the scene. Syntax No. of bits Mnemonic

[0149] In the above structure the following applies integerld This value represents the newly derived integer from the string identifier. All integerld values shall be unique. In case of audio sources in the scene, the integerld represents the audio source identifier (objectsource, channelsource, hoaSource) for the respect

[0150] To enable interworking with to enable mapping of audio sources (Objectsources, ChannelSources and HOASources) with the audio elements MPEG-H 3DA bitstream the following constraints are proposed:

[0151] If audioSourceEntityType is equal to 0, this entity does not correspond to an audio source in the audio scene being rendered.

[0152] If audioSourceEntityType is equal to 1 , mpeghStreamld value shall be equal to the MHASPacketLabel denoting the MPEG-H stream identifier. This assumes each stream carries a single audio source in the MPEG-I scene (object, channel, HOA).

[0153] If audioSourceEntityType is equal to 2, mpeghStreamld value shall be equal to the MHASPacketLabel denoting the MPEG-H 3D audio stream identifier. This assumes each MPEG-H 3DA stream may carry more than one audio source in the MPEG-I audio scene (object, channel, HOA).

[0154] If audioSourceEntityType is equal to 3, the audio for this will arrive as communication audio directly to the renderer and not as MPEG-H 3DA audio bitstream. The mpeghStreamld and mae_Elementld shall be 0. The mapping between the audio sources in the audio scene and the related communication audio is taken care of by the player or presentation engine. mpeghStreamld indicates the value of the MPEG-H 3DA stream carrying the audio data corresponding to the index i of type audioSourceEntityType. mae_Elementld indicates the audio element identifier (elemldx in ISO / IEC 23008-3) in the MPEG-H 3DA stream. This attribute is relevant when audioSourceEntityType is equal to 2, since multiple audio elements are carried in the single MPEG-H 3D audio stream. This can be ignored if there is only one audio source per MPEG-H 3DA stream (e.g., if audioSourceEntityType is equal to 1). In case of coded audio streams with multiple elements per stream, the element index within the stream is necessary to disambiguate.

[0155] The values 5-15 for audioSourceEntityType are reserved.

[0156] In another example embodiment, a more generic structure can be implemented as described in the following example:

[0157] In this example the syntax of the earlier structure applies but here

[0158] Streamlds indicates the value of the MPEG-H 3DA stream carrying the audio data corresponding to the index i of type audioSourceEntityType. Elementldx indicates the audio element identifier (elemldx in ISO / IEC 23008-3) in the MPEG-H 3DA stream. In a similar manner to the above structure this can be ignored where there is only one audio source per MPEG-H 3DA stream. In case of coded audio streams with multiple elements per stream, the element index within the stream is necessary to disambiguate.

[0159] In some embodiments the audioStreamsQ structure can be in the following format

[0160] In this example the following can apply: audioStreamsCount This value is the number of audio streams in this payload audioStreamld This value represents the MPEG-H 3D audio stream identifier for this audio stream. inputChannelsCount This value represents the number of Input Channels for this audio stream.

[0161] In case of MPEG-H 3D audio stream, the different channels in case of multi-channel Objectsource or ChannelSource or HOASource is carried as part of the same MPEG-H 3D audio stream. In an implementation embodiment these are carried as part of the mainstream or a substream of the MPEG-H 3D audio delivery with multiple streams, but not separated across different substreams or partially mainstream and partially substreams. i np utChan nel I ndex This value is the channel index or elemldx of the MPEG-H 3D audio stream.

[0162] In another implementation embodiment, the audio data or coded audio data identifier In this example the following can apply: audioStreamsCount This value is the number of audio streams in this payload audioStreamld This value represents the unique identifier for this audio stream in the immersive audio scene. inputChannelsCount This value represents the number of Input Channels for this audio stream. In case of MPEG-H 3D audio stream, the different channels in case of multi-channel Objectsource or ChannelSource or HOASource is carried as part of the same MPEG-H 3D audio stream. In an implementation embodiment these are carried as part of the mainstream or a substream of the MPEG-H 3D audio delivery with multiple streams, but not separated across different substreams or partially mainstream and partially substreams. i np utChan nel I ndex This value is the channel index or elemldx of the MPEG-H 3D audio stream. audioStreamBitstreamldentifier This value is the audio data identifier. This can be bitstream identifier for MPEG-H 3D audio MHAS stream. This can be the track identifier related to the audio data carried in a File Format compliant file or file fragment.

[0163] In some embodiments, the correspondence information can also carry granular sound data base and bitstreams.

[0164] In some embodiments the following structure can be applied If audioSourcePresentFlag is equal to 0, this entity does not correspond to an audio source in audio scene being rendered.

[0165] If audioSourcePresentFlag is equal to 1 , this entity does not correspond to an audio source in the audio scene being rendered and the audio source is identified by universal unique identifier.

[0166] Furthermore in some embodiments the following structure can be applied

[0167] If audioSourcePresentFlag is equal to 0, this entity does not correspond to an audio source in audio scene being rendered. If audioSourcePresentFlag is equal to 1 , this entity does not correspond to an audio source in the audio scene being rendered and the audio source is identified by uniform resource identifier, whre uri could be defined for a give audio codec / bitstream that carry MIA config package.

[0168] For example, when mpegiSceneConfig is carried by MPEG-H bitstream

[0169] • scheme: mpegh

[0170] • authority: bitstream • query: streamldx

[0171] • fragmeng: elementldx

[0172] URI = mpegh: / / bitstream?streameldx#elementldx

[0173] In case the mpegiSceneConfig is provided by other means and audio source are expected by real time media then the following scheme could be utilized (defined in ISO / IEC 23090-14:Annex C) For example, a first entity could assign URI equal to rtmedia: / / @?label=1 , and a second entity could assign URI equal to rtmedia: / / @?label=2, during session initiation the following SDP would allow to map the MPEG-I audio source to real time audio source v=0 o=bob 280744730 28977631 IN IP4 host.example.com s= i=A Seminar on the session description protocol c=IN IP4 192.0.2.2 t=0 0 m=audio 6886 RTP / AVP 0 a=label:1 m=audio 22334 RTP / AVP 0 a=label:2

[0174] With respect to Figure 8 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 2000 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder or the renderer or any functional block as described above.

[0175] In some embodiments the device 2000 comprises at least one processor or central processing unit 2007. The processor 2007 can be configured to execute various program codes such as the methods described herein.

[0176] In some embodiments the device 2000 comprises a memory 2011 . In some embodiments the at least one processor 2007 is coupled to the memory 2011 . The memory 2011 can be any suitable storage means. In some embodiments the memory 2011 comprises a program code section for storing program codes implementable upon the processor 2007. Furthermore, in some embodiments the memory 2011 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 2007 whenever needed via the memory-processor coupling.

[0177] In some embodiments the device 2000 comprises a user interface 2005. The user interface 2005 can be coupled in some embodiments to the processor 2007. In some embodiments the processor 2007 can control the operation of the user interface 2005 and receive inputs from the user interface 2005. In some embodiments the user interface 2005 can enable a user to input commands to the device 2000, for example via a keypad. In some embodiments the user interface 2005 can enable the user to obtain information from the device 2000. For example, the user interface 2005 may comprise a display configured to display information from the device 2000 to the user. The user interface 2005 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 2000 and further displaying information to the user of the device 2000. In some embodiments the user interface 2005 may be the user interface for communicating.

[0178] In some embodiments the device 2000 comprises an input / output port 2009. The input / output port 2009 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 2007 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.

[0179] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example, in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).

[0180] The input / output port 2009 may be configured to receive the signals.

[0181] In some embodiments the device 2000 may be employed as at least part of the renderer. The input / output port 2009 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar.

[0182] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0183] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.

[0184] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory- and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.

[0185] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0186] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

[0187] As used in this application, the term “circuitry” may refer to one or more or all of the following:

[0188] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and

[0189] (b) combinations of hardware circuits and software, such as (as applicable):

[0190] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and

[0191] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and

[0192] I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0193] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).

[0194] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.

Claims

CLAIMS:

1. A method for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the method comprising: obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source; obtaining a coded audio bitstream for the immersive audio scene; determining at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; and including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

2. The method as claimed in claim 1, wherein obtaining a coded audio bitstream for the immersive audio scene comprises: obtaining at least one scene relevant audio signal; and generating the coded audio bitstream based on the at least one scene relevant audio signal.

3. The method as claimed in claim 2, wherein determining audio bitstream identifiers with correspondence to audio source identifiers comprises generating correspondence information between the at least one scene relevant audio signal and the at least one audio source within the obtained audio scene information; and encoding the correspondence information within a bitstream further comprising the audio scene information and the at least one scene relevant audio signal.

4. The method as claimed in any of claims 1 to 3, further comprising at least one of: storing the bitstream for retrieval by a renderer; transmitting the bitstream to a renderer.

5. The method as claimed in any of claims 1 to 4, wherein obtaining immersive audio scene information, for defining an audio scene comprising at least one audio source comprises obtaining or generating the immersive audio scene information within an encoder input format.

346. The method as claimed in claim 2 or any claim dependent on claim 2, wherein obtaining at least one scene relevant audio signal comprises: determining a coding method for the at least one scene relevant audio signal; and encoding the at least one scene relevant audio signal based on the determined coding method.

7. The method as claimed in claim 6, wherein determining at least one audio bitstream identifier with correspondence to at least one audio source identifier comprises determining at least one audio bitstream identifier based on the coding method for the at least one scene relevant audio signal.

8. The method as claimed in any of claims 1 to 7, wherein obtaining a coded audio bitstream for the immersive audio scene comprises: determining a delivery method for the at least one scene relevant audio signal; and encoding an audio bitstream based on the determined delivery method.

9. The method as claimed in claim 8, wherein determining the at least one audio bitstream identifier with correspondence to at least one audio source identifier comprises determining the at least one audio bitstream identifier based on the delivery method for the at least one scene relevant audio signal.

10. The method as claimed in any of claims 7 or 9, wherein the at least one audio bitstream identifier based on the coding method or delivery method for the at least one scene relevant audio signal is of a MPEG- H 3D stream for the delivery method for the at least one scene relevant audio signal.11 . The method as claimed in any of claims 1 to 10, wherein the at least one audio source comprises at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

12. The method as claimed in any of claims 1 to 11 , wherein including the at least one audio bitstream identifier with correspondence to the at least one audio source identifier comprises including the correspondence information.

13. A method for assisting immersive audio rendering, the method comprising:35obtaining a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtaining the immersive audio scene information from the bitstream; obtaining the coded audio bitstream from the bitstream; obtaining the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; mapping the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configuring a renderer based on the mapping; and rendering an output audio signal based on the configured renderer.

14. The method as claimed in claim 13, further comprising at least one of: retrieving the bitstream; and receiving the bitstream.

15. The method as claimed in any of claims 13 or 14, wherein the audio scene information is an encoder input format.

16. The method as claimed in any of claims 13 to 15, wherein the at least one audio bitstream identifier is of a MPEG-H 3D stream.

17. The method as claimed in any of claims 13 to 16, wherein the at least one audio source comprises at least one of: at least one audio object source; at least one channel source; and at least one higher order Ambisonics source.

18. The method as claimed in any of claims 13 to 17, wherein configuring a renderer based on the mapping comprises configuring at least one audio data path based on the mapping.

19. An apparatus for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain immersive audio scene information, for defining an audio scene comprising at least one audio source; obtain a coded audio bitstream for the immersive audio scene; determine at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; and include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

20. An apparatus for assisting immersive audio rendering, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to: obtain a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtain the immersive audio scene information from the bitstream; obtain the coded audio bitstream from the bitstream; obtain the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; map the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configure a renderer based on the mapping; and render an output audio signal based on the configured renderer.

21. A computer program comprising instructions, which, when executed by an apparatus, cause the apparatus to perform the method of any of claims 1 to 18.

22. An apparatus for generating correspondence information between an immersive audio bitstream and a coded audio bitstream for assisting immersive audio rendering, the apparatus comprising means configured to: obtain immersive audio scene information, for defining an audio scene comprising at least one audio source; obtain a coded audio bitstream for the immersive audio scene; determine at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; and include the at least one audio bitstream identifier with correspondence to the at least one audio source identifier within a bitstream, the bitstream further comprising the immersive audio scene information and the coded audio bitstream.

23. An apparatus for assisting immersive audio rendering, the apparatus comprising means configured to:: obtain a bitstream, the bitstream comprising: immersive audio scene information, the immersive audio scene information defining an audio scene comprising at least one audio source; a coded audio bitstream for the immersive audio scene; and at least one audio bitstream identifier with correspondence to at least one audio source identifier, the at least one audio bitstream identifier determined based on the at least one audio source in the audio scene; obtain the immersive audio scene information from the bitstream; obtain the coded audio bitstream from the bitstream; obtain the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; map the obtained coded audio bitstream and the at least one audio source within the obtained immersive audio scene information based on the at least one audio bitstream identifier with correspondence to the at least one audio source identifier; configure a renderer based on the mapping; and render an output audio signal based on the configured renderer.38