Immersive audio format selection
By transmitting initialization information within transport protocol payloads, the immersive audio decoder is pre-initialized, addressing compatibility and performance issues, ensuring optimal rendering across diverse scenarios.
Patent Information
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-06
AI Technical Summary
Existing immersive audio codecs lack efficient mechanisms for initializing immersive audio decoders prior to decoding, leading to suboptimal performance and compatibility issues across various operating scenarios.
The apparatus and method involve transmitting initialization information related to immersive audio frames within transport protocol payloads to initialize the decoder before decoding, including details on coded formats, sub-formats, and channel configurations, enabling pre-initialization of decoders.
This approach enhances decoder initialization efficiency, improves compatibility across diverse scenarios, and ensures optimal audio rendering by aligning decoder settings with encoder configurations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field The present application relates to apparatus and methods for immersive audio output format selection, but not exclusively for implementing immersive audio output format selection when employing immersive voice and audio services (IVAS) real-time transport protocol (RTP) payloads. Background Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the immersive voice and audio services (IVAS) codec which has been designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing. This audio codec handles the encoding, decoding and rendering of speech, music and generic audio. It furthermore supports channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. The codec operates with low latency to enable conversational services as well as support high error robustness under various transmission conditions. For example, the input signals can be presented to an IVAS encoder in one of the supported formats (and in some allowed combinations of the formats). Similarly, the decoder can be configured to output the audio in any of the supported formats. The main supported input formats for IVAS are stereo, multichannel (MC), object-based audio (ISM), scene-based audio (SBA), and Metadata-assisted spatial audio (MASA). In addition, the following combinations are supported: Objects with MASA (OMASA) and Objects with SBA (OSBA). IVAS furthermore includes the EVS codec for mono input operation. The IVAS output formats include mono, stereo, multi-channel (including custom loudspeaker layouts), FOA, HOA2, HOA3, and binaural. This flexibility is because as a spatial audio codec supporting at least three degrees of rotation freedom (yaw, pitch, roll) for all spatial inputs, the IVAS codec is expected to be used in a variety of scenarios, all of which cannot be known beforehand. In addition, a so-called pass-through operation is possible allowing, e.g., MASA output for MASA input and ISM output for ISM input, in other words where the audio could be provided in its original format after transmission (encoding / decoding). Additionally RTP (Real-Time Transport Protocol) is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data transfer to multiple destinations through IP multicast or to a specific destination through IP unicast. The majority of the RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols may also be utilized. RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol (RTSP). The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transfer of multimedia data, and its companion protocol (RTCP) is used to periodically send control information and QoS (Quality of Service) parameters. RTP sessions are typically initiated between client and server or between client and another client (or a multi-party topology) using a signalling protocol, such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols typically use the Session Description Protocol (SDP), such as defined by RFC 8866 to specify parameters for the sessions. Summary There is provided according to a first aspect an apparatus for configuring an immersive audio decoder, the apparatus comprising means configured to: obtain initialization information based on an established conversational immersive audio session; and include the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. The means configured to include the initialization information relating to the immersive audio frame in at least one transport protocol payload may be configured to transmit the initialization information such that the initialization information is received by at least one further apparatus and employed to initialize the immersive audio decoder prior to decoding the immersive audio frame. The initialization information relating to the immersive audio frame may comprise at least one of: information identifying a coded format of the immersive audio frame; information identifying a coded sub-format of the immersive audio frame; information identifying an encoder input sampling rate of the immersive audio frame; and information identifying a number of channels with respect to an at least one coded format of the immersive audio frame. The means configured to include the initialization information relating to an immersive audio frame in the at least one transport protocol payload may be configured to generate at least one processing information frame comprising the initialization information. The means configured to include the initialization information relating to an immersive audio frame in the at least one transport protocol payload may be configured to include the initialization information in at least one of: a real-time transport protocol payload; a real-time transport protocol header extension; and a real-time transport control protocol payload. The means may be further configured to transmit the initialization information in the at least one transport protocol payload to the further apparatus. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a mono channel format; a stereo channel format; a multichannel format; a Metadata-Assisted Spatial Audio format; an Independent Streams with Metadata format; a Scene-Based Audio format; an Object-based audio with Metadata-Assisted Spatial Audio format; and an Objectbased audio with Scene-Based Audio format. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a multichannel 5.1 sub-format; a multichannel 5.1.2 sub-format; a multichannel 5.1.4 sub-format; a multichannel 7.1 sub-format; a multichannel 7.1.4 sub-format; a Metadata-Assisted Spatial Audio one transport channel sub-format; a Metadata-Assisted Spatial Audio two transport channels sub-format; an Independent Streams with Metadata single stream sub-format; an Independent Streams with Metadata two streams sub-format; an Independent Streams with Metadata three streams sub-format; an Independent Streams with Metadata four streams sub-format; a Scene-Based Audio First order Ambisonics sub-format; a Scene-Based Audio Second, higher, order Ambisonics sub-format; a Scene-Based Audio Third, higher, order Ambisonics sub-format; a Scene-Based Audio planar First order Ambisonics sub-format; a Scene-Based Audio planar Second, higher, order Ambisonics sub-format; a Scene-Based Audio planar Third, higher, order Ambisonics sub-format; an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format; an Objectbased audio with Metadata-Assisted Spatial Audio two independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format; an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-format. The means may be further configured to negotiate with the further apparatus support for the initialization information. The means configured to negotiate with the further apparatus may be configured to negotiate employing a session description file. The means may be further configured to: encode at least one immersive audio signal to an immersive audio frame based on the at least one coded format and / or sub-format; include the immersive audio frame in at the at least one transport protocol payload or in a further at least one transport protocol payload. The means configured to obtain initialization information relating to at least one immersive audio frame may be configured to obtain initialization information relating to at least one immersive audio frame for at least one of: during a session setup; and after a session setup. The initialization information may be associated to the immersive audio encoder configuration. The immersive audio encoder may be based on the session. The means may further encode the immersive audio frame. The means may further be configured to include the encoded immersive audio frame. The means configured to perform including the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by an immersive audio decoder may be configured to perform one of: including in the at least one transport protocol payload the initialization information and the encoded immersive audio frame; including in the at least one transport protocol payload only the initialization information. According to a second aspect there is provided an apparatus for initializing an immersive audio decoder, the apparatus comprising means configured to: obtain initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initialize the immersive audio decoder based on at least part of the obtained initialization information. The means may be further configured to: obtain the at least one immersive audio frame; and decode the at least one immersive audio frame at the initialized immersive audio decoder. The means configured to obtain initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may be further configured to obtain, the initialization information from an earlier part of an at least one transport protocol payload with respect to a later part of the at least one transport protocol comprising the immersive audio frame. The means configured to obtain initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may be further configured to obtain the initialization information in a separate earlier at least one transport protocol payload with respect to a later at least one transport protocol payload comprising the at least one immersive audio frame. The initialization information relating to the immersive audio frame may comprise at least one of: information identifying a coded format; information identifying a coded sub-format; information identifying an encoder input sampling rate; and information identifying a number of channels with respect to the at least one coded format. The means configured to obtain initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may be configured to obtain the initialization information from at least one of: a real-time transport protocol payload; a real-time transport protocol header extension; and a real-time transport control protocol payload. The means configured to obtain initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may be further configured to: decode at a further decoder at least part of the at least one transport protocol payload to obtain the initialization information; and initialize the immersive audio frame decoder based at least on the initialization information. The means is further configured to perform pre-initializing the immersive audio decoder or at the further decoder prior to decoding at the immersive audio decoder or at a further decoder at least part of the at least one immersive audio frame to obtain the initialization information. The further decoder may be a lightweight decoder configured to obtain the initialization information. The further decoder may be a lightweight decoder configured to only obtain the initialization information. The means may be further configured to: pre-initialize at least two immersive audio decoders, each of the at least two immersive audio decoders pre-initialized to decode the at least one immersive audio frame based on a different one of at least one coded format and / or different output sampling rate, wherein the means configured to obtain initialization information from an immersive audio frame or prior to the decoding immersive audio frame of an established conversational immersive audio session may be configured to: decode at each of the pre-initialized at least two immersive audio decoders the at least one immersive audio frame to generate at least two decoder outputs; determine the initialization information relating to the immersive audio frame based on the at least two decoder outputs, and wherein the means configured to initialize the immersive audio decoder based on at least part of the obtained initialization information may be configured to: select one of the at least two decoders based on the initialization information; and finalize an initialization of the selected one of the at least two decoders. The means configured to determine the initialization information relating to the immersive audio frame based on the at least two decoder outputs may be configured to: analyse the output of each of the at least two decoders; and determine which of the at least two decoders to select based on the analysis. The means configured to analyse the output of each of the at least two decoders may be configured to detect at least one of: an error output; and a lower quality output measured against one or more of the other outputs or a threshold value. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a mono channel format; a stereo channel format; a multichannel format; a Metadata-Assisted Spatial Audio format; an Independent Streams with Metadata format; a Scene-Based Audio format; an Object-based audio with Metadata-Assisted Spatial Audio format; and an Objectbased audio with Scene-Based Audio format. Initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a multichannel 5.1 sub-format; a multichannel 5.1.2 sub-format; a multichannel 5.1.4 sub-format; a multichannel 7.1 sub-format; a multichannel 7.1.4 sub-format; a Metadata-Assisted Spatial Audio one transport channel sub-format; a Metadata-Assisted Spatial Audio two transport channels sub-format; an Independent Streams with Metadata single stream sub-format; an Independent Streams with Metadata two streams sub-format; an Independent Streams with Metadata three streams sub-format; an Independent Streams with Metadata four streams sub-format; a Scene-Based Audio First order Ambisonics sub-format; a Scene-Based Audio Second, higher, order Ambisonics sub-format; a Scene-Based Audio Third, higher, order Ambisonics sub-format; a Scene-Based Audio planar First order Ambisonics sub-format; a Scene-Based Audio planar Second, higher, order Ambisonics sub-format; a Scene-Based Audio planar Third, higher, order Ambisonics sub-format; an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format; an Object-based audio with Metadata-Assisted Spatial Audio two independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format; an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-. The means may be further configured to negotiate with a further apparatus support for the initialization information. The means configured to negotiate with the further apparatus may be configured to negotiate employing a session description file. The means may be further configured to obtain further initialization information with respect to an output decoded format. The means configured to initialize the immersive audio frame decoder may be further configured to initialize the immersive audio frame decoder based on the output decoded format. The information identifying the output decoded format may comprise at least one of: information identifying the output format and / or sub-format; information identifying a decoder output sampling rate; and information identifying a number of channels to output with respect to the immersive audio frame decoder. The means configured to include the initialization information relating to the immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted separately from the immersive audio bitstream may be configured to packetize the initialization information. The means configured to obtain, from the at least one transport protocol payload, initialization information relating to an immersive audio frame may be configured to obtain initialization information relating to at least one immersive audio frame for at least one of: during a session setup; and after a session setup. The immersive audio frame may be encoded by an encoder configured for a session. According to a third aspect there is provided a method for an apparatus for configuring an immersive audio decoder, the method comprising: obtaining initialization information based on an established conversational immersive audio session; and including the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. Including the initialization information relating to the immersive audio frame in at least one transport protocol payload may comprise transmitting the initialization information such that the initialization information is received by at least one further apparatus and employed to initialize the immersive audio decoder prior to decoding the immersive audio frame. The initialization information relating to the immersive audio frame may comprise at least one of: information identifying a coded format of the immersive audio frame; information identifying a coded sub-format of the immersive audio frame; information identifying an encoder input sampling rate of the immersive audio frame; and information identifying a number of channels with respect to an at least one coded format of the immersive audio frame. Including the initialization information relating to an immersive audio frame in the at least one transport protocol payload may comprise generating at least one processing information frame comprising the initialization information. Including the initialization information relating to an immersive audio frame in the at least one transport protocol payload may comprise including the initialization information in at least one of: a real-time transport protocol payload; a real-time transport protocol header extension; and a real-time transport control protocol payload. The method may further comprise transmitting the initialization information in the at least one transport protocol payload to the further apparatus. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a mono channel format; a stereo channel format; a multichannel format; a Metadata-Assisted Spatial Audio format; an Independent Streams with Metadata format; a Scene-Based Audio format; an Object-based audio with Metadata-Assisted Spatial Audio format; and an Objectbased audio with Scene-Based Audio format. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a multichannel 5.1 sub-format; a multichannel 5.1.2 sub-format; a multichannel 5.1.4 sub-format; a multichannel 7.1 sub-format; a multichannel 7.1.4 sub-format; a Metadata-Assisted Spatial Audio one transport channel sub-format; a Metadata-Assisted Spatial Audio two transport channels sub-format; an Independent Streams with Metadata single stream sub-format; an Independent Streams with Metadata two streams sub-format; an Independent Streams with Metadata three streams sub-format; an Independent Streams with Metadata four streams sub-format; a Scene-Based Audio First order Ambisonics sub-format; a Scene-Based Audio Second, higher, order Ambisonics sub-format; a Scene-Based Audio Third, higher, order Ambisonics sub-format; a Scene-Based Audio planar First order Ambisonics sub-format; a Scene-Based Audio planar Second, higher, order Ambisonics sub-format; a Scene-Based Audio planar Third, higher, order Ambisonics sub-format; an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format; an Objectbased audio with Metadata-Assisted Spatial Audio two independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format; an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-format. The method may further comprise negotiating with the further apparatus support for the initialization information. Negotiating with the further apparatus may comprise negotiating employing a session description file. The method may further comprise: encoding at least one immersive audio signal to an immersive audio frame based on the at least one coded format and / or sub-format; including the immersive audio frame in at the at least one transport protocol payload or in a further at least one transport protocol payload. Obtaining initialization information relating to at least one immersive audio frame may comprise obtaining initialization information relating to at least one immersive audio frame for at least one of: during a session setup; and after a session setup. The initialization information may be associated to the immersive audio encoder configuration. The immersive audio encoder may be based on the session. The method may further comprise encoding the immersive audio frame. The method may further comprise including the encoded immersive audio frame. Including the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by an immersive audio decoder may comprise one of: including in the at least one transport protocol payload the initialization information and the encoded immersive audio frame; including in the at least one transport protocol payload only the initialization information. According to a fourth aspect there is provided a method for an apparatus for initializing an immersive audio decoder, the method comprising: obtaining initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initializing the immersive audio decoder based on at least part of the obtained initialization information. The method may further comprise: obtaining the at least one immersive audio frame; and decoding the at least one immersive audio frame at the initialized immersive audio decoder. Obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may further comprise obtaining, the initialization information from an earlier part of an at least one transport protocol payload with respect to a later part of the at least one transport protocol comprising the immersive audio frame. Obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may further comprise obtaining the initialization information in a separate earlier at least one transport protocol payload with respect to a later at least one transport protocol payload comprising the at least one immersive audio frame. The initialization information relating to the immersive audio frame may comprise at least one of: information identifying a coded format; information identifying a coded sub-format; information identifying an encoder input sampling rate; and information identifying a number of channels with respect to the at least one coded format. Obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may comprise obtaining the initialization information from at least one of: a real-time transport protocol payload; a real-time transport protocol header extension; and a real-time transport control protocol payload. Obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may further comprise: decoding at a further decoder at least part of the at least one transport protocol payload to obtain the initialization information; and initializing the immersive audio frame decoder based at least on the initialization information. The method further comprises pre-initializing the immersive audio decoder or at the further decoder prior to decoding at the immersive audio decoder or at a further decoder at least part of the at least one immersive audio frame to obtain the initialization information. The further decoder may be a lightweight decoder configured to obtain the initialization information. The further decoder may be a lightweight decoder configured to only obtain the initialization information. The method may further comprise: pre-initializing at least two immersive audio decoders, each of the at least two immersive audio decoders pre-initialized to decode the at least one immersive audio frame based on a different one of at least one coded format and / or different output sampling rate, wherein obtaining initialization information from an immersive audio frame or prior to the decoding immersive audio frame of an established conversational immersive audio session may comprise: decoding at each of the pre-initialized at least two immersive audio decoders the at least one immersive audio frame to generate at least two decoder outputs; determining the initialization information relating to the immersive audio frame based on the at least two decoder outputs, and wherein initializing the immersive audio decoder based on at least part of the obtained initialization information may comprise: selecting one of the at least two decoders based on the initialization information; and finalizing an initialization of the selected one of the at least two decoders. Determining the initialization information relating to the immersive audio frame based on the at least two decoder outputs may comprise: analysing the output of each of the at least two decoders; and determining which of the at least two decoders to select based on the analysis. Analysing the output of each of the at least two decoders may comprise detecting at least one of: an error output; and a lower quality output measured against one or more of the other outputs or a threshold value. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a mono channel format; a stereo channel format; a multichannel format; a Metadata-Assisted Spatial Audio format; an Independent Streams with Metadata format; a Scene-Based Audio format; an Object-based audio with Metadata-Assisted Spatial Audio format; and an Objectbased audio with Scene-Based Audio format. Initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a multichannel 5.1 sub-format; a multichannel 5.1.2 sub-format; a multichannel 5.1.4 sub-format; a multichannel 7.1 sub-format; a multichannel 7.1.4 sub-format; a Metadata-Assisted Spatial Audio one transport channel sub-format; a Metadata-Assisted Spatial Audio two transport channels sub-format; an Independent Streams with Metadata single stream sub-format; an Independent Streams with Metadata two streams sub-format; an Independent Streams with Metadata three streams sub-format; an Independent Streams with Metadata four streams sub-format; a Scene-Based Audio First order Ambisonics sub-format; a Scene-Based Audio Second, higher, order Ambisonics sub-format; a Scene-Based Audio Third, higher, order Ambisonics sub-format; a Scene-Based Audio planar First order Ambisonics sub-format; a Scene-Based Audio planar Second, higher, order Ambisonics sub-format; a Scene-Based Audio planar Third, higher, order Ambisonics sub-format; an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format; an Object-based audio with Metadata-Assisted Spatial Audio two independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format; an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-. The method may further comprise negotiating with a further apparatus support for the initialization information. Negotiating with the further apparatus may comprise negotiating employing a session description file. The method may further comprise obtaining further initialization information with respect to an output decoded format. Initializing the immersive audio frame decoder may further comprise initializing the immersive audio frame decoder based on the output decoded format. The information identifying the output decoded format may comprise at least one of: information identifying the output format and / or sub-format; information identifying a decoder output sampling rate; and information identifying a number of channels to output with respect to the immersive audio frame decoder. Including the initialization information relating to the immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted separately from the immersive audio bitstream may comprise packetizing the initialization information. Obtaining, from the at least one transport protocol payload, initialization information relating to an immersive audio frame may comprise obtaining initialization information relating to at least one immersive audio frame for at least one of: during a session setup; and after a session setup. The immersive audio frame may be encoded by an encoder configured for a session. According to a fifth aspect there is provided an apparatus for configuring an immersive audio decoder, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtaining initialization information based on an established conversational immersive audio session; and including the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. The apparatus caused to perform including the initialization information relating to the immersive audio frame in at least one transport protocol payload may be caused to perform transmitting the initialization information such that the initialization information is received by at least one further apparatus and employed to initialize the immersive audio decoder prior to decoding the immersive audio frame. The initialization information relating to the immersive audio frame may comprise at least one of: information identifying a coded format of the immersive audio frame; information identifying a coded sub-format of the immersive audio frame; information identifying an encoder input sampling rate of the immersive audio frame; and information identifying a number of channels with respect to an at least one coded format of the immersive audio frame. The apparatus caused to perform including the initialization information relating to an immersive audio frame in the at least one transport protocol payload may be caused to perform generating at least one processing information frame comprising the initialization information. The apparatus caused to perform including the initialization information relating to an immersive audio frame in the at least one transport protocol payload may be caused to perform including the initialization information in at least one of: a real-time transport protocol payload; a real-time transport protocol header extension; and a real-time transport control protocol payload. The apparatus may be further caused to perform transmitting the initialization information in the at least one transport protocol payload to the further apparatus. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a mono channel format; a stereo channel format; a multichannel format; a Metadata-Assisted Spatial Audio format; an Independent Streams with Metadata format; a Scene-Based Audio format; an Object-based audio with Metadata-Assisted Spatial Audio format; and an Objectbased audio with Scene-Based Audio format. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a multichannel 5.1 sub-format; a multichannel 5.1.2 sub-format; a multichannel 5.1.4 sub-format; a multichannel 7.1 sub-format; a multichannel 7.1.4 sub-format; a Metadata-Assisted Spatial Audio one transport channel sub-format; a Metadata-Assisted Spatial Audio two transport channels sub-format; an Independent Streams with Metadata single stream sub-format; an Independent Streams with Metadata two streams sub-format; an Independent Streams with Metadata three streams sub-format; an Independent Streams with Metadata four streams sub-format; a Scene-Based Audio First order Ambisonics sub-format; a Scene-Based Audio Second, higher, order Ambisonics sub-format; a Scene-Based Audio Third, higher, order Ambisonics sub-format; a Scene-Based Audio planar First order Ambisonics sub-format; a Scene-Based Audio planar Second, higher, order Ambisonics sub-format; a Scene-Based Audio planar Third, higher, order Ambisonics sub-format; an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format; an Objectbased audio with Metadata-Assisted Spatial Audio two independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format; an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-format. The apparatus may be further caused to perform negotiating with the further apparatus support for the initialization information. The apparatus caused to perform negotiating with the further apparatus may be caused to perform negotiating employing a session description file. The apparatus may be further caused to perform: encoding at least one immersive audio signal to an immersive audio frame based on the at least one coded format and / or sub-format; including the immersive audio frame in at the at least one transport protocol payload or in a further at least one transport protocol payload. The apparatus caused to perform obtaining initialization information relating to at least one immersive audio frame may be further caused to perform obtaining initialization information relating to at least one immersive audio frame for at least one of: during a session setup; and after a session setup. The initialization information may be associated to the immersive audio encoder configuration. The immersive audio encoder may be based on the session. The method may further comprise encoding the immersive audio frame. The method may further comprise including the encoded immersive audio frame. The apparatus caused to perform including the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by an immersive audio decoder may be further caused to perform one of: including in the at least one transport protocol payload the initialization information and the encoded immersive audio frame; including in the at least one transport protocol payload only the initialization information. According to a sixth aspect there is provided an apparatus for initializing an immersive audio decoder, the apparatus comprising: at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtaining initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initializing the immersive audio decoder based on at least part of the obtained initialization information. The apparatus may further be caused to perform: obtaining the at least one immersive audio frame; and decoding the at least one immersive audio frame at the initialized immersive audio decoder. The apparatus caused to perform obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may further be caused to perform obtaining, the initialization information from an earlier part of an at least one transport protocol payload with respect to a later part of the at least one transport protocol comprising the immersive audio frame. The apparatus caused to perform obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may further be caused to perform obtaining the initialization information in a separate earlier at least one transport protocol payload with respect to a later at least one transport protocol payload comprising the at least one immersive audio frame. The initialization information relating to the immersive audio frame may comprise at least one of: information identifying a coded format; information identifying a coded sub-format; information identifying an encoder input sampling rate; and information identifying a number of channels with respect to the at least one coded format. The apparatus caused to perform obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may be caused to perform obtaining the initialization information from at least one of: a realtime transport protocol payload; a real-time transport protocol header extension; and a real-time transport control protocol payload. The apparatus caused to perform obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame may further be caused to perform: decoding at a further decoder at least part of the at least one transport protocol payload to obtain the initialization information; and initializing the immersive audio frame decoder based at least on the initialization information. The apparatus may further be caused to perform pre-initializing the immersive audio decoder or at the further decoder prior to decoding at the immersive audio decoder or at a further decoder at least part of the at least one immersive audio frame to obtain the initialization information. The further decoder may be a lightweight decoder configured to obtain the initialization information. The further decoder may be a lightweight decoder configured to only obtain the initialization information. The apparatus may further be caused to perform: pre-initializing at least two immersive audio decoders, each of the at least two immersive audio decoders pre-initialized to decode the at least one immersive audio frame based on a different one of at least one coded format and / or different output sampling rate, wherein The apparatus caused to perform obtaining initialization information from an immersive audio frame or prior to the decoding immersive audio frame of an established conversational immersive audio session may further be caused to perform: decoding at each of the pre-initialized at least two immersive audio decoders the at least one immersive audio frame to generate at least two decoder outputs; determining the initialization information relating to the immersive audio frame based on the at least two decoder outputs, and wherein the apparatus caused to perform initializing the immersive audio decoder based on at least part of the obtained initialization information may further be caused to perform: selecting one of the at least two decoders based on the initialization information; and finalizing an initialization of the selected one of the at least two decoders. The apparatus caused to perform determining the initialization information relating to the immersive audio frame based on the at least two decoder outputs may further be caused to perform: analysing the output of each of the at least two decoders; and determining which of the at least two decoders to select based on the analysis. The apparatus caused to perform analysing the output of each of the at least two decoders may further be caused to perform detecting at least one of: an error output; and a lower quality output measured against one or more of the other outputs or a threshold value. The initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a mono channel format; a stereo channel format; a multichannel format; a Metadata-Assisted Spatial Audio format; an Independent Streams with Metadata format; a Scene-Based Audio format; an Object-based audio with Metadata-Assisted Spatial Audio format; and an Objectbased audio with Scene-Based Audio format. Initialization information relating to the immersive audio frame may comprise an identifier for at least one of: a multichannel 5.1 sub-format; a multichannel 5.1.2 sub-format; a multichannel 5.1.4 sub-format; a multichannel 7.1 sub-format; a multichannel 7.1.4 sub-format; a Metadata-Assisted Spatial Audio one transport channel sub-format; a Metadata-Assisted Spatial Audio two transport channels sub-format; an Independent Streams with Metadata single stream sub-format; an Independent Streams with Metadata two streams sub-format; an Independent Streams with Metadata three streams sub-format; an Independent Streams with Metadata four streams sub-format; a Scene-Based Audio First order Ambisonics sub-format; a Scene-Based Audio Second, higher, order Ambisonics sub-format; a Scene-Based Audio Third, higher, order Ambisonics sub-format; a Scene-Based Audio planar First order Ambisonics sub-format; a Scene-Based Audio planar Second, higher, order Ambisonics sub-format; a Scene-Based Audio planar Third, higher, order Ambisonics sub-format; an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format; an Object-based audio with Metadata-Assisted Spatial Audio two independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format; an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format; an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format; an Object based audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format; an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format; an Objectbased audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format; an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; and an Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-. The apparatus may further be caused to perform negotiating with a further apparatus support for the initialization information. The apparatus caused to perform negotiating with the further apparatus may further be caused to perform negotiating employing a session description file. The apparatus may further be caused to perform obtaining further initialization information with respect to an output decoded format. The apparatus caused to perform initializing the immersive audio frame decoder may further be caused to perform initializing the immersive audio frame decoder based on the output decoded format. The information identifying the output decoded format may comprise at least one of: information identifying the output format and / or sub-format; information identifying a decoder output sampling rate; and information identifying a number of channels to output with respect to the immersive audio frame decoder. The apparatus caused to perform including the initialization information relating to the immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted separately from the immersive audio bitstream may further be caused to perform packetizing the initialization information. The apparatus caused to perform obtaining, from the at least one transport protocol payload, initialization information relating to an immersive audio frame may further be caused to perform obtaining initialization information relating to at least one immersive audio frame for at least one of: during a session setup; and after a session setup. The immersive audio frame may be encoded by an encoder configured for a session. According to a seventh aspect there is provided an apparatus for configuring an immersive audio decoder, the apparatus comprising: obtaining circuitry configured to obtain initialization information based on an established conversational immersive audio session; and including circuitry configured to include the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. According to an eighth aspect there is provided an apparatus for initializing an immersive audio decoder, the apparatus comprising: obtaining circuitry configured to obtain initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initializing circuitry configured to initialize the immersive audio decoder based on at least part of the obtained initialization information. According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for configuring an immersive audio decoder, to perform at least the following: obtain initialization information based on an established conversational immersive audio session; and include the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. According to a tenth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for initializing an immersive audio decoder, to perform at least the following: obtain initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initialize the immersive audio decoder based on at least part of the obtained initialization information. According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for configuring an immersive audio decoder, to perform at least the following: obtain initialization information based on an established conversational immersive audio session; and include the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for initializing an immersive audio decoder, to perform at least the following: obtain initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initialize the immersive audio decoder based on at least part of the obtained initialization information. According to a thirteenth aspect there is provided an apparatus for configuring an immersive audio decoder, the apparatus comprising: means for obtaining initialization information based on an established conversational immersive audio session; and means for including the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. According to a fourteenth aspect there is provided an apparatus for initializing an immersive audio decoder, the apparatus comprising: means for obtaining initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and means for initializing the immersive audio decoder based on at least part of the obtained initialization information. According to a fifteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for configuring an immersive audio decoder, to perform at least the following: obtain initialization information based on an established conversational immersive audio session; and include the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder. According to a sixteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for initializing an immersive audio decoder, to perform at least the following: obtain initialization information from an immersive audio frame or prior to decoding the immersive audio frame; and initialize the immersive audio decoder based on at least part of the obtained initialization information. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Fig.1 shows schematically example server and peer-to-peer teleconferencing systems within which embodiments may be implemented; Fig.2 shows schematically an example packet format with IVAS payload structure; Figs.3a to 3c show schematically decoder and renderer configurations according to some embodiments; Fig.4 shows schematically an example decoder initializer according to some embodiments; Fig.5 shows schematically a further example decoder initializer according to some embodiments; Fig.6 shows schematically another example decoder initializer configuration according to some embodiments; Fig.7 shows schematically a further example decoder initializer configuration according to some embodiments; Fig.8 shows schematically a flow diagram of example encoder and decoder initialized by a negotiation phase according to some embodiments; Figs.9a and 9b show schematically flow diagrams of sender and receiver based operations according to some embodiments; Figs. 10a and 10b show schematically further flow diagrams of sender and receiver based operations according to some embodiments; Figs. 11 a and 11 b show schematically another set of flow diagrams of sender and receiver based operations according to some embodiments; Fig. 12 shows an example device suitable for implementing the apparatus described herein; and Fig. 13 shows an example decoder initialization scenario. Embodiments of the Application The following describes in further detail suitable apparatus and possible mechanisms for the provision of efficient IVAS audio. An example system within which embodiments may be implemented is shown in Fig. 1. Fig. 1, for example, shows an example teleconferencing system within which some embodiments can be implemented. In this example there is shown two sites or rooms, Room A 100 and Room B 102. Room A 100 comprises a ‘talker’ or user, Talker 1 103. Room B 102 comprises one ‘talker’ or user, Talker RX 141. In the following example within room A is a suitable teleconference apparatus (or more generally telecommunications apparatus 110) configured to spatially capture and encode the audio environment and furthermore is configured to render a spatial audio signal to the room. The apparatus can in some embodiments be implemented by a user equipment (UE) operating within a cellular communications system or accessing any suitable access network. Within each of the other rooms may be a suitable teleconference apparatus (or more generally telecommunications apparatus such as apparatus 120 within room B) configured to render a spatial audio signal to the room and furthermore is configured to capture and encode at least a mono audio and optionally configured to spatially capture and encode the audio environment. In the following examples each room is provided with the means to spatially capture, encode spatial audio signals, receive spatial audio signals and render these to a suitable listener. It would be understood that there may be other embodiments where the system comprises some apparatus configured to only capture and encode audio signals (in other words the apparatus is a ‘transmit’ only apparatus), and other apparatus configured to only receive and render audio signals (in other words the apparatus is a ‘receive’ only apparatus). In such embodiments the system within which embodiments may be implemented may comprise apparatus with varying abilities to capture / render audio signals. The teleconference apparatus (for each site or room) 110, 120 can be configured to call into a teleconference controlled by and implemented over a server 111. In some embodiments the communications or teleconferencing system comprises a (peer-to-peer) communications system (rather than the server based system shown in Fig.1) within which some embodiments can be implemented. Thus, for example, two or more UEs can be configured to interact directly with each other (for example to implement an immersive audio phone call between users). In such a scenario one of the UEs can be configured to deliver spatial ambience as one stream and employ a close-up microphone (for example a Lavalier microphone) to capture the speech as an audio object or audio source. The sender UE can be configured to encode the spatial ambience audio signals in a MASA format stream and the close-up microphone audio signal as an object format stream. The two audio streams can then be delivered as separated IVAS streams. The sender UE, in addition, can be configured to encode processing information during the encoding to deliver the PI frames together with the IVAS frames to the receiver UE. The teleconference apparatus can be configured to spatially capture and encode the audio environment and furthermore can be configured to render a spatial audio signal to the room. In this example only the communications or signalling path from the Room A 100 to the Room B 102 is shown for simplicity but a duplex or multipoint communication system comprising multiple signalling paths can be implemented using the methods as described herein without significant inventive input. The teleconference apparatus (for each site or room) 110, 120 is further configured to communicate with each other to implement a teleconference function. As shown in Fig.1, the apparatus 110, 120 and server 111 can comprise suitable encoder and decoder functionality. For example, the apparatus 110 is shown comprising an (IVAS) encoder 101, the server 111 is shown comprising a (IVAS) decoder and encoder 121 and the apparatus 120 is shown comprising an (IVAS) decoder 131. In such a manner an object 120 (the audio signals representing the user or talker 1111) can be encoded by the encoder 101 which generates a bitstream 106 to be passed to a server 111. The server 111 can then decode, (optionally then mix with other objects and otherwise process the audio signals) and encode then to generate the bitstream 108 to be passed to the apparatus 120. The apparatus 120 can then decode the audio signals and present them to the user or talker ‘Talker RX’ 141. Although this example shows a teleconference application the encoder / decoder functionality can be applied to the streaming of any suitable media. The IVAS decoder / renderer for each of the teleconference apparatus 102 can be furthermore configured to handle multiple input streams that may each originate from a different encoder. The IVAS codec algorithm which is employed in the above example is described in 3GPP TS 26.253 (Codec for Immersive Voice and Audio Services; Detailed Algorithmic Description incl. RTP payload format and SDP parameter definitions), currently at v2.0.0 (SP-240030). Furthermore the IVAS codec floating-point C code is provided in 3GPP TS 26.258. The IVAS codec bitstream contains information on the contents of each bitstream frame. This information includes the aforementioned input format and additional information on the present format, i.e., sub format signaling, (e.g., order of Ambisonics, channel layout of MC, number of transport channels, etc.). This is in addition to the encoded audio signal and possible metadata. Furthermore, IVAS bitstream frames are designed to be standalone in such way that IVAS decoder can start decoding from any valid IVAS frame (containing full data of for the format specified in the bitstream) and produce good quality output. As discussed previously RTP is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP is furthermore designed to carry a multitude of multimedia formats, which permit the transport of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may therefore require a profile and payload format specifications. The profile is configured to define the codec used to encode the payload data and the mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header. For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codec such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798. The IVAS RTP payload format is currently being developed, and the latest state is described in TS 26.253 Annex A (v2.0.0, SP-240030). Recent proposed changes are also described in a CR document S4-241325. An RTP session can be established for each multimedia stream. Audio and video streams may be implemented which use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification can furthermore be configured to recommend port numbers for RTP, and furthermore to recommend the use of the next odd port number for the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols. Each RTP stream can comprise RTP packets, and the RTP packet in turn can comprise a RTP header and payload pair. Fig.2 shows an example IVAS RTP packet structure (following S4-241325) 200. The RTP Header (and possible RTP Header Extension) 201 follow a known RTP design. The structure 100 further comprises a IVAS payload 211. The IVAS payload 211 comprises a payload header 203 section, a (IVAS) frame data 205 section and an optional PI (processing information) data 207 section. The payload header 203 section can comprise different types of header bytes, for example: ToC (Table of Content) and E-bytes (Extra bytes, or further bytes which can be used to define or signal aspects of the payload). The ToC bytes can be employed to describe the content of the frame data section (by indicating the size / bitrate for the data frames). The ToC bytes also can differentiate IVAS frames from EVS frames, in situations where the IVAS is operating in mono EVS mode. The E-bytes can signal additional information, such as Codec Mode Requests (CMR) which indicate a request to change the bitrate (and possibly other configurations like format and bandwidth) of an incoming IVAS stream. E-bytes may also be used to explicitly indicate the presence of the PI data section at the end of the payload. The frame data 205 section can include the IVAS data frames (and EVS data frames in a case of employing IVAS in mono mode). The data frames represent the encoded IVAS bitstreams. The bitstream includes the encoded IVAS audio data with possible additional metadata. The bitstream can also include initialization data or information, for example information about the input / coded format and sub-format of the encoded data (e.g., multichannel 5.1 format). Any EVS frames do not include such format data. The PI data 207 section can include (PI) processing information data and related headers. PI data can be used to transmit any (non-audio) data that can be used to assist the rendering or processing of the audio data, such as, for example, scene and device orientation data. The PI data can also be used to request something from the other session participant, such as, for example to mute an incoming stream or increase noise suppression. Feedback data can also be transmitted, for example, the head orientation of a listener. In IVAS, the coded format (e.g., MASA, SBA or MC) of an IVAS bitstream is only indicated within the IVAS bitstream. Consequently, for a receiver to know the coded format of a received bitstream, the receiver has to first receive and then decode the bitstream. In order to decode the bitstream, the receiver initializes a full IVAS decoder instance. The initialization process includes setting an output format for the decoder. In other words the output format for the decoder is set before any knowledge of the coded format of the incoming bitstream. The aim of the embodiments presented herein is an attempt to overcome this requirement, as the coded format for the incoming bitstream affects which output format should be chosen for maximum quality of experience and maximum decoder interoperability. Thus, the embodiments describe a mechanism to indicate the coded format for an incoming bitstream to a receiver before the need to initialize a full IVAS decoder instance at the receiver. The current implementations of the IVAS codec standards are limited in support for switching between transmitted IVAS formats within an encoder (and especially within a single IVAS session and without initialization) and furthermore switching between output formats within a decoder (similarly within a single IVAS session and without initialization). The embodiments as described herein describe practical implementations of an IVAS decoder which is pre-initialized with a configuration that requests a specific output format. The decoder can then be fully initialized when the transmitted IVAS format is known. This can, in some embodiments, be determined when the first ‘good’ IVAS frame is read from the bitstream and is provided to the decoder. After this, the decoder is fully locked or defined or determined for this IVAS format and an output format combination and no switching is supported for this decoder instance. The first ‘good’ IVAS frame is a frame that contains sufficient information of the IVAS format to allow full initialization of the decoder. Such frames may be at least a normal IVAS frame and a silence descriptor (SID) frame of DTX operation. Any other frame types carrying the required information are also sufficient. The following examples demonstrate the flexibility of the embodiments described herein. In a first example, a sender device (containing at least an IVAS encoder) and a receiver device (containing at least an IVAS decoder) are configured to negotiate IVAS format options and receiver may also present possible output formats. The negotiation, in this example, ends in offering the sender the possibility to use IVAS formats “mono” or “SBA”. The sender device, completely independent from the receiver device, selects an IVAS format for transmission. The information of the IVAS format is then included only in the bitstream provided by the IVAS encoder and is not known by the receiver device (other than one or more formats selected as part of SDP negotiation). In the situation where there is SDP negotiation with multiple preferences for IVAS formats listed in preference order, for example, listed as (SBA, mono) in a preference order (selected by the receiver), the sender is likely to select the first format. However, this behaviour is not guaranteed, and the sender device could select whichever format is supported by the negotiation so the problem of selection of transmitted format by the sender device independent of the configuration of the receiver device remains a possibility. At the same time, the receiver device prepares for decoding the incoming bitstream. This can be implemented by pre-initializing the decoder with the desired output format and other configuration parameters. This enables the decoder to be configured such that it is able to decode a frame from the bitstream. Once the decoder receives a first ‘good’ frame, then the decoder is fully initialized after decoding the IVAS format (and possible sub details or other information) from the bitstream. The receiver device, however, has to select an output format and pre-initialize with it before knowledge of what IVAS format is in the following bitstream. Furthermore there is no way of determining this information without decoding a ‘good’ IVAS frame from the bitstream. In other words, the initialization of the IVAS decoder must be delayed until arrival of the first ‘good’ IVAS frame. Hence any network congestion, where the first ‘good’ IVAS frame is delayed, can result in further temporal delay being added with regards to the initialization of the decoder. Thus, in a first scenario, the receiver device selects an immersive output format, for example a binaural output format as this is the intended main use case for IVAS as indicated above. However as mono is an allowed IVAS format for the bitstream, the sender device selects this mode and generates a mono format bitstream which is provided to the receiver device. However a mono format to binaural format rendering is not supported by the current IVAS algorithm and reference implementation. The decoder in the receiver device according to this scenario could at a minimum generate an error or crash completely. The receiver device can attempt to prevent this error or crash (if it detects it), by clearing the existing decoder instance (or resetting the decoder configuration) and creating a new decoder instance, at the cost of time reconfiguring a new decoder instance and wasting all of the preparation work on the existing decoder instance. In a second scenario, the receiver device prepares or is configured to allow an output format that is supported by any negotiated IVAS format. In this scenario, this would be a mono output as a mono encoded bitstream only supports this output format. In such a situation, any received IVAS frame is able to be decoded successfully but the output format is suboptimal. In a third scenario a passthrough output is selected at the receiver device (an EXT output in IVAS configuration) where the output is what is received as a signaled input of the IVAS encoder. This would result in a mono output format for a received mono IVAS format and a specific order Ambisonics output format for a received SBA IVAS format. However, this could cause complexity problems as decoding a SBA IVAS format to generate Ambisonics format output can be significantly more complex than generating a binaural format output, when only a binaural format output is required. This can be a problem for resource constrained receiver devices which may be equipped to handle all the possible output formats. In a further example, a decoder is firstly initialized with a stereo output format. In IVAS, all the immersive coded or input formats support rendering to stereo output. Additionally, mono input to stereo output can be handled by trivial dual-mono playback, which is already done in practice with mono codecs. Moreover, IVAS is an immersive audio codec which is intended to be used with immersive input and output formats. Stereo output would result in a higher immersion for the listener than mono output, and as such is more desirable for the listener. Although stereo provides a common output format, it cannot be considered truly immersive output format and immersive transmissions (e.g., MASA, SBA, and MC) do not achieve optimal output quality with it. In a further example, a sender device intends to send a multi-channel (MC) IVAS format bitstream to a receiver device. A negotiation is performed between the devices to include a MC IVAS format as the only possible bitstream format but it is not known beforehand what is the channel layout of the MC format at the sender device. The receiver device furthermore intends to employ loudspeaker playback for the MC format as this provides optimal quality. However, the receiver in this example has a horizontal loudspeaker configuration (in other words no elevated loudspeakers that would be indicated as +X in setup configuration) and can in practice correctly support 5.1 or 7.1 multichannel setups. In some embodiments, formats with elevated speakers are indicated with a third number. For example 7.1+4, 7.1.4, and 7_1_4 may indicate identical format with 7 horizontal level loudspeakers, 1 subwoofer or LFE channel, and 4 elevated (i.e., non-horizontal level) loudspeakers. Thus the receiver prepares or configures a decoder instance for decoding bitstream. To pre-initialize, the receiver device needs to supply the decoder with an output format control. If both receiver device and sender device select the same layout, then an optimal quality is received. However, if this does not happen, then there can be a loss in quality. For example: Receiver device selects 5.1, and sender device selects 7.1. A downmix is performed and the resultant quality is not optimal. The receiver device could have selected 7.1 output format for optimal quality but did not know to do it beforehand and cannot switch later; Receiver device selects 7.1, and sender device selects 5.1. An upmix is performed and quality is not optimal. Moreover, the decoding complexity is higher than necessary as 8 output channels are created, whereas 6 would have been enough. The Receiver device could have selected 5.1 output format for optimal quality and complexity but did not know to do it beforehand and cannot switch later; Receiver device selects 7.1, and sender device selects 5.1+4. A downmix is performed and the resultant quality is not optimal as it would be better to downmix to 5.1 which is a direct subset of 5.1 +4. Again, the receiver device did not know this beforehand; Receiver device selects EXT output (to get passthrough operation for 5.1 and 7.1), and sender selects 7.1+4. The decoder therefore produces extra output channels and the receiver device needs to either perform separate downmix (increased complexity) or discard some of the channels (not optimal quality). Thus, if the receiver does not know what channel layout to initialize for MC output beforehand, then there can be a loss in quality of the output and / or an increase in complexity. The receiver device could delay the initialization of the decoder until the reception of the IVAS frames. However this prevents using the IVAS frames as part of an integrated media processing pipeline. Therefore configuration based on the bitstream requires some form of decoding the bitstream which is not optimal for efficient implementations. The embodiments as disclosed in the following relates to apparatus or methods for initializing the immersive audio codec where there is provided initialization information during or after session setup (for example during or immediately after a call between two parties) aiming to achieve early initialization of the immersive audio decoder without the need to wait for the arrival of the immersive audio frame from the media sender to the media receiver. This can be achieved by one of the following approaches: providing decoder initialization information as a processing information packet which is transmitted via RTP which precedes the IVAS frame delivered via RTP; providing a pre-processing stage at the receiver device to initialize an IVAS decoder based on the received IVAS frame; providing media sender codec properties via SDP during session setup to enable initialization of the decoder prior to reception of IVAS frames from the media sender. In some embodiments these can be applied to the 3GPP IVAS standard and more specifically the standard definition, the RTP payload definition of the standard, and implementations of the standard. The concept as discussed further in the following embodiments with respect to IVAS and RTP can be applied to other transmission codecs that negotiate format options. In some embodiments, a parameter is defined and received in IVAS RTP signaling which contains the IVAS format and specifics of the format necessary for successfully initializing the decoder before the IVAS bitstream frame is decoded. If this signaling or data is received sufficiently early, then any computational spike caused by the ‘full’ initialization operation can furthermore be mitigated. To achieve this early initialization the IVAS decoder is provided with at least one IVAS frame of predefined “initialization data” from a store (for example stored on the receiver device in memory) or received from a transmission. In some embodiments, a process or method can be defined for the receiver device, where the receiver device comprises a (non-standard lightweight) decoder of IVAS bitstreams. In this approach, the receiver device is paused or waits to apply or instigate initialization until the lightweight decoder provides full format information. Once this is full format information is received and decoded by the (lightweight) decoder, then the full initialization (pre-initialization included) is implemented. In some embodiments the lightweight decoder is a separate decoder to the ‘main’ decoder or is implemented as a reduced complexity version of the decoder or decoders. In some embodiments, a process or method is defined for the receiver device, where the receiver device pre-initializes multiple output format decoders and using a (lightweight non-standard) decoder or IVAS RTP signaling, the correct one of the pre-initialized decoders is then selected for decoding the bitstream (once the information is available). In some further embodiments, all of the pre-initialized decoders are configured to receive one (or more than one) frame of incoming bitstream to determine which output formats are usable and then an optimal decoder output is selected for use. To initialize an IVAS decoder, the coded format for the received bitstream needs to be known. In some embodiments initialization data or configuration data can be transmitted which comprises the coded format. The coded format can be defined by a major format defining a specific type format and the m inor format which 5 can aid or assist in not only defining a format type but also configuration relating to the format type. For example the following table can define major formats IVAS input / coded format (major) Mono Stereo MC MASA ISM SBA OMASA OSBA And the following table can define minor formats (for example indicate the 10 channel configuration of the multichannel MC format type. IVAS input / coded format (minor) Mono Stereo Binaural MC 5.1 MC 5.1.2 MC 5.1.4 MC 7.1 MC 7.1.4 MASA1 MASA2 ISM1 ISM2 ISM3 ISM4 SBA (FOA) SBA (HOA2) SBA (HOA3) SBA (planar FOA) SBA (planar HOA2) SBA (planar HOA3) OMASA (1 ISM) OMASA (2 ISM) OMASA (3 ISM) OMASA (4 ISM) OSBA (1 ISM, FOA) OSBA (2 ISM, FOA) OSBA (3 ISM, FOA) OSBA (4 ISM, FOA) OSBA (1 ISM, HOA2) OSBA (2 ISM, HOA2) OSBA (3 ISM, HOA2) OSBA (4 ISM, HOA2) OSBA (1 ISM, HOA3) OSBA (2 ISM, HOA3) OSBA (3 ISM, HOA3) OSBA (4 ISM, HOA3) OSBA (1 ISM, planar FOA) OSBA (2 ISM, planar FOA) OSBA (3 ISM, planar FOA) OSBA (4 ISM, planar FOA) OSBA (1 ISM, planar HOA2) OSBA (2 ISM, planar HOA2) OSBA (3 ISM, planar HOA2) OSBA (4 ISM, planar HOA2) OSBA (1 ISM, planar HOA3) OSBA (2 ISM, planar HOA3) OSBA (3 ISM, planar HOA3) OSBA (4 ISM, planar HOA3) Additionally, the output configuration (e.g., binaural or stereo output) is defined, configured or set during the initialization phase of the decoder. The aim of the embodiments as discussed herein is to counter or avoid the issues that were presented in the above example cases. The embodiments presented herein therefore either introduce RTP level signaling to prepare the decoder (in an early manner) correctly, or perform a flexible initialization of the decoder based on the IVAS bitstream and avoiding the decoder to being locked into a non-optimal output format. With respect to Fig.3a is shown a schematical representation of current decoder configurations. The current decoder comprises an IVAS frame (as payload) input 300 which is configured to receive or otherwise receive an IVAS frame 211 as a payload of RTP signalling (such as shown with respect to Fig.2). In other words, the IVAS frame is received or otherwise obtained as an IVAS coded bitstream. The IVAS frame 211 as the payload of the RTP is passed to a IVAS decoder 301. Additionally the IVAS decoder 301 is configured to receive the application initialization 302. The IVAS decoder 301 then is configured to decode the IVAS frame 211 based on being configured by the application initialization 302 information to generate the rendered audio signals output 304. However as the IVAS frame 211 as the payload of RTP and the application initialization 302 information are not aware of each other or have no prior communication then this can lead to the issues in the ‘locked’ configuration of the decoder creating a sub-optimal quality audio output or a sub-optimal use of resources. In other words a disconnect between the application initialization 302 information and received IVAS bitstream 211 can be a drawback for scenarios where the twin requirements are addressed: Precise alignment between the bitstream and the reproduction to maximize consistency of reproduction quality; and Optimal computational resource utilization. With respect to Fig.4 is shown example operations with respect to the sender 410 and receiver 420 based on the system shown in Fig.3a. Thus for example as shown by 401, the decoder is pre-initialized with a specific format output, in this example a 7.1 multichannel (MC) output format. Then the decoder within the receiver 420 is configured, as shown by 402, to receive from the sender 410 a first ‘good’ IVAS frame with encoded audio in a specific input format. In this example the input format is a 5.1 multichannel (MC) input format. Furthermore, as shown by 403, the decoder is then configured to receive the 5.1 MC input and produce 7.1 MC output (for example by upmixing the 5.1 MC input). Hence as shown by 404, the IVAS Decoder 301 is configured to provide a 7.1 MC output. This results in a sub-optimal approach, as a 5.1 MC output would have been sufficient in terms of quality and better in terms of complexity. Fig.3b shows schematically an example system according to some embodiments. In this example the system comprises an IVAS frame (as payload) input 300 which is configured to receive or otherwise receive an IVAS frame 211 as a payload of RTP signalling (such as shown with respect to Fig.2). In other words, the IVAS frame is received or otherwise obtained as an IVAS coded bitstream. The IVAS frame 211 as the payload of the RTP is passed to a bitstream decoder 311. The bitstream decoder 311 is configured to extract from the bitstream information about the decoded bitstream 310 and pass this information 310 to a renderer configurator 313. The renderer configurator 313 having received the information 310 is configured to generate output format configuration information 312 and pass this to a renderer 315. The renderer 315 is configured to receive the output format configuration information 312 and from this configured itself such that it can generate a suitable rendered audio signals output 304 from the decoded bitstream generated by the bitstream decoder 311. In other words, an intermediate entity (between input and renderer) in the media processing pipeline is employed to read the bitstream information and subsequently configure the output parameters to satisfy the twin requirements listed above. The benefit of these embodiments is that the system can be realized without any additional signaling of information from the encoder within the bitstream. However, the media processing pipeline in these embodiments requires an additional tap for at least part of the decoded bitstream for the configurator. Fig.3c shows schematically a further example system according to some embodiments. In this example the system comprises an IVAS frame (as payload) input 300 which is configured to receive or otherwise receive an IVAS frame 211 as a payload of RTP signalling (such as shown with respect to Fig.2). In other words, the IVAS frame is received or otherwise obtained as an IVAS coded bitstream. The IVAS frame 211 as the payload of the RTP is passed to a IVAS decoder (initialization data or information) 321. The IVAS decoder (initialization data or information) 321 is configured to extract from the bitstream initialization related information 322 from the bitstream and pass this information 322 to a renderer configurator 323. In some embodiments this bitstream initialization related information 322 is provided as additional signaling information within the RTP header or RTP payload. The renderer configurator 323 having received the information 322 is configured to generate output format configuration information 324 and pass this to a bitstream decoder and renderer 325. The bitstream decoder and renderer 325 is configured to receive the output format configuration information 324 and from this information configure itself such that it can generate a suitable rendered audio signals output 304 from the partially decoded bitstream generated by the IVAS decoder (initialization data or information) 321. In other words, the additional signaling information is included together with the bitstream payload in such a manner to provide easy access to the ‘compact’ information for initialization of the decoder in accordance with the bitstream. In such embodiments there can be a closely coupled decoder and renderer without the need for any additional tap for the decoded bitstream. As discussed previously although the examples discussed in Fig.3b and Fig.3c demonstrate embodiments with respect to IVAS frame bitstreams it would be appreciated that these examples and the following methods can be employed on other bitstream format where the codec may receive a bitstream which is not mutually exclusive (or where the decoding of one format is not different from another format). In other words, these embodiments can be extended to apparatus and methods where the (coded or input) format of the bitstream is encoded within the bitstream itself and the format is only obtainable by decoding the bitstream. In some embodiments the above examples can be implemented by introducing RTP signaling (or equivalent signaling outside of IVAS bitstream) to signal the IVAS format and the related sub format. This signaling can be included, for example, in an IVAS RTP transmission before the IVAS codec bitstream (containing the audio signals) is transmitted. Once the receiving device reads from the IVAS RTP packet the information (for example as the initialization data or information 322 or the decoded bitstream information 310) for the incoming IVAS format / sub-format, the receiving device can successfully initialize the IVAS decoder for the combination of IVAS format, sub format information, and output format. The initialization of the IVAS decoder can be achieved by providing the IVAS decoder (implemented as the bitstream decoder 311 or IVAS decoder- initialization data 321 as shown in Fig.3b or Fig.3c respectively) with a single frame of IVAS bitstream preconstructed or defined for the specific format. This preconstructed bitstream frame can be obtained from a suitable storage (memory) or constructed on demand by a deterministic algorithm for the purpose. In other words the pre-constructed init-frames are stored on the receiver and used when signalling indicates the format (but not yet the full frame). However in some embodiments similar data to the pre-constructed init-frames can be sent to the receiver. In some embodiments the signaling further enables any processed output (based on this initialization information call to the IVAS decoder) to be discarded. In other words the signaling is identified as not comprising any suitable audio signal information and only comprises information initializing the decoder (for example the format / sub-format information). In some embodiments, the preconstructed bitstream frame advantageously would also match the expected bitrate for the upcoming IVAS bitstream to avoid reconfiguration of the decoder. However, in some situations the bitrates between the IVAS bitstream and the (initialization) IVAS bitstream differ. The RTP signaling for the purpose of this embodiment can comprise (directly) the format signaling bits of the IVAS bitstream. In some other embodiments the format signaling bits can be constructed as follows. The initialization data (information) or configuration data which can comprise the input or coded format indication can be included, for example, in the PI data section 207 of an IVAS RTP payload 211. In some embodiments a PI type of INPUT_FORMAT or CODED_FORMAT or MAJOR-FORMAT or MINOR_FORMAT or any other suitable PI type name could be used to transmit or signal the input / coded format / major format / minor format / sub-format information. The format / sub-format information can, in some examples, comprise the major IVAS format such as shown in the table above or as a more detailed sub-format can be described (such as the minor format table shown also above). In some embodiments, the format / sub-format can be described as a combination of the major and minor formats. The values in the tables are examples and can be even more detailed (in terms of format) or less detailed in some examples. The tabulated formats can also be named or defined differently, for example, a mono format could be referred to as a EVS format. In some embodiments the initialization data (information) or configuration data which can comprise the input / coded format can be transmitted in any suitable format, for example, as a RTP header extension and / or as a RTCP. In some embodiments, the initialization data (information) or configuration data which can comprise the input / coded format can be transmitted or otherwise included as part of decoder / renderer / receiver initialization or decoder / renderer / receiver configuration packets or payloads. For example, the initialization or configuration packets could be transmitted in the PI data section of an IVAS RTP payload. In some embodiments the PI type for the initialization or configuration packets could be named as DECODER_INIT or DECODER_INITIALIZATION or DECODER_CONFIG or DECODER_CONFIGURATION or any other suitable PI type name. In some embodiments the initialization or configuration PI type could have a flag or other indication indicating that the associated PI data should be processed and taken into account before decoding the IVAS frames in the payload. The flag or indication could be embedded into the PI header section or it could be part of the PI data. In some embodiments the flag or indication could be indicated in the E-bytes section. For example, some of the reserved bits in the E-byte indicating the presence of PI data could be used to indicate that the PI data should be processed before decoding the IVAS frames. In some embodiments, the initialization or configuration packets could be identified in the IVAS RTP payload header section. For example, some bits in the subsequent E-bytes could be reserved to signal the presence of initialization or configuration packets in the IVAS payload. The initialization or configuration packets could follow the E-bytes section or the packets could be positioned otherwise suitably in the IVAS payload. In some embodiments, the initialization or configuration data could be transmitted during an active session. For example, if the input / coded format of the transmitted bitstream changes during the session, an initialization or configuration data packet could be transmitted to a receiver to signal the change in the configuration data. The initialization or configuration data packet could be transmitted prior the coded / input format change is active to allow the receiver side to adapt to the change more robustly. For example, the initialization or configuration data packet could indicate X number of frames or seconds or any other time value when the change takes place in the future (a similar approach is described in GB patent application GB2317940.1 for input format change). In some embodiments, the transmitted encoded frames can be in an EVS format (as IVAS supports EVS encoding and decoding for mono audio). In these cases, it is beneficial to set the decoder to produce mono output, as the EVS frames represent mono encoded audio and there is no spatial or stereo output defined for them. In such embodiments the EVS format can be explicitly signaled similarly as for IVAS formats (with the PI frames, RTP HE or RTCP, e.g.). However, the ToC header bytes in the IVAS RTP payload can also differentiate between IVAS and EVS frames. The EVS format can also be determined from the ToC byte and the decoder can be properly / more optimally initialized based on this information. One key additional benefit in this approach is that the decoder can be fully initialized before there is a need to decode a bitstream for an output. As initialization impacts on the processing requirements (or has a significant complexity impact), this approach can reduce the overall maximum complexity experienced by the receiving device. With respect to Fig.5 is shown example implementation operations for this type of signaling embodiment, and which can be employed using the embodiments shown in Fig.3b and Fig.3c. Thus, for example, the sender 510 is configured to signal the initialization or configuration data comprising the format / sub-format information to the receiver 520 which can comprise a decoder 551 which can comprise, as shown in Fig.3b, a bitstream decoder 311 / renderer configurator 313 / renderer 315 system or with respect to Fig.3c, a IVAS decoder (initialization data or information) 321 / renderer configurator 323 / bitstream decoder and renderer 325 system. As shown by 501, the decoder 551 receives an IVAS RTP packet which comprises initialization or configuration data which in turn explicitly indicates the used IVAS input / coded format (for example a 5.1 multichannel format or 5.1 MC). Then as shown by 502, the decoder 551 is initialized with a specific format output, in this example a 5.1 multichannel (MC) output format based on the received initialization or configuration data and the indicated or signaled 5.1 MC input format information. In other words, the decoder 551 is initialized with both 5.1 MC input and output. Then the decoder 551 is configured, as shown by 503, to receive from the sender 510 the IVAS frames with encoded audio in the indicated specific input format. In this example the IVAS frames are 5.1 multichannel (MC) input format frames. Hence as shown by 504, the decoder 551 is configured to provide a 5.1 MC output which results in an optimal approach, as the 5.1 MC output is sufficient in terms of quality and better in terms of complexity. In some embodiments, the initialization or configuration data or information transmitted in the IVAS RTP packet can be transmitted in initialization or configuration frames. The initialization frames may only contain the initialization information without any audio data. In some embodiments, the initialization frames can be distinguished or detected in an IVAS payload with a ToC (Table of Content) byte or with an E-byte in the IVAS payload header. The initialization or configuration information can also be delivered in some embodiments as another IVAS frame type. For example, as a new type of TOC (Table of Content). Furthermore, in some embodiments, the initialization information can be delivered prior to sending a full conventional IVAS frame. This can enable the decoder to be initialized and ready to decode the IVAS frame. Also in some embodiments, the initialization information can be sent as early media to enable initialization of the decoder and make it ready to decode IVAS frames. In some embodiments the initialization information can be obtained from a decoded immersive audio frame. In some embodiments, as mentioned previously, rather than signal using a separate signaling operation a standard IVAS signaling and bitstream is employed but a purpose-build lightweight (non-standard) IVAS decoder is employed in a manner such as shown in Fig.3c. The IVAS decoder (initialization data or information) 321 is designed such that it can receive any standard compliant IVAS bitstream frame but is configured to only decode the bitstream to obtain the necessary data for determining the IVAS format and the related sub format signaling. In these embodiments, the receiver is configured to not pre-initialize the IVAS decoder (Bitstream decoder and Renderer 325) rather the receiver runs the lightweight decoder (IVAS decoder - initialization data or information, 321) for each arriving IVAS bitstream frame until IVAS format and sub format signaling information can be decoded from it. In some embodiments where the receiver obtains IVAS frames that are valid, the lightweight decoder can be configured to decode IVAS format and subformat signalling information from the IVAS frame. This can assist a subsequent initialization of the IVAS decoder to decode and render IVAS frames. Such a format decoder can be included within any part of the IVAS decoder. At this point, the receiver, for example the renderer configurator 323, selects or decides the suitable output format based on the initialization or configuration data comprising the IVAS format and (possibly) sub format signaling and fully initializes the standard IVAS decoder (Bitstream decoder and Renderer 325) using the optimal output format and the received IVAS bitstream frame. With respect to Fig.6 is shown example implementation operations for this type of lightweight decoder (or hybrid or split decoder) embodiment, and which can be employed using the embodiments shown in Fig.3c. Thus, for example, the sender 610 is configured to send, as shown by 601, to the receiver 620 and the lightweight decoder 649 which can be represented with respect to Fig.3c, a IVAS decoder (initialization data or information) 321 ‘bad’ IVAS frames which do not comprise any input / coded format indication. Then the receiver 620 and the lightweight decoder 649 receive a first ‘good’ IVAS frame which does comprise initialization or configuration data comprising the input / coded format indication (for example a 5.1 multichannel format or 5.1 MC). The lightweight decoder 649 can, as shown by 603, then inform the actual decoder 651 (which can comprise the renderer configurator 323 and bitstream decoder and renderer 325) system about the input / decoded format (MC 5.1) obtained from the initialization or configuration data and also passes the collected IVAS frames to the decoder. The decoder can then be initialized with a 5.1 MC input and output and any collected IVAS frames can be decoded as shown by 604. Then, as shown by 605, the decoder 651 receives IVAS frames with encoded audio in the indicated specific input format. In this example the received frames are 5.1 multichannel (MC) input format frames. Hence as shown by 606, the decoder 651 is configured to provide a 5.1 MC output which results in an optimal approach, as the 5.1 MC output is sufficient in terms of quality and better in terms of complexity based on the 5.1 MC format input. In some embodiments the light-weight decoding could be integrated as part of a standard IVAS decoder as shown in Fig.3a or bitstream decoder for example as shown in the apparatus shown in Fig.3b. In such implementations the (bitstream) decoder can comprise a hybrid light-weight / full decoder wherein the light-weight decoder is able to extract from the IVAS frame the IVAS format and sub format signaling information and pass this to a renderer configurator which is able to configure the renderer while the full decoder extracts the IVAS frame audio signals and passes these to the renderer. In some embodiments, this light-weight decoder is implemented as a partial decoding algorithm within the full decoder. In this case, the light-weight decoder may be part of the standard. In some further embodiments there can be employed multiple decoder instances, each of which are pre-initialized at the receiver device to cover all the supported (or expected) output formats. These multiple decoder instances are then maintained until information is received about which IVAS format and related sub format is employed in the received IVAS bitstream. The source for this information can be the information integrated within the RTP signaling as described in the first example or from the light-weight (or hybrid) decoder example. Once this information is determined, one of the decoder instances (a ‘correct’ or optimal decoder instance) can be selected based on the desired output format, IVAS input format, and the IVAS input sub-format combination. In some embodiments the other decoder instances can then be removed or deactivated at this point. These embodiments aim to solve the problem of optimal format selection as an output format can be selected based on the IVAS format and (possibly) sub format. However, these embodiments require additional memory or processing resources to pre-initialize possible decoders and then maintaining these possible decoders at the same time. However, this resource impact can be minimized as pre-initialized codec instances do not have renderer memory reserved which is more significant. In some embodiments all of the pre-initialized (possible) decoders are employed to decode the incoming bitstream. In some embodiments this can be implemented as a parallel decoder configuration (where one or more instances are employed to decode the input IVAS bitstream) or as a serial decoder (where a preference order or other order of decoder instances are employed). Then a decoder instance is selected based on which instance successfully produces an output (without generating an error or crashing) or selecting a decoder instance which produces an output without any up- or downmixing or conversion operations. This instance selection based on the analysis of the output can result in significant resources being required due to multiple full initializations of the decoder before the selection occurs. An example for these decoder instance selection embodiments is presented with respect to Fig.7, and which can be employed using the embodiments shown in either of the example configurations of Fig.3a, Fig.3b or Fig.3c. For example, as shown by 701, the receiver 720 is configured to employ multiple decoders instances each of which are pre-initialized to output a different output format. Thus as shown in Fig.7 there is shown a decoder 761 which is configured to output a 5.1 MC output based on a 5.1 MC output configuration 760, a decoder 771 which is configured to output a 5.1+4 MC output based on a 5.1+4 MC output configuration 770, a decoder 781 which is configured to output a 7.1 MC output based on a 7.1 MC output configuration 780, and a decoder 791 which is configured to output a 7.1+4 MC output based on a 7.1+4 MC output configuration 790. Then, the sender 710 is configured to send, as shown by 702, to the receiver 720 and the decoder instances 761, 771, 781, 791, a ‘good’ IVAS frame which comprises initialization or configuration data comprising the input / coded format indication (for example a 5.1 multichannel format or 5.1 MC). The decoder instance which has been pre-initialized with the ‘optimal’ output format, for example the 5.1 MC output format pre-initialized decoder 761 is then selected as shown by 703. Furthermore, the other decoder instances 771,781,791 can be discarded. The selected decoder 761 thus receives IVAS frames with encoded audio in the indicated specific input format and output the pre-initialized 5.1 MC output format. Then, as shown by 704, the selected (5.1 MC output format) decoder instance 761 is configured to receive IVAS frames with initialization or configuration data indicating the encoded audio in the indicated specific input format. In this example the received frames are 5.1 multichannel (MC) input format frames. Then as shown by 705, the selected (5.1 MC output format) decoder instance 761 is configured to provide a 5.1 MC output which results in an optimal approach, as the 5.1 MC output is sufficient in terms of quality and better in terms of complexity based on the 5.1 MC format input. To initialize an IVAS decoder, the output format and output sampling rate are required to be set. As described in this invention, the coded format of the incoming IVAS frames can be taken into account when determining the output format for the decoder. For optimizing the decoder processing further, also the sampling rate used at the sender can be taken into account when initializing the decoder. The encoder input sampling rate determines the maximum frequency limit for the encoded audio. For example, if the used input sampling rate is 16 kHz, based on Nyquist theorem, the maximum frequency limit for the encoded audio is 8 kHz. Audio content above 8 kHz will have artifacts and is thus not usable. The input sampling rate used at the encoding is not explicitly signalled in the IVAS bitstream. I.e., the receiver of an IVAS bitstream does not know which sampling rate was used for the encoder input. If a sampling rate of 16 kHz was used for the encoder input, the decoder output should choose the same sampling rate for optimal decoder utilization. If a higher sampling rate is chosen for the output, rendering resources are wasted as the higher frequencies (> 8 kHz in this case) do not have any meaningful content. Rendering uses less resources for lower output sampling rates. To optimize the decoder processing further, the initialization or configuration data can comprise signalling or information identifying an input sampling rate used at the encoding which could be signalled to the receiver before the decoder is initialized. Similarly as for the coded format signalling, the input sampling rate could be signalled via IVAS RTP packets within initialization or configuration data. The sampling rate information could be included in a PI type of INPUT_SAMPLING_RATE or ENCODER_SAMPLING_RATE or ENCODER_INPUT_SAMPLING_RATE or any other suitably named PI type. The encoder input sampling rate could also be included in the decoder / renderer / receiver initialization or configuration packet. In some embodiments, as shown in Fig.7, there can be additional pre-initialized decoder instances with varying output sampling rates (in addition to varying output formats). The decoder instance with a matching output sampling rate with the encoder input sampling (and a matching output format) can then be selected to be used following 703. In some embodiments, the encoding sampling rate could be indicated within the received IVAS packet or payload and used in the process of determining the matching output sampling rate. In some other embodiments, the matching output sampling rate could be determined by inspecting the produced output signals of the different decoder instances. For example, if a low encoder input sampling rate (e.g., 16 kHz) is used, higher output sampling rates should not produce higher quality output. Therefore, a low output sampling rate (e.g., 16 kHz) can be determined to be used. In another example, if the encoder input uses a high sampling rate (e.g., 48 kHz), the output should degrade when using a lower output sampling rate (e.g., 16 kHz). Therefore, a high output sampling rate (e.g., 48 kHz) can be determined to be used. Fig. 13 shows an example decoder initialization scenario, where initialization / configuration data is fed to the decoder 1380. Some of the initialization or configuration data 1360 is obtained from the sender 1310. This includes the IVAS coded format 1361 of the incoming bitstream and the encoder input sampling rate 1363. The initialization or configuration data or information 1370 from the receiver 1320 includes output format 1371 and output sampling rate 1373. Both the output format and output sampling rate can therefore depend on initialization or configuration data or information 1360 from the sender 1310. Fig.8 shows an example flow diagram showing IVAS encoder and decoder initialization and processing operations. The flow diagram is shown separated by the vertical line between encoder / sender device 810 and decoder / receiver device 820. Furthermore time flow is shown as a vertical line down Fig.8. Thus, there is shown by 801 a negotiation phase wherein IVAS coded format for the session is negotiated. This can involve the session parameters (including coded format) being negotiated for the IVAS session. After the negotiation is finished, the encoder is initialized, as shown by 803. Then following the receiving of the audio input as shown by 805 the encoding process is started as shown by 807. The encoded bitstream is then output as shown by 809 to the decoder / receiver 820. The encoded bitstream is transmitted to the receiver device, which uses the encoded bitstream to initialize an IVAS decoder as shown by 813. After the decoder initialization, the decoder 820 can as shown by 815 decoder or otherwise process the incoming IVAS encoded bitstreams and as shown by 817 and produce a decoder output. As presented by the time flow in Fig.8 where the decoder is initialized as shown by 813 only after a first bitstream is received from the sender may result in an sub-optimal initialization because there is no prior knowledge of the incoming bitstream format before decoding the first IVAS encoded bitstream, as explained previously. However as discussed in the embodiments above if the decoder / receiver has prior knowledge of the coded format of the incoming IVAS bitstream, for example from initialization or configuration data, the decoder can, as shown by 811, be initialized before receiving a first IVAS bitstream. This is presented as initialization a) (as shown by 811), rather than initialization b) (as shown by 813) in Fig.8. For example after the negotiation phase, the initialization or configuration data comprising agreed initial coded format for the session can be indicated to the receiver, for example, through RTP transmission. This allows the receiver to start the decoder initialization immediately and even before receiving any IVAS bitstream. This reduces the peak complexity at the beginning of the decoding. Additionally, some processing time is saved, because the decoder is ready to decode a received IVAS bitstream as soon as it is received, whereas in the known implementations the initialization step is only started when the first IVAS bitstream is received. The light-weight pre-decoder and multiple pre-initialized decoder embodiments can still require the IVAS bitstream to be received before the decoder can be initialized. However, these embodiments avoid the need for re-initializing the decoder, if the original output format was chosen sub-optimally when considering the received coded format for the bitstream. With respect to Fig.9a and Fig.9b are shown flow diagrams of example operations describing the RTP (or pre-IVAS input format messaging) examples described above. For example with respect to Fig.9a is shown the sender side operations. Therefore, as shown by 901, is an operation of determine an IVAS initialization information, for example a coded format to be used at the start of the session (e.g., through SDP negotiation). In some embodiments this information could comprise any of: information identifying a coded sub-format of the immersive audio bitstream; information identifying an immersive audio payload format of the immersive audio bitstream; information identifying an immersive audio payload sub-format of the immersive audio bitstream; information identifying an encoder input sampling rate of the immersive audio bitstream; and information identifying a number of channels with respect to an at least one coded format of the immersive audio bitstream. Then, as shown by 903, is packetize the coded format information or indication into an IVAS RTP packet (without audio data). In some embodiments any suitable manner can be employed for including the information relating to the immersive audio bitstream in at least one transport protocol payload such that the initialization information is transmitted separately from the immersive audio bitstream. Furthermore, as shown by 905, is transmit the IVAS RTP packet comprising the information to a receiver. For example with respect to Fig.9b is shown the complimentary receiver side operations. Therefore, as shown by 911, is an operation of receive and de-packetize (or otherwise extract) an IVAS RTP packet (or other example such as discussed above) with the initialization or configuration data, for example, comprising the coded format information. Then, as shown by 913, is initialize an IVAS decoder by taking into account the initialization or configuration data and in this example the coded format information (and / or the sampling rate or other information type as discussed above in some embodiments). With respect to Fig. 10a and Fig. 10b are shown flow diagrams of example operations describing the light-weight or hybrid decoder examples described above. For example with respect to Fig. 10a is shown the sender side operations. Therefore, as shown by 1001, is an operation of obtaining audio data and encoding it into an IVAS bitstream. Then, as shown by 1003, is the operation of packetize the IVAS bitstream into an IVAS RTP packet. Furthermore, as shown by 1005, is transmit the RTP packet to a receiver. For example with respect to Fig. 10b is shown the complimentary receiver side operations. Therefore, as shown by 1011, is an operation of receive and de-packetize an IVAS RTP packet. Then, as shown by 1013, is an operation to decode the received IVAS bitstream with a light-weight (or hybrid) decoder to extract the initialization or configuration data comprising the coded format information (and / or sampling rate or other information type) from the bitstream. Following on, as shown by 1015, is an operation of initialize an IVAS decoder by taking into account the initialization or configuration data comprising the coded format information. With respect to Fig.11a and Fig.11 b are shown flow diagrams of example operations describing the multi-decoder selection examples described above. For example with respect to Fig. 11 a is shown the sender side operations. Therefore, as shown by 1101, is an operation of obtaining audio data and encoding it into an IVAS bitstream. Then, as shown by 1103, is the operation of packetize the IVAS bitstream into an IVAS RTP packet. Furthermore, as shown by 1105, is the operation of transmitting the RTP packet to a receiver. With respect to Fig.11b is shown the complimentary receiver side operations. Therefore, as shown by 1111, is an operation of pre-initializing multiple IVAS decoders with different output format and / or output sampling rate configurations. Then, as shown by 1113, is an operation of receive and de-packetize an IVAS RTP packet. Additionally, as shown by 1115, is an operation to feeding the received IVAS bitstream to the pre-initialized decoders. Following on, as shown by 1117, is an operation of determining the coded format and / or sampling rate of the received bitstream at the decoders. As described above this can be detecting an error in the decoding operations or the rendering of the output audio signals or determining a poor quality or erroneous output audio signal. Finally as shown by 1119, is an operation of selecting the most suitable decoder based on the determined coded format and / or sampling rate information (discard the other decoders) and finalize the initialization process of the selected decoder. With respect to Fig. 12 an example electronic device is shown. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 1900 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. In some embodiments the device 1900 comprises at least one processor or central processing unit 1907. The processor 1907 can be configured to execute various program codes such as the methods such as described herein. In some embodiments the device 1900 comprises a memory 1911. In some embodiments the at least one processor 1907 is coupled to the memory 1911. The memory 1911 can be any suitable storage means. In some embodiments the memory 1911 comprises a program code section for storing program codes implementable upon the processor 1907. Furthermore in some embodiments the memory 1911 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1907 whenever needed via the memory-processor coupling. In some embodiments the device 1900 comprises a user interface 1905. The user interface 1905 can be coupled in some embodiments to the processor 1907. In some embodiments the processor 1907 can control the operation of the user interface 1905 and receive inputs from the user interface 1905. In some embodiments the user interface 1905 can enable a user to input commands to the device 1900, for example via a keypad. In some embodiments the user interface 1905 can enable the user to obtain information from the device 1900. For example the user interface 1905 may comprise a display configured to display information from the device 1900 to the user. The user interface 1905 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1900 and further displaying information to the user of the device 1900. In some embodiments the device 1900 comprises an input / output port 1909. The input / output port 1909 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1907 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The transceiver input / output port 1909 may be configured to receive the signals and in some embodiments obtain the focus parameters as described herein. In some embodiments the device 1900 may be employed to generate a suitable audio signal using the processor 1907 executing suitable code. The input / output port 1909 may be coupled to any suitable audio output for example to a multichannel speaker system and / or headphones (which may be a headtracked or a non-tracked headphones) or similar. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. The foregoing description has provided by way of exemplary and nonlimiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims. 3GPP 3rd Generation Partnership Project CMR Codec mode request EVS Enhanced Voice Services ISM Independent Streams with Metadata (i.e., type of Object-Based Audio) IVAS Immersive Voice and Audio Services MASA Metadata-Assisted Spatial Audio MC Multichannel OMASA Object-based audio with MASA (combined input format) OSBA Object-based audio with SBA (combined input format) PI Processing information (audio) RTCP Real-Time Transport Control Protocol RTP Real-Time Transport Protocol RTP HE Real-Time Transport Protocol Header Extension SBA Scene-Based Audio SDP Session Description Protocol 5 ToC Table of Content
Claims
1. An apparatus for configuring an immersive audio decoder, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform:obtaining initialization information based on an established conversational immersive audio session; andincluding the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by the immersive audio decoder.
2. The apparatus as claimed in claim 1, caused to perform including the initialization information relating to the immersive audio frame in at least one transport protocol payload is caused to further perform transmitting the initialization information such that the initialization information is received by at least one further apparatus and employed to initialize the immersive audio decoder prior to decoding the immersive audio frame.
3. The apparatus as claimed in any of claims 1 or 2, wherein the initialization information relating to the immersive audio frame comprises at least one of: information identifying a coded format of the immersive audio frame; information identifying a coded sub-format of the immersive audio frame; information identifying an encoder input sampling rate of the immersive audio frame; andinformation identifying a number of channels with respect to an at least one coded format of the immersive audio frame.
4. The apparatus as claimed in any of claims 1 to 3, caused to perform including the initialization information in a transport protocol payload is caused to further perform generating at least one processing information frame comprising the initialization information.
5. The apparatus as claimed in any of claims 1 to 4, caused to perform including initialization information in a transport protocol payload is caused to further perform including the initialization information in at least one of:a real-time transport protocol payload;a real-time transport protocol header extension; and a real-time transport control protocol payload.
6. The apparatus as claimed in any of claims 1 to 5, wherein the initialization information relating to the immersive audio frame comprises an identifier for at least one of:a mono channel format;a stereo channel format;a multichannel format;a Metadata-Assisted Spatial Audio format;an Independent Streams with Metadata format;a Scene-Based Audio format;an Object-based audio with Metadata-Assisted Spatial Audio format;an Object-based audio with Scene-Based Audio format;a multichannel 5.1 sub-format;a multichannel 5.1.2 sub-format;a multichannel 5.1.4 sub-format;a multichannel 7.1 sub-format;a multichannel 7.1.4 sub-format;a Metadata-Assisted Spatial Audio one transport channel sub-format;a Metadata-Assisted Spatial Audio two transport channels sub-format;an Independent Streams with Metadata single stream sub-format;an Independent Streams with Metadata two streams sub-format;an Independent Streams with Metadata three streams sub-format;an Independent Streams with Metadata four streams sub-format;a Scene-Based Audio First order Ambisonics sub-format;a Scene-Based Audio Second, higher, order Ambisonics sub-format;a Scene-Based Audio Third, higher, order Ambisonics sub-format;a Scene-Based Audio planar First order Ambisonics sub-format;a Scene-Based Audio planar Second, higher, order Ambisonics sub-format;a Scene-Based Audio planar Third, higher, order Ambisonics sub-format;an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format;an Object-based audio with Metadata-Assisted Spatial Audio two independent streams sub-format;an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format;an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format;an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; andan Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-format.
7. The apparatus as claimed in any of claims 1 to 6, further caused to perform negotiating with the further apparatus support for the initialization information.
8. The apparatus as claimed in claim 7, caused to perform negotiating with the further apparatus is further caused to perform negotiating employing a session description file.
9. The apparatus as claimed in claim 3 or any of claims 4 to 8 when dependent on claim 3, further caused to perform:encoding at least one immersive audio signal to an immersive audio frame based on the at least one coded format and / or sub-format;including the immersive audio frame in at the at least one transport protocol payload or in a further at least one transport protocol payload.
10. The apparatus as claimed in any of claims 1 to 9, caused to perform obtaining initialization information relating to at least one immersive audio frame is caused to perform obtaining initialization information relating to at least one immersive audio frame for at least one of:during a session setup; and after a session setup.
11. An apparatus for initializing an immersive audio decoder, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform:obtaining initialization information from an immersive audio frame or prior to decoding the immersive audio frame of an established conversational immersive audio session; andinitializing the immersive audio decoder based on at least part of the obtained initialization information.
12. The apparatus as claimed in claim 11, further caused to perform: obtaining the at least one immersive audio frame; anddecoding the at least one immersive audio frame at the initialized immersive audio decoder.
13. The apparatus as claimed in any of claims 11 to 12, caused to perform obtaining initialization information from the immersive audio frame or prior to the decoding the immersive audio frame of an established conversational immersiveaudio session is further caused to perform, obtaining the initialization information from an earlier part of an at least one transport protocol payload with respect to a later part of the at least one transport protocol comprising the immersive audio frame.
14. The apparatus as claimed in any of claims 11 to 13, caused to perform obtaining initialization information from an immersive audio frame or prior to the decoding the immersive audio frame of an established conversational immersive audio session is further caused to perform obtaining the initialization information in a separate earlier at least one transport protocol payload with respect to a later at least one transport protocol payload comprising the at least one immersive audio frame.
15. The apparatus as claimed in any of claims 11 to 14, wherein the initializationinformation relating to the immersive audio frame comprises at least one of: information identifying a coded format;information identifying a coded sub-format;information identifying an encoder input sampling rate; andinformation identifying a number of channels with respect to the at least one coded format.
16. The apparatus as claimed in any of claims 11 to 15, caused to perform obtaining initialization information from an immersive audio frame or prior to the decoding the immersive audio frame of an established conversational immersive audio session is further caused to perform obtaining the initialization information from at least one of:a real-time transport protocol payload;a real-time transport protocol header extension; and a real-time transport control protocol payload.
17. The apparatus as claimed in any of claims 11 to 16, caused to perform obtaining initialization information from an immersive audio frame or prior todecoding the immersive audio frame of an established conversational immersive audio session is further caused to perform:decoding at the immersive audio decoder or at a further decoder at least part of the at least one immersive audio frame to obtain the initialization information; andinitializing the immersive audio decoder based at least on the initialization information.
18. The apparatus as claimed in claim 17, wherein the further decoder is a lightweight decoder configured to obtain the initialization information.
19. The apparatus as claimed in any of claims 11 to 18, is further caused to perform:pre-initializing at least two immersive audio decoders, each of the at least two immersive audio decoders pre-initialized to decode the at least one immersive audio frame based on a different one of at least one output format and / or different output sampling rate,wherein the apparatus caused to perform obtaining initialization information from an immersive audio frame or prior to the decoding the immersive audio frame of an established conversational immersive audio session is caused to perform:decoding at each of the pre-initialized at least two immersive audio decoders the at least one immersive audio frame to generate at least two decoder outputs;determining the initialization information relating to the immersive audio frame based on the at least two decoder outputs, andwherein the apparatus caused to perform initializing the immersive audio decoder based on at least part of the obtained initialization information is further caused to perform:selecting one of the at least two decoders based on the initialization information; andfinalizing an initialization of the selected one of the at least two decoders.
20. The apparatus as claimed in claim 19, caused to perform determining the initialization information relating to the immersive audio frame based on the at least two decoder outputs is further caused to perform:analysing the output of each of the at least two decoders; anddetermining which of the at least two decoders to select based on the analysis.
21. The apparatus as claimed in claim 20, caused to perform analysing the output of each of the at least two decoders is caused to further perform detecting at least one of:an error output; anda lower quality output measured against one or more of the other outputs or a threshold value.
22. The apparatus as claimed in any of claims 11 to 20, wherein the initialization information relating to the immersive audio frame comprises an identifier for at least one of:a mono channel format;a stereo channel format;a multichannel format;a Metadata-Assisted Spatial Audio format;an Independent Streams with Metadata format;a Scene-Based Audio format;an Object-based audio with Metadata-Assisted Spatial Audio format; an Object-based audio with Scene-Based Audio formata multichannel 5.1 sub-format;a multichannel 5.1.2 sub-format;a multichannel 5.1.4 sub-format;a multichannel 7.1 sub-format;a multichannel 7.1.4 sub-format;a Metadata-Assisted Spatial Audio one transport channel sub-format;a Metadata-Assisted Spatial Audio two transport channels sub-format;an Independent Streams with Metadata single stream sub-format;an Independent Streams with Metadata two streams sub-format;an Independent Streams with Metadata three streams sub-format;an Independent Streams with Metadata four streams sub-format;a Scene-Based Audio First order Ambisonics sub-format;a Scene-Based Audio Second, higher, order Ambisonics sub-format;a Scene-Based Audio Third, higher, order Ambisonics sub-format;a Scene-Based Audio planar First order Ambisonics sub-format;a Scene-Based Audio planar Second, higher, order Ambisonics sub-format;a Scene-Based Audio planar Third, higher, order Ambisonics sub-format;an Object-based audio with Metadata-Assisted Spatial Audio single independent stream sub-format;an Object-based audio with Metadata-Assisted Spatial Audio two independent streams sub-format;an Object-based audio with Metadata-Assisted Spatial Audio three independent streams sub-format;an Object-based audio with Metadata-Assisted Spatial Audio four independent streams sub-format;an Object-based audio with Scene-Based Audio single independent stream and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and planar First order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and planar second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and planar second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio four independent streams and planar second, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio single independent stream and planar third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio two independent streams and planar third, higher, order Ambisonics sub-format;an Object-based audio with Scene-Based Audio three independent streams and planar third, higher, order Ambisonics sub-format; andan Object-based audio with Scene-Based Audio four independent streams and planar third, higher, order Ambisonics sub-format.
23. The apparatus as claimed in any of claims 11 to 22, further caused to perform negotiating with a further apparatus support for the initialization information.
24. The apparatus as claimed in claim 23, caused to perform negotiating with the further apparatus is further caused to perform negotiating employing a session description file.
25. A method for an apparatus for configuring an immersive audio decoder, the method comprising:obtaining initialization information based on an established conversational immersive audio session; andincluding the initialization information relating to an immersive audio frame in at least one transport protocol payload such that the initialization information is transmitted prior to decoding the immersive audio frame by an immersive audio decoder.
26. A method for an apparatus for initializing an immersive audio decoder, the method comprising:obtaining initialization information from an immersive audio frame or prior to decoding the immersive audio frame of an established conversational immersive audio session; andinitializing the immersive audio decoder based on at least part of the obtained initialization information.
Citation Information
Patent Citations
Apparatus and methods
GB2633769A
A method and apparatus for negotiation of conversational immersive audio session
WO2024179766A1