Representing descriptive metadata in spatial audio

WO2026162311A1PCT designated stage Publication Date: 2026-08-06NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2026-01-16
Publication Date
2026-08-06

Smart Images

  • Figure EP2026051003_06082026_PF_FP_ABST
    Figure EP2026051003_06082026_PF_FP_ABST
Patent Text Reader

Abstract

There is disclosed an apparatus configured to receive a transport protocol payload comprising at least one metadata parameter and a parameter type of the at least one metadata parameter; determine, from the parameter type of the at least one metadata parameter, that the at least one metadata parameter belongs to a group of metadata parameters, wherein the group of metadata parameters is related to an encoded immersive audio frame; and collate the at least one metadata parameter with other metadata parameters of the group of metadata parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Representing Descriptive Metadata in Spatial Audio

[0002] Field

[0003] The present application relates to apparatus and methods for implementing the distribution of descriptive metadata into packets, in particular for implementing the distribution of descriptive metadata information into an immersive audio real-time transport protocol (RTP) payload.

[0004] Background

[0005] Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the immersive voice and audio services (IVAS) codec (3GPP TS 26.253) which is designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing. This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources as audio objects. The codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.

[0006] The input signals are presented to the IVAS encoder in one of the supported formats (and in some allowed combinations of the formats). Similarly, it is expected that the decoder can output the audio in supported formats.Additionally, RTP (Real-Time Transport Protocol) is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data transfer to multiple destinations through IP multicast or to a specific destination through IP unicast. The majority of the RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols may also be utilized. RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol (RTSP).

[0007] The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transfer of multimedia data, and its companion protocol (RTCP) is used to periodically send control information and QoS (Quality of Service) parameters.

[0008] RTP sessions are typically initiated between client and server or between client and another client (or a multi-party topology) using a signalling protocol, such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols typically use the Session Description Protocol (SDP), such as defined by RFC 8866 to specify parameters for the sessions.

[0009] Summary

[0010] According to a first aspect there is an apparatus configured to: separate at least one metadata parameter from a group of metadata parameters, wherein the group of metadata parameters are related to an encoded immersive audio frame; include a parameter type of the at least one metadata parameter in a transport protocol payload; and include the at least one metadata parameter in the transport protocol payload.

[0011] The transport protocol payload may contain the encoded immersive audio frame.The transport protocol payload may comprise Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, wherein the PI header comprises a PI type field associated with the PI data frame and where the apparatus configured to include a parameter type of the at least one metadata parameter in a transport protocol payload may be configured to: include the parameter type of the at least one metadata parameter in the PI type field of the PI header; and where the apparatus configured to include the at least one metadata parameter in the transport protocol payload may be configured to: include the at least one metadata in the PI data frame.

[0012] The PI type field of the PI header may indicate one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

[0013] The group of metadata parameters may be a MASA descriptive metadata.

[0014] The transport protocol payload may comprise one of: a real-time transport protocol payload; a real-time transport protocol payload header extension; a real-time transport control protocol payload; or a data channel transport protocol payload.

[0015] The apparatus may be further configured to include the parameter type of the at least one metadata parameter in a pi-types parameter of the session description protocol (SDP).

[0016] The encoded immersive audio frame may be an encoded IVAS frame.

[0017] According to a second aspect there is an apparatus configured to: receive a transport protocol payload comprising at least one metadata parameter and a parameter type of the at least one metadata parameter; determine, from the parameter type of the at least one metadata parameter, that the at least one metadata parameter belongs to a group of metadata parameters, wherein thegroup of metadata parameters is related to an encoded immersive audio frame; and collate the at least one metadata parameter with other metadata parameters of the group of metadata parameters.

[0018] Where each metadata parameter of the group of metadata parameters may have a unique parameter type.

[0019] The apparatus may be further configured to: determine a further metadata parameter of the group of metadata parameters, based on the parameter type of the at least one metadata parameter, wherein a parameter type of the further metadata parameter is different from the parameter type of the at least one metadata parameter; and collate the further metadata parameter with the at least one metadata parameter and the other metadata parameters of the group of metadata parameters.

[0020] Where the parameter type of the further metadata parameter may be a source format.

[0021] The apparatus may be further configured to: set at least one of the other metadata parameters to a predetermined value based on the parameter type of the at least one metadata parameter.

[0022] Where the transport protocol payload may comprise Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, wherein the PI header comprises a PI type field associated with the PI data frame, wherein the at least one metadata parameter is contained in the PI data frame, and wherein the parameter type of the at least one metadata parameter is contained in the PI type field of the PI header.Where the PI type field of the PI header may indicate one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

[0023] Where the group of metadata parameters may be MASA descriptive metadata.

[0024] Wherein the transport protocol payload may comprise one of: a real-time transport protocol payload; a real-time transport protocol payload header extension; a realtime transport control protocol payload; or a data channel transport protocol payload.

[0025] Wherein the encoded immersive audio frame may be an encoded IVAS frame.

[0026] According to a third aspect there is a method comprising: separating at least one metadata parameter from a group of metadata parameters, wherein the group of metadata parameters are related to an encoded immersive audio frame; including a parameter type of the at least one metadata parameter in a transport protocol payload; and including the at least one metadata parameter in the transport protocol payload.

[0027] Where the transport protocol payload may contain the encoded immersive audio frame.

[0028] Where the transport protocol payload comprises Processing Information (PI) data may comprise a PI header and a PI data frame corresponding to the PI header, where the PI header may comprise a PI type field associated with the PI data frame and where the method comprising including a parameter type of the at least one metadata parameter in a transport protocol payload may comprise: including the parameter type of the at least one metadata parameter in the PI type field of the PI header; and where the method comprising including the at least one metadataparameter in the transport protocol payload may comprise: including the at least one metadata in the PI data frame.

[0029] Where the PI type field of the PI header may indicate one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

[0030] Where the group of metadata parameters may be MASA descriptive metadata.

[0031] Where the transport protocol payload may comprise one of: a real-time transport protocol payload; a real-time transport protocol payload header extension; a realtime transport control protocol payload; or a data channel transport protocol payload.

[0032] The method may further comprise including the parameter type of the at least one metadata parameter in a pi-types parameter of the session description protocol (SDP).

[0033] Where the encoded immersive audio frame may be an encoded IVAS frame.

[0034] According to a fourth aspect there is a method comprising: receiving a transport protocol payload comprising at least one metadata parameter and a parameter type of the at least one metadata parameter; determining, from the parameter type of the at least one metadata parameter, that the at least one metadata parameter belongs to a group of metadata parameters, wherein the group of metadata parameters is related to an encoded immersive audio frame; and collating the at least one metadata parameter with other metadata parameters of the group of metadata parameters.

[0035] Where each metadata parameter of the group of metadata parameters may have a unique parameter type.The method may further comprise: determining a further metadata parameter of the group of metadata parameters, based on the parameter type of the at least one metadata parameter, where a parameter type of the further metadata parameter may be different from the parameter type of the at least one metadata parameter; and collating the further metadata parameter with the at least one metadata parameter and the other metadata parameters of the group of metadata parameters.

[0036] Where the parameter type of the further metadata parameter maybe a source format.

[0037] The method may further comprise setting at least one of the other metadata parameters to a predetermined value based on the parameter type of the at least one metadata parameter.

[0038] Wherein the transport protocol payload may comprise Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, where the PI header may comprise a PI type field associated with the PI data frame, where the at least one metadata parameter may be contained in the PI data frame, and where the parameter type of the at least one metadata parameter may be contained in the PI type field of the PI header.

[0039] Where the PI type field of the PI header may indicate one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

[0040] Where the group of metadata parameters may be MASA descriptive metadata.

[0041] Where the transport protocol payload may comprise one of: a real-time transport protocol payload; a real-time transport protocol payload header extension; a real-time transport control protocol payload; or a data channel transport protocol payload.

[0042] Where the encoded immersive audio frame may be an encoded IVAS frame.

[0043] According to a fifth aspect there is an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: separate at least one metadata parameter from a group of metadata parameters, wherein the group of metadata parameters are related to an encoded immersive audio frame; include a parameter type of the at least one metadata parameter in a transport protocol payload; and include the at least one metadata parameter in the transport protocol payload.

[0044] According to a sixth aspect there is an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive a transport protocol payload comprising at least one metadata parameter and a parameter type of the at least one metadata parameter; determine, from the parameter type of the at least one metadata parameter, that the at least one metadata parameter belongs to a group of metadata parameters, wherein the group of metadata parameters is related to an encoded immersive audio frame; and collate the at least one metadata parameter with other metadata parameters of the group of metadata parameters.

[0045] An apparatus comprising means for performing the actions of the method as described above.

[0046] An apparatus configured to perform the actions of the method as described above.

[0047] A computer program comprising program instructions for causing a computer to perform the method as described above.A computer program product stored on a medium may cause an apparatus to perform the method as described herein.

[0048] An electronic device may comprise apparatus as described herein.

[0049] A chipset may comprise apparatus as described herein.

[0050] Embodiments of the present application aim to address problems associated with the state of the art.

[0051] Summary of the Figures

[0052] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:

[0053] Figure 1 shows schematically example server and peer-to-peer teleconferencing systems within which embodiments may be implemented;

[0054] Figure 2 shows schematically an example RTP packet structure for IVAS;

[0055] Figure 3 shows schematically an example audio processing information (PI) data section for IVAS;

[0056] Figure 4 shows schematically an example PI header for IVAS;

[0057] Figure 5 depicts a scenario of transmitting descriptive metadata;

[0058] Figure 6 depicts a scenario of receiving descriptive metadata;Figure 7 depicts a scenario of determining different source formats from the various combinations of received descriptive metadata parameters;

[0059] Figure 8 shows schematically example PI data frames for various single descriptive metadata parameters;

[0060] Figure 9 shows schematically example PI data frames for various combined descriptive metadata parameters;

[0061] Figure 10 shows schematically an example encoder configuration as shown in Figure 1 according to some embodiments;

[0062] Figure 11 shows a flow diagram showing a method of operation for the encoder as shown in Figure 10 according to some embodiments;

[0063] Figure 12 shows schematically an example decoder configuration as shown in Figure 1 according to some embodiments;

[0064] Figure 13 shows a flow diagram showing a method of operation for the decoder as shown in Figure 12 according to some embodiments; and

[0065] Figure 14 shows an example device suitable for implementing the apparatus shown.

[0066] Embodiments of the Application

[0067] The following describes in further detail suitable apparatus and possible mechanisms for the provision of efficient IVAS audio.

[0068] The IVAS codec is configured to obtain or receive a multi audio signal such as the signals received from a microphone array, or a multi-channel (MC) audio format input such as channel layouts 5.1, 7.1, 5.1+2 etc. Additionally, the IVAS codec canalso receive other forms of input such as a Scene Based Audio (SBA) ambisonics format input and one or more audio object signals. In the following description an audio object can be referred to as a an independent stream with metadata - ISM format. Furthermore, in some situations the codec is configured to handle more than one input format at a time. This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats. An example of two different audio input formats is the combination of the Metadata-Assisted Spatial Audio (MASA) format with audio object format. MASA is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.

[0069] The MASA format for IVAS uses audio (transport) signal(s) together with corresponding spatial metadata. The spatial metadata comprises parameters which define the spatial aspects of the audio signals and which may contain for example, directions and direct-to-total energy ratios in frequency bands. The MASA stream can, for example, be obtained by capturing spatial audio with microphones of a suitable capture device. The MASA stream can be obtained also from other sources, such as specific spatial audio microphones (such as Ambisonics or arraymicrophones), studio mixes (for example, a 5.1 audio channel mix) or other content by means of a suitable format conversion. An audio signal input to an immersive voice codec (such as IVAS) can be simultaneously encoded as 1 - N audio signals to give a transport audio stream and analysed to give a MASA metadata stream.

[0070] An example system within which embodiments may be implemented is shown in Figure 1.

[0071] Figure 1 , for example, shows an example teleconferencing system within which some embodiments can be implemented. In this example there is shown two sites or rooms, Room A 100 and Room B 102. Room A 100 comprises several ‘talkers’ or users, in this instance there is shown four ‘talkers’ Talker 1 103, Talker 2 112, Talker 3 113 and Talker 4 114. Room B 102 comprises one ‘talker’ or user, Talker RX 141.In the following example within room A is a suitable teleconference apparatus (or more generally telecommunications apparatus 110) configured to spatially capture and encode the audio environment and furthermore is configured to render a spatial audio signal to the room. Within each of the other rooms may be a suitable teleconference apparatus (or more generally telecommunications apparatus such as apparatus 120 within room B) configured to render a spatial audio signal to the room and furthermore is configured to capture and encode at least a mono audio and optionally configured to spatially capture and encode the audio environment.

[0072] In the following examples each room is provided with the means to spatially capture, encode spatial audio signals, receive spatial audio signals and render these to a suitable listener. It would be understood that there may be other embodiments where the system comprises some apparatus configured to only capture and encode audio signals (in other words the apparatus is a ‘transmit’ only apparatus), and other apparatus configured to only receive and render audio signals (in other words the apparatus is a ‘receive’ only apparatus). In such embodiments the system within which embodiments may be implemented may comprise apparatus with varying abilities to capture / render audio signals.

[0073] The teleconference apparatus (for each site or room) 110, 120 can be configured to call into a teleconference controlled by and implemented over a server 111.

[0074] In some embodiments the communications or teleconferencing system comprises a (peer-to-peer) communications system (rather than the server-based system shown in Figure 1) within which some embodiments can be implemented. Thus, for example, two or more UEs can be configured to interact directly with each other (for example to implement an immersive audio phone call between users). In such a scenario one of the UEs can be configured to deliver spatial ambience as one stream and employ a close-up microphone (for example a Lavalier microphone) to capture the speech as an audio object or audio source. The sender UE can beconfigured to encode the spatial ambience audio signals in a MASA format stream and the close-up microphone audio signal as an object format stream (or ISM stream). The two encoded audio streams can then be delivered as separated IVAS encoded audio streams, rather than a combined IVAS format stream of MASA format with audio object (ISM format). The sender UE, in addition, can be configured to encode processing information (PI) during the encoding and deliver the PI frames together with the IVAS frames to the receiver UE.

[0075] This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats. An example of two different audio input formats is the combination of the MASA format with audio object format.

[0076] The teleconference apparatus can be configured to spatially capture and encode the audio environment and furthermore can be configured to render a spatial audio signal to the room. In this example only the communications or signalling path from the Room A 100 to the Room B 102 is shown for simplicity but a duplex or multipoint communication system comprising multiple signalling paths can be implemented using the methods as described herein without significant inventive input.

[0077] The teleconference apparatus (for each site or room) 110, 120 is further configured to communicate with each other to implement a teleconference function.

[0078] As shown in Figure 1 , the apparatus 110, 120 and server 111 can comprise suitable encoder and decoder functionality. For example, the apparatus 110 is shown comprising an (IVAS) encoder 101, the server 111 is shown comprising a (IVAS) decoder and encoder 121 and the apparatus 120 is shown comprising an (IVAS) decoder 131. In such a manner audio signals capturing the spatial audio environment together with audio objects (an audio object being audio signals which can capture a specific talker, Figure 1 depicts speaker 1 talking which is captured as the audio object 140) can be encoded by the encoder 101 which generates a bitstream 106 to be passed to a server 111. The server 111 can then decode,(optionally then mix with other objects and otherwise process the audio signals) and encode then to generate the bitstream 108 to be passed to the apparatus 120. The apparatus 120 can then decode the audio signals and present them to the user or talker ‘Talker RX’ 141.

[0079] Although this example shows a teleconference application the encoder / decoder functionality can be applied to the streaming of any suitable media.

[0080] The IVAS decoder / renderer for each of the teleconference apparatus 120 can be furthermore configured to handle multiple input streams that may each originate from a different encoder.

[0081] As discussed previously RTP is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP is furthermore designed to carry a multitude of multimedia formats, which permit the transport of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may therefore require a profile and payload format specifications.

[0082] The profile is configured to define the codec used to encode the payload data and the mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header. For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP).An RTP session can be established for each multimedia stream. Audio and video streams may be implemented which use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification can furthermore be configured to recommend port numbers for RTP, and furthermore to recommend the use of the next odd port number for the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.

[0083] Each RTP stream can comprise RTP packets, and the RTP packet in turn can comprise a RTP header and payload pair.

[0084] One proposal for an IVAS RTP packet structure is described in the 3GPP document S4-241325 and is shown in Figure 2. The 3GPP document proposes that the IVAS RTP packet structure 201 will comprise an RTP Header 203 (with a possible header extension HDREXT) together with an IVAS payload 205. The format of the RTP header will conform to the conventional norms of RTP design. Figure 2 depicts the proposed IVAS payload structure according to S4-241325 as comprising a payload header 207, encoded IVAS frame data 209 and (Processing Information) (PI) data 211.

[0085] For the payload header it is proposed that the header can comprise a Table of Content (ToC) byte which accompanies the IVAS frame data and is used to convey information relating to the size and bitrate of the corresponding IVAS frame. Consequently, when an RTP payload comprises more than one encoded IVAS frame it will also comprise more than one ToC bytes, with each ToC byte being associated with a particular encoded IVAS frame. In addition, the ToC byte may also comprise information relating to whether the IVAS encoder is operating in an EVS mode of operation. Additionally, it is suggested in the 3GPP document S4-241325 document that the IVAS payload header can contain extra bytes (referred to as E-bytes for want of a better term) in order to assist in the conveyance of further IVAS related information to the decoder. For instance, one proposal suggests that the E-bytes can house the Codec Mode Request (CMR) which is used to request a change to the bitrate to the IVAS stream (and other suggested coding parameters such as coding format and bandwidth) at the decoder. Furthermore, the E-bytes field of the IVAS payload header can also be used to explicitly signal the presence of the PI data section at the end of the IVAS payload. The PI data can be sent as separate payload frames, in addition to the IVAS frames as part of the IVAS RTP packet stream. The PI data comprise PI frame data which can be identified from the coded IVAS frames with the help of the E-byte signalling and a PI header located prior to or in front of the PI frame data.

[0086] The concept of processing information (PI) data was introduced in GB GB2314333.2, where a thorough description can be found. In short, PI data are supplementary information frames that can carry information to supplement the rendering or consumption of the IVAS audio bit stream at the receiver. These frames can be added to RTP packets carrying the IVAS streams. The presence of PI data in an RTP packet can be signalled using the ToC byte methodology, where each PI data has an associated ToC byte in the RTP packet, or as discussed above using the E-bytes filed. A specific bit code within the E-bytes can be used to indicate PI data within the IVAS payload. Furthermore, the concepts discussed in GB2314333.2 allow for different types of PI data.

[0087] The presence of PI data in an RTP packet can be signalled using the E-byte in the RTP header which indicates there is PI data at the end of the payload. PI headers are then used to identify PI data frames within the IVAS payload, where a PI header is associated with a PI data frame and contains information identifying the type and size of the PI data frame. The PI header also contains information which associates a specific IVAS encoded frame (of the frame data section 209 of the IVAS payload 205) to the PI data frame. Therefore, in IVAS payloads which comprise multiple IVAS encoded frames, there can be multiple PI headers, where each PI header is coupled with their respective PI data frame to a respective IVAS encoded frame.Figure 3 shows a possible structure for the PI data section of the IVAS payload. For example, the PI data section 301 is one in which comprises PI header section 303 followed by a PI frame data section 305. The PI frame data section 305 comprises the PI data frames for the IVAS payload, and the PI header section comprises the PI headers in which each PI header corresponds to PI frame. To note, each PI data frame can also comprise its own header within the frame data. A PI header can have the structure shown in Figure 4. For example, an example PI header 400 is one in which the header comprises a PF bit 401 at the beginning of the frame, which is followed by a PM frame marker field 402 which is then followed by the PI type field 403 and finally followed by the PI size field 404. The PF bit 401 indicates if another PI header follows by the use of the binary indicator (1 ) or (0). PM frame marker field 402 is configured to indicate the IVAS encoded audio frame associated with the corresponding PI frame. The PI frame type identifier 403 can be configured to indicate the type of data carried by the PI frame associated with the PI header. The type could be, for example, orientation data, priority of audio objects in ISM mode, scene information, or any suitable non-audio information. The PI size field 404 can be configured to indicate the size of the PI frame associated with the PI header.

[0088] Annex A of the IVAS standard document 3GPP TS 26.258 specifies MASA descriptive metadata which can accompany but is separate to the encoded IVAS frame. The MASA descriptive metadata are parameters which describe the audio environment from which the IVAS encoded MASA frame (i.e. MASA metadata and audio transport signal(s)) is captured. They can be included into the RTP IVAS payload in the form of audio processing information (PI) as ancillary information to be used as ancillary information at the receiver to assist in the rendering of the audio signal.

[0089] The MASA descriptive metadata parameters can entail parameters such as source format which describes the format of the audio signals from which the MASA stream was captured and can comprise values which indicate whether the source format is microphone grid, channel-based, ambisonics or default / other in which case theaudio signals originate from unknown format(s) including mixed sources. Other examples of MASA descriptive metadata parameters include a parameter which gives the channel distance for cases when the audio scene is captured using the microphone grid or when the source format is default / other. A further example of a descriptive metadata parameter is the number of channels parameter which is used to indicate the number of audio transport channels in the IVAS encoded stream and in the captured MASA stream. This information is given as a single bit where 0 indicates one audio transport channel and 1 to indicates two audio transport channels. A full list of MASA descriptive metadata can be found in annex A.4 of the 3GPP standard document TS 26.258.

[0090] Currently, 3GPP TS 26.258 Annex A.4 specifies that the MASA descriptive metadata parameter can comprise a 16-bit value, where the first four bits are reserved for encoding the number of channels, the number of directions and source format. The final 12 bits are assigned to represent a variable description whose values represent various further descriptive metadata parameters, the values of which are dependent on the first four bits. That is the specific combination of number of channels, number of directions and source format dictates the type of descriptive metadata held by the final 12 bits.

[0091] There already exists a mechanism for transmitting the above 16 bit MASA descriptive metadata parameter as a single PI data frame with an assigned PI type of MASA_DESCRPIPTIVE_META. According to the document, the 16 bit descriptive value is encoded as a PI data frame according to Table 1 below.

[0092]

[0093]

[0094] Table 1

[0095] However, the above mechanism for transmitting the MASA descriptive metadata has a drawback in that all parameters that make up the MASA descriptive metadata being transmitted together as one entity. There is currently no method of transmitting the descriptive metadata in a piecemeal fashion where individual parameters of the descriptive metadata are transmitted to the receiver. For instance, some decoding and rendering applications may only require a subset of the MASA descriptive metadata parameters, therefore in these situations it would be advantageous to only transmit the descriptive metadata parameters required at the receiver rather than transmitting all the MASA descriptive metadata parameters. Not only would this result in a saving of transmitted bits, the separation of the descriptive metadata into individual transmittable elements can also reduce the duplication of data. For instance, the current version of the IVAS standard TS 26.253 stipulates that the Number of Channels and Number of Directions MASA parameters are part of the encoded IVAS frame data. However, these two parameters are also included in the above MASA descriptive metadata. Clearly in this situation it would be advantageous to not to have to send the Number of Channels and the Number of Directions when transmitting MASA descriptive metadata as PI data.

[0096] As explained above the IVAS encoder is capable of handling a multitude of different input format signals in addition to the MASA format. These additional formats can comprise scene-based audio (SBA) (which is based on ambisonics), multichannel (MC), audio objects and a combination thereof. It has been noticed that it would be advantageous that if some of the formats (other than MASA) has access to particularparameters of the MASA descriptive metadata set whilst rendering the audio signal. To this end it would be beneficial to be able to only send the MASA descriptive metadata data parameters which are useful to the process of rendering of particular parameters rather than sending the whole set of MASA descriptive metadata parameters.

[0097] Embodiments aim to provide the above benefits.

[0098] One mechanism for sending parameters of MASA descriptive metadata can comprise sending some of the descriptive metadata parameters as part of the MASA format IVAS encoded frame and sending any remaining parameters of the MASA descriptive metadata as PI data. The remaining parameters of the MASA descriptive metadata can be sent as either individual PI data frames which will be discussed later on, or as one PI data frame. To this end, Figure 5 depicts such a scenario at the encoder in which the number of directions and number of channels (depicted as 510 in Figure 5) of the MASA descriptive metadata 501 are transmitted as part of the encoded IVAS frame 503 within the RTP IVAS payload 505. The other parameters of the MASA descriptive metadata (depicted as 507 in Figure 5) are transmitted in the RTP payload using PI data 509, which as discussed previously can be in the form of individual PI data frames or combined together as one PI data frame.

[0099] Figure 6, depicts a decoder receiving the IVAS payload 505. The number of directions and number of channels (depicted as 610 in Figure 6) are received as part of the encoded IVAS frame 503 within the received RTP IVAS payload 505. The other parameters of the MASA descriptive metadata (depicted as 607 in Figure 6) are received in the RTP payload as PI data 509. The parameters of MASA descriptive metadata 601 are shown in Figure 6 as being collated from the elements 607 of the received PI data 509 and the two elements 610 from the received encoded IVAS frame 503.Therefore, at the decoder the MASA descriptive metadata can be formed into a whole entity by collating the various parameters that constitute the MASA descriptive metadata, when the parameters are transmitted via different means from one another. Furthermore, some of the parameters that constitute the MASA descriptive metadata need not be transmitted, rather these parameters can be derived or inferred from other received metadata parameters .

[0100] In some situations, a specific parameter of the MASA descriptive metadata may be inferred by a particular combination of received other parameters of the MASA descriptive metadata, thereby negating the need to transmit information relating to the specific parameter. For example, one parameter of the MASA descriptive metadata which fits such a criterium is the Source Format. This parameter indicates the format of source signals into the IVAS encoder as a two bit value according to Table 2 below. However, the value of the Source Format can be inferred by the receipt of specific combinations of other parameters of the MASA descriptive metadata

[0101]

[0102] Table 2

[0103] The specific example of the Source Format is illustrated by way of Figures 5, 6 and 7. In Figure 5 the Source Format parameter 506 is depicted as not being transmitted and in Figure 6 the Source Format parameter 606 is shown as being determined from the following received parameters 607; Transport Definition, Channel Angle, Channel Distance and Channel Layout. Figure 7 depicts how the Source Formatparameter can be determined from various combinations of; Transport definition, Channel angle, Channel Distance and Channel Layout. For example, the Source Format can be determined to be a microphone grid, when the Transport definition, Channel angle and Channel distance parameters are received as part of the descriptive metadata set. This scenario is depicted as 701 in Figure 7. The Source Format can be determined as Ambisonics when Transport definition and Channel angle parameters are received as part of the descriptive metadata set, this scenario is depicted as 703 in Figure 7. The Source Format can be determined to be channelbased (i.e. , a multi-channel format), when the Channel layout element is present in the received descriptive metadata, this scenario is depicted as 705 in Figure 7. Finally, when none of the four parameters are received, the source format can be determined to be Unknown or Other Format, this scenario is depicted as 707 in Figure 7.

[0104] In general, the Unknown or Other category of format can be used as a default case in situations when the Source Format is indeterminable from the received MASA descriptive metadata parameters. For instance, in some situations, some of the MASA descriptive metadata parameters can be missing such that the receiver is unable to unambiguously determine the Source Format. For example, when the received IVAS RTP payload contains only the Transport Definition parameter, there are not enough parameters to enable the receiver to determine whether the Source Format is a Microphone grid or Ambisonics. On the other hand, the receiver may not be able to determine the Source Format as a result of the received metadata parameters. For example, when the received payload includes Transport Definition, Channel Angle and Channel Layout, the receiver is unable to distinguish (for the Source Format) between Ambisonics or Channel-based format. In these situations, the Source Format for the received MASA descriptive metadata can be set as “Other” or “Unknown” or some other similar value, that indicates that the source format is not known.In some embodiments, the Source Formats can be listed in an order or preference. For example, a microphone grid is more preferred than Ambisonics which is more preferred than channel-based format etc. Therefore, if the received IVAS payload contains metadata parameters which result in an indeterminable Source Format, the format can be determined based on the order of preference. For example, if as a result of processing the received payload the Source Format is indicated to be either a microphone grid or a channel-based format (i.e. , the payload contains Transport definition, Channel angle, Channel distance and Channel layout parameters), the receiver can determine the Source Format as a microphone grid as a result of its higher order of preference.

[0105] As discussed previously, a descriptive metadata parameter can be transmitted separately rather than being packaged together as set of descriptive metadata parameters and transmitted using a PI type for combined metadata parameters, such as described above for the set of MASA descriptive metadata parameters. To that end, each individual descriptive metadata parameter can be transmitted as a single entity via a PI data frame under a PI type which is specific to the individual descriptive metadata parameter.

[0106] In embodiments, some descriptive metadata parameters can be individually transmitted within an IVAS payload as an individual PI data frame containing the value of the parameter.

[0107] For example, the Transport Definition parameter whose purpose is to indicate the configuration of either the two transport channels within the IVAS codec or the two transport channels obtained as part of the audio capture process and is encoded as a 3-bit value according to Table 3 below. This value can be transmitted as a single element in a PI data frame. The PI data frame can be structured so that the first three bits hold the parameter value, and the following five bits are zero-padded in order to complete byte. This form of the Transport Definition parameter can take the PI type moniker which is descriptive of the parameter such asTRANSPORT-DEFINITION or TRANSPORT-CONFIGURATION. The structure of the PI data frame is shown as 801 in Figure 8.

[0108]

[0109] Table 3

[0110] In some embodiments the PI type associated with the Transport Definition parameter can be used to distinguish between the Source Formats of Microphone grid and Default / Other. This can be accomplished by having PI types which distinguish between Microphone grid and other source format such as TRANSPORT_DEFINITION_MICROPHONE_GRID and TRANSPORT_DEFINITION_OTHER_GRID.

[0111] When the Transport Definition parameter is used as part of a MASA input to the IVAS encoder, the parameter can be set to the same value as that would be used by the above MASA descriptive metadata when indicating the transport definition for the audio transport signals.

[0112] The parameter can also be applicable to other input formats such as a stereo format where a similar indication can be conveyed.

[0113] For Multichannel and audio object (ISM) input format the Transport Definition parameter can be used to signal information relevant to the capture system.A further example of an individual descriptive metadata parameter is the Channel Angle parameter whose purpose is to indicate the symmetric angle positions of the directivity patterns with respect to the audio transport signals. The Channel Angle parameter can take the angular values of ±90 deg, ±70 deg, ±55 deg, ±45 deg, ±30 deg, ±0 deg. Where, the zero angle indicates a frontal direction, the positive angles indicate directions on the right hand side and the negative angles indicate directions on the left hand side. However, it is also to be understood that the above angles may be determined according to other predetermined schemes.

[0114] In embodiments the Channel Angle parameter can be encoded as a 3-bit value according to Table 4 below and transmitted as a single element in a PI data frame as shown by 803 in Figure 8. As before, the PI data frame can be structured so that the first three bits hold the parameter value, and the following five bits are zero-padded in order to complete byte. This form of the Channel Angle parameter can take a PI type moniker which is descriptive of the parameter such as CHANNEL-ANGLE or OPENING-ANGLE.

[0115]

[0116] Table 4

[0117] In some embodiments the PI type associated with the Channel Angle parameter can also be used to distinguish between the Source Formats of Microphone grid and Default / Other. As before this can be accomplished by having PI types which distinguish between Microphone grid and other source format such as CHANNEL_ANGLE_MICROPHONE_GRID and CHANNEL_ANGLE_OTHER_GRID.When the Channel Angle parameter is used as part of a MASA input to the IVAS encoder, the parameter can be set to the same value as that would be used by the above MASA descriptive metadata when indicating the channel angle for the audio transport signals.

[0118] For a Multichannel input format, the Channel Angle parameter can be used to signal information relevant opening of the front left-right pair during capture.

[0119] Another example of an individual descriptive metadata parameter is the Channel Distance parameter which indicates the distance between two capturing points, that is the distance between two microphones in the audio scene. The parameter may encompass a range of values from Om, <0.01 m to over 1m with as many as 60 individual values between 0.01 m to 1m.

[0120] The Channel Distance parameter can be encoded as a 6-bit value according to Table 5 below and transmitted as a single element in a PI data frame as shown by 805 in Figure 8. The PI data frame for the Channel Distance parameter can be structured so that the first six bits hold the parameter value, and the following two bits are zero-padded in order to complete byte. This form of the Channel Distance parameter can take a PI type moniker such as CHANNEL_DISTANCE, MIC_DISTANCE, CHANNEL_SEPERATION, CHANNEL-SPACING, MIC_SPACING, or GRID_SPCING, that is a moniker which is descriptive of the Channel Distance parameter

[0121] <

[0122]

[0123] | 111 111 | > 1 m ~~ |

[0124] Table 5

[0125] When the Channel Distance parameter is used as part of MASA input to the IVAS encoder, the parameter can be set to the same value as that would be used by the above MASA descriptive metadata when indicating the channel distance associated with the audio transport signals.

[0126] In relation to a stereo input format the Channel Distance parameter can be used for a similar purpose as for the case of the equivalent MASA Channel Distance parameter.

[0127] With respect to audio object (ISM) input format, the Channel Distance parameter can be used to indicate the distance between individual audio object which could prove to be useful in helping in echo cancellation.

[0128] With respect to MC (multichannel) input format to the IVAS encoder, the Channel Distance parameter can be used to indicate the distances between capture microphones.

[0129] Furthermore, specific PI types such as CHANNEL_DISTANCE_PI_0_1 and CHANNEL_DISTANCE_PI_4_5 can be used indicate the distance between front left and right channels and the distance between the surround left and right channels respectively, of a 5.1 channel configuration. In these instances, the PI data frame for the PI type CHANNEL_DISTANCE_PI_0_1 can hold the distance between the front left and right channels, and the PI data frame for the PI type CHANNEL_DISTANCE_PI_4_5 can be used to hold the distance between the surround left and right channels.

[0130] Another example of an individual descriptive metadata parameter is the Channel Layout parameter which indicates the generic layout of the multichannel (MC) input to the IVAS encoder. The value of this parameter encodes the channel layouts as a3-bit value according to Table 6 below and is transmitted as a single element in a PI data frame as shown by 807 in Figure 8. As before, the PI data frame can be structured so that the first three bits hold the parameter value, and the following five bits are zero-padded in order to complete byte. The Channel Layout parameter can take a PI type moniker which is descriptive of the parameter such as CHANNEL-LAYOUT.

[0131]

[0132] Table 6

[0133] When the Channel Layout parameter is used as part of MASA input, the parameter can be set to the same value as that would be used by the above MASA descriptive metadata when indicating the Channel Layout associated with the audio transport signals.

[0134] As discussed above for a MC input format, the Channel Layout can be used to convey the multichannel layout of the encoded IVAS frame.

[0135] In some instances, a MC signal may be conveyed using ambisonics by creating a beam pattern for a series of azimuth angles from the MC signal and then determining an Ambisonic signal of a defined order for each beam pattern. The MC signal can be retrieved from the Ambisonics by forming beam patterns in the directions of the MC signal with the aid of the CHANNEL-LAYOUT parameter.

[0136] Another example of an individual descriptive metadata parameter is the Source Format parameter which indicates the format of the source signals from which theinput signal stream to the IVAS encoder is generated. The Source Format parameter is arranged to have a value which indicates one of; microphone grid, channelbased, Ambisonics or other.

[0137] In embodiments, the Source Format parameter can be encoded using 2 bits according to Table 7 below and transmitted as a single element in a PI data frame. The structure of the PI data frame is shown as 809 in Figure 8 where it can be seen that the first two bits hold the parameter value, and the following 6 bits are zero-padded in order to complete byte. The Source Format parameter can take a PI type moniker such as SOURCE_FORMAT.

[0138]

[0139] Table 7

[0140] When the Source Format parameter is used as part of MASA input, the parameter can be set to the same value as that would be used by the above MASA descriptive metadata when indicating the source for the audio transport signals.

[0141] In the context of SBA and MC input formats to the IVAS encoder, the Source Format parameter can be used to indicate the source that was used for the audio transport signals when the source was anything other than Ambisonics for SBA and multichannel for MC.

[0142] In further aspects of the invention, various combinations of the above descriptive metadata parameters can be combined into groups in accordance with the prescribed descriptive metadata parameters for a particular Source Format. For example, in the case of a Source Format of microphone grid, the prescribed descriptive metadata parameters can comprise: Transport Definition, ChannelAngle and Channel distance. These metadata parameters can be combined into a single PI data frame (as shown by 901 in Figure 9) and given a specific PI type such as MICROPHONE_GRID for example.

[0143] A similar approach can be used for the Ambisonics Format. In this case the combined group of prescribed descriptive metadata parameters can comprise Transport Definition and Channel Angle. These metadata parameters can be combined into a single PI data frame (as depicted in Figure 9 as 903) and given a specific PI type moniker of AMBISONICS_SOURCE_FORMAT.

[0144] The transmission of combined groups of descriptive metadata parameters associated with a particular source format can be advantageous. For instance, using the combined descriptive metadata parameters approach can result in less bits being transmitted than by sending the parameters as single PI data frames due to less zero padding of the parameters. Furthermore, the above combined approach inherently leads to the signalling of the Source Format.

[0145] Alternatively, it some aspects of the invention it may be preferable to combine the above combined PI types with single PI types. This scenario may be preferable when the descriptive metadata parameters associated with a particular Source Format are required to be updated at different intervals. Therefore, rather than sending a combined descriptive metadata parameter PI type when only one of the parameters need updating, it would be more preferable to just send a single PI type containing the updated descriptive metadata parameter.

[0146] With regards to the case of Binaural audio. Binaural audio can be signalled to the receiver by one of several ways such as the Transport definition parameter as discussed above; a separate binaural PI type to indicate that the audio is binaural, or using the E-bytes in the IVAS payload header.In some embodiments the parameters of the MASA descriptive metadata can be populated at the receiver in response to the receipt of a subset of “key” descriptive metadata parameters. In other words, it may not be necessary to transmit all the parameters of the MASA descriptive metadata in order to recreate the whole set of parameters at the receiver. For example, for the case of Binaural audio, the receipt of the Transport definition parameters is sufficient enough to generate the MASA descriptive metadata set, providing of course that the T ransport definition parameter indicates Binaural audio. For in this case, Source format can be set to “default / other”, Channel Angle can be set to “Unspecified” and the Channel Distance value can be set to “Unspecified” or to other preset values such as 20cm which is the approximate distance between human ears.

[0147] At the receiver, when the parameters of the MASA descriptive metadata are received via separate means such as by separate PI data frames, the MASA descriptive metadata can be assembled by collating the various parameters which make up the MASA descriptive metadata.

[0148] The (IVAS) encoder according to some embodiments, such as the example encoder 101, is shown in Figure 10.

[0149] The encoder is configured to obtain the audio processing information 1000, which in this context can comprise for example a descriptive metadata parameter or a combined group of descriptive metadata parameters. The encoder is also configured to obtain audio processing type information 1002, which can be the PI types of the descriptive metadata parameter or combined group of descriptive metadata parameters 1000. Additionally, the encoder comprises a RTP generator 1001 which is configured to receive the audio processing information 1000 and the audio processing type information 1002 and generate the RTP packets containing the PI payload.The RTP generator 1001, in some embodiments, comprises an audio processing information packetizer (where the PI header identifies the type of PI data frame) 1003.

[0150] The RTP generator 1001, in some embodiments, further comprises an PI type header appender 1005 configured to append the PI type information to the packet.

[0151] The RTP generator 1001, in some embodiments, further comprises an PI payload appender 1007. The PI payload appender 1007 is configured to append the PI payload to the packet prior to sending the generated packets.

[0152] Figure 11 shows the operation of the encoder 101 from the perspective of encoding a MASA format input.

[0153] The first operation is one of obtaining the data for the RTP payload which can comprise obtaining MASA metadata and accompanying audio transport signals. This can be accomplished by an audio analysis front end to the encoder which analyses the input audio signal to provide the aforementioned transport signals and MASA metadata. Additionally, the encoder is also configured to obtain the descriptive metadata parameters as discussed above. These operations are shown as steps 1101 and 1113 in Figure 11.

[0154] Next the encoder can be configured, depending on the decoding and rendering process at the receiver, to at least one of: separate at least some of the obtained descriptive metadata parameters into separate descriptive metadata parameters for transmission to the receiver; or group at least some of the obtained descriptive metadata parameters into a combined group of descriptive metadata parameters for transmission to the receiver. This operation is shown as step 1103 in Figure 11. The separated / grouped descriptive metadata parameters can be then packaged into a format required for a PI data frame as depicted by processing step 1105.Also shown in Figure 11 is the processing step 1107 of encoding the audio transport signals and MASA metadata into an IVAS frame.

[0155] Then for the RTP packet the IVAS payload can be generated as shown in in Figure 11 by step 1109.

[0156] The IVAS payload generation step of 1109 can for example comprise the operations of appending the PI header including the PI type to the PI data frame and then appending the encoded IVAS frame.

[0157] Then the RTP packet can be generated by packetizing the IVAS payload into the RTP packet. The RTP packet generation is shown as step 1111 in Figure 11.

[0158] When it comes to the negotiation between transmitter and receiver the existing pi-types parameter in the SDP protocol can be configured to provide information as to the above descriptive metadata PI types. The pi-types parameter is defined for IVAS in Annex A of TS 26.253 specification. This would allow for instance a negotiation phase prior to the sending of the IVAS RTP packets whereby it may be checked whether specific descriptive metadata PI types corresponding to particular metadata parameters are supported by the decoder. This then would allow the receiver to determine whether it is capable of supporting a specific metadata descriptive type for the duration of the session.

[0159] Additionally, in some embodiments the descriptive metadata parameter itself may be transmitted to the receiver as a value within the SDP protocol. For instance, the descriptive metadata parameter may be passed to the receiver as par to the SDP negotiation phase.

[0160] In some embodiments, the separate descriptive metadata elements can be transmitted as part of RTP Header Extension, via RTCP or via RTP Data Channel.The (IVAS) decoder according to some embodiments, such as the example decoder 131, is shown in Figure 12.

[0161] The decoder in some embodiment is configured to receive or otherwise obtain the RTP packets 1200. Additionally, the decoder comprises a RTP extractor 1201 which is configured to receive the RTP packets 1200 and extract from these RTP packets the PI payload and enable the decoder to render the audio signals in the RTP packet (IVAS) payload based on the metadata descriptive parameters.

[0162] The RTP extractor 1201, in some embodiments, comprises a payload type determiner 1203. The payload type determiner 1203 is configured to determine from the RTP packets the PI information type and thus be able to determine the PI information and apply the PI information based on the determined PI information type.

[0163] Figure 13 shows the operations of the example RTP extractor shown in Figure 12 for the case of extracting metadata descriptive parameters for a MASA encoded audio signal or other IVAS coded format.

[0164] The first operation is one of obtaining or receiving and then parsing RTP packets as shown in Figure 13 by step 1301.

[0165] Then the PI types in the IVAS payload is decoded which can then be used to parse the contents of the PI frames associated with the PI type. In embodiments there can be several different PI types present, each PI type being associated with either a separate descriptive metadata parameter or a combined descriptive metadata comprising more than one parameter, This is shown as processing step 1303 in Figure 13.

[0166] Then the PI frame(s) is / are decoded in view of the PI type(s) in order to extract the descriptive metadata parameters. This is shown by step 1305 in Figure 13. Inembodiments this step may also involve assembling the various received descriptive metadata parameter, whether they are received as individual parameters or as a combined group of parameters into a set of descriptive metadata parameters.

[0167] The step of decoding the received IVAS frame data can take the form of decoding the IVAS frame such that the output is in the MASA format of two (or one) transport audio channel signals with accompanying MASA metadata.

[0168] In terms of the architecture of the IVAS decoder the above mode of operation is called the EXT output format. Details of the EXT output format can be found in the patent application PCT / EP2024 / 077653. In summary, the EXT output format is a “pass through” mode where the format of the (spatial) audio output at the decoder / renderer is of the same form as the input to the IVAS encoder. For instance, if the input to the IVAS encoder is the MASA format of transport audio signal channels and MASA metadata. Then an IVAS codec operating in “pass through” mode would produce an output at the decoder / renderer which adheres to the same format as the audio input, which in this case would be an output in the MASA format.

[0169] In EXT mode the MASA format (output) frame would typically be accompanied with descriptive metadata comprising the number of directions and number of transport channels.

[0170] Step 1307 in Figure 13 depicts the decoding of the encoded IVAS frame data, and as explained above, the output can be a MASA format frame.

[0171] As shown by step 1309, the set of descriptive metadata parameters decoded in step 1305 can be combined with the descriptive metadata accompanying the MASA format frame (from step 1307) to produce the full descriptive metadata.

[0172] The full descriptive metadata can be used together with the MASA format frame to produce a rendered output as shown by step 13011 in Figure 13.Other examples of descriptive metadata parameters can be related to capturing the audio scene such as a Channel direction which indicates the direction in terms of an azimuth and elevation of a channel. A Channel direction parameter can be assigned a specific PI type with a moniker such as CHANNEL_DIRECTION or CHANNEL_DIRECTIONS. The PI data frame for the Channel direction can then have a single channel direction or multiple channel directions.

[0173] The use of a Channel Direction descriptive metadata parameter can be used in the conjunction with the previously discussed Channel Distance parameter in order to indicate the positions of audio objects when the IVAS codec is operating in an ISM format input mode.

[0174] The use of a channel direction descriptive metadata parameter can be useful in the scenario of a microphone grid with more than two microphones, where the channel direction element can be used to convey information about the positions of the microphone in the grid. For example, when considering a handheld device such as a mobile phone with a microphone positioned two or more comers of the device. A Channel Direction parameter can provide directional information for a corner relative to the centre of the device. Furthermore, when the Channel Direction parameter is combined with the Channel Distance parameter the resulting combination of parameters can define microphone grid for the device.

[0175] Another example of descriptive metadata parameter can be an indication of a predetermined capturing device, where the microphone grid of the device is known in advance or predetermined. In such cases the descriptive metadata would simply indicate the capturing device, and from this indication the parameters of the microphone grid associated with the capturing device can be obtained from a local store or database.With respect to Figure 14 an example electronic device is shown. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.

[0176] In some embodiments the device 1400 comprises at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes such as the methods such as described herein.

[0177] In some embodiments the device 1400 comprises a memory 1411.

[0178] In some embodiments the at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage means. In some embodiments the memory 1411 comprises a program code section for storing program codes implementable upon the processor 1407. Furthermore, in some embodiments the memory 1411 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1407 whenever needed via the memory-processor coupling.

[0179] In some embodiments the device 1400 comprises a user interface 1405. The user interface 1405 can be coupled in some embodiments to the processor 1407. In some embodiments the processor 1407 can control the operation of the user interface 1405 and receive inputs from the user interface 1405. In some embodiments the user interface 1405 can enable a user to input commands to the device 1400, for example via a keypad. In some embodiments the user interface 1405 can enable the user to obtain information from the device 1400. For example, the user interface 1405 may comprise a display configured to display information from the device 1400 to the user. The user interface 1405 can in some embodiments comprise a touchscreen or touch interface capable of both enabling information to be entered to the device 1400 and further displaying information to the user of the device 1400.

[0180] In some embodiments the device 1400 comprises an input / output port 1409. The input / output port 1409 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1407 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.

[0181] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802. X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).

[0182] In some embodiments the device 1400 may be employed to generate a suitable audio signal using the processor 1407 executing suitable code. The input / output port 1409 may be coupled to any suitable audio output for example to a multichannel speaker system and / or headphones (which may be a headtracked or a non-tracked headphones) or similar.

[0183] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described asblock diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0184] The embodiments of this invention may be implemented by computer software executable by a data processor of a mobile device for example, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.

[0185] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.

[0186] Embodiments of this invention can be practised as a computer software product in the form of an app. The app can either reside on the electronic / computing device or in an app repository such as an “App Store” which is typically sited remotely from the electronic / computing device. When the app is sited in an “App Store,” the app istypically downloaded from the “App Store” to an electronic / computing device over a communication network, such as an IP based network. The downloaded app comprising the invention can then execute as a computer software product on the electronic / computing device.

[0187] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0188] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

[0189] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.

Claims

CLAIMS:

1. An apparatus configured to:separate at least one metadata parameter from a group of metadata parameters, wherein the group of metadata parameters are related to an encoded immersive audio frame;include a parameter type of the at least one metadata parameter in a transport protocol payload; andinclude the at least one metadata parameter in the transport protocol payload.

2. The apparatus as claimed in Claim 1, wherein the transport protocol payload contains the encoded immersive audio frame.

3. The apparatus as claimed in Claims 1 and 2, wherein the transport protocol payload comprises Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, wherein the PI header comprises a PI type field associated with the PI data frame and wherein the apparatus configured to include a parameter type of the at least one metadata parameter in a transport protocol payload is configured to:include the parameter type of the at least one metadata parameter in the PI type field of the PI header; and wherein the apparatus configured to include the at least one metadata parameter in the transport protocol payload is configured to: include the at least one metadata in the PI data frame.

4. The apparatus as claimed in Claim 3, wherein the PI type field of the PI header indicates one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

5. The apparatus as Claimed in Claims 1 to 4, wherein the group of metadata parameters is MASA descriptive metadata.

6. The apparatus as claimed in Claims 1 to 5, wherein the transport protocol payload comprises one of:a real-time transport protocol payload;a real-time transport protocol payload header extension;a real-time transport control protocol payload; ora data channel transport protocol payload.

7. The apparatus as claimed in Claim 1 is further configured to include the parameter type of the at least one metadata parameter in a pi-types parameter of the session description protocol (SDP).

8. The apparatus as claimed in Claims 1 to 7, wherein the encoded immersive audio frame is an encoded IVAS frame.

9. An apparatus configured to:receive a transport protocol payload comprising at least one metadata parameter and a parameter type of the at least one metadata parameter;determine, from the parameter type of the at least one metadata parameter, that the at least one metadata parameter belongs to a group of metadata parameters, wherein the group of metadata parameters is related to an encoded immersive audio frame; andcollate the at least one metadata parameter with other metadata parameters of the group of metadata parameters.

10. The apparatus as claimed in Claim 9, wherein each metadata parameter of the group of metadata parameters has a unique parameter type.

11. The apparatus as claimed in Claims 9 and 10, is further configured to: determine a further metadata parameter of the group of metadata parameters, based on the parameter type of the at least one metadata parameter,wherein a parameter type of the further metadata parameter is different from the parameter type of the at least one metadata parameter; andcollate the further metadata parameter with the at least one metadata parameter and the other metadata parameters of the group of metadata parameters.

12. The apparatus as claimed in 11 , wherein the parameter type of the further metadata parameter is a source format.

13. The apparatus as claimed in 9, is further configured to:set at least one of the other metadata parameters to a predetermined value based on the parameter type of the at least one metadata parameter.

14. The apparatus as claimed in Claims 9 to 13, wherein the transport protocol payload comprises Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, wherein the PI header comprises a PI type field associated with the PI data frame, wherein the at least one metadata parameter is contained in the PI data frame, and wherein the parameter type of the at least one metadata parameter is contained in the PI type field of the PI header.

15. The apparatus as claimed in Claim 14, wherein the PI type field of the PI header indicates one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

16. The apparatus as Claimed in Claims 9 to 15, wherein the group of metadata parameters is MASA descriptive metadata.

17. The apparatus as claimed in Claims 9 to 16, wherein the transport protocol payload comprises one of:a real-time transport protocol payload;a real-time transport protocol payload header extension;a real-time transport control protocol payload; ora data channel transport protocol payload.

18. The apparatus as claimed in Claims 9 to 17, wherein the encoded immersive audio frame is an encoded IVAS frame.

19. A method comprising:separating at least one metadata parameter from a group of metadata parameters, wherein the group of metadata parameters are related to an encoded immersive audio frame;including a parameter type of the at least one metadata parameter in a transport protocol payload; andincluding the at least one metadata parameter in the transport protocol payload.

20. The method as claimed in Claim 19, wherein the transport protocol payload contains the encoded immersive audio frame.

21. The method as claimed in Claims 19 and 20, wherein the transport protocol payload comprises Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, wherein the PI header comprises a PI type field associated with the PI data frame and wherein the method comprising including a parameter type of the at least one metadata parameter in a transport protocol payload is comprises:including the parameter type of the at least one metadata parameter in the PI type field of the PI header; and wherein the method comprising including the at least one metadata parameter in the transport protocol payload is comprises:including the at least one metadata in the PI data frame.

22. The method as claimed in Claim 21 , wherein the PI type field of the PI header indicates one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

23. The method as Claimed in Claims 19 to 22, wherein the group of metadata parameters is MASA descriptive metadata.

24. The method as claimed in Claims 19 to 23, wherein the transport protocol payload comprises one of:a real-time transport protocol payload;a real-time transport protocol payload header extension;a real-time transport control protocol payload; ora data channel transport protocol payload.

25. The method as claimed in Claim 19 further comprises including the parameter type of the at least one metadata parameter in a pi-types parameter of the session description protocol (SDP).

26. The method as claimed in Claims 19 to 25, wherein the encoded immersive audio frame is an encoded IVAS frame.

27. A method comprising:receiving a transport protocol payload comprising at least one metadata parameter and a parameter type of the at least one metadata parameter;determining, from the parameter type of the at least one metadata parameter, that the at least one metadata parameter belongs to a group of metadata parameters, wherein the group of metadata parameters is related to an encoded immersive audio frame; andcollating the at least one metadata parameter with other metadata parameters of the group of metadata parameters.

28. The method as claimed in Claim 27, wherein each metadata parameter of the group of metadata parameters has a unique parameter type.

29. The method as claimed in Claims 27 and 28, further comprises:determining a further metadata parameter of the group of metadata parameters, based on the parameter type of the at least one metadata parameter, wherein a parameter type of the further metadata parameter is different from the parameter type of the at least one metadata parameter; andcollating the further metadata parameter with the at least one metadata parameter and the other metadata parameters of the group of metadata parameters.

30. The method as claimed in 29, wherein the parameter type of the further metadata parameter is a source format.

31. The method as claimed in 27, further comprises:setting at least one of the other metadata parameters to a predetermined value based on the parameter type of the at least one metadata parameter.

32. The method as claimed in Claims 27 to 31, wherein the transport protocol payload comprises Processing Information (PI) data comprising a PI header and a PI data frame corresponding to the PI header, wherein the PI header comprises a PI type field associated with the PI data frame, wherein the at least one metadata parameter is contained in the PI data frame, and wherein the parameter type of the at least one metadata parameter is contained in the PI type field of the PI header.

33. The method as claimed in Claim 32, wherein the PI type field of the PI header indicates one of; Source Format, Channel Layout, Channel Distance, Transport Definition, Channel Angle or Channel Direction.

34. The method as Claimed in Claims 27 to 33, wherein the group of metadata parameters is MASA descriptive metadata.

35. The method as claimed in Claims 27 to 34, wherein the transport protocol payload comprises one of:a real-time transport protocol payload;a real-time transport protocol payload header extension;a real-time transport control protocol payload; ora data channel transport protocol payload.

36. The method as claimed in Claims 27 to 34, wherein the encoded immersive audio frame is an encoded IVAS frame.