Rendering support in immersive conversational audio
The split rendering approach in immersive audio codecs reduces computational load and latency on lightweight devices by initial rendering on capable devices and packetizing RTP payloads with split rendering parameters, addressing rendering challenges on lightweight devices and enabling seamless format changes.
Patent Information
- Application Number
- PCT/EP2024/085928
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-03
AI Technical Summary
Existing immersive audio codecs face challenges in efficiently rendering spatial audio on lightweight devices with limited processing capabilities, particularly in split rendering scenarios, where computational load and latency are high, and session renegotiation for format changes is slow.
The proposed solution involves creating a split rendering approach where initial rendering is performed on a capable device, generating a header to identify the rendering payload, and packetizing it for transmission using RTP, allowing seamless switching between split and non-split rendering without session renegotiation, by incorporating split rendering parameters in the RTP payload format.
This approach reduces computational load on lightweight devices by offloading rendering processes to capable devices, lowers motion-to-sound latency, and enables seamless adaptation to network conditions and listener changes during ongoing audio sessions.
Smart Images

Figure EP2024085928_03072025_PF_FP_ABST
Abstract
Description
[0001] RENDERING SUPPORT IN IMMERSIVE CONVERSATIONAL AUDIO
[0002] Field
[0003] The present application relates to apparatus and methods for lower complexity rendering support in immersive conversational audio, but not exclusively for lower complexity rendering support in immersive conversational audio employing immersive voice and audio services (IVAS) real-time transport protocol (RTP) payload.
[0004] Background
[0005] Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the immersive voice and audio services (IVAS) codec which is being designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network. Such immersive services include uses for example in immersive voice and audio for applications such as virtual reality (VR), augmented reality (AR) and mixed reality (MR) as well as spatial voice communication including teleconferencing. This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.
[0006] The input signals are presented to the IVAS encoder in one of the supported formats (and in some allowed combinations of the formats). Similarly, it is expected that the decoder can output the audio in supported formats. A pass-through mode has been proposed, where the audio could be provided in its original format after transmission (encoding / decoding).
[0007] The immersive audio can be rendered at the receiving device (e.g., a mobile phone) to generate a spatial audio scene to the end-user. In some cases, however, the receiving device might be a lightweight device with little processing capabilities. In such cases, to lower the complexity of the rendering at the receiver, the rendering process can be divided to multiple parts where the different renderings can be processed at different devices or processing units. For example, with two rendering parts, the first rendering or pre-rendering can be performed on a separate device (or the same device) as the receiving end-user device. The first rendering device receives (and decodes if needed) the input signal to be rendered and renders the input signal into an intermediate rendering format. The intermediate rendering format is then passed (transmitted / stored / moved) to a second rendering or postrendering device which renders the intermediate rendering format into the output signals. This division to first / second or pre / post renderings can also be called split rendering.
[0008] ISAR (Immersive Audio for Split Rendering Scenarios) targets development of solutions for immersive binaural audio on head-tracked devices that are compatible with the split architectures envisaged in 3GPP, especially for low- complexity and lightweight devices (e.g., AR glasses), and that demonstrate operational benefits over solutions with full decoding and rendering in the end-user device. Such split rendering solutions can relate to pre-rendering operations on a capable device or entity resulting in an intermediate representation including, e.g., audio with (or without) post-rendering control metadata, transmission of said intermediate representation, and post-rendering operations that generate the audio for playback on the low-complexity and lightweight (end-user) device. Split rendering capability can be considered to be in the scope of the IVAS codec.
[0009] Additionally RTP (Real-Time Transport Protocol) is intended for an end-to- end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data transfer to multiple destinations through IP multicast or to a specific destination through IP unicast. The majority of the RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols may also be utilized. RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol (RTSP).
[0010] The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transfer of multimedia data, and its companion protocol (RTCP) is used to periodically send control information and QoS (Quality of Service) parameters. RTP sessions are typically initiated between client and server or between client and another client (or a multi-party topology) using a signalling protocol, such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols typically use the Session Description Protocol (SDP), such as defined by RFC 8866 to specify parameters for the sessions.
[0011] There is provided according to a first aspect an apparatus for transmission of conversational immersive audio, the apparatus comprising means configured to: receive audio data, during the transmission of conversational immersive audio; generate a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generate a header configured to identify the rendering payload; generate a conversational immersive audio payload comprising the header and the rendering payload; and packetize the conversational immersive audio payload for transmission.
[0012] The means configured to receive audio data, during the transmission of conversational immersive audio may be configured to receive an immersive audio bitstream.
[0013] The immersive audio bitstream may be an IVAS frame.
[0014] The means configured to generate a rendering payload based on the audio data may be configured to: decode immersive audio data from the immersive audio bitstream; render at least partially the decoded immersive audio data by generating a partially rendered interim format audio data suitable for further rendering; and encode the partially rendered interim format audio data to generate a rendering payload.
[0015] The header may be an IVAS RTP payload header.
[0016] The header may comprise a ToC byte defining that the payload comprises the partially rendered interim format audio data.
[0017] The means may be configured to packetize the conversational immersive audio payload for transmission as a real-time transport protocol packet payload. The means may be further configured to receive information identifying at least one further rendering external to the apparatus format, wherein the means configured to generate the header may be configured to generate the header based on the information identifying the at least one further rendering external to the apparatus format.
[0018] The means may be further configured to receive from the further apparatus information identifying the further apparatus is capable of further rendering external to the apparatus, wherein the means configured to generate the rendering payload based on the audio data may be configured to generate the rendering payload as the initial rendering for a split rendering, based on the information identifying the further apparatus is capable of further rendering external to the apparatus.
[0019] The means may be further configured to receive from the further apparatus at least one of: listener position information; and listener orientation information; select a degree-of-freedom parameter based on the at least one of: listener position information; and listener orientation information, wherein the means configured to generate the header may be configured to generate a header based on the selected degree-of-freedom parameter.
[0020] The means may be further configured to negotiate with a further apparatus support for the conversational immersive audio payload.
[0021] The means configured to negotiate with the further apparatus may be configured to negotiate employing a session description file.
[0022] According to a second aspect there is provided an apparatus for rendering of conversational immersive audio, the apparatus comprising means configured to: receive a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determine based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0023] The means configured to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering may be further configured to transmit the rendering payload for further rendering to a further apparatus. The means configured to transmit the rendering payload for further rendering to the further apparatus may be configured to transmit the rendering payload for further rendering to the further apparatus based on at least one of: a determination that the apparatus is unable to perform the further rendering of the rendering payload comprising the initial rendering; and a selection for transmitting the rendering payload for further rendering to the further apparatus.
[0024] The means configured to receive a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload may be configured to: receive at least one real-time transport protocol packet; parse the at least one real-time transport protocol packet to obtain the packetized conversational immersive audio payload.
[0025] The at least one real-time transport protocol packet may comprise a realtime transport protocol payload.
[0026] The header may indicate that the payload comprises the initial rendering.
[0027] The means configured to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering may be further configured to: decode the rendering payload based on the header; and further render the rendering payload.
[0028] The means may be further configured to generate and transmit to an additional apparatus information identifying at least one further rendering format for the apparatus, wherein the received packetized conversational immersive audio payload may be based on the information identifying the at least one further rendering format.
[0029] The means may be further configured to generate and transmit to an additional apparatus information identifying the apparatus is capable of further rendering.
[0030] The means may be further configured to obtain and transmit to the additional apparatus at least one of: listener position information; and listener orientation information.
[0031] The means may be further configured to negotiate with the additional apparatus support for further rendering.
[0032] The means configured to negotiate with the additional apparatus may be configured to negotiate employing a session description file. According to a third aspect there is provided a method for an apparatus for transmission of conversational immersive audio, the method comprising: receiving audio data, during the transmission of conversational immersive audio; generating a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generating a header configured to identify the rendering payload; generating a conversational immersive audio payload comprising the header and the rendering payload; and packetizing the conversational immersive audio payload for transmission.
[0033] Receiving audio data, during the transmission of conversational immersive audio may comprise receiving an immersive audio bitstream.
[0034] The immersive audio bitstream may be an IVAS frame.
[0035] Generating a rendering payload based on the audio data may comprise: decoding immersive audio data from the immersive audio bitstream; rendering at least partially the decoded immersive audio data by generating a partially rendered interim format audio data suitable for further rendering; and encoding the partially rendered interim format audio data to generate a rendering payload.
[0036] The header may be an IVAS RTP payload header.
[0037] The header may comprise a ToC byte defining that the payload comprises the partially rendered interim format audio data.
[0038] The method may comprise packetizing the conversational immersive audio payload for transmission as a real-time transport protocol packet payload.
[0039] The method may further comprise receiving information identifying at least one further rendering external to the apparatus format, wherein generating the header may comprise generating the header based on the information identifying the at least one further rendering external to the apparatus format.
[0040] The method may further comprise receiving from the further apparatus information identifying the further apparatus is capable of further rendering external to the apparatus, wherein generating the rendering payload based on the audio data may comprise generating the rendering payload as the initial rendering for a split rendering, based on the information identifying the further apparatus is capable of further rendering external to the apparatus.
[0041] The method may further comprise receiving from the further apparatus at least one of: listener position information; and listener orientation information; select a degree-of-freedom parameter based on the at least one of: listener position information; and listener orientation information, wherein generating the header may comprise generating a header based on the selected degree-of-freedom parameter.
[0042] The method may further comprise negotiating with a further apparatus support for the conversational immersive audio payload.
[0043] Negotiating with the further apparatus may comprise negotiating employing a session description file.
[0044] According to a fourth aspect there is provided a method for an apparatus for rendering of conversational immersive audio, the method comprising: receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determining based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0045] Providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering may further comprise transmitting the rendering payload for further rendering to a further apparatus.
[0046] Transmitting the rendering payload for further rendering to the further apparatus may comprise transmitting the rendering payload for further rendering to the further apparatus based on at least one of: a determination that the apparatus is unable to perform the further rendering of the rendering payload comprising the initial rendering; and a selection for transmitting the rendering payload for further rendering to the further apparatus.
[0047] Receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload may comprise: receiving at least one real-time transport protocol packet; parsing the at least one real-time transport protocol packet to obtain the packetized conversational immersive audio payload.
[0048] The at least one real-time transport protocol packet may comprise a realtime transport protocol payload.
[0049] The header may indicate that the payload comprises the initial rendering.
[0050] Providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering may further comprise: decoding the rendering payload based on the header; and further rendering the rendering payload.
[0051] The method may further comprise generating and transmitting to an additional apparatus information identifying at least one further rendering format for the apparatus, wherein the received packetized conversational immersive audio payload may be based on the information identifying the at least one further rendering format.
[0052] The method may further comprise generating and transmitting to an additional apparatus information identifying the apparatus is capable of further rendering.
[0053] The method may further comprise obtaining and transmitting to the additional apparatus at least one of: listener position information; and listener orientation information.
[0054] The method may further comprise negotiating with the additional apparatus support for further rendering.
[0055] Negotiating with the additional apparatus may comprise negotiating employing a session description file.
[0056] According to a fifth aspect there is provided an apparatus for transmission of conversational immersive audio, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive audio data, during the transmission of conversational immersive audio; generate a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generate a header configured to identify the rendering payload; generate a conversational immersive audio payload comprising the header and the rendering payload; and packetize the conversational immersive audio payload for transmission.
[0057] The apparatus caused to receive audio data, during the transmission of conversational immersive audio may be caused to receive an immersive audio bitstream.
[0058] The immersive audio bitstream may be an IVAS frame.
[0059] The apparatus caused to generate a rendering payload based on the audio data may be caused to: decode immersive audio data from the immersive audio bitstream; render at least partially the decoded immersive audio data by generating a partially rendered interim format audio data suitable for further rendering; and encode the partially rendered interim format audio data to generate a rendering payload.
[0060] The header may be an IVAS RTP payload header.
[0061] The header may comprise a ToC byte defining that the payload comprises the partially rendered interim format audio data.
[0062] The apparatus caused to packetize the conversational immersive audio payload for transmission as a real-time transport protocol packet payload.
[0063] The apparatus may be further caused to receive information identifying at least one further rendering external to the apparatus format, wherein apparatus caused to generate the header may be caused to generate the header based on the information identifying the at least one further rendering external to the apparatus format.
[0064] The apparatus may be further caused to receive from the further apparatus information identifying the further apparatus is capable of further rendering external to the apparatus, wherein the apparatus caused to generate the rendering payload based on the audio data may be caused to generate the rendering payload as the initial rendering for a split rendering, based on the information identifying the further apparatus is capable of further rendering external to the apparatus.
[0065] The apparatus may be further caused to receive from the further apparatus at least one of: listener position information; and listener orientation information; select a degree-of-freedom parameter based on the at least one of: listener position information; and listener orientation information, wherein the apparatus caused to generate the header may be caused to generate a header based on the selected degree-of-freedom parameter.
[0066] The apparatus may be further caused to negotiate with a further apparatus support for the conversational immersive audio payload.
[0067] The apparatus caused to negotiate with the further apparatus may be further caused to negotiate employing a session description file.
[0068] According to a sixth aspect there is provided an apparatus for rendering of conversational immersive audio, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive a packetized conversational immersive audio payload; wherein the conversational immersive audio payload comprises a header and a payload, determine based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0069] The apparatus caused to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering may be further caused to transmit the rendering payload for further rendering to a further apparatus.
[0070] The apparatus caused to transmit the rendering payload for further rendering to the further apparatus may be caused to transmit the rendering payload for further rendering to the further apparatus based on at least one of: a determination that the apparatus is unable to perform the further rendering of the rendering payload comprising the initial rendering; and a selection for transmitting the rendering payload for further rendering to the further apparatus.
[0071] The apparatus caused to receive a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload may be caused to: receive at least one real-time transport protocol packet; parse the at least one real-time transport protocol packet to obtain the packetized conversational immersive audio payload.
[0072] The at least one real-time transport protocol packet may comprise a realtime transport protocol payload. The header may indicate that the payload comprises the initial rendering.
[0073] The apparatus caused to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering may be further caused to: decode the rendering payload based on the header; and further render the rendering payload.
[0074] The apparatus may be further caused to generate and transmit to an additional apparatus information identifying at least one further rendering format for the apparatus, wherein the received packetized conversational immersive audio payload may be based on the information identifying the at least one further rendering format.
[0075] The apparatus may be further caused to generate and transmit to an additional apparatus information identifying the apparatus is capable of further rendering.
[0076] The apparatus may be further caused to obtain and transmit to the additional apparatus at least one of: listener position information; and listener orientation information.
[0077] The apparatus may be further caused to negotiate with the additional apparatus support for further rendering.
[0078] The apparatus caused to negotiate with the additional apparatus may be caused to negotiate employing a session description file.
[0079] According to a seventh aspect there is provided an apparatus for transmission of conversational immersive audio, the apparatus comprising: receiving circuitry configured to receive audio data, during the transmission of conversational immersive audio; generating circuitry configured to generate a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generating circuitry configured to generate a header configured to identify the rendering payload; generating circuitry configured to generate a conversational immersive audio payload comprising the header and the rendering payload; and packetizing circuitry configured to packetize the conversational immersive audio payload for transmission. According to an eighth aspect there is provided an apparatus for rendering of conversational immersive audio, the apparatus comprising: receiving circuitry configured to receive a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determining circuitry configured to determine based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and providing circuitry configured to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0080] According to a ninth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for transmission of conversational immersive audio to perform at least the following: receiving audio data, during the transmission of conversational immersive audio; generating a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generating a header configured to identify the rendering payload; generating a conversational immersive audio payload comprising the header and the rendering payload; and packetizing the conversational immersive audio payload for transmission.
[0081] According to a tenth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus for rendering of conversational immersive audio to perform at least the following: receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determining based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0082] According to an eleventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for transmission of conversational immersive audio to perform at least the following: receiving audio data, during the transmission of conversational immersive audio; generating a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generating a header configured to identify the rendering payload; generating a conversational immersive audio payload comprising the header and the rendering payload; and packetizing the conversational immersive audio payload for transmission.
[0083] According to a twelfth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for rendering of conversational immersive audio to perform at least the following: receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determining based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0084] According to a thirteenth aspect there is provided an apparatus for transmission of conversational immersive audio comprising: means for receiving audio data, during the transmission of conversational immersive audio; means for generating a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; means for generating a header configured to identify the rendering payload; means for generating a conversational immersive audio payload comprising the header and the rendering payload; and means for packetizing the conversational immersive audio payload for transmission.
[0085] According to a fourteenth aspect there is provided an apparatus for rendering of conversational immersive audio comprising: means for receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; means for determining based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and means for providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial.
[0086] According to a fifteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for transmission of conversational immersive audio to perform at least the following: receiving audio data, during the transmission of conversational immersive audio; generating a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generating a header configured to identify the rendering payload; generating a conversational immersive audio payload comprising the header and the rendering payload; and packetizing the conversational immersive audio payload for transmission.
[0087] According to a sixteenth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for rendering of conversational immersive audio to perform at least the following: receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determining based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
[0088] An apparatus comprising means for performing the actions of the method as described above.
[0089] An apparatus configured to perform the actions of the method as described above.
[0090] A computer program comprising program instructions for causing a computer to perform the method as described above.
[0091] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.
[0092] An electronic device may comprise apparatus as described herein.
[0093] A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art.
[0094] Summary of the Figures
[0095] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:
[0096] Figure 1 shows schematically example server and peer-to-peer teleconferencing systems within which embodiments may be implemented;
[0097] Figure 2 shows schematically an example packet format with Enhanced Voice Services (EVS) Header-Full payload structure;
[0098] Figure 3 shows schematically an example Table of Content (ToC) byte structure for a EVS codec;
[0099] Figure 4 shows schematically an example Table of Content (ToC) byte structure for an IVAS frame according to some embodiments;
[0100] Figure 5 shows schematically an example Table of Content (ToC) byte structure for an IVAS split rendering frame according to some embodiments;
[0101] Figure 6 shows schematically an example frame structure with a ToC SSB bits indicating a SR frame, and BLI bits employed for bitrate indication according to some embodiments;
[0102] Figure 7 shows schematically an example frame structure with a ToC byte (SSB and BLI combination) indicating a SR frame according to some embodiments;
[0103] Figure 8 shows schematically an example frame structure with a ToC byte indicating a SR frame type for the SR frame according to some embodiments;
[0104] Figure 9 shows schematically an example frame structure with a ToC byte indicating an extension frame type according to some embodiments;
[0105] Figure 10 shows schematically an example frame structures where the SR header is integrated as part of the SR frame according to some embodiments;
[0106] Figures 11 a and 11 b show schematically example PI frames indicating a SR frame according to some embodiments;
[0107] Figure 12 shows a schematic view of an example exchange of different PI frames and IVAS frames between two UEs in an IVAS session;
[0108] Figure 13 shows schematically a flow diagram session negotiation offer and answer operations according to some embodiments; Figure 14 shows schematically a flow diagram session for packetization and de-packetization of split rendered frames according to some embodiments;
[0109] Figures 15a, 15b to 15c show schematically example end to end systems of IVAS split rendering negotiation, initiation and indication in RTP streams according to some embodiments; and
[0110] Figure 16 shows an example device suitable for implementing the apparatus shown.
[0111] Embodiments of the Application
[0112] The following describes in further detail suitable apparatus and possible mechanisms for the provision of efficient IVAS audio.
[0113] An example system within which embodiments may be implemented is shown in Figure 1 .
[0114] Figure 1 , for example, shows an example teleconferencing system within which some embodiments can be implemented. In this example there is shown two sites or rooms, Room A 100 and Room B 102. Room A 100 comprises a ‘talker’ or user, Talker 1 103. Room B 102 comprises one ‘talker’ or user, Talker RX 141.
[0115] In the following example within room A is a suitable teleconference apparatus (or more generally telecommunications apparatus 110) configured to spatially capture and encode the audio environment and furthermore is configured to render a spatial audio signal to the room. The apparatus can in some embodiments be implemented by a user equipment (UE) operating within a cellular communications system or accessing any suitable access network. Within each of the other rooms may be a suitable teleconference apparatus (or more generally telecommunications apparatus such as apparatus 120 within room B) configured to render a spatial audio signal to the room and furthermore is configured to capture and encode at least a mono audio and optionally configured to spatially capture and encode the audio environment.
[0116] In the following examples each room is provided with the means to spatially capture, encode spatial audio signals, receive spatial audio signals and render these to a suitable listener. It would be understood that there may be other embodiments where the system comprises some apparatus configured to only capture and encode audio signals (in other words the apparatus is a ‘transmit’ only apparatus), and other apparatus configured to only receive and render audio signals (in other words the apparatus is a ‘receive’ only apparatus). In such embodiments the system within which embodiments may be implemented may comprise apparatus with varying abilities to capture / render audio signals.
[0117] The teleconference apparatus (for each site or room) 110, 120 can be configured to call into a teleconference controlled by and implemented over a server 111.
[0118] In some embodiments the communications or teleconferencing system comprises a (peer-to-peer) communications system (rather than the server based system shown in Figure 1 ) within which some embodiments can be implemented. Thus, for example, two or more UEs can be configured to interact directly with each other (for example to implement an immersive audio phone call between users). In such a scenario one of the UEs can be configured to deliver spatial ambience as one stream and employ a close-up microphone (for example a Lavalier microphone) to capture the speech as an audio object or audio source. The sender UE can be configured to encode the spatial ambience audio signals in a MASA format stream and the close-up microphone audio signal as an object format stream. The two audio streams can then be delivered as separated IVAS streams. The sender UE, in addition, can be configured to encode processing information during the encoding to deliver the PI frames together with the IVAS frames to the receiver UE.
[0119] The teleconference apparatus can be configured to spatially capture and encode the audio environment and furthermore can be configured to render a spatial audio signal to the room. In this example only the communications or signalling path from the Room A 100 to the Room B 102 is shown for simplicity but a duplex or multipoint communication system comprising multiple signalling paths can be implemented using the methods as described herein without significant inventive input.
[0120] The teleconference apparatus (for each site or room) 110, 120 is further configured to communicate with each other to implement a teleconference function.
[0121] As shown in Figure 1 , the apparatus 110, 120 and server 111 can comprise suitable encoder and decoder functionality. For example, the apparatus 110 is shown comprising an (IVAS) encoder 101 , the server 111 is shown comprising a (IVAS) decoder and encoder 121 and the apparatus 120 is shown comprising an (IVAS) decoder 131. In such a manner an object 120 (the audio signals representing the user or talker 1 111 ) can be encoded by the encoder 101 which generates a bitstream 106 to be passed to a server 111. The server 111 can then decode, (optionally then mix with other objects and otherwise process the audio signals) and encode then to generate the bitstream 108 to be passed to the apparatus 120. The apparatus 120 can then decode the audio signals and present them to the user or talker ‘Talker RX’ 141 .
[0122] Although this example shows a teleconference application the encoder / decoder functionality can be applied to the streaming of any suitable media.
[0123] The IVAS decoder / renderer for each of the teleconference apparatus 102 can be furthermore configured to handle multiple input streams that may each originate from a different encoder.
[0124] Furthermore as discussed above a 3GPP Work Item ISAR (Immersive Audio for Split Rendering Scenarios) targets development of solutions for immersive binaural audio on head-tracked devices that are compatible with the split architectures envisaged in 3GPP, especially for low-complexity and lightweight devices (e.g., AR glasses), and that demonstrate operational benefits over solutions with full decoding and rendering in the end-user device. Furthermore, the apparatus in Room B may be such that it is computationally constrained and hence prefers to perform the IVAS stream decoding and rendering in as low computational resources as possible. In such a scenario, the server can deliver the IVAS stream with split render processing. This results in delivery of split rendered frames from the server to room B and the room B user’s head orientation information is delivered to the server to perform split rendering taking into account the latest head orientation information of the listener in room B.
[0125] As discussed previously RTP is intended for an end-to-end, real-time transfer of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP is furthermore designed to carry a multitude of multimedia formats, which permit the transport of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may therefore require a profile and payload format specifications.
[0126] The profile is configured to define the codec used to encode the payload data and the mapping to payload format codes in the protocol field Payload Type (PT) of the RTP header.
[0127] For example, the RTP profile for audio and video conferences with minimal control is defined in RFC 3551 . The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codec such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798.
[0128] An RTP session can be established for each multimedia stream. Audio and video streams may be implemented which use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification can furthermore be configured to recommend port numbers for RTP, and furthermore to recommend the use of the next odd port number for the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.
[0129] Each RTP stream can comprise RTP packets, and the RTP packet in turn can comprise a RTP header and payload pair.
[0130] Enhanced Voice Services (EVS) is a mono voice codec standardized in 3GPP and described in the TS 26.445 specification document. The codec can have two operating modes: EVS Primary and EVS AMR-WB IO (Adaptive Multi Rate Wideband Inter-Operable).
[0131] The IVAS codec can be considered to be an extension to the EVS codec and as such the IVAS and EVS codecs can have some similarities in terms of design and implementation. The RTP payload structure, while not having been specified for the IVAS codec, is envisioned to have similarities to the RTP payload structure in EVS. The RTP payload format of EVS is described in 3GPP TS 26.445 Annex A. In EVS, the RTP payload format is divided into two different embodiments: a Compact format and a Header-Full format. In the EVS Compact payload format, a RTP packet includes only a single EVS speech frame for EVS Primary mode. For EVS AMR-WB IO mode, the compact RTP packet also includes a 3-bit Codec Mode Request (CMR) field in front of the speech frame. In the EVS Compact format, the different modes and bitrates for the speech frames are identified by the size of the RTP payload. For example, an RTP packet of size 328 bits is assigned for EVS Primary mode with 16.4 kbps bitrate, as is shown in Table A.1 in TS 26.445 Annex A.
[0132] In EVS Header-Full format, the RTP payload consists of the speech frame(s) accompanied by an optional CMR byte and Table of Content (ToC) bytes. The CMR byte is used to request a change in bitrate or coding mode that the receiver wants to receive. The request is sent as a CMR byte as part of the Header-Full EVS packet. In EVS AMR-WB IO Compact format, the CMR functionality is also present as a 3-bit signalling at the beginning of the packet.
[0133] The ToC bytes in EVS Header-Full format describe the mode and bitrate for the accompanied EVS coded speech frames. RTP payload structures for EVS Header-Full format are furthermore shown in Figure 2.
[0134] Figure 2 shows, for example, a payload structure of Header-Full format with ToC single frame 201 , a payload structure of Header-Full format with ToC multiple frames 203, a payload structure of Header-Full format with CMR and ToC single frame 205 and a payload structure of Header-Full format with CMR and ToC multiple frames 207.
[0135] Each of the payload structures comprises ToC / CMR bytes where for each of these bytes there is a first bit 211 , 221 , 231 at the beginning of the ToC / CMR bytes which can be used to differentiate the bytes between ToC (where the bit value is 0) and CMR (where the bit value is 1 ).
[0136] In the payload structure 201 where there is a single speech frame in the payload, the speech data 215 is preceded by a single ToC byte 213 (with optional zero padding 217 at the end of the payload).
[0137] In examples where the single speech frame payload structure can have an optional CMR byte, such as shown in payload structure 205. The payload structure 205 differs from the payload structure 201 in that there is a CMR byte 232 (with associated indicator bit 231 value of 1 ) positioned before the ToC byte 213.
[0138] In some example structures there can be two speech frames in the payload structure 203, 207. In this example an additional ToC byte 223 (with associated indicator bit 221 value of 0) and the ToC byte 213 is positioned at the beginning of the payload followed by the speech data frames 215 and 225 respectively to the order of the ToC bytes.
[0139] Furthermore the multiple speech frame payload structure can have an optional CMR byte, such as shown in payload structure 207. The payload structure 207 differs from the payload structure 203 in that there is a CMR byte 232 (with associated indicator bit 231 value of 1 ) positioned before the ToC bytes 213, 223.
[0140] With respect to Figure 3 the EVS Header-Full frame ToC byte structure 301 according to some embodiments is shown. The Header Type (H) identification bit 303 is used to differentiate between CMR and ToC bytes, the (F) bit 305 is used to indicate if another ToC byte follows this frame and the Frame type (FT) index 307 is used to indicate the EVS bitrate and mode of the frame.
[0141] Split rendering as discussed above and previously is an approach of reducing the processing complexity and possibly also deployment footprint on the final rendering device that renders received bitstream into output audio signals. Split rendering can be implemented by splitting the rendering task into two steps employed on separate devices (or the same device) a first “prerenderer” device usually receives (and decodes if needed) the input signal to be rendered and then renders the input signal into an intermediate split Tenderer format. This split Tenderer format is then passed (transmitted / stored / moved) to a second “postrenderer” device which renders the split Tenderer format into the output signals.
[0142] The aim of the approach is to reduce complexity at the “postrenderer” device with the cost of increased complexity at the “prerenderer” device while maintaining the quality of the rendering. Depending on the implementation, it is also possible that the overall total complexity increases. This is a valid sacrifice when the benefit is that the “postrenderer” (and rendering in general) can be deployed to a low power device such as wireless headphones. Please note that the term Tenderer and post Tenderer are contextual for a given deployment setup. In some scenarios, the network connectivity between the Tenderer device and the post Tenderer device may be a cellular network connection (e.g., an MCU performing split rendering of incoming IVAS stream from other UEs before delivering the split rendered audio frames to a mobile device connected to the network via a 4G / 5G bearer) whereas in some cases it can be a BT / WIFI (e.g., a mobile device before delivering split rendered audio frames to the earbuds) or any other fixed wireless connection.
[0143] The proposed split Tenderer is used for head-tracked binaural Tenderer where the split rendering approach also allows reducing motion-to-sound delay as the “postrenderer” is intended to be located in head-tracked headphones.
[0144] The operational flow in an example split Tenderer is following:
[0145] Renderer or Prerenderer (if read in context of postrenderer) (and IVAS decoder):
[0146] Receive IVAS bitstream;
[0147] Decode IVAS bitstream to decoded format;
[0148] Render decoded format using binaural rendering to a default head pose;
[0149] Render decoded format to multiple alternative head pose;
[0150] Determine from default head pose binaural rendered signal and alternative head pose binaural rendered signals split rendering metadata that represents how the default head pose binaural renderer signals should be altered to obtain the alternative head pose binaural rendered signals;
[0151] Construct and transmit a split renderer bitstream from the default head pose binaural rendered signals and the determined split rendering metadata.
[0152] Postrenderer:
[0153] Receive and decode split renderer bitstream;
[0154] Obtain “postrenderer” device head pose;
[0155] Render the decoded split renderer signals into binaural output using the actual “postrenderer” head pose, the decoded default head pose binaural rendered signals and the split rendering metadata. The EVS ToC byte indicating IVAS bitrates and frame contents in some embodiments is presented in the following table.
[0156] As discussed split rendering is a processing approach between two processing entities (e.g., a server and a mobile phone or a phone and wireless headphones) where the rendering process is divided between the two processing entities. In a traditional rendering approach, e.g., in an IVAS session, a sender UE encodes an IVAS coded audio frame, transmits an IVAS coded bitstream frame, a receiver receives a IVAS coded bitstream, decodes the bitstream, and renders the decoded bitstream to some output playback system. In the split rendering approach, part of the rendering is already processed at the sender side (or by an intermediate processing entity) before transmitting any bitstream to the receiver. This way the computational load and motion to sound latency of the rendering at the final receiver side can be reduced.
[0157] For the IVAS codec, a split rendering approach has been proposed to be part of the codec. In the proposed split rendering approach, a receiver decodes and renders an IVAS bitstream to some output format, e.g., to binaural format. Then, the receiver re-encodes (or transcodes) the rendered output into a format that is referred to as split rendering bitstream. This split rendering bitstream can be further transmitted, e.g., to a further playback device such as headphones. The further playback device would then decode and render the received split rendering bitstream to the listener.
[0158] The example method reduces the rendering load and motion to sound latency at the end device (such as headphones) as part of the rendering is processed at the intermediate device or processing entity (such as a mobile phone). Furthermore, the processing load of the mobile phone could be diminished by shifting the split rendering processing from the phone to the network (e.g., to an edge server). In this approach, an IVAS bitstream would be decoded, rendered, and re-encoded to a split rendering bitstream already at the edge server. The split rendered bitstream could then be transmitted to the mobile phone over a suitable network, which could further pass the bitstream to the headphones for final rendering.
[0159] The split rendering method requires that the split rendered bitstream is transmitted from a network based processing element (e.g., an edge server) to a receiver (e.g., a mobile phone). The split rendered bitstream needs to be transported with minimal end-to-end latency. For example, the split rendering bitstream is transmitted as an IVAS RTP payload over RTP transport. However, the current proposed IVAS RTP payload format supports only a regular (non-split rendered) IVAS bitstream. In other words there is currently no defined way of carrying split rendered data within the IVAS RTP payload. Additionally, in order to ensure seamless adaptation over different network environments, the payload format design is, as defined in the following examples, configured to support both split rendered and regular (non-split rendered) IVAS bitstreams in the same IVAS session. From another viewpoint in order to ensure seamless IVAS end-to-end capability, there needs to be additional functionality to carry both types of data in the IVAS RTP payload.
[0160] In addition to the edge server approach, a split rendering approach within a personal area network (e.g., Bluetooth or BT network) or local network (such as a mesh or WIFI network) (e.g., from mobile phone to headphones) is a valid processing approach in an IVAS session. In some cases, it might be beneficial for the sender to know if the receiver (e.g., a mobile phone) is using split rendering output processing. Currently, there is no such signaling available and signaling about the subsequent usage is discussed in the following examples to be added to an IVAS session.
[0161] The concept as discussed in the following embodiments relates to apparatus and methods for carriage of split rendered immersive conversational audio in an immersive audio transmission session (between sender UE and receiver UE) where there is provided a method for identifying the split rendered payloads in an immersive audio transport payload to enable transmission and rendering of the split rendered and non-split rendered immersive audio interchangeably in an ongoing immersive audio transmission session without the need for session re-negotiation. This results in seamlessly offloading computational complexity to the sender UE and taking it back if needed during an ongoing immersive audio transmission session. This overcomes the problem of slow session re-negotiation.
[0162] There is also provided apparatus and methods for identifying at least one split rendering output format of the receiving device during the session negotiation phase. The split rendering output format or a change to use split rendering output format may also be indicated during the session via bi-directional (or feedback providing) processing information frames. Similar functionality may also be used to request to start to use split rendering during the session by one or more parties in the session.
[0163] In some embodiments, the sender UE is able to switch from transmission from immersive audio coded bitstream to a split rendered bitstream as RTP payload based on an indication from the receiver UE.
[0164] In some embodiments, the sender UE transmitting split rendered bitstream as RTP payload is configured to select the degree of freedom parameter based on the received predicted pose information. This results in automatic adaptation implicitly based on listener pose changes.
[0165] In some further embodiments, the inclusion of split-rendered bitstream (e.g., split rendered and coded audio frame derived from rendering of IVAS frame) alternatively to the immersive audio coded bitstream (e.g., IVAS frame) is negotiated during session setup. This ensures that the sender UE and the receiver UE are enabled or capable to generate and transmit split-rendered audio frames while the receiver UE is able to parse, process and render the split-rendered audio frames.
[0166] In some embodiments, a split rendering capability parameter is included in the session description file for session negotiation. The presence of this parameter in the session offer indicates the readiness of the session offer creating UE for supporting split rendering immersive audio coded bitstream delivery. The receiver UE (of the offer) can retain the capability in the answer if it wishes to support or consume split rendering, which comes with the additional task of sending head orientation signaling.
[0167] In yet further embodiments, the split rendering attribute in the session description carries additional switch to indicate whether the split rendering is receiver UE initiated only or sender UE initiated only or either of the UEs.
[0168] This split rendered bitstream support in an IVAS session can in some embodiments be achieved as follows:
[0169] Sender (delivering split rendered audio frames):
[0170] Collect audio data from a source;
[0171] Encode the audio data using IVAS split rendering approach;
[0172] Create a split rendering header to indicate necessary split rendering related information and attach it to the split rendered bitstream resulted from the encoding;
[0173] Create an identifier header to recognize split rendered payload and attach it to the split rendering data resulted from the previous steps to create an RTP payload;
[0174] Packetize the RTP payload and transmit it to the receiver.
[0175] Receiver (receiving split rendered audio frames): Receive a RTP packet;
[0176] Read the identifier header and recognize a split rendered payload type
[0177] Either:
[0178] Read the split rendering header information to process the split rendered frame correctly;
[0179] Decode the split rendered frame accordingly;
[0180] Transmit the decoded data to a rendering entity which can render the data accordingly;
[0181] Or:
[0182] Transmit (or pass-forward) the split rendering payload to a further rendering entity;
[0183] As shown in Figure 2, the Header-Full IVAS frame comprises one or several Table-of-Content (ToC) indicators 213, 223 and their associated IVAS speech / audio frames 215, 225. With respect to Figure 4 is shown an example Table-of-Content (ToC) byte structure 401 for an IVAS frame.
[0184] The (F) bit 403 indicates whether more frames follow this entry. The F (1 bit) 403: If set to 1 , the bit indicates that the corresponding frame is followed by another speech or PI frame in this payload, implying that another ToC byte follows this entry. If set to 0, the bit indicates that this frame is the last frame in this payload and no further header entry follows this entry.
[0185] The Bitrate Level Indication (BLI) 409 describes the content of the associated frame, e.g., the bitrate of an IVAS frame. The Supplemental Signaling Bits (SSB) 405 can be used to describe various things, for example IVAS input format specific signaling. The BLI (4-5 bits) 409 can, for example, be configured to indicate the bitrate or other frame content indication for the frame. From the content indication, the receiver can determine the size of the received frame either directly from the bitrate or from pre-determined frame sizes, e.g., for the SPEECH_LOST and NO_DATA frames. The BLI (or in some implementations the Frame Type or FT) bits can indicate, for example, the bitrate of an IVAS speech frame, a SPEECH_LOST, NO_DATA or comfort noise (SID) frames. An example for BLI bit values (for 5 bits) are presented in the following table. At this point of the codec development, it is not decided if the aforementioned three types (SPEECH_LOST, NO_DATA, SID) are supported in the final codec. The SPEECH_LOST, NO_DATA and SID frames are part of the EVS codec and at it is likely that at least the SID and NO_DATA frame types are being incorporated into the IVAS specification. It is understood that the final IVAS codec specification, however, might have different frame types present than presented here.
[0186] In some embodiments, 4 bits are reserved for the BLI part. The frame type bit values and content indications when using 4 bits are presented below. In these embodiments the BLI part identifies the bitrate where the BLI bits are bit values other than a defined value (for example 1111 ) and can be used to identify other aspects, for example SPEECH_LOST, NO_DATA and SID frames with the combination of the BLI indicator OTHER (the defined value) and the extra bits.
[0187] As shown above there are 15 available bit allocations for the BLI bits reserved for future use (bit allocations 10001 - 11111 ), when 5 bits are reserved for the BLI indicator. In some embodiments when 4 bits are reserved for the BLI indicator, there are 5 available bit allocations for future use, when BLI value 1111 is used. BLI value 1110 provides an additional 8 bit allocations for future use when combined with the extra bits. These available bit allocations can be used in the future to indicate frame contents that will be defined, such as PI frames described hereafter.
[0188] These bit allocation values are examples only and it would be appreciated that in some embodiments the bit allocation values can be otherwise configured.
[0189] The number of bits in the SSB 405 can vary between 2-3 depending on how many bits are reserved for the BLI bits 409 in the IVAS ToC byte 401 . The SSB (2- 3 bits) 405 can be bits reserved for future use. If 4 bits are reserved for the BLI indicator, the extra 3 bits can be partly used to identify other frame content than bitrates / frame-size (e.g., SPEECH_LOST, NO_DATA, SID) as demonstrated in further detail in GB2313472.9, where the BLI bits are referred to as FT bits (frame type index) and SSB as extra bits.
[0190] With respect to the following table, there are presented example split rendering parameters that could be transmitted with an IVAS split rendering frame to be able to render the frame correctly at the receiver side.
[0191] In some embodiments, the parameters could differ based on the implementation choices. These parameters can comprise:
[0192] DOF: the DOF (degrees of freedom) parameter or field which can indicate the current degrees of freedom associated with the audio scene;
[0193] HQMODE: the HQMODE parameter can indicate whether a high-quality mode is being implemented;
[0194] CODEC: the CODEC parameter can indicate the audio codec employed in the current system; and
[0195] RENDERER: the RENDERER parameter can indicate the Tenderer mode employed in the current system.
[0196] In some embodiments the HQMODE parameter could be combined with the DOF parameter. These transmitted parameters can form a split rendering header that is mentioned in the later example embodiments.
[0197] The split rendering parameters described above are examples only and other parameters or a selection of the split rendering parameters can be implemented for other split rendering implementations. In some embodiments the split rendering parameters can be signaled as part of the split Tenderer bitstream.
[0198] The following examples and embodiments describe how the split rendering data (header + frame) could be integrated to an IVAS RTP payload. Figure 5 shows an example how the split rendering information header 501 could be incorporated directly into an IVAS ToC byte. The F-bit 503 would be similar to the F-bit 403 in a non-split rendering IVAS ToC byte such as described earlier with respect to the example shown in Figure 4. In some embodiments the SSB bits 405 can comprise one SSB bit reserved for indicating split rendering and non-split rendering frames, for example the first SSB bit 504. Alternatively, in some embodiments, a 2- or 3-bit SSB sequence could be reserved for indicating a split rendering frame, for example (10) or (110) SSB sequence.
[0199] The HQ bit 505 indicates the presence of HQ mode in the SR frame. For example the bit value 1 could be used to indicate that the HQ is present and a value of 0 is used to indicate that it is not present.
[0200] The DOF bits 511 in some embodiments can be used to indicate the degrees of freedom for the SR frame. In some embodiments the HQ and DOF bits can be merged into a 3-bit sequence, where one possible value is 3 DOF HQ (the HQ mode is currently only proposed to be used in 3 DOF mode).
[0201] The merging of 3 bits for DOF enables the four 3DOF operating points (including 3DoF HQ mode) and also future extension for 6DOF.
[0202] The BLI bits 409 indicate the bitrate for the SR frame. The following examples show possible bitrate values.
[0203] Currently, bitrates 256, 384, 512 and 798 kbps are expected to be proposed for an IVAS related split rendering. However, to support more bitrates more bits can be reserved from the SSB part 405. For example, in some embodiments only 1 bit is used for split rendering frame signaling.
[0204] In this example some parameters, the optional CODEC and RENDERER split rendering parameters, are not transmitted as part of the split rendering ToC header byte.
[0205] In some embodiments, only some of the split rendering parameters signaled to a receiver with the split rendering bitstream. Furthermore as discussed above the split rendering parameters could be labelled using different labels, different or combined parameters. In either case, in some cases, the split rendering parameters (or some of the parameters) can fit into the traditional (non-split rendering) ToC byte. For example, the split rendering data could fit into the 4 BLI bits. In this case, a 3-bit SSB sequence (110 for example) could be reserved for indicating split rendering frame, and the BLI bits could be used to carry the split rendering parameters to render the frame correctly in the receiver.
[0206] This approach would avoid sending a separate split rendering header with the parameters as they would be already carried with the (split rendering) IVAS ToC byte or the SR header could have less data if part of the data is already part of the ToC byte.
[0207] With respect to Figure 6 there is shown an example embodiment 601 showing how a split rendering (SR) frame is integrated to an IVAS RTP payload.
[0208] In this example the IVAS ToC byte 603 comprises a F bit 611 , which can be similar to the F bit examples as discussed above, and SSB bits 613 (for example with a value of 110 or any other available sequence) of the IVAS ToC byte configured indicate a split rendering payload. Additionally the BLI bits 615 in the IVAS ToC byte 603 can be configured to indicate the bitrate for the split rendering frame. The bitrate indication (BLI bits 615) could, in some embodiments, have values which are the same as those used for IVAS frames. For example the BLI bits could indicate a bitrate range of 13.2 - 512 kbps. In some embodiments the bitrate range supported by the indications could be a range of values specific for the split rendering frame (for example if the split rendering supports different bitrates).
[0209] The ToC byte 603, in some embodiments, is followed by a split rendering (SR) header 605. The split rendering header 605 in these embodiments contains the necessary information for processing the split rendering frame. The split rendering header 605 can therefore indicate one or more of the split rendering parameters, DOF; HQMODE; CODEC; and RENDERER. The size of the split rendering frame could, based on the implementation, be determined based on the bitrate indicated in the ToC byte or the size could be explicitly indicated in the SR header.
[0210] In some embodiments the split rendering (SR) header 605 is followed by a SR frame 607. The SR frame 607 comprises the SR frame data which is to be processed by the receiver / player device.
[0211] The example in Figure 7 shows an example embodiment 701 showing how a split rendering (SR) frame is integrated to an IVAS RTP payload. The example is similar to the example described with respect to Figure 6, but a split rendering payload is indicated with the combination of the SSB and BLI bits (e.g., a combination of SSB 101 and BLI 1110 or any other available sequences).
[0212] For example the example 701 shows a SR ToC byte 703 which comprises a F bit 711 , which can be similar to the F bit examples as discussed above. Furthermore as indicated above it is the combination of a specific value for the SSB bits 713 (for example with a value of 101 or any other available sequence) and BLI bits 715 (with an example value of 1110 or any suitable available value) in the SR ToC byte 703 which identifies that the payload is a split rendering payload. The combination sequence identifies a new frame type: split rendering payload (e.g., SPLIT_RENDERING) such as indicated by the table below.
[0213] The ToC (SR) byte 703, in some embodiments, is followed by a split rendering (SR) header 605. The split rendering header 605 in these embodiments contains the necessary information for processing the split rendering frame. The split rendering header 605 can therefore indicate one or more of the split rendering parameters, DOF; HQMODE; CODEC; and RENDERER. The size of the split rendering frame could, based on the implementation, be determined based on the bitrate explicitly indicated in the SR header. In some implementations, the SR frame carries the SR header information within the SR frame. Consequently, the SR frame can be self-contained in terms of information required for decoding and / or rendering with the specified degrees of freedom and other parameters (e.g., different codecs, operating modes, etc.) (during the split rendered frame creation).
[0214] In some embodiments the split rendering (SR) header 605 is followed by a SR frame 607. The SR frame 607 comprises the SR frame data which is to be processed by the receiver / player device.
[0215] Figure 11 a shows another example 1100 on how a split rendering frame could be integrated to an IVAS RTP payload. In this example, an IVAS ToC byte 1102 indicates a processing information (PI) frame 1104 which consequently indicates a split rendering payload. For example the PI frame type can be SPLIT_RENDING or another suitable label identifying the PI frame type is a split rendering frame. The PI frame 1104, which follows the IVAS ToC byte 1102 furthermore contains information for processing the split rendering frame, for example the split rendering parameters such as defined above (including bitrate).
[0216] In other words the SR header information as described above can be integrated to the PI frame 1104. In such a manner there is no requirement for defining or creating any additional header definitions (SR header) in the IVAS RTP payload description.
[0217] In some embodiments, a SR header could follow the PI frame indicating SR payload, the SR frame 1106. In these embodiments, the PI frame 1104 is configured to indicate the split rendering payload (e.g., with a SPLIT_RENDERING type) and the split rendering related information can then be identified through the SR header.
[0218] Figure 8 shows another example embodiment how a split rendering frame could be integrated to an IVAS RTP payload. The approach in this embodiment is similar to the example shown with respect to Figure 7 where the IVAS ToC byte indicates a split rendering payload, followed by a SR header and frame. In this example however the ToC byte follows the design presented in the EVS specification.
[0219] For example the example 801 shows a SR ToC byte 803 which comprises a header Type (H) identification bit 811 is used to differentiate between CMR and ToC bytes, the (F) bit 813 is used to indicate if another ToC byte follows this frame and the Frame type (FT) index 815 identifies a new frame type: split rendering payload (e.g., SPLIT_RENDERING), for example as shown in the following table.
[0220] The FT bits sequence could be reserved from any freely available sequence in the ToC byte (for example one of the values marked as “for future use”). In some embodiments the IVAS RTP payload format, the IVAS bitrates are identified with the EVS ToC byte with FT sequences 01 +XXXX, where the XXXX indicates any freely available bit sequence with respect to the bit sequence allocation is 3GPP SA4 specification TS 26.445, Annex specifying the EVS RTP payload format.
[0221] In some embodiments the IVAS split rendering payload 607 should be identified in the same FT section 815. For example sequences starting with 00, 10 or 11 are already reserved for EVS codec, thus a FT sequence identifying an IVAS SR payload could be, for example, reserved as 01 +1110 or any other freely available sequence. Figure 9 shows another example embodiment on how a split rendering frame could be integrated to an IVAS RTP payload. In this example 901 , the ToC byte 903 is configured according to the EVS style as described in Figure 2 and 3 and in the example presented previously with respect to Figure 8. For example the example 901 shows a SR ToC byte 903 which comprises a header Type (H) identification bit 911 used to differentiate between CMR and ToC bytes, the (F) bit 913 is used to indicate if another ToC byte follows this frame and the Frame type (FT) index 915 identifies a new frame type: extension frame or extension byte type (e.g., EXTENSION), for example using the value such as shown in the above table.
[0222] The SR ToC byte 903 is furthermore followed by an extension frame (SR) or split rendering extension frame 905. The extension frame 905 can be employed to indicate any frame content that is not able to fit in the ToC indication. For example the extension frame 905 can comprise parameters such as SID, SPEECH_LOST, NO_DATA, PI_FRAME, SPLIT-RENDERING), for example as shown by the following table:
[0223] The extension frame 905 can be a single byte or any suitable number of bits for indicating the extension content. The extension frame 905 can also indicate a split rendering payload, followed by a SR header 605 and SR frame 607 (the contents of which can be similar to those described above). In some embodiments the extension frame 905 is an IVAS specific addition to the payload header, and the FT sequence in the (EVS-like) ToC byte indicating the extension frame should ideally be identified in the same FT section as the other IVAS related contents (e.g., bitrate). In some embodiments, the extension frame 905 comprises bitrate information about the SR frame 607 (or payload). For example, the bitrate for the split rendering frame 607 could be indicated with the extension frame 905 in which case the split rendering header 605 is not configured to carry bitrate information about the SR frame 607.
[0224] With respect to Figure 10, there is shown some alternative embodiments based on the earlier presented examples. In this approach, the split rendering header is integrated to the SR frame. In other words there is no separate SR header in the payload.
[0225] Thus as shown in Figure 10 there can be a first example (similar to that shown in Figure 6) where there is a SR ToC 1001 with bitrate indication through ToC (BLI bits), SR indicated with SSB bits and a combined SR header and frame 1011 , a second example (similar to that shown in Figure 7) where there is a ToC indicates SR frame, bitrate indication in the SR header 1003 and a combined SR header and frame 1011 , a third example (similar to that shown in Figure 11 a) where there is a PI frame which indicates SR frame, and bitrate either in PI(SR) frame or in the SR header 1005 and a combined SR header and frame 1011 , a fourth example (similar to that shown in Figure 8) where there is a EVS style ToC indicating a SR frame (with FT bits in ToC), and with bitrate indication through SR header 1007 and a combined SR header and frame 1011 , and a fifth example (similar to that shown in Figure 9) where there is a EVS style ToC indicating an extension frame (with FT bits in ToC), with an extension frame indicates an SR frame, and bitrate indication in the SR header and a combined SR header and frame 1011.
[0226] The combined SR header and frame 1011 is one where the SR header is integrated as part of the SR frame. The header part comprises the SR information. The bitrate can be identified with the SR header or with the ToC byte. The SR frame furthermore comprises the split rendering data.
[0227] Figure 11 b furthermore shows another example 1101 how a split rendering frame could be integrated to an IVAS RTP payload. This approach is similar to the example shown in Figure 11 a, but the bitrate of the SR frame is indicated with a separate ToC byte. In this example, an IVAS ToC byte 1103 indicates a processing information (PI) frame which consequently indicates a split rendering payload. For example the PI frame type can be SPLIT_RENDING or another suitable label identifying the PI frame type is a split rendering frame.
[0228] The PI frame 1107 indicates SR payload and includes all the SR related information except the bitrate which is indicated with the separate ToC byte 1105. In this embodiment, the receiver of the payload detects two ToC bytes: one 1103 to indicate a PI frame and another 1105 to indicate a bitrate of some data frame. The PI frame 1107 is detected to be a split rendering PI frame, in which case the second ToC byte is determined to indicate the bitrate of the associated split rendering frame.
[0229] In some embodiments, a SR header could follow the PI frame 1107 indicating SR payload. In this case, the PI frame 1107 is configured to indicate the split rendering payload 1109 (e.g., with a SPLIT_RENDERING type) and the split rendering related information (except the bitrate) would be identified through the SR header.
[0230] Figure 12 shows an example scenario according to some embodiments where two UEs 110, 120 are communicating through an IVAS session. These examples focus on aspects where the encoder / sender UE 110 or UE A is sending IVAS frames 1201 to the decoder / receiver UE 120 or UE B. The PI forward frames 1203 can be interpreted as processing information data indication or indicators from the sender UE 1 10 to the receiver UE 120. The type of data within the PI forward frames 1203 can in some embodiments relate to the capturing device orientation and labelling the data, for example, as mentioned above and in further detail as described in GB2313472.9.
[0231] Split rendering has the benefit of dividing the processing load between the session parties. During an IVAS session, some party might prefer to start receiving a split rendered bitstream, for example, to lower the computational load. Thus as shown in Figure 12, PI request frames can be an action request transmitted from the receiver UE 120 to the sender UE 110. The action requests can be, for example, a request to start to receive a split rendered bitstream from the sender UE 110. This can for example be in the form as shown in the following table:
[0232] A request type of SPLIT-RENDERING could be used to indicate a request to start using split rendering in both receive and send directions. The request could be immediate, or the request could indicate timing information (e.g., a number of processing frames or seconds) after the rendering should be switched to split rendering.
[0233] In some embodiments the SPLIT_RENDERING request type could only indicate a request to start using split rendering stream from sender to receiver (i.e. , from A to B) or the other way around (from B to A). If the SPLIT_RENDERING request type would indicate a request for both receive and send directions, the individual directions could be requested with separate PI frame request types. SPLIT_RENDERING_SEND request type could be used to request to start using split rendering in the direction from UE 120 B to UE 110 A. SPLIT_RENDERING_RECV request type could be used to request to start using split rendering in the direction from A to B. The sender A can respond to split rendering request with a PI response frame or the sender A can start using split rendering which the receiver B notices from the received SR payloads.
[0234] Thus, with the above approaches, it is possible for the UE and an edge server to switch the responsibility of decoding the IVAS bitstream and construction the split Tenderer bitstream. For example, the receiving mobile device or UE can first decode the IVAS bitstream and construct a split Tenderer bitstream. Then, due to any reason, the receiving mobile device or UE may request the sending edge server to start providing the split Tenderer bitstream, e.g., after 50 IVAS frames (1 second). The request to switch the server to use split rendering may be transmitted, for example, with the SPLIT_RENDERING PI frame type described above. At the same time, mobile device or UE starts providing the head-orientation feedback data from the final UE, e.g., headphones, to the sending edge server to prepare it for the change. After 50 IVAS frames, edge server switches to transmitting split Tenderer bitstream and the receiving mobile device can adjust accordingly and pass the split Tenderer bitstream to the final UE. The transformation in this case will be almost seamless.
[0235] Furthermore, there can comprise PI feedback frames 1207, as shown in Figure 12 transmitted from the receiver UE 120 to the sender UE 110. The feedback frames can comprise information from the receiver UE 120 to the sender UE 110. The PI feedback frames 1207 can differ from the PI request frames 1205 in that the PI feedback frames 1207 do not cause the sender UE 110 to switch format but can for example comprise information relating to the receiver UE which the sender UE can choose whether or not to process. For example the feedback information can be the playback device orientation or the head rotation of the listener at the receiver UE which the sender UE 110 can, in some circumstances use to improve the encoding or processing of the capture audio signals.
[0236] In some embodiments, as shown in Figure 12, there can also be PI response frames 1209. The PI response frames 1209 can be PI frames in response to the PI request frames 1205. For example, if the receiver UE 120 is requesting the capture device orientation of the sender A, the sender UE 110 can send the orientation to the receiver UE 120 as a PI response frame 1209.
[0237] In some embodiments each request can be identified with an ID and the same ID can also be present in the response. Additionally in some embodiments the request is identified by a unique ID which is echoed by the same ID in the PI response frame 1209.
[0238] Split rendering can also have another meaning in an IVAS session: it can indicate a type of output format. For example, the receiver UE 120 B can be a mobile phone that is connected to wireless headphones. The connection from the phone to the wireless headphones may use a type of split rendering, where most of the processing is computed by the mobile phone. This split rendering connection can be informed to the sender UE 110 A as an output format for the receiver UE 120 B (i.e., the receiver UE is then configured to output split rendering data). The split rendering output format can be added as one of the IVAS supported output format SDP parameters.
[0239] In some embodiments, the output format or output format change to split rendering by the receiver UE 120 B can be informed to the sender UE 110 A with the PI feedback frames. For example using a OUTPUT_FORMAT and OUTPUT_FORMAT_CHANGE feedback PI frame type, which can be used to inform the used output format or to indicate an output format change from UE 120 B to UE 110 A. An indication for split rendering output could be added to these PI frame types.
[0240] Additionally the requests can indicate output format changes through SSB and BLI combinations in the IVAS ToC byte. In a similar approach, the split rendering output format change could also be added as one of the indicated output format change options with the SSB and BLI bits.
[0241] Figure 13, for example, shows flow diagrams of the session negotiation support for split rendering payload types using an SDP parameter split_rendering (or split_rend or any other suitable label or parameter value) according to some embodiments. The parameter in some embodiment can have permissive values of 0 (disable) and 1 (enable) for indicating the support for split rendering payload during the session in both send and receive directions.
[0242] Separate parameters, for example, split_rendering_send and split_rendering_recv could be added additionally to indicate split rendering payload support in the separate directions. These parameters can have similar permissive values, 0 (disable) and 1 (enable), to indicate the support for split rendering payload.
[0243] However, the SDP negotiation phase of split rendering payload support is not necessary to enable the SR payload support. All the described embodiments to enable split rendering payloads in an IVAS session are self-contained and do not require SDP negotiation to be recognized as split rendering payloads among IVAS RTP packets. With respect to the session negotiation creation process there is shown, by 1301 , of the operation of obtaining information if IVAS split rendering is supported by the sender device.
[0244] Following this, as shown by 1302, is including a parameter to enable split rendered payloads in the session description file.
[0245] Then is shown, by 1303, is generating session negotiation offer to a receiver user equipment.
[0246] With respect to the session negotiation answer process there is shown, by 1311 , of the operation of receiving a session negotiation offer.
[0247] Following this, as shown by 1312, is parsing IVAS split rendering indication in the received session negotiation offer.
[0248] Then is shown, by 1313, obtaining information if IVAS split rendering is supported by the receiver device.
[0249] After which is shown, by 1314, including a split rendering parameter indicating support for SR within a session negotiation answer session description file.
[0250] Then is shown, by 1315, is transmitting the session negotiation answer to the sender user equipment which delivered the session negotiation offer.
[0251] In some embodiments, in addition to indicating the capability and readiness to support split rendering, the session negotiation can also indicate the initiation possibility for the receiver UE, the sender UE or either of the UEs in a conversational immersive audio unidirectional or bidirectional transmission session.
[0252] In such embodiments the Media format parameter “sr” can be an additional parameter for IVAS. This indicates capability for split rendering of IVAS. The presence of sr=1 in the offer and answer indicates bidirectional support for split rendering of IVAS.
[0253] In some further implementations, there may be sr_send and sr_recv media format parameters which can also be present to indicate split rendering support in only one direction.
[0254] In some further implementation embodiments, there may be another switch to the split rendering media format parameter to indicate initiation capability. A value of 1 can indicate receiver UE initiation, 2 can indicate sender UE initiation; absence of any number or switch to the sr parameter can indicate that either of the UEs can initiate split rendering.
[0255] In some further embodiments, split rendering can be a separate attribute as a media line attribute, with the same switches as described above (for enabling / disabling and initiation role).
[0256] Figure 14 for example shows flow diagrams of the packetization and depacketization of the split rendered IVAS payload according to some embodiments.
[0257] With respect to the packetization of the split rendered IVAS payload there is shown, by 1401 , the operation of obtaining split rendered IVAS frame.
[0258] Following this, as shown by 1402, is generating split rendering header for the SR IVAS frame. Then is shown, by 1403, generating split rendering indication header and / or IVAS RTP payload header for the SR payload.
[0259] After which is shown, by 1404, appending the above SR IVAS frame and SR header to the above indication and / or IVAS payload header.
[0260] There follows, as shown by 1405, including the resulting bitstream comprising PI frame header and / or split rendering indication, SR header and IVAS SR frame in the RTP payload.
[0261] Then is the operation of, as shown by 1406, transmit RTP packet to the decoder / receiver user equipment B.
[0262] With respect to the depacketization of the split rendered IVAS payload there is shown, by 1411 , of the operation of receiving a RTP packet.
[0263] Following this, as shown by 1412, is extracting RTP payload and determining presence of split rendered IVAS payload in the RTP payload.
[0264] Then is shown, by 1413, extracting split rendering header and frame.
[0265] After which is shown, by 1414, delivering split rendered IVAS frames to the corresponding IVAS SR stream receiver user equipment processing pipeline.
[0266] Figures 15a to 15c show example end-to-end systems capable of implementing IVAS split rendering negotiation, initiation and indication in RTP data streams.
[0267] Figure 15a for example shows an IVAS end-to-end configuration where there is no split rendering.
[0268] The UE 110 comprises an IVAS encoder 1501 configured to encode the spatial audio captured signals, which are passed to an IVAS / RTP packetizer 1503.
[0269] The UE 1 110 and specifically the IVAS / RTP packetizer 1503 is configured to transmit the RTP stream 1583 to the internet 1581 which is then passed via the RTP stream 1585 to the UE 2 120.
[0270] Additionally there is a session negotiation SDP carries split rendering negotiation 1561 between the UE 1 110 and UE2 120.
[0271] The UE 2 comprises a IVAS RTP depacketizer 1551 configured to receive the RTP stream 1585 and obtain the IVAS data which is passed to an IVAS decoder 1553 which is then configured to output the audio signals.
[0272] Figure 15b for example shows an IVAS end-to-end configuration where the split rendering is initiated by the receiver UE 120 by sending an appropriate feedback message. The IVAS end to end scenario is shown with split rendering support. This message can be delivered in band with the IVAS audio bitstream in the payload or the RTP header extension or delivered as an RTCP feedback message.
[0273] The UE 110 comprises an IVAS encoder 1501 configured to encode the spatial audio captured signals, which are passed to an IVAS decoder 1511.
[0274] The IVAS decoder 1511 then is configured to pass the decoded signals to a IVAS split Tenderer 1512 which is configured to output the partially rendered audio to the RTP packetizer with SR support 1523.
[0275] The UE 1 110 and specifically the IVAS / RTP packetizer with SR support 1523 is configured to transmit the RTP stream 1583 to the internet 1581 which is then passed via the RTP stream 1585 to the UE 2 120.
[0276] Additionally there is a session negotiation SDP carries split rendering negotiation 1561 between the UE 1 110 and UE2 120.
[0277] Furthermore is shown the SR initialization request and orientation feedback 1571 between the UE 1 110 and UE2 120
[0278] The UE 2 comprises a IVAS RTP depacketizer 1561 configured to receive the RTP stream 1585 and obtain the split rendered IVAS data which is passed to an IVAS decoder 1553 which is then configured to output the audio signals.
[0279] Figure 15c for example shows an IVAS end to end scenario the split rendering initiated by the sender UE. This can be done by switching to split rendering format frame delivery. Upon receipt of the SR frames, the receiver UE initiates delivering the head orientation information to the sender UE.
[0280] The UE 110 comprises an IVAS encoder 1501 configured to encode the spatial audio captured signals, which are passed to an IVAS decoder 1511.
[0281] The IVAS decoder 1511 then is configured to pass the decoded signals to a IVAS split Tenderer 1512 which is configured to output the partially rendered audio to the RTP packetizer with SR support 1523.
[0282] The UE 1 110 and specifically the IVAS / RTP packetizer with SR support 1523 is configured to transmit the RTP stream 1583 to the internet 1581 which is then passed via the RTP stream 1585 to the UE 2 120.
[0283] Additionally there is a session negotiation SDP carries split rendering negotiation 1561 between the UE 1 110 and UE2 120. Furthermore is shown the SR orientation feedback 1573 between the UE 1 110 and UE2 120
[0284] The UE 2 comprises a IVAS RTP depacketizer 1561 configured to receive the RTP stream 1585 and obtain the IVAS data which is passed to an IVAS decoder 1553 which is then configured to output the audio signals.
[0285] With respect to Figure 16 an example electronic device is shown. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 1900 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.
[0286] In some embodiments the device 1900 comprises at least one processor or central processing unit 1907. The processor 1907 can be configured to execute various program codes such as the methods such as described herein.
[0287] In some embodiments the device 1900 comprises a memory 1911. In some embodiments the at least one processor 1907 is coupled to the memory 1911. The memory 1911 can be any suitable storage means. In some embodiments the memory 1911 comprises a program code section for storing program codes implementable upon the processor 1907. Furthermore in some embodiments the memory 1911 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1907 whenever needed via the memory-processor coupling.
[0288] In some embodiments the device 1900 comprises a user interface 1905. The user interface 1905 can be coupled in some embodiments to the processor 1907. In some embodiments the processor 1907 can control the operation of the user interface 1905 and receive inputs from the user interface 1905. In some embodiments the user interface 1905 can enable a user to input commands to the device 1900, for example via a keypad. In some embodiments the user interface 1905 can enable the user to obtain information from the device 1900. For example the user interface 1905 may comprise a display configured to display information from the device 1900 to the user. The user interface 1905 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1900 and further displaying information to the user of the device 1900.
[0289] In some embodiments the device 1900 comprises an input / output port 1909. The input / output port 1909 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1907 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
[0290] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802. X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA).
[0291] The transceiver input / output port 1909 may be configured to receive the signals and in some embodiments obtain the focus parameters as described herein.
[0292] In some embodiments the device 1900 may be employed to generate a suitable audio signal using the processor 1907 executing suitable code. The input / output port 1909 may be coupled to any suitable audio output for example to a multichannel speaker system and / or headphones (which may be a headtracked or a non-tracked headphones) or similar.
[0293] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0294] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
[0295] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
[0296] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0297] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0298] The foregoing description has provided by way of exemplary and nonlimiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.
[0299] 3GPP 3rdGeneration Partnership Project
[0300] AMR-WB IO Adaptive Multi Rate Wideband Inter-Operable
[0301] BLI Bitrate Level Indication
[0302] CMR Codec mode request
[0303] CNG Comfort Noise Generator
[0304] DTX Discontinuous Transmission
[0305] EVS Enhanced Voice Services
[0306] FOA First-Order Ambisonics
[0307] FT Frame type (index)
[0308] HOA2 2ndorder Higher-Order Ambisonics
[0309] HOA3 3rdorder Higher-Order Ambisonics
[0310] ISM Independent Streams with Metadata
[0311] (i.e. , type of Object-Based Audio)
[0312] IVAS Immersive Voice and Audio Services kbps kilobits per second
[0313] MASA Metadata-Assisted Spatial Audio
[0314] MC Multichannel
[0315] MCU Multipoint Conferencing Unit
[0316] OMASA Object-based audio with MASA (combined input format) OSBA Object-based audio with SBA (combined input format)
[0317] PI Processing information (audio)
[0318] PLC Packet Loss Concealment RTCP Real-Time Transport Control Protocol
[0319] RTP Real-Time Transport Protocol
[0320] SBA Scene-Based Audio
[0321] SDP Session Description Protocol SSB Supplemental Signaling Bits
[0322] ToC Table of Content
[0323] UE User equipment
Claims
CLAIMS:
1. An apparatus for transmission of conversational immersive audio, the apparatus comprising means configured to: receive audio data, during the transmission of conversational immersive audio; generate a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generate a header configured to identify the rendering payload; generate a conversational immersive audio payload comprising the header and the rendering payload; and packetize the conversational immersive audio payload for transmission.
2. The apparatus as claimed in claim 1 , wherein the means configured to receive audio data, during the transmission of conversational immersive audio is configured to receive an immersive audio bitstream.
3. The apparatus as claimed in claim 2, wherein the immersive audio bitstream is an IVAS frame.
4. The apparatus as claimed in any of claims 1 to 3, wherein the means configured to generate a rendering payload based on the audio data is configured to: decode immersive audio data from the immersive audio bitstream; render at least partially the decoded immersive audio data by generating a partially rendered interim format audio data suitable for further rendering; and encode the partially rendered interim format audio data to generate a rendering payload.
5. The apparatus as claimed in any of claims 1 to 4, wherein the header is an IVAS RTP payload header.
6. The apparatus as claimed in claim 5, wherein the header comprises a ToC byte defining that the payload comprises the partially rendered interim format audio data.
7. The apparatus as claimed in any of claims 1 to 6, wherein the means configured to packetize the conversational immersive audio payload for transmission as a real-time transport protocol packet payload.
8. The apparatus as claimed in any of claims 1 to 7, wherein the means is further configured to receive information identifying at least one further rendering external to the apparatus format, wherein the means configured to generate the header is configured to generate the header based on the information identifying the at least one further rendering external to the apparatus format.
9. The apparatus as claimed in any of claims 1 to 8, wherein the means is further configured to receive from the further apparatus information identifying the further apparatus is capable of further rendering external to the apparatus, wherein the means configured to generate the rendering payload based on the audio data is configured to generate the rendering payload as the initial rendering for a split rendering, based on the information identifying the further apparatus is capable of further rendering external to the apparatus.
10. The apparatus as claimed in any of claims 1 to 9, wherein the means is further configured to: receive from the further apparatus at least one of: listener position information; and listener orientation information; select a degree-of-freedom parameter based on the at least one of: listener position information; and listener orientation information, wherein the means configured to generate the header is configured to generate a header based on the selected degree-of-freedom parameter.
11. The apparatus as claimed in any of claims 1 to 10, wherein the means is further configured to negotiate with a further apparatus to at least one of: support for the conversational immersive audio payload; and negotiate employing a session description file.
12. An apparatus for rendering of conversational immersive audio, the apparatus comprising means configured to: receive a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determine based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
13. The apparatus as claimed in claim 12, wherein the means configured to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering is further configured to transmit the rendering payload for further rendering to a further apparatus.
14. The apparatus as claimed in claim 13, wherein the means configured to transmit the rendering payload for further rendering to the further apparatus is configured to transmit the rendering payload for further rendering to the further apparatus based on at least one of: a determination that the apparatus is unable to perform the further rendering of the rendering payload comprising the initial rendering; and a selection for transmitting the rendering payload for further rendering to the further apparatus.
15. The apparatus as claimed in any of claims 12 to 14, the means configured to receive a packetized conversational immersive audio payload; wherein the conversational immersive audio payload comprises a header and a payload is configured to: receive at least one real-time transport protocol packet;parse the at least one real-time transport protocol packet to obtain the packetized conversational immersive audio payload.
16. The apparatus as claimed in claim 15, wherein the at least one real-time transport protocol packet comprises a real-time transport protocol payload.
17. The apparatus as claimed in any of claims 12 to 16, wherein the header indicates that the payload comprises the initial rendering.
18. The apparatus as claimed in any of claims 12 to 16, wherein the means configured to provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering is further configured to: decode the rendering payload based on the header; and further render the rendering payload.
19. The apparatus as claimed in any of claims 12 to 18, wherein the means is further configured to generate and transmit to an additional apparatus information identifying at least one further rendering format for the apparatus, wherein the received packetized conversational immersive audio payload is based on the information identifying the at least one further rendering format.
20. The apparatus as claimed in claim 19, wherein the means is further configured to at least one of: generate and transmit to an additional apparatus information identifying the apparatus is capable of further rendering; and obtain and transmit to the additional apparatus at least one of: listener position information; and listener orientation information.21 . The apparatus as claimed in any claim 19 or 20, wherein the means is further configured to negotiate with the additional apparatus to at least one of: support for further rendering; and negotiate employing a session description file.
22. A method for an apparatus for transmission of conversational immersive audio, the method comprising: receiving audio data, during the transmission of conversational immersive audio; generating a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus; generating a header configured to identify the rendering payload; generate a conversational immersive audio payload comprising the header and the rendering payload; and packetizing the conversational immersive audio payload for transmission.
23. A method for an apparatus for rendering of conversational immersive audio, the method comprising: receiving a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determining based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and providing the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.
24. An apparatus for transmission of conversational immersive audio, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive audio data, during the transmission of conversational immersive audio; generate a rendering payload based on the audio data, wherein the rendering payload comprises an initial rendering inside the apparatus, and the rendering payload is configured to be used for further rendering external to the apparatus;generate a header configured to identify the rendering payload; generate a conversational immersive audio payload comprising the header and the rendering payload; and packetize the conversational immersive audio payload for transmission.
25. An apparatus for rendering of conversational immersive audio, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive a packetized conversational immersive audio payload, wherein the conversational immersive audio payload comprises a header and a payload; determine based on the header, if the payload comprises a rendering payload, wherein the rendering payload comprises an initial rendering; and provide the rendering payload for further rendering, based on the header indicating the rendering payload comprises the initial rendering.