Backwards compatible RTP payload format, encoding and decoding for split rendering support in immersive audio
Patent Information
- Application Number
- AU2025226880
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-26
- Filing Date
- 2025-02-27
- Publication Date
- 2026-08-27
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 763,901, filed February 26, 2025, U.S. Provisional Application No. 63 / 713,709, filed October 30, 2024, U.S. Provisional Application No. 63 / 713,545, filed October 29, 2024, U.S. Provisional Application No. 63 / 559,132, filed February 28, 2024, and U.S. Provisional Application No. 63 / 558,651, filed February 28, 2024, the entire contents of each of which is hereby incorporated by reference. TECHNICAL FIELD
[0002] This application relates generally to audio processing, and more specifically to audio coding, transmission, and rendering performed between two or more separate devices (“split rendering”). BACKGROUND
[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.
[0004] Voice and video encoder / decoder (codec) standards are recently focused on developing codecs for immersive audio services, such as for immersive voice and audio services (IVAS). The immersive audio services are expected to support a range of audio service capabilities, including but not limited to mono to stereo upmixing and fully immersive audio encoding, decoding, and rendering.
[0005] A wide range of devices, including endpoints and network nodes, are expected to support immersive audio services (e.g., IVAS). Some example devices include, but are not limited to, mobile devices such as smart phones and electronic tablets, personal computers, video and audio conferencing, home theater devices, television and video gaming devices, extended reality (XR) devices, and other suitable devices. These devices, end points, and network nodes may have various acoustic interfaces for sound capture, transmission, relay, and rendering.
[0006] Extended reality (XR) applications, such as augmented reality (AR), mixed reality (MR), and virtual reality (VR), may be implemented for use on any variety of physical devices, where such XR applications support immersive audio. More specifically, XR devices may leverage immersive audio services to enhance the user experience with a fully immersive interactive experience. In such XR applications, the audio rendering may be adjusted responsive to the motion of the user. For example, a user’s head position and / or head movement may be tracked, and the audio rendering may be adjusted in response to the user’s tracked motion. Thus, an immersive audio experience may process head movements using models with three degrees of freedom (3DoF) or six degrees of freedom (6DoF).
[0007] Various immersive audio services, such as IVAS, may be used to render high quality audio renditions at the XR device that include awareness of user position or pose information. However, making such adjustments according to pose information may require significant computational processing capabilities to achieve a high-quality immersive audio experience.
[0008] It is with respect to these and other considerations that the disclosure made herein is presented. BRIEF SUMMARY OF THE DISCLOSURE
[0009] Techniques for coding, signaling, and decoding audio signals are described herein. In some examples, these techniques may be applied to split-rendering of audio between an upstream device and a downstream device, such as for an IVAS split-rendering topology.
[0010] Briefly stated, systems and methods for adapting an immersive audio pay load for RTP transport are disclosed. One example provides a method for adapting an immersive audio payload for RTP transport between a near-end device and an end device. The example method includes receiving split-rendering control information from the end device over the communication link and adapting a rendering operation or mode of the near-end device based on the received split-rendering control information. The example method includes preparing, responsive to the received splitrendering control information, a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device and the end device. The first RTP payload is associated with a split rendering mode for the end device. The example method includes transmitting the first RTP payload to the end device through the communication link.
[0011] In some described embodiments, a method for adapting an immersive audio payload for RTP transport between a near-end device and an end device is described, the method for the near-end device including: receiving split-rendering control information from the end device (101) over the communication link (111); adapting a rendering operation or mode of the near-end device (102) based on the received split-rendering control information; preparing, responsive to the received split-rendering control information, a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device (102) and the end device (101), wherein the first RTP payload is associated with a split rendering mode for the end device (101); and transmitting the first RTP pay load to the end device (101) through the communication link (111). In some alternatives, the first RTP payload with audio data corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.”
[0012] In some embodiments, a method for processing an immersive audio payload for RTP transport received from an upstream device by a downstream device is described, the method for the downstream device including: receiving a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device by a communication link, wherein the first RTP payload is associated with a split rendering mode for the upstream device and the downstream device; applying a first split rendering operating mode of the downstream device based on the first RTP payload from the upstream device; and transmitting splitrendering control information from the downstream device in a reverse direction over the communication link to the upstream device. In some alternatives, the first RTP payload with audio data corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.”
[0013] In some embodiments, a method for generating an immersive audio payload for RTP transport is described, the method including: generating a ToC byte for inclusion in a first RTP payload; setting a bit rate (BR) field included in the ToC byte to indicate that the first RTP payload is a split rendering payload; generating a split rendering ToC byte (SR ToC) for inclusion in the first RTP payload sequential to the ToC byte; setting a split rendering bit rate (SR-BR) field included in the SR ToC byte to indicate an IVAS split rendering bit rate; and preparing the first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between an upstream device and a downstream device. In some alternatives, the first RTP payload with audio data corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata”.
[0014] In some embodiments, a method for RTP transport for an adaptive immersive audio pay load using a common payload format is described, the method including: preparing a first RTP payload with audio data for at least one stream of coded immersive audio for communication between an upstream device, a first downstream device, and a second downstream device, wherein the first RTP payload is associated with a first immersive audio coding mode; transmitting the first RTP payload to the second downstream device through a first communication link, processing the first RTP payload by the second downstream device to generate an adapted first RTP payload, wherein the first RTP payload is decoded according to the first immersive audio coding mode, pre-rendered and re-encoded according to a split rendering mode to a split rendering representation comprising coded binaural audio and related split rendering metadata; preparing the adapted first RTP payload for RTP transport for communication between the second downstream device and the first downstream device using the same pay load format, wherein the adapted first RTP pay load is associated with the used split rendering coding mode, transmitting the adapted first RTP payload to the first downstream device through a second communication link; and receiving and processing the adapted first RTP payload by the first downstream device, wherein the adapted first payload is decoded and postrendered according to the associated split rendering mode. In some alternatives, the first RTP payload is decoded, pre-rendered and re-encoded according to a split rendering mode to a split rendering representation comprising coded two channel audio and related split rendering metadata.
[0015] In some embodiments, an IVAS configuration request (IVAS-cr) byte is encoded in the first RTP payload to signal an audio bandwidth request and an audio format request. In some embodiments, an extra (E) byte is encoded in the first RTP payload to signal the presence of processing information for rendering the at least one stream of coded binaural audio. Other bytes may also be included within the RTP payload for controlling rendering of audio and / or setting control operations for split rendering modes. In some alternatives, the extra byte (E) is encoded in the first RTP payload to signal the presence of processing information for rendering the at least one stream of coded two channel audio.
[0016] Various aspects of the present disclosure provide for processing of audio signals, and effect improvements in at least the technical fields of audio processing, audio encoding, audio decoding, virtual reality, and the like.
[0017] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.
[0018] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims. DESCRIPTION OF THE DRAWINGS
[0019] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:
[0020] FIG. 1 is a block diagram of an example IVAS RTP framework that leverages a real-time transport protocol (RTP) to implement split rendering capabilities in which various aspects of the present disclosure can be practiced.
[0021] FIG. 2 illustrates a table providing the structures of Table of Contents (ToC) bytes indicating PI data blocks according to some aspects of the present disclosure.
[0022] FIG. 3 illustrates a table providing PI types for forward direction signaling according to some aspects of the present disclosure.
[0023] FIG. 4 illustrates a table for providing PI types for reverse direction signaling according to some aspects of the present disclosure.
[0024] FIG. 5 illustrates a table for providing OUTF bits and indicated IVAS output formats according to some aspects of the present disclosure.
[0025] FIG. 6 illustrates a table for metadata types specified in the metadata task force and estimated PI data frame sizes in bytes according to some aspects of the present disclosure.
[0026] FIG. 7 illustrates an example PI data block included in an IVAS RTP payload according to some aspects of the present disclosure.
[0027] FIG. 8 illustrates an example RTP header with an IVAS payload according to some aspects of the present disclosure.
[0028] FIG. 9 illustrates an example IVAS payload header according to some aspects of the present disclosure.
[0029] FIG. 10 illustrates a ToC byte for an IVAS frame and a corresponding bit rate index table according to some aspects of the present disclosure.
[0030] FIG. 11 illustrates an SR ToC byte and a corresponding split rendering bit rate table according to some aspects of the present disclosure.
[0031] FIG. 12 illustrates a codec mode request (CMR) byte and a corresponding CMR coding table according to some aspects of the present disclosure.
[0032] FIG. 13 illustrates an example IVAS-cr byte, a corresponding first IVAS-cr coding table, and a corresponding second IVAS-cr coding table according to some aspects of the present disclosure.
[0033] FIG. 14 illustrates an example extra (E) byte according to some aspects of the present disclosure.
[0034] FIG. 15 illustrates an example initial E byte according to some aspects of the present disclosure.
[0035] FIG. 16 illustrates a subsequent E byte for a bandwidth request, a first subsequent E byte coding table, and a second subsequent E byte coding table according to some aspects of the present disclosure.
[0036] FIG. 17 illustrates a subsequent E byte for a CMR and a subsequent E byte coding table according to some aspects of the present disclosure.
[0037] FIG. 18 illustrates a subsequent E byte for indicating PI data presence in the payload according to some aspects of the present disclosure.
[0038] FIG. 19 illustrates a subsequent E byte for an SR configuration request according to some aspects of the present disclosure.
[0039] FIG. 20 illustrates a block diagram of various example methods for adapting an immersive audio payload for RTP transport between an upstream device and a downstream device, according to some aspects of the present disclosure.
[0040] FIG. 21 illustrates a block diagram of various example methods for processing an immersive audio payload for RTP transport received from an upstream device by a downstream device, according to some aspects of the present disclosure.
[0041] FIG. 22 illustrates a block diagram of various example methods for generating an immersive audio payload for RTP transport, according to some aspects of the present disclosure.
[0042] FIG. 23 illustrates a block diagram of various example methods for generating an immersive audio payload using a common payload format, according to some aspects of the present disclosure.
[0043] FIGS. 24A-24J illustrate various example pay load structures according to some aspects of the present disclosure.
[0044] FIG. 25A illustrates a schematic block diagram of an example device architecture (e.g., an apparatus) that may be used to implement various aspects of the present disclosure.
[0045] FIG. 25B illustrates a schematic block diagram of an example CPU implemented in the device architecture of FIG. 25A that may be used to implement various aspects of the present disclosure. DETAILED DESCRIPTION
[0046] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments with reference to the accompanying drawings. The illustrative embodiments in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes made, without departing from the spirit or scope of the present disclosure. In light of the present disclosure, it will be apparent to one of ordinary skill in the art that the various described features and implementations may be practiced without many of these specific details. In some instances, well-known methods, procedures, components, and circuits, have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described hereafter that can each be used independently of one another or with any combination of other features. Thus, the features may be arranged, substituted, combined, separated, or designed into other configurations, which is contemplated in light of the present disclosure.
[0047] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0048] Throughout this disclosure, various terms are used to describe upstream and downstream devices that may cooperatively provide a split-rendering solution for immersive audio experiences. A typical implementation for split rendering solution may include an upstream device that may be considered a “heavyweight processing device”, while the downstream device may be considered a “lightweight processing device.” The terms “heavyweight” and “lightweight” may take on a number of meanings based on the context presented. In some examples, a “lightweight processing device" or “heavyweight processing device” may refer to the physical weight of the device. In other examples, a “lightweight processing device" and “heavyweight processing device” may refer to the form factor of being portable versus less portable in terms of size. In still other examples, a “lightweight processing device" and “heavyweight processing device” may refer to different battery or power requirements, while in other examples the distinctions are based on computational speed and / or complexity.
[0049] Various Acronyms that may appear throughout this disclosure and in the associated claims and / or drawings are listed below. Other commonly used acronyms and terms of art may be excluded from this list in the interest of brevity. Thus, a short list of acronyms is provided below as an easy reference for the reader. IVAS - Immersive Voice and Audio Services ISAR - Immersive Audio for Split Rendering Scenarios BS - Bitstream EVS - Enhanced Voice Services RTP - Real-Time Transport Protocol UE - User Equipment SR - Split Rendering PI - Processing Information ToC - Table of Contents WB - Wideband SWB - Super-Wideband FB - Fullband SBA - Scene-Based Audio ISM - Independent Streams with Metadata DTX - Discontinuous Transmission CMR - Codec Mode Request CONV - Convention INTRP - Interpolation AE - Acoustic Environment MASA - Metadata-Assisted Spatial Audio FMT - Format
[0050] An immersive voice and audio services (IVAS) system is expected to support a range of audio service capabilities, including but not limited to fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including but not limited to mobile and smart phones, electronic tablets, personal computers, conference phones, conference rooms, extended reality (XR) devices such as virtual reality (VR) and augmented reality (AR) devices, home theatre devices, and other suitable devices. In addition to support for a large number of different devices, IVAS is also expected to support a broad and varied set of communication topologies, including but not limited to a mix of wired communications and wireless communications, where the wireless communications may include a mix of cellular, Wi-Fi, and Bluetooth technologies. The real-time transport protocol (RTP) for IVAS includes support for the broad variety of devices and connectivity options.
[0051] IVAS systems described herein may include one to four independent (mono) streams of audio with metadata (ISM) or may include stereo audio (including binaural audio). IVAS systems may process wideband (WB), super-wideband (SWB), and fullband (FB) audio bandwidths at 16 kHz, 32 kHz, and 48 kHz sampling rates, respectively. IVAS systems may be capable of multichannel audio in 5.1, 7.1, 5.1+2, 5.1+4, and 7.1+4 configurations. IVAS systems may support scene-based audio (SBA), such as Ambisonics, up to order 3. IVAS systems may also support metadata-assisted spatial audio (MASA), as well as combinations of ISM and MASA and combinations of ISM and SBA. IVAS system bitrates may range from 13.2 kbps to 512 kbps.
[0052] Examples described herein may also be implemented using Enhanced Voice Services (EVS) systems. EVS systems may support all EVS operation modes (e.g., mono modes) of the IVAS codec, including the EVS primary and AMR-WB IO modes using a payload syntax compatible to the header-full format. EVS systems may process WB, SWB, and FB audio bandwidths at 8 kHz, 16 kHz, 32 kHz, and 48 kHz sampling rates, respectively. EVS system bitrates may range from 5.9 to 128 kbps. EVS systems may operate using a 20 ms frame duration and handle multiple frames per RTP payload. EVS systems may be capable of discontinuous transmission (DTX), transmission of Processing Information, and switching between EVS and IVAS operations in the same pay load type.
[0053] FIG. 1 is a block diagram of an example IVAS RTP framework 100 that leverages a realtime transport protocol (RTP) to implement split rendering capabilities. The example IVAS RTP framework includes three devices: end device 101, near-end user equipment 102, and far-end user equipment 103. End device 101 is an end device where RTP communications reach an end point via communication link 111, such as for a user operated mobile device. Near-end user equipment 102 is a near-end device, such as a network node, which is communicatively interposed between end device 101 and far-end user equipment 103, where near-end user equipment 102 is communicatively linked to both end device 101 and far-end user equipment 103 via communication links 111 and 112, respectively. Far-end user equipment 103 is a far-end device that is communicatively coupled to near-end user equipment 102 via a communication link 112. End device 101 and near-end user equipment 102 are considered to be a downstream device, while far-end user equipment 103 is considered upstream in the network. FIG. 1 illustrates merely one example RTP communication scenario, and others are also contemplated.
[0054] In the example illustrated in FIG. 1, an IVAS communication session may be established between two devices (e.g., 5G user equipment or UE) such as end device 101 and far-end user equipment 103. Far-end user equipment 103 is a far-end device, which is at a remote location with respect to the end device 101 which is an end device. Near-end user equipment 102, the near-end device, has a connection established to an end device. In many cases, the end device 101 is a lightweight device such as earbuds or AR glasses, which is too constrained (e.g., limited in power, heat dissipation, processing capabilities, weight, etc.) to run a full IVAS decoder with immersive binaural rendering. For such constrained cases, a split rendering session may be established between the two devices so that the processing capabilities from outside of the end-device may be leveraged. For example, the far-end user equipment 103 may include a full IVAS decoder, while the end device 101 may include a limited IVAS decoder. IVAS RTP packets may be communicated between the far-end user equipment 103 and the near-end user equipment 102 via communication link 112, where the near-end user equipment 102 may provide pre-rendering before the RTP packets are communicated to the end device 101 via communication link 111. In such a way, rendering and pre-rendering capabilities of the upstream devices are leveraged by the end device 101 to achieve a split-rendered solution.
[0055] The communication links between the various devices may be established through a variety of means, including but not limited to Ethernet, WiFi, 5G, to name a few. For the purposes of the below discussion, the communication links will be referred to simply as a link or session, which may use IP / UDP / RTP packets (IVAS RTP packets). The near-end user equipment 102 receives IVAS RTP packets from the far-end user equipment 103, which contains IVAS coded audio frames. These incoming packets are de-packetized / de-multiplexed by the near-end user equipment 102. A portion of the pay load is then fed to an IVAS decoder / split rendering pre-renderer and encoder instance at the near-end user equipment 102. Another portion of the pay load corresponds to control information for the IVAS encoder, to control the operation in the end device 101. The SR prerenderer and encoder instance at near-end user equipment 102 transcodes the incoming IVAS audio frames to the SR immediate audio representation comprising coded binaural audio and SR metadata. Frames of this SR intermediate audio representation are then re-packetized together by the near-end user equipment 102 with the control information for the IVAS encoder operation in the end device 101, whereby the latter may have been modified by the near-end user equipment 102. Such re-packetized packets are then transmitted by the near-end user equipment 102 to the end device 101. In parallel to that operation, the near-end user equipment 102 receives IVAS RTP packets from the end device 101. The near-end user equipment 102 de-packetizes / de-multiplexes these received packets, whereby some portions of the packet payload may be information controlling the operation of the SR pre-renderer and decoder of the near-end user equipment 102, which is routed and used accordingly. Other portions of the received packets, which may be IVAS encoded audio data produced by the end device 101 are re-packetized together with control information for the far-end IVAS encoder operation by near-end user equipment 102. Such packets are then sent upstream to the far-end user equipment 103.
[0056] Operations at the end device 101 may include receiving IVAS RTP packets from the nearend user equipment 102, depacketizing the received packets, and routing a portion of the received payload to the IVAS SR decoder and post-renderer instance on the end device 101. Another portion of the received packets may be control information for the IVAS encoder operation at the end device 101, which is routed and used accordingly. In parallel, audio may be captured by the end device 101, and encoded in frames using an IVAS encoding mode (e.g. EVS AMRWB-IO) selected according to the received control information. The coded audio data is then packetized together with information controlling the operation of the SR pre-renderer and decoder, and sent upstream from the end device 101 to the near-end user equipment 102.
[0057] In the following, a more detailed functional description is provided of the packetization and de-packetization operations at the lightweight device or end device 101, and the near-end user equipment 102, which can be carried out efficiently using packetization conventions of the suggested RTP payload format. Processing Information
[0058] Processing Information (PI) is the mechanism of payload formats described herein to carry additional information that is consumed by the IVAS receiver in addition to the IVAS frame data. PI frame data is not required for the IVAS decoding process but provides additional information that may be used for rendering the audio signal. Details regarding PI payloads may be found in Annex A of the 3GPP TS 26.253 standard.
[0059] FIG. 2 illustrates a table 200 providing the structures of Table of Contents (ToC) bytes indicating PI data blocks. PI data blocks may be indicated with the IVAS ToC byte with SSB sequence 110. The three bits after the SSB indicate the PI frame type. Pre-defined PI frame types are indicated with bit sequences 000 through 111. Sequence Illis reserved for PI frames where the type of the PI frame is indicated in the PI header. The last bit in the ToC byte acts as a “pre-defined size flag”. The flag indicates whether the size of the PI frame is known beforehand (e.g., stored in memory) (1) or if the size is explicitly indicated in the PI header (0).
[0060] In some instances, the PI header precedes the PI data frame. The PI header includes a PI type field (8 bits) and a size field (8 bits). The type field indicates the type for the PI data frame. The size field indicates the size for the PI data frame section in bytes. A size of zero is valid. For example, a PI data frame may have a size of zero if the type of the PI frame is enough to interpret the request (for example, VOICE_INDICATION).
[0061] If the type of the PI data frame is known, the PI header contains only the size indication (8 bits), where the size field indicates the size for the PI Data frame section in bytes. If the size of the PI data block is pre-defined, the PI header contains only the type indication (8 bits). In this case, the pre-defined size indicates the size of the whole PI data block including the PI header and data frame. A PI data frame may be valid at the receiver until another PI data frame of the same type is received.
[0062] FIG. 3 illustrates a table 300 providing PI types for forward direction signaling. FIG. 4 illustrates a table 400 for providing PI types for reverse direction signaling.
[0063] Each orientation PI data frame may be preceded by an orientation header. Each orientation header includes a 3-bits long convention field (CONV) to signal the used convention in the PI data section. For example, the CONV bits 000 may correspond to the orientation convention Quaternions, 001 may correspond to Yaw, 010 may correspond to Yaw + Pitch, 011 may correspond to Yaw + Pitch + Roll, 100 may correspond to Azimuth, 101 may correspond to Elevation, and 110 may correspond to Azimuth + Elevation. The full structure of the orientation header may depend on the orientation type.
[0064] Orientation types supported herein may include scene orientation and device orientation. For feedback signaling, the playback device orientation and the listener head orientation are supported. The scene orientation PI data frames may be used to transmit the scene orientation of the sending side. Additionally, information on when to apply the scene orientation may be transmitted in a form 13 of interpolation frame count. The count indicates the number of audio frames (20 ms) after the associated scene orientation is valid. The interpolation frame count may be indicated with three bits (INTRP) in the scene orientation header. The scene orientation header may be zero-padded to force byte-alignment. An INTRP bits value of 000 may indicate zero interpolation frames (E.g., the scene orientation is valid immediately). An INTRP bits value of 001 may indicate 50 interpolation frames, corresponding to one second of audio. An INTRP bits value of 010 may indicate 100 interpolation frames, an INTRP bits value of 011 may indicate 150 interpolation frames, an INTRP bits value of 100 may indicate 200 interpolation frames, an INTRP bits value of 101 may indicate 300 interpolation frames, an INTRP bits value of 110 may indicate 400 interpolation frames, and an INTRP bits value of 111 may indicate 500 interpolation frames.
[0065] The device orientation PI data frames may be used to transmit the orientation of some device on the sender side, such as the orientation of an audio capturing device. The orientation of a capturing device can be used to stabilize the captured spatial audio content by, for example, removing undesirable orientation changes. The device orientation header includes a flag (C) to indicate if the orientation is already compensated in the transmitted audio (C==l) or is not already compensated in the transmitted audio (C==0). The device orientation header may be zero-padded to force byte alignment.
[0066] For feedback signaling, the playback device orientation and the listener head orientation may both be supported. With reference to table 400 of FIG. 4, the PLAYBACK_DEVICE_ORIENTATION describes the orientation of the device used for playback of the received IVAS stream. The HEAD_ORIENTATION describes the head orientation of the listener. The feedback orientations may be preceded by a header including the 3-bits long CONV field.
[0067] PI orientation data may be structured as quaternions, Euler angles, or spherical coordinates. In quaternions, each angle may have 16 bits reserved for the value. The quaternion component values range from -1 to 1, and with 116 bits values, the resolution for a single component is 2 * (216-l)'1 « 0.00003. In Euler angles, the angles yaw, pitch, and roll are present with 16 bits reserved for the value of each angle. The angles may range from -180° (exclusive) to 180°, and with 16 bits values, the angle resolution is 360° * 2'16 « 0.0055°. In spherical coordinates (azimuth and elevation), each angle has 16 bits reserved for the value. Azimuth angles range from -180° to 180° and the elevation angles range from -90° to 90°. With 16 bits values, the angle resolution is 360° * 2'16 « 0.0055° and 181° * 2'16 « 0.0028° for elevation angles.
[0068] The size of the PI data frame may be explicitly stated in the PI header. This allows the use of dynamic sized PI data frames and a dynamic resolution for orientations, for example. For orientation PI data frames, the size of the PI orientation data section may be used to determine the size for each individual angle component. For example, if the size of the PI orientation frame is 96 bits (not including the PI header and the orientation header), the bits are divided equally to the angle representations. For example, with all the Euler angles present, 96 / 3 = 32 bits would be assigned for each component. If only the yaw and pitch would be present, those components would be assigned 96 / 2 = 48 bits each. The same bit division policy may apply to all supported orientation conventions. The size of the PI frame data section may be evenly divisible by the number of orientation components present in the PI data frame.
[0069] Default orientation PI data frames (scene or device) with known sizes and type indicated in the IVAS ToC byte (sequences 110 + 000+ 1 SCENE_ORIENTATION and 110 + 001 + 1 DEVICE_ORIENTATION as presented in table 200 and table 300) may use quaternion convention for the orientation data. The default size for scene and device orientations may be 72 bits with an 8 bit orientation header (with CONV set to 000) followed by a 64 bits orientation in quaternions where 16 bits are reserved for each component.
[0070] Acoustic environment (AE) PI data frames may be used to transmit acoustic environment data. The AE data may be transmitted as full or compact packets, where the compact packets contain only a low-resolution version of the AE data. Furthermore, the full AE packets can be divided into smaller data fragments which can be transmitted individually. By combining all the fragmented AE packets, the receiver can reconstruct the full acoustic environment. A compact acoustic environment PI data frame may consist of an identification field (ID, 7 bits), an RT60 field (15 bits), and a DSR field (18 bits) summing up to 40 bits size. The RT60 field may indicate the time that it takes for reflections to reduce 6 OdB in energy level, per frequency band. The DSR field may indicate a diffuse to source signal energy ratio, per frequency band. The compact AE PI data frame may be used when the AE PI type is indicated in the IVAS ToC (sequence 110 + 010 + 1, ACOUSTIC_ENVIRONEMENT_COMPACT) and the size is pre-defined. ACOUSTIC_ENVIRONMENT_FULL may describe the full acoustic environment including the ID, a frequency grid, the RT60 field, the DSR field, a pre-delay, a room size, absorption coefficients, and a listener location, totaling over 500 bits. Pre-delay may indicate a delay at which the computation of DSR values was performed. Early reflection parameters may include 3D rectangular virtual room dimensions and broadband energy absorption coefficient per wall. ACOUSTIC_ENVIRONMENT_FRAGMENT contains a fragment of the full AE.
[0071] LABEL_ID PI data frames may be used to label (or identify) the transmitted audio content with an 8-bit ID field. For example, 0000 0000 may correspond to Speech, 0000 0001 may correspond to Music, and 0000 0010 may correspond to Ambience. The label ID PI data frame may be used when the label PI type is indicated in the IVAS ToC (sequence 110 + 011 + 1, LABEL_ID) and the size is pre-defined. Label data may also be transmitted as user defined labels. For example, a LABEL_UTF8 PI data frame includes four characters. Each character is encoded with UTF-8 coding where 8 bits are used per character. Multiple labels may be transmitted with a single LABEL_UTF8 PI data frame by separating the labels in the data section with a special character (e.g., a null or tab ACII character), allowing for sending labels for different, for example, ISM objects.
[0072] VOICE_INDICATION PI data frames may be used to indicate if the transmitted audio contains voice data. When the voice indication PI type is indicated in the IVAS ToC (sequence 110 + 100 + 1, VOICE_INDICATION) and the size is pre-defined, no extra PI data section is transmitted. The voice indication is recognized from the PI type alone.
[0073] VOICE_INDICATION_ISM PI data frames (1 byte) may be used to indicate voice for individual ISM objects. Each bit at the beginning of the ISM voice indication byte represents a single audio object. A bit value of 1 indicates voice data for the associated audio object. The last four bits are used for zero padding to force the PI data frame to be byte-aligned. With VOICE_INDICATION_ISM, the size of the PI data frame is known (1 byte) even if the pre-defined size flag is set to 1 in the ToC byte.
[0074] COMMON_AUDIO_SCENE_ISM PI data frames may be used to specify which audio objects originate from a common audio scene. For example, with multiple ISM objects, some objects might be of virtual origin and not part of a common scene shared by the rest of the ISM objects. Each bit at the beginning of the common audio scene indication byte represents a single audio object. All objects marked as 1 indicate that they originate from the same audio scene. Objects marked as 0 indicate that they do not originate from the same audio scene as the objects marked with 1. The last four bits are used for zero padding to force the PI data frame to be byte-aligned. With COMMON_AUDIO_SCENE_ISM, the size of the PI data frame may be known (1 byte) even if the pre-defined size flag is set to 1 in the ToC byte.
[0075] SEPARATE_AUDIO_SCENE PI data frames may be used to indicate that some of the audio inputs originate from a separate audio scene. The division of the common and separate audio may not be specified between inputs. No extra data section is transmitted with SEPARATE_AUDIO_SCENE PI data (i.e., the PI data frame section is empty), as the separate audio scene indication is recognized from the PI type alone.
[0076] MASA_DESCRIPTIVE_META PI data frames may be used to convey the MASA descriptive metadata as described in Annex A of the TS 26.258 standard. The descriptive metadata may be upwards of two bytes.
[0077] MULTISTREAM PI data frames may be used to transmit identification data for multiple streams in a single IVAS RTP pay load, which is described below in more detail. In an example including two streams, each stream has an 8-bit section consisting of an ID field (3 bits) identifying the stream, an S-flag (1 bit) identifying the streams originating from a common audio scene, an F-flag (1 bit) determining if another stream follows this entry (i.e., if another 8-bit multistream section follows), and an NT field (3 bit) that describes how many IVAS ToC bytes are associated with the stream.
[0078] MULTISTREAM_CMR PI data frames may be used to transmit multiple CMR bytes in an IVAS multistreaming session. In an example with two CRMs, each stream has a 16-bit section assigned to them. The ID and F-flags are as described with respect to the MUTLISTREAM PI data frame. The ID field is used to identify the stream that is affected by the associated CMR (i.e., to identify a received stream). Additionally, the identifier section (ID, F-flag, E-flag) may be zero-padded to force byte-alignment, the CMR byte fields in the PI data frame may be identical to IVAS or EVS CMR bytes depending on the value of the E-flag. The E-flag determines whether the CMR byte is in IVAS (1) or in EVS (0) format.
[0079] FIG. 5 illustrates a table 500 for providing OUTF bits and indicated IVAS output formats. OUTPUT_FORMAT PI data frames may be used to indicate feedback information for IVAS output format for playback. The first four bits in an 8-bit OUTPUT_FORMAT PI data frame indicates the IVAS output format. The last four bits may be used for zero-padding.
[0080] MUTE_STREAM PI data frames may be used to request to mute an incoming IVAS stream in a multistreaming session. For example, the requested stream may send NO_DATA to the receiver after detecting a MUTE_STREAM request. A MUTE_STREAM PI data frame may include two stream identifiers which indicate a request to mute the incoming streams with ID1 and ID2. If a MUTE_STREAM request is targeted for only a single stream, only that stream identifier shall be present in the PI data frame. If a MUTE_STREAM request is targeted for more streams, the MUTE_STREAM PI data frame may be extended with additional bytes to include all necessary stream identifiers. The MUTE_STREAM PI data frame may be zero-padded.
[0081] MUTE_ISM PI data frames may be used to request to mute individual audio objects. Each of the first four bits in the MUTE_ISM PI data frame represents a single audio object. All objects marked as 1 indicate that they should be muted. The last four bits may be used for zero padding.
[0082] FIG. 6 illustrates a table 600 for metadata types specified in the metadata task force and estimated PI data frame sizes in bytes. Some of the metadata are already included in the supported PI types (for example, reverb as part of acoustic environments, voice indication, device and scene orientations, and the like). The remaining metadata types may also be supported and transmitted through PI frames.
[0083] FIG. 7 illustrates an example PI data block included in an IVAS RTP payload 700. For example, a PI data block is identified with an IVAS ToC byte. The PI header may contain information about the size and type of the PI data in addition to identification flags. In the example IVAS RTP payload 700, the IVAS RTP payload 700 includes one PI data block and one IVAS frame. The PI data block is identified with a ToC(PI) byte. The F-bit in the ToC(PI) byte indicates that another ToC byte follows (the ToC byte for identifying the IVAS frame). After the two ToC bytes, the PI data block follows. The IVAS frame follows the PI data block. Multiple PI data blocks may be identified in a single payload with an identifying ToC byte for each PI data block.
[0084] When IVAS is implemented, RTP packets may include both PI data and audio data, and the PI data may need to be synchronized with the audio data. When forward direction PI data is included in RTP packets, the PI data may be associated with an audio frame and may use the media time of the audio data. However, if PI data is transmitted and no audio frame is available, such as during DTX periods, then a NO_DATA frame may be included in the packet containing the PI data. Alternatively, PI data may be transmitted with SID frames.
[0085] In some instances, if PI data is not needed to be transmitted when there is no audio frame available, such as during DTX periods where no audio frames are transmitted, then the transmission of PI data may be delayed until an audio frame is available. If there are multiple PI data frames with the same type available after a delay period, the most recent PI data may be selected for transmission (e.g., the most recent device orientation may be transmitted). The other (older) PI data frames with the same type may be ignored. Depending on the type of the delayed PI data frames, in some instances, all of the PI data frames with the same type may be transmitted.
[0086] When receiving an RTP packet with both PI data and several audio frames, the media time of the first audio frame may be calculated from the RTP time stamp. The media time of a subsequent audio frame may be calculated using the media time of a preceding audio frame and adding 20 ms. PI data may be sent in a separate RTP packet from the audio frame and the media time calculated from the RTP time stamp. RTP Payload
[0087] Examples described herein provide RTP payloads for indicating whether split rendering is available. For example, a new indicator may be added to the IVAS ToC byte indicating whether split rendering is available. When the ToC byte is set for split rendering, a split rendering frame is provided in the payload.
[0088] The RTP pay load structure may be similar to the EVS pay load structure for header-full format, as described in the TS 26.445 standard. The first byte of an IVAS RTP payload may be a ToC byte or a CMR / EXT byte, both structured as the respective ToC and CMR bytes of the TS 26.445 standard. Features of EVS payload format may be retained. IVAS features may be accessible through use of reserved or unused code points in EVS payload format. It should be noted that the use of an RTP payload to transmit split rendering data is merely an example, and other payload types may be implemented for transmitting split rendering data.
[0089] FIG. 8 illustrates an example RTP header 800 with an IVAS payload 805. The frame data consists of one or more IVAS or EVS coded frames. The PI data, when included, provides additional multi-purpose metadata to support rendering, as previously described. The IVAS payload 805 may include zero-padding bits in addition and at the end of the IVAS pay load 805.
[0090] FIG. 9 illustrates an example IVAS payload header 900. The IVAS payload header 900 consists of ToC bytes and Extra (E) bytes, described below in more detail. The first bit of the IVAS pay load header 900 is a Header Type identification bit (H), which indicates whether a header byte is a ToC or E byte. If the H bit is set to 0, the corresponding byte is a ToC byte. If the H bit is set to 1, the corresponding byte is an E byte.
[0091] The ToC bytes define the content of the frame data in the IVAS payload following the IVAS payload header. For each IVAS or EVS frame and for each NO_DATA frame (i.e., a frame that has zero size frame data) in the payload, there shall be one ToC byte to signal the IVAS mode and bit rate. The ToC bytes and the respective frame data may be in the same order.
[0092] FIG. 10 illustrates a ToC byte 1000 for an IVAS frame and a corresponding bit rate index table 1050. The ToC byte 1000 includes the following fields: F (1 bit) and BR (4 bits). The F field indicates whether the header byte is followed by another header byte. For example, if set to 1, the F field indicates that the header byte is followed by another header byte. If set to 0, the F field indicates that the header byte is the last header byte in the payload and is not followed by another header byte. The BR field indicates the IVAS bit rate corresponding to an IVAS mode and bit rate.
[0093] The BR field may also be used to indicate whether split rendering is available. For example, when the BR field is set to 1110, the BR field may indicate that split rendering is enabled. In another example, when the BR field is set to 1111, the BR field may indicate that split rendering is disabled. When the ToC byte 1000 is set for split rendering, a split rendering ToC (SR ToC) is implemented to indicate a split rendering frame in the payload. In the case of several SR ToC bytes, the split rendering frames follow in the same sequence as the ToC bytes.
[0094] FIG. 11 illustrates an SR ToC byte 1100 and a corresponding split rendering bit rate table 1105. The SR ToC byte 1100 includes the following fields: F (1 bit), res (2 bit), and SR-BR (4 bits). The F field may mirror the F field of the preceding ToC byte 1000, and may be ignored during processing. The res field indicates reserved bits, and may be set to zero. The SR-BR field indicates an IVAS split rendering bit rate. The split rendering bit rate table 1105 provides example values of the SR-BR field and corresponding IVAS split rendering bit rates. In some instances, the SR ToC byte 1100 includes an MDC field indicating whether the split rendering mode configuration is stereo, O-DoF, 3-DoF, or 6-DoF.
[0095] FIG. 12 illustrates a codec mode request (CMR) byte 1200 and a corresponding CMR coding table 1205. The CMR byte 1200 is compliant with CMR of TS 26.445, using the reserved code points for T = “ 111”. Using the CMR byte 1200, all EVS modes and rates may be requested. The CMR byte 1200 includes the following fields: H (1 bit), T (3 bit), and D (4 bit). The IVAS bit rate is encoded in the D field of the CMR byte 1200. The CMR byte with IVAS-br (bit rate) request is followed by an IVAS-cr (configuration request) byte 1300 (described below with respect to FIG. 13). One code point in the D field is reserved to indicate an extension with an E-byte is to follow. Unlike EVS, several CMRs may follow in sequence, indicated by the H field. Each CMR may correspond to a respective IVAS stream out of multiple incoming IVAS streams. Example values of the H field, the T field, and the D field are provided in the CMR coding table 1205.
[0096] FIG. 13 illustrates an example IVAS-cr byte 1300, a corresponding first IVAS-cr coding table 1305, and a corresponding second IVAS-cr coding table 1310. The IVAS-cr byte 1300 follows a CMR byte 1200 with an IVAS-br request. The IVAS-cr byte 1300 has the following fields: U (1 bit), BW (3 bit), and FMT (4 bit). The U bit is a reserved bit, and may be set to zero or otherwise ignored. The BW field indicates an audio bandwidth request. The first IVAS-cr coding table 1305 provides example audio bandwidths and the respective coding values for the BW field. The FMT field indicates an audio format request. The second IVAS-cr coding table 1310 provides example audio format requests and the respective coding values for the FMT field.
[0097] FIG. 14 illustrates an example extra (E) byte 1400. An E byte 1400 may contain extra information and may precede the ToC bytes 1000 of the coded frames that they relate to. Multiple E bytes 1400 may precede a ToC byte 1000. After the initial E-byte with a codec mode request (CMR) (described below with respect to FIG. 16), there may be multiple subsequent E bytes preceding ToC bytes 1000. Subsequent E bytes 1400 may be extended by another E byte 1400 of the same type. E bytes 1400 may precede any ToC byte 1000.
[0098] FIG. 15 illustrates an example initial E byte 1500. If a CMR is sent in the current RTP packet, the initial E byte may follow the structure of a CMR byte. The initial E byte 1500 includes the following fields: H (1 bit) and BR (4 bit). The H field is a header type identification bit, and may be set to 1 for any E byte. The BR field indicates the IVAS bit rate. With reference to CMR coding table 1205 of FIG. 12, the CMR code-point “NOREQ” (when D=1 111) may be considered equivalent to no CMR value being sent, and is ignored by the receiver. It should be understood that, when operating in IVAS immersive mode, a received EVS CMR having a T value that is not equal 21 to 111 is a request to switch to EVS operation mode. Alternatively, when operating in EVS mode, a received IVAS CMR having a T value that is equal to 111 is a request to switch to IVAS immersive operation mode. Additionally, in a split rendering session, the CMR byte may be respective of the first stream. Rate requests for a second stream use a second initial E byte for that stream.
[0099] FIG. 16 illustrates a subsequent E byte 1600 for a bandwidth request, a first subsequent E byte coding table 1605, and a second subsequent E byte coding table 1610. Subsequent E bytes 1600 may follow an initial E byte 1500 to request bandwidth or coded format, or to indicate the presence of PI data in the payload. The E byte 1600 includes the following fields: H (1 bit), ET (2 bits), and BW (2 bits). The res field indicates reserved bits, and may be set to zero. The H field is a header type identification bit, and may be set to 1 for any E byte. The ET field indicates a type of subsequent E byte. The types of E bytes and the respective coding values are provided in the first subsequent E byte coding table 1605. The BW field indicates a requested bandwidth. The possible requested bandwidths and the respective coding values are provided in the second subsequent E byte coding table 1610.
[0100] FIG. 17 illustrates a subsequent E byte 1700 for a CMR and a subsequent E byte coding table 1705. The subsequent E byte 1700 includes the following fields: H (1 bit), ET (2 bits), and FMT (3 bits). The res field indicates reserved bits, and may be set to zero. The H field and the ET field are as previously described with respect to the subsequent E byte 1600. The FMT field indicates a requested coded format. The possible formats and the respective coding values are provided in the subsequent E byte coding table 1705.
[0101] FIG. 18 illustrates a subsequent E byte 1800 for indicating PI data presence in the payload. The subsequent E byte 1800 includes the following fields: H (1 bit) and ET (2 bits). The res field indicates reserved bits, and may be set to zero. The H field and the ET field are as previously described with respect to the subsequent E byte 1600. When indicating PI data presence, the ET field may be set to a value of 10.
[0102] Split Renderer configuration requests may also be provided in E bytes. An SR subsequent E byte may request diegetic or non-diegetic support indicating whether audio is a headtrackable stream or a non-headtrackable stream. For example, a diegetic value of 0 may correspond to requesting a non-headtrackable stereo, O-DoF binaural, or non-diegetic stream. A diegetic value of 1 may correspond to requesting a headtrackable diegetic stream with pose correction at the receiver. An SR subsequent E byte may also request pose correction around specific axes or request disabling of pose correction around specific axes.
[0103] FIG. 19 illustrates a subsequent E byte 1900 for an SR configuration request. The subsequent E byte 1900 includes the following fields: H (1 bit), ET (2 bit), D (1 bit), Y (1 bit), P (1 bit), and R (1 bit). The res field indicates reserved bits, and may be set to zero. The H field and the ET field are as previously described with respect to the subsequent E byte 1600.
[0104] If the stream is a diegetic split rendering stream, the D field indicates whether the stream should have diegetic support. For example, the D field having a value of 1 indicates a request for enabling of diegetic support, while the D field having a value of 0 indicates a request for disabling of diegetic support. If the stream is a non-diegetic split rendering, the D field may be ignored or set to 0.
[0105] The Y field indicates a request to generate (and transmit) pose correction metadata around the yaw axis (e.g., the z-axis). For example, the Y field having a value of 1 enables creation of post correction metadata around the yaw axis, while the Y field having a value of 0 disables creation of pose correction metadata around the yaw axis. The P field indicates a request to generate (and transmit) pose correction metadata around the pitch axis (e.g., the y-axis). For example, the P field having a value of 1 enables creation of pose correction metadata around the pitch axis, while the P field having a value of 0 disables creation of pose correction metadata around the pitch axis. The R field indicates a request to generate (and transmit) pose correction metadata around the roll axis (e.g., the x-axis). For example, the R field having a value of 1 enables creation of pose correction metadata around the roll axis, while the R field having a value of 0 disables creation of pose correction metadata around the roll axis.
[0106] FIG. 20 illustrates a block diagram of various example methods 2000 for adapting an immersive audio payload for RTP transport between an upstream device and a downstream device, which may be performed by the near-end user equipment 102. The methods 2000 may be performed by one or more processors, which may be configured to perform methods 2000 via machine-executable instructions. The methods 2000 may be broken into various blocks or partitions, such as blocks 2002, 2004, 2006, 2008, 2010, 2012, and 2014. The various process blocks illustrated in FIG. 20 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 2002.
[0107] At block 2002, “Receiving Split-Rendering Control Information From The End Device Over The Communication Link Responsive,” an example method 2000 may include receiving splitrendering control information from the end device over the communication link. For example, with reference to FIG. 1, the near-end user equipment 102 receives split-rendering control information over the communication link 111 from the end device 101. The split-rendering control information may be, for example, the CMR byte 1200 or an initial E byte 1500 indicating a split-rendering bit rate. Processing may proceed from block 2002 to block 2004.
[0108] At block 2004, “Adapting A Rendering Operation Or Mode Of The Near-End Device Based On The Received Split-Rendering Control Information,” an example method 2000 may include adapting a rendering operation or mode of the near-end device based on the received split-rendering control information. For example, with respect to FIG. 1 and responsive to the CMR byte 1200, the near-end user equipment 102 updates a rendering operation or mode for rendering audio content. Processing may proceed from block 2004 to block 2006.
[0109] At block 2006, “Preparing, Responsive To The Received Split-Rendering Control Information, A First RTP Payload With Audio Data For At Least One Stream Of Coded Binaural Audio And Related Split Rendering Metadata For Communication Between The Near-End Device And The End Device,” an example method 2000 may include preparing, responsive to the received split-rendering control information, a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device and the end device. For example, with reference to FIG. 1, the near-end user equipment 102 generates a first RTP payload for communication between the near-end user equipment 102 and the end device 101. The first RTP payload may be associated with a split rendering mode for the end device 101. The audio data may be indicated by a ToC byte 1000 or a SR ToC byte 1100. In some alternative embodiments, the first RTP payload with audio data at block 2006 corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.” Processing may proceed from block 2006 to block 2008.
[0110] At block 2008, “Transmitting The First RTP Payload To The End Device Through The Communication Link,” an example method 2000 may include transmitting the first RTP payload to the end device through the communication link. For example, with reference to FIG. 1, the near-end user equipment 102 transmits the first RTP payload to the end device 101 via the communication link 111. Processing may proceed from block 2008 to block 2010.
[0111] At block 2010, “Receiving Second Split-Rendering Control Information From The End Device In Response To The Transmitted First RTP Payload,” an example method 2000 may include receiving second split-rendering control information from the end device in response to the transmitted first RTP payload. For example, with reference to FIG. 1, the near-end user equipment 102 receives second split-rendering control information from the end device 101 via the communication link 111. The second split-rendering control information may be received in response to the transmitted first RTP pay load. Processing may proceed from block 2010 to block 2012.
[0112] At block 2012, “Preparing, Based On The Received Second Split-Rendering Control Information, A Second RTP Payload With Audio Data For At Least One Stream Of Coded Binaural Audio And Related Split Rendering Metadata For Communication Between The Near-End Device And The End Device,” an example method 2000 may include preparing, based on the received second split-rendering control information, a second RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device and the end device. For example, with respect to FIG. 1, the near-end user equipment 102 generates a second RTP payload including coded binaural audio and split rendering metadata. In some alternative embodiments, the second RTP payload with audio data at block 2012 corresponds to “a second RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.” The second RTP payload is responsively pre-rendered based on the received second split-rendering control information and is associated with the split rendering mode for both the near-end user equipment 102 and the end device 101. Processing may proceed from block 2012 to block 2014.
[0113] At block 2014, “Transmitting The Second RTP Payload To The End Device Through The Communication Link To Enable Split-Rendering On The End Device,” an example method 2000 may include transmitting the second RTP payload to the end device through the communication link to enable split-rendering on the end device. For example, with reference to FIG. 1, the near-end user 25 equipment 102 transmits the second RTP pay load to the end device 101 via the communication link 111.
[0114] FIG. 21 illustrates a block diagram of various example methods 2100 for processing an immersive audio payload for RTP transport received from an upstream device by a downstream device, which may be performed by the end device 101. The methods 2100 may be performed by a one or more processors, which may be configured to perform methods 2100 via machine-executable instructions. The methods 2100 may be broken into various blocks or partitions, such as blocks 2102, 2104, 2106, 2108, and 2110. The various process blocks illustrated in FIG. 21 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 2102.
[0115] At block 2102,” Receiving A First RTP Payload With Audio Data For At Least One Stream Of Coded Binaural Audio And Related Split Rendering Metadata From The Upstream Device By A Communication Link,” an example method 2100 may include receiving a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device by a communication link. For example, with reference to FIG. 1, the near-end end device 101 receives a first RTP payload from the near-end user equipment 102. In some alternative embodiments, the first RTP payload with audio data at block 2102 corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.” The first RTP pay load is associated with a split rendering mode for the end device 101 and the near-end user equipment 102. Processing may proceed from block 2102 to block 2104.
[0116] At block 2104, “Applying A First Rendering Operating Mode Of The Downstream Device Based On The Received Split-Rendering Metadata From The Upstream Device,” an example method 2100 may include applying a first rendering operating mode of the downstream device based on the received split-rendering metadata from the upstream device. For example, with reference to FIG. 1, the end device 101 adjusts its rendering operating mode based on the first RTP pay load. In some instances, the first RTP pay load contains a SR ToC byte 1100 indicating a splitrendering bit rate. Processing may proceed from block 2104 to block 2106.
[0117] At block 2106, “Transmitting Split-Rendering Control Information From The Downstream Device In A Reverse Direction Over The Communication Link To The Upstream Device,” an example method 2100 may include transmitting split-rendering control information from the downstream device in a reverse direction over the communication link to the upstream device. For example, with reference to FIG. 1, the end device 101 transmits split-rendering control information to the near-end user equipment 102 via the communication link 111. The split-rendering control information may be transmitted alongside encoded upstream audio data. Processing may proceed from block 2106 to block 2108.
[0118] At block 2108, “Receiving A Second RTP Payload With Audio Data For At Least One Stream Of Coded Binaural Audio And Related Split Rendering Metadata From The Upstream Device By The Communication Link,” an example method 2100 may include receiving a second RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device by the communication link. For example, with reference to FIG. 1, the end device 101 receives a second RTP payload from the near-end user equipment 102 via the communication link 111. In some alternative embodiments, the second RTP payload with audio data at block 2108 corresponds to “a second RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.” Processing may proceed from block 2108 to block 2110.
[0119] At block 2110, “Rendering Audio From The Second RTP Payload Based On The Adjusted Rendering Mode,” an example method 2100 may include rendering audio from the second RTP payload based on the adjusted rendering mode. For example, with reference to FIG. 1, the end device 101 renders audio from the second RTP pay load based on the adjusted rendering mode.
[0120] FIG. 22 illustrates a block diagram of various example methods 2200 for generating an immersive audio pay load for RTP transport, which may be performed by the end device 101 and / or the far-end user equipment 103. The methods 2200 may be performed by one or more processors, which may be configured to perform methods 2200 via machine-executable instructions. The methods 2200 may be broken into various blocks or partitions, such as blocks 2202, 2204, 2206, 2208, and 2210. The various process blocks illustrated in FIG. 22 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 2202.
[0121] At block 2202, “Generating A ToC Byte For Inclusion In A First RTP Payload,” an example method 2200 may include generating a ToC byte for inclusion in a first RTP payload. For example, the far-end user equipment 103 or the near-end user equipment 102 may generate a ToC byte 1000 of FIG. 10. Processing may proceed from block 2202 to block 2204.
[0122] At block 2204, “Setting A Bit Rate (BR) Field Included In The ToC Byte To Indicate That The First RTP Payload Is A Split Rendering Payload,” an example method 2200 may include setting a BR field included in the ToC byte to indicate that the first RTP payload is a split rendering payload. For example, with reference to FIG. 10, the BR field of the ToC byte 1000 may be set to 1110, as shown in the corresponding bit rate index table 1050. Processing may proceed from block 2204 to block 2206.
[0123] At block 2206, “Generating A Split Rendering ToC Byte (SR ToC) For Inclusion In The First RTP Payload Sequential To The ToC Byte,” an example method 2200 may include generating an SR ToC Byte for inclusion in the first RTP payload sequential to the ToC byte. For example, the far-end user equipment 103 or the near-end user equipment 102 may generate a SR ToC byte 1100 of FIG. 11. Processing may proceed from block 2206 to block 2208.
[0124] At block 2208, “Setting A Split Rendering Bit Rate (SR-BR) Field Included In The SR ToC Byte To Indicate An IVAS Split Rendering Bit Rate,” an example method 2200 may include setting an SR-BR field included in the SR ToC Byte to indicate an IVAS split rendering bit rate. For example, with reference to FIG. 11, the SR-BR field of the SR ToC Byte may be set to indicate the selected IVAS split rendering bit rate, as shown in the corresponding split rendering bit rate table 1105. Processing may proceed from block 2208 to block 2210.
[0125] At block 2210, “Preparing The First RTP Payload With Audio Data For At Least One Stream Of Coded Binaural Audio And Related Split Rendering Metadata For Communication Between An Upstream Device And A Downstream Device,” an example method 2200 may include preparing the first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between an upstream device and a downstream device. In some alternative embodiments, the first RTP payload with audio data at block 2210 corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.” For example, the far-end user equipment 103 or the near-end user equipment 102 generates the first RTP pay load.
[0126] FIG. 23 illustrates a block diagram of various example methods 2300 for RTP transport of an adaptive immersive audio payload using a common payload format, which may be performed by the end device 101, the near-end user equipment 102, and / or the far-end user equipment 103. The methods 2300 may be performed by one or more processors, which may be configured to perform methods 2300 via machine-executable instructions. The methods 2300 may be broken into various blocks or partitions, such as blocks 2302, 2304, 2306, 2308, 2310, and 2312. The various process blocks illustrated in FIG. 23 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 2302.
[0127] At block 2302, “Preparing A First RTP Payload With Audio Data For At Least One Stream Of Coded Binaural Audio And Related Split Rendering Metadata For Communication Between The Upstream Device, The First Downstream Device, And The Second Downstream Device,” an example method 2300 may include preparing a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the upstream device, the first downstream device, and the second downstream device. In some alternative embodiments, the first RTP payload with audio data at block 2302 corresponds to “a first RTP payload with audio data for at least one stream of coded two channel audio and related split rendering metadata.” For example, the far-end user equipment 103 generates the first RTP payload. Processing may proceed from block 2302 to block 2304.
[0128] At block 2304, “Transmitting The First RTP Payload To The Second Downstream Device Through A First Communication Link,” an example method 2300 may include transmitting the first RTP payload to the second downstream device through a first communication link. For example, the far-end user equipment 103 transmits the first RTP pay load to the near-end user equipment 102 through the communication link 112. Processing may proceed from block 2304 to block 2306.
[0129] At block 2306, “Processing The First RTP Payload By The Second Downstream Device To Generate An Adapted First RTP Payload,” an example method 2300 may include processing the first RTP payload by the second downstream device to generate an adapted first RTP payload. For example, the near-end user equipment 102 processes the first RTP payload. Processing may proceed from block 2306 to block 2308.
[0130] At block 2308, “Preparing The Adapted First RTP Payload For RTP Transport For Communication Between The Second Downstream Device And The First Downstream Device Using The Same Payload Format,” an example method 2300 may include preparing the adapted first RTP payload for RTP transport for communication between the second downstream device and the first downstream device using the same payload format. For example, the near-end user equipment 102 prepares the adapted first RTP pay load for communication to the end device 101. Processing may proceed from block 2308 to block 2310.
[0131] At block 2310, “Transmitting The Adapted First RTP Payload To The First Downstream Device Through A Second Communication Link,” an example method 2300 may include transmitting the adapted first RTP payload to the first downstream device through a second communication link. For example, the near-end user equipment 102 transmits the adapted first RTP payload to the end device 101 through the communication link 111. Processing may proceed from block 2310 to block 2312.
[0132] At block 2312, “Receiving And Processing The Adapted First RTP Payload By The First Downstream Device,” an example method 2300 may include receiving and processing the adapted first RTP payload by the first downstream device. For example, the end device 101 receives and processes the adapted first RTP pay load.
[0133] In some instances, the near-end user equipment 102 operates to extract split-rendering control information from the end device 101 that is intended for the far-end user equipment 102. For example, the end device 101 transmits a second RTP payload including split-rendering control information to the near-end user equipment 102 through the communication link 111. The near-end user equipment 102 receives the second RTP pay load from the end device 101 and extracts the splitrendering control information from the second RTP payload. At least a first portion of the splitrendering control information may be intended for the near-end user equipment 102 and is implemented by the near-end user equipment 102. A second portion of the RTP payload may be intended for the far-end user equipment 103, e.g. upstream coded audio data generated by end device 101. Near-end user equipment 102 may also add control information for downstream coding operation by far-end user equipment 103. Accordingly, the near-end user equipment 102 may generate an adapted second RTP payload including the second portion of the RTP payload for the far-end user equipment 103 and transmit the adapted second RTP payload to the far-end user equipment 103 through the communication link 112. Further Embodiments
[0134] FIGS. 24A-24J illustrate various example payload structures. For example, FIG. 24A illustrates an example single IVAS frame in a packet at 96 kbps. FIG. 24B illustrates an example of two IVAS frames in a packet at 13.2 and 64 kbps. FIG. 24C illustrates an example of a single IVAS frame in a packet at 96 kbps. The payload structure of FIG. 24C includes a CMR for an EVS operating mode. FIG. 24D illustrates an example of a single IVAS frame in a packet at 96 kbps. The payload structure of FIG. 24D includes a CMR for an IVAS 80 kbps operating mode, an SWB operating bandwidth, and ISM-2.
[0135] FIG. 24E illustrates an example of a single IVAS frame in a packet at 96 kbps. The payload structure of FIG. 24E includes a CMR for an IVAS 80 kbps operating mode, format unchanged, and PI data for ACOUSTIC_ENVIRONMENT_COMPACT. FIG. 24F illustrates an example of three IVAS frames in a packet, each at 512 kbps. The payload structure of FIG. 24F includes a CMR for an IVAS 80 kbps operating mode, format unchanged, and PI data for ACOUSTIC_ENVIRONMENT_COMPACT. As the payload structure is greater than a size limit, the pay load is split into three RTP packets.
[0136] FIG. 24G illustrates an example of a single EVS-AMRWB-IO frame in a packet at 23.85 kbps. The payload structure of FIG. 24G includes an IVAS CMR for one IVAS stream at 384 kbps and another IVAS stream at 256 kbps without any configuration requests. FIG. 24H illustrates a single IVAS-SR frame in a packet at 512 kbps transmitted to a lightweight device (e.g., to the end device 101).
[0137] FIG. 241 illustrates two diegetic (3-DoF) IVAS-SR frames at 512 kbps each and one non-diegetic frame at 256 kbps transmitted to a lightweight device. The payload structure of FIG. 241 includes a CMR for EVS-AMRWB-IO 23.85 kbps on the return link. FIG. 24J illustrates a single EVS-AMRWB-IO frame in a packet at 23.85 kbps transmitted from a lightweight device. The payload structure of FIG. 24J includes an IVAS-SR CMR for one diegetic (3-DoF) stream at 384 kbps and one non-diegetic stream at 256 kbps, and includes head tracking data. The examples of FIG. 24H-24J may be examples of split-rendering payloads.
[0138] As an example of packetization operation at the end device 101, the end device 101 may create a payload portion with IVAS encoded audio representing the captured audio signal. Implementers may use the format that is compatible with the EVS RTP pay load format to generate the ToC byte of the payload header and apply formatting of the audio payload. In addition, to packetize the information controlling the operation of the SR pre-renderer and encoder, the following may be performed.
[0139] In a first step, the end device 101 may check if any of the following control information for the SR pre-renderer and encoder should be sent: M, the number of requested SR audio streams, MDC, the requested split rendering configuration, and SR-BR, the requested total bit rate of SR coding in any of the audio streams. If true, the end device 101 continues to the second step. Else, the end device 101 continues to the third step.
[0140] In a second step, the end device 101 sets a stream counter] to 1. The end device 101 creates a CMR / EXT byte with code with an E-byte identifier: (H, T, D) = “11111110” to signal a subsequent E-byte. The end device 101 creates an E-byte with code “00010000” to identify a subsequent IVAS-SR CMR byte. The end device 101 creates an IVAS-SR CMR byte with code (M, MDC, SR-BR) with sub-code words for M, MDC, and SR-BR, where MDC and SR-BR encodes the respective control parameters for the j-th stream. The second step is repeated for all stream counters (e.g., if stream counter] equals M, the end device 101 continues to the third step).
[0141] In a third step, the end device 101 checks if head-tracker data needs to be sent. If headtracker data should be sent, the end device 101 proceeds to the fourth step. Otherwise, the end device 101 proceeds to the fifth step.
[0142] In a fourth step, the end device 101 creates a CMR / EXT byte with code with an E-byte identifier: (H, T, D) = “11111110” to signal a subsequent E-byte. The end device 101 creates an Ebyte with code “00010110” to identify the subsequent insertion of PI bytes encoding head-tracker data (for example, head orientation). The end device 101 inserts PI frame data, such as a header and head-tracker data consisting of 9 bytes of data.
[0143] In a fifth step, the end device 101 checks whether any IVAS audio should be sent upstream. If yes, the end device 101 creates a ToC byte and audio payload according to the existing EVS or IVAS RTP pay load format specifications.
[0144] As an example of de-packetization operations at the near-end user equipment 102 of packets arriving from the end device 101, a portion of the received pay load may be used for controlling the operation of the SR pre-renderer and encoder. Another portion of the received payload may be intended for onward transmission to the far-end device (e.g., the far-end user equipment 103).
[0145] In a first step for pay load identification, the near-end user equipment 102 initializes byte index k to 0 such that it points to the first byte of the RTP packet payload.
[0146] In a second step for payload identification, the near-end user equipment 102 reads the byte from the RTP packet payload with index k. If index k is not at the end of the RTP packet payload, the near-end user equipment 102 continues to the third step. Otherwise, the near-end user equipment 102 terminates payload identification processing.
[0147] In a third step for payload identification, the near-end user equipment 102 parses the leftmost bit (MSB) of the byte at index k, if the leftmost bit is 0, a ToC byte is identified and the parsing operation terminates. If the leftmost bit is a 1, a CRM / EXT byte is identified. If the 7 left bits (corresponding to T, D of the CMR / EXT byte) are equal to the IVAS-Eb identifier of 1111110, the byte index k is incremented and an E-byte parser is initiated. Otherwise, depending on whether the sub-code portion T is equal to or unequal to 111, either the EVS CMR parser or the IVAS-cr parser are initiated.
[0148] In a first step for E byte parsing, the near-end user equipment 102 reads the byte from the RTP packet payload with index k.
[0149] In a second step for E byte parsing, the near-end user equipment 102 parses the four PI frame type identifier bits (PFT) of that byte. If the bits equal 1000, 1001, or 1010, an E-byte is identified. The near-end user equipment 102 identifies the length in bytes of the PI, and the corresponding bytes are routed to a processing unit prepared to control the operation of the SR pre-renderer and encoder. The byte index k is increased by the PI length.
[0150] If the PFT bits equal 1000, the near-end user equipment 102 increments k and reads the byte from the RTP packet payload with index k. The near-end user equipment 102 parses the SR-BR sub-code to identify the requested SR bit rate. If the SR-BR field is 111, the previously used or negotiated bit rate is retained. The near-end user equipment 102 parses the next 2 bits of the byte, the MDC sub-code, to identify the requested SR metadata configuration. If these bits equal 11, the previously used or negotiated configuration is retained. This sub-code is an identification of the used SR mode, e.g., whether the respective stream to be encoded is diegetic and including pose correction metadata allowing the post-renderer to adjust the received binaural audio (or other audio such as two channel audio) in response to the actual head position available at the end device.
[0151] Next, the near-end user equipment 102 parses the M sub-code to identify the requested number of SR audio streams. If these bits equal 111, the previously used or negotiated number is retained. For the case with multiple audio streams sent by the SR encoded, i.e., M>1, for instance enumerated by m ranging from m=l.. .M, there may be up to M IVAS-SR codec mode requests, each signaled through the IVAS-Eb identifier (111, 1110) and followed by an E-byte set to 00010000 (with PFT bits set to 1000) and followed by an IVAS-SR CMR byte. In such a sequence of up to M IVAS-SR codec mode requests, the m-th of such a codec mode request is respective of the m-th audio stream. The sub-code for M may only be parsed in the IVAS-SR CMR byte of the respective first audio stream.
[0152] If the PFT bits do not equal 1000, the recipient of the PI is assumed to be the far-end user equipment 103. In such an instance, the length in bytes of that PI is identified and the corresponding bytes are routed to a processing unit prepared to handle re-packetization and onward transmission to the far-end user equipment 103. The byte index k is increased by the PI length.
[0153] As an example of operations related to generation of IVAS CMR and PI data originating in the near-end user equipment 102 and directed to the far-end user equipment 103, the near-end user equipment 102 may currently receive from the far-end user equipment 103 packets of an IVAS bitstream for which a bit rate or coding mode should be changed. In addition, there may be extra information that the end device 101 or the near-end user equipment 102 wishes to be sent to the far-end user equipment 103.
[0154] In a first step of creating a codec mode request for the far-end user equipment 103, the nearend user equipment 102 checks whether control information for the far-end user equipment 103 IVAS encoder should be sent: D (IVAS bit rate), BW (audio bandwidth) or FMT (audio format).
[0155] In a second step, the near-end user equipment 102 creates a 4-bit sub-code word D encoding the requested bit rate. The near-end user equipment 102 creates a 5-bit code word (H, T) = 11111, where H=1 signals the CMR / EXT byte and T=1111 signals that the byte is a CMR for an IVAS codec mode. The near-end user equipment 102 creates the CMR / EXT byte by combining the D and H,T sub-code words (H, T, D). The near-end user equipment 102 creates an IVAS-cr byte with subcode words for BW and FMT.
[0156] The term “create” may comprise inserting the created byte or sequence of bytes into a buffer prepared to store that byte sequence. The buffer stores the portion of information to be sent upstream that originates in the near-end user equipment 102.
[0157] In the following, processing operations are described related to combining and packetization of pay load portions intended for onward transmission to the far-end user equipment 103. The information to be sent upstream is the previously described portion of respective IVAS, CMR, and PI data originating in the near-end user equipment 102 and directed to the far-end user equipment 103, and the portion received from the end device 101 and intended for onward transmission to the far-end user equipment 103.
[0158] The following example provides parsing and de-multiplexing operations related to depacketization operations at the near-end user equipment 102 of packets arriving from the far-end user equipment 103. A portion of the received pay load is IVAS coded audio to be IVAS decoded and subsequently fed into the SR pre-renderer and encoder. Another portion is intended to be used for controlling the operation of the upstream audio encoder in the end device 101. The packet may be combined with control information for the encoder generated by the near-end user equipment 102. After combining, the packet may be sent downstream to the end device 101.
[0159] The portions are identified and separated from each other for subsequent routing to the relevant respective processing units. In a first step, the near-end user equipment 102 initializes byte index k to 0 such that the index k points to the first byte of the RTP packet payload.
[0160] In a second step, the near-end user equipment 102 reads the byte from the RTP packet payload with index k.
[0161] In a third step, the near-end user equipment 102 parses the leftmost bit (MSB) of that byte. If the bit equals 0, the byte is identified as a ToC byte. In this case, the ToC byte and the remainder of the pay load are identified as IVAS audio data intended for routing to the processing unit prepared to handle IVAS decoding, SR pre-rendering and encoding. As the payload format incorporates the existing EVS or IVAS RTP payload format, the near-end user equipment 102 terminates parsing operations. However, if the bit equals 1, the byte is identified as a CMR / EXT byte. In this case, the next 7 bits are parsed. If the bits (corresponding to the T,D portion of the CMR / EXT byte) are equal to the IVAS-Eb identifier (1111110), the byte index k is incremented and the E-byte parser is initiated. Depending on the parsing result, the relevant number of bytes belonging to the E-byte and potential extra information (PI data) are stored in a buffer and byte index k is increased by the corresponding number of bytes. Otherwise, depending on whether the sub-code portion T is equal to or unequal to 111, the byte is identified as either an EVS CMR or as an IVAS CMR. In the case of an EVS CMR, the byte is stored in a buffer and index k is incremented. In the case of an IVAS CMR, the byte and the successor byte are stored in a buffer and index k is increased by 2. The next byte is selected until the byte index k points beyond the end of the payload, at which point parsing terminates.
[0162] In the following, processing operations are described related to combining information obtained after parsing of the payload data arriving from the far-end user equipment 103 with corresponding information generated at the end device 101. Such information to be combined may be CMRs for the upstream audio encoded in the end device 101 or PI data. As an example, based on link quality metrics of the incoming links from the end device 101, a certain maximum bit rate for the end device audio encoder may be found suitable to optimize the received quality. A similar measurement may have been performed at the far-end user equipment 103, e.g., based on the quality of the decoded audio originating from the end device 101. This may lead to a codec mode request obtained from the far-end user equipment 103. The combination of both requests could potentially lead to the selection of the minimum of the bit rate requests of both codec mode requests. After combining the two codec mode requests into one, it is ready to become part of the payload to be sent to the end device 101. A similar combination of PI data may occur depending on its nature. In some instances, both PI data entities are sent to the end device 101.
[0163] A further processing step of the downlink audio processing in the near-end user equipment 102 is IVAS decoding, SR pre-rendering, and encoding. Such operations may be controlled by the control information (e.g., requested number of SR streams M, bit rate of the streams, SR metadata configuration) received upstream from the end device or downstream from the far-end device 103 (that may also control the SR configuration). An entity at the near-end user equipment 102 may combine these control requests and operate the SR pre-renderer and encoder accordingly. The SR encoder may generate audio payload data portions of a number of SR audio streams, with or without SR metadata.
[0164] The following processing operations may be done to format the mentioned data portions to be transmitted to the end device 101, for example, to bring it into the RTP pay load format to be used for transmission to the end device 101. The example is similar to RTP packet operations at the end device 101.
[0165] At a first step, the near-end user equipment 102 determines whether any of the following control information is to be sent to the upstream audio encoder: EVS CMR (codec mode request for EVS coding mode to be used) or IVAS CMR (codec mode request for IVAS coding mode to be used), which may include D (IVAS bit rate), BW (audio bandwidth), and FMT (audio format) fields.
[0166] At a second step, if the EVS CMR is to be sent, the near-end user equipment 102 creates a CMR byte according to the EVS RTP pay load format. If an IVAS CMR is to be sent, the near-end user equipment 102 creates a 4-bit sub-code word D encoding the requested bit rate. The near-end user equipment 102 creates a five-bit sub-code word (H, T) = 11111, where H=1 signals the CMR / EXT byte and T=1 111 signals that the byte is a CMR for an IVAS codec mode. The near-end user equipment 102 creates a CMR / EXT byte by combining D and (H,T) sub-code words. The near-end user equipment 102 creates an IVAS-cr byte with code (U,BW,FMT) with sub-code words for BW and FMT. The field U is set to 0.
[0167] At a third step, the near-end user equipment 102 determines whether any PI data elements are to be sent to the end device. If yes, the near-end user equipment 102 creates a CMR / EXT byte with and E-byte identifier (H, T, D) = 11111110 to signal a subsequent E-byte. The near-end user equipment 102 creates an E-byte where the sub-codes PFT and S identify information to be transmitted with the PI data. The near-end user equipment 102 creates the PI data bytes, and if there are further PI data elements to be transmitted that are not present in the RTP payload buffer, the near-end user equipment 102 repeats the third step with the next PI data element.
[0168] At a fourth step, the near-end user equipment 102 creates SR audio ToC and audio data payload. For example, the near-end user equipment 102 creates a ToC byte with code 00011111 indicating that one or several SR ToC bytes will follow. For each SR audio stream out of M SR audio streams, the near-end user equipment 102 creates an SR ToC byte composed of F, MDC, and SR-BR sub-codes involving a stream counter k running from 1 to M. For each SR audio stream, the near-end user equipment 102 creates the corresponding audio data payload.
[0169] The following example provides de-packetization operations at the end device 101 of packets arriving from the near-end user equipment 102 comprising parsing and de-multiplexing operations. A portion of the received pay load may be used for controlling the operation of an upstream audio encoder. A portion of the received payload is SR coded audio to be processed by an SR decoder and post-renderer. Another portion may be used for controlling the operation of the upstream audio encoder. A further portion may be auxiliary information transported as PI data. Each portion may be identified and separated for subsequent routing to the relevant processing units.
[0170] In a first step, the end device 101 initializes byte index k to 0 such that the byte index k points to the first byte of the RTP packet payload.
[0171] In a second step, the end device 101 reads the byte from the RTP packet payload with index k.
[0172] In a third step, the end device 101 parses the leftmost bit (MSB) of that byte. If the bit is equal to 1, the byte is identified as a CMR / EXT byte. In such an instance, the end device 101 applies parsing procedures as previously described. EVS / IVAS related codec mode request information may be propagated as control information to the upstream audio encoded. PI data may be propagated to the appropriate processor. The index k is subsequently increased by the number of bytes read from the RTP packet payload. If the bit is 0, the byte is identified as a ToC byte. In such an instance, the ToC byte is parsed. If the ToC byte is equal to 00011111, several SR ToC bytes are indicated to follow. Otherwise, the ToC byte may indicate an EVS or IVAS payload.
[0173] An SR ToC counter m may be initialized to 0 to parse the SR ToC bytes. Parsing the SR ToC bytes includes incrementing counters m and k, reading the byte from the RTP packet payload with index k, and parsing MDC bits of the byte, which encode the split rendering configuration of the pertinent SR audio stream. The sub-code is stored with an association to the m-th SR audio stream. The end device 101 parses SR-BR bits of the SR ToC byte, which encode the total bit rate of the pertinent SR audio stream. The end device 101 may calculate the corresponding number of bytes of the SR audio stream by dividing the indicated SR bit rate by 400, which is stored with an association to the m-th SR audio stream. The end device 101 may parse the leftmost bit (MSB) of that byte to determine whether another SR ToC byte follows.
[0174] In a fourth step, the end device 101 may read SR audio data from the RTP packet payload, beginning with initializing an SR audio stream counter] to 1. The end device 101 looks up a stored number of bytes for the j-th SR audio stream and reads the identified number of bytes from the RTP packet payload starting from byte index k into a buffer to store the audio frame data of the j-th SR audio stream. The end device 101 increments k by the number of read bytes. If j equals the total number of parsed SR ToC bytes m, the parsing procedure ends. Otherwise, j is incremented and the process is repeated for the next SR audio stream.
[0175] It is notable that the RTP payload size may in some cases exceed certain system limits like a maximum transfer unit size (MTU). For example, the MTU size for Ethernet is 1500 bytes. To deal with this limitation an RTP payload fragmentation mechanism is provided.
[0176] At a sending end, this involves checking if the total amount of data to be transmitted in a packet exceeds the limit. If this is the case, a CMR / EXT byte is created and inserted into the payload buffer before the ToC section, to signal the subsequent insertion of an E-byte that signals the use of RTP packet fragmentation. As an example, according to the tables in FIG. 3, this may involve creating and inserting a CMR / EXT byte with code (H,T,D) = “1,111,1110” followed by an E-byte with code (U*3,PFT,S) = “000,0111,1” or “000,0111,0”. In the former case, the used RTP packet size in case of fragmentation may be some pre-agreed size that may, e.g., have been agreed during session setup (SDP negotiation). In the latter case, a code indicating the RTP packet size may follow as part of PI data.
[0177] The RTP packetization is then modified such that the complete payload data is distributed to several RTP packets. Hereby, the first packet will carry the RTP payload header and a first portion of the audio data up to the size limit. Subsequently, the remaining payload data will be inserted in further RTP packets of size not exceeding the limit until the complete RTP payload is distributed to RTP packets. Note that subsequent RTP packets carrying RTP payload fragments are header-less.
[0178] At receiving end, the CMR / EXT byte parsing procedure is extended by the ability to detect CMR / EXT bytes with code (H,T,D) = “1,111,1110” followed by a E-bytes with code (U*3,PFT,S) = “000,0111,1” or “000,0111,0”. In the latter case, the procedure is further extended to read the size limit for fragmentation. Otherwise, the size limit is pre-agreed and thus already known. In addition, 39 when parsing the first RTP packet containing the RTP payload header, the total payload size is calculated based on the information available in the payload header. Subsequently, the complete RTP payload of the current (first fragmented packet) is read and concatenated with the RTP payloads of the subsequent fragmented RTP packets until the total payload has been collected. After that, parsing and de-multiplexing operations on the total payload buffer can take place as described above. RTP packets arriving after a last fragmented RTP packet are subsequently again parsed for their RTP payload header information.
[0179] A feature useful for split rendering is round trip delay estimation between end device and near-end UE, where near-end UE contains the SR pre-renderer and encoder, and end device contains the SR decoder and post-renderer. Round trip delay measurement is based on the principle that one of these entities sends a characteristic bit sequence to the other entity, which is then mirrored and sent back to the first entity. The first entity may then correlate the transmitted bit sequence with the mirrored sequence. This allows to calculate the time it takes to get back the transmitted sequence, which is the round-trip delay. One component of such a round-trip delay measurement is the capability to transmit characteristic bit sequences. According to one example (see FIG. 4), an E-byte code point PFT sub-code = “1001” is provided allowing to transmit a single bit. According to another example, an E-byte code point PFT sub-code = “1010” is provided allowing to transmit a bit sequence of arbitrary size using the PI mechanism. Note that the characteristic bit sequence originating from the first entity may be any bit sequence (preferably a pseudo random sequence) that is in any case transmitted to the other entity, e.g., a specific portion of the audio payload. In case a single bit (or flag) is used, the round-trip delay estimation may require more measurement time corresponding to more transmitted bits before reliable estimates can be expected. Example Device Architecture
[0180] FIG. 25A illustrates a schematic block diagram of an example device architecture 2500 (e.g., an apparatus 2500) that may be used to implement various aspects of the present disclosure. Architecture 2500 includes but is not limited to servers and client devices, systems, and methods as described in reference to FIGS. 1-24J. As shown, the architecture 2500 includes central processing unit (CPU) 2501 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 2502 or a program loaded from, for example, storage unit 2508 to random access memory (RAM) 2503. The CPU 2501 may be, for example, an electronic processor 2501. In RAM 2503, the data required when CPU 2501 performs the various 40 processes is also stored, as required. CPU 2501, ROM 2502, and RAM 2503 are connected to one another via bus 2504. Input / output interface 2505 is also connected to bus 2504.
[0181] The following components are connected to I / O interface 2505: input unit 2506, that may include a keyboard, a mouse, or the like; output unit 2507 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 2508 including a hard disk, or another suitable storage device; and communication unit 2509 including a network interface card such as a network card (e.g., wired or wireless).
[0182] In some implementations, input unit 2506 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats.
[0183] In some implementations, output unit 2507 include systems with various number of speakers. Output unit 2507 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats.
[0184] In some embodiments, communication unit 2509 is configured to communicate with other devices (e.g., via a network). Drive 2510 is also connected to I / O interface 2505, as required. Removable medium 2511, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 2510, so that a computer program read therefrom is installed into storage unit 2508, as required. A person skilled in the art would understand that although apparatus 2500 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0185] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 2509, and / or installed from the removable medium 2511, as shown in FIG. 25A.
[0186] FIG. 25B illustrates a schematic block diagram of an example CPU 2501 implemented in the device architecture 2500 of FIG. 25A that may be used to implement various aspects of the present disclosure. The CPU 2501 includes an electronic processor 2520 and a memory 2521. The electronic processor 2520 is electrically and / or communicatively connected to the memory 2521 for bidirectional communication. The memory 2521 stores an RTP coding software 2522. In some examples, memory 2521 may be located internal to the electronic processor 2520, such as in an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 2521 may be located external to the electronic processor 2520, such as in a ROM 2502, a RAM 2503, flash memory or a removable medium 2511, or another non-transitory computer readable medium that is contemplated for device architecture 2500. In some instances, the electronic processor 2520 may implement the immersive audio coding software 2522 stored in the memory 2521 to perform, among other things, any of the methods 2000 of FIG. 20, methods 2100 of FIG. 21, methods 2200 of FIG. 22, and / or methods 2300 of FIG. 23.
[0187] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 2501 in combination with other components of FIG. 25A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as nonlimiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. Payload Parameters
[0188] The following section provides parameters of the IVAS pay load format. These parameters cover real-time transfer via RTP and non-real-time transfers via stored files.
[0189] The following parameters apply to RTP transfer only.
[0190] mode-set: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0191] ptime: see IETF RFC 4566 (2006): “SDP: Session Description Protocol”, M. Handley, V. Jacobson and C. Perkins.
[0192] maxptime: see IETF RFC 4566 (2006): “SDP: Session Description Protocol”, M. Handley, V. Jacobson and C. Perkins.
[0193] dtx / dtx-recv: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0194] max-red: see IETF RFC 4867 (2007): “RTP Payload Format and File Storage Format for the Adaptive Multi-Rate (AMR) and Adaptive Multi-Rate Wideband (AMR-WB) Audio Codecs”, Sjoberg, J., Westerlund, M., Lakeniemi, A., and Q. Xie.
[0195] channels: The number of audio channels shall not be present. The use of the channels parameter as defined in IETF RFC 3551 (2003): “RTP Profde for Audio and Video Conferences with Minimal Control”, Schulzrinne, H. and S. Casner does not permit signaling all IVAS Immersive mode coded formats; formats need to be derived from the cf / cf-send / cf-recv parameters.
[0196] ivas-mode-switch: This parameter defines the mode at the start or update of the session for the send and the receive directions. Permissible values are 0 and 1. If ivas-mode-switch is 0 or not present, IVAS Immersive mode is used. If ivas-mode-switch is 1, depending on the setting of evs-mode-switch, EVS Primary or AMR-WB IO mode is used. The mode initially used in the session may later be modified.
[0197] cmr: As defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description” for the EVS Primary and AMRWB-IO modes. For IVAS Immersive modes the bit rate, bandwidth and format requests are disabled when cmr is -1. The bitrate, bandwidth and format requests are enabled when cmr is 0 or the cmr parameter is not present. When cmr is 1 the bit rate requests using the initial E byte shall be present in every packet (but may be NO_REQ); format and bandwidth requests for IVAS Immersive modes are optional when cmr is 1.
[0198] The following parameters are applicable to only IVAS Immersive operation: IVAS computational complexity and memory demands depend on the setting of the following parameters for course codec bit rate, audio bandwidth, and coded format; in addition, factors beyond the signaling, such as complexity of a specific implementation and the (rendered) output format may be significant.
[0199] ibr: Specifies the range of source codec bitrate for IVAS Immersive mode in the session, in kilobits per second, for the direction specified by the session directionality attribute or the suffix. The br parameter can either have: a single bitrate (ibrl); or a hyphen-separated pair of two bitrates (ibrl-ibr2). If a single value is included, this bitrate, ibrl, is used. If a hyphen-separated pair of two bitrates is included, ibrl and ibr2 are used as the minimum bitrate and the maximum bitrate respectively, ibrl shall be smaller than ibr2. ibrl and ibr2 have a value from the set in Table 4.2-2 of 3GPP TS 26.445. If none of these parameters is present, all bitrates consistent with the IVAS codec capabilities are allowed in the session.
[0200] ibr-send / ibr-recv: ibr parameter in send or receive direction.
[0201] ibw: Specifies the audio bandwidth for IVAS Immersive modes to be used in the session, for the direction specified by the session directionality attribute or the suffix, ibw has a value from the set: wb, swb, fb, wb-swb, and wb-fb. wb, swb, and fb represent wideband, super-wideband, and fullband respectively, and wb-swb, and wb-fb represent all bandwidths from wideband to super-wideband, and fullband respectively. If none of these parameters is present, all bandwidths consistent with the negotiated bitrate(s) are allowed in the session.
[0202] ibw-send / ibw-recv: ibw parameter in send or receive direction.
[0203] cf: Specifies the IVAS Immersive mode coded-format (cf) transmitted in the IVAS Immersive mode frames in the session. IVAS coded format corresponds to the format represented in the IVAS Immersive mode coded frames, which is generally the input format to the encoder. The cf parameter is a list of supported comma-separated IVAS Immersive mode coded formats in the order of preference, using the identifiers from Table A.4.1-1 of 3GPP TS 26.445 (column "Identifier"). Selection of the format is application-specific and out of scope of this document. EVS frames in the session are in mono format; switching to mono shall be possible. For SR format, the following applies: While the formats offered by the offerer may be a list containing SR and other formats, the answer shall either exclusively contain SR or a set of the other offered formats excluding SR. A combination of SR with other formats is not permissible.
[0204] IVAS Immersive mode coded-formats include stereo, SB A, MASA, ISM, MC, OMASA, OSBA, and SR operations. Mono operation may be supported using EVS. IVAS payloads are selfcontained for all IVAS coded formats except SR and mono (i.e., they require no additional signaling for decoding than the pay load size).
[0205] cf-send / cf-recv: cf parameter in send or receive direction.
[0206] pi-types: Specifies the supported PI data types for the session. The pi-types parameter is a list of supported comma-separated PI data types using SPD indications. If the pi-types parameter is not present, PI data is not enabled for the session.
[0207] pi-types-send / pi-types-recv: pi-types parameter in send or receive direction.
[0208] pi-br: Specifies the maximum peak bitrate for the PI data section (excluding the E-bytes for indication) for each packet in the session in kilobits per second. Bitrate calculation for PI data shall take the packet interval, i.e. value of ptime into account. The parameter indicates the maximum bitrate for the PI data. If pi-br parameter is not present, a default value of 0 shall be used.
[0209] pi-br-send / pi-br-recv: pi-br parameter in send or receive direction.
[0210] sr-dof: Specifies the number of degrees of freedom supported in the head-tracked split rendering session. Permissive values are -1, 0, 1, 2, 3. A value of -1 means that respective stream will be a non-diegetic stream in which the pre-renderer does not expect head-tracker data / does not take such data into account during pre-rendering. A value in the range of 0 - 3 means that the prerenderer expects head-tracker data, conveyed by PI frames or some other mechanism. A value of D > 0 means that metadata is generated and transmitted in the SR bitstream allowing the post-renderer to make pose corrections of the binaural audio (or other audio) in D degrees of freedom. D=1 means support of pose corrections around a single axis only, D=2 means support of pose corrections around two axes , D=3 means support of pose corrections around all 3 axes. If the first value is greater than -1, it may be followed separated by a comma by a second value respective a second non-diegetic stream. The only permissible value for that stream is -1. If sr-dof is not present in a negotiated SR session, it defaults to 3. A positive value of D may additionally be appended with the optional suffix *, indicating high-efficiency rather than high-quality split renderer metadata calculation to be used in the session.
[0211] sr-tc: Specifies the codec format for the transport channels. This parameter must be present in a negotiated split rendering session. Permissible values are ‘LCLD’ and ‘LC3plus’ followed by an optional value specifying the maximum allowed split rendering bitrate, in kilobits per second, for the SR stream(s) to be used in the session. If the bit rate value is left empty, the maximum of 512 is assumed. Following, the bitrate value, a frame length indicator, in ms, may be provided. If no value is specified, the frame length defaults to 20 ms. The frame length indicator may be 5, 10 or 20; it has effect only for non-diegetic streams or streams without pose correction metadata (D=-l or D=0 in sr-dof parameter). Table 1105 of FIG. 11 specifies the bitrates that can be specified as maximum bitrates. When LC3plus is used in a session with split rendering, two further values, ‘fdi’ and ‘bwr’ shall additionally be supplied, ‘fdi’ is specified in [Ref TS 103 634]; it shall be set to "1" or "2", depending on the split rendering configuration, ‘bwr’ is specified in [Ref TS 103 634]; it shall be set to "fb". All value fields shall be separated by commas even if left empty.
[0212] The following parameters are applicable only to EVS Primary and AMR-WB IO modes.
[0213] evs-mode-switch: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”. If ivas-mode-switch is 0 or not present, evs-mode-switch should not be present and shall be ignored.
[0214] hf-only: as specified in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description” except that the default and only allowed value of hf-only shall be 1 in this payload format. As the only allowed value for this parameter is 1 it is not required to include this parameter.
[0215] ch-send: Shall not be present. The EVS modes in this payload format shall be mono-only.
[0216] ch-recv: Shall not be present. The EVS modes in this payload format shall be mono-only.
[0217] The following parameters are applicable only to EVS Primary modes:
[0218] br: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0219] br-send: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0220] br-recv: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0221] bw: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”. Narrow-band is not supported for IVAS operation.
[0222] bw-send: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0223] bw-recv: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0224] ch-aw-recv: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0225] The following parameters are applicable only to EVS AMR-WB IO modes: mode-change-period: see IETF RFC 4867 (2007): “RTP Payload Format and File Storage Format for the Adaptive Multi-Rate (AMR) and Adaptive Multi-Rate Wideband (AMR-WB) Audio Codecs”, Sjoberg, J., Westerlund, M., Lakeniemi, A., and Q. Xie.
[0226] mode-change-capability: as defined in Annex A of 3GPP TS 26.445: “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”.
[0227] mode-change-neighbor: see IETF RFC 4867 (2007): “RTP Payload Format and File Storage Format for the Adaptive Multi-Rate (AMR) and Adaptive Multi-Rate Wideband (AMR-WB) Audio Codecs”, Sjoberg, J., Westerlund, M., Lakeniemi, A., and Q. Xie.
[0228] The information carried in the media type specification has a specific mapping to fields in the Session Description Protocol (SDP), which may be used to describe RTP sessions. When SDP is used to specify sessions employing the IVAS codec, the mapping is as follows: The media type (“audio”) goes in SDP “m=” as the media name. The media subtype (payload format name) goes in SDP “a=rtpmap” as the encoding name. The RTP clock rate in “a=rtpmap” may be 16000, and the encoding parameters (number of channels) may be omitted. The parameters “ptime” and “maxptime” go in the SDP “a=ptime” and “a=maxptime” attributes, respectively. Any remaining parameters go in the SDP “a=fmtp” attribute by copying them directly from the media type parameter string as a semicolon-separated list of parameter=value pairs.
[0229] The following considerations apply when using SDP Offer-Answer procedures to negotiate the use of IVAS payload in RTP.
[0230] hf-only: Shall not be included in the SDP offer. The answerer shall include this parameter only if it is set to 1 in the SDP offer. If the value in the SDP offer is not equal to 1, the payload type shall be rejected.
[0231] ivas-mode-switch: When ivas-mode-switch is not offered for a payload type, the answerer may include ivas-mode-switch for the payload type in the SDP answer. When ivas-mode-switch is offered for a payload type and the payload type is accepted, the answerer shall not modify or remove ivas-mode-switch for the payload type in the SDP answer.
[0232] cmr: When cmr is not offered for a payload type, the answerer may include cmr for the payload type in the SDP answer. When cmr is offered for a payload type and the payload type is accepted, the answerer shall not modify or remove cmr for the payload type in the SDP answer.
[0233] ibr: When the same bitrate or bitrate range is defined for the send and the receive directions, ibr should be used but ibr-send and ibr-recv may also be used, ibr can be used even if the session is negotiated to be sendonly, recvonly, or inactive. For sendonly session, ibr and ibr-send can be interchangeably used. For recvonly session, ibr and ibr-recv can be interchangeably used. When ibr is not offered for a payload type, the answerer may include ibr for the payload type in the SDP answer. When ibr is offered for a payload type and the payload type is accepted, the answerer shall include ibr in the SDP answer which shall be identical to or a subset of ibr for the payload type in the SDP offer.
[0234] ibr-send: When ibr-send is not offered for a pay load type, the answerer may include ibr-recv for the pay load type in the SDP answer. When ibr-send is offered for a payload type and the pay load type is accepted, the answerer shall include ibr-recv in the SDP answer, and the ibr-recv shall be identical to or a subset of ibr-send for the payload type in the SDP offer.
[0235] ibr-recv: When ibr-recv is not offered for a payload type, the answerer may include ibr-send for the payload type in the SDP answer. When ibr-recv is offered for a payload type and the payload type is accepted, the answerer shall include ibr-send in the SDP answer, and the ibr-send shall be identical to or a subset of ibr-recv for the payload type in the SDP offer.
[0236] ibw: When the same bandwidth or bandwidth range is defined for the send and the receive directions, ibw should be used but ibw-send and ibw-recv may also be used, ibw can be used even if the session is negotiated to be sendonly, recvonly, or inactive. For sendonly session, ibw and ibw-send can be interchangeably used. For recvonly session, ibw and ibw-recv can be interchangeably used. When ibw is not offered for a payload type, the answerer may include ibw for the payload type in the SDP answer. When ibw is offered for a payload type and the payload type is accepted, the answerer shall include ibw in the SDP answer, which shall be identical to or a subset of ibw for the payload type in the SDP offer.
[0237] ibw-send: When ibw-send is not offered for a payload type, the answerer may include ibw-recv for the payload type in the SDP answer. When ibw-send is offered for a payload type and the payload is accepted, the answerer shall include ibw-recv in the SDP answer, and the ibw-recv shall be identical to or a subset of ibw-send for the payload type in the SDP offer.
[0238] ibw-recv: When ibw-recv is not offered for a payload type, the answerer may include ibw-send for the payload type in the SDP answer. When ibw-recv is offered for a payload type and the pay load is accepted, the answerer shall include ibw-send in the SDP answer, and the ibw-send shall be identical to or a subset of ibw-recv for the payload type in the SDP offer.
[0239] cf: The SDP offer [shall] list at least one but may list several IVAS Immersive mode coded formats. The SDP answer shall include at least one IVAS Immersive mode coded format and should respond with the one most preferred coded format from the list in the SDP offer. If more than one format is present in the SDP answer, the first format shall be used at the start of a session and may only be modified by the adaptation mechanisms present in this specification. When the same IVAS Immersive mode coded formats are defined for the send and the receive directions, cf should be used but cf-send and cf-recv may also be used. For sendonly session, cf and cf-send can be interchangeably used. For recvonly session, cf and cf-recv can be interchangeably used.
[0240] cf-send: When cf-send is offered for a pay load type and the pay load type is accepted, the answerer shall include cf-recv in the SDP answer, and the cf-recv shall be identical to or a subset of the cf-send parameter for the pay load type in the SDP offer.
[0241] cf-recv: When cf-recv is offered for a payload type and the payload type is accepted, the answerer shall include cf-send in the SDP answer, and the cf-send shall be identical to or a subset of the cf-recv parameter for the payload type in the SDP offer.
[0242] pi-types: The SDP offer shall list at least one but may list several supported pi types when pi data is enabled in the offer. When one or more of the offered pi types are supported, the SDP answer shall be identical to or a subset of the pi types listed in the SDP offer. When the same pi types are defined for the send and the receive directions, pi-types should be used but pi-types-send and pi-types-recv may also be used. For sendonly session, pi-types and pi-types-send can be interchangeably used. For recvonly session, pi-types and pi-types-recv can be interchangeably used. When none of the offered pi-types is supported, the answerer shall not include pi-types in the SDP answer.
[0243] pi-types-send: When pi-types-send is offered in the SDP offer and it is accepted, the answerer shall include pi-types-recv in the SDP answer, and the pi-types-recv shall be identical to or a subset of the pi-types-send parameter in the SDP offer.
[0244] pi-types-recv: When pi-types-recv is offered in the SDP offer and it is accepted, the answerer shall include pi-types-send in the SDP answer, and the pi-types-send shall be identical to or a subset of the pi-types-recv parameter in the SDP offer.
[0245] pi-br: When the same bitrate is defined for the send and the receive directions, pi-br should be used but pi-br-send and pi-br-recv may also be used, pi-br can be used even if the session is negotiated to be sendonly, recvonly, or inactive. For sendonly session, pi-br and pi-br-send can be interchangeably used. For recvonly session, pi-br and pi-br-recv can be interchangeably used. When pi-br is not offered in the SDP offer, the answerer shall not include pi-br in the SDP answer. When pi-br is offered in the SDP offer and it is accepted, the answerer shall include pi-br in the SDP answer which shall be identical or lower than pi-br in the SDP offer.
[0246] pi-br-send: When pi-br-send is offered in the SDP offer and it is accepted, the answerer shall include pi-br-recv in the SDP answer, and the pi-br-recv shall be identical or lower than pi-br-send in the SDP offer.
[0247] pi-br-recv: When pi-br-recv is offered in the SDP offer and it is accepted, the answerer shall include pi-br-send in the SDP answer, and the pi-br-send shall be identical or lower than pi-br-recv in the SDP offer.
[0248] cf-recv: When cf-recv is offered for a payload type (typically by lightweight end device), it must list at least SR as one IVAS Immersive mode coded formats. To accept the offer with split rendering used in the subsequent session, the answer shall contain cf_send with SR as the only IVAS Immersive mode coded format.
[0249] cf-send: When cf-send is offered for a pay load type (typically by a pre-rendering node / device other than lightweight end device), it must list at least SR as one IVAS Immersive mode coded formats. To accept the offer with split rendering used in the subsequent session, the answer shall contain cf_recv with SR as the only IVAS Immersive mode coded format.
[0250] sr-dof: When cf-recv or cf-send is offered for a pay load type with SR listed, the offer may additionally contain the parameter sr-dof. In that case and if the SR session is accepted, the answerer shall include sr-dof in the SDP answer, and the sr-dof shall be identical or lower than sr-dof in the SDP offer. If the first value in the answer is reduced to -1, the second value shall not be present. The answerer may add but shall not remove a * suffix unless the value is smaller than 1.
[0251] sr-tc: When cf-recv or cf-send is offered for a payload type with SR listed, the offer shall additionally contain the parameter sr-tc. In that case and if the SR session is accepted, the answerer shall include sr-tc in the SDP answer, and the sr-tc parameter shall be identical except for the field specifying the bitrate value. If a bitrate value was specified in the offer, the same or a lower bitrate value out of the set of available bitrates may be used in the answer. If not present, the answer may leave that field open or specify a bitrate.
[0252] As split rendering sessions are typically limited to one direction between two directly connected nodes / end points of an IVAS codec session, only directional parameters shall be used to negotiate an IVAS session with split rendering on that connection. When a SR session is accepted, the SDP answer shall not contain any other session parameters for the direction towards the lightweight end device than the split rendering related parameters listed above, PI data related parameters and cmr. The other direction (from the lightweight device) remains unconstrained with the only exception that directional parameters shall be used to specify the session on it.
[0253] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0254] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a randomaccess memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0255] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0256] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.
[0257] EEE1. A method for adapting an immersive audio pay load for RTP transport between a nearend device and an end device, the method for the near-end device comprising: receiving splitrendering control information from the end device (101) over the communication link (111); adapting a rendering operation or mode of the near-end device (102) based on the received splitrendering control information; preparing, responsive to the received split-rendering control information, a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device (102) and the end device (101), wherein the first RTP payload is associated with a split rendering mode for the end device (101); and transmitting the first RTP pay load to the end device (101) through the communication link (111).
[0258] EEE2. The method of EEE1, further comprising: receiving second split-rendering control information from the end device (101); preparing, based on the received second split-rendering control information, a second RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device and the end device, wherein the second RTP payload is responsively pre-rendered based on the received second split-rendering control information and is associated with the split rendering mode for both the near-end device and the end device; and transmitting the second RTP pay load to the end device through the communication link to enable split-rendering on the end device.
[0259] EEE3. The method of any one of EEE1 to EEE2, wherein a table of contents (ToC) byte is provided in the first RTP payload to signal a split rendering mode.
[0260] EEE4. The method of EEE3, wherein a split rendering ToC (SR ToC) byte is provided sequentially to the ToC byte to signal a bit rate for the split rendering mode.
[0261] EEE5. The method of EEE4, wherein the SR ToC byte includes one or more bit fields to indicate one or more of a split rendering configuration mode, a bit rate for split rendering, and whether any additional SR Toe bytes will follow.
[0262] EEE6. The method of any one of EEE1 to EEE5, wherein an initial extra (E)-byte is provided in the first RTP payload to signal a codec mode request (CMR) associated with a codec mode used in upstream direction from the end device.
[0263] EEE7. The method of any one of EEE1 to EEE6, wherein a subsequent E-byte is provided in the first RTP pay load to signal an audio bandwidth request or an audio format request associated with a codec mode used in upstream direction from the end device.
[0264] EEE8. The method of any one of EEE1 to EEE7, wherein a subsequent E-byte is provided in the first RTP pay load to signal a presence of processing information in the RTP pay load.
[0265] EEE9. The method of any one of EEE1 to EEE8, wherein the split-rendering control information includes at least one selected from the group consisting of a bit rate request (1200, 1500), a split renderer configuration request (1900), and an indication of whether to enable diegetic support for the at least one stream of coded binaural audio.
[0266] EEE10. The method of any one of EEE1 to EEE9, further comprising: receiving, from the far-end device (103), an immersive audio payload, which is processed in response to the received split-rendering control information in preparation of the first RTP payload including the at least one stream of coded binaural audio and the related split rendering metadata.
[0267] EEE11. A method for processing an immersive audio payload for RTP transport received from an upstream device by a downstream device, the method for the downstream device comprising: receiving a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device by a communication link, wherein the first RTP payload is associated with a split rendering mode for the upstream device and the downstream device; applying a first split rendering operating mode of the downstream device based on the first RTP payload from the upstream device; and transmitting split-rendering control information from the downstream device in a reverse direction over the communication link to the upstream device.
[0268] EEE12. The method of EEE11, further comprising: receiving a second RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device by the communication link, wherein the second RTP payload is associated with the split rendering mode for the downstream device and is responsive to pre-rendering based on the split-rendering control information from the downstream device; and by the downstream device rendering audio data from the second RTP pay load based on the adjusted rendering operating mode.
[0269] EEE13. The method of any one of EEE10 to EEE12, wherein a table of contents (ToC) byte is provided in the first RTP payload to signal a split-rendering mode.
[0270] EEE14. The method of EEE13, wherein a split rendering ToC (SR ToC) byte is provided sequentially to the ToC byte to signal a bit rate for the split-rendering mode.
[0271] EEE15. The method of EEE14, wherein the SR ToC byte includes one or more bit fields to indicate one or more of a split-rendering configuration mode, a bit rate for split-rendering, and whether any additional SR Toe bytes will follow.
[0272] EEE16. The method of any one of EEE13 to EEE15, wherein an initial E-byte is provided in the first RTP pay load to signal a codec mode request (CMR) associated with a codec mode used in an upstream direction from the downstream device.
[0273] EEE17. The method of any one of EEE11 to EEE16, wherein a CMR byte is provided in the first RTP payload to signal an IVAS bit rate in the upstream direction.
[0274] EEE18. The method of any one of EEE11 to EEE17, wherein an IVAS configuration request (IVAS-cr) byte is provided in the first RTP payload to signal an audio bandwidth request and an audio format request associated with a codec mode used in upstream direction from the downstream device.
[0275] EEE 19. The method of any one of EEE11 to EEE18, wherein an extra (E) byte is provided in the first RTP pay load to signal a presence of processing information in the RTP pay load.
[0276] EEE20. The method of any one of EEE11 to EEE 19, wherein the split-rendering control information includes at least one selected from the group consisting of a bit rate request, a split renderer configuration request, and an indication of whether to enable diegetic support for the at least one stream of coded binaural audio.
[0277] EEE21. The method of any one of EEE1 to EEE20, wherein the RTP pay loads are fully compatible with the EVS codec RTP pay load format.
[0278] EEE22. A method for generating an immersive audio payload for RTP transport, the method comprising: generating a ToC byte for inclusion in a first RTP payload; setting a bit rate (BR) field included in the ToC byte to indicate that the first RTP payload is a split rendering payload; generating a split rendering ToC byte (SR ToC) for inclusion in the first RTP payload sequential to the ToC byte; setting a split rendering bit rate (SR-BR) field included in the SR ToC byte to indicate an IVAS split rendering bit rate; and preparing the first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between an upstream device and a downstream device.
[0279] EEE23. The method of EEE22, wherein the SR ToC byte includes one or more bit fields to indicate one or more of a split rendering configuration mode and the SR-BR field.
[0280] EEE24. A method for RTP transport for an adaptive immersive audio payload using a common payload format, the method comprising: preparing a first RTP payload with audio data for at least one stream of coded immersive audio for communication between an upstream device, a first downstream device, and a second downstream device, wherein the first RTP payload is associated with a first immersive audio coding mode; transmitting the first RTP payload to the second downstream device through a first communication link; processing the first RTP payload by the second downstream device to generate an adapted first RTP payload, wherein the first RTP payload is decoded according to the first immersive audio coding mode, pre-rendered and re-encoded according to a split rendering mode to a split rendering representation comprising coded binaural audio and related split rendering metadata; preparing the adapted first RTP payload for RTP transport for communication between the second downstream device and the first downstream device using the same pay load format, wherein the adapted first RTP pay load is associated with the used split rendering coding mode; transmitting the adapted first RTP pay load to the first downstream device through a second communication link; and receiving and processing the adapted first RTP payload by the first downstream device, wherein the adapted first payload is decoded and post-rendered according to the associated split rendering mode.
[0281] EEE25. The method of EEE24, further comprising transmitting a second RTP pay load including split-rendering control information from the first downstream device in a reverse direction over the second communication link to the second downstream device.
[0282] EEE26. The method of EEE25, further comprising: receiving the second RTP payload from the first downstream device over the second communication link; extracting the split-rendering control information from the second RTP payload; implementing, by the second downstream device, at least a first portion of the split-rendering control information; processing the second RTP payload by the second downstream device to generate an adapted second RTP payload including at least a second portion of the second RTP payload including coded upstream audio data; and transmitting the adapted second RTP payload to the upstream device through the first communication link.
[0283] EEE27. The method of EEE26, wherein processing the second RTP payload further includes adding additional control information used in a downstream operation from the upstream device.
[0284] EEE28. The method of EEE24, further comprising: preparing a second RTP payload for RTP transport for communication between the second downstream device and the upstream device, wherein the second RTP payload includes control information for the upstream device that was included in a third RTP payload transmitted from the first downstream device to the second downstream device; and transmitting the second RTP payload to the upstream device through the first communication link.
[0285] EEE29. An apparatus comprising: an electronic processor configured to perform operations including the method of any one of EEE1 to EEE28.
[0286] EEE30. A non-transitory computer-readable storage medium recording a program of instructions that is executable by a device to perform the method of any one of EEE1 to EEE28.
[0287] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0288] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0289] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0290] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
Claims
1. A method (2000) for adapting an immersive audio payload for RTP transport between a near-end device (102) and an end device (101), the method (2000) for the near-end device (102) comprising:receiving split-rendering control information from the end device (101) over the communication link (111);adapting a rendering operation or mode of the near-end device (102) based on the received split-rendering control information;preparing, responsive to the received split-rendering control information, a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device (102) and the end device (101), wherein the first RTP pay load is associated with a split rendering mode for the end device (101); andtransmitting the first RTP pay load to the end device (101) through the communication link (Hl).
2. The method (2000) of claim 1, further comprising:receiving second split-rendering control information from the end device (101);preparing, based on the received second split-rendering control information, a second RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between the near-end device (102) and the end device (101), wherein the second RTP payload is responsively pre-rendered based on the received second split-rendering control information and is associated with the split rendering mode for both the near-end device (102) and the end device (101); andtransmitting the second RTP payload to the end device (101) through the communication link (111) to enable split-rendering on the end device (101).
3. The method (2000) of claim 1 or claim 2, wherein a table of contents (ToC) byte (1000) in the first RTP payload is provided to signal a split rendering mode.
4. The method (2000) of claim 3, wherein a split rendering ToC (SR ToC) byte (1100) isprovided sequentially to the ToC byte (1000) to signal a bit rate for the split rendering mode.
5. The method (2000) of claim 4, wherein the SR ToC byte (1100) includes one or more bit fields to indicate one or more of a split rendering configuration mode, a bit rate for split rendering, and whether any additional SR Toe bytes (1100) will follow.
6. The method (2000) of any of one of claims 1-5, wherein an initial extra (E)-byte (1400) is provided in the first RTP pay load to signal a codec mode request (CMR) associated with a codec mode used in upstream direction from the end device (101).
7. The method (2000) of any one of claims 1-6, wherein a subsequent E-byte (1400) is provided in the first RTP payload to signal an audio bandwidth request or an audio format request associated with a codec mode used in upstream direction from the end device (101).
8. The method (2000) of any one of claims 1-7, wherein a subsequent E-byte (1400) is provided in the first RTP payload to signal a presence of processing information in the RTP payload.
9. The method (2000) of any one of claims 1-8, wherein the split-rendering control information includes at least one selected from the group consisting of a bit rate request (1200, 1500), a split renderer configuration request (1900), and an indication of whether to enable diegetic support for the at least one stream of coded binaural audio.
10. The method (2000) of any one of claims 1-9, further comprising: receiving, from the far-end device (103), an immersive audio payload, which is processed in response to the received split-rendering control information in preparation of the first RTP payload including the at least one stream of coded binaural audio and the related split rendering metadata.
11. A method (2100) for processing an immersive audio pay load for RTP transport received from an upstream device (102) by a downstream device (101), the method for the downstream device (101) comprising:receiving a first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device (102) by a communication link (111),wherein the first RTP payload is associated with a split rendering mode for the upstream device (102) and the downstream device (101);applying a first split rendering operating mode of the downstream device (101) based on the first RTP payload from the upstream device (102); andtransmitting split-rendering control information from the downstream device (101) in a reverse direction over the communication link (111) to the upstream device (102).
12. The method (2100) of claim 11, further comprising:receiving a second RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata from the upstream device (102) by the communication link (111), wherein the second RTP payload is associated with an adjusted split rendering mode for the downstream device (101) and is responsive to pre-rendering based on the split-rendering control information from the downstream device (101); andby the downstream device (101) rendering audio data from the second RTP pay load based on the adjusted split rendering operating mode.
13. The method (2100) of any one of claims 10-12, wherein a table of contents (ToC) byte (1000) is provided in the first RTP payload to signal a split-rendering mode.
14. The method (2100) of claim 13, wherein a split rendering ToC (SR ToC) byte (1100) is provided sequentially to the ToC byte (1000) to signal a bit rate for the split-rendering mode.
15. The method (2100) of claim 14, wherein the SR ToC byte (1100) includes one or more bit fields to indicate one or more of a split-rendering configuration mode, a bit rate for split-rendering, and whether any additional SR Toe bytes (1100) will follow.
16. The method (2100) of any one of claims 13-15, wherein an initial E-byte (1400) is provided in the first RTP payload to signal a codec mode request (CMR) associated with a codec mode used in an upstream direction from the downstream device (101).
17. The method (2100) of any one of claims 11-16, wherein a CMR byte (1200) is provided in the first RTP payload to signal an IVAS bit rate in the upstream direction.
18. The method (2100) of any one of claims 11-17, wherein an IVAS configuration request (IVAS-cr) byte (1300) is provided in the first RTP payload to signal an audio bandwidth request and an audio format request associated with a codec mode used in upstream direction from the downstream device (101).
19. The method (2100) of any one of claims 11-18, wherein an extra (E) byte (1400) is provided in the first RTP payload to signal a presence of processing information in the RTP payload.
20. The method (2100) of any one of claims 11-19, wherein the split-rendering control information includes at least one selected from the group consisting of a bit rate request (1200, 1500), a split renderer configuration request (1900), and an indication of whether to enable diegetic support for the at least one stream of coded binaural audio.
21. The method (2100) of any one of claims 1-20, wherein the RTP payloads are fully compatible with the EVS codec RTP pay load format.
22. A method (2200) for generating an immersive audio payload for RTP transport, the method comprising:generating a ToC byte (1000) for inclusion in a first RTP payload;setting a bit rate (BR) field included in the ToC byte (1000) to indicate that the first RTP payload is a split rendering payload;generating a split rendering ToC byte (SR ToC) (1100) for inclusion in the first RTP payload sequential to the ToC byte (1000);setting a split rendering bit rate (SR-BR) field included in the SR ToC byte (1100) to indicate an IVAS split rendering bit rate; andpreparing the first RTP payload with audio data for at least one stream of coded binaural audio and related split rendering metadata for communication between an upstream device (102) and a downstream device (101).
23. The method (2200) of claim 22, wherein the SR ToC byte (1100) includes one or more bit fields to indicate one or more of a split rendering configuration mode and the SR-BR field.
24. A method for RTP transport of an adaptive immersive audio payload using a common payload format, the method comprising:preparing a first RTP payload with audio data for at least one stream of coded immersive audio for communication between an upstream device (103), a first downstream device (101), and a second downstream device (102), wherein the first RTP pay load is associated with a first immersive audio coding mode;transmitting the first RTP pay load to the second downstream device (102) through a first communication link (112);processing the first RTP pay load by the second downstream device (102) to generate an adapted first RTP payload, wherein the first RTP payload is decoded according to the first immersive audio coding mode, pre-rendered and re-encoded according to a split rendering mode to a split rendering representation comprising coded binaural audio and related split rendering metadata;preparing the adapted first RTP payload for RTP transport for communication between the second downstream device (102) and the first downstream device (101) using the same pay load format, wherein the adapted first RTP payload is associated with the used split rendering coding mode;transmitting the adapted first RTP pay load to the first downstream device (101) through a second communication link (111); andreceiving and processing the adapted first RTP payload by the first downstream device (101), wherein the adapted first payload is decoded and post-rendered according to the associated split rendering mode.
25. The method of claim 24, further comprising transmitting a second RTP pay load including split-rendering control information from the first downstream device (101) in a reverse direction over the second communication link (111) to the second downstream device (102).
26. The method of claim 25, further comprising:receiving the second RTP pay load from the first downstream device (101) over the second communication link (111);extracting the split-rendering control information from the second RTP payload;implementing, by the second downstream device (102), at least a first portion of the splitrendering control information;processing the second RTP pay load by the second downstream device (102) to generate an adapted second RTP payload including at least a second portion of the second RTP payload including coded upstream audio data; andtransmitting the adapted second RTP payload to the upstream device (103) through the first communication link (112).
27. The method of claim 26, wherein processing the second RTP payload further includes adding additional control information used in a downstream operation from the upstream device (103).
28. The method of claim 24, further comprising:preparing a second RTP payload for RTP transport for communication between the second downstream device (102) and the upstream device (103), wherein the second RTP payload includes control information for the upstream device (103) that was included in a third RTP payload transmitted from the first downstream device (101) to the second downstream device (102); andtransmitting the second RTP payload to the upstream device (103) through the first communication link (112).
29. An apparatus (2500) comprising:an electronic processor (2520) configured to perform operations including the method of any one of claims 1-28.
30. A non-transitory computer-readable storage medium (2521) recording a program of instructions that is executable by a device to perform the method of any one of claims 1-28.