Network rendering and transcoding of augmented reality data

Network rendering techniques allow devices lacking advanced capabilities to participate in AR sessions by converting 3D content to 2D using an AR application server, overcoming rendering limitations.

JP2026505325APending Publication Date: 2026-02-13QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025545102
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-12
Filing Date
2024-02-13
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Certain devices, such as AR glasses and HMDs, lack the advanced rendering capabilities or require excessive power to handle complex augmented reality (AR) content, limiting their participation in AR communication sessions.

Method used

Implement network rendering (split rendering) using an AR application server to render AR data on behalf of devices that lack processing capabilities, enabling them to participate in AR sessions by converting 3D content to 2D.

Benefits of technology

Enables devices without advanced rendering capabilities to participate in AR sessions by offloading rendering tasks to a capable server, ensuring seamless AR communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505325000001_ABST
    Figure 2026505325000001_ABST
Patent Text Reader

Abstract

An exemplary first user equipment (UE) for communicating media data includes a memory configured to store the media data and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: send a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establish a media communication session with a transcoding device performing media functions or multimedia resource functions, the transcoding device being between the first UE and the second UE; receive transcoded media data from the transcoding device, the transcoding device transcoding the AR media data received from the second UE; and present the transcoded media data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Patent Application No. 18 / 438,992, filed February 12, 2024, U.S. Provisional Patent Application No. 63 / 484,564, filed February 13, 2023, and U.S. Provisional Patent Application No. 63 / 596,869, filed November 7, 2023, the entire contents of each of which are incorporated by reference. U.S. Patent Application No. 18 / 438,992, filed February 12, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 484,564, filed February 13, 2023, and U.S. Provisional Patent Application No. 63 / 596,869, filed November 7, 2023.

[0002] FIELD This disclosure relates to the storage and transport of encoded video data. [Background technology]

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radiotelephones, video teleconferencing devices, etc. Digital video devices implement video compression techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also referred to as High Efficiency Video Coding (HEVC)), and extensions to such standards, to more efficiently transmit and receive digital video information.

[0004] Video compression techniques perform spatial and / or temporal prediction to reduce or remove redundancy inherent in video sequences. In block-based video coding, video frames or slices may be partitioned into macroblocks. Each macroblock may be further partitioned. Macroblocks in intra-coded (I) frames or slices are coded using spatial prediction with respect to neighboring macroblocks. Macroblocks in inter-coded (P or B) frames or slices may use spatial prediction with respect to neighboring macroblocks in the same frame or slice, or temporal prediction with respect to other reference frames.

[0005] After the video data is encoded, it can be packetized for transmission or storage and assembled into a video file that conforms to any of a variety of standards, such as the International Organization for Standardization (ISO) Base Media File Format and its extensions, e.g., AVC. Summary of the Invention

[0006] Generally, this disclosure describes techniques for communicating augmented reality (AR) data between two or more user equipment (UE) devices. In particular, a UE may not be capable of rendering three-dimensional (3D) visual content, such as 3D virtual object data, into two-dimensional (2D) visual content (e.g., image or video data). Therefore, to participate in an AR communication session, a separate device, such as an application server (AS) in a 5G network, may be configured to render the 3D visual content into 2D visual content in accordance with the techniques of this disclosure. The UE may then present the 2D visual content.

[0007] In one embodiment, a method for communicating media data includes: sending, by a first user equipment (UE), a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding AR media data into two-dimensional video data; establishing, by the first UE, a media communication session with a transcoding device performing a media function or a multimedia resource function, the transcoding device being between the first UE and the second UE; receiving, by the first UE, transcoded media data from the transcoding device, the transcoding device transcoding from the AR media data received from the second UE; and presenting, by the first UE, the transcoded media data.

[0008] In another embodiment, a first user equipment (UE) for communicating media data includes a memory configured to store the media data and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: send a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establish a media communication session with a transcoding device performing media functions or multimedia resource functions, the transcoding device being between the first UE and the second UE; receive transcoded media data from the transcoding device, the transcoding device transcoding the AR media data received from the second UE; and present the transcoded media data.

[0009] In another embodiment, a method for communicating media data includes receiving, by a transcoding device performing a media function or a multimedia resource function, a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; in response to receiving, by the transcoding device, the AR media data from the second UE, rendering, by the transcoding device, the AR media data to form rendered 2D media data; and transmitting, by the transcoding device, the rendered 2D media data to the first UE.

[0010] In another embodiment, a transcoding device that performs a media function or a multimedia resource function includes a memory configured to store media data and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: receive a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; and, in response to receiving the AR media data from the second UE, render the AR media data to form rendered 2D media data; and transmit the rendered 2D media data to the first UE.

[0011] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram illustrating an example system implementing techniques for streaming media data over a network. [Figure 2] FIG. 2 is a block diagram illustrating elements of an exemplary video file. [Figure 3]FIG. 1 is a block diagram illustrating an example network including various devices for performing the techniques of this disclosure. [Figure 4] FIG. 10 is a flow diagram illustrating an example procedure for establishing an AR call in accordance with the techniques of this disclosure. [Figure 5] FIG. 1 is a conceptual diagram illustrating an exemplary IP Multimedia Subsystem (IMS) architecture. [Figure 6] FIG. 6 is a flow diagram illustrating an example AR call setup process that may be used by the architecture of FIG. 5. [Figure 7] FIG. 6 is a flow diagram illustrating an example AR call setup process that may be used by the architecture of FIG. 5. [Figure 8] FIG. 10 is a flow diagram illustrating an example call setup procedure that may enable remote rendering of AR content based on a request from a sender of the AR content. [Figure 9] FIG. 1 is a call flow diagram illustrating an example process for setting up an AR call in accordance with techniques of this disclosure. [Figure 10] 1 is a flowchart illustrating an example method that may be performed by a client device, e.g., a user equipment (UE) device, to request transcoding of AR media data in accordance with techniques of this disclosure. [Figure 11] 1 is a flowchart illustrating an example method for transcoding AR media data in accordance with techniques of this disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Augmented reality (AR) calls (or other extended reality (XR) media communication sessions, such as mixed reality (MR) or virtual reality (VR)) can require significant processing resources to render the content of the AR call scene, especially when multiple participants contribute to the creation of a complex AR call scene. These scenes can include virtual environments that can be anchored to real-world locations, as well as content from all participants in the call. Content from participants can include, for example, user avatars, slide materials, 3D virtual objects, etc.

[0014] Physically based rendering (PBR) can be included in rendering AR or other XR data. PBR generally involves rendering image data by emulating light transmission in the virtual world to reproduce real-world lighting, including user shadows and specular object reflections. Advanced rendering capabilities such as PBR may not be available for certain devices, such as AR glasses and head-mounted displays (HMDs), or may require too much power to operate on such devices.

[0015] Techniques of this disclosure include using signaling to invoke split rendering (also referred to herein as “network rendering”) for AR calls over an IP Multimedia Subsystem (IMS). Using such signaling, a client device (e.g., user equipment (UE)) may signal another device that the client device is requesting split rendering; the other device will render image data from AR data for the AR media communication session, and the client device will present the rendered image data. The client device may be, for example, an HMD, AR glasses, etc. The other device may be an AR application server (AS). In this way, such a device may be able to participate in the AR communication session even when such a device is not capable of rendering the AR data. Additionally, an upstream device capable of rendering AR data may receive a request to render the AR data on behalf of another device, such as a UE, HMD, or AR glasses, and render the AR data on behalf of the other device, thereby achieving split rendering.

[0016] This disclosure describes techniques that can be used to handle network rendering (e.g., split rendering) of AR call session data, for example, as a transcoding operation. Network rendering can be triggered by an IMS application server (AS) for client devices that do not have the necessary AR processing capabilities. A media function (MF) or multimedia resource function (MRF) can act on behalf of an endpoint (e.g., UE-A) to generate corresponding AR content, for example, to place 2D overlay videos of participants on a 3D screen in a virtual 3D scene.

[0017] 1 is a block diagram illustrating an example system 10 that implements techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. Client device 40 and server device 60 are communicatively coupled by a network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled by network 74 or another network, or may be communicatively coupled directly. In some examples, content preparation device 20 and server device 60 may comprise the same device.

[0018] 1 includes an audio source 22 and a video source 24. Audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by audio encoder 26. Alternatively, audio source 22 may include a storage medium storing previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. Video source 24 may include a video camera, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data that generates video data to be encoded by video encoder 28. Content preparation device 20 is not necessarily communicatively coupled to server device 60 in all embodiments; multimedia content may also be stored on a separate medium that is read by server device 60.

[0019] The raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may acquire audio data from a speaking participant while the speaking participant is speaking, and video source 24 may simultaneously acquire video data of the speaking participant. In other embodiments, audio source 22 may include a computer-readable storage medium containing stored audio data, and video source 24 may include a computer-readable storage medium containing stored video data. In this manner, the techniques described in this disclosure may be applied to live, streaming, real-time audio and video data, or to archived, pre-recorded audio and video data.

[0020] An audio frame corresponding to a video frame is generally an audio frame that includes audio data captured (or generated) by audio source 22 contemporaneously with the video data captured (or generated) by video source 24 contained within that video frame. For example, a speaking participant typically generates audio data by speaking while audio source 22 captures that audio data, and video source 24 captures video data of the speaking participant contemporaneously, i.e., while audio source 22 is capturing the audio data. Thus, an audio frame may correspond in time to one or more particular video frames. Thus, an audio frame corresponding to a video frame generally corresponds to a situation in which the audio data and video data are captured contemporaneously, and the audio frame and video frame contain the contemporaneously captured audio data and video data, respectively.

[0021] In some embodiments, audio encoder 26 may encode, in each encoded audio frame, a timestamp representing the time the audio data for that encoded audio frame was recorded, and similarly, video encoder 28 may encode, in each encoded video frame, a timestamp representing the time the video data for that encoded video frame was recorded. In such embodiments, an audio frame corresponding to a video frame may include the audio frame including a timestamp and the video frame including the same timestamp. Content preparation device 20 may include internal clocks that enable audio encoder 26 and / or video encoder 28 to generate timestamps, or audio source 22 and video source 24 may include internal clocks that can be used by audio encoder 26 and / or video encoder 28 to associate audio data and video data with timestamps, respectively.

[0022] In some embodiments, audio source 22 may send data to audio encoder 26 corresponding to the time the audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to the time the video data was recorded. In some embodiments, audio encoder 26 may encode a sequence identifier in the encoded audio data that indicates the relative temporal order of the encoded audio data, but not necessarily the absolute time the audio data was recorded; similarly, video encoder 28 may also use a sequence identifier to indicate the relative temporal order of the encoded video data. Similarly, in some embodiments, the sequence identifier may be mapped to or otherwise correlated to a timestamp.

[0023] The audio encoder 26 typically generates a stream of coded audio data, while the video encoder 28 generates a stream of coded video data. Each individual stream of data (whether audio or video) may be referred to as an elementary stream. An elementary stream is a single, digitally coded (and possibly compressed) component of a media presentation. For example, a coded video or audio portion of a media presentation may be an elementary stream. An elementary stream may be converted into a packetized elementary stream (PES) before being encapsulated within a video file. Within the same media presentation, a stream ID may be used to distinguish PES packets belonging to one elementary stream from others. The basic unit of data for an elementary stream is the packetized elementary stream (PES) packet. Therefore, coded video data typically corresponds to an elementary video stream. Similarly, audio data corresponds to one or more respective elementary streams.

[0024] 1, encapsulation unit 30 of content preparation device 20 receives an elementary stream including coded video data from video encoder 28 and an elementary stream including coded audio data from audio encoder 26. In some embodiments, video encoder 28 and audio encoder 26 may each include a packetizer for forming PES packets from the coded data. In other embodiments, video encoder 28 and audio encoder 26 may each interface with a corresponding packetizer for forming PES packets from the coded data. In still other embodiments, encapsulation unit 30 may include packetizers for forming PES packets from the coded audio and video data.

[0025] The video encoder 28 may encode the video data of the multimedia content in various manners to generate different representations of the multimedia content at various bit rates and with different characteristics, such as pixel resolution, frame rate, compliance with various coding standards, compliance with various profiles and / or levels of profiles for various coding standards, representations with one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. As used in this disclosure, a representation may include one of audio data, video data, text data (e.g., for closed captioning), or other such data. A representation may include an elementary stream, such as an audio elementary stream or a video elementary stream. Each PES packet may include a stream_id, which identifies the elementary stream to which the PES packet belongs. The encapsulation unit 30 is responsible for assembling the elementary streams into streamable media data.

[0026] Encapsulation unit 30 receives PES packets for the elementary streams of the media presentation from audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. Coded video segments can be organized into NAL units, which provide a "network-friendly" video representation that addresses applications such as video telephony, storage, broadcast, or streaming. NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may contain block-, macroblock-, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some embodiments, a coded picture at a time instance, typically presented as a primary coded picture, can be included within an access unit, which may contain one or more NAL units.

[0027] Non-VCL NAL units may include, among other things, parameter set NAL units and SEI NAL units. Parameter sets may contain sequence-level header information (in sequence parameter sets (SPS)) and picture-level header information that does not change frequently (in picture parameter sets (PPS)). Parameter sets (e.g., PPS and SPS) allow information that does not change frequently to not need to be repeated for each sequence or picture, thus improving coding efficiency. Furthermore, the use of parameter sets may avoid the need for redundant transmission for error resilience by enabling out-of-band transmission of important header information. In an embodiment of out-of-band transmission, parameter set NAL units may be transmitted on a different channel from other NAL units, such as SEI NAL units.

[0028] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding coded picture samples from VCL NAL units, but that can assist processes related to decoding, display, error resilience, and other purposes. SEI messages can be included in non-VCL NAL units. SEI messages are normative parts of some standard specifications and are therefore not necessarily mandatory for the implementation of standard-compliant decoders. SEI messages can be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information can be included in SEI messages, such as the scalability information SEI message in SVC embodiments and the view scalability information SEI message in MVC. These exemplary SEI messages can convey information about, for example, operation point extraction and operation point characteristics.

[0029] Server device 60 includes a Real-time Transport Protocol (RTP) transmission unit 70 and a network interface 72. In some embodiments, server device 60 may include multiple network interfaces. Furthermore, any or all of the features of server device 60 may be implemented on other devices in the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some embodiments, intermediate devices in the content delivery network may cache data for multimedia content 64 and may include components that substantially conform to those of server device 60. Generally, network interface 72 is configured to transmit and receive data over network 74.

[0030] In the example of FIG. 1 , according to techniques of this disclosure, server device 60 includes a rendering unit 66. Rendering unit 66 may be configured to receive extended reality (XR) media data, including, for example, augmented reality (AR) media data, mixed reality (MR) media data, or virtual reality (VR) media data, and to render image data on behalf of client device 40 using the XR media data and multimedia content 64. For example, client device 40 and other client devices participating in the AR call may provide XR media data such as pose information, avatars, documents (e.g., slides for a slide deck for a presentation), 3D virtual objects, etc. to server device 60. Multimedia content 64 may correspond to other virtual object data or pre-rendered image data. Finally, rendering unit 66 may combine various XR media data from multiple client devices and render image or video data using the XR media data for client device 40.

[0031] Although shown as forming part of server device 60, in other embodiments, rendering unit 66 may be included in content preparation device 20, e.g., as part of video source 24. For example, video source 24 may include scene data for a virtual scene, such as a background, virtual objects (e.g., tables, chairs, walls, a presentation screen, etc.), and virtual light sources, as well as avatars from other users (e.g., users of client device 40), user virtual objects, pose information about the users, etc. Ultimately, the rendering unit of content preparation device 20 may render 2D images from these virtual objects and provide the 2D images to video encoder 28 to be encoded. In this manner, content preparation device 20 may render images for presentation by client device 40 using split rendering.

[0032] The RTP transmission unit 70 is configured to deliver media data, including rendered XR / AR / MR / VR media content, to the client device 40 over the network 74 in accordance with RTP, which is standardized in Request for Comment (RFC) 3550 by the Internet Engineering Task Force (IETF). The RTP transmission unit 70 may also implement protocols related to RTP, such as the RTP Control Protocol (RTCP), the Real-time Streaming Protocol (RTSP), the Session Initiation Protocol (SIP), and / or the Session Description Protocol (SDP). The RTP transmission unit 70 may transmit the media data over a network interface 72, which may implement the Uniform Datagram Protocol (UDP) and / or the Internet Protocol (IP). Thus, in some embodiments, the server device 60 may transmit the media data over RTP and RTSP over UDP using the network 74.

[0033] The RTP sending unit 70 may receive an RTSP description request, for example, from the client device 40. The RTSP description request may include data indicating what types of data are supported by the client device 40. The RTP sending unit 70 may respond to the client device 40 with data indicating media streams, such as media content 64, that may be sent to the client device 40 along with corresponding network location identifiers, such as uniform resource locators (URLs) or uniform resource names (URNs).

[0034] The RTP sending unit 70 may then receive an RTSP setup request from the client device 40. The RTSP setup request may generally indicate how the media stream should be transported. The RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. The RTP sending unit 70 may reply to the RTSP setup request with a confirmation and data indicating the port on the server device 60 to which the RTP data and control data will be sent. The RTP sending unit 70 may then receive an RTSP play request to “play” the media stream, i.e., to send the media stream to the client device 40 over the network 74. The RTP sending unit 70 may also receive an RTSP teardown request to terminate the streaming session, in response to which the RTP sending unit 70 may stop sending media data to the client device 40 for the corresponding session.

[0035] RTP receiving unit 52 may similarly initiate a media stream by first sending an RTSP description request to server device 60. The RTSP description request may indicate the type of data supported by client device 40. RTP receiving unit 52 may then receive a reply from server device 60 specifying available media streams, such as media content 64, that may be sent to client device 40, along with corresponding network location identifiers, such as uniform resource locators (URLs) or uniform resource names (URNs).

[0036] RTP receiving unit 52 may then generate an RTSP setup request and send the RTSP setup request to server device 60. As described above, the RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on client device 40. In response, RTP receiving unit 52 may receive a confirmation from server device 60 that includes the port of server device 60 that server device 60 will use to transmit the media data and control data.

[0037] After establishing a media streaming session between server device 60 and client device 40, an RTP sending unit 70 of server device 60 may transmit media data (e.g., packets of media data) to client device 40 according to the media streaming session. Server device 60 and client device 40 may exchange control data (e.g., RTCP data), for example, indicating reception statistics by client device 40, thereby enabling server device 60 to perform congestion control or otherwise diagnose and address transmission failures.

[0038] Network interface 54 may receive and provide media for a selected media presentation to RTP receiving unit 52, which in turn may provide the media data to de-encapsulation unit 50. De-encapsulation unit 50 may de-encapsulate elements of the video file into constituent PES streams, de-packetize the PES streams to remove encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending, for example, on whether the encoded data is part of an audio stream or a video stream, as indicated by the PES packet headers for that stream. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of a stream, to video output 44.

[0039] Each of the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiving unit 52, and decapsulation unit 50, as applicable, may be implemented as any of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. Each of the video encoder 28 and video decoder 48 may be included within one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and audio decoder 46 may be included within one or more encoders or decoders, any of which may be integrated as part of a combined CODEC. An apparatus including the video encoder 28, the video decoder 48, the audio encoder 26, the audio decoder 46, the encapsulation unit 30, the RTP receiving unit 52, and / or the decapsulation unit 50 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular telephone.

[0040] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate in accordance with the techniques of this disclosure. For illustrative purposes, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may be configured to perform these techniques instead of (or in addition to) server device 60.

[0041] Encapsulation unit 30 can form NAL units, which include a header that identifies the program to which the NAL unit belongs and a payload, e.g., audio data, video data, or data describing the transport stream or program stream to which the NAL unit corresponds. For example, in H.264 / AVC, an NAL unit includes a one-byte header and a variable-sized payload. NAL units that include video data within their payloads can contain video data at various levels of granularity. For example, an NAL unit can include a block of video data, multiple blocks, a slice of video data, or an entire picture of video data. Encapsulation unit 30 can receive encoded video data from video encoder 28 in the form of PES packets of elementary streams. Encapsulation unit 30 can associate each elementary stream with a corresponding program.

[0042] Encapsulation unit 30 can also assemble access units from multiple NAL units. Generally, an access unit may include one or more NAL units for representing a frame of video data, as well as audio data corresponding to that frame, if such audio data is available. An access unit generally includes all NAL units for an output time instance, e.g., all audio and video data for a time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, a particular frame for all views of the same access unit (same time instance) can be rendered simultaneously. In one embodiment, an access unit may include a coded picture at a time instance, which may be presented as a primary coded picture.

[0043] Thus, an access unit may include all audio and video frames of a common time instance, e.g., all views corresponding to time X. This disclosure also refers to the coded pictures of a particular view as a "view component." That is, a view component may include coded pictures (or frames) for a particular view at a particular time. Thus, an access unit may be defined as including all view components of a common time instance. The decoding order of access units is not necessarily the same as the output order or display order.

[0044] After encapsulation unit 30 assembles NAL units and / or access units into a video file based on the received data, encapsulation unit 30 passes the video file to output interface 32 for output. In some embodiments, encapsulation unit 30 may store the video file locally or transmit the video file to a remote server via output interface 32 rather than transmitting the video file directly to client device 40. Output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium such as an optical drive, a magnetic media drive (e.g., a floppy drive), a universal serial bus (USB) port, a network interface, or other output interface. Output interface 32 outputs the video file to a computer-readable medium such as, for example, a transmission signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable medium.

[0045] Network interface 54 may receive NAL units or access units via network 74 and provide the NAL units or access units to de-encapsulation unit 50 via RTP receiving unit 52. De-encapsulation unit 50 may de-encapsulate elements of the video file into constituent PES streams, de-packetize the PES streams to extract encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending, for example, on whether the encoded data is part of an audio stream or a video stream, as indicated by the stream's PES packet headers. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of a stream, to video output 44.

[0046] FIG. 2 is a block diagram illustrating elements of an exemplary video file 150. As described above, video files according to the ISO Base Media File Format and its extensions store data in a series of objects called "boxes." In the example of FIG. 2, video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. While FIG. 2 represents one example of a video file, it should be understood that other media files may contain other types of media data (e.g., audio data, timed text data, etc.) that are structured similarly to the data in video file 150 according to the ISO Base Media File Format and its extensions.

[0047] The file type (FTYP) box 152 generally represents the file type for the video file 150. The file type box 152 may contain data identifying specifications that describe best use for the video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the movie fragment box 164, and / or the MFRA box 166.

[0048] 2, MOOV box 154 includes a movie header (MVHD) box 156, a track (TRAK) box 158, and one or more movie extends (MVEX) boxes 160. In general, MVHD box 156 may describe general characteristics of video file 150. For example, MVHD box 156 may include data describing when video file 150 was originally created, data describing when video file 150 was last modified, data describing a timescale for video file 150, data describing a playback duration for video file 150, or other data that generally describes video file 150.

[0049] TRAK box 158 may contain data about a track of video file 150. TRAK box 158 may include a track header (TKHD) box that describes characteristics of the track corresponding to TRAK box 158. In some embodiments, TRAK box 158 may contain coded video pictures, while in other embodiments, the coded video pictures for that track may be contained within a movie fragment 164 that may be referenced by data in TRAK box 158 and / or sidx box 162.

[0050] In some embodiments, video file 150 may include two or more tracks. Thus, MOOV box 154 may include a number of TRAK boxes equal to the number of tracks in video file 150. TRAK box 158 may describe characteristics of the corresponding track in video file 150. For example, TRAK box 158 may describe temporal and / or spatial information about the corresponding track. When encapsulation unit 30 (FIG. 1) includes a parameter set track in a video file, such as video file 150, a TRAK box similar to TRAK box 158 in MOOV box 154 may describe characteristics of the parameter set track. Encapsulation unit 30 may signal, within a TRAK box describing a parameter set track, the presence of a sequence-level SEI message in that parameter set track.

[0051] MVEX box 160 may describe the characteristics of the corresponding movie fragment 164, if present, to signal, for example, that video file 150 includes movie fragment 164 in addition to the video data contained in MOOV box 154. In the context of streaming video data, coded video pictures may be contained in movie fragment 164 rather than in MOOV box 154. Thus, all coded video samples may be contained in movie fragment 164 rather than in MOOV box 154.

[0052] The MOOV box 154 may contain a number of MVEX boxes 160 equal to the number of movie fragments 164 in the video file 150. Each of the MVEX boxes 160 may describe the characteristics of a corresponding one of the movie fragments 164. For example, each MVEX box may contain a movie extends header box (MEHD) that describes the duration for the corresponding one of the movie fragments 164.

[0053] As described above, encapsulation unit 30 may store a sequence data set within a video sample that does not contain the actual coded video data. A video sample may generally correspond to an access unit, which is a representation of a coded picture at a particular time instance. In the context of AVC, a coded picture includes one or more VCL NAL units that contain information for constructing all pixels of the access unit and other associated non-VCL NAL units, such as SEI messages. Accordingly, encapsulation unit 30 may include a sequence data set, which may include a sequence-level SEI message, within one of the movie fragments 164. Encapsulation unit 30 may further signal the presence of the sequence data set and / or the sequence-level SEI message within one of the movie fragments 164 in one of the MVEX boxes 160 corresponding to that movie fragment 164.

[0054] The SIDX box 162 is an optional element of video file 150; that is, a video file conforming to the 3GPP file format ("3GPP" is a registered trademark) or other such file formats does not necessarily include a SIDX box 162. According to an embodiment of the 3GPP file format, the SIDX box can be used to identify a subsegment of a segment (e.g., a segment contained within video file 150). The 3GPP file format defines a subsegment as "a self-contained set of one or more consecutive Movie Fragment boxes with corresponding Media Data boxes, where a Media Data box containing data referenced by a Movie Fragment box must follow that Movie Fragment box and precede the next Movie Fragment box containing information about the same track." The 3GPP file format also indicates that a SIDX box "contains a sequence of references to subsegments of the (sub)segment documented by that box. The referenced subsegments are contiguous in presentation time. Similarly, bytes referenced by a Segment Index box are always contiguous within a segment. The referenced size gives a count of the number of bytes in the referenced material."

[0055] SIDX box 162 generally provides information describing one or more subsegments of a segment contained within video file 150. For example, such information may include the playback time at which the subsegment begins and / or ends, a byte offset for the subsegment, whether the subsegment contains (e.g., starts with) a stream access point (SAP), the type for the SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), the location of the SAP (in terms of playback time and / or byte offset) within the subsegment, etc.

[0056] A movie fragment 164 may include one or more coded video pictures. In some embodiments, a movie fragment 164 may include one or more groups of pictures (GOPs), each of which may include several coded video pictures, e.g., frames or pictures. Furthermore, as noted above, a movie fragment 164 may include a sequence data set in some embodiments. Each movie fragment 164 may include a movie fragment header box (MFHD, not shown in FIG. 2). The MFHD box may describe characteristics of the corresponding movie fragment, such as a sequence number for that movie fragment. Movie fragments 164 may be included in video file 150 in sequence number order.

[0057] MFRA box 166 can describe random access points within movie fragments 164 of video file 150. This can assist in performing trick modes, such as performing a seek to a particular temporal position (i.e., playback time) within a segment encapsulated by video file 150. MFRA box 166 is generally optional in some embodiments and need not be included within a video file. Similarly, a client device, such as client device 40, does not necessarily need to reference MFRA box 166 to correctly decode and display video data in video file 150. MFRA box 166 may include a number of track fragment random access (TFRA) boxes (not shown) equal to the number of tracks in video file 150, or, in some embodiments, equal to the number of media tracks (e.g., non-hint tracks) in video file 150.

[0058] In some embodiments, movie fragment 164 may include one or more stream access points (SAPs), such as an IDR picture. Similarly, MFRA box 166 may provide an indication of the location of those SAPs within video file 150. Thus, a temporal sub-sequence of video file 150 may be formed from the SAPs of video file 150. This temporal sub-sequence may also include other pictures, such as P-frames and / or B-frames, that depend on the SAPs. Frames and / or slices of a temporal sub-sequence may be arranged within segments such that frames / slices of the temporal sub-sequence that depend on other frames / slices of the sub-sequence can be properly decoded. For example, in a hierarchical arrangement of data, data used for prediction with respect to other data may also be included within the temporal sub-sequence.

[0059] 3 is a block diagram illustrating an example network 170 including various devices for performing the techniques of this disclosure. In this example, the network 170 includes user equipment (UE) devices 172, 174, a call session control function (CSCF) 176, a multimedia telephony application server (MMTel AS) 178, a data channel control function (DCCF) 180, a multimedia resource function (MRF) 186, and an augmented reality application server (AR AS) 182.

[0060] UEs 172, 174 represent examples of UEs that may participate in an AR communication session 188. That is, UEs 172, 174 may exchange AR media data related to a virtual scene represented by a scene description. A user of UE 172, 174 may view the virtual scene including virtual objects and user AR data such as an avatar, a shadow cast by the avatar, user virtual objects, user-provided documents such as slides, images, videos, or other such data. Finally, a user of UE 172, 174 may experience an AR call from the perspective (first or third person) of their own avatar and the corresponding virtual objects and avatars in the scene.

[0061] The UEs 172, 174 may each collect pose data about the user of the UE 172, 174. For example, the UEs 172, 174 may collect pose data including the user's position corresponding to a position in a virtual scene and a viewport orientation, such as the direction the user is looking (i.e., the orientation of the UE 172, 174 in the real world, corresponding to the orientation of the virtual camera). The UEs 172, 174 may provide this pose data to the AR AS 182 and / or to each other.

[0062] Each of the UEs 172, 174 may perform various functions generally attributed to the content preparation device 20, the server device 60, and the client device 40 of FIG. 1 . However, in accordance with the techniques of this disclosure, one of the UEs 172, 174, e.g., the UE 172, is not capable of performing 3D virtual object rendering and would not include a rendering unit for rendering 3D virtual objects into 2D image or video data. Thus, the UE 172 would not perform the functionality of the rendering unit 66 of FIG. 1 . Instead, the rendering unit 184 of the AR AS 182 may perform the functionality of the rendering unit 66 of FIG. 1 on behalf of the UE 172, as described in more detail below.

[0063] CSCF 176 may be a Proxy CSCF (P-CSCF), an Interrogating CSCF (I-CSCF), or a Serving CSCF (S-CSCF). CSCF 176 may generally authenticate users of UEs 172 and / or 174, inspect signaling for appropriate use, provide Quality of Service (QoS), provide policy enforcement, participate in Session Initiation Protocol (SIP) communications, provide session control, direct messages to appropriate application servers, provide routing services, etc. CSCF 176 may represent one or more I / S / P CSCFs.

[0064] The MMTel AS 178 represents an application server for providing voice, video, and other telephony services over a network, such as a 5G network. The MMTel AS 178 may provide telephony applications and multimedia capabilities to the UEs 172, 174.

[0065] The DCCF 180 may act as an interface between the MMTel AS 178 and the MRF 186 to request data channel resources from the MRF 186 and to confirm that the data channel resources have been allocated. The MRF 186 may, in some examples, be an enhanced MRF (eMRF). Generally, the MRF 186 generates a scene description for each participant in an AR communication session.

[0066] The AR AS 182 may participate in an AR communication session 188 according to the techniques of this disclosure. In particular, the AR AS 182 includes a rendering unit 184. For illustrative purposes, even if the UE 172 is capable of displaying / presenting two-dimensional (2D) image or video data, it may be assumed that the UE 172 is not capable of rendering virtual object data to form such data. Thus, according to the techniques of this disclosure, the rendering unit 184 of the AR AS 182 may render virtual object data such as scene data, avatar data, pose information for the avatar, and for the viewport of the UE 172 (i.e., the direction in which the user of the UE 172 is facing and / or rotated). In this manner, the UE 172 and the AR AS 182 may perform split rendering. The AR AS 182 may be an Edge AS that meets the requirements of an AR call.

[0067] According to the techniques of this disclosure, AR communication session data may be transcoded into 2D overlay video data. A data channel for the AR communication session 188 may deliver a scene description for the AR communication session 188. The scene description may be used to configure a scene that will serve as a shared space for all participants (e.g., users of UEs 172, 174) in the AR communication session 188 (sometimes referred to as an “AR call”). Each participant may declare support for the AR call as well as the rendering capabilities of their respective UEs 172, 174 in an invitation to the AR call. The UEs 172, 174 may receive respective scene descriptions tailored to their rendering capabilities. The scene description may offer alternative representations from which the UEs 172, 174 may select.

[0068] 4 is a flow diagram illustrating an example procedure for establishing an AR call in accordance with the techniques of this disclosure. The various steps of FIG. 4 are described with respect to the components of FIG.

[0069] Initially, UE 172 ("UE1" in FIG. 4) may request to initiate an AR call or join an ongoing AR call / conference with UE 174 ("UE2" in FIG. 4) (200). Accordingly, UE 172 may generate a Session Description Protocol (SDP) offer indicating that UE 172 can receive only 2D content. The offer may indicate that UE 172 can send pause information and other 2D / 3D media. UE 174 sends the invite to the I / S / P-CSCF, i.e., CSCF 176 in FIG. 3.

[0070] The CSCF 176 identifies the AR call being offered (202). The CSCF 176 then forwards the invitation to the MMTel AS 178 (204).

[0071] The MMTel AS 178 identifies the capabilities of the UE 172 and decides to invoke split rendering functionality for the call. The MMTel AS 178 receives the 3D content and potentially rewrites the SDP offer to offer 3D content as well (206). The MMTel AS 178 then forwards the invitation to the AR AS 182 (208). The invitation can be a regular Session Initiation Protocol (SIP) invitation, or the MMTel AS 178 can use a service-based architecture (SBA) interface (e.g., a RESTful interface). The MMTel AS 178 also sends a request to the DCCF 180 to service data channel resources to the AR call application (210).

[0072] The DCCF 180, in response, sends a request to the MRF 186 to allocate the necessary data channel resources for the AR call (212). The DCCF 180 may indicate the type of application to enable the MRF 186 to generate an appropriate scene description for the AR call. In response, the MRF 186 indicates to the DCCF 180 when the resources have been allocated. The DCCF 180 confirms the allocation of the data channel resources to the MMTel AS 178 (214).

[0073] Additionally, the MMTel AS 178 forwards 216 the invitation generated in step 4 to the UE 174 .

[0074] In response, the UE 174 notifies the MMTel AS 178 that the UE 174 has accepted the call (218).

[0075] After receiving the reply from the UE 174, the MMTel AS 178 passes the reply to the AR AS 182 (220).

[0076] The AR AS 182 sets up resources for the AR call and for converting 3D scene objects to 2D image and / or video data (222). The AR AS 182 may then accept the invitation from the UE 172 (224), thereby establishing an AR communication session between the AR AS 182 and the UE 172, including pre-rendered AR call content on the downlink channel.

[0077] Both the AR AS 182 and the UE 174 establish connections to the MRF 186 to receive the scene description of the call (226).

[0078] Finally, connections are established between the UE 172 and the AR AS 182, and between the AR AS 182 and the UE 174 (228).

[0079] The UE 172 may send 230 the media stream data and pause information to the AR AS 182.

[0080] The AR AS 182 may forward the media data and pose information received from the UE 172 to the UE 174, or the AR AS 182 may generate 3D media and send the generated 3D media to the UE 174 (232).

[0081] The AR AS 182 may receive the media generated by the UE 174 (234). The AR AS 182 may use the rendering unit 184 to render the scene as described by the scene description along with the media received from the UE 172 and the UE 174 (236). Finally, the AR AS 182 may stream the rendered media to the UE 172 for display (238). Thus, the UE 172 may display the rendered media.

[0082] This procedure allows a network element to transparently invoke split rendering for an AR call (e.g., by the AR AS 182) without explicit intervention of the UE 172. The MMTel AS 178 may be responsible for selecting an appropriate split rendering server, which is executed by the AR AS 182.

[0083] In some embodiments, the UE 172 may instead receive an invitation to join an AR call. In some embodiments, the UE 172 may instead participate in an AR conference with multiple participants instead of only two UEs.

[0084] Appropriate signaling may be used to identify the AR call and convey, for example, the rendering capabilities of the UE 172 to the AR AS 182. The UE 172 may indicate whether it supports AR calls during the registration process or call setup with the CSCF 176. The UE 172 may also indicate its own rendering capabilities. These capabilities may include display configurations (e.g., access to an HMD), OpenXR support (supported view configurations and projection layers), and / or GPU capabilities such as support for different rendering pipelines and supported scene complexities.

[0085] New Session Description Protocol (SDP) level attributes can be used to indicate these capabilities for the session setup phase. An exemplary Augmented Backus-Naur Form (ABNF) syntax for such an SDP level attribute is shown below:

[0086] [Table 1]

[0087] The MMTel AS 178 may detect this attribute and, accordingly, decide to invoke split rendering capabilities for the AR call. The absence of the attribute may indicate that the device does not have AR capabilities. If the UE 174 is offering 3D content, the MMTel AS 178 may also invoke split rendering capabilities to convert from 3D to 2D for the UE 172.

[0088] 5 is a conceptual diagram illustrating an example IP Multimedia Subsystem (IMS) architecture 250. In this example, the IMS architecture 250 includes an augmented reality (AR) application server (AS) 252, a network exposure function (NEF) 254, a data channel (DC) signaling function (DCSF) 256, an IP Multimedia Subsystem (IMS) home subscriber system (HSS) 258, an IMS AS 260, a media function / multimedia resource function (MF / MRF) 262, an I / S call session control function (CSCF) 264, an interconnection border control function (IBCF) 266, a P-CSCF 268, an IMS access gateway (AGW) 270, a transition gateway (TrGW) 272, a user equipment 274, and a DC application repository (DCAR) 276. The IMS architecture may support data channel services in the IMS. AR calls may be enabled through the exchange of scene descriptions and scene description updates over the data channel. AR calls may be supported by a multimedia resource function (MRF) or a media function. The MF / MRF 262 may use a service-based interface to interact with a data channel application server (AS).

[0089] As described in more detail below, the UE 274 may establish an AR media session / call with a second UE (not shown), which may be coupled to a remote IMS as shown in FIG. 5. The UE 274 may be unable to render AR data or may choose not to render the AR data and instead request that the AR data be network-rendered, for example, by the MF / MRF 262. Accordingly, the UE 274 may send a request to the MF / MRF 262 via the I / S-CSCF 264 for network rendering of the AR data. In response, the MF / MRF 262 may receive AR media data from the second UE, then transcode the AR media data into 2D video data, and transmit the transcoded 2D video data (transcoded media data) to the UE 274. The UE 274 may then receive the transcoded media data from the MF / MRF 262 and present the transcoded media data.

[0090] Thus, UE 274 represents one embodiment of a first user equipment (UE) for communicating media data, including a memory configured to store media data and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: send a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establish a media communication session with a transcoding device performing a media function or multimedia resource function; the transcoding device is between the first UE and the second UE; receive transcoded media data from the transcoding device that the transcoding device transcodes from the AR media data received from the second UE; and present the transcoded media data.

[0091] Similarly, MF / MRF 262 represents one embodiment of a transcoding device that performs media functions or multimedia resource functions, including a memory configured to store media data and a processing system including one or more processors implemented in circuitry, where the processing system is configured to: receive a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; and, in response to receiving the AR media data from the second UE, render the AR media data to form rendered 2D media data; and transmit the rendered 2D media data to the first UE.

[0092] 6 and 7 are flow diagrams illustrating an example AR call setup process that may be used by the architecture of FIG. 5. Initially, a UE (such as UE 274 in FIG. 5) sends an invitation for an AR session to an IP Multimedia Subsystem (IMS) Application Server (AS) (280), such as IMS AS 260 in FIG. 5. The invitation may include an audio / video offer, a Session Description Protocol (SDP) offer for a bootstrap DC, etc. The IMS AS may perform DC routing determination and DCSF discovery (282). The IMS AS may then send a session event control notification message (e.g., Nimsas_SessionEventControl_Notify) to a DCSF, such as DSCF 256 in FIG. 5 (284).

[0093] The DCSF may then determine whether DC is provisioned and, if so, determine the DC control policy (286). The DCSF may also create originating and terminating DC media information (288). The DCSF may then return a media control and media command request message (e.g., Nimsas_MediaControl_Media_information) to the IMS AS (290).

[0094] The IMS AS may then perform DCMF / enMRF discovery (292). The IMS AS may request that the MF / MRF (such as MF / MRF 262 in FIG. 5) reserve originating and terminating media resources, causing the MF / MRF to allocate resources for originating and terminating MDC1 (294). The IMS AS may also send a media control media instruction response (e.g., Nimsas_MediaControl_MediaInstruction_response) to the DCSF (296). The DCSF may respond with a session event control notification response (e.g., Nimsas_SessionEventControl_Notify response) (298).

[0095] The IMS AS may then send the Invite message to an I / S-CSCF (300), such as I / S-CSCF 264 in Figure 5. The I / S-CSCF may then send the Invite message to a second UE (not shown in Figure 5), which may be communicatively coupled to the remote IMS in Figure 5 (302). The remote IMS (also referred to as the "terminating network") and the second UE may perform terminating network negotiation (304) to establish an AR session between the first UE (UE #1 in Figure 6) and the second UE. This may result in an exchange of 18X / PRACK / 200 OK (PRACK) / update messages (306).

[0096] 7, the second UE may send a 200 OK message to the I / S-CSCF (320). The I / S-CSCF may send the 200 OK message to the IMS AS (322). The IMS AS may send a session event control notify message (e.g., Nimsas_SessionEventControl_Notify) to the DCSF (324). The DCSF may respond with a session event control notify response message (e.g., Nimsas_SessionEventControl_Notify response) to the IMS AS (326). The IMS AS may then send a 200 OK message to the first UE (328).

[0097] The first UE may perform DC1 establishment with the MF / MRF (e.g., may bootstrap DC using stream ID / dcmap:0,10) (330). The first UE may also download a DC application list (332). The first UE may send a dedicated DC application request to download from the originating DCSF (334).

[0098] The MF / MRF may perform DC2 establishment with the second UE (e.g., may bootstrap DC using stream ID / dcmap:100, 110) (336). The second UE may download a DC application list from the DCSF (338). The second UE may also send a dedicated DC application request to download from the originating DCSF (340).

[0099] The first UE and the second UE may then perform DC3 establishment and bootstrap the DC using, for example, stream ID / dcmap:100 / 110 (342). The first UE and the second UE may download a DC application list (344). The first UE and the second UE may also perform a DC application request and download from the incoming DCSF (346). The second UE may then perform DC4 establishment (e.g., bootstrap the DC using stream ID / dcmap:0,10), download a DC application list, and perform a dedicated DC application download (348). The first UE and the second UE may then engage in an AR session as another subsequent procedure (350).

[0100] 8 is a flow diagram illustrating an example call setup procedure that may enable remote rendering of AR content based on a request from a sender of the AR content. Initially, two UEs (UE A and UE B in this example) perform an audio / video establishment procedure to bootstrap data channel establishment for the first and second UEs (360). The first UE may then determine to request network media rendering based on its status (362). According to techniques of this disclosure, the first UE may negotiate a request to perform AR media rendering with the IMS network (364). The first UE and the IMS network may establish an application data channel (366). The first UE and the IMS network may further perform media renegotiation to anchor the first UE's audio and / or video to the MF / MRF (368).

[0101] The IMS network and the second UE may also perform media renegotiation to anchor the second UE's audio and / or video data to the MF / MRF (370). The first UE may initiate AR media rendering (372). The first UE may send AR data to the MF / MRF for network-assisted rendering (374). The MF / MRF may perform AR media rendering (376). The MF / MRF may then send the rendered AR media data (audio and / or video) to the second UE, for example, via RTP (378).

[0102] 9 is a call flow diagram illustrating an example process for setting up an AR call in accordance with the techniques of this disclosure. In the example of FIG. 9, UE1 first decides to initiate an AR call with UE2 or join an ongoing AR conference (400). UE1 generates an SDP offer indicating that UE1 can receive only 2D content. The offer may indicate that UE1 can send pose information and other 2D / 3D media. UE1 sends an invite to the I / S / P-CSCF, which forwards the invite to the IMS AS.

[0103] The IMS AS may then identify the capabilities of UE1 and decide to invoke a network rendering function for the AR call (402). The IMS AS may further rewrite the SDP offer to redirect the 3D media and scene content to the media function (MF).

[0104] The IMS AS may then negotiate data channel resources for the session with the MF / MRF (404).

[0105] If network rendering is not required, the IMS AS may forward the invite including the data channel information to the I / S / P-CSCF (406).

[0106] The IMS AS may allocate resources for transcoding the 3D content via the data channel DCSF (408).

[0107] The IMS AS may then forward the updated invitation to the I / S / P-CSCF (410). This invitation may indicate that the 3D content should be routed to the MF.

[0108] The I / S / P-CSCF may then forward 412 the invitation generated in step 410 to UE2.

[0109] UE2 may then send data to the IMS AS that UE2 has accepted the call invitation (414).

[0110] The second UE may connect to the AR AS (416). The AR AS may send a scene update to UE2 and the MF (418).

[0111] The second UE may then transmit 420 the AR media data to the MF / MRF.

[0112] The MF / MRF may transcode the AR media data from the second UE of the 3D scene by performing network rendering (422).

[0113] The MF / MRF may stream the rendered media resulting from the transcoding and network rendering to UE1 (424).

[0114] This procedure allows the network to invoke remote rendering for an AR call without explicit intervention of the UE. The IMS AS is in this embodiment responsible for selecting an appropriate network MF that will perform the network rendering operation for the session.

[0115] The AR AS is the central entity for the AR call in this embodiment. In this embodiment, the AR AS manages the scene description for the session and performs scene composition. Endpoints can send scene updates and pause information to the AR AS over the data channel.

[0116] UE1 may determine that UE1 is participating in an AR call, or may not receive information indicating that UE1 is participating in an AR call. If UE1 determines that UE1 is participating in an AR call, UE1 may share information about its display capabilities and its current viewer pose for XR remote rendering. To share this information with the IMS AS, UE1 may use SDP attributes in accordance with the techniques of this disclosure. The IMS AS may detect signaling from UE1 and use the information to configure an MF remote rendering session. The SDP attributes may include any or all of the following: Display configuration, e.g., access to HMD OpenXR support, i.e. supported view configurations and projection layers GPU capabilities, such as support for different rendering pipelines, as well as supported scene complexity

[0117] SDP session-level attributes may indicate these capabilities for the session setup phase, in accordance with the techniques of this disclosure. The ABNF syntax of an SDP session-level attribute may be as follows:

[0118] [Table 2]

[0119] The IMS AS may detect this attribute and decide to invoke a remote rendering capability for the AR call. The absence of the attribute may indicate that the device does not have AR capabilities. If other participants are offering 3D content, the IMS AS may invoke a network rendering capability to convert from 3D to 2D.

[0120] FIG. 10 is a flowchart illustrating an example method that may be executed by a client device, such as a user equipment (UE) device, to request transcoding of AR media data in accordance with the techniques of this disclosure. Initially, the UE may determine that the UE is not capable of rendering AR data (450). Assuming that a user of the UE requests to participate in an AR media session with a second, different UE, the UE may form a request for network transcoding of the AR media data (452). The request may be a request that the AR media data be transcoded into two-dimensional video data. The UE may send the request to the CSCF (454). The CSCF may forward the request to the transcoding device. Thus, the UE can send the request to the CSCF to cause the CSCF to send the request to the transcoding device.

[0121] Finally, the UE may establish a session with the transcoding device (456). The UE may receive transcoded media data from the transcoding device (458). That is, the second UE may send AR media data to the transcoding device, which may render and transcode the AR media data into 2D video data and then send the 2D video data (i.e., the transcoded media data) to the UE. The UE may then present the transcoded media data (460).

[0122] In some examples, despite not being able to render AR media data (or not actually participating in the AR media data rendering process), the UE may still be able to generate AR-related data, such as pose data. The pose data may represent, for example, the relative position of a user of the UE in a 3D virtual scene and the direction / orientation in which the user of the UE is looking. Thus, the UE may generate AR-related data, such as pose data, based on, for example, sensor data collecting the position and orientation of the UE (462). The UE may then transmit the AR-related data to a transcoding device (or to a second UE via the transcoding device).

[0123] Thus, the method of FIG. 10 represents one embodiment of a method for communicating media data, the method including: sending, by a first user equipment (UE), a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding AR media data into two-dimensional video data; establishing, by the first UE, a media communication session with a transcoding device performing a media function or a multimedia resource function, the transcoding device being between the first UE and the second UE; receiving, by the first UE, transcoded media data from the transcoding device, the transcoding device transcoding the AR media data received from the second UE; and presenting, by the first UE, the transcoded media data.

[0124] 11 is a flowchart illustrating an example method for transcoding AR media data in accordance with techniques of this disclosure. The method of FIG. 11 may be performed by a transcoding device, such as a device that executes a media function (MF) and / or a multimedia rendering function (MRF). The transcoding device may correspond, for example, to the AR AS 182 or MRF 186 of FIG. 3 or the MF / MRF 262 of FIG. 5.

[0125] Initially, the transcoding device may receive a request from a first UE to transcode AR data for the first UE (480). The transcoding device may initiate transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) application server (AS) to which the first UE is communicatively coupled. Thus, when the transcoding device receives AR data from a second UE (482) destined for the first UE, the transcoding device may render the AR data to form 2D media data (484). The AR data may be encoded using a first encoding scheme, such that the transcoding device may decode the AR data prior to rendering, and then, after rendering the AR data to form the 2D video data, the transcoding device may encode the 2D video data using, for example, ITU-T H.264 / Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), or other such video coding standard (e.g., AV1). In this manner, the media data may be transcoded. The transcoding device may then transmit the rendered 2D media data to the first UE (486).

[0126] In some examples, the transcoding device may receive AR-related data, such as pose data, from the first UE (488). In response, the transcoding device may transmit the AR-related data to the second UE (490).

[0127] Thus, the method of FIG. 11 represents one embodiment of a method for communicating media data, the method including receiving, by a transcoding device performing a media function or a multimedia resource function, a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; in response to receiving, by the transcoding device, the AR media data from the second UE, rendering, by the transcoding device, the AR media data to form rendered 2D media data; and transmitting, by the transcoding device, the rendered 2D media data to the first UE.

[0128] The following clauses represent various examples of the techniques of this disclosure.

[0129] Clause 1: A method for communicating media data, comprising: sending, by a first user equipment (UE), a request to initiate an augmented reality (AR) media communication session with a second UE, the request including data indicating that the first UE is capable of providing AR-related data; establishing, by the first UE, the media communication session with an AR application server (AS), the AR AS being between the first UE and the second UE; and transmitting, by the first UE, the AR-related data to the second UE via the AR AS.

[0130] Clause 2: The method of clause 1, further including receiving, by the first UE, from the AR AS, AR media data, the AR media data including AR data rendered by the AR AS corresponding to the second UE; and presenting, by the first UE, the AR media data.

[0131] Clause 3: A method for communicating media data, comprising: receiving, by a first user equipment (UE), an invitation to initiate an augmented reality (AR) media communication session from a second UE; establishing, by the first UE, a connection to an enhanced multimedia resource function (eMRF); receiving, by the first UE, a scene description from the eMRF via the connection to the eMRF; establishing, by the first UE, a connection to an AR application server (AR AS); and transmitting, by the first UE, AR-related data to the second UE via the AR AS.

[0132] Clause 4: The method of clause 1, further including receiving, by the first UE, from the AR AS, AR media data, the AR media data including AR data rendered by the AR AS corresponding to the second UE; and presenting, by the first UE, the AR media data.

[0133] Clause 5: A method for communicating media data, the method including: receiving, by one or more processors, a request from a second user equipment (UE) to initiate an augmented reality (AR) media communication session with the first UE, the request including data indicating that the first UE is capable of providing AR-related data; determining, by the one or more processors, that the first UE is capable of receiving two-dimensional (2D) video data and not three-dimensional (3D) video data; sending, by the one or more processors, an invitation to the AR media communication session to the second UE; establishing, by the one or more processors, a first media communication session with the first UE and a second media communication session with the second UE; transmitting, by the one or more processors, the AR-related data to the second UE in response to receiving the AR-related data from the first UE; rendering, by the one or more processors, the AR media data to form rendered 2D media data in response to receiving the AR media data from the second UE; and transmitting the rendered 2D media data to the first UE.

[0134] Clause 6: A method for communicating media data, the method including: receiving, by one or more processors, a request from a second user equipment (UE) to initiate an augmented reality (AR) media communication session with the first UE, the request including data indicating that the first UE is capable of providing AR-related data; determining, by the one or more processors, that the first UE is capable of receiving two-dimensional (2D) video data and not three-dimensional (3D) video data; sending, by the one or more processors, an invitation to the AR media communication session to the second UE; establishing, by the one or more processors, a first media communication session with the first UE and a second media communication session with the second UE; in response to receiving the AR media data from the second UE, rendering, by the one or more processors, the AR media data to form rendered 2D media data; and transmitting, by the one or more processors, the rendered 2D media data to the first UE.

[0135] Clause 7: A system for communicating media data, the system comprising one or more means for performing the method according to any of clauses 1 to 6.

[0136] Clause 8: The system of clause 7, wherein the one or more means include one or more processors implemented in the circuit and a memory configured to store the AR media data.

[0137] Clause 9: A first user equipment (UE) device for communicating media data, comprising: means for sending a request to initiate an augmented reality (AR) media communication session with a second UE, the request including data indicating that the first UE is capable of providing AR-related data; an AR application server (AS), the AS being between the first UE and the second UE; means for establishing a media communication session with the AR AS; and means for transmitting the AR-related data to the second UE via the AR AS.

[0138] Clause 10: A first user equipment (UE) device for communicating media data, comprising: means for receiving an invitation to initiate an augmented reality (AR) media communication session from a second UE; means for establishing a connection to an enhanced multimedia resource function (eMRF); means for receiving a scene description from the eMRF via the connection to the eMRF; means for establishing a connection to an AR application server (AR AS); and means for transmitting AR-related data to the second UE via the AR AS.

[0139] Clause 11: A method for communicating media data, comprising: sending, by a first user equipment (UE), a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establishing, by the first UE, a media communication session with a transcoding device that performs a media function or a multimedia resource function and is between the first UE and the second UE; receiving, by the first UE, transcoded media data from the transcoding device, the transcoding device transcoding the AR media data received from the second UE; and presenting, by the first UE, the transcoded media data.

[0140] Clause 12: The method of clause 11, wherein sending the request further includes indicating in the request that the first UE can provide AR-related data, and the method further includes transmitting, by the first UE, the AR-related data to the second UE.

[0141] Clause 13: The method of clause 12, wherein the AR-related data includes pose information for the first UE.

[0142] Clause 14: The method of clause 11, wherein sending a request to initiate an AR media call includes sending a request to initiate an AR media call to an IP Multimedia Subsystem (IMS) Application Server (AS) via a CSCF.

[0143] Clause 15: The method of clause 14, further comprising: receiving, by the CSCF, an updated invitation to the AR media call from the IMS AS and, in response to sending the updated invitation to the second UE, receiving an accept message for the AR media call from the second UE.

[0144] Clause 16: A first user equipment (UE) for communicating media data, comprising: a memory configured to store the media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: send a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establish a media communication session with a transcoding device performing a media function or multimedia resource function, the transcoding device being between the first UE and the second UE; receive from the transcoding device transcoded media data that the transcoding device transcodes from the AR media data received from the second UE; and present the transcoded media data.

[0145] Clause 17: The first UE described in Clause 16, wherein to send the request, the processing system is further configured to indicate in the request that the first UE can provide AR-related data, and the processing system is further configured to send the AR-related data to the second UE.

[0146] Clause 18: The first UE of clause 17, wherein the AR-related data includes pose information for the first UE.

[0147] Clause 19: The first UE described in Clause 16, wherein the processing system is configured to send a request to initiate an AR media call to an IP Multimedia Subsystem (IMS) Application Server (AS) via the CSCF to send a request to initiate an AR media call.

[0148] Clause 20: The first UE of clause 19, wherein the processing system is further configured to receive an accept message for the AR media call from the second UE in response to the CSCF receiving an updated invitation for the AR media call from the IMS AS and sending the updated invitation to the second UE.

[0149] Clause 21: The first UE of clause 16, further comprising a display configured to display the transcoded media data.

[0150] Clause 22: The first UE of Clause 16, wherein the first UE comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0151] Clause 23: A method for communicating media data, comprising: receiving, by a transcoding device performing a media function or a multimedia resource function, a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; in response to receiving, by the transcoding device, the AR media data from the second UE, rendering, by the transcoding device, the AR media data to form rendered 2D media data; and transmitting, by the transcoding device, the rendered 2D media data to the first UE.

[0152] Clause 24: The method of clause 23, further comprising determining that the first UE is capable of receiving two-dimensional (2D) video data and that the first UE has requested transcoding of the AR media data into 2D video data.

[0153] Clause 25: The method of clause 23, wherein the request includes data indicating that the first UE is capable of providing the AR-related data.

[0154] Clause 26: The method of clause 25, further comprising transmitting, by the transcoding device, the AR-related data to the second UE in response to receiving the AR-related data from the first UE.

[0155] Clause 27: The method of clause 23, further comprising: sending, by the transcoding device, an invitation to the AR media call to the second UE; and establishing, by the transcoding device, a first media communication session with the first UE and a second media communication session with the second UE.

[0156] Clause 28: The method of clause 23, further comprising initiating AR transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) Application Server (AS).

[0157] Clause 29: A transcoding device that performs a media function or a multimedia resource function, comprising: a memory configured to store media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: receive a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; and, in response to receiving the AR media data from the second UE, render the AR media data to form rendered 2D media data; and transmit the rendered 2D media data to the first UE.

[0158] Clause 30: The transcoding device described in Clause 29, wherein the processing system is further configured to determine that the first UE is capable of receiving two-dimensional (2D) video data and that the first UE has requested transcoding of the AR media data into the 2D video data.

[0159] Clause 31: The transcoding device of clause 29, wherein the request includes data indicating that the first UE is capable of providing AR-related data.

[0160] Clause 32: The transcoding device of clause 31, wherein the processing system is further configured to transmit the AR-related data to the second UE in response to receiving the AR-related data from the first UE.

[0161] Clause 33: The transcoding device of clause 29, wherein the processing system is further configured to: send an invitation to the AR media call to the second UE and establish a first media communication session with the first UE and a second media communication session with the second UE.

[0162] Clause 34: The transcoding device of clause 29, wherein the processing system is further configured to initiate AR transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) Application Server (AS).

[0163] Clause 35: A method for communicating media data, the method comprising: sending, by a first user equipment (UE), a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establishing, by the first UE, a media communication session with a transcoding device performing a media function or a multimedia resource function, the transcoding device being between the first UE and the second UE; receiving, by the first UE, transcoded media data from the transcoding device, the transcoding device transcoding the AR media data received from the second UE; and presenting, by the first UE, the transcoded media data.

[0164] Clause 36: The method of clause 35, wherein sending the request further includes indicating in the request that the first UE is capable of providing AR-related data, and the method further includes transmitting, by the first UE, the AR-related data to the second UE.

[0165] Clause 37: The method of clause 36, wherein the AR-related data includes pose information for the first UE.

[0166] Clause 38: The method of any of clauses 35 to 37, wherein sending a request to initiate an AR media call includes sending a request to initiate an AR media call to an IP Multimedia Subsystem (IMS) Application Server (AS) via a CSCF.

[0167] Clause 39. The method of clause 38, further comprising: receiving, by the CSCF, an updated invitation to the AR media call from the IMS AS and, in response to sending the updated invitation to the second UE, receiving an accept message for the AR media call from the second UE.

[0168] Clause 40. A first user equipment (UE) for communicating media data, comprising: a memory configured to store the media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: send a request to a call session control function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establish a media communication session with a transcoding device performing media functions or multimedia resource functions, the transcoding device being between the first UE and the second UE; receive from the transcoding device transcoded media data that the transcoding device transcodes from the AR media data received from the second UE; and present the transcoded media data.

[0169] Clause 41: The first UE described in Clause 40, wherein to send the request, the processing system is further configured to indicate in the request that the first UE is capable of providing AR-related data, and the processing system is further configured to transmit the AR-related data to the second UE.

[0170] Clause 42: The first UE of clause 41, wherein the AR-related data includes pose information for the first UE.

[0171] Clause 43: A first UE described in any of clauses 40 to 42, wherein the processing system is configured to send a request to initiate an AR media call to an IP Multimedia Subsystem (IMS) Application Server (AS) via the CSCF to send a request to initiate an AR media call.

[0172] Clause 44: The first UE of clause 43, wherein the processing system is further configured to receive an accept message for the AR media call from the second UE in response to the CSCF receiving an updated invitation for the AR media call from the IMS AS and sending the updated invitation to the second UE.

[0173] Clause 45: The first UE of any of clauses 40 to 44, further comprising a display configured to display the transcoded media data.

[0174] Clause 46: The first UE described in any of Clauses 40 to 45, wherein the first UE comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0175] Clause 47. A method of communicating media data, comprising: receiving, by a transcoding device performing a media function or a multimedia resource function, a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; in response to receiving, by the transcoding device, the AR media data from the second UE, rendering, by the transcoding device, the AR media data to form rendered 2D media data; and transmitting, by the transcoding device, the rendered 2D media data to the first UE.

[0176] Clause 48: The method of clause 47, further comprising determining that the first UE is capable of receiving two-dimensional (2D) video data and that the first UE has requested transcoding of the AR media data into 2D video data.

[0177] Clause 49: The method of clause 47 or 48, wherein the request includes data indicating that the first UE is capable of providing the AR-related data.

[0178] Clause 50: The method of clause 49, further comprising transmitting, by the transcoding device, the AR-related data to the second UE in response to receiving the AR-related data from the first UE.

[0179] Clause 51: The method of any of clauses 47 to 50, further comprising: sending, by the transcoding device, an invitation to the AR media call to the second UE; and establishing, by the transcoding device, a first media communication session with the first UE and a second media communication session with the second UE.

[0180] Clause 52: The method of any of clauses 47 to 51, further comprising initiating AR transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) Application Server (AS).

[0181] Clause 53: A transcoding device that performs a media function or a multimedia resource function, comprising: a memory configured to store media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: receive a request from a first user equipment (UE) to transcode AR media data for an AR call with a second UE; and, in response to receiving the AR media data from the second UE, render the AR media data to form rendered 2D media data; and transmit the rendered 2D media data to the first UE.

[0182] Clause 54: The transcoding device described in Clause 53, wherein the processing system is further configured to determine that the first UE is capable of receiving two-dimensional (2D) video data and that the first UE has requested transcoding of the AR media data into the 2D video data.

[0183] Clause 55: The transcoding device of clause 53 or 54, wherein the request includes data indicating that the first UE is capable of providing the AR-related data.

[0184] Clause 56: The transcoding device of Clause 55, wherein the processing system is further configured to transmit the AR-related data to the second UE in response to receiving the AR-related data from the first UE.

[0185] Clause 57: A transcoding device as described in any of clauses 53 to 56, wherein the processing system is further configured to: send an invitation to the AR media call to the second UE and establish a first media communication session with the first UE and a second media communication session with the second UE.

[0186] Clause 58: A transcoding device described in any of clauses 53 to 57, wherein the processing system is further configured to initiate AR transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) Application Server (AS).

[0187] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communications protocol. As such, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0188] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0189] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. It is also possible for these techniques to be implemented entirely in one or more circuits or logic elements.

[0190] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units have been described in this disclosure to highlight functional aspects of devices configured to implement the disclosed techniques, but they do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0191] Various embodiments have been described. These and other embodiments are within the scope of the following claims.

Claims

1. 1. A method for communicating media data, comprising: sending, by a first user equipment (UE), a request to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data, to a call session control function (CSCF); establishing, by the first UE, a media communication session with a transcoding device that performs a media function or a multimedia resource function, the transcoding device being between the first UE and the second UE; receiving, by the first UE, transcoded media data from the transcoding device, the transcoding device transcoding the AR media data received from the second UE; presenting, by the first UE, the transcoded media data; and A method comprising:

2. 2. The method of claim 1, wherein transmitting the request further comprises indicating in the request that the first UE can provide AR-related data, and the method further comprises transmitting, by the first UE, AR-related data to the second UE.

3. The method of claim 2 , wherein the AR-related data includes pause information for the first UE.

4. 2. The method of claim 1, wherein sending the request to initiate the AR media call comprises sending the request to initiate the AR media call to an IP Multimedia Subsystem (IMS) Application Server (AS) via the CSCF.

5. 5. The method of claim 4, further comprising receiving, by the CSCF, an updated invitation to the AR media call from the IMS AS and, in response to sending the updated invitation to the second UE, an accept message for the AR media call from the second UE.

6. A first user equipment (UE) for communicating media data, comprising: a memory configured to store media data; a processing system including one or more processors implemented in circuitry, said processing system comprising: sending a request to a Call Session Control Function (CSCF) to initiate an augmented reality (AR) media call with a second UE, the request including data indicating a request for transcoding the AR media data into two-dimensional video data; establishing a media communication session with a transcoding device performing a media function or a multimedia resource function, the transcoding device being between the first UE and the second UE; receiving transcoded media data from the transcoding device, the transcoding device having transcoded the AR media data received from the second UE; presenting the transcoded media data; a first user equipment (UE) configured to:

7. 7. The first UE of claim 6, wherein, to transmit the request, the processing system is further configured to indicate in the request that the first UE can provide AR-related data, and the processing system is further configured to transmit the AR-related data to the second UE.

8. The first UE of claim 7 , wherein the AR-related data includes pause information for the first UE.

9. 7. The first UE of claim 6, wherein to send the request to initiate the AR media call, the processing system is configured to send the request to initiate the AR media call to an IP Multimedia Subsystem (IMS) Application Server (AS) via the CSCF.

10. 10. The first UE of claim 9, wherein the processing system is further configured to receive, from the second UE, an accept message for the AR media call in response to the CSCF receiving an updated invitation for the AR media call from the IMS AS and sending the updated invitation to the second UE.

11. The first UE of claim 6 , further comprising a display configured to display the transcoded media data.

12. The first UE of claim 6 , wherein the first UE comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

13. 1. A method for communicating media data, comprising: receiving, by a transcoding device performing a media function or a multimedia resource function, a request from a first user equipment (UE) to transcode augmented reality (AR) media data for an AR call with a second UE; in response to receiving, by the transcoding device, AR media data from the second UE, rendering, by the transcoding device, the AR media data to form rendered 2D media data; transmitting, by the transcoding device, the rendered 2D media data to the first UE; A method comprising:

14. 14. The method of claim 13, further comprising: determining that the first UE is capable of receiving two-dimensional (2D) video data and that the first UE has requested transcoding of AR media data into the 2D video data.

15. The method of claim 13 , wherein the request includes data indicating that the first UE is capable of providing AR-related data.

16. The method of claim 15 , further comprising transmitting, by the transcoding device, the AR-related data to the second UE in response to receiving the AR-related data from the first UE.

17. sending, by the transcoding device, an invitation to the AR media call to the second UE; establishing, by the transcoding device, a first media communication session with the first UE and a second media communication session with the second UE; The method of claim 13 further comprising:

18. 14. The method of claim 13, further comprising initiating AR transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) Application Server (AS).

19. 1. A transcoding device that performs a media function or a multimedia resource function, comprising: a memory configured to store media data; a processing system including one or more processors implemented in circuitry, said processing system comprising: receiving a request from a first user equipment (UE) to transcode augmented reality (AR) media data for an AR call with a second UE; In response to receiving AR media data from the second UE, rendering the AR media data to form rendered 2D media data; transmitting the rendered 2D media data to the first UE; A transcoding device that is configured as follows:

20. 20. The transcoding device of claim 19, wherein the processing system is further configured to determine that the first UE is capable of receiving two-dimensional (2D) video data and that the first UE has requested transcoding of AR media data into the 2D video data.

21. The transcoding device of claim 19 , wherein the request includes data indicating that the first UE can provide AR-related data.

22. 22. The transcoding device of claim 21, wherein the processing system is further configured to transmit the AR-related data to the second UE in response to receiving the AR-related data from the first UE.

23. the processing system comprising: sending an invitation to the AR media call to the second UE; establishing a first media communication session with the first UE and a second media communication session with the second UE; 20. The transcoding device of claim 19, further configured to:

24. 20. The transcoding device of claim 19, wherein the processing system is further configured to initiate AR transcoding resources based on configuration data received from an IP Multimedia Subsystem (IMS) Application Server (AS).