Mixed Media Data Formats and Transport Protocols

By employing transport protocols and payload formats that support various media types, the synchronization of media streams in real-time communication sessions is improved, addressing the challenge of synchronized media presentation in mixed media environments.

JP2025532523APending Publication Date: 2025-10-01QUALCOMM INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2025514435
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-24
Filing Date
2023-08-28
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

Existing technologies face challenges in synchronizing the presentation of multiple media streams during real-time communication sessions, particularly in mixed media environments involving virtual avatars, leading to suboptimal user experiences.

Method used

The implementation of transport protocols and payload formats that support various media types, such as WebRTC over RTP, to synchronize media presentations by including header data indicating media types and attributes in packets transmitted during media communication sessions.

Benefits of technology

Enhances user experience by ensuring synchronized presentation of media streams, particularly in immersive environments like augmented and mixed reality, by accurately representing audio and visual data in real-time communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532523000001_ABST
    Figure 2025532523000001_ABST
Patent Text Reader

Abstract

An exemplary device for processing media data includes a memory configured to store media data and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: determine one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets including data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Patent Application No. 18 / 454,992, filed August 24, 2023, U.S. Provisional Patent Application No. 63 / 484,571, filed February 13, 2023, and U.S. Provisional Patent Application No. 63 / 376,409, filed September 20, 2022, the entire contents of each of which are incorporated herein by reference.

[0002] FIELD This disclosure relates to the storage and transport of encoded video data. [Background technology]

[0003] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, video conferencing devices, etc. Digital video devices implement video compression techniques, such as those described in standards established by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also known as High Efficiency Video Coding, HEVC), and extensions to such standards, to more efficiently transmit and receive digital video information.

[0004] Video compression techniques perform spatial and / or temporal prediction to reduce or remove redundancy inherent in video sequences. In block-based video coding, a video frame or slice may be partitioned into macroblocks. Each macroblock may be further partitioned. Macroblocks in an intra-coded (I) frame or slice are coded using spatial prediction with respect to neighboring macroblocks. Macroblocks in an inter-coded (P or B) frame or slice may use spatial prediction with respect to neighboring macroblocks in the same frame or slice or temporal prediction with respect to other reference frames.

[0005] After the video data is encoded, it may be packetized for transmission or storage and assembled into a video file that conforms to any of a variety of standards, such as the International Organization for Standardization (ISO) base media file format, such as AVC, and its extensions. Summary of the Invention

[0006] Generally, this disclosure describes technologies related to real-time communication protocols, such as WebRTC (Web Real-Time Communication). WebRTC can be run over the Real-time Transport Protocol (RTP). A WebRTC session can be used to exchange multiple streams of media data, such as audio, video, and / or extended reality (XR), such as augmented reality (AR), mixed reality (MR), or virtual reality (VR). For example, a user may desire to have a communication session involving mixed media types, where virtual avatars are represented in a virtual environment and each user can be represented by one of the virtual avatars. In such a context, synchronizing the presentation of media of two or more media streams during playback can improve the user experience. For example, it may be desirable to animate an avatar's face to accurately represent audio data from a corresponding user, and therefore it may be important to synchronize the animation data with the sound data. This disclosure describes various technologies that can be used, for example, with RTP, to support various media types. In particular, this disclosure describes transport protocols and payload formats that can support various media types, e.g., for real-time communication, and in this manner, synchronization between various media presentations transported using, e.g., RTP, can be achieved.

[0007] In one embodiment, a method for processing media data includes determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmitting the packets to devices involved in the media communication session.

[0008] In another example, a device for processing media data includes: a memory configured to store media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: determine one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets including data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.

[0009] In another embodiment, a device for processing media data includes: means for determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; means for constructing packets including data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and means for transmitting the packets to devices involved in the media communication session.

[0010] In another embodiment, a computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to determine one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.

[0011] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram illustrating an example system that implements techniques for streaming media data over a network. [Figure 2] FIG. 2 is a block diagram illustrating elements of an exemplary video file. [Figure 3] FIG. 1 is a conceptual diagram illustrating an example Web Real-Time Communications (WebRTC) protocol stack. [Figure 4] FIG. 1 is a conceptual diagram illustrating a Real-time Transport Protocol (RTP) header format. [Figure 5] FIG. 1 is a conceptual diagram illustrating an exemplary RTP header extension. [Figure 6] FIG. 10 is a conceptual diagram illustrating one embodiment of a two-byte header form. [Figure 7] FIG. 1 is a conceptual diagram illustrating an embodiment of a 2-byte RTP header extension. [Figure 8] FIG. 1 is a conceptual diagram illustrating an example common packet format for feedback messages. [Figure 9] FIG. 1 is a conceptual diagram illustrating an example Stream Control Transmission Protocol (SCTP) chunk field format. [Figure 10] FIG. 1 is a conceptual diagram illustrating an exemplary SCTP data chunk format. [Figure 11] FIG. 1 is a conceptual diagram illustrating an example media transmission architecture capable of implementing the techniques of this disclosure. [Figure 12] FIG. 1 is a conceptual diagram illustrating exemplary RTP extensions and payloads for media types in accordance with the techniques of this disclosure. [Figure 13] FIG. 1 is a conceptual diagram illustrating an example RTP extension element format for interactivity media type data. [Figure 14] FIG. 1 is a conceptual diagram illustrating an exemplary immersive media profile (IMP). [Figure 15] FIG. 1 is a conceptual diagram illustrating an example SCTP chunk format for an immersive media type. [Figure 16] FIG. 10 is a conceptual diagram illustrating another embodiment of an SCTP chunk format for immersive media type data. [Figure 17] FIG. 1 is a conceptual diagram illustrating an exemplary timed metadata payload format in accordance with the techniques of this disclosure. [Figure 18] FIG. 10 is a block diagram illustrating an example action set header format in accordance with the techniques of this disclosure. [Figure 19] FIG. 10 is a block diagram illustrating an exemplary action metadata payload format in accordance with the techniques of this disclosure. [Figure 20] FIG. 10 is a block diagram illustrating an example gaze interaction metadata action payload format in accordance with the techniques of this disclosure. [Figure 21] FIG. 1 is a block diagram illustrating an example spatial location metadata payload format in accordance with the techniques of this disclosure. [Figure 22]FIG. 2 is a block diagram illustrating an exemplary spatial rate metadata payload format in accordance with the techniques of this disclosure. [Figure 23] FIG. 10 is a block diagram illustrating an exemplary view metadata payload format in accordance with the techniques of this disclosure. [Figure 24] 1 is a flowchart illustrating an example method for transmitting user interaction data and associated metadata to another device involved in an XR multimedia communication session, in accordance with techniques of this disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Generally, this disclosure describes technologies related to the transport of immersive metadata. The Third Generation Partnership Project (3GPP®) Media Capability for Augmented Reality (MeCAR) permanent document (3GPP TSG SA WG4 S4-221150 “MeCAR Permanent Document v3.0”, August 2022) and the 5G_RTP permanent document (3GPP TSG SA WG4 S4-221209 “5G_RTP Permanent Document v0.0.2”, August 2022) define numerous augmented reality (AR) and mixed reality (MR) media types (metadata). Some interactivity media types, such as user pose, viewport, gesture, body movement, and facial expression, are exchanged between users in real time with low latency requirements. Some of the media types, such as viewport and body movement for gaming applications, can be synchronized with each other or with other media in real time. Mechanisms for transporting these interactive media types are important for immersive real-time communication (RTC), which is being investigated in 3GPP SA4.

[0014] More specifically, this disclosure describes technologies related to real-time communication protocols, such as WebRTC (Web Real-Time Communications). WebRTC can be run over the Real-Time Transport Protocol (RTP). A WebRTC session can be used to exchange multiple streams of media data, such as audio, video, and / or extended reality (XR), such as augmented reality (AR), mixed reality (MR), or virtual reality (VR). For example, a user may desire to have a communication session involving mixed media types, where virtual avatars are represented in a virtual environment and each user can be represented by one of the virtual avatars. In such a context, synchronizing the presentation of media of two or more media streams during playback can improve the user experience. For example, it may be desirable to animate an avatar's face to accurately represent audio data from a corresponding user, and therefore, synchronizing the animation data with the sound data may be important. This disclosure describes various techniques that can be used, for example, with RTP, to support various media types. Specifically, this disclosure describes transport protocols and payload formats that can support various media types, for example, for real-time communication. In this way, synchronization between different media presentations transported using, for example, RTP can be achieved.

[0015] The techniques of this disclosure may be applied to video files that conform to video data encapsulated according to any of the ISO Base Media File Format, Scalable Video Coding (SVC) file format, Advanced Video Coding (AVC) file format, 3rd Generation Partnership Project (3GPP) file format, and / or Multiview Video Coding (MVC) file format, or other similar video file formats.

[0016] 1 is a block diagram illustrating an example system 10 that implements techniques for streaming media data over a network. In this example, system 10 includes a content preparation device 20, a server device 60, and a client device 40. Client device 40 and server device 60 are communicatively coupled by a network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled by network 74 or another network, or may be communicatively coupled directly. In some examples, content preparation device 20 and server device 60 may comprise the same device.

[0017] In some embodiments, the techniques of this disclosure may be accomplished by, for example, two user equipment (UE) devices communicating over network 74. In such cases, both UE devices may include components similar to those of content preparation device 20, server device 60, and client device 40. When one UE device is transmitting media data, it may perform functionality attributed to content preparation device 20 and server device 60, and the other UE device receiving the media data may perform functionality attributed to client device 40. Both transmission and reception of media data may be performed in parallel by each of the UE devices.

[0018] 1 , content preparation device 20 comprises audio source 22 and video source 24. Audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data to be encoded by audio encoder 26. Alternatively, audio source 22 may comprise a storage medium that stores previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. Video source 24 may comprise a video camera that generates video data to be encoded by video encoder 28, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data. Content preparation device 20 is not necessarily communicatively coupled to server device 60 in all examples, but may store multimedia content on a separate medium that is read by server device 60.

[0019] The raw audio and video data may include analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may acquire audio data from a speaking participant while that participant is speaking, and video source 24 may simultaneously acquire video data of the speaking participant. In other examples, audio source 22 may comprise a computer-readable storage medium containing stored audio data, and video source 24 may comprise a computer-readable storage medium containing stored video data. Thus, the techniques described in this disclosure may be applied to live, streaming, real-time audio and real-time video data, or to archived, pre-recorded audio and video data.

[0020] An audio frame corresponding to a video frame is generally an audio frame that includes audio data captured (or generated) by audio source 22 contemporaneously with video data captured (or generated) by video source 24 that is included in the video frame. For example, audio source 22 captures audio data while a speaking participant generally generates the audio data by speaking, and video source 24 captures video data of the speaking participant contemporaneously, i.e., while audio source 22 captures the audio data. Thus, an audio frame may correspond in time to one or more particular video frames. Thus, an audio frame corresponding to a video frame generally corresponds to a situation in which the audio data and video data were captured contemporaneously, for which the audio frame and video frame, respectively, include contemporaneously captured audio data and video data.

[0021] In some examples, audio encoder 26 may encode a timestamp in each encoded audio frame representing the time the audio data for that encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp in each encoded video frame representing the time the video data for that encoded video frame was recorded. In such examples, audio frames corresponding to a video frame may include an audio frame including a timestamp and a video frame including the same timestamp. Content preparation device 20 may include internal clocks from which audio encoder 26 and / or video encoder 28 may generate timestamps, or which audio source 22 and video source 24 may use to associate audio data and video data, respectively, with timestamps.

[0022] In some examples, audio source 22 may send data to audio encoder 26 corresponding to the time the audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to the time the video data was recorded. In some examples, audio encoder 26 may encode a sequence identifier in the encoded audio data to indicate the relative temporal order of the encoded audio data, although the sequence identifier does not necessarily indicate the absolute time the audio data was recorded; similarly, video encoder 28 may also use the sequence identifier to indicate the relative temporal order of the encoded video data. Similarly, in some examples, the sequence identifier may be mapped with or correlated to a timestamp.

[0023] The audio encoder 26 generally generates a stream of encoded audio data, while the video encoder 28 generates a stream of encoded video data. Each individual stream of data (whether audio or video) may be referred to as an elementary stream. An elementary stream is a single digitally coded (and possibly compressed) component of a media presentation. For example, a coded video or audio portion of a media presentation may be an elementary stream. An elementary stream may be converted into a packetized elementary stream (PES) before being encapsulated within a video file. Within the same media presentation, a stream ID may be used to distinguish PES packets belonging to one elementary stream from others. The basic unit of data for an elementary stream is the packetized elementary stream (PES) packet. Thus, coded video data generally corresponds to an elementary video stream. Similarly, audio data corresponds to one or more respective elementary streams.

[0024] 1, encapsulation unit 30 of content preparation device 20 receives an elementary stream including coded video data from video encoder 28 and an elementary stream including coded audio data from audio encoder 26. In some examples, video encoder 28 and audio encoder 26 may each include a packetizer for forming PES packets from the coded data. In other examples, video encoder 28 and audio encoder 26 may each interface with a respective packetizer for forming PES packets from the coded data. In yet other examples, encapsulation unit 30 may include packetizers for forming PES packets from the coded audio data and the coded video data.

[0025] Video encoder 28 may encode the video data of the multimedia content in various ways to generate different representations of the multimedia content at various bit rates and with various characteristics, such as pixel resolution, frame rate, compliance with various coding standards, compliance with various profiles and / or profile levels for various coding standards, representations having one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. A representation, as used in this disclosure, may include one of audio data, video data, text data (e.g., for closed captioning), or other such data. The representation may include an elementary stream, such as an audio elementary stream or a video elementary stream. Each PES packet may include a stream_id that identifies the elementary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling the elementary streams into streamable media data.

[0026] Encapsulation unit 30 receives PES packets for the elementary streams of the media presentation from audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. Coded video segments may be organized into NAL units, which provide a "network-friendly" video representation compatible with applications such as video telephony, storage, broadcast, or streaming. NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may include block-, macroblock-, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some examples, a coded picture at one time instance may be contained within an access unit, which is typically presented as a primary coded picture and may contain one or more NAL units.

[0027] Non-VCL NAL units may include, among other things, parameter set NAL units and SEI NAL units. Parameter sets may contain sequence-level header information (in sequence parameter sets (SPS)) and picture-level header information that does not change frequently (in picture parameter sets (PPS)). With parameter sets (e.g., PPS and SPS), infrequently changing information does not need to be repeated for each sequence or picture, thus improving coding efficiency. Furthermore, the use of parameter sets may enable out-of-band transmission of important header information, eliminating the need for redundant transmission for error recovery. In an example of out-of-band transmission, parameter set NAL units may be transmitted on a different channel from other NAL units, such as SEI NAL units.

[0028] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding coded picture samples from VCL NAL units, but may assist processes related to decoding, display, error resilience, and other purposes. SEI messages may be included in non-VCL NAL units. SEI messages are normative parts of some standard specifications and therefore are not always mandatory in the implementation of standard-compliant decoders. SEI messages may be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information may be included in SEI messages, such as the scalability information SEI message in the example of SVC and the view scalability information SEI message in MVC. These exemplary SEI messages may convey information about, for example, operation point extraction and operation point characteristics.

[0029] Server device 60 includes a Real-time Transport Protocol (RTP) transmission unit 70 and a network interface 72. In some examples, server device 60 may include multiple network interfaces. Furthermore, any or all of the functionality of server device 60 may be implemented on other devices in the content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices in the content delivery network may cache data for multimedia content 64 and include components that substantially conform to those of server device 60. Generally, network interface 72 is configured to send and receive data over network 74.

[0030] The RTP unit 70 is configured to deliver media data to the client device 40 over the network 74 in accordance with RTP, standardized by the Internet Engineering Task Force (IETF) in Request for Comment (RFC) 3550. The RTP unit 70 may also implement protocols related to RTP, such as the RTP Control Protocol (RTCP), the Real-time Streaming Protocol (RTSP), the Session Initiation Protocol (SIP), and / or the Session Description Protocol (SDP). The RTP unit 70 may transmit the media data via a network interface 72, which may implement the Uniform Datagram Protocol (UDP) and / or the Internet Protocol (IP). Thus, in some examples, the server device 60 may transmit media data via RTP and RTSP over UDP using the network 74.

[0031] RTP unit 50 can receive an RTSP description request, for example, from client device 40. The RTSP description request can include data indicating what types of data are supported by client device 40. RTP unit 50 can respond to client device 40 with data indicating media streams, such as media content 64, that can be sent to client device 40 along with corresponding network location identifiers, such as uniform resource locators (URLs) or uniform resource names (URNs).

[0032] Next, the RTP unit 50 may receive an RTSP setup request from the client device 40. The RTSP setup request may generally indicate how the media stream is to be transported. The RTSP setup request may include a transport specifier, such as a network location identifier for the requested media data (e.g., media content 64) and a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. The RTP unit 50 may reply to the RTSP setup request with a confirmation and data indicating the port on the server device 60 to which the RTP data and control data will be sent. Next, the RTP unit 50 may receive an RTSP play request to “play” the media stream, i.e., send it over the network 74 to the client device 40. The RTP unit 50 may also receive an RTSP teardown request to terminate the streaming session, and in response, the RTP unit 50 may stop sending media data to the client device 40 for the corresponding session.

[0033] RTP receiving unit 52 may similarly initiate a media stream by first sending an RTSP description request to server device 60. The RTSP description request may indicate the type of data supported by client device 40. RTP receiving unit 52 may then receive a reply from server device 60 specifying available media streams, such as media content 64, that may be sent to client device 40 along with corresponding network location identifiers, such as uniform resource locators (URLs) or uniform resource names (URNs).

[0034] RTP receiving unit 52 may then generate an RTSP setup request and send the RTSP setup request to server device 60. As described above, the RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on client device 40. In response, RTP receiving unit 52 may receive a confirmation from server device 60 that includes the port of server device 60 that server device 60 will use to send the media data and control data.

[0035] After establishing the media streaming session between server device 60 and client device 40, RTP unit 50 of server device 60 may transmit media data (e.g., packets of media data) to client device 40 according to the media streaming session. Server device 60 and client device 40 may exchange control data (e.g., RTCP data), for example, indicating reception statistics by client device 40, which allows server device 60 to perform congestion control or otherwise diagnose and address transmission impairments.

[0036] Network interface 54 may receive media for a selected media presentation and provide it to RTP receiving unit 52, which may provide the media data to de-encapsulation unit 50. De-encapsulation unit 50 may de-encapsulate elements of the video file into constituent PES streams, de-packetize the PES streams to extract the encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, as indicated by the stream's PES packet headers, for example. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of the stream, to video output 44.

[0037] The video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiving unit 52, and decapsulation unit 50 may each be implemented as any of a variety of suitable processing circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware, or any combination thereof, where applicable. Each of the video encoder 28 and video decoder 48 may be included in one or more encoders or decoders, any of which may be integrated as part of a composite video encoder / decoder (codec). Similarly, each of the audio encoder 26 and audio decoder 46 may be included in one or more encoders or decoders, any of which may be integrated as part of a composite codec. An apparatus including video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiving unit 52, and / or decapsulation unit 50 may comprise an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular telephone.

[0038] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate in accordance with the techniques of this disclosure. By way of example, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may be configured to perform these techniques instead of (or in addition to) server device 60.

[0039] Encapsulation unit 30 may form NAL units that include a header that identifies the program to which the NAL unit belongs and a payload, e.g., audio data, video data, or data describing the transport or program stream to which the NAL unit corresponds. For example, in H.264 / AVC, an NAL unit includes a one-byte header and a variable-sized payload. NAL units that include video data within their payload may include various levels of granularity of the video data. For example, an NAL unit may include a block of video data, multiple blocks, a slice of video data, or an entire picture of video data. Encapsulation unit 30 may receive encoded video data from video encoder 28 in the form of PES packets of elementary streams. Encapsulation unit 30 may associate each elementary stream with a corresponding program.

[0040] Encapsulation unit 30 can also assemble access units from multiple NAL units. Generally, an access unit may comprise one or more NAL units for representing a frame of video data and, when such audio data is available, audio data corresponding to that frame. An access unit generally includes all NAL units for one output time instance, e.g., all audio and video data for one time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, specific frames for all views of the same access unit (same time instance) may be rendered simultaneously. In one example, an access unit may include a coded picture within one time instance, which may be presented as a coded primary picture.

[0041] Thus, an access unit may include all audio and video frames of a common temporal instance, e.g., all views corresponding to time X. This disclosure also refers to the coded pictures of a particular view as a "view component." That is, a view component may include coded pictures (or frames) for a particular view at a particular time. Thus, an access unit may be defined as including all view components of a common temporal instance. The decoding order of access units is not necessarily the same as the output or display order.

[0042] After encapsulation unit 30 assembles NAL units and / or access units into a video file based on the received data, encapsulation unit 30 passes the video file to output interface 32 for output. In some examples, encapsulation unit 30 may store the video file locally or transmit the video file to a remote server via output interface 32 instead of transmitting the video file directly to client device 40. Output interface 32 may comprise, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium such as an optical drive, a magnetic media drive (e.g., a floppy drive), a universal serial bus (USB) port, a network interface, or other output interface. Output interface 32 outputs the video file to a computer-readable medium such as, for example, a transmission signal, a magnetic medium, an optical medium, a memory, a flash drive, or other computer-readable medium.

[0043] Network interface 54 may receive NAL units or access units via network 74 and provide the NAL units or access units to de-encapsulation unit 50 via RTP receiving unit 52. De-encapsulation unit 50 may de-encapsulate elements of the video file into constituent PES streams, de-packetize the PES streams to extract the encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, as indicated by the stream's PES packet headers, for example. Audio decoder 46 decodes the encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes the encoded video data and sends the decoded video data, which may include multiple views of the stream, to video output 44.

[0044] In accordance with the techniques of this disclosure, client device 40 and server device 60 can be involved in a media communication session, such as an XR session. Client device 40 can represent a UE, and server device 60 can represent another UE or an edge server or other server device. As part of the media communication session, a user of client device 40 can interact with the virtual scene, for example, by moving, changing viewport / gaze direction, interacting with objects, speaking, etc. The various interactions can correspond to different types of media data, for example, audio, visual, changing virtual scene objects, or changing the user's viewpoint of the virtual scene.

[0045] Generally, to communicate such interactions, client device 40 can form packets, e.g., RTP packets, that include media data representing the interaction as well as header information indicating the media type for the interaction. In this manner, server device 60 can use the header information to determine the number of sets of media data to be included in the packet, the type of media data, and then extract the media data from the packet. The header information can correspond to RTP header extensions, as described in more detail below. The header information can further include attributes for each media type.

[0046] FIG. 2 is a block diagram illustrating elements of an exemplary video file 150. As described above, video files according to the ISO Base Media File Format and its extensions store data in a series of objects called "boxes." In the example of FIG. 2, video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. While FIG. 2 represents one example of a video file, it should be understood that other media files may contain other types of media data (e.g., audio data, timed text data, etc.) that are similarly structured to the data in video file 150 according to the ISO Base Media File Format and its extensions.

[0047] The file type (FTYP) box 152 generally describes the file type of the video file 150. The file type box 152 may contain data identifying specifications that represent the best use of the video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the movie fragment box 164, and / or the MFRA box 166.

[0048] 2, MOOV box 154 includes a Movie Header (MVHD) box 156, a Track (TRAK) box 158, and one or more Movie Extension (MVEX) boxes 160. Generally, MVHD box 156 may describe general characteristics of video file 150. For example, MVHD box 156 may include data describing when video file 150 was originally created, data describing when video file 150 was last modified, data describing the timescale of video file 150, data describing the duration of playback of video file 150, or other data that generally describes video file 150.

[0049] TRAK box 158 may contain data for a track of video file 150. TRAK box 158 may contain a track header (TKHD) box that describes characteristics of the track corresponding to TRAK box 158. In some examples, TRAK box 158 may contain coded video pictures, while in other examples, the coded video pictures for a track may be contained within a movie fragment 164 that may be referenced by data in TRAK box 158 and / or data in sidx box 162.

[0050] In some examples, video file 150 may include two or more tracks. Thus, MOOV box 154 may include a number of TRAK boxes equal to the number of tracks in video file 150. TRAK box 158 may describe characteristics of the corresponding track in video file 150. For example, TRAK box 158 may describe temporal and / or spatial information of the corresponding track. If encapsulation unit 30 (FIG. 1) includes a parameter set track in a video file such as video file 150, a TRAK box similar to TRAK box 158 in MOOV box 154 may describe characteristics of the parameter set track. Encapsulation unit 30 may signal, within the TRAK box describing the parameter set track, the presence of a sequence-level SEI message for the parameter set track.

[0051] MVEX box 160 may, for example, describe the characteristics of the corresponding movie fragment 164 to signal that video file 150 includes movie fragment 164 in addition to the video data contained in MOOV box 154, if any. In the context of streaming video data, coded video pictures may be contained in movie fragment 164 rather than in MOOV box 154. Thus, all coded video samples may be contained in movie fragment 164 rather than in MOOV box 154.

[0052] The MOOV box 154 may contain a number of MVEX boxes 160 equal to the number of movie fragments 164 in the video file 150. Each MVEX box 160 may describe the characteristics of a corresponding movie fragment 164. For example, each MVEX box may contain a Movie Extension Header (MEHD) box that describes the duration of the corresponding movie fragment 164.

[0053] As mentioned above, encapsulation unit 30 may store a sequence data set within a video sample that does not contain actual coded video data. A video sample may generally correspond to an access unit, which is a representation of a coded picture at a particular time instance. In the context of AVC, a coded picture includes one or more VCL NAL units containing information for configuring all pixels of the access unit, as well as other associated non-VCL NAL units, such as SEI messages. Accordingly, encapsulation unit 30 may include a sequence data set, which may include a sequence-level SEI message, within one of movie fragments 164. Encapsulation unit 30 may further signal the presence of the sequence data set and / or the sequence-level SEI message as being present within one of movie fragments 164 within one of MVEX boxes 160 corresponding to one of movie fragments 164.

[0054] The SIDX box 162 is an optional element of the video file 150; that is, a video file conforming to the 3GPP file format or other such file formats does not necessarily include the SIDX box 162. According to an example 3GPP file format, the SIDX box can be used to identify a subsegment of a segment (e.g., a segment contained within the video file 150). The 3GPP file format defines a subsegment as "a self-contained set of one or more consecutive Movie Fragment boxes corresponding to a Media Data box(es), where the Media Data box containing the data referenced by a Movie Fragment box follows that Movie Fragment box and precedes the next Movie Fragment box containing information about the same track." The 3GPP file format also indicates that a SIDX box "contains a sequence of references to subsegments of the (sub)segment documented by the box. The referenced subsegments are contiguous in presentation time. Similarly, the bytes referenced by a Segment Index box are always contiguous within the segment. The referenced size gives a count of the number of bytes in the referenced material."

[0055] SIDX box 162 generally provides information describing one or more subsegments of a segment contained within video file 150. For example, such information may include the playback time at which the subsegment begins and / or ends, a byte offset for the subsegment, whether the subsegment contains (e.g., begins with) a stream access point (SAP), the type of SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), the location of the SAP (in terms of playback time and / or byte offset) within the subsegment, etc.

[0056] A movie fragment 164 may include one or more coded video pictures. In some examples, a movie fragment 164 may include one or more groups of pictures (GOPs), each of which may include multiple coded video pictures, e.g., frames or pictures. Additionally, as described above, a movie fragment 164 may include a sequence data set in some examples. Each movie fragment 164 may include a Movie Fragment Header Box (MFHD, not shown in FIG. 2). The MFHD box may describe characteristics of the corresponding movie fragment, such as the movie fragment's sequence number. Movie fragments 164 may be included in video file 150 in sequence number order.

[0057] MFRA box 166 can describe random access points within movie fragments 164 of video file 150. This can assist in performing trick modes, such as performing a search for a specific temporal location (i.e., playback time) within a segment encapsulated by video file 150. MFRA box 166 is generally optional in some examples and need not be included in a video file. Similarly, a client device, such as client device 40, does not necessarily need to reference MFRA box 166 to correctly decode and display video data in video file 150. MFRA box 166 may include multiple track fragment random access (TFRA) boxes (not shown) equal in number to the number of tracks in video file 150, or in some examples, may include multiple TFRA boxes equal in number to the number of media tracks (e.g., non-hint tracks) in video file 150.

[0058] In some examples, movie fragment 164 may include one or more stream access points (SAPs), such as an IDR picture. Similarly, MFRA box 166 may provide an indication of the location of the SAPs within video file 150. Thus, a temporal subsequence of video file 150 may be formed from the SAPs of video file 150. The temporal subsequence may also include other pictures, such as P-frames and / or B-frames, that depend on the SAPs. Frames and / or slices of a temporal subsequence may be ordered within a segment such that frames / slices of the temporal subsequence that depend on other frames / slices of the subsequence can be properly decoded. For example, in a hierarchical organization of data, data used for prediction for other data may also be included within a temporal subsequence.

[0059] FIG. 3 is a conceptual diagram illustrating an exemplary Web Real-Time Communications (WebRTC) protocol stack 180. WebRTC is a protocol suite that supports direct, interactive RTC between web browsers or between browsers and other entities using audio, video, collaboration, gaming, and the like. FIG. 3 illustrates an exemplary WebRTC protocol stack 180. In the WebRTC framework, communications between parties consist of media (e.g., audio and video) and non-media data. Media is transmitted using the Secure Real-time Transport Protocol (SRTP), and non-media data is handled by using the Stream Control Transmission Protocol (SCTP). Priorities associated with media or data flows are categorized in the API as “very low,” “low,” “medium,” or “high.” WebRTC implementations may attempt to set QoS on transmitted packets according to the guidelines in IETF RFC 8837, “Differentiated Services Code Point (DSCP) Packet Markings for WebRTC QoS.”

[0060] FIG. 4 is a conceptual diagram illustrating an RTP header format 190. The RTP header format 190 of FIG. 4 is defined in IETF RFC 3550, "RTP: A Transport Protocol for Real-Time Applications," July 2003. As shown in FIG. 4, the RTP header format 190 includes a payload type (PT) field 192. According to RFC 3550, the PT field 192 has a value that identifies the format of the RTP payload and determines its interpretation by an application. For example, a static PT value of 31 indicates H.261 video coding.

[0061] Dynamic payload types can be defined through other protocols. For example, an RTP packet stream can be associated with an SDP "m=" line by comparing the RTP payload type number used by the RTP packet stream with the payload type signaled in the "a=rtpmap:" line in the media section of the SDP. An example of a media representation in SDP is shown below, where the PT value 98 indicates ITU-T H.264 / Advanced Video Coding (AVC).

[0062] [Table 1]

[0063] Figure 5 is a conceptual diagram illustrating an example RTP header extension 200. RTP provides an extension mechanism for carrying additional information in the RTP packet header. If the extension (X) bit in the RTP header is 1, a variable-length header extension must be added to the RTP header following the CSRC list (if present). Two types of RTP header extensions are defined in IETF RFC 8285, "A General Mechanism for RTP Header Extensions," October 2017. The one-byte header form of the extension has a fixed first 16-bit bit pattern of 0xBEDE, and each extension element begins with a byte containing an ID and length. The four-bit length is the number of data bytes of this header extension element following the one-byte header minus one.

[0064] 5 illustrates one embodiment of a one-byte header extension having three extension elements 202A, 202B, and 202C and padding data 204A and 204B. Specifically, extension element 202A includes an ID 206A, a length value 208A, and data 210A, extension element 202B includes an ID 206B, a length value 208B, and data 210B, and extension element 202C includes an ID 206C, a length value 208C, and data 210C.

[0065] Figure 6 is a conceptual diagram illustrating one embodiment of a two-byte header form 220. Specifically, the embodiment of Figure 6 is a two-byte (16-bit) RTP header field. In this embodiment, a four-bit "appbits" field 222 can have an application-dependent value and can be defined to have any value or meaning.

[0066] 7 is a conceptual diagram illustrating one embodiment of a two-byte RTP header extension 230. Each RTP header extension element may begin with a byte of data that contains an identifier (ID) 232A, 232B, 232C and another byte of data that contains a length value 234A, 234B, 234C.

[0067] According to RFC 3550, RTP header extensions are to be used for data that can be safely ignored by the receiver without affecting interoperability.

[0068] WebRTC supports extensions such as client-to-mixer audio levels, mixer-to-client audio levels, media stream identification, and video orientation adjustment. WebRTC transmits all types of media in a single RTP session and uses RTP and RTCP multiplexing to further reduce the number of required UDP ports (IETF RFC8834 "Media Transport and Use of RTP in WebRTC", January 2021). RTP allows for the definition of modified or additional fixed capability fields in the RTP header, which immediately follow the SSRC field of the existing fixed header in profile specifications.

[0069] The RTP Control Protocol (RTCP) is based on the periodic transmission of control packets to all endpoints in a session. The percentage of session bandwidth added for RTCP can be fixed at 5%. Low-latency RTCP feedback (FB) messages are classified into three categories according to IETF RFC4585, "Extended RTP Profile for Real-time Transport Control Protocol (RTCP)-Based Feedback (RTP / AVPF)," July 2006: transport-layer FB messages, payload-specific FB messages, and application-layer FB messages.

[0070] 8 is a conceptual diagram illustrating an exemplary common packet format 240 for feedback messages. Application layer FB messages provide a mechanism for transparently conveying feedback from a receiver (e.g., client device 40 of FIG. 1) to a sender (e.g., server device 60 of FIG. 1). Data exchanged between two application instances executed by such devices is typically defined in an application protocol specification.

[0071] In the embodiment of Figure 8, there is a 5-bit FMT field 242 that indicates the associated type of FB message. The embodiment of Figure 8 also illustrates an 8-bit PT field 244 that indicates the RTCP packet type, such as a transport layer FB message (RTPFB) or a payload-specific FB message (PSFB). An application layer FB message is a specific case of a payload-specific message and can be identified by PT = PSFB and FMT = 15. There may be exactly one application layer FB message contained in the FCI field, unless the application layer FB message structure itself allows for stacking.

[0072] 9 is a conceptual diagram illustrating an example Stream Control Transmission Protocol (SCTP) chunk field format 250. WebRTC supports multiple simultaneous reliable and unreliable data channels. Reliable data channels can transport important real-time game control information or non-real-time text chat. SCTP provides the following features for transporting non-media data between browsers, in accordance with IETF RFC8831, "WebRTC Data Channels," January 2021: support for multiple unidirectional streams, ordered and unordered delivery of user messages, and reliable and partially reliable transport of user messages.

[0073] In SCTP, a sender cannot put more than one application message into an SCTP user message, and when message interleaving is not supported, the maximum message size is 16 KB. SCTP ensures that messages are delivered to SCTP users in order within a stream. While one stream is blocked waiting for the next in-order user message, delivery from other streams can proceed. SCTP can also bypass the ordered delivery service and allow user messages to be delivered to the user as soon as a particular user message is received.

[0074] An SCTP packet includes a common header and one or more chunks. The common header includes a source port number, a destination port number, a verification tag, and a checksum. Each chunk can contain either control information or user data. Figure 9 illustrates an exemplary chunk field format 250. In this example, chunk field format 250 includes a chunk type field 252, a chunk flag 254, a chunk length field 256, and a chunk value 258.

[0075] Figure 10 is a conceptual diagram illustrating an example SCTP data chunk format 260. If the chunk type is payload data, the chunk format 260 as illustrated in Figure 10 can be used for the data chunk. The U bit 262 indicates an unordered data chunk. The B bit 264 and E bit 266 indicate a fragmented user message. The payload protocol identifier 268 represents an application-specific protocol identifier passed by upper layers to SCTP for transmission to its peers.

[0076] SCTP Partial Reliability Extension (PR-SCTP) (IETF RFC3758 "Stream Control Transmission Protocol (SCTP) Partial Reliability Extension", May 2004, The SCTP standard (pertaining to the RFC 2448 standard) provides an ordered, unreliable data transfer service and allows an SCTP endpoint to signal its peer to move a cumulative acknowledgement (ack) point forward when any missing data should not be sent or retransmitted by the sender (e.g., server device 60).

[0077] 3GPP SA4 defines a number of media types for immersive RTC. These types can be categorized as follows: Device information data specifies device capabilities such as FOV, sensor information, camera information, and projection information. Spatial and object description data describes spatial or object information such as scene descriptions, spatial descriptions, and 3D visual models. Interactivity data represents user interactions such as user pose with four quaternions for orientation, three vectors for position, and a 64-bit timestamp; viewport with six parameters carried via RTCP in IMS telephony (3GPP TS 26.114 "IP Multimedia Subsystem (IMS) Multimedia Telephony Media handling and interaction", June 2022), gestures, body movements, facial expressions, and AR anchor points.

[0078] Device information data may be static. Device information can be exchanged at the start of a session or WebRTC interactive connection establishment (IETF RFC8445 "Interactive Connectivity Establishment (ICE)", July 2018), and the data can be carried in the Session Description Protocol (SDP) (IETF RFC4566 "SDP: Session Description Protocol", July 2008). Device information updates can also be carried in RTP or RTCP application layer FB messages.

[0079] The spatial and object description information can be exchanged from time to time by events when the scene or 3D model is updated or changed. The data size can be large and does not need to change frequently.

[0080] Interactivity data is important in immersive RTC applications. Some interactivity data, such as pose, viewport, or gesture, can trigger events on the receiving side, and the sender may expect a response as soon as possible. Therefore, very low latency transportation is desirable. Some interactivity data, such as facial expressions or body movements, can manipulate a 3D avatar model on the receiving side and be synchronized with specific media data (e.g., mouth movements may need to be synchronized with audio data). Therefore, a synchronization scheme may be required when transporting interactivity data. Depending on the application, several media types can be delivered together. Some media types may be delivered reliably, while others may be delivered unreliably. Some media types may require burst transport, while others may be transported continuously or periodically. The following table lists the characteristics of the media types. A transport protocol and payload format design that supports these media types is desirable.

[0081] [Table 2]

[0082] 11 is a conceptual diagram illustrating an example media transmission architecture capable of implementing the techniques of this disclosure. The example of FIG. 11 illustrates a high-level architecture for metadata transmission. In this example, the architecture includes a UE 300 and a UE / server 290. That is, the UE / server 290 may correspond to a UE or a server such as an edge server. In this example, the UE / server 290 includes a clock 292, an OpenXR runtime 294, and a metadata parser 296, and the UE 300 includes a clock 302, an OpenXR runtime 304, and a metadata aggregator 306.

[0083] The OpenXR runtime 304 of the UE 300 can collect and / or derive metadata from sensor input using specific application programming interfaces (APIs), e.g., the OpenXR API. The metadata aggregator 306 can extract essential metadata from the OpenXR runtime 304 and transport the metadata with associated timing information to the remote UE / server 290 over connection 310, e.g., in the form of data contained in a data channel or RTP / RTCP header extensions. The UE / server 290 can then receive and parse the metadata carried in the input stream using a metadata parser 296 and pass the metadata and corresponding timing information to the OpenXR runtime 294 for various operations, such as rendering.

[0084] FIG. 12 is a conceptual diagram illustrating an example RTP extension 320 and payload for media types according to the techniques of this disclosure. According to the techniques of this disclosure, the server device 60 and the client device 40 (FIG. 1) can exchange device information using the SDP framework for immersive RTC. As an example, extensions to RTP can be used to convey additional information, such as device information, to enhance immersive RTC. In this example, the RTP header extension 320 can include a series of sets of header values ​​324A, 324B, and 324C, each of which includes an ID value, a length value, a media type value, a media attribute value, and a media length value for each of the series of sets of media data payloads 322A, 322B, and 322C. SDP attributes, such as 3gpp_irtcw and / or 3gpp_ibacs, can be defined to indicate an immersive WebRTC stream or an IMS-based AR conversation stream, respectively. Values ​​of these attributes can include camera or display information, such as FOV, projection, resolution, refresh rate, default pose, and orientation.

[0085] Spatial and object information data can be transported over RTP, and the payload format can use existing RTP header fields specified in RFC 3550. A mark bit (M) can indicate the last packet in the transmission order of the spatial and object information data. The payload type can indicate a specific media type, such as a scene description, a spatial description, or a 3D visual model. The assignment of RTP payload types can be performed in a dynamic manner and indicated in the SDP media description lines.

[0086] In some embodiments, the RTP header extension 320 can be used to indicate each media type and data size when multiple media types are bundled in an RTP packet. The data field of each extension element can contain, following the ID and length fields, values ​​that specify the media type and media data size carried in the RTP payload. Additional indications or flags may be included in the header extension element data field to specify media type attributes.

[0087] One exemplary attribute is a synchronization group. Media types that belong to the same synchronization group can be synchronized at the receiving end (e.g., client device 40). Another exemplary attribute can indicate whether the media type data should be synchronized with other media data, such as audio or video. Media types that do not require synchronization can be assigned to a default group (e.g., group 0).

[0088] In the example of Figure 12, it is assumed that there are three data types (Type 0, Type 1, and Type 2). In the example of Figure 12, there are three sets of RTP header extensions (324A, 324B, 324C), one for each media type. Attributes can specify, for example, whether the corresponding media type is to be synchronized with another media type, and if so, which of the other media types.

[0089] 13 is a conceptual diagram illustrating an example RTP extension element format 330 for interactivity media type data. The RTP extension element format 330 includes an ID field 332, a length field 334, a media type field 336, a media type attribute field 338, a media type data length field 340, media type data 342, and padding data 344. In this embodiment, the RTP extension can be used to carry media data types. Each extension element can carry one media type data. The media type attribute field 338 can include data specifying attributes such as priority, synchronization indication, timestamp, and / or reliability information. The extension element can also specify the media type data length in bytes in the media type data length field 340.

[0090] For some latency-tolerant media types, such as viewports, interactivity media type data can be transported over RTCP. Application layer FB messages can be used to carry media type data. The contents of a feedback control information (FCI) entry in an RTCP packet can include a media type identifier, a media type data structure, a sequence number, a timestamp, and media type attributes. When multiple media types are bundled together in a single RTCP packet, indicators to indicate the number of media types and the length of each media type data may be included in the FCI entry.

[0091] OpenXR specifies an action set, XrActionSet, that contains one or more actions (e.g., XrAction). An RTP / RTCP header extension can carry one XrActionSet, which contains an action set header followed by one or more action metadata payloads.

[0092] The action set header may include a 16-bit field indicating the action set metadata type, a 16-bit field indicating the action set attributes, a 64-bit field indicating the action set name, a 32-bit field indicating the action set priority, a 32-bit field indicating the number of actions associated with the action set, and a 16-bit length field indicating the length of one action metadata payload.

[0093] The action metadata payload can include a 16-bit field indicating the action type, a 16-bit field indicating the action attributes, and a 32-bit field indicating the action type as an XrActionType as specified in OpenXR. The action metadata field can carry a corresponding OpenXR extension data structure.

[0094] In some embodiments, the RTP header extension can carry one XR action due to RTP data length constraints. The extension element data can carry the action metadata payload, as described above. An additional priority field can be defined to indicate the priority of the action set with which the action is associated.

[0095] In some embodiments, the OpenXR action set metadata can be converted to JavaScript Object Notation (JSON) format. The JSON formatted metadata may be further compressed and carried in RTP / RTCP header extension element data or media type data.

[0096] FIG. 14 is a conceptual diagram illustrating an example immersive media profile (IMP) 350. The IMP profile may be a profile of RTP. The IMP may define immersive media-specific interpretations of profile-dependent fields. In the IMP, the RTP header fields X, CC, and M may be merged into attribute fields. One attribute field may indicate whether a timestamp is present. Another attribute field may indicate whether two or more media types are included in the packet. When multiple media types are bundled into a common packet, the number of media types may be signaled.

[0097] In the example of Figure 14, three media types are included in the RTP packet. In one example, the fields have the following meanings: T 352: 1 bit. If the T bit is set, a timestamp must be present in the header. N 354: 4 bits. This field indicates the total number of media types present in the packet. A value of 0 is not valid. · PT 356A, 356B, 356C: This field indicates the media type of the corresponding media data 360A, 360B, 360C contained in this packet. Length 358A, 358B, 358C: 16 bits. This field indicates the media type data length in 16 bits.

[0098] In some embodiments, additional media type attribute fields such as synchronization, reliability, and priority may be specified for each media type in a packet. The synchronization field may indicate whether associated media types share the same timestamp of the packet. The reliability field may indicate whether associated media type data should be delivered, and packet loss may be reported to the sender to request retransmission.

[0099] FIG. 15 is a conceptual diagram illustrating an example SCTP chunk format 370 for an immersive media type. According to the techniques of this disclosure, immersive media type data can be carried using a WebRTC data channel over SCTP. Each interactivity media type can be transmitted in its own SCTP stream. Alternatively, multiple media types, such as synchronized media types, can be bundled into a single SCTP stream. For example, an application executed by client device 40 or server device 60 can configure data channel properties, such as reliable or unreliable transmission, in-order or out-of-order delivery, and priority for transporting particular media types.

[0100] Each media type data can be assigned to a data payload chunk within an SCTP packet. The payload protocol identifier can indicate the media type specified by a higher application layer. The media type data can also be stored in the payload user data.

[0101] Since SCTP does not provide a timestamp in the payload data chunk header, an attribute field can be added to indicate whether a timestamp is present in the chunk. The timestamp 372 can be signaled in a fixed position in the chunk for the associated media type data, for example, in a 32-bit field following the payload protocol identifier.

[0102] In some embodiments, the payload protocol identifier field of the SCTP payload data chunk type (type=0) can indicate a dedicated application protocol, and the chunk user data can carry the protocol data. The application protocol header can carry attributes such as data boundaries, media type, timestamp, sequence number, synchronization indication, and data length, and the application protocol payload data can carry the media type data structure.

[0103] In some embodiments, a new type of chunk can be defined for each interaction media type, and the chunk flags field can include media type attributes such as a sync attribute to indicate whether an associated timestamp is present in the chunk. An unordered attribute can indicate whether the data chunk is ordered or unordered.

[0104] In the example of Figure 15, the S bit 374 is the synchronization bit. When set to "1", this bit indicates that this is a synchronized data chunk and that a timestamp has been assigned to this data chunk. Otherwise, the receiver can ignore the value of the timestamp field.

[0105] 16 is a conceptual diagram illustrating another embodiment of an SCTP chunk format 380 for immersive media type data. In this embodiment, a specific type can be defined for every immersive media type. A payload type field 382 may be assigned to the immersive media type data chunk to indicate the specific media type.

[0106] FIG. 17 is a conceptual diagram illustrating an exemplary timed metadata payload format 390 in accordance with the techniques of this disclosure. In this example, the timed metadata payload format 390 includes a data channel payload header 392, an action set header 394, and an action set 398 that includes action metadata 396A-396N. In OpenXR, actions are created at initialization and are then used to request input device state, create action spaces, or control haptic events. An action set is a collection of actions. FIG. 17 illustrates an exemplary data chunk structure for action set timed metadata.

[0107] In one embodiment, the data chunk payload header is the first 16 bytes of the data chunk, and the payload protocol identifier (PPID) is the value "53" for "WebRTC Binary". The user data can include an action set, XrActionSet, defined in OpenXR. The action set can include one or more actions, XrAction.

[0108] FIG. 18 is a block diagram illustrating an exemplary action set header format 400 in accordance with the techniques of this disclosure. In this embodiment, the action set header format 400 includes a metadata type value 402, an action set attribute 404, an action set name 406, an action set priority 408, a number of actions 410, and action lengths 412A-412N for each of the number of actions. The metadata type value 402 can indicate that the payload is for an action set. The action set attribute 404 can be a 16-bit field indicating the action set attribute. The action set name 406 can be a 64-bit field indicating the name of the action set. The action set priority 408 can be a 32-bit field indicating the priority of the action set. The number of actions 410 can be a 32-bit field indicating the total number of actions included in the action set. Each of the action lengths 412A-412N can be a 16-bit field indicating the length of a corresponding action among the actions included in the action set.

[0109] 19 is a block diagram illustrating an exemplary action metadata payload format 420 in accordance with the techniques of this disclosure. In this example, action metadata payload format 420 includes metadata type 422, action attributes 424, action type 426, and action metadata 428. Metadata type 422 may be a 16-bit field indicating a registered extension number of an action defined in OpenXR. Action attributes 424 may be a 16-bit field indicating attributes of the action. Action type 426 may be a 32-bit field indicating an XrActionType enumeration value specified in OpenXR. Action metadata 428 may include timed metadata associated with the action.

[0110] If the data length of the action set is large, each data chunk may carry one action metadata, and the payload format may be the same as the action metadata payload format. An additional priority field can be defined to indicate the priority of the action set with which the action is associated. An additional action set ID field can be used to indicate the corresponding action set with which the action is associated.

[0111] FIG. 20 is a block diagram illustrating an example eye-gaze interaction metadata action payload format 430 in accordance with the techniques of this disclosure. The eye-gaze interaction metadata action payload format 430 can be used for OpenXR's XR_EXT_eye_gaze_interaction. The eye-gaze interaction metadata action payload format 430 of FIG. 20 includes a metadata type 432, a rev 434, a flag 436, an action type 438, an XR gaze sample time ext:time 440, and an XR pose f 442. The metadata type field 432 can have a value of 31. The flag 436 may be a 4-bit field that maps to XrSpaceLocationFlags defined in OpenXR. The action type 438 can have a value of "XR_ACTION_TYPE_POSE_INPUT" based on OpenXR.

[0112] XR Pose f 442 can represent seven floating-point elements, including XrQuaternionf(x, y, z, w) and XrVector3f(x, y, z). That is, each component of XrQuaternion(x, y, z, and w) and XrVector3f(x, y, and z) can be represented as a respective floating-point value. XrQuaternion can represent the user's three-dimensional rotation in virtual space, and XrVector3 values ​​can represent the user's three-dimensional position in virtual space. Each element can be represented by a 32-bit single-precision floating-point number or a 64-bit double-precision floating-point number, as specified in IETF RFC4506.

[0113] In some embodiments, additional timing information, such as predicted time, can be carried in addition to OpenXR XrTime. For example, hand tracking predictions at a predicted time of 100 milliseconds and a predicted time of 1 second can both be delivered to a server. The server can select the appropriate tracking prediction based on varying end-to-end delay.

[0114] In some embodiments, the required elements or fields of OpenXR action sets and actions can be converted to JSON format. The JSON-formatted metadata payload can be further compressed and carried in the data chunk user data payload.

[0115] 21 is a block diagram illustrating an example spatial location metadata payload format 450 in accordance with the techniques of this disclosure. An OpenXR application uses the xrLocateSpace function to find the pose of an XrSpace origin within a base XrSpace at a given historical or predicted time. Location-specific spatial elements include a base space, a space, a time, and a spatial location as defined in OpenXR.

[0116] The spatial location XrSpaceLocation includes XrSpaceLocationFlagBits, XrSpaceVelocity, and XrPosef. The spatial velocity includes a velocity flag, a linear velocity in meters per second, and an angular velocity in radians per second. These values ​​may be contained in a data chunk structure such as that shown in Figure 21.

[0117] Spatial location metadata payload format 450 of FIG. 21 includes metadata type 452, rev field 454, flags 456, XR space:space 458, XR space:basespace 460, XR time:time 462, and XR pose f 464. Metadata type 452 may include values ​​defined for XR spatial locations. Flags 456 may be a 4-bit field indicating XR spatial location flags as defined in OpenXR. XR space:space 458 and XR space:basespace 460 may be respective 32-bit fields that together indicate a spatial handle as defined in OpenXR. XR time:time 462 may be a sample timestamp expressed in nanoseconds.

[0118] XR Pose f 464 can represent seven floating-point elements, including XrQuaternionf(x,y,z,w) and XrVector3f(x,y,z). That is, each component of XrQuaternion(x,y,z, and w) and XrVector3f(x,y, and z) can be represented as a respective floating-point value. XrQuaternion can represent the user's three-dimensional rotation in virtual space at the sample time, and XrVector3 value can represent the user's three-dimensional position in virtual space at the sample time. Each element can be represented by a 32-bit single-precision floating-point number or a 64-bit double-precision floating-point number, as specified in IETF RFC4506.

[0119] FIG. 22 is a block diagram illustrating an exemplary spatial rate metadata payload format 470 in accordance with the techniques of this disclosure. The spatial rate metadata payload format 470 in the example of FIG. 22 includes metadata type 472, rev 474, flags 476, XR space:space 478, XR space:basespace 480, XR time:time 482, linear velocity 484, and angular velocity 486. Metadata type 472 can have values ​​defined for XR spatial rates. Flags 476 may be a 4-bit field indicating an XR spatial rate flag as defined in OpenXR. XR space:space 478 and XR space:basespace 480 may be respective 32-bit fields that together indicate a spatial handle as defined in OpenXR.

[0120] Linear velocity 484 can have three floating-point elements, including XrVector3f(x,y,z). Each element can be represented by a 32-bit single-precision floating-point number or a 64-bit double-precision floating-point number, as specified in IETF RFC 4506. The elements of an XrVector3 value can represent component velocities in specific directions, such as up / down, left / right, and forward / backward, so that together these elements represent velocity in a specific three-dimensional vector direction at the sample time.

[0121] Angular velocity 486 can have three floating-point elements, including XrVector3f(x,y,z). Each element can be represented by a 32-bit single-precision floating-point number or a 64-bit double-precision floating-point number, as specified in IETF RFC 4506. The elements of an XrVector3 value can represent component angular velocities in specific Euler angle directions, so that together they represent a rotation around a specific Euler angle at the sample time.

[0122] In some embodiments, additional timing information, such as predicted time, can be carried in addition to OpenXR XrTime. For example, spatial location predictions at predicted times of 100 milliseconds and 1 second can both be delivered to a server. The server can select an appropriate predicted location based on varying end-to-end delay.

[0123] In some embodiments, the association between the spatial velocity XrSpaceVelocity and the spatial location XrSpaceLocation may be indicated in the payload.

[0124] In some embodiments, the required elements or fields of an OpenXR spatial location or spatial velocity can be converted to JSON format. The JSON formatted payload can be further compressed and carried in a data chunk user data payload.

[0125] 23 is a block diagram illustrating an example view metadata payload format 490 in accordance with the techniques of this disclosure. OpenXR applications can use xrLocateViews to retrieve the viewer pose and projection parameters needed to render each view for use in the constituent projection layers. The xrLocateViews function returns view and projection information for a particular display time. XrView includes, for example, the user's pose and field of view (fov) at a particular time.

[0126] The view metadata payload format 490 of Figure 23 includes metadata type 492, rev 494, flags 496, XR space:space 498, XR space:display time 500, XR pose f:pose 502, and XR pose f:field of view (fov) 504. Metadata type 492 may have values ​​defined for XrView. Flags 496 may be a 4-bit field indicating XrViewStateFlags as defined in OpenXR. XR space:space 498 may indicate a space handle as defined in OpenXR. XR time:display time 500 may be a display time in nanoseconds.

[0127] XR Pose f: Pose 502 can include seven floating-point elements, including XrQuaternionf(x, y, z, w) and XrVector3f(x, y, z). That is, each component of XrQuaternion(x, y, z, and w) and XrVector3f(x, y, and z) can be represented as a respective floating-point value. XrQuaternion can represent the user's three-dimensional rotation in virtual space, and XrVector3 values ​​can represent the user's three-dimensional position in virtual space at display time. Each element can be represented by a 32-bit single-precision floating-point number or a 64-bit double-precision floating-point number, as specified in IETF RFC 4506.

[0128] XRpose:fov 504 may contain four floating-point elements: XrFovf (angleLeft, angleRight, angleUp, angleDown). Each element can be represented by a 32-bit single-precision floating-point number or a 64-bit double-precision floating-point number, as specified in IETF RFC4506. The elements of XRpose:fov 504 can define the user's field of view at display time.

[0129] Multiple views may be bundled into a data chunk, or each view may be carried in a separate data chunk. When bundling multiple views, an XrLocateViews header may be added to indicate the number of views, the length of each view element, the common space, and the display time shared by the view elements. In some embodiments, the required elements or fields of an OpenXR view element may be converted to JSON format, and the JSON-formatted metadata payload can be further compressed and carried in the data chunk user data payload.

[0130] 24 is a flowchart illustrating an example method for transmitting user interaction data and associated metadata to another device involved in an XR multimedia communication session in accordance with the techniques of this disclosure. The method of FIG. 24 is described with respect to the UE 300 of FIG. 11. It should be understood that other devices, such as the client device 40 of FIG. 1, may also perform this or similar methods.

[0131] Initially, the UE 300 may present a representation of a virtual scene to a user of the UE 300, for example, via a heads-up display (HUD). The UE 300 may also be communicatively coupled to various input devices, such as one or more controllers, gamepads, keyboards, mice, etc. The user may interact with the virtual scene in various ways, for example, by physically moving, rotating, or changing a viewport (e.g., using controller input), touching virtual objects (e.g., as determined by collision detection between a representation of the user's avatar and objects in the virtual scene, button presses, etc.), speaking, or other such interactions. The UE 300 may determine these user interactions with the virtual scene (510).

[0132] The UE 300 may determine a media type for each of the user interactions (512). For example, the media type may include animations associated with a virtual object being moved, audio data collected from the user, user poses, viewports, gestures, body movements, facial expressions, etc. The UE 300 may construct a packet containing a respective set of data for each of the user interactions (514). Similarly, the UE 300 may add header data to the packet indicating the media type and attributes of the user interaction (516), for example, as described above with respect to FIGS. 12-23. The UE 300 may then transmit the packet to the UE / server 290 representing another device involved in the XR multimedia communication session (518).

[0133] In this manner, the method of FIG. 24 represents one embodiment of a method that includes determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmitting the packets to devices involved in the media communication session.

[0134] The following clauses represent various embodiments of the techniques of this disclosure.

[0135] Clause 1: A method for receiving media data, the method comprising: receiving packets containing media data of a media communication session, the packets including header data indicating one or more media types of the media data and one or more attributes of each of the one or more media types; extracting the media data of each of the one or more media types; processing the media data of each of the one or more media types in accordance with the corresponding one or more attributes; and presenting the media data.

[0136] Clause 2: The method of clause 1, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, user default pose, user default orientation, media data timestamp, priority data, reliability indication, or whether two or more media types are to be presented synchronously.

[0137] Clause 3: The method of clause 1 or 2, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

[0138] Clause 4: The method of any one of clauses 1 to 3, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value that indicates whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

[0139] Clause 5: The method of any one of clauses 1 to 4, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value indicating whether the media data represents a pause, a gesture, or a viewport.

[0140] Clause 6: The method of any one of clauses 1 to 5, wherein the one or more media types include two or more media types, and the two or more media types are one of audio data, video data, augmented reality (AR) data, extended reality (XR) data, mixed reality (MR) data, or virtual reality (VR) data.

[0141] Clause 7: The method of any one of clauses 1 to 6, wherein the header data includes one or more length values ​​of one or more types of media data in the packet.

[0142] Clause 8: The method of any one of clauses 1 to 7, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which the X field, CC field, and M field are merged into one or more attribute fields.

[0143] Clause 9: The method of any one of clauses 1 to 8, wherein the header data includes a value indicating whether to include a timestamp value in the header data.

[0144] Clause 10: The method of any one of clauses 1 to 9, wherein the header data includes a value indicating the total number of media types contained in the packet.

[0145] Clause 11: The method of any one of clauses 1 to 10, wherein the packet includes a first packet, the media data conforming to media data of a first type, the header data indicating a timestamp of the media data, and the header data indicating that the media data of the first type is to be synchronized with media data of a second type, the method further comprising receiving a second packet including media data of the second type and second header data specifying the timestamp, and presenting the media data comprises presenting the media data of the first type in the first packet simultaneously with presenting the media data of the second type in the second packet in response to determining that the header data of the first packet and the second header data of the second packet each specify the timestamp.

[0146] Clause 12: The method of any one of clauses 1 to 11, wherein receiving the packets includes receiving the packets as part of a Real-time Transport Protocol (RTP) session.

[0147] Clause 13: The method of any one of clauses 1 to 12, wherein receiving packets includes receiving packets as part of a Web Real-Time Communications (WebRTC) session.

[0148] Clause 14: The method of any one of clauses 1 to 11, wherein receiving the packet includes receiving the packet as part of a Stream Control Transmission Protocol (SCTP) stream.

[0149] Clause 15: The method of clause 14, wherein the SCTP stream includes media data of each of one or more media types.

[0150] Clause 16: The method of clause 15, wherein each of the one or more media types of the SCTP streams is to be synchronized, and the method further comprises receiving at least one additional SCTP stream containing media data of an additional media type that is to be unsynchronized.

[0151] Clause 17: The method of any one of clauses 14 to 16, wherein one or more media types are each assigned to a respective data payload chunk of the packet.

[0152] Clause 18: The method of clause 14, wherein the SCTP stream contains one media data of one or more media types, and the method further comprises receiving one or more additional packets of one or more additional SCTP streams, each of the one or more additional SCTP streams containing a different one media data of one or more media types.

[0153] Clause 19: The method of any one of clauses 14 to 18, wherein a portion of the payload chunk of the packet indicates whether a timestamp is provided for the media data of the payload chunk.

[0154] Clause 20: The method of any one of clauses 14 to 19, wherein a portion of the payload chunk of the packet indicates whether synchronization information is provided for the media data of the payload chunk.

[0155] Clause 21: A device for receiving media data, the device comprising one or more means for performing the method of any one of clauses 1 to 20.

[0156] Clause 22: A device according to clause 21, wherein the one or more means include one or more processors implemented in circuitry.

[0157] Clause 23: A device according to clause 21 or 22, wherein the one or more means comprises a memory configured to store the media data.

[0158] Clause 24: A device for receiving media data, the device comprising: means for receiving packets containing media data of a media communication session, the packets including header data indicating one or more media types of the media data and one or more attributes of each of the one or more media types; means for extracting the media data of each of the one or more media types; means for processing the media data of each of the one or more media types according to the corresponding one or more attributes; and means for presenting the media data.

[0159] Clause 25: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to perform the method of any one of clauses 1 to 20.

[0160] Clause 26: A method for transmitting media data, the method comprising: generating packets containing media data of a media communication session, the packets including header data indicating one or more media types of the media data and one or more attributes of each of the one or more media types; and transmitting the packets to a client device.

[0161] Clause 27: A method for processing media data, the method comprising: determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmitting the packets to devices involved in the media communication session.

[0162] Clause 28: The method of clause 27, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, user default pose, user default orientation, media data timestamp, priority data, reliability indication, or whether two or more media types are to be presented synchronously.

[0163] Clause 29: The method of clause 27, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

[0164] Clause 30: The method of clause 27, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value that indicates whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

[0165] Clause 31: The method of clause 27, wherein the header data comprises an extension to a Real-time Transport Protocol (RTP) header, the RTP header comprising a payload type value indicating whether the media data represents a pause, a gesture, or a viewport.

[0166] Clause 32: The method of clause 27, wherein the one or more media types include an indication of an augmented reality (AR) anchor point, user pose data, viewport data, gesture data, body movement data, or facial expression data.

[0167] Clause 33: The method of clause 27, wherein the header data includes one or more length values ​​of one or more types of media data in the packet.

[0168] Clause 34: The method of clause 27, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which the X, CC, and M fields are merged into one or more attribute fields.

[0169] Clause 35: The method of clause 27, wherein the header data includes a value indicating whether the header data includes a timestamp value.

[0170] Clause 36: The method of clause 27, wherein the header data includes a value indicating the total number of media types contained in the packet.

[0171] Clause 37: The method of clause 27, wherein the packet includes a first packet, the media data conforming to media data of a first type, the header data indicating a timestamp of the media data, and the header data indicating that the media data of the first type is to be synchronized with media data of a second type, the method further including constructing a second packet including second header data specifying the media data of the second type and the timestamp, and transmitting the second packet to devices involved in the communication session.

[0172] Clause 38: The method of clause 27, wherein transmitting the packets includes transmitting the packets as part of a Real-time Transport Protocol (RTP) session.

[0173] Clause 39: The method of clause 27, wherein transmitting the packets includes transmitting the packets as part of a Web Real-Time Communications (WebRTC) session.

[0174] Clause 40: The method of clause 27, wherein transmitting the packets includes transmitting the packets as part of a Stream Control Transmission Protocol (SCTP) stream.

[0175] Clause 41: The method of clause 40, wherein the SCTP stream includes media data of each of one or more media types.

[0176] Clause 42: The method of clause 41, wherein each of the one or more media types of the SCTP streams is to be synchronized, and the method further comprises sending at least one additional SCTP stream containing media data of the additional media type(s) is to be unsynchronized.

[0177] Clause 43: The method of clause 40, wherein one or more media types are each assigned to a respective data payload chunk of the packet.

[0178] Clause 44: The method of clause 40, wherein the SCTP stream contains one media data of one or more media types, and the method further comprises sending one or more additional packets of one or more additional SCTP streams, each of the one or more additional SCTP streams containing a different one media data of one or more media types.

[0179] Clause 45: The method of clause 40, wherein a portion of the payload chunk of the packet indicates whether a timestamp is provided for the media data of the payload chunk.

[0180] Clause 46: The method of clause 40, wherein a portion of the payload chunk of the packet indicates whether synchronization information is provided for the media data of the payload chunk.

[0181] Clause 47: A device for processing media data, comprising: a memory configured to store media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: determine one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets including data for a media communication session, the packets including header data indicating the one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.

[0182] Clause 48: The device of clause 47, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, user default pose, user default orientation, timestamp of media data, priority data, reliability indication, or whether two or more media types are to be presented synchronously.

[0183] Clause 49: The device of clause 47, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

[0184] Clause 50: The device of clause 47, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value indicating whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

[0185] Clause 51: The device of clause 47, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, the RTP header including a payload type value indicating whether the media data represents a pose, a gesture, or a viewport.

[0186] Clause 52: The device of clause 47, wherein the one or more media types include an indication of an augmented reality (AR) anchor point, user pose data, viewport data, gesture data, body movement data, or facial expression data.

[0187] Clause 53: The device of clause 47, wherein the header data includes one or more length values ​​of one or more types of media data in the packet.

[0188] Clause 54: The device of clause 47, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which the X field, CC field, and M field are merged into one or more attribute fields.

[0189] Clause 55: A device for processing media data, comprising: means for determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; means for constructing packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and means for transmitting the packets to devices involved in the media communication session.

[0190] Clause 56: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to determine one or more user interactions with the virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.

[0191] Clause 57: A method for processing media data, the method comprising: determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmitting the packets to devices involved in the media communication session.

[0192] Clause 58: The method of clause 57, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, user default pose, user default orientation, media data timestamp, priority data, reliability indication, or whether two or more media types are to be presented synchronously.

[0193] Clause 59: The method of clause 57 or 58, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

[0194] Clause 60: A method according to any one of clauses 57 to 59, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value indicating whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

[0195] Clause 61: A method according to any one of clauses 57 to 60, wherein the header data comprises an extension to a Real-time Transport Protocol (RTP) header, the RTP header comprising a payload type value indicating whether the media data represents a pause, a gesture, or a viewport.

[0196] Clause 62: A method according to any one of clauses 57 to 61, wherein the one or more media types include an indication of an augmented reality (AR) anchor point, user pose data, viewport data, gesture data, body movement data, or facial expression data.

[0197] Clause 63: The method of any one of clauses 57 to 62, wherein the header data includes one or more length values ​​of one or more types of media data in the packet.

[0198] Clause 64: The method of any one of clauses 57 to 63, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which the X field, CC field, and M field are merged into one or more attribute fields.

[0199] Clause 65: The method of any one of clauses 57 to 64, wherein the header data includes a value indicating whether to include a timestamp value in the header data.

[0200] Clause 66: The method of any one of clauses 57 to 65, wherein the header data includes a value indicating the total number of media types contained in the packet.

[0201] Clause 67: A method according to any one of clauses 57 to 66, wherein the packet comprises a first packet, the media data conforming to media data of a first type, the header data indicating a timestamp of the media data, and the header data indicating that the media data of the first type is to be synchronized with media data of a second type, the method further comprising constructing a second packet comprising second header data specifying the media data of the second type and the timestamp, and transmitting the second packet to devices involved in the communication session.

[0202] Clause 68: The method of any one of clauses 57 to 67, wherein transmitting the packets includes transmitting the packets as part of a Real-time Transport Protocol (RTP) session.

[0203] Clause 69: The method of any one of clauses 57 to 68, wherein transmitting the packets includes transmitting the packets as part of a Web Real-Time Communications (WebRTC) session.

[0204] Clause 70: The method of any one of clauses 57 to 69, wherein transmitting the packets includes transmitting the packets as part of a Stream Control Transmission Protocol (SCTP) stream.

[0205] Clause 71: The method of clause 70, wherein the SCTP stream includes media data of each of one or more media types.

[0206] Clause 72: The method of clause 71, wherein each of the one or more media types of the SCTP streams is to be synchronized, and the method further comprises sending at least one additional SCTP stream containing media data of the additional media type(s) is to be unsynchronized.

[0207] Clause 73: The method of clause 71 or 72, wherein one or more media types are each assigned to a respective data payload chunk of the packet.

[0208] Clause 74: A method according to any one of clauses 71 to 73, wherein the SCTP stream contains one media data of one or more media types, and the method further comprises sending one or more additional packets of one or more additional SCTP streams, each of the one or more additional SCTP streams containing a different one media data of one or more media types.

[0209] Clause 75: The method of any one of clauses 71 to 74, wherein a portion of the payload chunk of the packet indicates whether a timestamp is provided for the media data of the payload chunk.

[0210] Clause 76: The method of any one of clauses 71 to 75, wherein a portion of the payload chunk of the packet indicates whether synchronization information is provided for the media data of the payload chunk.

[0211] Clause 77: A device for processing media data, comprising: a memory configured to store media data; and a processing system including one or more processors implemented in circuitry, wherein the processing system is configured to: determine one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets including data for a media communication session, the packets including header data indicating the one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.

[0212] Clause 78: The device of clause 77, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, user default pose, user default orientation, timestamp of media data, priority data, reliability indication, or whether two or more media types are to be presented synchronously.

[0213] Clause 79: A device as described in clause 77 or 78, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

[0214] Clause 80: A device described in any one of clauses 77 to 79, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value indicating whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

[0215] Clause 81: A device described in any one of clauses 77 to 80, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, and the RTP header includes a payload type value indicating whether the media data represents a pose, a gesture, or a viewport.

[0216] Clause 82: A device described in any one of clauses 77 to 81, wherein the one or more media types include an indication of an augmented reality (AR) anchor point, user pose data, viewport data, gesture data, body movement data, or facial expression data.

[0217] Clause 83: The device of any one of clauses 77 to 82, wherein the header data includes one or more length values ​​of one or more types of media data in the packet.

[0218] Clause 84: A device described in any one of clauses 77 to 83, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which the X field, CC field, and M field are merged into one or more attribute fields.

[0219] Clause 85: A device for processing media data, comprising: means for determining one or more user interactions with a virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; means for constructing packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and means for transmitting the packets to devices involved in the media communication session.

[0220] Clause 86: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to determine one or more user interactions with the virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; construct packets containing data for a media communication session, the packets including header data indicating one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; and transmit the packets to devices involved in the media communication session.

[0221] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communications protocol. As such, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0222] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically and discs reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0223] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. The techniques may also be implemented entirely in one or more circuits or logic elements.

[0224] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units have been described in this disclosure to highlight functional aspects of devices configured to implement the disclosed techniques, but they do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0225] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. 1. A method for processing media data, comprising: determining one or more user interactions with the virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating the one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; transmitting the packets to devices participating in the media communication session; A method comprising:

2. 2. The method of claim 1, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, a user's default pose, a default orientation of the user, a timestamp of the media data, priority data, a reliability indication, or whether two or more media types are to be presented synchronously.

3. 2. The method of claim 1, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, the RTP header including a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

4. 2. The method of claim 1, wherein the header data comprises an extension to a Real-time Transport Protocol (RTP) header, the RTP header comprising a payload type value indicating whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

5. 2. The method of claim 1, wherein the header data comprises an extension to a Real-time Transport Protocol (RTP) header, the RTP header comprising a payload type value indicating whether the media data represents a pose, a gesture, or a viewport.

6. The method of claim 1 , wherein the one or more media types include an indication of an augmented reality (AR) anchor point, user pose data, viewport data, gesture data, body movement data, or facial expression data.

7. The method of claim 1 , wherein the header data includes one or more length values ​​of the one or more types of media data in the packet.

8. 2. The method of claim 1, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which the X, CC, and M fields are merged into one or more attribute fields.

9. The method of claim 1 , wherein the header data includes a value indicating whether to include a timestamp value in the header data.

10. The method of claim 1 , wherein the header data includes a value indicating a total number of media types contained in the packet.

11. the packets include a first packet, the media data conforming to a first type of media data, the header data indicating a timestamp of the media data, and the header data indicating that the first type of media data is to be synchronized with a second type of media data, the method comprising: constructing a second packet including second header data specifying the second type of media data and the timestamp; transmitting the second packet to the devices involved in the communication session; Further comprising: The method of claim 1.

12. The method of claim 1 , wherein transmitting the packets comprises transmitting the packets as part of a Real-time Transport Protocol (RTP) session.

13. The method of claim 1 , wherein transmitting the packets comprises transmitting the packets as part of a Web Real-Time Communications (WebRTC) session.

14. 10. The method of claim 1, wherein transmitting the packets comprises transmitting the packets as part of a Stream Control Transmission Protocol (SCTP) stream.

15. The method of claim 14 , wherein the SCTP stream includes media data for each of the one or more media types.

16. 16. The method of claim 15, wherein each of the one or more media types of the SCTP streams is to be synchronized, the method further comprising transmitting at least one additional SCTP stream containing media data of an additional media type that is to be unsynchronized.

17. 15. The method of claim 14, wherein each of the one or more media types is assigned to a respective data payload chunk of the packet.

18. 15. The method of claim 14, wherein the SCTP stream includes media data of one of the one or more media types, the method further comprising transmitting one or more additional packets of one or more additional SCTP streams, each of the one or more additional SCTP streams including media data of a different one of the one or more media types.

19. 15. The method of claim 14, wherein a portion of the payload chunk of the packet indicates whether a timestamp is provided for media data in the payload chunk.

20. 15. The method of claim 14, wherein a portion of the payload chunk of the packet indicates whether synchronization information is provided in the media data of the payload chunk.

21. A device for processing media data, comprising: a memory configured to store media data; a processing system including one or more processors implemented in circuitry; wherein the processing system comprises: determining one or more user interactions with the virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating the one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; transmitting the packets to devices involved in the media communication session; It is configured as follows: device.

22. 22. The device of claim 21, wherein the one or more attributes indicate one or more of camera information, display information, field of view, projection, resolution, refresh rate, a user's default pose, a default orientation of the user, a timestamp of the media data, priority data, a reliability indication, or whether two or more media types are to be presented synchronously.

23. 22. The device of claim 21, wherein the header data includes an extension to a Real-time Transport Protocol (RTP) header, the RTP header including a mark bit (M bit) having a value indicating whether the packet is the last packet of interactive media data in transmission order.

24. 22. The device of claim 21, wherein the header data comprises an extension to a Real-time Transport Protocol (RTP) header, the RTP header comprising a payload type value indicating whether the media data corresponds to a scene description, a spatial description, or a three-dimensional (3D) visual model.

25. 22. The device of claim 21, wherein the header data comprises an extension to a Real-time Transport Protocol (RTP) header, the RTP header comprising a payload type value indicating whether the media data represents a pose, a gesture, or a viewport.

26. 22. The device of claim 21, wherein the one or more media types include an indication of an augmented reality (AR) anchor point, user pose data, viewport data, gesture data, body movement data, or facial expression data.

27. 22. The device of claim 21, wherein the header data includes one or more length values ​​of the one or more types of media data in the packet.

28. 22. The device of claim 21, wherein the header data conforms to a profile of the Real-time Transport Protocol (RTP) in which X, CC, and M fields are merged into one or more attribute fields.

29. A device for processing media data, comprising: means for determining one or more user interactions with the virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; means for constructing packets containing data for a media communication session, the packets including header data indicating the one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; means for transmitting said packets to devices participating in said media communication session; A device comprising:

30. A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to: determining one or more user interactions with the virtual scene, each of the one or more user interactions corresponding to a particular media type of the one or more media types; constructing packets containing data for a media communication session, the packets including header data indicating the one or more media types of the one or more user interactions and one or more attributes of each of the one or more media types; causing devices participating in the media communication session to transmit the packets; A computer-readable storage medium.

Citation Information

Cited By

  • Game live streaming system, game live streaming method, and game live streaming program

    JP7896949B1