Split rendering configuration for multimedia immersion and dialogue data in wireless communication systems
The proposed device and method for wireless communication systems address the challenges of transporting interactive and immersive multimedia data by employing segmented rendering configurations, ensuring efficient and synchronized delivery of multimedia data in XR applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LENOVO (SINGAPORE) PTE LTD
- Filing Date
- 2023-05-16
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies lack efficient and flexible solutions for transporting interactive and immersive multimedia data in wireless communication systems, particularly for XR applications, due to the rapid evolution of data formats, real-time sensitivity, and lack of standardized encoding and transport mechanisms.
A device and method for wireless communication systems that utilize a media configuration for segmented rendering of multimedia data, including parameters for immersive and interaction data, enabling effective transport and synchronization through a multimedia segmented rendering content delivery session.
Facilitates the real-time transport and synchronization of diverse multimedia data types with varying formats and sizes, supporting split rendering configurations and interactive metadata transport, enhancing the immersive experience in XR applications.
Smart Images

Figure 2026510623000001_ABST
Abstract
Description
Technical Field
[0001] The subject matter disclosed in this specification generally relates to the field of implementing split rendering configurations for multimedia immersion and interactive data in wireless communication systems. This specification defines apparatuses and methods for wireless communication in wireless communication systems.
Background Art
[0002] Interactive and immersive multimedia communication implies various information flows that carry potentially time-sensitive inputs from one terminal to be transported to a remote terminal over a network. Applications that rely on such multimedia communication modes are becoming increasingly common, along with massive online games, cloud gaming, and extended reality (XR) deployments across the market. Multimedia information flows often supersede traditional video and audio flows and include additional formats in the following categories, namely device capabilities, media descriptions, and spatial interaction information, respectively. These are communicated over heterogeneous networks between the graphics rendering engine and the user device. Thus, these media and data types underlie the successful implementation of truly immersive and interactive applications that process user input information and return responses partially based on stimulating user input under a set of latency constraints.
[0003] In accordance with 3GPP (Registered Trademark) Technical Report TR26.928 (v17.0.0 - April 2022), XR is used as an inclusive term for various types of sense of reality. These types of sense of reality include virtual reality (VR), augmented reality (AR), and mixed reality (MR).
[0004] VR is a rendered version of a delivered visual and audio scene. The rendering is designed to mimic real-world visual and audio perceptual stimuli as naturally as possible as the viewer or user moves within the limits defined by the application. Virtual reality typically requires, though not always, the user to wear a head-mounted display (HMD) to completely replace the user's field of view with simulated visual components, and headphones to provide the user with accompanying audio. Some forms of head tracking and motion tracking in VR are also typically necessary to ensure, from the user's perspective, that objects and sound sources remain in sync with the user's movements, allowing the simulated visual and audio components to be updated. Some implementations may, but are not strictly required, provide additional means for interacting with the virtual reality simulation.
[0005] Augmented reality (AR) is when additional information or artificially generated objects, or content overlaid on the user's current environment, is provided to the user. Such additional information or content is typically visual and / or acoustic, and the user's observation of the user's current environment may be direct, without intermediate sensing, processing, and rendering, or indirect, where the user's perception of the user's environment may be relayed, augmented, or processed via sensors.
[0006] Mixed Reality (MR) is an advanced form of augmented reality (AR) in which several virtual elements are inserted into a physical scene with the intention of creating the illusion that these elements are part of the real-world scene.
[0007] XR refers to all real and virtual environments and human-machine interactions generated by computer technology and wearables. XR includes representative forms such as AR, MR, and VR, as well as areas interpolated between them. The level of virtuality ranges from partially perceptual input to fully immersive VR. A major aspect of XR is the extension of human experience, particularly related to the perception of presence (represented by VR) and the acquisition of recognition (represented by AR).
[0008] Interactive and spatial computing related to XR application activities are central to the success of immersive XR experiences. This applies equally to other mainstream interactive applications such as cloud gaming (CG). Interactive data and associated spatial computing determine the XR rendering engine response or CG gaming engine response to user physical input, thus contributing to the cyber-physical illusion of immersion between the physical and virtual worlds.
[0009] The data that such applications carry and utilize to generate cyber-physical immersive illusions is categorized into several classes, as described in 3GPP Technical Document S4-221557 (November 2022) or, alternatively, Technical Report TR26.926 (v1.1.0 - February 2022). This includes device capability classes, media description classes, and interaction and immersion metadata classes.
[0010] In the device capability class, the format associated with this data class describes the physical and hardware capabilities of the end-user equipment (UE) and / or glass device. Some examples in this sense are camera subsystem capabilities and camera configuration (e.g., focal length, available zoom and depth calibration information, main camera attitude reference, etc.) and projection format (e.g., cubemap, equirectangular, fisheye, stereo, etc.). Device capability data is usually static and available before session establishment, and therefore its transfer and transport over the network is not a major issue, as it can be embedded within common session configuration procedures and protocols such as the Session Initiation Protocol (SIP) and / or Session Description Protocol (SDP). Device capability data is not thus real-time sensitive and does not have real-time transport requirements.
[0011] In the media description class, data describes the spatial and / or object content of a landscape. For example, this data could be a scene description used to detail the 3D composition of space, anchoring 2D and 3D objects in the scene (for example, typically as a tree or graph structure, usually in glTF2.0 or JSON syntax). Another possible representation is a spatial description used for spatial computing and mapping of the real world to its virtual relatives or vice versa. In some other examples, this data type may store 3D model descriptors of objects and their attributes (i.e., sets of intersections, edges, and faces), formatted as, for example, a mesh, or point cloud data formatted under the Polygon (PLY) syntax, to be consumed by a visual presentation device, i.e., a UE. Other data types may represent dynamic world graph representations, where selected trackables (e.g., geo-trackables such as geo-cached AR / QR codes, physical objects located at specified world locations, dynamic physical objects such as buses and subways) must dynamically enter and leave the world scene viewpoint and be communicated to the AR runtime in real time. The media description class of the data may be large (i.e., often exceeding 10 megabytes) and may be updated infrequently (within regularity of tens of seconds) under various event triggers (e.g., user viewport changes, new objects entering the scene, old objects leaving the scene, scene changes and / or updates, etc.). The media description data may be real-time sensitive as it is involved in completing the display of virtual renderings to presenting devices such as UEs, and as a result may benefit from real-time transport over a network.
[0012] In the Interaction and Immersion metadata class, this data type includes: User Viewport Description (i.e., encoding of motion orientation, altitude, tilt, and associated range, describing the projection of the user view onto the target display), User Field of View (FoV) (i.e., the range of the visible world from the observer's point of view, typically described in the angular domain, spanning vertical and horizontal planes, e.g., radians / degrees), User Pose / Orientation Tracking Data (i.e., 3D vectors time-stamped in microseconds / nanoseconds for position and quaternion representations for orientation, describing up to six DoFs), User Gesture Tracking Data (i.e., an array of tracked hands, each consisting of an array of wrist joint locations relative to base space), User Body Tracking Data (e.g., BioVision hierarchical BVH encoding of body and segmental movements), User Facial Expressions / Eye Movements (e.g., an array of keypoint / feature locations, or their encoding into a predetermined facial expression class). It stores user-space interaction information, such as tracking data, user-actionable inputs (e.g., OpenXR actions that capture user input to controllers or HW command units supported by AR / VR devices), segmented rendering pose and spatial information (e.g., pose information that stores pose information used by the segmented rendering server to pre-render XR scenes or, alternatively, scene projections as video frames), and application and AR anchor data and descriptions (i.e., metadata that determines the position of an object or point in user space as an anchor for positioning virtual 2D / 3D objects, such as text rendering, 2D / 3D photo / video content, etc.). [Prior art documents] [Patent Documents]
[0013] [Patent Document 1] U.S. Patent Application No. 63 / 420,885 [Patent Document 2] U.S. Patent Application No. 63 / 478,932 [Non-patent literature]
[0014] [Non-Patent Document 1] "RTP: A Transport Protocol for Real-Time Applications", IETF standard RFC3550 [Non-Patent Document 2] "The Secure Real-time Transport Protocol (SRTP)", IETF standard RFC3711 [Non-Patent Document 3] "WebRTC: Real-Time Communication in Browsers," W3C Standards Recommendation, March 6, 2023. [Non-Patent Document 4] "Frame Marking RTP Header Extension," November 2021, IETF Working Draft [Non-Patent Document 5] "Encryption of Header Extensions in the Secure Real-time Transport Protocol (SRTP)", IETF standard RFC6904 [Non-Patent Document 6] "A General Mechanism for RTP Header Extensions", IETF standard RFC:8285 [Non-Patent Document 7] Alliance for Open Media, "AV1 Bitstream & Decoding Process Specification," 182, de Rivaz, p. and Haughton (2018) [Non-Patent Document 8] "Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services-coding of moving video: versatile supplemental enhancement information messages for coded video bitstreams", ITU-T Series H Specification V8 (08 / 2020) [Non-Patent Document 9] "Terminal provider codes notification form: available information regarding the identification of national authorities for the assignment of ITU-T recommendation T.35 terminal provider codes", ITU-T recommendation T.35 [Non-Patent Document 10] "Split Rendering Media Service Enabler", 3GPP technical specification TS26.565 v0.3.0 [Non-Patent Document 11] "5G Real-time Media Communication Architecture", 3GPP Technical Specification 26.506 v1.1.0 [Non-Patent Document 12] "An offer / answer model with the Session Description Protocol (SDP)", RFC Standard 3264 [Non-Patent Document 13] "Real-time metadata transport over RTP", S4aR220052 [Non-Patent Document 14] "Real-time metadata transport over data channel", S4aR220053 [Non-Patent Document 15] "Real-time metadata transport over data channel", S4-221557
Non-Patent Document 16
Non-Patent Document 17
Summary of the Invention
Problems to be Solved by the Invention
[0015] Generally, interactive and immersive class data has several characteristics. These characteristics include a small data footprint, typically ranging from 32 bytes to approximately several hundred bytes per message, without an established codec for compression of the data source; a high sampling rate that varies, for example, between 60 Hz and 250 Hz for video FPS frequency, and in some cases where sample aggregation is not performed, raw sample reports can be sent even at a sampling frequency of 1000 Hz; the data can trigger responses with low latency requirements (e.g., end-to-end, up to 50 milliseconds from interaction to response as perceived by the user); the data can be synchronized with other media streams (e.g., video or audio media streams); the data can be synchronized with other interactive data (e.g., pose information may be synchronized with user actions or alternatively object actions); reliability is optional as determined by individual application requirements (e.g., in a split rendering scenario, the server may predict future pose estimations based on available pose information, and thus a high reliability with an error rate less than 1e-3 is not required); data encoding follows typically proprietary / non-standardized or rapidly evolving application-dependent and interaction-dependent formats (e.g., formats and transport containers for extremely large amounts of metadata are not yet fully defined and / or specified); it includes carrying privacy-sensitive interactive events and data such as user pose / gaze, user input / action, and / or user hand gestures or body orientation.
[0016] Therefore, the interactive and immersive metadata class is real-time sensitive and requires real-time transport over the network equivalent to existing solutions for established media flows related to video or alternatively audio codec.
[0017] Furthermore, in XR multimedia information flows, the format and syntax of such information flows are often application-dependent, platform-dependent, and / or hardware-dependent, and mainstream encoding, syntax, and semantics are not well established, in stark contrast to well-established media formats and codecs (e.g., audio or video codecs). Moreover, because such information flows and associated data formats are evolving rapidly, mainstream transport-specific solutions have not yet been established. The latter fact necessitates rapid adaptation to new versions at a faster pace than typical conventional media codec development cycles. This provides application developers with the necessary modern tools for rapidly emerging groundbreaking interactive applications in their search for network-based transport solutions for such information flows in a real-time, flexible, and encoding / syntax-agnostic manner.
[0018] Currently, media description and interaction data that benefit from real-time transmission and synchronization may be transmitted in several implementations based on at least three different options based on existing technologies. These technologies are the WebRTC data channel based on the Stream Control Transmission Protocol (SCTP), RTP header extensions that embed metadata information within the bandwidth of the RTP transport (e.g., U.S. Patent Application No. 63 / 420,885), and novel RTP payload formats generally dedicated to the transport of real-time metadata (e.g., U.S. Patent Application No. 63 / 478,932).
[0019] A possible solution for the first technology (described in 3GPP Tdoc S4-221557) involves using the WebRTC SCTP data channel to carry interaction metadata. A comprehensive data channel payload format for time-sensitive metadata, including timestamps, is added to the chunk user data section of the SCTP data channel. This is a flexible option for metadata transport, as it allows for the transport of metadata not directly related to the media and enables distinctions regarding reliability, priority, and ordering requirements by setting up data channels with different characteristics. It relies on protocols established by the IETF and technologies available in the market today. On the downside, it lacks essential timing, synchronization, and jitter support, and in the case of SCTP, it lacks the FEC mechanism.
[0020] In the case of the second technique (described in 3GPP Tdoc S4-221555 or alternatively in U.S. Patent Application No. 63 / 420,885), a possible solution involves an RTP header extension designed to carry limited-size interaction metadata while its associated media content is carried within the RTP payload. To enable scalability and flexibility, support for a single metadata type or multiple metadata types may be carried within the proposed header extension. This technique has the advantage that the transported metadata is time-synchronized with the media data. Furthermore, it includes all the robustness and timing mechanisms provided by RTP (e.g., synchronization, jitter, congestion control support, FEC mechanism, etc.). However, this only makes sense when a media stream exists, which is often the case for AR or VR-specific use cases supported by segmented rendering. If a media stream does not exist, however, the transmission of an RTP packet with an empty / dummy payload will be required. Depending on the metadata type, the potentially large size of the RTP header is another issue. RTP header extensions can also be silently ignored by the receiver if subsequent devices are unable to process them.
[0021] In the case of the third technology, it is possible to use a separate RTP stream in which the dialogue metadata is carried within the RTP payload (as proposed in U.S. Patent Application No. 63 / 478,932). In RTP, media coding details such as signal sampling rate, frame size, and timing are specified in the RTP payload format. Therefore, sending dialogue metadata in a separate RTP stream is based on defining a new RTP payload format for dialogue metadata, which thereby formalizes a comprehensive RTP metadata payload dedicated to the transport of various dialogue and immersion metadata. The advantage of this approach is that it allows the use of all RTP mechanisms (timing, synchronization, jitter management support, and FEC robustness, etc.) while providing a comprehensive format that can cover all types of dialogue metadata. This requires that the carryover in the IETF become a general-purpose transport standard. However, defining a new payload format typically takes at least two years in the IETF, meaning that for 3GPP or similar, the developed format may become useful for networked systems at the earliest at the end of the 3GPP Release 19 cycle, or alternatively in the midterm.
[0022] Therefore, solutions for transporting multimedia interactive and immersive data that accommodate diverse multimedia data types (e.g., media descriptions, interactive and immersive metadata classes, and their subcategories), including various syntaxes and formats, various data sizes (e.g., from hundreds of bytes to tens of megabytes), various time-dependent data generation (e.g., from periodically generated with strict timing for event-based data), and real-time synchronization constraints, are of interest.
[0023] Furthermore, split rendering requires, in addition, effective signaling capabilities for media configuration of various information flows that act as metadata for split rendering and media display. Therefore, solutions for media configuration signaling as an enabler for split rendering architectures, as well as for interactive and immersive metadata transport, are proposed herein.
[0024] Procedures for split-rendering configurations for multimedia immersion and interaction data in wireless communication systems are disclosed herein. These procedures may be carried out by apparatus and methods for wireless communication in wireless communication systems. [Means for solving the problem]
[0025] A device for wireless communication is provided, comprising a processor and memory coupled to the processor, wherein the processor is configured to determine a media configuration for segmented rendering of a video media flow, the media configuration comprising one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; to signal the media configuration to a second device; to establish a multimedia segmented rendering content delivery session comprising the video media flow with the second device, at least in part on the media configuration; and to use one or more data units of multimedia immersion and interaction data to segment render the video media flow.
[0026] A device for wireless communication is further provided, comprising a processor and memory coupled to the processor, wherein the processor is configured to receive a media configuration from a first device for segmented rendering of a video media flow, the media configuration comprising one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; configure the device to receive the video media flow using the media configuration; decode the encoded video stream of the video media flow, comprising decoding and extracting the non-video encoded metadata; and consume one or more data units of multimedia immersion and interaction data from the non-video encoded metadata.
[0027] A method for wireless communication is further provided, the method comprising determining a media configuration for segmented rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; signaling the media configuration to a second device; establishing a multimedia segmented rendering content delivery session comprising the video media flow with the second device, at least in part on the media configuration; and using one or more data units of multimedia immersion and interaction data to segment render the video media flow. A method for wireless communication is further provided, the method comprising: receiving a media configuration from a first device for segmented rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; configuring the device with the media configuration to receive the video media flow; decoding the encoded video stream of the video media flow, comprising extracting the non-video encoded metadata; and consuming one or more data units of multimedia immersion and interaction data from the non-video encoded metadata.
[0028] To illustrate the manner in which the advantages and features of this disclosure may be obtained, the description of this disclosure is expressed by reference to several apparatuses and methods shown in the accompanying drawings. Each of these drawings illustrates only a few aspects of this disclosure and should not be considered to limit its scope. The drawings may be simplified for clarity and may not necessarily be drawn to a constant scale.
[0029] A method and apparatus for a segmented rendering configuration for multimedia immersion and interaction data in a wireless communication system is described below, merely as an example, with reference to the attached drawings. [Brief explanation of the drawing]
[0030] [Figure 1] This figure shows one embodiment of a wireless communication system. [Figure 2] This is a diagram showing one embodiment of a user device. [Figure 3] This figure shows one embodiment of a network node. [Figure 4] This diagram shows the RTP and RTCP protocol stacks over an IP network. [Figure 5] This diagram shows the WebRTC (SRTP) protocol stack over an IP network. [Figure 6] This figure shows the RTP packet format and header information. [Figure 7] This figure shows the SRTP packet format and header information. [Figure 8] This figure shows the RTP / SRTP header extension format and syntax. [Figure 9] This is a simplified block diagram of a comprehensive video codec that performs spatial and temporal compression of video sources. [Figure 10] This figure shows the video-encoded base stream and the corresponding multiple NAL units. [Figure 11] This figure shows one embodiment of a wireless communication method in a wireless communication system. [Figure 12] This figure shows an alternative embodiment of a wireless communication method in a wireless communication system. [Figure 13a] This figure shows one embodiment of a split-rendering architecture for interactive and immersive multimedia applications. [Figure 13b]This figure shows one embodiment of a high-level partitioned rendering flow. [Figure 14] This figure shows the representation of multimedia interaction and immersive user data as metadata within the video-encoded base stream for the MPEG H-26x family of video codecs. [Figure 15] This figure shows the representation of multimedia interaction and immersion user data as metadata within the video encoding base stream for the AV1 video codec. [Modes for carrying out the invention]
[0031] As will be understood by those skilled in the art, aspects of this disclosure may be embodied as systems, apparatus, methods, or program products. Accordingly, the configurations described herein may be implemented entirely in hardware form, entirely in software form (including firmware, resident software, microcode, etc.), or in a combination of software and hardware forms.
[0032] For example, the disclosed methods and apparatus may be implemented as custom very large-scale integrated ("VLSI") circuits or as hardware circuits comprising off-the-shelf semiconductors such as gate arrays, logic chips, transistors, or other discrete components. The disclosed methods and apparatus may also be implemented within programmable hardware devices such as field-programmable gate arrays, programmable array logic, or programmable logic devices. As another example, the disclosed methods and apparatus may include one or more physical or logical blocks of executable code, which may be organized as objects, procedures, or functions, for example.
[0033] Furthermore, the methods and apparatus may take the form of a program product embodied in one or more computer-readable storage devices that store machine-readable code, computer-readable code, and / or program code, which are referred to below as code. The storage devices may be tangible, non-temporary, and / or non-transmitting. The storage devices may not embody signals. In some configurations, the storage devices employ signals only for accessing the code.
[0034] Any combination of one or more computer-readable media may be used. The computer-readable media may be computer-readable storage media. The computer-readable storage media may be a storage device that stores code. The storage device may be, for example, but not limited to, a system, apparatus, or device of electronic, magnetic, optical, electromagnetic, infrared, holographic, micromechanical, or semiconductor, or any suitable combination thereof.
[0035] More specific examples (a non-exclusive list) of storage devices include, namely, electrical connections having one or more wires, portable computer diskettes, hard disks, random access memory ("RAM"), read-only memory ("ROM"), erasable programmable read-only memory ("EPROM") or flash memory, portable compact disk read-only memory ("CD-ROM"), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the context of this specification, a computer-readable storage medium may be any tangible medium on which a program for use by or related to an instruction execution system, apparatus, or device may be stored or stored.
[0036] Any reference throughout this Specification to an example of a particular method or apparatus, or similar language, means that the particular feature, structure, or characteristic described in relation to that example is included in at least one implementation of the methods and apparatus described herein. Thus, any reference to an example of a particular method or apparatus, or similar language, means "one or more, but not all, examples," unless otherwise specified, and may or may not refer to all of the same example. The terms "including," "comprising," "having," and their variations mean "including, but not limited to," unless otherwise specified. Listings of items do not imply that any or all of the items are mutually exclusive unless otherwise specified. The terms "a," "an," and "the" also mean "one or more," unless otherwise specified.
[0037] As used herein, a list with the conjunction "and / or" includes any single item in the list, or any combination of items in the list. For example, the list A, B, and / or C includes A only, B only, C only, a combination of A and B, a combination of B and C, a combination of A and C, or a combination of A, B, and C. As used herein, a list using the term "one or more of" includes any single item in the list, or any combination of items in the list. For example, one or more of A, B, and C includes A only, B only, C only, a combination of A and B, a combination of B and C, a combination of A and C, or a combination of A, B, and C. As used herein, a list using the term "one of" includes one unique single item from any single item in the list. For example, "one of A, B, and C" includes A only, B only, or C only, and excludes the combination of A, B, and C. When used herein, “members selected from the group consisting of A, B, and C” includes one unique individual of A, B, or C, but excludes combinations of A, B, and C. When used herein, “members selected from the group consisting of A, B, and C, and combinations thereof” includes A only, B only, C only, a combination of A and B, a combination of B and C, a combination of A and C, or a combination of A, B, and C.
[0038] Furthermore, the features, structures, or properties described herein may be combined in any preferred manner. The following description provides numerous specific details, such as examples of programming, software modules, user selection, network transactions, database queries, database structures, hardware modules, hardware circuits, and hardware chips, in order to provide a full understanding of the disclosure. However, those skilled in the art will recognize that the disclosed methods and apparatus may be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not illustrated or described in detail to avoid obscuring aspects of the disclosure.
[0039] The methods and apparatuses disclosed are described below with reference to schematic flowcharts and / or schematic block diagrams of the methods, apparatuses, systems, and program products. It will be understood that each block in the schematic flowcharts and / or schematic block diagrams, as well as combinations of blocks in the schematic flowcharts and / or schematic block diagrams, can be implemented by code. This code may be provided to a general-purpose computer, a dedicated computer, or a processor of another programmable data processing device to generate a machine such that instructions executed via the processor of the computer or other programmable data processing device create means for performing the functions / operations specified in the schematic flowcharts and / or schematic block diagrams.
[0040] The code may also be stored in a storage device that can target a computer, other programmable data processing device, or other device to function in a particular way, such as to produce a product containing instructions that perform functions / operations specified in a schematic flowchart and / or schematic block diagram.
[0041] In order to generate a process to be performed by a computer, such that the code to be executed on a computer or other programmable device provides a process for performing a function / action specified in a schematic flowchart and / or schematic block diagram, the code may also be loaded onto a computer, another programmable device, or another device to cause the computer, another programmable device, or another device to perform a series of operational steps.
[0042] The schematic flowcharts and / or schematic block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, systems, methods, and program products. In this regard, each block in the schematic flowcharts and / or schematic block diagrams may represent a module, segment, or portion of code containing one or more executable instructions of code for performing a specified logical function.
[0043] It should also be noted that in some alternative implementations, the functions mentioned within a block may be performed in a different order than those shown in the diagram. For example, two blocks shown consecutively may actually be executed substantially in parallel, or blocks may sometimes be executed in reverse order depending on the functions involved. Other steps and methods may be devised that are equivalent in function, logic, or effect to one or more blocks or parts of the illustrative diagram.
[0044] The descriptions of elements in each drawing may refer to elements in preceding drawings. Similar numbers refer to the same elements across all drawings.
[0045] Figure 1 shows one embodiment of a wireless communication system 100 for a split-rendering configuration for multimedia immersion and interaction data in a wireless communication system. In one embodiment, the wireless communication system 100 includes a remote unit 102 and a network unit 104. A specific number of remote units 102 and network units 104 are shown in Figure 1, but those skilled in the art will recognize that any number of remote units 102 and network units 104 may be included in the wireless communication system 100. The wireless communication system may comprise a wireless communication network and at least one wireless communication device. The wireless communication device is typically a 3GPP user equipment (UE). The wireless communication network may comprise at least one network node. The network node may be a network unit.
[0046] In one embodiment, the remote unit 102 may include computing devices such as desktop computers, laptop computers, personal digital assistants ("PDAs"), tablet computers, smartphones, smart televisions (e.g., Internet-connected televisions), set-top boxes, game consoles, security systems (including security cameras), in-vehicle computers, network devices (e.g., routers, switches, modems), aerial vehicles, and drones. In some embodiments, the remote unit 102 includes wearable devices such as smartwatches, fitness bands, and optical head-mounted displays. Furthermore, the remote unit 102 may be referred to as a subscriber unit, mobile, mobile station, user, terminal, mobile terminal, fixed terminal, subscriber station, UE, user terminal, device, or by other terms used in the art. The remote unit 102 may communicate directly with one or more of the network units 104 via UL communication signals. In some embodiments, the remote unit 102 may communicate directly with other remote units 102 via side-link communication.
[0047] The network unit 104 may be distributed across geographical areas. In some embodiments, the network unit 104 may include access points, access terminals, bases, base stations, node B, eNB, gNB, home node B, relay nodes, devices, core network, airborne servers, radio access nodes, APs, NRs, network entities, access and mobility management functions ("AMF"), integrated data management functions ("UDM"), integrated data repository ("UDR"), UDM / UDR, policy control functions ("PCF"), radio access network ("RAN"), network slice selection functions ("NSSF"), operation, administration, and management ("OAM"), session management functions ("SMF"), user plane functions ("UPF"), and applications. The network unit 104 is generally part of a radio access network, which includes one or more controllers commutably coupled to one or more corresponding network units 104. The radio access network is generally commutably coupled to one or more core networks, which may be coupled to other networks such as the Internet and public switched telephone networks. These and other elements of the radio access and core networks are not illustrated but are generally well known to those skilled in the art.
[0048] In one implementation, the wireless communication system 100 conforms to the New Radio (NR) protocol standardized by 3GPP, with the network unit 104 transmitting on the downlink (DL) using orthogonal frequency division multiplexing ("OFDM") modulation, and the remote unit 102 transmitting on the uplink (UL) using single-carrier frequency division multiple access ("SC-FDMA") or OFDM. However, more generally, the wireless communication system 100 may implement several other open or proprietary communication protocols, such as WiMAX, IEEE 802.11 variants, GSM, GPRS, UMTS, LTE variants, CDMA2000, Bluetooth®, ZigBee, Sigfox, and LoraWAN. This disclosure is not intended to limit to any particular wireless communication system architecture or protocol implementation.
[0049] The network unit 104 may serve several remote units 102 within a serving area, for example, a cell or cell sector, via a wireless communication link. The network unit 104 transmits DL communication signals to serve the remote units 102 in the time domain, frequency domain, and / or spatial domain.
[0050] Figure 2 shows a user device 200 that may be used to carry out the methods described herein. The user device 200 is used to carry out one or more of the solutions described herein. The user device 200 is one or more of the user devices described in the embodiments herein. In particular, the user device 200 may comprise, for example, an application client 1311, an MSH 1312, a split rendering client 1313, and a UE 1310 in Figure 13a. The user device 200 includes a processor 205, memory 210, an input device 215, an output device 220, and a transceiver 225.
[0051] The input device 215 and output device 220 may be combined into a single device such as a touchscreen. In some implementations, the user equipment 200 does not include any input device 215 and / or output device 220. The user equipment 200 may include one or more of the processor 205, memory 210, and transceiver 225, and may not include the input device 215 and / or output device 220.
[0052] As shown in the figure, the transceiver 225 includes at least one transmitter 230 and at least one receiver 235. The transceiver 225 may communicate with one or more cells (or wireless coverage areas) supported by one or more base units. The transceiver 225 may be capable of operating on unlicensed spectrum. Furthermore, the transceiver 225 may include multiple UE panels supporting one or more beams. In addition, the transceiver 225 may support at least one network interface 240 and / or application interface 245. The application interface 245 may support one or more APIs. The network interface 240 may support 3GPP reference points such as Uu, N1, PC5, etc. Other network interfaces 240 may be supported as will be understood by those skilled in the art.
[0053] The processor 205 may include any known controller capable of executing computer-readable instructions and / or logical operations. For example, the processor 205 may be a microcontroller, microprocessor, central processing unit ("CPU"), graphics processing unit ("GPU"), auxiliary processing unit, field-programmable gate array ("FPGA"), or similar programmable controller. The processor 205 may execute instructions stored in memory 210 to perform the methods and routines described herein. The processor 205 is communicatively coupled to memory 210, input device 215, output device 220, and transceiver 225.
[0054] The processor 205 may control the user device 200 to perform the user device behavior described herein. The processor 205 may include an application processor (also called the "main processor") that manages application domain and operating system ("OS") functions, and a baseband processor (also called the "baseband radio processor") that manages radio functions.
[0055] Memory 210 may be a computer-readable storage medium. Memory 210 may include volatile computer storage media. For example, memory 210 may include RAM, including dynamic RAM ("DRAM"), synchronous dynamic RAM ("SDRAM"), and / or static RAM ("SRAM"). Memory 210 may include non-volatile computer storage media. For example, memory 210 may include a hard disk drive, flash memory, or any other suitable non-volatile computer storage device. Memory 210 may include both volatile and non-volatile computer storage media.
[0056] Memory 210 may store data relating to the implementation of traffic category fields as described herein. Memory 210 may also store program code and related data, such as an operating system or other controller algorithms running on device 200.
[0057] The input device 215 may include any known computer input device, including touch panels, buttons, keyboards, styluses, microphones, etc. The input device 215 may be integrated with the output device 220, for example, as a touchscreen or similar touch-sensitive display. The input device 215 may include a touchscreen on which text can be entered using a virtual keyboard displayed on the touchscreen and / or by handwriting on the touchscreen. The input device 215 may include two or more different devices, such as a keyboard and a touchscreen.
[0058] The output device 220 may be designed to output visual signals, acoustic signals, and / or tactile signals. The output device 220 may include an electronically controllable display or display device capable of outputting visual data to the user. For example, the output device 220 may include, but is not limited to, a liquid crystal display ("LCD"), a light-emitting diode ("LED") display, an organic LED ("OLED") display, a projector, or a similar display device capable of outputting images, text, etc., to the user. As another non-limiting example, the output device 220 may include a wearable display, such as a smartwatch, smart glasses, or a head-up display, that is separate from the rest of the user equipment device 200 but communicatively coupled to it. Furthermore, the output device 220 may be a component of a smartphone, personal digital assistant, television, table computer, notebook (laptop) computer, personal computer, vehicle dashboard, etc.
[0059] The output device 220 may include one or more speakers for generating sound. For example, the output device 220 may generate an audible alarm or notification (e.g., a beep or chime). The output device 220 may include one or more haptic devices for generating vibration, motion, or other tactile feedback. All or part of the output device 220 may be integrated with the input device 215. For example, the input device 215 and the output device 220 may form a touchscreen or similar touch-sensitive display. The output device 220 may be located near the input device 215.
[0060] The transceiver 225 communicates with one or more network functions of a mobile communication network via one or more access networks. The transceiver 225 operates under the control of the processor 205 to transmit messages, data, and other signals, and to receive messages, data, and other signals. For example, the processor 205 may selectively activate the transceiver 225 (or a portion thereof) at certain times to send and receive messages.
[0061] The transceiver 225 includes at least one transmitter 230 and at least one receiver 235. One or more transmitters 230 may be used to provide uplink communication signals to a base unit of a wireless communication network. Similarly, one or more receivers 235 may be used to receive downlink communication signals from the base unit. Although only one transmitter 230 and one receiver 235 are illustrated, the user equipment 200 may have any preferred number of transmitters 230 and receivers 235. Furthermore, the transmitters 230 and receivers 235 may be any preferred type of transmitter and receiver. The transceiver 225 may include a first transmitter / receiver pair used to communicate with a mobile communication network over a licensed radio spectrum, and a second transmitter / receiver pair used to communicate with a mobile communication network over an unlicensed radio spectrum.
[0062] A first transmitter / receiver pair may be used to communicate with a mobile communications network over licensed radio spectrum, and a second transmitter / receiver pair, used to communicate with a mobile communications network over unlicensed radio spectrum, may be combined into a single transceiver unit, e.g., a single chip, that performs the functions for use with both licensed and unlicensed radio spectrum. The first and second transmitter / receiver pairs may share one or more hardware components. For example, several transceivers 225, transmitters 230, and receivers 235 may be implemented as physically separate components that access shared hardware and / or software resources, such as a network interface 240.
[0063] One or more transmitters 230 and / or one or more receivers 235 may be implemented and / or integrated into a single hardware component, such as a multi-transceiver chip, a system-on-a-chip, an application-specific integrated circuit ("ASIC"), or other types of hardware components. One or more transmitters 230 and / or one or more receivers 235 may be implemented and / or integrated into a multi-chip module. Other components, such as a network interface 240 or other hardware components / circuits, may be integrated into a single chip along with any number of transmitters 230 and / or receivers 235. Transmitters 230 and receivers 235 may be logically configured as a transceiver 225 using another common control signal, or as modular transmitters 230 and receivers 235 implemented in the same hardware chip or multi-chip module.
[0064] Figure 3 shows further details of a network node 300 that may be used to implement the method described herein. The network node 300 may be one implementation of an entity in a wireless communication network, for example, in one or more of the wireless communication networks described herein. The network node 300 may comprise network nodes such as the edge network 1320 in Figure 13a (e.g., SRAF 1321 and configuration function 1321a and provisioning function 1321b, RTC AS 1322 and signaling servers 1322a and SRS 1322b), or it may comprise a node in the data network 1330 such as ASP 1331. The network node 300 includes a processor 305, memory 310, input device 315, output device 320, and transceiver 325.
[0065] The input device 315 and output device 320 may be combined into a single device such as a touchscreen. In some implementations, the network node 300 does not include any input device 315 and / or output device 320. The network node 300 may include one or more of the processor 305, memory 310, and transceiver 325, and may not include the input device 315 and / or output device 320.
[0066] As shown in the figure, the transceiver 325 includes at least one transmitter 330 and at least one receiver 335, where the transceiver 325 communicates with one or more remote units 200. In addition, the transceiver 325 may support at least one network interface 340 and / or application interface 345. The application interface 345 may support one or more APIs. The network interface 340 may support 3GPP reference points such as Uu, N1, N2, and N3. Other network interfaces 340 may be supported as will be understood by those skilled in the art.
[0067] The processor 305 may include any known controller capable of executing computer-readable instructions and / or logical operations. For example, the processor 305 may be a microcontroller, microprocessor, CPU, GPU, auxiliary processing unit, FPGA, or similar programmable controller. The processor 305 may execute instructions stored in memory 310 to perform the methods and routines described herein. The processor 305 is communicatively coupled to memory 310, input device 315, output device 320, and transceiver 325.
[0068] Memory 310 may be a computer-readable storage medium. Memory 310 may include volatile computer storage media. For example, memory 310 may include RAM, including dynamic RAM ("DRAM"), synchronous dynamic RAM ("SDRAM"), and / or static RAM ("SRAM"). Memory 310 may include non-volatile computer storage media. For example, memory 310 may include a hard disk drive, flash memory, or any other suitable non-volatile computer storage device. Memory 310 may include both volatile and non-volatile computer storage media.
[0069] Memory 310 may store data relating to establishing a multipath unicast link and / or mobile operation. For example, memory 310 may store parameters, configurations, resource allocations, policies, etc., as described herein. Memory 310 may also store program code and related data, such as operating systems or other controller algorithms running on network node 300.
[0070] The input device 315 may include any known computer input device, including touch panels, buttons, keyboards, styluses, microphones, etc. The input device 315 may be integrated with the output device 320, for example, as a touchscreen or similar touch-sensitive display. The input device 315 may include a touchscreen on which text can be entered using a virtual keyboard displayed on the touchscreen and / or by handwriting on the touchscreen. The input device 315 may include two or more different devices, such as a keyboard and a touchscreen.
[0071] The output device 320 may be designed to output visual signals, acoustic signals, and / or tactile signals. The output device 320 may include an electronically controllable display or display device capable of outputting visual data to a user. For example, the output device 320 may include, but is not limited to, an LCD display, an LED display, an OLED display, a projector, or a similar display device capable of outputting images, text, etc., to a user. As another non-limiting example, the output device 320 may include a wearable display, such as a smartwatch, smart glasses, or a head-up display, that is separate from the rest of the network node 300 but communicatively coupled to it. Furthermore, the output device 320 may be a component of a smartphone, personal digital assistant, television, table computer, notebook (laptop) computer, personal computer, vehicle dashboard, etc.
[0072] The output device 320 may include one or more speakers for generating sound. For example, the output device 320 may generate an audible alarm or notification (e.g., a beep or chime). The output device 320 may include one or more haptic devices for generating vibration, motion, or other haptic feedback. All or part of the output device 320 may be integrated with the input device 315. For example, the input device 315 and the output device 320 may form a touchscreen or similar touch-sensitive display. The output device 320 may be located near the input device 315.
[0073] The transceiver 325 includes at least one transmitter 330 and at least one receiver 335. One or more transmitters 330 may be used to communicate with a UE, as described herein. Similarly, one or more receivers 335 may be used to communicate with network functions in the PLMN and / or RAN, as described herein. Although only one transmitter 330 and one receiver 335 are illustrated, a network node 300 may have any preferred number of transmitters 330 and receivers 335. Furthermore, the transmitters 330 and receivers 335 may be any preferred type of transmitter and receiver.
[0074] There are several standardized, real-time-appropriate transport architectures and protocols, including the Real-Time Transport Protocol (RTP), defined in IETF standard RFC3550 titled "RTP: A Transport Protocol for Real-Time Applications"; the Secure Real-Time Transport Protocol (SRTP), a securely provisioned real-time transport protocol, defined in IETF standard RFC3711 titled "The Secure Real-time Transport Protocol (SRTP)"; and WebRTC, a web-targeted stacked web real-time communication protocol, defined in the W3C standard recommendation dated March 6, 2023, titled "WebRTC: Real-Time Communication in Browsers."
[0075] RTP is a media codec-agnostic network protocol with application layer framing used to deliver multimedia (e.g., audio, video, etc.) data in real time over an IP network. RTP is used with its sister protocol for control, namely the Real-Time Transport Control Protocol (RTCP), to provide end-to-end functionality such as jitter compensation, packet loss and out-of-order delivery detection, synchronization, and source stream multiplexing. Figure 4 shows an overview of the RTP and RTCP protocol stacks. IP layer 405 carries signaling from the media session data plane 410 and from the media session control plane 450. The data plane 410 stack contains functions for User Datagram Protocol (UDP) 412, RTP 416, RTCP 414, media codecs 420, and quality control 422. The control plane 450 stack contains functions for UDP 452, Transmission Control Protocol (TCP) 454, Session Initiation Protocol (SIP) 462, and Session Description Protocol (SDP) 464.
[0076] SRTP is a secure version of RTP that provides encryption (primarily through payload secrecy), message authentication and integrity protection (by signing the PDU, i.e., the header and payload), and replay attack protection. Like RTP, the SRTP sister protocol is SRTCP, which provides the same functionality as its RTCP counterpart. Therefore, in the simplified SRTP version, RTP header information is still accessible but immutable, while the payload is encrypted. These security provisions are shown in Figure 7. Furthermore, the key exchange and additional security parameters required to use SRTP are based on the Datagram Transport Layer Security (DTLS) key exchange procedure. For these reasons, SRTP is used as the transport protocol for media within the WebRTC stack, ensuring secure RTC multimedia communication over a web browser interface.
[0077] Figure 5 shows an overview of the WebRTC (i.e., SRTP-based) protocol stack. As illustrated, IP layer 505 carries signaling from the data plane 510 and control plane 550. The data plane 510 stack includes functions for UDP 512, Interactive Connectivity Establishment (ICE) 524, Datagram Transport Layer Security (DTLS) 526, SRTP 517, SRTCP 515, media codecs 520, quality control 522, and SCTP 528. ICE 524 may use the Session Traversal Utilities for NAT (STUN) protocol and Traversal Using Relays around NAT (TURN) to handle real-time media content delivery across heterogeneous networks and NAT rules and firewalls. The SCTP528 data plane is primarily dedicated to application data channels and does not need to be time-constrained, while the SRTP517-based stack, i.e., SRTCP515, which includes control elements, encoding, i.e., media codec520, and quality of service (QoS), i.e., quality control522, is dedicated to time-constrained transport. The control plane550 is shown as comprising TCP554, TLS556, HTTP558, SSE / XHR / etc.568, XMPP / etc.570, SDP564, and SIP562.
[0078] The RTP and SRTP header information share the same format, as shown in Figures 6 and 7, respectively. Figure 6 shows an RTP packet 630, and Figure 7 shows an SRTP packet 760. A brief overview of the fixed header information for packets 630 and 760 is provided below.
[0079] The "V" 641 and 761 are two bits that indicate the protocol version being used.
[0080] "P" 643, 763 is a 1-bit field that indicates the presence of one or more zero-padding octets at the end of the payload, thereby indicating that padding may be necessary, in particular for fixed-size encryption blocks or for carrying multiple RTP / SRTP packets over lower-layer protocols.
[0081] The "X" 634, 764 is a single bit that indicates that an RTP header extension, typically associated with a specific data / profile that carries more information about the data (for example, an RTP header extension for video data, as described in an IETF working draft dated November 2021 titled "Frame Marking RTP Header Extension," or a frame that marks a comprehensive RTP header extension, such as an RTP / SRTP extension protocol, as described in the IETF standard RFC6904 titled "Encryption of Header Extensions in the Secure Real-time Transport Protocol (SRTP)") follows a standard fixed RTP / SRTP header.
[0082] The "CC" 636 and 766 are 4 bits that indicate the number of contributing media sources (CSRCs) following the fixed header.
[0083] The bits "M" 638 and 768 are single bits intended to mark information frame boundaries within a packet stream, and their behavior is strictly specified by the RTP profile (e.g., H.264, H.265, H.266, AV1, etc.).
[0084] "PT" 640, 770 are 7 bits (e.g., 96 for H.264, 97 for H.265, 98 for AV1, etc.) that indicate the payload type, which is dynamic and often negotiated by SDP in the case of audio and video codec profiles. Payload profiles rely on IETF profiles registered with IANA that describe how data transmissions are encapsulated within the payload of an RTP PDU. Current IANA-registered payload profiles, such as those specified in the ITU-T standard H.265 V8 (08 / 2021), describe audio / video codec and application-based forward error correction (FEC) coded media content, but do not yet address uncoded non-audiovisual data formats.
[0085] The "sequence numbers" 642 and 772 are 16 bits that represent sequence numbers that increment by 1 each time an RTP data packet is sent through the session.
[0086] The "timestamp" 644, 774 is a 32-bit timestamp in tick units of the payload type clock, reflecting the sampling moment of the first octet of the RTP data packet (associated with a video frame for a video stream), and the first timestamp of the first RTP packet is randomly selected.
[0087] The "Synchronization Source (SSRC) Identifiers" 646 and 776 are 32-bit fields that represent random identifiers for the source of a stream of RTP packets that form part of the same timing and sequence number space, allowing the receiver to group packets based on the synchronization source for playback.
[0088] The "Contributing Source (CSRC) identifiers" 648, 778 are a list of up to 16 32-bit CSRC entries, each given the amount of CSRC mixed by the RTP mixer in the current payload, as signaled by the CC bit. The list identifies the contributing source for the payload contained in this packet, given the contributing source's CSRC identifier.
[0089] A brief summary of the remaining aspects of the complete header information for packets 630 and 760 is described below.
[0090] The "RTP Header Extensions" 650, 780 are variable-length fields present when X bits 634, 764 are marked. The header extensions are appended to the RTP fixed header information, after the CSRC lists 648, 778, if present. The RTP header extensions 650, 780 consist of the following fields: a 16-bit extension identifier defined by the profile and usually negotiated and determined via the Session Description Protocol (SDP) signaling mechanism; a 16-bit length field describing the extension header length in 32-bit multiples, except for the first 32 bits corresponding to the 16-bit extension identifier and the 16-bit length field itself; and a 32-bit aligned header extension raw data field formatted according to the format of some RTP header extension identifier designations.
[0091] The RTP header extensions 650 and 780 format and syntax are similar to those of SRTP. The format and syntax are shown in Figure 8. In addition, in both RTP and SRTP, only one RTP extension header 650 or 780 may be appended to the fixed header information, as described in the IETF standard RFC 3550, titled "RTP: A Transport Protocol for Real-Time Applications". However, for both RTP and SRTP, there are extensions to the base protocol that allow multiple RTP header extensions of a given type to be appended to the protocol's fixed header information, as described in the IETF standard RFC 8285, titled "A General Mechanism for RTP Header Extensions".
[0092] In some embodiments, RTP header extensions generated at the source may be ignored by the destination endpoint, which does not have the knowledge to interpret and process RTP header extensions sent by the source endpoint.
[0093] Starting with modern hybrid video coding, the video coding domain and metadata support are then briefly explained.
[0094] Interactivity and immersion in modern and future multimedia XR applications require assurances regarding the packet error rate (PER) and packet delay budget (PDB) for Quality of Experience (QoE). Video source jitter and wireless channel probabilistic characteristics of mobile communication systems make it difficult for previous systems to meet, particularly for high-rate-specific digital video transmissions such as 4K, 3D video, and 2x2K eye-buffered video.
[0095] Current video source information is encoded based on the 2D, 2D+depth, or alternatively, 3D representation of the video content. The encoded base stream video content is generally organized into two abstraction layers, regardless of the source encoder, intended to separate the storage and video coding domains—namely, network transport packetization and formatting—from the codec's video coding-related syntax and associated semantics, respectively. The first determines the bitstream format, and the latter specifies the content of the video-encoded bitstream.
[0096] For example, the MPEG video codec family (e.g., H.264, H.265, H.266) relies on Network Abstraction Layer (NAL) units to packetize and store bitstreams in a byte-aligned format for transport or storage over various media (including networks). A NAL unit (NALU) may encapsulate both Video Coding Layer (VCL) information, i.e., video-coded content (e.g., frames, slices, tiles, etc.), and non-VCL information, i.e., parameter sets, Supplemental Enhancement Information (SEI) messages, respectively. The NAL syntax thus encapsulates the VCL and non-VCL information and provides an abstraction containerization mechanism for coded streams in transit, i.e., for disk storage / caching / transmission and parsing / decoding.
[0097] In another example, open-source video codec alternatives (e.g., VP8 / VP9 or AV1) employ techniques similar to MPEG video codecs for packetization, storage, and communication across various media. For example, an AV1 bitstream may comprise open bitstream units (OBUs), each of which may store one or more video-coded frames, video-coded tiles, non-video-coded padding, and non-video-coded metadata as OBU_METADATA.
[0098] Based on the above, NALU, or alternatively OBU, provides a mechanism that can be utilized for transporting metadata unrelated to video-encoded bitstream chroma and luminance representations.
[0099] On the other hand, VCL, or alternatively video coding information, encapsulates the encoder's video coding procedure and compresses the source coded video information based on several entropy coding methods, such as context-adaptive binary arithmetic coding (CABAC) and context-adaptive variable-length coding (CAVLC).
[0100] A simplified description of the VCL procedure for comprehensively encoding video content is provided below. Pictures in a video sequence are divided into coding units of a configured size (e.g., macroblocks, coding tree units, blocks, or variations thereof). Coding units may then be divided under several tree partitioning structures or similar hierarchical structures, as described in ITU-T standards H.264 V8(08 / 2021), ITU-T standards H.265 V8(08 / 2021), and ITU-T standards H.266 V4(04 / 2022). For example, such a tree partitioning structure may have, for instance, a 10-way partition under a binary / ternal / quaternal tree or under some predetermined geometrically motivated 2D segmentation pattern, such as that described by de Rivaz, p. and Haughton (2018) in the paper entitled "AV1 Bitstream & Decoding Process Specification", 182 from the Alliance for Open Media.
[0101] The encoder uses visual references of such coding units to encode picture content using a residual-based difference scheme. The residuals are determined based on a prediction mode related to the reconstruction of information. Two prediction modes are universally available: intra-prediction (also simply called intra) or inter-prediction (or inter in short form). The intra mode is based on deriving and predicting residuals by calculating the residuals of the current coding unit based on the content of other coding units in the current picture, i.e., assuming the coded content of those adjacent coding units. The inter mode, on the other hand, is based on deriving and predicting residuals by calculating the residuals of the current coding unit based on the content of coding units from other pictures, i.e., assuming the coded picture content of those adjacent pictures.
[0102] The residuals are then further transformed for compression using several multi-dimensional (2D / 3D) spatial multimodal transforms, e.g., frequency-based (i.e., discrete cosine transform or similar) or wavelet-based linear transforms (e.g., Walsh-Hadamard transform or equivalent discrete wavelet transform), to extract the most prominent frequency components of the coding unit's residuals. Small high-frequency contributions of the residuals are reduced, and the floating-point transformed representation of the remaining residuals is further quantized based on several parametric quantization procedures to a selected number of bits per sample, e.g., 8 / 10 / 12 bits. Finally, the transformed and quantized residuals, as well as their associated motion vectors to their predicted references in either intra-mode or inter-mode, are encoded using an entropy coding mechanism to compress the information based on the probabilistic distribution of the source bit content. The output of this operation is a bitstream of the coded residual content in the VCL.
[0103] A simplified, comprehensive diagram of the blocks of a modern hybrid video codec (which applies both temporal and spatial compression via intra-prediction / inter-prediction) is shown in Figure 9.
[0104] Figure 9 shows a simplified block diagram 900 of a comprehensive video codec that performs both spatial and temporal (motion) compression of a video source. The encoder block is contained within domain 910 tagged “Encoder”. The decoder block is contained within domain 920 tagged “Decoder”. Those skilled in the art may associate the comprehensive diagram above representing the hybrid codec with a very large number of state-of-the-art video codecs, including, but not limited to, H.264, H.265, H.266 (collectively referred to as H.26x) or VP8 / VP9 / AV1. Accordingly, concepts used herein are considered in a general sense unless specifically clarified and unless the scope is narrowed below to some codec embodiments.
[0105] Block diagram 900 shows that the raw input video frame (picture) 901 is input to the picture block division function block 911 of the encoder 910. The subsequent function block 912 is illustrated as "spatial transformation". The subsequent function block 913 is illustrated as "quantization". The subsequent function block 914 is illustrated as "entropy coding". This function block 914 outputs to the video coded bitstream 902, but also to the motion estimation 915 of the encoder 910. The motion estimation 915 outputs to the interpretation prediction block 921 of the decoder 920. This block 921 outputs to a buffer 920 which itself outputs to the restored video frame video (picture) 903. The interpretation prediction block 921 may be switched to connect to an additive coupling unit that supplies to the spatial transformation 912. Furthermore, it is illustrated in block diagram 900 that the quantization block 913 may output to the inverse quantization block 926 of the decoder 920. The inverse quantization block 926 is illustrated as receiving the entropy decoder 927 of the video-encoded bitstream 902. The inverse quantization block 926 outputs to the inverse spatial transform block 925, which in turn feeds via an additive coupler to the loop and visual filtering block 923, which itself feeds to the buffer 922. As described above, block diagram 900 is illustrated as an example to convey the various functional blocks of a modern hybrid video codec in terms of both encoder and decoder operation.
[0106] The coded residual bitstream 902 is encapsulated within the base stream as a NAL unit, or equivalently as an OBU, ready for storage or transmission over the network. NAL units, or alternatively OBUs, are the primary syntactic elements of the video codec, and they may encapsulate coded video parameters (e.g., video parameter sets / sequence parameter sets / picture parameter sets (VPS / SPS / PPS)), one or more supplemental enhancement information (SEI) messages, or alternatively, an OBU metadata payload, as well as coded video headers and residual data (e.g., pictures, or equivalently, slices as segments of video frames or video tiles). The encapsulation generic syntax carries information described by codec-specific semantics intended to determine the use of metadata, non-video coded data, and video coded data, and assists the decoding process.
[0107] In one example referring to the MPEG video codec family (e.g., H.264, H.265, H.266), the NAL unit encapsulation syntax consists of a header portion that determines the beginning of the NAL units and their type, and a raw byte payload sequence that stores NAL unit-related information. The NAL unit payload may then be formed from a payload syntax or payload-specific header and associated payload-specific syntax. An important subset of NAL units consists of parameter sets, e.g., VPS, SPS, PPS, SEI messages, and constituent NAL units (also collectively referred to as non-VCL NAL units), as well as picture slice NAL units that store video encoded data (e.g., entropy-based arithmetic coding) as VCL information. These concepts are illustrated in Figure 10 for the context of a basic stream that is comprehensively applicable to the H.264, H.265, and H.266 MPEG families of video codecs.
[0108] Figure 10 shows a video-encoded base stream and its corresponding multiple NAL units 1000. Each NAL unit 1000 consists of a header 1010 and a payload 1020. The header 1010 stores information about the type, size, and video coding attributes, as well as parameters for information enclosed in the NAL unit data. The NAL unit data may be a non-VCL NAL with a video / sequence / picture parameter payload 1021 and a supplemental enhancement information payload 1022, or a VCL NAL with a frame / picture / slice payload 1023 having a header 1023a and a video-encoded payload 1023b. A non-VCL NAL may include one or more SEI messages 1022a in the supplemental enhancement information payload 1022. The NAL header 1010 is illustrated as comprising the NAL unit type, NAL unit byte length, video coding layer ID, and time video coding layer ID.
[0109] Therefore, the decoder implementation may implement a bitstream parser that extracts necessary metadata information and VCL-related metadata from the NAL unit sequence 1000, decode the VCL residual coded data sequence and convert it into its transformed and quantized values, apply an inverse linear transform to reconstruct residual significant content, perform intra-prediction or inter-prediction, apply additional filtering and error hiding procedures to reconstruct the luminance and chromatic representation of each coding unit, and play back the raw picture sequence representation as video playback.
[0110] These operations and procedures may occur sequentially or in parallel, as enumerated, depending on the decoder-specific implementation. Those skilled in the art should recognize that similar high-level operations are applicable to other families of video codecs, such as the AV1 codec.
[0111] To support the disclosures in this specification, metadata support in video codecs is briefly described below.
[0112] Modern video codecs, such as H.264, H.265, AV1, or H.266 as an alternative, provide a byte-aligned transport mechanism for metadata within the video-encoded base stream or, alternatively, the bitstream. Since the information contained within is not relevant to modifying the rumor or chroma of the decoded frame or, alternatively, the picture, such non-video-encoded data is referred to as metadata in the video coding context. This metadata is encapsulated within the NALU as SEI messages for the H.26x MPEG family of codecs, and within the OBU as OBU metadata for AV1.
[0113] For each codec, the metadata is found in the codec specification, for example, (ITU-T standard H.264 (08 / 2021), ITU-T standard H.265 V8 (08 / 2021), ITU-T standard H.266 V4 (04 / 2022), "AV1 Bitstream & Decoding Process Specification" from Alliance for Open Media, Rivaz, p and Haughton (2018) in paper titled 182), or alternatively, for example, the ITU-T Series H specification V8 (08 / 2020) titled "Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services-coding of moving video: versatile supplemental enhancement information messages for coded video bitstreams," or "Terminal provider codes notification form: available information regarding the identification of national authorities for the assignment of ITU-T recommendation T.35 terminal provider" Within the ITU specifications for metadata types, such as in ITU-T Recommendation T.35 titled "codes," there may be additional types related to different syntax and semantics. However, none of the specifications specify a user data SEI message or, as an alternative, an OBU metadata user private data format. Furthermore, all video codecs described herein, as well as their associated encoders and decoders, allow exposure of user data metadata types to the application layer via a dedicated interface, either through the passage of a NALU SEI message or, as an alternative, an OBU that carries the OBU metadata to the application layer.
[0114] The current WebRTC specification requires AVC / H.264 with its restricted baseline profile and VP8, along with additional optional support for VP9. Given the increasing support for AV1 and H.265 in major browser engines (e.g., Chromium, used by Google Chrome and Microsoft Edge, or alternatively, Mozilla, used by Mozilla Firefox), the WebRTC specification is expected to evolve soon to include additional optional support for AV1 and potentially a restricted profile of H.265. Meanwhile, in 3GPP up to release 18, H.264 and H.265 are the designated and supported default video codecs, with support for H.266 and AV1 being considered for further consideration in subsequent releases.
[0115] This disclosure leverages these facts regarding user data metadata and video codec support to propose a novel real-time in-band transport mechanism for interactive and immersive metadata related to XR applications and their video streams, which are associated with segmented rendering.
[0116] Split rendering and split rendering architectures, such as the 3GPP technical specification TS26.565 v0.3.0 titled "Split Rendering Media Service Enabler" and the 3GPP technical specification 26.506 v1.1.0 titled "5G Real-time Media Communication Architecture," are being considered by 3GPP. Split rendering offers advantages such as energy savings and high performance, in addition to enabling interactive and immersive XR applications. Split rendering enhances the user experience by providing access to highly sophisticated renderings that would otherwise be impossible or would require extremely high energy consumption in AR / VR glasses or 5G tethered UEs.
[0117] In split rendering, all or part of a 3D scene is rendered remotely on an edge application server (EAS), or alternatively, a split rendering server (SRS) located within the edge data network (EDN). The results of the split rendering process are streamed down to a 5G UE (if a tethered AR architecture is considered) or XR glasses for display. The spectrum of split rendering operations can be broad, ranging from full pre-rendering at the edge to offloading partial, more extensive rendering operations to the edge. This diversity extends from frames supporting 2D with single-eye buffer rendering, 2D with dual-eye buffer rendering, 2D with depth information rendering, and 3D scene rendering.
[0118] In the case of 2D using single-eye buffer rendering, the edge application server will generate a single 2D video rendering of the visual scene. Depending on the UE configuration, 2D rendering with two eye buffers and appropriate projection (e.g., equirectangular) may be required. Other features that support the encoding of additional information such as depth or transparency in the stream may also be added to the 2D rendering. On the UE, the scene synthesizer layers the available streams and constructs a primitive image buffer, which is exchanged with the XR runtime across the swap chain for display on AR glasses. On the other hand, in the case of partial rendering of 3D content, offloading means delegating some rendering operations to the EAS while still receiving the 3D scene in the UE. One example is offloading the light sintering of scene textures to the edge, which can be done using techniques such as ray tracing.
[0119] Regarding traffic, the segmented rendering UL traffic pattern can be summarized as the UE, or alternatively the AR glasses, streaming pose predictions to the SRS in the EDN. This traffic may also store relevant AR UL video streams, depending on the application type (e.g., multi-party AR conference, interactive and immersive classroom AR).
[0120] Regarding traffic, the split-rendering DL traffic pattern can be summarized as the UE or, alternatively, AR glasses receiving the rendered video media for display.
[0121] Typically, XR runtimes benefit from the rendered media being passed along with the associated poses used for rendering in order to perform proper scene composition and display. For example, an XR runtime may need to perform pose correction based on late-stage reprojection or, alternatively, asynchronous time warping to align with the current display time. Additional information, such as XR spatial processing or references and XR timestamps related to the SRS rendering operation for spatial and temporal synchronization, may be required by the XR runtime.
[0122] Therefore, the determination of the segmented rendering transport and signaling mechanism is crucial for XR interactive and immersive applications, and these are addressed herein.
[0123] This solution deals with the use of user data SEI messages (e.g., H.264, H.265, H.266) or alternatively OBU metadata (e.g., AV1) for transporting interaction and immersion metadata associated with XR applications in either UL or DL. The format used for such user data payloads is inclusive and may comprise at least two fields. The first field is an identifier that determines the syntax and semantics of the user data format, and the second field stores the user data information as a payload encoded according to the format identified by the first field. In some embodiments, the first field is a UUID that needs to be signaled to the corresponding remote receiver by the source sender so that the receiver can parse and process the user data information. This disclosure, in detail, specifies three signaling mechanisms that enable the use of such user data via SEI messages or alternatively OBU metadata for interactive and immersive applications sponsored by segmented rendering at the edge. For user data identifiers, such as UUIDs, the proposed signaling mechanisms are application-based signaling via the application plane, network-supported signaling via network application functions through the control plane, and SDP attribute-based in-band signaling via the user plane.
[0124] In some embodiments, a user data identifier, such as a UUID, may depend on at least one of the media capabilities of the sender, receiver, and split-rendering EAS, and thus its negotiation and signaling is performed for each media session of an interactive and immersive application, either during session initialization or, alternatively, during session renewal.
[0125] Disclosures herein provide a device for wireless communication comprising a processor and memory coupled to the processor, wherein the processor is configured to determine a media configuration for segmented rendering of a video media flow, the media configuration comprising one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; to signal the media configuration to a second device; to establish a multimedia segmented rendering content delivery session comprising the video media flow with the second device, at least in part on the media configuration; and to use one or more data units of multimedia immersion and interaction data to segment render the video media flow.
[0126] In some embodiments, the non-video encoded metadata comprises a first field having an identifier for the syntax and semantic representation format of one or more data units of multimedia immersion and interaction data, and a second field having one or more data units of multimedia immersion and interaction data encoded according to the syntax and semantic representation format corresponding to the identifier in the first field.
[0127] In some embodiments, the identifier is represented as a universally unique identifier (UUID).
[0128] In some embodiments, the UUID may be specific to a particular application or session, or it may be globally unique. The UUID may conform to the ISO-IEC-11578 ANNEX A format and ISO-IEC9834-8 version 4 UUID, i.e., a randomly generated UUID, or something similar. The term “session” comprises a temporary, interactive, i.e., updatable set of configurations and rules that determine the exchange of information, including media content, between two or more endpoints connected over a network, such as media formats and codecs, and network configuration.
[0129] In some embodiments, the processor is configured to cause the device to determine a media configuration by causing the device to determine a media configuration for a plurality of video media flows, each having an encoded video stream and each unencoded metadata, wherein one or more parameters of the media configuration map identifiers of a first field of each unencoded metadata to their respective video media flows.
[0130] In some embodiments, one or more parameters of the media configuration map a first field of the respective non-video-encoded metadata to its respective video media flow, based on at least one of application-specific stream mappings, 5-tuple representations, and media flow description attributes. For example, a 5-tuple may describe an IP flow comprising an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. A media description under SDP may be a media flow association to one or more UUIDs, as provided in the media description attributes defined by SDP.
[0131] In some embodiments, the processor is configured to cause the device to signal a media configuration via at least one of the following: an application interface (preferably between the device's application service provider "ASP" and the second device's split rendering-aware application); a control plane interface (preferably between the device's real-time communication application function "RTC AF" and the second device's media session handler "MSH"); and a user plane interface (preferably between the device's split rendering server "SRS" and the second device's split rendering client "SRC"). The control plane interface may be the RTC-5 reference interface in 5GS, and the user plane interface may be SR-4, RTC-4, or use SR-4m media-centric interfacing.
[0132] In some embodiments, the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs, consisting of the H.264 video codec standard, the H.265 video codec standard, the H.266 video codec standard, and the AV1 video codec standard. Other video codec standards that partially rely on the above standards, such as OMAF, V-PCC, or similar, may also be applied. OMAF encodes omnidirectional video content and comprises at least one video stream encoded using AVC and HEVC, while V-PCC encodes 2D projected 3D video and comprises a video stream that encodes projected 2D flat frames using AVC and HEVC.
[0133] In some embodiments, the video codec is an H.264, H.265, or H.266 codec, and the non-video coded metadata is encapsulated as a user data payload in one or more supplemental enhancement information "SEI" messages of type "User Data Unregistered", and / or the video codec is an AV1 codec, and the non-video coded metadata is encapsulated as a user data payload in one or more metadata open bitstream units "OBU" of type "Unregistered User Private Data". For SEI messages, the payload type may be equal to 5, and the SEI messages may be prepended or prepended to the NALU. For OBUs, the OBU metadata_type may be "X", where X can be any of 6 to 31.
[0134] In some embodiments, one or more data units of multimedia immersion and interaction data comprises immersion and interaction data selected from a list of augmented reality object representations, which include at least one of user viewpoint data, user field of view data, user posture / orientation data, user gesture tracking data, user body tracking data, user facial feature tracking data, user action and / or user input data, segmented rendering posture and spatial information, and graphical descriptions of objects and object position anchors.
[0135] In some embodiments, one or more data units of multimedia immersion and interaction data comprise augmented reality "XR" multimedia immersion and interaction data.
[0136] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources selected from a list of media sources, each consisting of one or more physical dedicated controllers, one or more red-green-blue "RGB" cameras, one or more RGB depth "RGBD" cameras, one or more infrared "IR" cameras, one or more microphones, and one or more haptic transducers.
[0137] In some embodiments, for the purposes of matching one or more rendered frames, displaying the latest user pose or orientation information, late-stage reprojection in a segmented rendering client "SRC", partial or complete pre-rendering of one or more frames in a segmented rendering server "SRS", and / or estimation of the most likely user pose and orientation in the SRS related to the expected display time for one or more frames to be partially or completely pre-rendered, the processor is configured to cause the device to use one or more data units of multimedia immersion and interaction data to segment-render a video media flow.
[0138] In some embodiments, the processor is configured to signal the device a media configuration when a multimedia segmented rendering content delivery session begins and / or when the multimedia segmented rendering content delivery session is updated.
[0139] In some embodiments, the processor is further configured to cause the device to determine a media configuration based on an application configuration of an application hosted on a second device, which includes one or more media capabilities of the second device, and / or at least one of an application service provider "ASP" configuration provisioned to at least one of a real-time communication application function "RTC AF", a provisioning function, a segmented rendering application function "SR AF", and / or a configuration function.
[0140] In some embodiments, the application configuration may be pre-configured with supported UUIDs, and / or the transport may be automatically enabled when a UUID is received.
[0141] In some embodiments, the first device is a network entity / node, and the second device is a UE.
[0142] Figure 11 shows one embodiment 1100 of a wireless communication method in a wireless communication system.
[0143] The first step 1110 comprises determining a media configuration for segmented rendering of a video media flow, the media configuration comprising one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata.
[0144] A further step 1120 comprises signaling the media configuration to a second device.
[0145] A further step 1130 comprises establishing a multimedia segmented rendering content distribution session with a video media flow, together with a second device, at least in part, based on the media configuration.
[0146] A further step 1140 comprises using one or more data units of multimedia immersion and interaction data to segment-render the video media flow.
[0147] In some embodiments, method 1100 may be executed by a processor that executes program code, such as a microcontroller, microprocessor, CPU, GPU, auxiliary processing unit, FPGA, etc.
[0148] In some embodiments, the non-video encoded metadata comprises a first field having an identifier for the syntax and semantic representation format of one or more data units of multimedia immersion and interaction data, and a second field having one or more data units of multimedia immersion and interaction data encoded according to the syntax and semantic representation format corresponding to the identifier in the first field.
[0149] In some embodiments, the identifier is represented as a universally unique identifier (UUID).
[0150] In some embodiments, the UUID may be specific to a particular application or session, or it may be globally unique. The UUID may conform to the ISO-IEC-11578 ANNEX A format and ISO-IEC9834-8 version 4 UUID, i.e., a randomly generated UUID, or something similar. The term “session” comprises a temporary, interactive, i.e., updatable set of configurations and rules that determine the exchange of information, including media content, between two or more endpoints connected over a network, such as media formats and codecs, and network configuration.
[0151] In some embodiments, determining a media configuration involves determining a media configuration for a plurality of video media flows, each having an encoded video stream and its own unencoded metadata, wherein one or more parameters of the media configuration map identifiers in a first field of the respective unencoded metadata to their respective video media flows.
[0152] In some embodiments, one or more parameters of the media configuration map a first field of the respective non-video-encoded metadata to its respective video media flow, based on at least one of application-specific stream mappings, 5-tuple representations, and media flow description attributes. The 5-tuple may describe an IP flow, for example, comprising an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. A media description under SDP may comprise a media flow association to one or more UUIDs, such as those provided in the SDP media description attribute.
[0153] In some embodiments, signaling the media configuration includes signaling via at least one of the following: an application interface (preferably between the device's application service provider "ASP" and the second device's split rendering-aware application); a control plane interface (preferably between the device's real-time communication application function "RTC AF" and the second device's media session handler "MSH"); and a user plane interface (preferably between the device's split rendering server "SRS" and the second device's split rendering client "SRC"). The control plane interface may be the RTC-5 reference interface in 5GS, and the user plane interface may be SR-4, RTC-4, or use SR-4m media-centric interfacing.
[0154] In some embodiments, the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs, consisting of the H.264 video codec standard, the H.265 video codec standard, the H.266 video codec standard, and the AV1 video codec standard. Other video codec standards that partially rely on the above standards, such as OMAF, V-PCC, or similar, may also be applied. OMAF encodes omnidirectional video content and comprises at least one video stream encoded using AVC and HEVC, while V-PCC encodes 2D projected 3D video and comprises a video stream that encodes projected 2D flat frames using AVC and HEVC.
[0155] In some embodiments, the video codec is an H.264, H.265, or H.266 codec, and non-video coded metadata is encapsulated as a user data payload in one or more supplemental enhancement information "SEI" messages of type "User Data Unregistered", and / or the video codec is an AV1 codec, and non-video coded metadata is encapsulated as a user data payload in one or more metadata open bitstream units "OBU" of type "Unregistered User Private Data". For SEI messaging, the payload type may be 5, and the SEI message may be prepended or prepended to the NALU. For OBU, the OBU metadata_type may be "X", where X can be any of 6 to 31.
[0156] In some embodiments, one or more data units of multimedia immersion and interaction data comprises immersion and interaction data selected from a list of augmented reality object representations, which include at least one of user viewpoint data, user field of view data, user posture / orientation data, user gesture tracking data, user body tracking data, user facial feature tracking data, user action and / or user input data, segmented rendering posture and spatial information, and graphical descriptions of objects and object position anchors.
[0157] In some embodiments, one or more data units of multimedia immersion and interaction data comprise augmented reality "XR" multimedia immersion and interaction data.
[0158] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources selected from a list of media sources, each consisting of one or more physical dedicated controllers, one or more red-green-blue "RGB" cameras, one or more RGB depth "RGBD" cameras, one or more infrared "IR" cameras, one or more microphones, and one or more haptic transducers.
[0159] For the purposes of displaying one or more rendered frame alignment, the latest user pose or orientation information, late-stage reprojection in a segmented rendering client "SRC", partial or complete pre-rendering of one or more frames in a segmented rendering server "SRS", and / or estimation of the most accurate user pose and orientation in the SRS related to the expected display time for one or more frames to be partially or completely pre-rendered, some embodiments include using one or more data units of multimedia immersion and interaction data to segment-render a video media flow.
[0160] Some embodiments include signaling the media configuration at the start of a multimedia split-rendered content delivery session and / or when the multimedia split-rendered content delivery session is updated.
[0161] In some embodiments, determining the media configuration involves determining the media configuration based on an application configuration of an application hosted on a second device, which includes one or more media capabilities of the second device, and / or an application service provider "ASP" configuration provisioned to at least one of the following: a real-time communication application function "RTC AF", a provisioning function, a segmented rendering application function "SR AF", and / or a configuration function. The application configuration may be pre-configured with supported UUIDs, and / or the transport may be automatically enabled when a UUID is received.
[0162] In some embodiments, the method is performed by a network entity / node, and the second device is a UE.
[0163] Disclosures herein further provide an apparatus for wireless communication in a wireless communication system, comprising a processor and memory coupled to the processor, wherein the processor is configured to receive a media configuration from a first apparatus for segmented rendering of a video media flow, the media configuration comprising one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; configure the apparatus to receive the video media flow; decode the encoded video stream of the video media flow, comprising decoding the non-video encoded metadata and consuming one or more data units of multimedia immersion and interaction data from the non-video encoded metadata.
[0164] In some embodiments, the non-video encoded metadata comprises a first field having an identifier for the syntax and semantic representation format of one or more data units of multimedia immersion and interaction data, and a second field having one or more data units of multimedia immersion and interaction data encoded according to the syntax and semantic representation format corresponding to the identifier in the first field.
[0165] In some embodiments, the identifier is represented as a universally unique identifier (UUID).
[0166] In some embodiments, the UUID may be specific to a particular application or session, or it may be globally unique. The UUID may conform to the ISO-IEC-11578 ANNEX A format and ISO-IEC9834-8 version 4 UUID, i.e., a randomly generated UUID, or something similar. The term “session” comprises a temporary, interactive, i.e., updatable set of configurations and rules that determine the exchange of information, including media content, between two or more endpoints connected over a network, such as media formats and codecs, and network configuration.
[0167] In some embodiments, the media configuration is for multiple video media flows, each having an encoded video stream and its own unencoded video metadata, and one or more parameters of the media configuration map identifiers in the first field of each unencoded video metadata to their respective video media flows.
[0168] In some embodiments, one or more parameters of the media configuration map a first field of the respective non-video-encoded metadata to its respective video media flow, based on at least one of application-specific stream mappings, 5-tuple representations, and media flow description attributes. The 5-tuple may describe an IP flow comprising an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. The media description under SDP may be a media flow association to one or more UUIDs, such as those provided in the SDP media description attribute.
[0169] In some embodiments, the processor is configured to cause the device to receive media configurations via at least one of the following: an application interface (preferably between the device's application service provider "ASP" and the second device's split rendering-aware application); a control plane interface (preferably between the device's real-time communication application function "RTC AF" and the second device's media session handler "MSH"); and a user plane interface (preferably between the device's split rendering server "SRS" and the second device's split rendering client "SRC"). The control plane interface may be the RTC-5 reference interface in 5GS. The user plane interface may be SR-4, RTC-4, or use SR-4m media-centric interfacing.
[0170] In some embodiments, the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs, consisting of the H.264 video codec standard, the H.265 video codec standard, the H.266 video codec standard, and the AV1 video codec standard. Other video codec standards that partially rely on the above standards, such as OMAF, V-PCC, or similar, may also be applied. OMAF encodes omnidirectional video content and comprises at least one video stream encoded using AVC and HEVC, while V-PCC encodes 2D projected 3D video and comprises a video stream that encodes projected 2D flat frames using AVC and HEVC.
[0171] In some embodiments, the video codec is an H.264, H.265, or H.266 codec, and non-video coded metadata is encapsulated as a user data payload in one or more Supplemental Enhancement Information "SEI" messages of type "Unregistered User Data", and / or the video codec is an AV1 codec, and non-video coded metadata is encapsulated as a user data payload in one or more metadata open bitstream units "OBU" of type "Unregistered User Private Data". For SEI messaging, the payload type may be 5, and the SEI message may be prepended or prepended to the NALU. For OBU, the OBU metadata_type may be X, where X can be any of 6 to 31.
[0172] In some embodiments, one or more data units of multimedia immersion and interaction data comprises immersion and interaction data selected from a list of augmented reality object representations, which include at least one of user viewpoint data, user field of view data, user posture / orientation data, user gesture tracking data, user body tracking data, user facial feature tracking data, user action and / or user input data, segmented rendering posture and spatial information, and graphical descriptions of objects and object position anchors.
[0173] In some embodiments, one or more data units of multimedia immersion and interaction data comprise augmented reality "XR" multimedia immersion and interaction data.
[0174] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources selected from a list of media sources, each consisting of one or more physical dedicated controllers, one or more red-green-blue "RGB" cameras, one or more RGB depth "RGBD" cameras, one or more infrared "IR" cameras, one or more microphones, and one or more haptic transducers.
[0175] In some embodiments, for the purpose of late-stage reprojection in a split-rendering client "SRC" for displaying one or more rendered frame alignments, up-to-date user pose or orientation information, the processor is configured to cause the device to use one or more data units of multimedia immersion and interaction data to split-render a video media flow.
[0176] In some embodiments, the processor is configured to cause the device to receive a media configuration when a multimedia segmented rendering content delivery session begins and / or when the multimedia segmented rendering content delivery session is updated.
[0177] In some embodiments, the media configuration is an application configuration of an application hosted on the device, comprising one or more media capabilities of the device, and / or at least one of an application service provider "ASP" configuration provisioned to at least one of a real-time communication application function "RTC AF", a provisioning function, a segmented rendering application function "SR AF", and / or a configuration function.
[0178] The application configuration may be pre-configured using supported UUIDs, and / or the transport may be automatically enabled when a UUID is received.
[0179] In some embodiments, the first device is a network entity / node, and the device is a UE.
[0180] Figure 12 shows one embodiment of method 1200 for wireless communication.
[0181] The first step 1210 comprises receiving a media configuration for segmented rendering of a video media flow from a first device, the media configuration comprising one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata.
[0182] A further step 1220 comprises configuring the device using a media configuration to receive a video media flow.
[0183] A further step 1230 comprises decoding the encoded video stream of the video media flow, the decoding comprising extracting non-video encoded metadata.
[0184] A further step 1240 comprises consuming one or more data units of multimedia immersion and interaction data from non-video encoded metadata.
[0185] In some embodiments, method 1200 may be executed by a processor that executes program code, such as a microcontroller, microprocessor, CPU, GPU, auxiliary processing unit, FPGA, etc.
[0186] In some embodiments, the non-video encoded metadata comprises a first field having an identifier for the syntax and semantic representation format of one or more data units of multimedia immersion and interaction data, and a second field having one or more data units of multimedia immersion and interaction data encoded according to the syntax and semantic representation format corresponding to the identifier in the first field.
[0187] In some embodiments, the identifier is represented as a universally unique identifier (UUID).
[0188] In some embodiments, the UUID is specific to a particular application or session, or it is globally unique.
[0189] The UUID may conform to the ISO-IEC-11578 ANNEX A format and ISO-IEC9834-8 version 4 UUID, i.e., a randomly generated UUID, or something similar. The term “session” refers to a temporary, interactive, i.e., updatable set of configurations and rules that determine the exchange of information, including media content, between two or more endpoints connected over a network, comprising, for example, media formats and codecs, network configurations.
[0190] In some embodiments, the media configuration is for multiple video media flows, each having an encoded video stream and its own unencoded video metadata, and one or more parameters of the media configuration map identifiers in the first field of each unencoded video metadata to their respective video media flows.
[0191] In some embodiments, one or more parameters of the media configuration map a first field of the respective non-video-encoded metadata to its respective video media flow based on at least one of application-specific stream mappings, 5-tuple representations, and media flow description attributes.
[0192] A 5-tuple may describe an IP flow comprising an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. Media descriptions under SDP may include media flow associations to description attributes, such as those provided in SDP media description attributes.
[0193] In some embodiments, receiving a media configuration involves receiving it via at least one of the following: an application interface (preferably between the device's application service provider "ASP" and the second device's split rendering-aware application); a control plane interface (preferably between the device's real-time communication application function "RTC AF" and the second device's media session handler "MSH"); and a user plane interface (preferably between the device's split rendering server "SRS" and the second device's split rendering client "SRC").
[0194] In some embodiments, the control plane interface may be the RTC-5 reference interface in 5GS. The user plane interface may be SR-4, RTC-4, or SR-4m media-centric interfacing.
[0195] In some embodiments, decoding uses at least one video codec selected from a list of video codecs, consisting of the H.264 video codec standard, the H.265 video codec standard, the H.266 video codec standard, and the AV1 video codec standard. Several other video codec standards that partially rely on the aforementioned standards, such as OMAF, V-PCC, or similar, may also be applied. OMAF encodes omnidirectional video content and comprises at least one video stream encoded using AVC and HEVC, while V-PCC encodes 2D projected 3D video and comprises a video stream that encodes projected 2D flat frames using AVC and HEVC.
[0196] In some embodiments, the video codec is an H.264, H.265, or H.266 codec, and non-video coded metadata is encapsulated as a user data payload in one or more supplemental enhancement information "SEI" messages of type "User Data Unregistered", and / or the video codec is an AV1 codec, and non-video coded metadata is encapsulated as a user data payload in one or more metadata open bitstream units "OBU" of type "Unregistered User Private Data". For SEI messaging, the payload type may be 5, and the SEI message may be prepended or prepended to the NALU. For OBU, the OBU metadata_type may be "X", where X can be any of 6 to 31.
[0197] In some embodiments, one or more data units of multimedia immersion and interaction data comprises immersion and interaction data selected from a list of augmented reality object representations, which include at least one of user viewpoint data, user field of view data, user posture / orientation data, user gesture tracking data, user body tracking data, user facial feature tracking data, user action and / or user input data, segmented rendering posture and spatial information, and graphical descriptions of objects and object position anchors.
[0198] In some embodiments, one or more data units of multimedia immersion and interaction data comprise augmented reality "XR" multimedia immersion and interaction data.
[0199] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources selected from a list of media sources, each consisting of one or more physical dedicated controllers, one or more red-green-blue "RGB" cameras, one or more RGB depth "RGBD" cameras, one or more infrared "IR" cameras, one or more microphones, and one or more haptic transducers.
[0200] For the purpose of late-stage reprojection in a split-rendering client "SRC" for displaying one or more rendered frame alignments, up-to-date user pose or orientation information, some embodiments include using one or more data units of multimedia immersion and interaction data to split-render a video media flow.
[0201] Some embodiments include receiving a media configuration at the start of a multimedia split-rendered content delivery session and / or upon updating the multimedia split-rendered content delivery session.
[0202] In some embodiments, the media configuration is an application configuration of an application hosted on the device, comprising one or more media capabilities of the device, and / or at least one of an application service provider "ASP" configuration provisioned to at least one of a real-time communication application function "RTC AF", a provisioning function, a segmented rendering application function "SR AF", and / or a configuration function.
[0203] In some embodiments, the application configuration may be pre-configured with supported UUIDs, and / or the transport may be automatically enabled when a UUID is received.
[0204] In some embodiments, the first device is a network entity / node, and the device is a UE.
[0205] Embodiments of the transport of interactive and immersive user data via a video base stream as non-video-encoded metadata are described in more detail below, with reference to the video source of the XR application. In the following sections, the terms interactive and immersive user data, interactive and immersive metadata, non-video-encoded user data, non-video-encoded embedded user data, non-video-encoded embedded metadata, or simply metadata are used interchangeably for this purpose.
[0206] Figure 13a provides a diagram of one embodiment 1300 of a split rendering architecture for interactive and immersive multimedia applications supported by split rendering at the edge. Embodiment 1300 shows that UE1310, EN1320, and DN1330 interface through various interfaces.
[0207] UE1310 is illustrated as comprising a split-rendering-aware application client 1311, an MSH1312 which may include an EEC1312a, a split-rendering client (SRC) 1313, and an XR runtime 1314. The split-rendering-aware application client 1311 interfaces with the XR runtime 1314. The split-rendering-aware application client 1311 interfaces with the SRC1313, for example, via SR-7. The split-rendering-aware application client 1311 may interface with the MSH1312, for example, via SR-6, and may interface with the EEC1312a via EDGE-5. The MSH1312 interfaces with the SRC1313, for example, via SR-7 and SR-6. The SRC1313 interfaces with the XR runtime 1314.
[0208] EN1320 is illustrated as comprising an SR AF1321 having configuration function 1321a, provisioning function 1321b, and EES1321c. The EES may be distributed across multiple instances 1321c that interface with each other, for example, via EDGE-9. EN1320 is further illustrated as comprising an RTC AS1322 having a signaling server 1322a and an SRS1322b which may have an EAS. SR AF1321 interfaces with RTC AS1322, for example, via SR-3. The EES1321c of SR AF1321 interfaces with SRS1322b which has an EAS, via EDGE-3.
[0209] The DN1330 is illustrated as having an ASP1331. The ASP1331 interfaces with the provisioning function 1321b of the SR AF1321 via SR-1. The ASP1331 interfaces with the SRS1322b, which has an EAS, via SR-2, for example.
[0210] The UE1310's segmented rendering-aware application client 1311 interfaces with the DN1330's ASP1331, for example, via SR-8. The UE1310's MSH1312 interfaces with the EN1320's SRAF1321, for example, via SR-5. The MSH1312's EEC1312a may further interface with the SRAF1321's EES1321c via EDGE1 and / or EDGE4 (for example, via ECS). The UE1310's SRC1313 interfaces with the RTC AS1322's SRS1322b and signaling server 1322a in the EN1320, for example, via SR-4 (for example, equipped with SR-4m, SR-4s).
[0211] Embodiment 1300, where Figure 13 is a split-rendering architecture for an interactive and immersive multimedia application, illustrates the architecture, functionality, and interface. The relevant split-rendering flow (e.g., provisioning, split management, and media delivery) is described below with respect to Figure 13b.
[0212] Figure 13b provides a diagram of a high-level partitioned rendering flow in one embodiment. The partitioned rendering-aware application 1311, MSH 1312, SRC 1313 (illustrated as comprising a scene manager, XR source manager, and media access functions), SRS 1322b, SR AF 1321, application service provider 1331, and XR runtime 1314 are shown in the figure.
[0213] The main steps and associated interfaces for segmented rendering are described below, thereby the segmented rendering application function (SR AF) is, in some embodiments, an instance of the comprehensive real-time communication application function (RTC AF) (see 3GPP Technical Specification 26.506 v1.1.0 entitled “5G Real-time Media Communication Architecture”).
[0214] In the first step 1301, ASP1331 provisions SR AF1321 via a provisioning request for a split rendering management session. Provisioning may, in some embodiments, be performed via an SR-1 reference point or, alternatively, via an SR-1 API. SR AF-1321 determines the split rendering configuration, partly based on the information received from ASP1331. In some embodiments, the edge-enabled SR AF1321 may additionally support an Edge Enabler Server (EES). In some implementations, the EES logic may be distributed across multiple edge DNs, thereby coordinating the context and coordination of the distributed EES instances via the EDGE-9 interface. The EES may further communicate with and register with the EAS split rendering function. In some embodiments, the split rendering server 1322b may represent an instantiation of the EAS for split rendering and may communicate with SR AF1321 via the EDGE-3 interface. This may imply, for example, registration / deregistration of EAS instances, access network capability exposure, and QoS management notifications and reports (e.g., bitrate compliance notifications). Step 1301 is illustrated as “Split Rendering Provisioning”.
[0215] In a further step 1302, the split-rendering-aware application 1311 obtains Service Access Information (SAI) from ASP 1331. SAI is a set of parameters and addresses required by the client to activate the reception of one or more DL / UL media sessions, perform dynamic policy calls, consumption / metric reporting, and request SR AF 1321 assistance. The acquisition of SAI may be performed in some applications by control plane signaling from SR AF 1321 to MSH 1312 via SR-5 (or similarly RTC-5), followed by exposure of the acquired SAI to application 1311 by MSH 1312 via the SR-6 (or similarly RTC-6) interface. In some embodiments, the SR-5 interface may include an EDGE-1 function that assists edge enabler clients instantiated by MSH1312 along with registration with the EES in SR AF1321, and an EDGE-4 function that allows the EEC to access the configuration of edge resources from the Edge Configuration Server (ECS) (e.g., EAS configuration information), respectively. In other embodiments, the SAI may be directly exposed by ASP1331 via an application-specific SR-8 (or similar RTC-8) interface (to be acquired after an SR-1 provisioning call). This step 1302 is illustrated as “Acquiring Service Access Information (SAI)”.
[0216] The partition management procedure follows. This can be triggered in some embodiments by the client, i.e., by the partition rendering-aware application 1311, as illustrated in step 1303a, and in some other embodiments by the ASP 1331, as illustrated in step 1303b.
[0217] In a first embodiment where partition management is triggered by the client, step 1303a proceeds with application 1311 requesting partition rendering assistance from UE SRC1313 via the SR-7 API. SRC1313 communicates with MSH1312 via the SR-6 API to determine the available SRS instances as well as the client media capabilities and capabilities (e.g., XR runtime API, XR runtime rendering capability, etc.) and the SAI configuration. SRC1313 then negotiates a media-centric partition configuration with SRS1322b at the edge. In some embodiments, this may imply partition management negotiations via an SDP procedure or dedicated partition rendering management signaling (e.g., as an add-on to WebRTC-specific signaling). For example, in the case of SDP-based signaling, the split is negotiated between SRC1313 and SRS1322b based on the SDP protocol, thereby using the SDP Proposal / Response Procedure to determine the media streams, media formats, and multiplexing supported by SRC1313 and SRS1322b during the split rendering session. After the split management negotiation, SRS1322b acknowledges and responds to SRC1313 with the split configuration via the user plane interface SR-4 (or similarly RTC-4), and the split rendering media delivery session is ready to begin. As a result, SRC1313 completes the application request for split rendering via the SR-7 interface. In this first embodiment, step 1303a is illustrated as “Client-Driven Split Management”.
[0218] In a second embodiment in which partition management is triggered by SRS1322b, step 1303b proceeds with SRS1322b requesting a partition rendering session from SRC1313 via the SR-4 (or RTC-4) interface. In some embodiments, this may be achieved in part via a signaling server function (i.e., carried out by a dedicated signaling server or similar), and in other embodiments, this may be left to SDP media session negotiation based on a proposal / response procedure. Upon receiving the request, SRC1313 queries MSH1312 for UE media capabilities, assuming the provisioned partition rendering configuration and available SAIs. MSH1312 responds to SRC1313 via SR-6 with available media capabilities and information. SRC1313 responds to the SRS1322b request via the SR-4 (or RTC-4) interface and negotiates media-centric partition management. Once an agreement is reached in the negotiations, the determined media configuration, media format, and multiplexing associated with the split configuration profile are returned to SRS1322b as a recognized split rendering configuration. SRC1313 then notifies the split rendering-aware application via the SR-7 interface (for example, by WebSocket or a similar asynchronous mechanism) of the split rendering configuration and the fact that the split rendering session is ready to start. In this second embodiment, step 1303b is illustrated as "network-driven split management".
[0219] A further step 1304 comprises media distribution. In some embodiments, this includes the establishment of a media distribution split rendering session by SRC1313. In such cases, SRC1313 may request SRS1322b to create a split rendering session based on the split rendering configuration profile determined from a previous step. In some embodiments, SRC1313 is comprised of at least a scene manager shim function, an XR source manager, and a media access function. Such SRC1313 may thus trigger a request to SRS1322b to create a split rendering session, either by the scene manager function or, alternatively, by the MAF function. SRS1322b responds to the request with a recognized response using the split rendering description of the split rendering configuration profile for the session. SRC1313 establishes a connection to SRS1322b and the split rendering session is initiated. In some embodiments, this may involve applying additional media-centric signaling via SR-4s as a signaling subinterface of SR-4, assuming a signaling server (e.g., WebRTC via WebSockets, or other signaling mechanism). Once a session is established, media is served as configured between SRC1313 and SRS1322b via SR-4 in both the UL and DL directions. This step 1304 is illustrated as “Media Delivery Setup”.
[0220] A further step 1305, involving a rendering loop, follows. A typical description of the rendering loop for split rendering, in some embodiments, includes SRC1313 receiving or acquiring pose information and user actions from the XR runtime. In other embodiments, the latter information is complemented by UL video and / or audio streams (for example, for AR conferencing or AR immersive and interactive applications). The pose, user actions, and any media streams are transmitted in the UL to SRS1322b via SR-4m (i.e., the SR-4's user plane media-centric subinterface). SRS1322b processes the received information, the information in its own service buffer, and renders at least partially the content of the next one or more frames for a split rendering-aware application. In one embodiment, SRS1322b may merge one or more poses and user actions together to estimate the pose information as close as possible to the expected display time of the frame. In such an embodiment, this estimation is used to render or pre-render the next frame as an alternative. In other embodiments, SRS1322b may select the most recent pose information for the expected display time of the next frame for rendering / pre-rendering operations. SRS1322b then transmits the rendered frame to SRC1313 via SR-4m. In some embodiments, the rendered frame may additionally include interaction and immersion metadata (e.g., pose information, user actions / gestures) used by SRS1322b to render the frame. This data is transmitted along with the video-encoded rendered frame as part of the relevant video base stream for the split rendering session. In one embodiment, this interaction and immersion metadata may be embedded in the video base stream as user-specific metadata, for example, as SEI messages for H.264, H.265, H.266, or alternatively, as OBU metadata for AV1.The media stream is then transmitted to the SRC1313 by the SRS1322b. The MAF function and video codec decode the media stream and expose the user data embedded in the base stream to the XR runtime 1314, or alternatively to the split rendering-aware application 1311, via the SR-7 interface. In one embodiment, the XR runtime 1314 displays the rendered frames and uses the embedded interaction and immersion metadata for late-stage reprojection and asynchronous time warping when correcting any errors between the current pose information to the XR runtime 1314 at display time and the pose and user action information estimated / used by the SRS1322b during the rendering operation.
[0221] In some embodiments, the payload of interaction and immersion metadata being communicated may contain at least two fields. In a first embodiment, the first field acts as a unique type identifier, i.e., a UUID, which determines the syntax and semantics of the corresponding information payload located in the second field. In a second embodiment, the second field carries the interaction and immersion data, data embedded as metadata within a video-encoded base stream. In further embodiments, the syntax and semantics of the second field may be determined at least in part based on the first field. This may imply index-based lookup (e.g., in a list of supported formats for interaction and immersion metadata by a specific communication endpoint such as a UE or alternatively an AS), data repository or alternatively registry lookup (e.g., querying an internet-based registry for one or more formats for interaction and immersion metadata), or selection of a pre-configured resource (e.g., application-determined format selection for interaction and immersion metadata).
[0222] Exemplary implementations of interactive and immersive metadata for various video codecs (e.g., H.264 / H.265 / H.266) base streams are outlined at a high level in Figures 14 and 15.
[0223] Figure 14 shows a representation of multimedia interaction and immersion user data 1400 as metadata within a video-encoded base stream for the MPEG H-26x family of video codecs. A NAL header 1411 and a NAL payload (raw bytes) 1412 are illustrated for a given NAL unit 1410. The NAL payload 1412 comprises a NAL SEI raw byte sequence payload 1413 which itself comprises a first SEI message 1413a and a second SEI message 1413b. The first SEI message 1413a is illustrated as "UUID=76994094-c7bd-436b-ac8e-c5205da905cc" and "SEI message" and "interaction and immersion data". The second SEI message 1413b is illustrated as "UUID=8b6d5df6-be48-40f1-814d-20ae061d078d", "SEI message", and "Interaction and immersion data".
[0224] Figure 15 shows a representation of multimedia interaction and immersion user data as metadata within the video encoding base stream for the AV1 video codec. An OBU header 1511 and an OBU payload 1512 are shown relative to a given OBU 1510 as a metadata OBU, the OBU payload 1512 having a UUID illustrated as "UUID=76994094-c7bd-436b-ac8e-c5205da905cc", and further illustrated as "Interaction and Immersion Data".
[0225] In some embodiments, for example, a video decoder in a MAF within an SRC located in the UE for DL split-rendering traffic, or alternatively, an ASP located in an SRS for UL split-rendering traffic, may pass user data metadata (containing interaction and immersion data) embedded within a video base stream (e.g., H.264, H.265, H.266, AV1, or similar). In further embodiments, the video decoder may expose user data to other functional blocks (e.g., XR runtime, split-rendering-aware application, SRS). This may be done based on a functional hook, handler, or callback, or alternatively, another dedicated interface (e.g., raw buffer, event bus, etc.), which in one embodiment exposes the interaction and immersion metadata payload, partly based on a set of configured UUID filters that extract relevant payload information. In one example, the SRC is configured for a split-rendering session using a media format via an H.264 video base stream that stores interaction and immersion data of type UUID=76994094-c7bd-436b-ac8e-c5205da905cc. This SRC filters and exposes the corresponding user data SEI messages that have been passed through by the MAF H.264 decoding instance. In one example, this may be done by the SRC via the SR-7 interface, or alternatively, via an API for any registered consumer, such as a split-rendering-aware application, or alternatively, via the XR runtime. This may occur in some embodiments if at least one of one or more UUIDs is configured and the functionality of interaction and immersion metadata transport via the video-encoded base stream is enabled.
[0226] Signaling of UUIDs and associated metadata formats, used by split rendering and transported via a media-centric user plane, is therefore required to enable the filtering and output of interactive and immersive user data corresponding to metadata from the video base stream.
[0227] A UUID may be schematically represented as a reference to an identifier related to the determination of the interaction and immersion user data type and payload format. Unless explicitly specified by convention, a UUID shall be treated herein as a general identifier.
[0228] Several embodiments are described below with reference to the signaling used. In detail, application-based signaling embodiments, network-supported signaling embodiments, and SDP-based signaling embodiments are described. The first instance describes application-based signaling.
[0229] In one embodiment, the ASP may signal to a split-rendering-aware application to enable the transport of interactive and immersive user data as metadata via one or more video base streams. The media streams are part of user-plane media-centric traffic related to DL split-rendering video traffic (e.g., split-rendered video streams with single or dual-eye buffering for an XR application) or, alternatively, UL video traffic (e.g., one or more UL video streams related to an AR interactive and immersive application).
[0230] In one example, an application may activate this functionality based on several application-specific configurations (e.g., JSON / YAML / XML media format descriptions and configurations) shared between the server, i.e., the ASP application infrastructure, and the client, the split-rendering-aware application. In one example, such configurations may contain options such as interaction-metadata-over-video-es=true to mark the enabled state for the transport of interaction and immersive user data as metadata over a video-encoded base stream. Furthermore, the interface used to communicate this configuration may, in some examples, comprise a proprietary implementation or signaling protocol served over a trusted communication channel such as the SR-8 interface outlined in Figures 13a and 13b. Some exemplary implementations may use WebSockets, HTTP, SCTP, or other trusted, recognized, and responsive messaging protocols to communicate this information.
[0231] In another embodiment, the configuration may further store a list of one or more UUIDs supported by the application, each corresponding to an interactive and immersive user data payload format. In some embodiments, these formats may depend on the XR runtime used by the UE, i.e., corresponding to the SRC in a split rendering setup. For example, this would be the case between the OpenXR hand tracking extension format and the Microsoft® HoloLens 2 specific hand tracking format. In other embodiments, the format may be encoded to correspond to one or more XR runtimes, e.g., a set of OpenXR abstract formats valid for one or more devices. In other embodiments, the UUIDs may correspond to one or more types of interactive and immersive user data payloads. For example, in some examples, a UUID may correspond to simple pose information corresponding to user head tracking, e.g., OpenXR's XrPosef, and another UUID may correspond to a set of complex pose, location, and velocity, each corresponding to a wrist joint related to user hand tracking, e.g., OpenXR's XR_EXT_hand_tracking.
[0232] In other embodiments, the application logic may be static with respect to the UUIDs and XR runtime supported by the application. In such embodiments, the split-rendering-aware application is configured using a static configuration of a supported list of UUIDs.
[0233] Furthermore, in another embodiment, the ability to transport interactive and immersive user data for a segmented rendering application as metadata via a video-encoded base stream is automatically enabled when at least one UUID is provided for filtering.
[0234] A segmented rendering-aware application, enabled and configured using interaction and immersion user data UUIDs, may, in some embodiments, further share its configuration with the SRC. The SRC may then apply the received configuration to filter the corresponding payloads of interaction and immersion user data by UUID after video decoding. These data payloads may be further exposed through other interfaces. In one example, the application may use the SR-7 API to configure the SRC with appropriate configurations for interaction and immersion user data UUIDs or, alternatively, identifiers.
[0235] In another embodiment relating to UL AR segmented rendering traffic, the SRS may request configuration information from the SR AF regarding the enablement of transport of interactive and immersive user data as metadata over the video base stream, as well as the corresponding UUIDs of such metadata types and payloads, which may be directly indicated by the ASP, for example, via the SR-1 interface.
[0236] In one embodiment, ASP signaling to an application of configuration information regarding the enablement of transport of interactive and immersive user data as metadata over a video base stream, as well as the corresponding UUIDs of such metadata types and payloads, may additionally include a mapping of UUIDs to the video base stream. In one example, this may be done by application-specific stream mapping (for example, based on an application media flow identifier), and in another example, this may be done by a 5-tuple information of available media flows (src addr, dst addr, srd port, dst port, protocol identifier).
[0237] Embodiments relating to network-supported signaling are described below. In one embodiment, one or more application functions may signal to the application to enable the transport of interaction and immersive user data as metadata over one or more video base streams. In a further embodiment, the signaling of this configuration may further include the UUID determining the type and payload format of the interaction and immersive user data to be embedded as metadata over the video base streams. In some embodiments, the configuration includes at least the following information, namely, a Boolean flag indicating the enablement of transport of interaction and immersion user data as metadata over one or more video base streams (for example, interaction-metadata-over-video-es=true indicates that the feature is enabled, and interaction-metadata-over-video-es=false indicates that the feature is disabled), one or more UUIDs or alternatively a list of identifiers determining the type and format of the payload of interaction and immersion user data to be transported over the video-encoded base stream (for example, viewport pose tracking information (in radians with respect to the horizontal and vertical axes as FoV), user pose information (for example, floating-point representations of 3D vectors for orientation and floating-point representations of 4D quaternions), and user hand tracking information (for example, objects representing two hands and a set of corresponding wrist joint locations, velocity vectors), e.g., UUID=[ea17c3e6-9be4-4d8d-98bf-b5772b273307, The information is transported via a container comprising [afcd396b-9c96-4c8c-b453-85200febfc09, c32dfb22-86fa-46b5-a584-6603a2d7a556]), thereby each piece of information having, in addition, their corresponding XR spatial references and timestamps to support XR spatial and temporal synchronization with the XR runtime.
[0238] In one exemplary implementation, such a container may be indicated by the SR AF from the split rendering edge DN to the MSH on the UE via the SR-5 interface. In such a case, this information is used to partially indicate the media configuration (e.g., the video codec used, the interaction and immersion type and format of metadata embedded in the video-encoded base stream). In another example, the SR-5 may have features specific to UL or DL media streaming corresponding to the M5d or M5u interface of the 5GMS architecture, and in yet another example, the SR-5 may have features specific to the RTC-5 interface corresponding to the 5GS Real-Time Communications (RTC) architecture.
[0239] In one embodiment, the MSH, which stores configuration information regarding the enablement of transport of interactive and immersive user data as metadata over one or more video base streams, and the corresponding UUID of such metadata, may expose this configuration to the SRC via an interface. The SRC, in turn, may request this information when starting or updating a segmented rendering media delivery session for proper processing, and may further expose the interactive and immersive user data to other functional blocks. In one example, this communication may occur via an SR-6 API that exposes the MSH configuration information regarding the transport and format of interactive and immersive metadata to the SRC. In some examples, such an interface may be implemented by a typical HTTP method, e.g., GET, or any other request-response based protocol.
[0240] In another embodiment relating to UL AR segmented rendering traffic, the SRS may request configuration information regarding the enablement of transport of interaction and immersive user data as metadata over the video base stream, as well as the corresponding UUIDs of such metadata types and payloads, from the SR AF via an available API, for example, SR-3.
[0241] In one embodiment, SR AF signaling to MSH of configuration information regarding the enablement of transport of interactive and immersive user data as metadata via the video base stream, as well as the corresponding UUIDs of such metadata types and payloads, may additionally include a mapping of UUIDs to the video base stream. In one example, this may be done by 5-tuple information of the available media flow (src addr, dst addr, srd port, dst port, protocol identifier).
[0242] Embodiments related to SDP-based signaling are described below.
[0243] In one embodiment, a split-rendering server may signal to an application a list of one or more UUIDs corresponding to the transport activation of one or more types and formats of interactive and immersive user data as metadata across one or more video base streams. The split-rendering server signals this configuration information for this purpose for each video base stream, indicating which UUIDs are supported by each video-encoded base stream. Furthermore, the split-rendering server signals the latter in-band to its corresponding split-rendering client via a media-centric user plane.
[0244] In one embodiment, the protocol used to signal this configuration is SDP, and the signaling of supported UUIDs and media negotiation between the server and client is based on an SDP proposal / response procedure. The SDP proposal / response is performed before the establishment or, as an alternative, the update of the split-rendered content delivery media session. In some embodiments, the SRS provides the SDP proposal to the SRC. The SRC parses the proposal, identifies the SDP attributes related to the interaction and immersion user data, and determines, based on the parsed values, whether the latter is supported. This may be based in part on information obtained from the MSH before the media session establishment regarding the UE media capabilities and / or the split-rendered-aware application media capabilities and configuration. To confirm support, the SRC copies the supported UUID configuration corresponding to the interaction and immersion user data types and formats into the SDP response to the SRS. This configuration is then applied to the DL traffic of the split-rendered content delivery session.
[0245] In some other embodiments, where UL segmented rendering traffic may thereby include a video stream as an additional element (e.g., an AR application), the SDP proposal / response procedure is round-trip. For this purpose, the SRC provides the SDP proposal to the SRS, along with supported UUIDs for interaction and immersion metadata embedded in the video base stream. The SRS processes the SDP proposal and, in turn, provides its response in the SDP response. If the SRS accepts the SDP UUID attribute proposal, it copies these into the corresponding SDP response which is sent back to the SRC. The segmented rendering media delivery session begins with the determined SDP proposal / response negotiation result, in which the supported UUID types and corresponding payloads are embedded in the video-encoded base stream within the UL as interaction and immersion metadata.
[0246] In one example, the SDP proposal / response procedure in DL or UL as an alternative between the SRC and SRS is performed via the SR-4 or RTC-4 user plane interface. In another example, the SDP proposal / response procedure in DL or UL as an alternative between the SRC and SRS is performed via the SR-4m media-centric user plane interface via 5GS. In yet another example, the SDP proposal / response procedure in DL or UL as an alternative between the SRC and a dedicated signaling server interfaced with the SRS is performed via the SR-4s signaling-centric user plane interface via 5GS.
[0247] In one embodiment, 3GPP-specific SDP attributes may be used via the 5GS implementation form to convey a UUID that determines the type and format of the interaction and immersion metadata payload via the video base stream as an SDP attribute. For this purpose, one or more SDP attributes may be used to enumerate one or more UUIDs for the video stream. In one example, the UUIDs may be assigned to two or more video media streams. In one example, the SDP-derived markings of such UUIDs are enumerated in the Extended Backus-Naur (ABNF) format as per RFC4122, and are repeated below for completeness. interaction-metadata-uuid=''a=3gpp-interaction-metadata-uuid:'' uuid CRLF uuid=field-1 ''-'' field-2 ''-'' field-3 ''-'' field-4 ''-'' field-5 field-1=4hexOctet Field - 2 = 2hex Octet field-3=2hexOctet field-4=2hexOctet field-5=6hexOctet hexOctet=hexDigit hexDigit hexDigit=''0'' / ''1'' / ''2'' / ''3'' / ''4'' / ''5'' / ''6'' / ''7'' / ''8'' / ''9'' / ``a'' / ``b'' / ``c'' / ``d'' / ``e'' / ``f'' / ``A'' / ``B'' / ``C'' / ``D'' / ``E'' / ``F''
[0248] The marking resulting from the SDP of a UUID as illustrated above may correspond to an SDP attribute of the form, for example, "a=3gpp-interaction-metadata-uuid:20354d7a-e4fe-47af-8ff6-187bca92f3f9". Therefore, the UUID format may conform to ISO-IEC-11578 ANNEX A or, alternatively, the ISO-IEC9834-8 format in some implementations. The UUID fields listed in the above ABNF description may, in one example, be derived as UUIDs of ISO-IEC9834-8 version 4, i.e., randomly.
[0249] In another example, a valid SDP enumeration corresponding to either an SDP proposal or a response is provided below. m=video 49230 RTP / AVP 96 a=rtpmap:96 H.264 / 90000 a=3gpp-interaction-metadata-uuid: ea17c3e6-9be4-4d8d-98bf-b5772b273307 / / Viewport a=3gpp-interaction-metadata-uuid: afcd396b-9c96-4c8c-b453-85200febfc09 / / User posture a=3gpp-interaction-metadata-uuid: c32dfb22-86fa-46b5-a584-6603a2d7a556 / / Hand tracking m=video 49231 RTP / AVP 97 a=rtpmap:97 H.265 / 90000 a=3gpp-interaction-metadata-uuid: ea17c3e6-9be4-4d8d-98bf-b5772b273307 / / Viewport a=3gpp-interaction-metadata-uuid: afcd396b-9c96-4c8c-b453-85200febfc09 / / User posture
[0250] In the example above, the UUIDs shown are determined for each media stream. Therefore, the H.264 media stream on port 49230 embeds UUIDs for viewport pose tracking information, user pose information, and user hand tracking information within the video base stream, while the H.265 media stream on port 49231 embeds only UUIDs for viewport pose tracking information and user pose information within the video base stream.
[0251] In one embodiment, once the SDP Proposal / Response configuration determines the UUID and associated interaction and immersion metadata type and format for the media stream, the SRC, or alternatively the SRS, processes the decrypted user data and exposes it to other functional blocks (e.g., XR runtime, split-rendering-aware application, SRS media processing function).
[0252] The disclosure herein proposes the necessary configuration signaling to enable XR application split rendering, along with the transport of interactive and immersive user data over the video base stream as SEI messages (H.264, H.265, H.266), or alternatively as OBU metadata (AV1). The solution supports multiple types of metadata based on UUIDs that identify the type and format of each metadata payload to be carried over the video base stream. The proposed signaling mechanism for media configuration is based on three methods: application layer signaling from the ASP to the application, followed by updating the media configuration to the split rendering client for the split rendering media delivery session; control plane signaling using the split rendering application functionality to indicate the media configuration to the media session handler, followed by updating the media configuration to the split rendering client for the split rendering media delivery session; and user plane signaling by SDP when establishing a split rendering media delivery session based on a new SDP attribute, where the new SDP attribute can configure a list of UUIDs that identify the metadata carried over the video base stream, per video media flow.
[0253] The problem addressed by this disclosure is media configuration signaling for real-time transport of interactive and immersive multimedia data for segmented rendering of interactive and immersive XR applications. User interaction and immersion data (e.g., pose information, FoV tracking, user actions) serve as input to segmented rendering. Furthermore, for efficient display processing in the UE, the XR runtime requires metadata information about the user pose and input used by the segmented rendering server to render the frames. This is required for late-stage reprojection. For this purpose, a real-time transport mechanism for rendered interaction and immersion metadata and media configuration signaling is needed for efficient display of segmented-rendered frames.
[0254] The present invention solves the problem by utilizing SEI messages and OBU metadata as a transport mechanism for interactive and immersive metadata in a split-rendering architecture. The transport relies on the video-encoded base stream grouping rendered frames with their associated pose information used for rendering at the edges. To assist the split-rendering client, the media configuration needs to include signaling that identifies the type of metadata carried over the video base stream, as well as its mapping to the video media flow (e.g., a 5-tuple). Three methods are proposed for signaling: i) an application signaling path, ii) a control plane signaling path for a real-time communication system AF, and iii) user plane signaling based on SDP proposal / response during RTP session establishment.
[0255] The proposed solution surpasses data channel approaches using the WebRTC SCTP stack because it benefits from the advantages of RTP / SRTP, namely timing, synchronization, jitter management, and reliability based on FEC. Furthermore, the proposed method essentially synchronizes the interaction and immersion metadata with one or more video streams used for piggybacking. Any kind of signaling is outside the scope of WebRTC.
[0256] The proposed solution is superior to RTP header extension techniques because it does not limit the maximum payload size for interactive and immersive multimedia data types. Furthermore, the proposed solution is transported as part of the RTP payload and thus benefits from RTP synchronization, jitter management, and reliability provided by FEC. RTP header extension signaling relies exclusively on the SDP proposal / response procedure. The SDP is the candidate signaling in the proposed disclosure, and introduces a new SDP attribute that is orthogonal to any RTP header extension-related signaling solution.
[0257] The proposed solution involves trade-offs regarding the method for defining new IETF RTP payloads for interactive and immersive multimedia data. The trade-offs primarily target the avoidance of the need to provide a complete RTP payload type specification.
[0258] More specifically, this disclosure provides an embodiment that utilizes application-based signaling. Signaling of renderer interactions and immersive metadata, as well as subsequent mapping to the video media flow, is performed by an application service provider that informs the application client of the media configuration. The application client then modifies the split rendering client and media session handler media configuration accordingly via APIs exposed in 5GS of the latter two functional blocks.
[0259] Further embodiments utilize network-supported signaling. Signaling of renderer interactions and immersive metadata, as well as subsequent mapping to the video media flow, is performed by a split-rendering AF that informs the media session handler of the media configuration. The media session handler exposes this configuration to application clients via its API, but also to split-rendering clients.
[0260] Further embodiments utilize SDP-based signaling. Signaling of renderer interaction and immersive metadata, as well as subsequent mapping to the video media flow, is performed by an SDP proposal / response procedure immediately before the establishment of the media session. A new media description attribute a=3gpp-interaction-metadata-uuid: is used to map metadata types (identified by UUIDs) to one or more RTP media streams. <uuid>is proposed.
[0261] From the perspective of a split rendering server, a method for configuring a multimedia split rendering content delivery session over a network is provided by the disclosure herein. The method includes determining a media configuration comprising mapping one or more interactive and immersive multimedia data types to one or more video media flows that store non-video encoded data from one or more media sources; signaling the determined media configuration to a remote endpoint; establishing a multimedia split rendering content delivery session configuration that includes at least one video media flow of at least one video encoded base stream that includes one or more mapped interactive and immersive multimedia data as a non-video encoded metadata payload, based in part on the signaled media configuration; and utilizing a non-video encoded metadata payload corresponding to one or more interactive and immersive multimedia data provided by at least one video encoded base stream for split rendering related to at least one video media flow of at least one video encoded base stream.
[0262] In some embodiments, the non-video encoded metadata payload comprises at least two fields.
[0263] In some embodiments, the first field indicates a syntax and semantic identifier representation format of the interactive and immersive multimedia data type, and the second field encodes the interactive and immersive multimedia data according to the identifier representation format determined by the first field.
[0264] In some embodiments, the first field comprises a Universal Unique Identifier (UUID) representation.
[0265] Some embodiments further comprise a configuration mapping that aligns to one or more identifier representation formats corresponding to one or more interactive and immersive multimedia data types to one or more video media flows.
[0266] Some embodiments further comprise signaling of media configuration being performed based on at least one of signaling executed by a split rendering aware application by an application service provider (ASP) via an application interface, signaling between a real-time communication application function (RTC AF) and a media session handler via a control plane interface (e.g., the RTC-5 reference interface in 5GS between RTC AF and MSH), and signaling between a split rendering server (SRS) and a split rendering client (SRC) via a user plane interface (e.g., SR-4 between SRS and SRC or alternatively the RTC-4 reference interface in 5GS).
[0267] Some embodiments further comprise determination of media configuration, based in part on at least one of an application service provider (ASP) configuration that provisions at least one of a real-time communication application function (RTC AF), a provisioning function, a split rendering application function (SR AF), and a configuration function, and an application configuration of media capabilities determined semi-statically and in part based on the hardware capabilities of the corresponding device that processes the logic of the application.
[0268] In some embodiments, at least one video encoded base stream is determined according to at least one of the H.264 video codec standard, the H. 265 video codec standard, the H.266 video codec standard, and the AV1 video codec standard.
[0269] In some embodiments, the non-video-encoded metadata payload provided within at least one video-encoded base stream includes at least one of the following: a Supplemental Enhancement Information (SEI) message of type Unregistered User Data, and a metadata open bitstream unit (OBU) of type Unregistered User Private Data.
[0270] In some embodiments, the interactive and immersive multimedia data type comprises one of the user pose data representations in XR space (e.g., as timestamped 3D position vectors and quaternion representations) describing up to six DoF pose object orientations. Such pose objects include user body parts or segments such as the head, joints, hands, or combinations thereof; user gesture tracking data representations (i.e., arrays of one or more hands tracked according to their poses, with each hand tracking additionally consisting of arrays of wrist joint locations and velocities against base XR space and XR runtime timestamps, for example, as per the OpenXR OpenXR_EXT_hand_tracking API specification); user body tracking data representations (e.g., BioVision Hierarchical (BVH) coding of body and segment movements and associated pose objects); user facial feature tracking data representations (e.g., arrays of keypoint / feature locations, poses, or their codings into predetermined facial expression classes); and a set of one or more user actions (i.e., OpenXR taking diverse user input to a controller or HW command unit supported by an OpenXR compliant AR / VR device). The XrAction handle may correspond to user actions and inputs to physical or logical controllers defined in the XR space, and a set of one or more augmented reality (AR) object representations, each object comprising at least one of a graphical description and associated position anchors (i.e., metadata that determines the position of an object or point in the XR space as anchors for positioning virtual 2D / 3D objects such as text renderings, static video content, and 2D / 3D video content).
[0271] Some embodiments further include split-rendering utilization of interactive and immersive data for at least one of the following: one or more rendered frame alignment; late-stage reprojection in a split-rendering client (SRC) for displaying the latest user pose and orientation information; partial or complete pre-rendering of the next one or more frames in a split-rendering server (SRS); and a split-rendering server (SRS) estimation of the most probable user pose and orientation related to the expected display time of the next one or more frames to be partially or completely pre-rendered.
[0272] In some embodiments, the mapping of interactive and immersive multimedia data types to media flows is based on at least one of a 5-tuple representation (for example, as a 5-tuple describing an IP flow, comprising (IP source address, IP destination address, source port, destination port, protocol identifier)) and a media flow identifier (for example, a media flow unique name, such as provided in a media description under an SDP, an SDP media description attribute, etc.).
[0273] The disclosure herein also provides a method for configuring a multimedia split-rendering content delivery session over a network, from the perspective of a split-rendering client, the method comprising: receiving over a network a media configuration comprising a mapping of one or more interactive and immersive multimedia data types to one or more video media flows, each of which stores non-video-coded data from one or more media sources; controlling a video decoder to generate a set of one or more information payloads corresponding to a non-video-coded payload, each of which thereby comprises at least two fields; and processing one or more information payloads as one or more samples of interactive and immersive multimedia data generated by one or more media sources.
[0274] In some embodiments, the non-video encoded metadata payload comprises at least two fields.
[0275] In some embodiments, the first field indicates the syntax and semantic identifier representation format of the interactive and immersive multimedia data type, and the second field encodes the interactive and immersive multimedia data according to the identifier representation format determined by the first field.
[0276] In some embodiments, the first field includes a Universally Unique Identifier (UUID) representation.
[0277] Some embodiments further include configuration mappings to one or more video media flows that are consistent with one or more identifier representation formats corresponding to one or more interactive and immersive multimedia data types.
[0278] Some embodiments further comprise the reception of media configuration based on at least one of the following: signaling performed by an Application Service Provider (ASP) to a split-rendering-aware application via an application interface; signaling between a Real-Time Communication Application Function (RTC AF) and a media session handler via a control plane interface (e.g., an RTC-5 reference interface in 5GS between the RTC AF and MSH); and signaling between a split-rendering server (SRS) and a split-rendering client (SRC) via a user plane interface (e.g., an SR-4 or alternatively, an RTC-4 reference interface in 5GS between the SRS and SRC).
[0279] Some embodiments further further include a media configuration partially based on an application configuration that provisions an application service provider (ASP) configuration to at least one of a real-time communication application function (RTC AF), a provisioning function, a segmented rendering application function (SR AF), and a configuration function, as well as an application configuration of media capabilities that are determined semi-statically and partially on the hardware capabilities of the corresponding device that processes the logic of the application.
[0280] In some embodiments, at least one video encoding base stream is determined according to at least one of the H.264 video codec standard, the H.265 video codec standard, the H.266 video codec standard, and the AV1 video codec standard.
[0281] In some embodiments, the non-video-encoded metadata payload provided within at least one video-encoded base stream includes at least one of the following: a Supplemental Enhancement Information (SEI) message of type Unregistered User Data, and a metadata open bitstream unit (OBU) of type Unregistered User Private Data.
[0282] In some embodiments, the interactive and immersive multimedia data type comprises one of the user pose data representations (for example, as a timestamped 3D position vector and quaternion representation in XR space describing the orientation of up to six DoF pose objects). Such posture objects include user body parts or segments (such as the head, joints, hands, or combinations thereof), user gesture tracking data representations (i.e., arrays of one or more hands tracked according to their postures, with each hand tracking additionally consisting of arrays of wrist joint locations and velocities relative to the base XR space and XR runtime timestamp, for example, as per the OpenXR_EXT_hand_tracking API specification), user body tracking data representations (e.g., BioVision hierarchical (BVH) coding of body and segment movements, as well as associated posture objects), user facial feature tracking data representations (e.g., arrays of keypoint / feature locations, their coding to postures or predetermined facial expression classes), and sets of one or more user actions (i.e., taking diverse user input to a controller or HW command unit supported by an OpenXR-compliant AR / VR device, OpenXR The XrAction handle may correspond to user actions and inputs to physical or logical controllers defined in the XR space, and a set of one or more augmented reality (AR) object representations, each object comprising at least one of a graphical description and associated position anchors (i.e., metadata that determines the position of an object or point in the XR space as anchors for positioning virtual 2D / 3D objects such as text renderings, static video content, and 2D / 3D video content).
[0283] Some embodiments further include split-rendering utilization of interactive and immersive data for at least one of the following: one or more rendered frame alignment; late-stage reprojection in a split-rendering client (SRC) for displaying the latest user pose and orientation information; partial or complete pre-rendering of the next one or more frames in a split-rendering server (SRS); and a split-rendering server (SRS) estimation of the most probable user pose and orientation related to the expected display time of the next one or more frames to be partially or completely pre-rendered.
[0284] In some embodiments, the mapping of interactive and immersive multimedia data types to media flows is based on at least one of a 5-tuple representation (for example, as a 5-tuple describing an IP flow, comprising (IP source address, IP destination address, source port, destination port, protocol identifier)) and a media flow identifier (for example, a media flow unique name, such as provided in a media description under an SDP, an SDP media description attribute, etc.).
[0285] This disclosure relates to the use of SEI message / OBU metadata for the transport of interaction and immersion metadata related to XR applications supported by segmented rendering and signaling required to enable such applications via 5GS.
[0286] Note that the above methods and apparatuses are illustrative rather than limiting the present invention, and that those skilled in the art can design many alternative configurations without departing from the scope of the appended claims. The term "comprising" does not exclude the presence of elements or steps other than those listed in the claims, "a" or "an" does not exclude a plurality, and the functions of several units described in the claims may be implemented by a single processor or other unit. Any reference signs in the claims should not be construed as limiting their scope.
[0287] Furthermore, although examples are given in the context of specific communication standards, these examples are not intended to limit the communication standards to which the disclosed methods and apparatuses may be applied. For example, although specific examples are given in the context of 3GPP, the principles disclosed herein can also be surely applied to other wireless communication systems, and any communication system using routing rules.
[0288] The method may also be embodied as a set of instructions stored on a computer-readable medium that, when loaded into a computer processor, a digital signal processor (DSP), or the like, causes the processor to execute the methods described above.
[0289] Where referenced, the OpenXR specification is described in git reference release 1.0.26, the RTP payload format media type is available in IANA standard RFC4855, and the SDP offer / answer model is described in RFC 3264 entitled “An offer / answer model with the Session Description Protocol (SDP)”. Furthermore, the following 3GPP technical specifications Tdocs, namely S4aR220052 entitled “Real-time metadata transport over RTP”, S4aR220053 entitled “Real-time metadata transport over data channel”, S4-221557 entitled “Real-time metadata transport over RTP”, S4-221555 entitled “Real-time metadata transport over RTP”, and S4-230359 entitled “Signaling the render pose and other related information”, are relevant to the disclosures herein.
[0290] The methods and apparatus described may be practiced in other specific forms. The methods and apparatus described should be considered merely illustrative and not limiting in any respect. Accordingly, the scope of the invention is indicated not by the above description but by the appended claims. All modifications that fall within the equivalent meaning and scope of the claims should be encompassed within those scopes.
[0291] The following abbreviations are used: 3GPP, Third Generation Partnership Project; 5G, Fifth Generation; 5GS, 5G System; 5QI, 5G QoS Identifier; AF, Application Function; AMF, Access and Mobility Function; AR, Augmented Reality; DL, Downlink; DTLS, Datagram Transport Layer Security; NAL, Network Abstraction Layer; NALU, NAL Unit; OBU, Open Bitstream Unit; PCF, Policy Control Function; PDU, Packet Data Unit; PPS, Picture Parameter Set; QoE, Experience Quality; QoS, Quality of Service; RAN, Radio Access Network; RTCP, Real-Time Control Protocol; RTP, Real-Time Protocol; SDAP, Service Data Adaptation Protocol; SEI, Supplemental Enhancement Information; SMF, Session Management Function; SRTCP, Secure Real-Time Control Protocol; SRTP, Secure Real-Time Protocol; SR AF, Split Rendering AF; SRC, Split Rendering Client; SRS, Split Rendering Server; TLS, Transport layer security; UE (User Equipment); UL (Uplink); UPF (User Plane Function); VCL (Video Coding Layer); VMAF (Video Multiplexing Method Assessment Function); VPS (Video Parameter Set); VR (Virtual Reality); WebRTC (Web Real-Time Communication); XR (Augmented Reality); XR AS (XR Application Server); and XRM (XR Media) are important in the fields addressed herein. [Explanation of symbols]
[0292] 100 Wireless Communication Systems 102 Remote Unit 102 UE 104 Network Units 200 User Equipment 205 Processor 210 memory 215 Input Devices 220 Output Devices 225 Transceiver 230 Transmitter 235 Receiver 240 network interfaces 245 Application Interfaces 300 network nodes 305 Processor 310 memory 315 Input Devices 320 Output Devices 325 Transceiver 330 Transmitter 335 Receiver 340 Network Interfaces 345 Application Interfaces 405 IP Layer 410 Media Session Data Plane 412 User Datagram Protocol (UDP) 414 RTCP 416 RTP 420 Media Codecs 422 Quality Control 450 Media Session Control Plane 452 UDP 454 Transmission Control Protocol (TCP) 462 Session Initiation Protocol (SIP) 464 Session Description Protocol (SDP) 505 IP Layer 510 Data Plane 512 UDP 515 SRTCP 517 SRTP 520 Media Codecs 522 Quality Control 524 Interactive Connectivity Establishment (ICE) 526 Datagram Transport Layer Security (DTLS) 528 SCTP 550 Control Plane 554 TCP 556 TLS 558 HTTP 562 SIP 564 SDP 568 SSE / XHR / etc. 570 XMPP / etc. 630 RTP packets Sequence number 642 644 Timestamp 646 Synchronization Source (SSRC) Identifier 648 Contributing Source (CSRC) Identifier 650 RTP Header Extensions 760 SRTP packets 772 Sequence number 774 Timestamp 776 Synchronization Source (SSRC) Identifier 778 Contributing Source (CSRC) Identifier 780 RTP Header Extension 901 Raw Input Video Frames 902 Video Encoded Bitstream 903 Recover video frames video 910 encoder 911 Picture Block Classification Function Block 912 Spatial transformation 913 Quantization 914 Entropy coding 915 Motion Estimation 920 Decoder, Buffer 921 Interpretation Block 922 buffers 923 Visual Filter Processing Block 925 Inverse Space Transformation Block 926 Inverse Quantization Block 927 Entropy Decoding 1000 NAL units 1010 NAL Header 1020 payload 1021 Video / Sequence / Picture Parameter Payload 1022 Supplemental Enhancement Information Payload 1022a SEI message 1023 Frame / Picture / Slice Payload 1023a Header 1023b Video Encoded Payload 1310 UE 1311 Split Rendering Aware Application Client 1312 MSH 1312a EEC 1313 Subdivision Rendering Client (SRC) 1314 XR Runtime 1320 EN 1321 SR AF 1321a Configuration Function 1321b Provisioning Function 1321c EES 1322 RTC AS 1322a Signaling Server 1322b SRS 1330 DN 1331 ASP 1400 Multimedia Interaction and Immersive User Data Representation 1410 NAL Unit 1411 NAL Header 1412 NAL payload 1413 NAL SEI Raw Byte Sequence Payload 1413a First SEI message 1413b Second SEI message 1500 Multimedia Interactions and Immersive User Data Representation 1510 OBU 1511 OBU Header 1512 OBU payload< / uuid>
Claims
1. A device for wireless communication, Processor and The device comprises a memory coupled to the processor, and the processor provides the device, Determining a media configuration for segmented rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating that the encoded video stream of the video media flow includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata. Signaling the aforementioned media configuration to the second device, Establishing a multimedia segmented rendering content distribution session with the aforementioned video media flow together with the second device, at least partially based on the media configuration, The system is configured to use one or more data units of multimedia immersion and interaction data to segment-render the aforementioned video media flow. Device.
2. The aforementioned non-video encoded metadata, A first field comprising identifiers for the syntax and semantic representation formats of one or more data units of multimedia immersion and interaction data, A second field comprising one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantic representation format corresponding to the identifier of the first field, The apparatus according to claim 1.
3. The apparatus according to claim 2, wherein the identifier is represented as a universally unique identifier "UUID".
4. The apparatus according to claim 3, wherein the UUID is specific to a particular application or session, or the UUID is globally unique.
5. The processor in the device The system is configured to determine the media configuration by causing the device to determine the media configuration for a plurality of video media flows, each having an encoded video stream and each having unencoded video metadata. The one or more parameters of the media configuration map the identifier of the first field of each of the non-video encoded metadata to their respective video media flows. The apparatus according to any one of claims 2 to 4.
6. The one or more parameters of the media configuration are Application-specific stream mapping, 5-tuple display, and Media flow description attribute Based on at least one of the above, the first field of each of the non-video encoded metadata is mapped to their respective video media flows. The apparatus according to claim 5.
7. The processor in the device Application interface, Control plane interface, and User plane interface The media configuration is configured to signal the media configuration via at least one of the following: The apparatus according to any one of claims 1 to 6.
8. The encoded video stream is H.264 video codec standard, H.265 video codec standard, H.266 video codec standard, and AV1 video codec standard Encoded / decoded using at least one video codec selected from a list of video codecs, The apparatus according to any one of claims 1 to 7.
9. The video codec comprises the H.264, H.265, or H.266 codec, and the non-video coded metadata is encapsulated as the user data payload in one or more supplemental enhancement information "SEI" messages of type "User Data Not Registered," and / or The video codec comprises the AV1 codec, and the non-video encoded metadata is encapsulated as a user data payload in one or more metadata open bitstream units (OBUs) of type "unregistered user private data". The apparatus according to claim 8.
10. The one or more data units of multimedia immersion and interaction data, User-perspective data, User field of view data, User posture / orientation data, User gesture tracking data, User body tracking data, User facial feature tracking data, User actions and / or user input data, Subdivision rendering pose and spatial information, as well as An augmented reality object representation comprising at least one of a graphical description of an object and an object position anchor. It features immersive and interactive data selected from a list consisting of, The apparatus according to any one of claims 1 to 9.
11. The processor in the device An application configuration for an application hosted on the second device, the application configuration having one or more media capabilities of the second device, and / or Real-time communication application function "RTC AF", Provisioning function, The segmented rendering application function "SR AF", and / or Configuration Function Application Service Provider (ASP) configuration provisioned to at least one of the following Further configured to determine the media configuration based on at least one of the following, The apparatus according to any one of claims 1 to 10.
12. A device for wireless communication, Processor and The device comprises a memory coupled to the processor, and the processor provides the device, Receiving a media configuration for segmented rendering of a video media flow from a first device, wherein the media configuration comprises one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata. The device is configured using the media configuration to receive the video media flow, Decoding the encoded video stream of the video media flow, comprising extracting the non-video encoded metadata, It is configured to consume one or more data units of multimedia immersion and interaction data from the non-video encoded metadata, Device.
13. The aforementioned non-video encoded metadata, A first field comprising identifiers for the syntax and semantic representation formats of one or more data units of multimedia immersion and interaction data, A second field comprising one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantic representation format corresponding to the identifier of the first field, The apparatus according to claim 12.
14. The apparatus according to claim 13, wherein the identifier is represented as a universally unique identifier "UUID".
15. The apparatus according to claim 14, wherein the UUID is specific to a particular application or session, or the UUID is globally unique.
16. The apparatus according to any one of claims 13 to 15, wherein the media configuration is for a plurality of video media flows, each comprising an encoded video stream and each non-video encoded metadata, and one or more parameters of the media configuration map the identifier of the first field of each non-video encoded metadata to each of those video media flows.
17. The one or more parameters of the media configuration are Application-specific stream mapping, 5-tuple display, and Media flow description attribute Based on at least one of the above, the first field of each of the non-video encoded metadata is mapped to their respective video media flows. The apparatus according to claim 16.
18. The processor in the device Application interface, Control plane interface, and User plane interface The media configuration is configured to be received via at least one of the following: The apparatus according to any one of claims 12 to 17.
19. The encoded video stream is H.264 video codec standard, H.265 video codec standard, H.266 video codec standard, and AV1 video codec standard Encoded / decoded using at least one video codec selected from a list of video codecs, The apparatus according to any one of claims 12 to 18.
20. The video codec comprises the H.264, H.265, or H.266 codec, and the non-video coded metadata is encapsulated as the user data payload in one or more supplemental enhancement information "SEI" messages of type "User Data Not Registered," and / or The video codec comprises the AV1 codec, and the non-video encoded metadata is encapsulated as a user data payload in one or more metadata open bitstream units (OBUs) of type "unregistered user private data". The apparatus according to claim 19.
21. The one or more data units of multimedia immersion and interaction data, User-perspective data, User field of view data, User posture / orientation data, User gesture tracking data, User body tracking data, User facial feature tracking data, User actions and / or user input data, Subdivision rendering pose and spatial information, as well as An augmented reality object representation comprising at least one of a graphical description of an object and an object position anchor. It features immersive and interactive data selected from a list consisting of, The apparatus according to any one of claims 12 to 20.
22. The aforementioned media configuration is An application configuration for an application hosted on a second device, the application configuration having one or more media capabilities of the second device, and / or Real-time communication application function "RTC AF", Provisioning function, The segmented rendering application function "SR AF", and / or Configuration Function Application Service Provider (ASP) configuration provisioned to at least one of the following Based on at least one of the following: The apparatus according to any one of claims 12 to 21.
23. A method for wireless communication, A step of determining a media configuration for segmented rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating that the encoded video stream of the video media flow includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata, The steps include signaling the media configuration to a second device, The steps include establishing a multimedia segmented rendering content distribution session, which includes the video media flow, together with the second device, at least partially based on the media configuration, The steps include using one or more data units of multimedia immersion and interaction data to segment-render the aforementioned video media flow, and A method that includes this.
24. The aforementioned non-video encoded metadata, A first field comprising identifiers for the syntax and semantic representation formats of one or more data units of multimedia immersion and interaction data, A second field comprising one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantic representation format corresponding to the identifier of the first field, The method according to claim 23.
25. The method according to claim 24, wherein the identifier is represented as a universally unique identifier "UUID".
26. The method according to claim 25, wherein the UUID is specific to a particular application or session, or the UUID is globally unique.
27. The step of determining the media configuration is The steps include determining the media configuration for a plurality of video media flows, each comprising an encoded video stream and each non-video encoded metadata, wherein one or more parameters of the media configuration map the identifier of the first field of each non-video encoded metadata to their respective video media flows. The method according to any one of claims 24 to 26.
28. The one or more parameters of the media configuration are Application-specific stream mapping, 5-tuple display, and Media flow description attribute Based on at least one of the above, the first field of each of the non-video encoded metadata is mapped to their respective video media flows. The method according to claim 27.
29. The step of signaling the media configuration is Application interface, Control plane interface, and User plane interface The step includes signaling via at least one of the following: The method according to any one of claims 23 to 28.
30. The encoded video stream is H.264 video codec standard, H.265 video codec standard, H.266 video codec standard, and AV1 video codec standard Encoded / decoded using at least one video codec selected from a list of video codecs, The method according to any one of claims 23 to 29.
31. The video codec comprises the H.264, H.265, or H.266 codec, and the non-video coded metadata is encapsulated as the user data payload in one or more supplemental enhancement information "SEI" messages of type "User Data Not Registered," and / or The video codec comprises the AV1 codec, and the non-video encoded metadata is encapsulated as a user data payload in one or more metadata open bitstream units (OBUs) of type "unregistered user private data". The method according to claim 30.
32. The one or more data units of multimedia immersion and interaction data, User-perspective data, User field of view data, User posture / orientation data, User gesture tracking data, User body tracking data, User facial feature tracking data, User actions and / or user input data, Subdivision rendering pose and spatial information, as well as An augmented reality object representation comprising at least one of a graphical description of an object and an object position anchor. It features immersive and interactive data selected from a list consisting of, The method according to any one of claims 23 to 31.
33. The step of determining the media configuration is An application configuration for an application hosted on the second device, the application configuration having one or more media capabilities of the second device, and / or Real-time communication application function "RTC AF", Provisioning function, The segmented rendering application function "SR AF", and / or Configuration Function Application Service Provider (ASP) configuration provisioned to at least one of the following The step of determining the media configuration based on at least one of the following: The method according to any one of claims 23 to 32.
34. A method for wireless communication, A step of receiving a media configuration for segmented rendering of a video media flow from a first device, wherein the media configuration comprises one or more parameters indicating that the encoded video stream of the video media flow comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata, The steps include configuring the device using the media configuration to receive the video media flow, A step of decoding the encoded video stream of the video media flow, comprising extracting the non-video encoded metadata, The steps include: consuming one or more data units of multimedia immersion and interaction data from the non-video encoded metadata; A method that includes this.
35. The aforementioned non-video encoded metadata, A first field comprising identifiers for the syntax and semantic representation formats of one or more data units of multimedia immersion and interaction data, A second field comprising one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantic representation format corresponding to the identifier of the first field, The method according to claim 34.
Citation Information
Patent Citations
US63/478,932
US63/420,885