Split rendering configuration for multimedia immersion and interaction data in wireless communication system
By adopting a split rendering configuration method in the wireless communication system, using UUID identifiers and signaling mechanisms to transmit interactive and immersive metadata, the flexibility and real-time problems of transmitting multimedia data in the prior art are solved, efficient multimedia data transmission and synchronization are achieved, and user experience of XR and cloud games is improved.
Patent Information
- Application Number
- CN202380082281.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-15
- Filing Date
- 2023-05-16
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, solutions for transmitting multimedia interaction and immersive data lack flexible, real-time and coding-independent solutions and cannot meet the needs of different multimedia data types, various syntax and formats, various data sizes, and real-time synchronization constraints, especially in XR and cloud gaming applications.
A split rendering configuration method in a wireless communication system is proposed. By determining the media configuration of a video media stream, including non-video coded metadata, establishing a multimedia split rendering content delivery session, and using multimedia immersion and interaction data for split rendering, the interactive and immersion metadata is transmitted using UUID identifiers and signaling mechanisms.
It realizes efficient and flexible transmission of multimedia interaction and immersive data in wireless communication systems, meets the requirements of real-time synchronization and low latency, supports various data formats and sizes, adapts to the needs of rapidly evolving application, and improves user experience.
Smart Images

Figure CN120303944A_ABST
Abstract
Description
Technical Field
[0001] The subject matter disclosed herein generally relates to the field of implementing split rendering configurations for multimedia immersion and interactive data in a wireless communication system. This document defines apparatuses and methods for wireless communication in a wireless communication system. Background Art
[0002] Interactive and immersive multimedia communication refers to various information flows that carry potentially time-sensitive inputs from one terminal to a remote terminal over a network. With the expansion of large-scale online games, cloud games, and extended reality (XR) in the market, applications relying on such multimedia communication modes have become increasingly popular. Multimedia information flows typically go beyond traditional video and audio streams and separately include additional formats of the following categories: device capabilities, media descriptions, and spatial interaction information. These are communicated to and from a graphics rendering engine and user devices over heterogeneous networks. Thus, these media and data types are the basis for the successful implementation of truly immersive and interactive applications that process user input information under a set of latency constraints and return responses partly based on exciting user inputs.
[0003] According to 3GPP Technical Report TR 26.928 (v17.0.0, April 2022), XR is used as an umbrella term for different types of realities. These types of realities include: virtual reality (VR), augmented reality (AR), and mixed reality (MR).
[0004] VR is a rendered version of a delivered visual and audio scene. In this case, the rendering is designed to mimic the visual and auditory sensory stimuli of the real world as naturally as possible as the observer or user moves within the limits defined by the application. Virtual reality typically (but not necessarily) requires the user to: wear a head-mounted display (HMD) to completely replace the user's field of view with the simulated visual component; and wear headphones to provide the user with accompanying audio. In VR, some form of head and motion tracking of the user is also typically required to allow the simulated visual and audio components to be updated to ensure that, from the user's perspective, objects and sound sources remain consistent with the user's movement. In some implementations, additional means of interacting with the virtual reality simulation may be provided, but are not strictly necessary.
[0005] AR refers to the situation where the user is provided with additional information or artificially generated objects, or content superimposed on their current environment. Such additional information or content is typically visual and / or auditory, and its observation of the current environment can be direct, without intermediate sensing, processing, and rendering, or can be indirect, where its perception of its environment is relayed via sensors and can be enhanced or processed.
[0006] MR is an advanced form of AR, in which some virtual elements are inserted into the physical scene, aiming to provide the illusion that these elements are part of the real scene.
[0007] XR refers to all real and virtual combined environments and human-computer interactions generated by computer technology and wearable devices. XR includes representative forms such as AR, MR, and VR, as well as the areas in between. The level of virtuality ranges from partial sensory input to fully immersive VR. A key aspect of XR is the extension of human experience, which is particularly related to presence (represented by VR) and cognitive acquisition (represented by AR).
[0008] The core of a successful immersive XR experience is the interaction and spatial computing associated with XR application activities. This also applies to other mainstream interaction-driven applications, such as cloud gaming (CG). Interaction data and the associated spatial computing determine the response of the XR rendering engine or the CG game engine to the user's physical input, thus resulting in the cyber-physical illusion of immersion between the physical world and the virtual world.
[0009] As described in 3GPP technical document S4-221557 (November 2022) or alternatively in technical report TR 26.926 (v1.1.0, February 2022), the data carried and utilized by such applications to generate the cyber-physical immersion illusion is classified into multiple categories. This includes device capability categories, media description categories, and interaction and immersion metadata categories.
[0010] In the device capability category, the format associated with this data category describes the physical and hardware capabilities of the end-user device (UE) and / or the glasses device. Some examples in this sense are camera subsystem capabilities and camera configurations (e.g., focal length, available zoom, and depth calibration information, pose reference of the main camera, etc.), projection formats (e.g., cube map, equirectangular, fisheye, stereo, etc.). Device capability data is usually static and available before session establishment, so its transmission and transfer over the network are not very important because the device capability data can be embedded into typical session configuration processes and protocols, such as the Session Initiation Protocol (SIP) and / or the Session Description Protocol (SDP). Device capability data itself is not time-sensitive in real-time and has no real-time transmission requirements.
[0011] In the media description class, the spatial and / or object content of the data description view. For example, the data can be a scene description, which is used to elaborate on the 3D combination of 2D and 3D objects spatially anchored in the scene (e.g., a tree structure or graph structure, usually in the glTF2.0 or JSON syntax). Another possible representation is a spatial description for spatial computing and the mapping of the real world to its virtual counterpart (and vice versa). In some other examples, this data type can include: a 3D model descriptor of an object and its attributes, such as being formatted as a mesh (i.e., a collection of vertices, edges, and faces); or point cloud data formatted under the PoLYgon (PLY) syntax for use by a visual rendering device (i.e., the UE). Other data types can represent a dynamic world graph representation, whereby selected trackable objects (e.g., geographically cached AR / QR codes, geotrackable objects such as physical objects located at specified world positions, dynamic physical objects such as buses, subways, etc.) dynamically enter and leave the perspective of the world scene and need to be communicated to the AR runtime in real time. The media description class of the data may be of a relatively large size (i.e., usually even exceeding 10 megabytes), and this media description class may be updated at a relatively low frequency (within a time scale of more than a dozen seconds) under various event triggers (e.g., user viewport change, new object entering the scene, old object exiting the scene, scene change and / or update, etc.). The media description data can be real-time sensitive as it is involved in completing the display of virtual rendering to a rendering device (such as the UE), and thus can benefit from real-time transmission over the network.
[0012] In the interaction and immersion metadata class, this data type contains user spatial interaction information, such as: user viewport description (i.e., the encoding of azimuth, elevation, tilt angle, and the associated movement range, which describes the projection of the user's view onto the target display); user field of view (FoV) (i.e., the range of the visible world from the perspective of the observer, usually described in the angular domain, e.g., radians / degrees in the vertical and horizontal planes); user pose / orientation tracking data (i.e., 3D vectors with microsecond / nanosecond timestamps for position and quaternion representation for describing up to 6DoF orientation); user gesture tracking data (i.e., an array of tracked hands, each hand consisting of an array of hand joint positions relative to the base space);
[0013] User body tracking data (e.g., BioVision hierarchical BVH encoding of body motion and body segment motion); user facial expression / eye movement tracking data (e.g., as an array of key points / feature locations or its encoding into a predefined facial expression class); user manipulable input (e.g., OpenXR action capture of user input to controllers or hardware command units supported by an AR / VR device); split rendering pose and spatial information (e.g., pose information containing pose information used by a split rendering server to prerender an XR scene, or alternatively, a scene projection as a video frame); application and AR anchor data and descriptions (i.e., metadata determining the position of an object or point in user space, as an anchor for placing virtual 2D / 3D objects, such as text rendering, 2D / 3D photo / video content, etc.). Summary of the Invention
[0014] Generally, interactive and immersive data types have certain characteristics. These characteristics include: low data footprint, typically from 32 bytes to approximately several hundred bytes per message, with no established codec for compressing the data source; high sampling rates that vary between video FPS frequencies (e.g., 60Hz to 250Hz), and in some cases where sample aggregation is not performed, even at a sampling frequency of 1000Hz, raw sample reports can be sent; the data can trigger responses with low latency requirements (e.g., end-to-end of up to 50 milliseconds from interaction to response perceived by the user); can be synchronized with other media streams (e.g., video or audio media streams); can be synchronized with other interactive data (e.g., pose information can be synchronized with user actions, or alternatively with object actions); reliability is optional and determined by individual application requirements (e.g., in a split rendering scenario, the server can predict future pose estimates based on available pose information and thus does not require high reliability with an error rate below 10^(-3)); data encoding typically follows proprietary / non-standardized or rapidly evolving application-dependent and interaction-dependent formats (e.g., formats and transport containers for a large amount of metadata have not been fully defined and / or specified); carrying privacy-sensitive interaction events and data, such as: user pose / gaze, user input / action, and / or user gestures or body orientation.
[0015] Therefore, interactive and immersive metadata types are real-time sensitive and require real-time transmission over a network, comparable to existing solutions for established media streams related to video codecs or alternatively audio codecs.
[0016] In addition, in XR multimedia information streams, the format and grammar of such information streams typically depend on the application, platform, and / or hardware (HW), and are inconsistent with well-established media formats and codecs (e.g., audio or video codecs), and no mainstream coding, grammar, and semantics are well established. In addition, due to the rapid evolution of such information streams and related data formats, no mainstream transmission-specific solution has been established. This latter fact requires rapid adaptation to new versions at a higher rate than the typical traditional media codec development cycle. This has motivated the search for transmission solutions over the network of such information streams in a real-time, flexible, and codec / grammar-agnostic manner to provide application developers with the necessary modern tools for rapidly emerging and disruptive interactive applications.
[0017] Currently, media descriptions and interactive data that facilitate real-time transmission and synchronization can be sent in some implementations based on at least three different options according to the prior art. These techniques are: WebRTC data channels based on the Stream Control Transmission Protocol (SCTP); RTP header extensions that embed metadata information in the band in RTP transmissions (e.g., U.S. Patent Application #63 / 420,885); and new RTP payload formats generally dedicated to real-time metadata transmission (e.g., U.S. Patent Application #63 / 478,932).
[0018] Potential solutions for the first technique (discussed in 3GPP Tdoc S4-221557) include the SCTP data channel of WebRTC for carrying interactive metadata. A common data channel payload format for timing metadata including timestamps is added to the block user data section of the SCTP data channel. This is a flexible option for metadata transmission because the SCTP data channel allows carrying metadata not directly associated with the media and enables differentiation in terms of reliability, priority, and ordering requirements by establishing data channels with different attributes. The SCTP data channel relies on protocols established by the IETF and technologies available in today's market. The downside is that in the case of SCTP, there is no inherent timing, synchronization, and jitter support, and there is a lack of FEC mechanism.
[0019] In the case of the second technique, potential solutions (discussed in 3GPP Tdoc S4-221555 or alternatively in US Patent Application #63 / 420,885) include RTP header extensions, which are designed to carry a limited amount of interactive metadata while the associated media content is carried in the RTP payload. Support for a single metadata type or multiple metadata types can be carried in the proposed header extension to allow for scalability and flexibility. An advantage of this solution is that the transmitted metadata is time-synchronized with the media data. Additionally, all of the robustness and timing mechanisms provided by RTP are included (e.g., synchronization, jitter, congestion control support, FEC mechanisms, etc.). However, this only makes sense in the presence of a media stream, which is typically the case for AR-specific use cases or VR-specific use cases with the help of split rendering. If there is no media stream, RTP packets with an empty / dummy payload need to be transmitted. Another consideration is the potentially large size of the RTP header, which depends on the metadata type. If a receiver is unable to process the RTP header extension, the receiver can also silently ignore it.
[0020] In the case of the third technique, the use of a separate RTP stream is possible, where the interactive metadata is carried in the RTP payload (as proposed in US Patent Application #63 / 478,932). In RTP, the details of media encoding (such as signal sampling rate, frame size, and timing) are specified in the RTP payload format. Therefore, sending interactive metadata in a separate RTP stream is based on defining a new RTP payload format for the interactive metadata, thus creating a common RTP metadata payload dedicated to transmitting different interactive and immersive metadata. An advantage of this solution is that it enables the use of all RTP mechanisms (timing, synchronization, jitter management support, and FEC robustness, etc.), while providing a common format that can cover all types of interactive metadata. This needs to be advanced in the IETF to become a common transmission standard. However, in the IETF, the definition of a new payload format typically takes at least two years, meaning that the developed format will be useful for 3GPP or similar network systems at the earliest by the end of the 3GPP Release 19 cycle or alternatively in the mid-term.
[0021] Therefore, a solution for transmitting multimedia interactive and immersive data is of interest, which caters to different multimedia data types (e.g., media descriptions, interactive and immersive metadata classes and their subclasses), including various syntaxes and formats, various data sizes (e.g., from hundreds of bytes to dozens of megabytes), various timing data generations (e.g., from data with strict timing generated irregularly to event-based data), and real-time synchronization constraints.
[0022] In addition, split rendering additionally requires an efficient signaling function for media configurations of various information flows, as metadata for the split rendering function as well as media display. Therefore, a solution for media configuration signaling and interactive and immersive metadata transmission is proposed herein as an enabler for the split rendering architecture.
[0023] A process for split rendering configuration of multimedia immersive and interactive data in a wireless communication system is disclosed herein. The process can be implemented by an apparatus and method for wireless communication in a wireless communication system.
[0024] An apparatus for wireless communication is provided, the apparatus comprising: a processor; and a memory coupled to the processor, the processor being configured to cause the apparatus to: determine a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate that the encoded video stream of the video media stream includes one or more data units of multimedia immersive and interactive data as non-video-encoded metadata; signal the media configuration to a second apparatus; establish, at least in part based on the media configuration, a multimedia split rendering content delivery session including the video media stream with the second apparatus; and use the one or more data units of multimedia immersive and interactive data for split rendering the video media stream.
[0025] A further apparatus for wireless communication is provided, the apparatus comprising: a processor; and a memory coupled to the processor, the processor being configured to cause the apparatus to: receive from a first apparatus a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate that the encoded video stream of the video media stream includes one or more data units of multimedia immersive and interactive data as non-video-encoded metadata; configure the apparatus to receive the video media stream using the media configuration; decode the encoded video stream of the video media stream, wherein the decoding includes extracting non-video-encoded metadata; and consume the one or more data units of multimedia immersive and interactive data from the non-video-encoded metadata.
[0026] There is also provided a wireless communication method, the method comprising: determining a media configuration for split rendering of a video media stream, wherein the media configuration comprises one or more parameters that indicate that the encoded video stream of the video media stream comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; signaling the media configuration to a second device; establishing, at least in part based on the media configuration, a multimedia split rendering content delivery session with the second device that comprises the video media stream; and using the one or more data units of multimedia immersion and interaction data for split rendering the video media stream. There is also provided a method for wireless communication, the method comprising: receiving, from a first device, a media configuration for split rendering of a video media stream, wherein the media configuration comprises one or more parameters that indicate that the encoded video stream of the video media stream comprises one or more data units of multimedia immersion and interaction data as non-video encoded metadata; configuring the device to receive the video media stream using the media configuration; decoding the encoded video stream of the video media stream, wherein the decoding comprises extracting the non-video encoded metadata; and consuming the one or more data units of multimedia immersion and interaction data from the non-video encoded metadata. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To describe the manner in which the advantages and features of the present disclosure may be obtained, a description of the present disclosure is presented by reference to certain apparatus and methods illustrated in the accompanying drawings. Each of these drawings depicts only certain aspects of the present disclosure and should not, therefore, be considered as limiting its scope. For clarity, the drawings may have been simplified and are not necessarily drawn to scale.
[0028] A method and apparatus for split rendering configuration of multimedia immersion and interaction data in a wireless communication system will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0029] Figure 1 An embodiment of a wireless communication system is illustrated;
[0030] Figure 2 An embodiment of a user equipment device is illustrated;
[0031] Figure 3 An embodiment of a network node is illustrated;
[0032] Figure 4 An RTP and RTCP protocol stack over an IP network is illustrated;
[0033] Figure 5 A WebRTC (SRTP) protocol stack over an IP network is illustrated;
[0034] Figure 6 An RTP packet format and header information are illustrated;
[0035] Figure 7 Illustrates the SRTP packet format and header information;
[0036] Figure 8 Illustrates the RTP / SRTP header extension format and syntax;
[0037] Figure 9 Illustrates a simplified block diagram of a general video codec that performs spatial and temporal compression of a video source;
[0038] Figure 10 Illustrates a video coding elementary stream and corresponding multiple NAL units;
[0039] Figure 11 Illustrates an embodiment of a wireless communication method in a wireless communication system;
[0040] Figure 12 Illustrates an alternative embodiment of a wireless communication method in a wireless communication system;
[0041] Figure 13a Illustrates an embodiment of a split rendering architecture for interactive and immersive multimedia applications;
[0042] Figure 13b Illustrates an embodiment of an advanced split rendering stream;
[0043] Figure 14 Illustrates the representation of multimedia interaction and immersive user data as metadata within a video coding elementary stream for MPEG H-26x series video codecs; and
[0044] Figure 15 Illustrates the representation of multimedia interaction and immersive user data as metadata within a video coding elementary stream for the AV1 video codec. Detailed Description
[0045] Those skilled in the art will understand that aspects of the present disclosure may be embodied as a system, apparatus, method, or program product. Accordingly, the arrangements described herein may be implemented in a fully hardware form, a fully software form (including firmware, resident software, microcode, etc.), or a form combining software aspects and hardware aspects.
[0046] For example, the disclosed methods and apparatuses may be implemented as hardware circuits, including custom very large scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. The disclosed methods and apparatuses may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, and the like. As another example, the disclosed methods and apparatuses may include one or more physical or logical blocks of executable code, which may be organized, for example, as objects, procedures, or functions.
[0047] Additionally, the methods and apparatuses may take the form of a program product, which is embodied in one or more computer-readable storage devices that store machine-readable code, computer-readable code, and / or program code, hereinafter referred to as code. The storage device may be tangible, non-transitory, and / or non-transmission. The storage device may not embody a signal. In some arrangements, the storage device only takes the form of a signal for accessing the code.
[0048] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable storage medium. The computer-readable storage medium may be a storage device that stores the code. The storage device may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, holographic, micro-mechanical, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
[0049] More specific examples (a non-exhaustive list) of the storage device include the following: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0050] Throughout this specification, references to examples of a particular method or apparatus, or the like, mean that the particular features, structures, or characteristics described in connection with the example are included in at least one implementation of the methods and apparatuses described herein. Thus, unless otherwise expressly specified, references to features of examples of a particular method or apparatus, or the like, may, but do not necessarily, refer to the same example, but rather mean "one or more but not all examples". Unless otherwise expressly specified, the terms "including", "comprising", "having", and variations thereof mean "including but not limited to". Unless otherwise expressly specified, the listing of items does not imply that any or all of the items are mutually exclusive. Unless otherwise expressly specified, the terms "a", "an", and "the" also refer to "one or more".
[0051] As used herein, a list with the conjunction "and / or" includes any single item in the list or a combination of items in the list. For example, the list of A, B, and / or C includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C. As used herein, a list using the term "one or more of..." includes any single item in the list or a combination of items in the list. For example, one or more of A, B, and C includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C. As used herein, a list using the term "one of..." includes one and only one of any single item in the list. For example, "one of A, B, and C" includes only A, only B, or only C, and does not include the combination of A, B, and C. As used herein, "a member selected from the group consisting of A, B, and C" includes one and only one of A, B, or C, and does not include the combination of A, B, and C. As used herein, "a member selected from the group consisting of A, B, and C and their combinations" includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C.
[0052] Furthermore, the features, structures, or characteristics described herein may be combined in any suitable manner. In the following description, numerous specific details are provided, such as examples of programming, software modules, user selections, network transactions, database queries, database structures, hardware modules, hardware circuits, hardware chips, etc., to provide a thorough understanding of the present disclosure. However, those skilled in the relevant art will recognize that the disclosed methods and apparatuses may be practiced without one or more of the specific details, or may be practiced using other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0053] Aspects of the disclosed methods and apparatuses will now be described with reference to the schematic flowcharts and / or schematic block diagrams of methods, apparatuses, systems, and program products. It will be understood that each block of the schematic flowcharts and / or schematic block diagrams, and combinations of blocks in the schematic flowcharts and / or schematic block diagrams, can be implemented by code. This code can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / actions specified in the schematic flowcharts and / or schematic block diagrams.
[0054] The code can also be stored in a storage device that can direct a computer, other programmable data processing apparatus, or other devices to operate in a particular manner, such that the instructions stored in the storage device produce an article of manufacture including instructions for implementing the functions / actions specified in the schematic flowcharts and / or schematic block diagrams.
[0055] The code can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices, thereby producing a computer-implemented process such that the code executed on the computer or other programmable apparatus provides a process for implementing the functions / actions specified in the schematic flowcharts and / or schematic block diagrams.
[0056] The schematic flowcharts and / or schematic block diagrams in the figures illustrate the possible architectures, functions, and operations of apparatuses, systems, methods, and program products. In this regard, each block in the schematic flowcharts and / or schematic block diagrams can represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function(s).
[0057] It should also be noted that in some alternative implementations, the functions shown in the blocks may not occur in the order shown in the figures. For example, in fact, two consecutive blocks shown may be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functions involved. Other steps and methods may be conceived that are equivalent in function, logic, or effect to one or more blocks or portions thereof of the illustrated figures.
[0058] The description of the elements in each figure may refer to the elements of the subsequent figures. In all the figures, the same reference numerals refer to the same elements.
[0059] Figure 1An embodiment of a wireless communication system 100 for split rendering configuration of multimedia immersion and interactive data in a wireless communication system is depicted. In one embodiment, the wireless communication system 100 includes a remote unit 102 and a network unit 104. Although Figure 1 a specific number of remote units 102 and network units 104 are depicted, those skilled in the art will recognize that any number of remote units 102 and network units 104 may be included in the wireless communication system 100. A wireless communication system may include a wireless communication network and at least one wireless communication device. The wireless communication device is typically a 3GPP user equipment (UE). The wireless communication network may include at least one network node. The network node may be a network unit.
[0060] In one embodiment, the remote unit 102 may include a computing device, such as a desktop computer, laptop computer, personal digital assistant (PDA), tablet computer, smart phone, smart TV (e.g., a TV connected to the Internet), set-top box, gaming console, security system (including security cameras), in-vehicle computer, network device (e.g., router, switch, modem), aircraft, drone, etc. In some embodiments, the remote unit 102 includes a wearable device, such as a smart watch, fitness band, optical head-mounted display, etc. Additionally, the remote unit 102 may be referred to as a subscriber unit, mobile device, mobile station, user, terminal, mobile terminal, fixed terminal, subscriber station, UE, user terminal, device, or other terms used in the art. The remote unit 102 may communicate directly with one or more of the network units in the network unit 104 via UL communication signals. In certain embodiments, the remote unit 102 may communicate directly with other remote units 102 via sidelink communication.
[0061] The network element 104 can be distributed over a geographical area. In some embodiments, the network element 104 may also be referred to as an access point, access terminal, base station, base station, Node B, eNB, gNB, home Node B, relay node, device, core network, air server, radio access node, AP, NR, network entity, access and mobility management function (AMF), unified data management function (UDM), unified data repository (UDR), UDM / UDR, policy control function (PCF), radio access network (RAN), network slice selection function (NSSF), operation, maintenance and management (OAM), session management function (SMF), user plane function (UPF), application function, authentication server function (AUSF), security anchor function (SEAF), trusted non-3GPP gateway function (TNGF), application function, service enabler architecture layer (SEAL) function, vertical application enabler server, edge enabler server, edge configuration server, mobile edge computing platform function, mobile edge computing application, application data analysis enabling server, SEAL data delivery server, middleware entity, network slice capability management server, or any other term used in the art. The network element 104 is generally part of a radio access network that includes one or more controllers communicatively coupled to one or more corresponding network elements 104. The radio access network is generally communicatively coupled to one or more core networks, which may be coupled to other networks such as the Internet and the public switched telephone network, etc. These elements and other elements of the radio access network and the core network are not shown, but are generally well known to those of ordinary skill in the art.
[0062] In one implementation, the wireless communication system 100 complies with the new radio (NR) protocol standardized in 3GPP, where the network element 104 transmits on the downlink (DL) using an orthogonal frequency division multiple access (OFDM) modulation scheme, and the remote unit 102 transmits on the uplink (UL) using a single carrier frequency division multiple access (SC-FDMA) scheme or an OFDM scheme. However, more generally, the wireless communication system 100 may implement some other open or proprietary communication protocol, e.g., WiMAX, IEEE802.11 variants, GSM, GPRS, UMTS, LTE variants, CDMA2000, ZigBee, Sigfox, LoraWAN and other protocols. The present disclosure is not intended to be limited to the implementation of any particular wireless communication system architecture or protocol.
[0063] The network element 104 may serve multiple remote units 102 within a service area (e.g., a cell or a cell sector) via a wireless communication link. The network element 104 transmits DL communication signals to serve the remote units 102 in the time domain, frequency domain, and / or spatial domain.
[0064] Figure 2 The user equipment device 200 that may be used to implement the methods described herein is depicted. The user equipment device 200 is used to implement one or more of the solutions described herein. The user equipment device 200 conforms to one or more of the user equipment devices described in the embodiments herein. Specifically, the user equipment device 200 may include Figure 13a the UE 102 or UE 1310, such as including an application client 1311, an MSH 1312, and a split rendering client 1313. The user equipment device 200 includes a processor 205, a memory 210, an input device 215, an output device 220, and a transceiver 225.
[0065] The input device 215 and the output device 220 may be combined into a single device, such as a touch screen. In some implementations, the user equipment device 200 does not include any input device 215 and / or output device 220. The user equipment device 200 may include one or more of the following: a processor 205, a memory 210, and a transceiver 225, and may not include an input device 215 and / or an output device 220.
[0066] As depicted, the transceiver 225 includes at least one transmitter 230 and at least one receiver 235. The transceiver 225 may communicate with one or more cells (or wireless coverage areas) supported by one or more base station units. The transceiver 225 may be operable on an unlicensed spectrum. In addition, the transceiver 225 may include multiple UE panels supporting one or more beams. In addition, the transceiver 225 may support at least one network interface 240 and / or application interface 245. The (multiple) application interfaces 245 may support one or more APIs. The (multiple) network interfaces 240 may support 3GPP reference points, such as Uu, N1, PC5, etc. As will be understood by those of ordinary skill in the art, other network interfaces 240 may be supported.
[0067] The processor 205 may include any known controller capable of executing computer-readable instructions and / or capable of performing logical operations. For example, the processor 205 may be a microcontroller, a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), an auxiliary processing unit, a field-programmable gate array (FPGA), or a similar programmable controller. The processor 205 may execute instructions stored in the memory 210 to perform the methods and routines described herein. The processor 205 is communicatively coupled to the memory 210, the input device 215, the output device 220, and the transceiver 225.
[0068] The processor 205 may control the user equipment device 200 to implement the user equipment device behavior described herein. The processor 205 may include: an application processor (also referred to as the "main processor") that manages application domain functions and operating system (OS) functions, and a baseband processor (also referred to as the "baseband radio processor") that manages radio functions.
[0069] The memory 210 may be a computer-readable storage medium. The memory 210 may include volatile computer storage media. For example, the memory 210 may include RAM, which includes dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), and / or static RAM (SRAM). The memory 210 may include non-volatile computer storage media. For example, the memory 210 may include a hard disk drive, flash memory, or any other suitable non-volatile computer storage device. The memory 210 may include both volatile computer storage media and non-volatile computer storage media.
[0070] The memory 210 may store data related to implementing a service category field as described herein. The memory 210 may also store program code and related data, such as an operating system or other controller algorithms operating on the device 200.
[0071] The input device 215 may include any known computer input device, which includes a touchpad, buttons, a keyboard, a stylus, a microphone, etc. The input device 215 may be integrated with the output device 220, for example, as a touch screen or a similar touch-sensitive display. The input device 215 may include a touch screen such that text can be input using a virtual keyboard displayed on the touch screen and / or by handwriting on the touch screen. The input device 215 may include two or more different devices, such as a keyboard and a touchpad.
[0072] The output device 220 may be designed to output visual signals, auditory signals, and / or tactile signals. The output device 220 may include an electronically controllable display or display device capable of outputting visual data to a user. For example, the output device 220 may include, but is not limited to: a liquid crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED) display, a projector, or a similar display device capable of outputting images, text, etc. to a user. As another non-limiting example, the output device 220 may include a wearable display that is separate from but communicatively coupled to the remainder of the user equipment device 200, such as a smartwatch, smart glasses, a heads-up display, etc. Additionally, the output device 220 may be a component of a smartphone, a personal digital assistant, a television, a desktop computer, a notebook (laptop) computer, a personal computer, a vehicle dashboard, etc.
[0073] The output device 220 may include one or more speakers for generating sound. For example, the output device 220 may generate an audible alert or notification (e.g., a beep or a buzzer). The output device 220 may include one or more haptic devices for generating vibration, motion, or other tactile feedback. All or part of the output device 220 may be integrated with the input device 215. For example, the input device 215 and the output device 220 may form a touchscreen or a similar touch-sensitive display. The output device 220 may be located near the input device 215.
[0074] The transceiver 225 communicates with one or more network functions of a mobile communication network via one or more access networks. The transceiver 225 operates under the control of the processor 205 to transmit messages, data, and other signals, and also to receive messages, data, and other signals. For example, the processor 205 may selectively activate the transceiver 225 (or a portion thereof) at a particular time in order to transmit and receive messages.
[0075] The transceiver 225 includes at least one transmitter 230 and at least one receiver 235. One or more transmitters 230 may be used to provide uplink communication signals to a base station unit of a wireless communication network. Similarly, one or more receivers 235 may be used to receive downlink communication signals from the base station unit. Although only one transmitter 230 and one receiver 235 are illustrated, the user equipment device 200 may have any suitable number of transmitters 230 and receivers 235. Additionally, the (plural) transmitters 230 and (plural) receivers 235 may be any suitable type of transmitter and receiver. The transceiver 225 may include a first transmitter / receiver pair for communicating with the mobile communication network via a licensed radio spectrum, and a second transmitter / receiver pair for communicating with the mobile communication network via an unlicensed radio spectrum.
[0076] A first transmitter / receiver pair that can be used to communicate with a mobile communication network over a licensed radio spectrum and a second transmitter / receiver pair that can be used to communicate with a mobile communication network over an unlicensed radio spectrum can be combined into a single transceiver unit, such as a single chip that performs functions for use with both the licensed radio spectrum and the unlicensed radio spectrum. The first transmitter / receiver pair and the second transmitter / receiver pair can share one or more hardware components. For example, certain transceivers 225, transmitters 230, and receivers 235 can be implemented as physically separate components that access shared hardware resources and / or software resources, such as, for example, network interface 240.
[0077] One or more transmitters 230 and / or one or more receivers 235 can be implemented and / or integrated into a single hardware component, such as a multi-transceiver chip, a system-on-chip, an application-specific integrated circuit (ASIC), or other types of hardware components. One or more transmitters 230 and / or one or more receivers 235 can be implemented and / or integrated into a multi-chip module. Other components, such as network interface 240 or other hardware components / circuits, can be integrated with any number of transmitters 230 and / or receivers 235 into a single chip. Transmitters 230 and receivers 235 can be logically configured as transceivers 225 that use one or more common control signals, or as modular transmitters 230 and receivers 235 implemented in the same hardware chip or in a multi-chip module.
[0078] Figure 3 Additional details of a network node 300 that can be used to implement the methods described herein are depicted. Network node 300 can be an implementation of an entity in a wireless communication network (e.g., in one or more of the wireless communication networks described herein). Network node 300 can include network nodes such as Figure 13a nodes of the edge network 1320 in [e.g., SR AF 1321 and configuration function 1321a and provisioning function 1322b; RTC AS 1322 and signaling server 1322a and SRS 1322b], or can be a node of the data network 1330, such as ASP 1331. Network node 300 includes a processor 305, a memory 310, an input device 315, an output device 320, and a transceiver 325.
[0079] The input device 315 and the output device 320 can be combined into a single device, such as a touch screen. In some implementations, the network node 300 does not include any input device 315 and / or output device 320. The network node 300 can include one or more of the following: a processor 305, a memory 310, and a transceiver 325, and may not include an input device 315 and / or an output device 320.
[0080] As depicted, the transceiver 325 includes at least one transmitter 330 and at least one receiver 335. Here, the transceiver 325 communicates with one or more remote units 200. Additionally, the transceiver 325 can support at least one network interface 340 and / or an application interface 345. The application interface(s) 345 can support one or more APIs. The network interface(s) 340 can support 3GPP reference points, such as Uu, N1, N2, and N3. As understood by those of ordinary skill in the art, other network interfaces 340 can be supported.
[0081] The processor 305 can include any known controller capable of executing computer-readable instructions and / or capable of performing logical operations. For example, the processor 305 can be a microcontroller, a microprocessor, a CPU, a GPU, an auxiliary processing unit, an FPGA, or a similar programmable controller. The processor 305 can execute the instructions stored in the memory 310 to perform the methods and routines described herein. The processor 305 is communicatively coupled to the memory 310, the input device 315, the output device 320, and the transceiver 325.
[0082] The memory 310 can be a computer-readable storage medium. The memory 310 can include volatile computer storage media. For example, the memory 310 can include RAM, which includes dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), and / or static RAM (SRAM). The memory 310 can include non-volatile computer storage media. For example, the memory 310 can include a hard disk drive, flash memory, or any other suitable non-volatile computer storage device. The memory 310 can include both volatile computer storage media and non-volatile computer storage media.
[0083] The memory 310 can store data related to establishing multipath unicast links and / or mobility operations. For example, as described herein, the memory 310 can store parameters, configurations, resource allocations, policies, etc. The memory 310 can also store program code and related data, such as an operating system or other controller algorithms operating on the network node 300.
[0084] The input device 315 may include any known computer input device, including a touchpad, buttons, a keyboard, a stylus, a microphone, etc. The input device 315 may be integrated with the output device 320, for example, as a touch screen or a similar touch-sensitive display. The input device 315 may include a touch screen such that text can be input using a virtual keyboard displayed on the touch screen and / or by handwriting on the touch screen. The input device 315 may include two or more different devices, such as a keyboard and a touchpad.
[0085] The output device 320 may be designed to output visual signals, auditory signals, and / or tactile signals. The output device 320 may include an electronically controllable display or display device capable of outputting visual data to a user. For example, the output device 320 may include, but is not limited to: an LCD display, an LED display, an OLED display, a projector, or a similar display device capable of outputting images, text, etc. to a user. As another non-limiting example, the output device 320 may include a wearable display separated from but communicatively coupled to the rest of the network node 300, such as a smartwatch, smart glasses, a heads-up display, etc. Additionally, the output device 320 may be a component of a smartphone, a personal digital assistant, a television, a desktop computer, a notebook (laptop) computer, a personal computer, a vehicle dashboard, etc.
[0086] The output device 320 may include one or more speakers for generating sound. For example, the output device 320 may generate an audible alert or notification (e.g., a beep or a buzzer). The output device 320 may include one or more tactile devices for generating vibration, movement, or other tactile feedback. All or part of the output device 320 may be integrated with the input device 315. For example, the input device 315 and the output device 320 may form a touch screen or a similar touch-sensitive display. The output device 320 may be located near the input device 315.
[0087] The transceiver 325 includes at least one transmitter 330 and at least one receiver 335. One or more transmitters 330 may be used to communicate with a UE as described herein. Similarly, one or more receivers 335 may be used to communicate with network functions in a PLMN and / or a RAN as described herein. Although only one transmitter 330 and one receiver 335 are illustrated, the network node 300 may have any suitable number of transmitters 330 and receivers 335. Additionally, the (multiple) transmitters 330 and the (multiple) receivers 335 may be any suitable type of transmitter and receiver.
[0088] There are certain standardized real-time applicable transport architectures and protocols, such as the Real-Time Transport Protocol (RTP) defined in IETF standard RFC 3550 titled "RTP: A Transport Protocol for Real-Time Applications", the Secure Real-Time Transport Protocol (SRTP) for its secure provision defined in IETF standard RFC 3711 titled "Secure Real-Time Transport Protocol (SRTP)", and the Web Real-Time Communication WebRTC of its network target stack defined in the W3C standard recommendation of March 6, 2023, titled "WebRTC: Real-Time Communication in Browsers".
[0089] RTP is a media codec-independent network protocol that has application layer framing for real-time delivery of multimedia (e.g., audio, video, etc.) data over IP networks. RTP is used in conjunction with its sister protocol for control, i.e., the Real-Time Transport Control Protocol (RTCP), to provide end-to-end features such as jitter compensation, packet loss and out-of-order delivery detection, synchronization, and source stream multiplexing. Figure 4 An overview of the RTP and RTCP protocol stacks is illustrated. The IP layer 405 carries signaling from the media session data plane 410 as well as the media session control plane 450. The data plane 410 stack includes functions for: User Datagram Protocol (UDP) 412, RTP 416, RTCP 414, media codec 420, and quality control 422. The control plane 450 stack includes functions for: UDP 452, Transmission Control Protocol (TCP) 454, Session Initiation Protocol (SIP) 462, and Session Description Protocol (SDP) 464.
[0090] SRTP is a secure version of RTP that provides: encryption (primarily through payload confidentiality), message authentication and integrity protection (through PDU, i.e., header and payload, signature), and replay attack protection. Similar to RTP, the SRTP sister protocol is SRTCP. SRTCP provides the same functionality for its RTCP counterpart. Thus, in the normal SRTP version, the RTP header information is still accessible but not modifiable, while the payload is encrypted. These security provisions are Figure 7 illustrated in. Additionally, the key exchange and additional security parameters required for using SRTP are based on the Datagram Transport Layer Security (DTLS) key exchange process. For these reasons, SRTP is used as the transport protocol for media in the WebRTC stack to ensure secure RTC multimedia communication through the network browser interface.
[0091] Figure 5Illustrates an overview of the WebRTC (i.e., SRTP-based) protocol stack. As illustrated, the IP layer 505 carries signaling from both the data plane 510 and the control plane 550. The data plane 510 stack includes functions for: UDP 512, Interactive Connectivity Establishment (ICE) 524, Datagram Transport Layer Security (DTLS) 526, SRTP 517, SRTCP 515, media codec 520, quality control 522, and SCTP 528. ICE 574 can use the Session Traversal Utilities for NAT (STUN) protocol and Traversal Using Relays around NAT (TURN) to solve real-time media content delivery across heterogeneous networks and NAT rules and firewalls. The SCTP 528 data plane is mainly dedicated to application data channels and can be non-time-critical, while the SRTP 517-based stack (including control elements (i.e., SRTCP 515), encoding elements (i.e., media codec 520), and Quality of Service (QoS) elements (i.e., quality control 522)) is dedicated to time-critical transmission. The control plane 550 is shown to include: TCP 554, TLS 556, HTTP 558, SSE / XHR / others 568, XMPP / others 570, SDP 564, and SIP 562.
[0092] The RTP header information and the SRTP header information share the same format, as respectively shown in Figure 6 and Figure 7 illustrated. Figure 6 Illustrates an RTP packet 630, while Figure 7 illustrates an SRTP packet 760. A brief overview of the fixed header information for packets 630 and 760 will now be provided.
[0093] ‘V’ 641, 761 is 2 bits, which indicates the protocol version in use.
[0094] ‘P’ - 643, 763 is a 1-bit field, which indicates the presence of one or more zero-padding octets at the end of the payload, where among other things, padding can be necessary for fixed-size encryption blocks or for carrying multiple RTP / SRTP packets over lower-layer protocols.
[0095] ‘X’ bits 634, 764 are 1 bit, which indicates that an RTP header extension that will usually be associated with a specific data / profile will follow the standard fixed RTP / SRTP header, and this RTP header extension will carry more information about the data (e.g., marking frames of the RTP header extension for video data as described in the IETF working draft titled "Frame Marking RTP Header Extension" in November 2021; or a generic RTP header extension (such as the RTP / SRTP extension protocol) as described in the IETF standard RFC 6904 titled "Encryption of Header Extensions in the Secure Real-Time Transport Protocol (SRTP)").
[0096] ‘CC’ bits 636, 766 are 4 bits, which indicate the number of contributing media sources (CSRCs) that follow the fixed header.
[0097] ‘M’ bits 638, 768 are 1 bit, which is intended to mark the information frame boundaries in a packet stream, and its behavior is precisely specified by the RTP profile (e.g., H.264, H.265, H.266, AV1, etc.).
[0098] ‘PT’ bits 640, 770 are 7 bits, which indicate the payload type, which can be dynamic in the case of audio and video codec profiles and is negotiated via SDP (e.g., 96 for H.264, 97 for H.265, 98 for AV1, etc.). The payload profiles are registered with IANA and depend on the IETF profile, which describes how data transmission is included within the payload of an RTP PDU. For example, as specified in ITU-T standard H.265 V8 (08 / 2021), the currently IANA-registered payload profiles describe audio / video codecs and application-based forward error correction (FEC) encoded media content, but there is no applicability to unencoded non-audio-visual data formats.
[0099] ‘Sequence Number’ bits 642, 772 are 16 bits, which indicate the sequence number, which increments by 1 with each RTP packet sent on the session.
[0100] ‘Timestamp’ bits 644, 774 are 32 bits, which indicate the timestamp in ticks of the payload type clock, which reflects the sampling moment of the first octet of the RTP data packet (for a video stream, associated with a video frame), and the first timestamp of the first RTP packet is randomly selected.
[0101] The ‘Synchronization Source (SSRC) identifier’ 646, 776 is a 32-bit field that indicates a random identifier for the source of an RTP packet stream that forms part of the same timing and sequence number space, such that the receiver can group packets based on the synchronization source for playback.
[0102] The ‘Contributing Source (CSRC) identifier’ 648, 778, which is a list of up to 16 CSRC entries, each 32 bits, gives the number of CSRCs mixed by an RTP mixer within the current payload, as flagged by the CC bits. Given the SSRC identifier of a contributing source, this list identifies the contributing sources for the payload contained in the packet.
[0103] A brief overview of the remaining aspects of the complete header information for packet 630 and packet 760 will now be described.
[0104] The ‘RTP header extension’ 650, 780 is a variable-length field that is present when the X bits 634, 764 are marked. The header extension is appended to the RTP fixed header information after the CSRC list 648, 778 (if present). The RTP header extension 650, 780 is 32-bit aligned and is formed by the following fields: a 16-bit extension identifier, which is defined by the profile and is typically negotiated and determined via the Session Description Protocol (SDP) signaling mechanism; a 16-bit length field, which describes the extension header length in multiples of 32 bits, excluding the first 32 bits corresponding to the 16-bit extension identifier and the 16-bit length field itself; and a 32-bit aligned header extension raw data field, which is formatted according to a specified format for a particular RTP header extension identifier.
[0105] The format and syntax of the RTP header extension 650, 780 are similar to those of SRTP. The format 800 and syntax are as Figure 8 illustrated. Additionally, in both RTP and SRTP, only one RTP extension header 650, 780 can be appended to the fixed header information, as described in the IETF standard RFC 3550, “RTP: A Transport Protocol for Real-Time Applications”. However, for both RTP and SRTP, there is an extension to the base protocol, as described in the IETF standard RFC 8285, “Generic Mechanisms for RTP Header Extensions”, to allow multiple RTP header extensions 650, 780 of a predefined type to be appended to the fixed header information of the protocol.
[0106] In some embodiments, an RTP header extension generated at the source can be ignored by a destination endpoint that does not have the knowledge to interpret and process the RTP header extension sent by the source endpoint.
[0107] Now, starting from modern hybrid video coding, the video coding domain and metadata support will be briefly described.
[0108] The interactivity and immersion of modern and future multimedia XR applications require ensuring that the packet error rate (PER) and packet delay budget (PDB) for QoE are met. The video source jitter in mobile communication systems and the random characteristics of wireless channels make it challenging to meet the former, especially for high-rate specific digital video transmissions, such as 4K, 3D video, 2×2K eye buffer video, etc.
[0109] Current video source information is encoded based on 2D, 2D+depth, or alternatively 3D representations of video content. Regardless of the source encoder, the encoded elementary stream video content is generally organized into two abstract layers, aiming to separate the storage domain and the video coding domain (i.e., network transmission packetization and format, and correspondingly, the video coding-related syntax and related semantics of the codec). The former determines the bitstream format, while the latter specifies the content of the video coding bitstream.
[0110] In one example, the MPEG video codec family (e.g., H.264, H.265, H.266) relies on Network Abstraction Layer (NAL) units to packetize the bitstream and store it in a byte-aligned format for transmission or storage on various media, including the network. A NAL unit (NALU) can include Video Coding Layer (VCL) information (i.e., video coding content (e.g., frames, slices, tiles, etc.) NALU), and correspondingly includes non-VCL information (i.e., parameter sets, Supplemental Enhancement Information (SEI) messages, etc.). The NAL syntax thus encapsulates VCL information and non-VCL information, and provides an abstract containerization mechanism for the encoded stream in transmission, i.e., for disk storage / caching / transmission and parsing / decoding.
[0111] In another example, open-source video codec alternatives (e.g., VP8 / VP9, or similarly, AV1) adopt a scheme similar to the MPEG video codec for packetization, storage, and communication on various media. For example, the AV1 bitstream consists of Open Bitstream Units (OBUs), and each OBU can contain one or more video-coded frames, video-coded tiles, non-video-coded padding, and non-video-coded metadata as OBU_METADATA.
[0112] Based on the above, NALUs or alternatively OBUs provide a mechanism that can be utilized to transmit metadata independent of the chrominance and luminance representations of the video coding bitstream.
[0113] On the other hand, the VCL or alternatively the video coding information encapsulates the video coding process of the encoder and compresses the source-coded video information based on some entropy coding method, e.g., context-adaptive binary arithmetic coding (CABAC), context-adaptive variable length coding (CAVLC), etc.
[0114] A simplified description of the VCL process for general coding of video content will now be described. Pictures in a video sequence are segmented into coding units of a configured size (e.g., macroblocks, coding tree units, blocks, or variants thereof). The coding units can then be split under some tree partitioning structure or similar hierarchical structure, as described in ITU-T standard H.264 V8 (08 / 2021), ITU-T standard H.265 V8 (08 / 2021), ITU-T standard H.266 V4 (04 / 2022). For example, such a tree partitioning structure can include binary / ternary / quaternary trees, or under some predefined geometric-driven 2D tiling patterns described by de Rivaz, p & Haughton (2018) in the paper titled "AV1 Bitstream and Decoding Process Specification" of the Alliance for Open Media 182, e.g., 10-way splitting.
[0115] The encoder uses the visual reference between such coding units to encode the picture content in a differential manner based on the residuals. The residuals are determined according to the prediction mode associated with information reconstruction. Two prediction modes are generally available, i.e., intra prediction (referred to simply as intra) or inter prediction (referred to simply as inter). The intra mode is based on deriving and predicting the residuals according to the content of other coding units within the current picture, i.e., by calculating the residuals of the current coding unit given the coded content of the adjacent coding units of the current coding unit. On the other hand, the inter mode is based on deriving and predicting the residuals according to the content of coding units from other pictures, i.e., by calculating the residuals of the current coding unit given the coded content of the adjacent coded pictures of the current coding unit.
[0116] Then, some multi-dimensional (2D / 3D) spatial multi-modal transforms are used to further transform the residuals for compression. For example, frequency-based (i.e., discrete cosine transform, etc.) or wavelet-based linear transforms (e.g., Walsh Hadamard transform, or equivalent discrete wavelet transform) are used to extract the most prominent frequency components in the residuals of the coding unit. The tiny high-frequency contributions of the residuals are discarded, and the floating-point transform representation of the remaining residuals is further quantized to the selected number of bits per sample based on some parametric quantization process, e.g., 8 bits / 10 bits / 12 bits. Finally, an entropy coding mechanism is used to encode the transformed and quantized residuals and their motion vectors (associated with the prediction reference of the residuals) in the intra-mode or inter-mode to compress the information based on the random distribution of the source bit content. The output of this operation is the bitstream of the encoded residual content of the VCL.
[0117] A simplified general diagram of a modern hybrid (applying both temporal compression and spatial compression via intra-prediction / inter-prediction) video codec block is shown in Figure 9 FIG.
[0118] Figure 9 FIG. 900 shows a simplified block diagram of a general video codec that performs both spatial compression and temporal (motion) compression of a video source. The encoder block is embodied within the domain 910 labeled "encoder". The decoder block is embodied within the domain 920 labeled "decoder". Those skilled in the art can associate the above-described general diagram of the hybrid codec with a large number of state-of-the-art video codecs, such as but not limited to H.264, H.265, H.266 (collectively referred to as H.26x) or VP8 / VP9 / AV1. Therefore, unless specifically clarified below and narrowed to the scope of a certain codec embodiment, the concepts utilized herein should be considered as general concepts.
[0119] Block diagram 900 shows an original input video frame (picture) 901 input to the picture block segmentation functional block 911 of an encoder 910. Subsequent functional block 912 is illustrated as 'Spatial Transform'. Subsequent functional block 913 is illustrated as 'Quantization'. Subsequent functional block 914 is illustrated as 'Entropy Coding'. This functional block 914 outputs not only to a video coding bitstream 902 but also to the motion estimation 915 of the encoder 910. The motion estimation 915 outputs to the inter-frame prediction block 921 of a decoder 920. This block 921 outputs to a buffer 920 which itself outputs to a restored video frame video (picture) 903. The inter-frame prediction block 921 can be switched to connect to a summing node fed into the spatial transform 912. Further, as illustrated in block diagram 900, the quantization block 913 can output to the inverse quantization block 926 of the decoder 920. The inverse quantization block 926 is illustrated as receiving the entropy decoding 927 of the video coding bitstream 902. The inverse quantization block 926 outputs to an inverse spatial transform block 925 which is fed via a summing node to a loop and visual filtering block 923 which itself is fed to a buffer 922. As described above, block diagram 900 is illustrated by way of example to convey the various functional blocks of a modern hybrid video codec from the perspective of both encoder operation and decoder operation.
[0120] Accordingly, the encoded residual bitstream 902 is encapsulated as a base stream, as NAL units,
[0121] or equivalently, as an OBU ready to be stored or transmitted over a network. The NAL units or alternatively the OBUs are the main syntax elements of the video codec and the NAL or OBU can encapsulate encoded video parameters (e.g., video parameter set / sequence parameter set / picture parameter set (VPS / SPS / PPS)), one or more supplementary enhancement information (SEI) messages or alternatively OBU metadata payloads, and encoded video headers and residual data (e.g., slices as picture segmentation, or equivalent video frames or video tiles). The encapsulation common syntax carries information described by codec-specific semantics which are intended to determine the use of metadata, non-video-encoded data, and video-encoded data and to assist the decoding process.
[0122] In an example referring to the MPEG video codec family (e.g., H.264, H.265, H.266), the NAL unit encapsulation syntax consists of the following items: a header part that determines the start and type of the NAL unit, and a sequence of raw byte payloads containing information related to the NAL unit. The NAL unit payload can then be formed by a payload syntax or a payload-specific header and an associated payload-specific syntax. A key subset of the NAL units is formed by the following items: parameter sets (e.g., VPS, SPS, PPS), SEI messages, and configuration NAL units (parameter sets, SEI messages, and configuration NAL units are also collectively referred to as non-VCL NAL units), and picture slice NAL units that contain video coding data (e.g., entropy-based arithmetic coding) as VCL information. These concepts are illustrated in Figure 10 for the content of elementary streams that are generally applicable to the H.264, H.265, H.266 MPEG family of video codecs.
[0123] Figure 10 Illustrated is an elementary video stream and its corresponding plurality of NAL units 1000. The NAL unit 1000 is formed by a header 1010 and a payload 1020. The header 1010 contains information about the type, size, and video coding attributes and parameters of the NAL unit data encapsulation information. The NAL unit data can be a non-VCL NAL including a video / sequence / picture parameter payload 1021 and a supplementary enhancement information payload 1022, or a VCL NAL including a frame / picture / slice payload 1023 having a header 1023a and a video coding payload 1023b. The non-VCL NALU can include one or more SEI messages 1022a in the supplementary enhancement information payload 1022. The NAL header 1010 is illustrated as including the NAL unit type, the NAL unit byte length, the video coding layer ID, and the temporal video coding layer ID.
[0124] Accordingly, a decoder implementation can implement a bitstream parser that extracts necessary metadata information and VCL-associated metadata from the NAL unit sequence 1000; decodes the VCL residual coding data sequence into its transformed and quantized values; applies an inverse linear transform and restores the residual valid content; performs intra prediction or inter prediction to reconstruct the luminance and chrominance representations of each coded unit; applies additional filtering and error concealment processes; and reproduces the original picture sequence representation for video playback.
[0125] These operations and processes can occur sequentially in the order listed or in parallel, depending on the specific implementation of the decoder. Those skilled in the art should recognize that similar high-level operations apply to other video codec families, such as, for example, the AV1 codec.
[0126] To aid the disclosure in this document, metadata support in video codecs will now be briefly described.
[0127] Modern video codecs (e.g., H.264, H.265, AV1, or alternatively H.266) provide a byte-aligned transport mechanism for metadata within a video coding elementary stream or alternatively a bitstream. Such non-video coding data is alternatively referred to as metadata in the video coding context because the information included is not related to changing the luminance or chrominance of decoded frames or alternatively decoded pictures. This metadata is also encapsulated into NALUs as SEI messages of the H.26x MPEG codec family and into OBUs as OBU metadata for AV1.
[0128] For each codec, metadata can have different types associated with different syntax and semantics, which are additionally specified in: codec specifications such as (ITU-T Standard H.264 (08 / 2021), ITU-T Standard H.265 V8 (08 / 2021), ITU-T Standard H.266 V4 (04 / 2022), Rivaz, p & Haughton (2018) in the paper titled "AV1 Bitstream and Decoding Process Specification" in the Alliance for Open Media 182); or alternatively, ITU metadata type specifications such as in ITU-T H Series Specification V8 (08 / 2020) titled "H Series: Audiovisual and Multimedia Systems, Audiovisual Service Infrastructure - Moving Picture Coding: Generic Supplementary Enhancement Information Messages for Encoded Video Bitstreams", or in ITU-T Recommendation T.35 titled "Terminal Provider Code Notification Table: Information Available on National Bodies Identifying for ITU-T Recommendation T.35 Terminal Provider Code Allocations". However, in all specifications, the user data SEI message, or alternatively the OBU metadata user private data format, is not specified. Additionally, all video codecs discussed in this document and their associated encoders and decoders allow: exposing user data metadata types to the application layer via a dedicated interface by passing NALU SEI messages to the application layer, or alternatively passing OBUs carrying OBU metadata.
[0129] In the current WebRTC specification, the supported and mandated video media codecs are AVC / H.264 and its constrained baseline profile, VP8, and additional optional support for VP9. Given the increasing support for AV1 and H.265 in major browser engines (e.g., Chromium used by Google Chrome and Microsoft Edge, or alternatively Mozilla used by Mozilla Firefox), it is expected that the WebRTC specification will soon evolve to additionally include AV1 support and potentially provide optional support for a restricted profile of H.265. On the other hand, prior to 3GPP Release 18, H.264 and H.265 were the specified and supported default video codecs, while support for H.266 and AV1 was considered for further study in subsequent releases.
[0130] This disclosure leverages these facts regarding user data metadata and video codec support to facilitate a new real-time in-band transmission mechanism for interactive and immersive metadata associated with XR applications and their video streams, which is related to split rendering.
[0131] Split rendering and split rendering architectures have been studied in 3GPP, e.g., 3GPP Technical Specification TS26.565 v0.3.0 titled "Split Rendering Media Service Enabler" and 3GPP Technical Specification 26.506 v1.1.0 titled "5G Real-Time Media Communication Architecture". In addition to interactive and immersive XR applications, split rendering also has the advantages of enabling energy savings and high performance. Split rendering allows for enhancing the user experience by providing access to advanced and complex rendering that would otherwise not be possible, or alternatively would require very high energy consumption on AR / VR glasses, or alternatively would require a 5G tethered UE.
[0132] In split rendering, all or part of the 3D scene is remotely rendered on an edge application server (EAS) or alternatively on a split rendering server (SRS) placed within an edge data network (EDN). The result of the split rendering process is streamed to a 5G UE (if a tethered AR architecture is considered) or streamed to XR glasses for display. The scope of split rendering operations can be wide, ranging from full pre-rendering at the edge to offloading partial, highly processed rendering operations to the edge. This diversity includes frames corresponding to 2D with single-eye buffer rendering, 2D with binocular buffer rendering, 2D with depth information rendering, and 3D scene rendering.
[0133] In the case of 2D with single-eye buffer rendering, the Edge Application Server will produce a single 2D video rendering of the visual scene. Depending on the UE's configuration, 2D rendering with two eye buffers and appropriate projection (e.g., equirectangular) may be required. Additional support streams encoding additional information such as depth or transparency can also be added to the 2D rendering. On the UE, the scene compositor layers the available streams and composes the raw image buffer, which is swapped with the XR runtime via a swap chain for display on the AR glasses. On the other hand, in the case of 3D content, partial render offloading delegates some rendering operations to the EAS while still receiving the 3D scene at the UE. An example is offloading the light baking of the scene textures to the edge, which can be performed using techniques such as ray tracing.
[0134] In terms of the service, the split rendering UL service mode can be summarized as: the UE or alternatively the AR glasses stream pose predictions to the SRS at the EDN. Depending on the type of application (e.g., multi-party AR conferencing, interactive and immersive classroom AR, etc.), the service may also include an associated AR UL video stream.
[0135] In terms of the service, the split rendering DL service mode can be summarized as: the UE or alternatively the AR glasses receive the rendered video media for display.
[0136] Typically, the XR runtime benefits from the rendered media being passed together with the associated poses used for rendering to perform proper scene composition and display. For example, the XR runtime may need to perform pose correction based on post-reprojection, or alternatively, perform asynchronous time warping to match the current display time. The XR runtime may require additional information such as the XR space handle or reference associated with the SRS rendering operation and the XR timestamp for spatial and temporal synchronization.
[0137] Therefore, determining the split rendering transport and signaling mechanisms is very important for XR interactive and immersive applications and will be discussed in this article.
[0138] This solution addresses the use of user data SEI messages (e.g., H.264, H.265, H.266) or alternatively OBU metadata (e.g., AV1) for transmitting interaction and immersion metadata associated with XR applications in UL or DL. The format of the payload for such user data is generic and can include at least two fields. The first field is an identifier that determines the syntax and semantics of the user data format, and the second field contains the user data information as a payload encoded according to the format identified by the first field. In some embodiments, the first field is a UUID, which needs to be signaled by the source sender to the corresponding remote receiver to allow the receiver to parse and process the user data information. This disclosure specifically specifies three signaling mechanisms that enable the use of such user data via SEI messages or alternatively OBU metadata for interactive and immersive applications supported by split rendering at the edge. The proposed signaling mechanisms for the user data identifier (e.g., UUID) are: application-based signaling on the application plane; network-assisted signaling via the network application function on the control plane; and in-band signaling based on SDP attributes on the user plane.
[0139] In some embodiments, the user data identifier (e.g., UUID) can depend on at least one of the sender, the receiver, and the media capabilities of the split rendering EAS, and thus, during session initialization or alternatively during session update, the negotiation and signaling of the user data identifier are performed for each media session of the interactive and immersive application.
[0140] Disclosed herein is a device for wireless communication, the device comprising: a processor; and a memory coupled to the processor, the processor being configured to cause the device to: determine a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate: the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata; signal the media configuration to a second device; establish, at least in part based on the media configuration, a multimedia split rendering content delivery session including the video media stream with the second device; and use one or more data units of multimedia immersion and interaction data for split rendering the video media stream.
[0141] In some embodiments, the non-video coding metadata includes: a first field that includes an identifier for a syntax and semantic representation format of one or more data units for multimedia immersive and interactive data; and a second field that includes one or more data units of multimedia immersive and interactive data, where the one or more data units of multimedia immersive and interactive data are encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
[0142] In some embodiments, the identifier is indicated by a Universally Unique Identifier ‘UUID’.
[0143] In some embodiments, the UUID is unique for a specific application or session, or the UUID is globally unique. The UUID can conform to the ISO-IEC-11578 Annex A format and the ISO-IEC 9834-8 version 4 UUID, i.e., a randomly generated UUID, etc. The term ‘session’ includes a set of temporary and interactive (i.e., updatable) configurations and rules, e.g., media formats and codecs, network configurations, where the set of configurations and rules determines the exchange of information (including media content) between two or more endpoints connected via a network connection.
[0144] In some embodiments, the processor is configured to cause the device to determine a media configuration by causing the device to perform the following: determine a media configuration for a plurality of video media streams, the plurality of video media streams having corresponding encoded video streams and corresponding non-video coding metadata, where one or more parameters of the media configuration map the identifier of the first field of the corresponding non-video coding metadata to the corresponding video media stream of the corresponding non-video coding metadata.
[0145] In some embodiments, one or more parameters of the media configuration map the first field of the corresponding non-video coding metadata to the corresponding video media stream of the corresponding non-video coding metadata based on at least one of the following: application-specific stream mapping; 5-tuple indication; and media stream description attributes. For example, the 5-tuple can describe an IP stream that includes an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. The media description under SDP can be associated with a media stream of one or more UUIDs provided in the media description attributes defined by SDP.
[0146] In some embodiments, the processor is configured to cause the device to signal media configuration through at least one of the following: an application interface (which preferably can be located between the application service provider 'ASP' of the device and the split-rendering-aware application of the second device); a control plane interface (which preferably is located between the real-time communication application function 'RTC AF' of the device and the media session handler 'MSH' of the second device); and a user plane interface (which preferably is located between the split-rendering server 'SRS' of the device and the split-rendering client 'SRC' of the second device). The control plane interface can be the RTC-5 reference interface in 5GS; the user plane interface can be SR-4, RTC-4, or the SR-4m media center interface can be used.
[0147] In some embodiments, an encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs that includes: the H.264 video codec specification; the H.265 video codec specification; the H.266 video codec specification; and the AV1 video codec specification. Other video codec specifications that rely on a portion of the said specifications can also apply, for example, OMAF, V-PCC, etc. OMAF encodes omnidirectional video content, and OMAF includes at least one video stream encoded using AVC and HEVC. V-PCC encodes 3D video with 2D projections, and V-PCC includes a video stream encoding projected 2D plane frames using AVC and HEFC.
[0148] In some embodiments, the video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more supplementary enhancement information 'SEI' messages of the 'user data not registered' type; and / or the video codec includes an AV1 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more metadata open bitstream units 'OBU' of the 'unregistered user private data' type. For SEI messages, the payload type can be equal to 5. In addition, the SEI message can be prefixed or suffixed to the NALU. For OBU, the OBU metadata type can be 'X', where X can be any one of 6 - 31.
[0149] In some embodiments, one or more data units of multimedia immersion and interaction data include: immersion and interaction data selected from a list including: user viewpoint data; user field of view data; user pose / orientation data; user gesture tracking data; user body tracking data; user facial feature tracking data; user action and / or user input data; split rendering pose and spatial information; and augmented reality object representation, which includes at least one of the following: a graphical description of the object, and an object position anchor.
[0150] In some embodiments, one or more data units of multimedia immersion and interaction data include extended reality 'XR' multimedia immersion and interaction data.
[0151] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources selected from a list of media sources including: one or more physical dedicated controllers; one or more red, green, blue 'RGB' cameras; one or more RGB depth 'RGBD' cameras; one or more infrared 'IR' cameras; one or more microphones; and one or more tactile sensors.
[0152] In some embodiments, the processor is configured to cause the device to use one or more data units of multimedia immersion and interaction data for split rendering a video media stream for the purposes of: performing post - reprojection at a split rendering client 'SRC' for displaying one or more rendered frames that match the latest user pose or orientation information; partially or fully pre - rendering one or more frames at a split rendering server 'SRS'; and / or estimating the most likely user pose and orientation at the SRS, which is associated with the expected display time for one or more frames that are partially or fully pre - rendered.
[0153] In some embodiments, the processor is configured to cause the device to signal a media configuration at the initiation of a multimedia split - rendering content delivery session and / or when updating a multimedia split - rendering content delivery session.
[0154] In some embodiments, the processor is further configured to cause the device to determine a media configuration based on at least one of the following: an application configuration of an application hosted on a second device, where the application configuration includes one or more media capabilities of the second device; and / or an application service provider 'ASP' configuration supplied to at least one of the following: a real - time communication application function 'RTC AF'; a provisioning function; a split - rendering application function 'SR AF'; and / or a configuration function;
[0155] In some embodiments, the application configuration may be pre-configured with supported UUIDs; and / or if a UUID is received, the transmission may be automatically enabled.
[0156] In some embodiments, the device is a network entity / network node, and the second device is a UE.
[0157] Figure 11 Embodiment 1100 of a wireless communication method in a wireless communication system is illustrated.
[0158] The first step 1110 includes: determining a media configuration for split rendering of a video media stream, where the media configuration includes one or more parameters that indicate: the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata.
[0159] Another step 1120 includes: signaling the media configuration to a second device.
[0160] Another step 1130 includes: establishing, at least in part based on the media configuration, a multimedia split rendering content delivery session with the second device that includes the video media stream.
[0161] Another step 1140 includes: using one or more data units of multimedia immersion and interaction data for split rendering the video media stream.
[0162] In certain embodiments, method 1100 may be executed by a processor that executes program code, such as, for example, a microcontroller, a microprocessor, a CPU, a GPU, an auxiliary processing unit, an FPGA, etc.
[0163] In some embodiments, the non-video encoded metadata includes: a first field that includes: an identifier for the syntax and semantic representation format of one or more data units of multimedia immersion and interaction data; and a second field that includes: one or more data units of multimedia immersion and interaction data, where the one or more data units of multimedia immersion and interaction data are encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
[0164] In some embodiments, the identifier is indicated by a universal unique identifier 'UUID'.
[0165] In some embodiments, the UUID is unique for a particular application or session, or the UUID is globally unique. The UUID can conform to the ISO-IEC-11578 Annex A format and ISO-IEC 9834-8 version 4 UUID, i.e., a randomly generated UUID, etc. A'session' includes a set of temporary and interactive (i.e., updatable) configurations and rules, such as media formats and codecs, network configurations, and the set of configurations and rules determines the exchange of information (including media content) between two or more endpoints connected via a network connection.
[0166] In some embodiments, determining a media configuration includes: determining a media configuration for a plurality of video media streams, the plurality of video media streams having corresponding encoded video streams and corresponding non-video encoded metadata, wherein one or more parameters of the media configuration map an identifier of a first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata.
[0167] In some embodiments, one or more parameters of the media configuration map the first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata based on at least one of the following: application-specific stream mapping; 5-tuple indication; and media stream description attributes. For example, the 5-tuple can describe an IP stream, which includes an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. The media description under SDP can include an association with a media stream of one or more UUIDs provided in the SDP media description attributes, etc.
[0168] In some embodiments, signaling the media configuration includes signaling through at least one of the following: an application interface (preferably located between the application service provider 'ASP' of the device and the split-rendering-aware application of the second device); a control plane interface (which is preferably located between the real-time communication application function 'RTC AF' of the device and the media session handler 'MSH' of the second device); and a user plane interface (which is preferably located between the split-rendering server 'SRS' of the device and the split-rendering client 'SRC' of the second device). The control plane interface can be the RTC-5 reference interface in 5GS; the user plane interface can be SR-4, RTC-4, or the SR-4m media center interface can be used.
[0169] In some embodiments, an encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs that includes: the H.264 video codec specification; the H.265 video codec specification; the H.266 video codec specification; and the AV1 video codec specification. Other video codec specifications that rely on a portion of the foregoing specifications may also be applicable, for example, OMAF, V-PCC, etc. OMAF encodes omnidirectional video content and OMAF includes at least one video stream encoded using AVC and HEVC, and V-PCC encodes 3D video with 2D projections and V-PCC includes a video stream that encodes projected 2D planar frames using AVC and HEFC.
[0170] In some embodiments, the video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more supplementary enhancement information 'SEI' messages of the 'user data unregistered' type; and / or the video codec includes an AV1 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more metadata open bitstream units 'OBU' of the 'unregistered user private data' type. For SEI messages, the payload type may be 5. Additionally, SEI messages may be prefixed or suffixed to NALUs. For OBUs, the OBU metadata type may be 'X', where X may be any one of 6 - 31.
[0171] In some embodiments, one or more data units of multimedia immersive and interactive data include: immersive and interactive data selected from a list including the following: user viewpoint data; user field of view data; user pose / orientation data; user gesture tracking data; user body tracking data; user facial feature tracking data; user actions and / or user input data; split rendering pose and spatial information; and augmented reality object representations, which augmented reality object representations include at least one of the following: a graphical description of the object, and an object location anchor.
[0172] In some embodiments, one or more data units of multimedia immersive and interactive data include extended reality 'XR' multimedia immersive and interactive data.
[0173] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources, which are selected from a media source list that includes: one or more physical dedicated controllers; one or more red, green, blue 'RGB' cameras; one or more RGB-depth 'RGBD' cameras; one or more infrared 'IR' cameras; one or more microphones; and one or more haptic sensors.
[0174] Some embodiments include using one or more data units of multimedia immersion and interaction data for split rendering of a video media stream, for the purposes of: performing post-reprojection at a split rendering client 'SRC' for displaying one or more rendered frames that match the latest user pose or orientation information; partially or fully pre-rendering one or more frames at a split rendering server 'SRS'; and / or estimating at the SRS the most likely user pose and orientation that is associated with the expected display time for one or more frames that are partially or fully pre-rendered.
[0175] Some embodiments include signaling media configuration at the initiation of, and / or when updating, a multimedia split rendering content delivery session.
[0176] In some embodiments, determining the media configuration includes determining the media configuration based on at least one of: an application configuration of an application hosted on a second device, where the application configuration includes one or more media capabilities of the second device; and / or an application service provider 'ASP' configuration that is supplied to at least one of: a real-time communication application function 'RTC AF'; a provisioning function; a split rendering application function 'SR AF'; and / or a configuration function. The application configuration may be pre-configured with a supported UUID; and / or if a UUID is received, transmission may be automatically enabled.
[0177] In some embodiments, the method is performed by a network entity / node, and the second device is a UE.
[0178] The present invention also provides a device for wireless communication in a wireless communication system, the device comprising: a processor; and a memory coupled to the processor, the processor being configured to cause the device to: receive, from a first device, a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate: the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata; use the media configuration to configure the device to receive the video media stream; decode the encoded video stream of the video media stream, wherein the decoding includes: extracting the non-video encoded metadata; and consuming one or more data units of multimedia immersion and interaction data from the non-video encoded metadata.
[0179] In some embodiments, the non-video encoded metadata includes: a first field that includes: an identifier for the syntax and semantic representation format of one or more data units of multimedia immersion and interaction data; and a second field that includes: one or more data units of multimedia immersion and interaction data, the one or more data units of multimedia immersion and interaction data being encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
[0180] In some embodiments, the identifier is indicated by a Universally Unique Identifier 'UUID'.
[0181] In some embodiments, the UUID is unique for a particular application or session, or the UUID is globally unique. The UUID can conform to the ISO-IEC-11578 Annex A format and the ISO-IEC 9834-8 version 4 UUID, i.e., a randomly generated UUID, etc. A'session' includes a set of temporary and interactive (i.e., updatable) configurations and rules, e.g., media formats and codecs, network configurations, the set of configurations and rules that determine the exchange of information (including media content) between two or more endpoints connected via a network connection.
[0182] In some embodiments, the media configuration is for a plurality of video media streams, the plurality of video media streams including: corresponding encoded video streams, and corresponding non-video encoded metadata, wherein one or more parameters of the media configuration map the identifier of the first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata.
[0183] In some embodiments, one or more parameters of the media configuration map a first field of the corresponding non-video-encoded metadata to the corresponding video media stream of the non-video-encoded metadata based on at least one of the following: application-specific stream mapping; 5-tuple indication; and media stream description attribute. The 5-tuple may describe an IP stream that includes an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. The media description under SDP may be associated with a media stream of one or more UUIDs provided in the media description attribute of the SDP, etc.
[0184] In some embodiments, the processor is configured to cause the device to receive media configuration through at least one of the following: an application interface (preferably located between the application service provider 'ASP' of the device and the split-rendering-aware application of the second device); a control plane interface (which is preferably located between the real-time communication application function 'RTC AF' of the device and the media session handler 'MSH' of the second device); and a user plane interface (which is preferably located between the split-rendering server 'SRS' of the device and the split-rendering client 'SRC' of the second device). The control plane interface may be the RTC-5 reference interface in 5GS. The user plane interface may be SR-4, RTC-4, or the SR-4m media center interface may be used.
[0185] In some embodiments, the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs that includes: the H.264 video codec specification; the H.265 video codec specification; the H.266 video codec specification; and the AV1 video codec specification. Other video codec specifications that rely on a portion of the above specifications may also apply, for example, OMAF, V-PCC, etc. OMAF encodes omnidirectional video content and OMAF includes at least one video stream encoded using AVC and HEVC. V-PCC encodes 3D video in 2D projection and V-PCC includes a video stream encoding projected 2D plane frames using AVC and HEFC.
[0186] In some embodiments, the video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more supplementary enhancement information 'SEI' messages of the 'user data unregistered' type; and / or the video codec includes an AV1 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more metadata open bitstream units 'OBU' of the 'unregistered user private data' type. For the SEI message, the payload type can be 5. Additionally, the SEI message can be prefixed or suffixed to the NALU. For the OBU, the OBU metadata type can be X, where X can be any one of 6 - 31.
[0187] In some embodiments, one or more data units of multimedia immersion and interaction data include: immersion and interaction data selected from the list including: user viewpoint data; user field of view data; user pose / orientation data; user gesture tracking data; user body tracking data; user facial feature tracking data; user actions and / or user input data; split rendering pose and spatial information; and augmented reality object representations, where the augmented reality object representations include at least one of the following: a graphical description of the object, and an object position anchor.
[0188] In some embodiments, one or more data units of multimedia immersion and interaction data include extended reality 'XR' multimedia immersion and interaction data.
[0189] In some embodiments, one or more data units of multimedia immersion and interaction data are from one or more media sources selected from a list of media sources including: one or more physical dedicated controllers; one or more red, green, blue 'RGB' cameras; one or more RGB-depth 'RGBD' cameras; one or more infrared 'IR' cameras; one or more microphones; and one or more tactile sensors.
[0190] In some embodiments, the processor is configured to cause the device to use one or more data units of multimedia immersion and interaction data for split rendering of a video media stream, for the purpose of: performing post-reprojection at the split rendering client 'SRC' for displaying one or more rendered frames that match the latest user pose or orientation information.
[0191] In some embodiments, the processor is configured to cause the device to receive a media configuration at the initiation of a multimedia split rendering content delivery session and / or when updating a multimedia split rendering content delivery session.
[0192] In some embodiments, the media configuration is based on at least one of the following: an application configuration of an application hosted on the device, where the application configuration includes one or more media capabilities of the device; and / or an Application Service Provider 'ASP' configuration, which is supplied to at least one of the following: a Real-Time Communication Application Function 'RTC AF'; a provisioning function;
[0193] a Split Rendering Application Function 'SR AF'; and / or a configuration function. The application configuration may be pre-configured with a supported UUID; and / or if a UUID is received, transmission may be automatically enabled.
[0194] In some embodiments, the first device is a network entity / node, and the device is a UE.
[0195] Figure 12 An embodiment of a method 1200 for wireless communication is illustrated.
[0196] A first step 1210 includes: receiving, from a first device, a media configuration for split rendering of a video media stream, where the media configuration includes one or more parameters that indicate that the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video-encoded metadata.
[0197] A further step 1220 includes: using the media configuration to configure the device to receive the video media stream.
[0198] A further step 1230 includes: decoding the encoded video stream of the video media stream, where the decoding includes extracting non-video-encoded metadata.
[0199] A further step 1240 includes: consuming one or more data units of multimedia immersion and interaction data from the non-video-encoded metadata.
[0200] In some embodiments, the method 1200 may be executed by a processor executing program code, such as, for example, a microcontroller, a microprocessor, a CPU, a GPU, a co-processing unit, an FPGA, etc.
[0201] In some embodiments, the non-video-encoded metadata includes: a first field that includes: an identifier for a syntax and semantic representation format of one or more data units of multimedia immersion and interaction data; and a second field that includes: one or more data units of multimedia immersion and interaction data, where the one or more data units of multimedia immersion and interaction data are encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
[0202] In some embodiments, the identifier is indicated by a Universally Unique Identifier 'UUID'.
[0203] In some embodiments, the UUID is unique for a particular application or session, or the UUID is globally unique.
[0204] The UUID can conform to the ISO-IEC-11578 Annex A format and the ISO-IEC 9834-8 version 4 UUID, i.e., randomly generated UUIDs, etc. A 'Session' includes a collection of temporary and interactive (i.e., updatable) configurations and rules, e.g., media formats and codecs, network configurations, and this collection of configurations and rules determines the exchange of information (including media content) between two or more endpoints connected via a network connection.
[0205] In some embodiments, the media configuration is for a plurality of video media streams, the plurality of video media streams including: corresponding encoded video streams, and media configurations for the plurality of video media streams of corresponding non-video encoded metadata, where one or more parameters of the media configuration map an identifier of a first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata.
[0206] In some embodiments, one or more parameters of the media configuration map the first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata based on at least one of the following: application-specific stream mapping; 5-tuple indication; and media stream description attributes.
[0207] The 5-tuple can describe an IP stream, which includes an IP source address, an IP destination address, a source port, a destination port, and a protocol identifier. The media description under SDP can include an association with a media stream of description attributes provided in the SDP media description attributes, etc.
[0208] In some embodiments, receiving the media configuration includes receiving via at least one of the following: an application interface (preferably located between the application service provider 'ASP' of the device and the split-rendering-aware application of the second device); a control plane interface (which is preferably located between the real-time communication application function 'RTC AF' of the device and the media session handler 'MSH' of the second device); and a user plane interface (which is preferably located between the split-rendering server 'SRS' of the device and the split-rendering client 'SRC' of the second device).
[0209] In some embodiments, the control plane interface can be the RTC-5 reference interface in 5GS. The user plane interface can be SR-4, RTC-4, or the SR-4m media center interface can be used.
[0210] In some embodiments, decoding uses at least one video codec selected from a list of video codecs, the list of video codecs including: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification. Some other video codec specifications that rely on a part of the above specifications may also be applicable, for example, OMAF, V-PCC, etc. OMAF encodes omnidirectional video content, and OMAF includes at least one video stream encoded using AVC and HEVC. V-PCC encodes 3D video of 2D projections, and V-PCC includes a video stream that encodes the projected 2D plane frames by means of AVC and HEFC.
[0211] In some embodiments, the video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more supplementary enhancement information 'SEI' messages of the 'user data unregistered' type; and / or the video codec includes an AV1 codec, and non-video coding metadata is encapsulated as the payload of user data in one or more metadata open bitstream units 'OBU' of the 'unregistered user private data' type. For the SEI message, the payload type can be 5. In addition, the SEI message can be prefixed or suffixed to the NALU. For the OBU, the OBU metadata type can be X, where X can be any one of 6 - 31.
[0212] In some embodiments, one or more data units of multimedia immersive and interactive data include: immersive and interactive data selected from a list including the following: user viewpoint data; user field of view data; user pose / orientation data; user gesture tracking data; user body tracking data; user facial feature tracking data; user actions and / or user input data; split rendering pose and spatial information; and augmented reality object representations, the augmented reality object representations including at least one of the following: a graphical description of the object, and an object location anchor.
[0213] In some embodiments, one or more data units of multimedia immersive and interactive data include extended reality 'XR' multimedia immersive and interactive data.
[0214] In some embodiments, one or more data units of multimedia immersive and interactive data come from one or more media sources selected from a list of media sources, the list of media sources including: one or more physical dedicated controllers; one or more red, green, blue 'RGB' cameras; one or more RGB-depth 'RGBD' cameras; one or more infrared 'IR' cameras; one or more microphones; and one or more tactile sensors.
[0215] Some embodiments include using one or more data units of multimedia immersion and interaction data for split rendering of a video media stream, for the purpose of: performing late reprojection at a split rendering client 'SRC' for displaying one or more rendered frames matching the latest user pose or orientation information.
[0216] Some embodiments include: receiving a media configuration at the initiation of, and / or when updating, a multimedia split rendering content delivery session.
[0217] In some embodiments, the media configuration is based on at least one of: an application configuration of an application hosted on the device, where the application configuration includes one or more media capabilities of the device; and / or an application service provider 'ASP' configuration, which is supplied to at least one of: a real-time communication application function 'RTC AF'; a supply function; a split rendering application function 'SR AF'; and / or a configuration function.
[0218] In some embodiments, the application configuration may be pre-configured with a supported UUID; and / or if a UUID is received, transmission may be automatically enabled.
[0219] In some embodiments, the first device is a network entity / node, and the device is a UE.
[0220] Aspects of embodiments of the transmission of interactive and immersive user data on a video elementary stream, as non-video-coded metadata, will now be described in more detail with reference to the video source of an XR application. Hereinafter, terms such as interactive and immersive user data, interactive and immersive metadata, non-video-coded user data, non-video-coded embedded user data, non-video-coded embedded metadata, or simply metadata may be used interchangeably for this purpose.
[0221] Figure 13a FIG. 1300 illustrates an embodiment of a split rendering architecture for interactive and immersive multimedia applications, which is supported by split rendering at the edge. Embodiment 1300 shows a UE 1310, an EN 1320, and a DN 1330 interfaced through various interfaces.
[0222] UE 1310 is illustrated as including: a split rendering aware application client 1311, a MSH 1312 that may include an EEC 1312a, a split rendering client (SRC) 1313, and an XR runtime 1314. The split rendering aware application client 1311 interfaces with the XR runtime 1314. The split rendering aware application client 1311 interfaces with the SRC 1313 through, for example, SR-7. The split rendering aware application client 1311 interfaces with the MSH 1312 through, for example, SR-6 and may interface with the EEC 1312a through EDGE-5. The MSH 1312 interfaces with the SRC 1313 through, for example, SR-7 and SR-6. The SRC 1313 interfaces with the XR runtime 1314.
[0223] EN 1320 is illustrated as including: an SR AF 1321, which includes a configuration function 1321a, a provisioning function 1321b, and an EES 1321c. The EES may be distributed across multiple instances 1321c that interface with each other through, for example, EDGE-9. EN 1320 is also illustrated as including: an RTC AS 1322, which includes a signaling server 1322a and an SRS 1322b, and the SRS 1322b may include an EAS. The SR AF 1321 interfaces with the RTC AS 1322 through, for example, SR-3. The EES 1321c of the SR AF 1321 interfaces with the SRS 1322b that includes an EAS through EDGE-3.
[0224] DN 1330 is illustrated as including an ASP 1331. The ASP 1331 interfaces with the provisioning function 1321b of the SR AF 1321 through SR-1. The ASP 1331 interfaces with the SRS 1322b that includes an EAS through, for example, SR-2.
[0225] The split rendering aware application client 1311 of the UE 1310 interfaces with the ASP 1331 of the DN 1330 through, for example, SR-8. The MSH 1312 of the UE 1310 interfaces with the SR AF 1321 of the EN 1320 through, for example, SR-5. The EEC 1312a of the MSH 1312 may also interface with the EES 1321c of the SR AF 1321 through EDGE 1 and / or EDGE 4 (e.g., via ECS). The SRC 1313 of the UE 1310 interfaces with the SRS 1322b and the signaling server 1322a of the RTC AS 1322 in the EN 1320 through, for example, SR-4 (e.g., composed of SR-4m, SR-4s).
[0226] If FIG. 13 is a split rendering architecture for interactive and immersive multimedia applications, Example 1300 illustrates the architecture, functions, and interfaces. The related split rendering processes (e.g., provisioning, split management, and media delivery) will now be described in conjunction with Figure 13b the description of the related split rendering processes (e.g., provisioning, split management, and media delivery).
[0227] Figure 13b A diagram of the high-level split rendering process in the example is provided. The diagram illustrates a split rendering-aware application 1311, an MSH 1312, an SRC 1313 (which is illustrated as including a scene manager, an XR source manager, and a media access function), an SRS 1322b, an SR AF 1321, an application service provider 1331, and an XR runtime 1314.
[0228] The main steps for split rendering and the associated interfaces are described below, where the split rendering application function (SRAF) is an instantiation of the general real-time communication application function (RTC AF) in some embodiments, see 3GPP Technical Specification 26.506 v1.1.0 titled "5G Real-Time Media Communication Architecture".
[0229] In a first step 1301, the ASP 1331 provisions the SRAF 1321 via a provisioning request for a split rendering management session. In some embodiments, the provisioning may be performed at the SR-1 reference point or alternatively the SR-1 API. The SR AF-1321 determines the split rendering configuration partly based on the information received from the ASP 1331. In some embodiments, the edge-enabled SR AF 1321 may additionally support an edge enabler server (EES). In some implementations, the EES logic may be distributed across multiple edge DNs, where the context and coordination of the distributed EES instances are coordinated via the EDGE-9 interface. The EES may also communicate with and register the EAS split rendering function. In some embodiments, the split rendering server 1322b may represent an instantiation of the EAS for split rendering and may communicate with the SR AF 1321 via the EDGE-3 interface. This may mean, for example, registration / deregistration of the EAS instance, exposure of access network capabilities, QoS management notifications and reports (e.g., bitrate adaptation notifications, etc.). This step 1301 is illustrated as "Split Rendering Provisioning".
[0230] In a further step 1302, the split rendering-aware application 1311 obtains service access information (SAI) from the ASP 1331. The SAI is a collection of parameters and addresses required for the client to: activate the reception of one or more DL / UL media sessions, perform dynamic policy calls, consume / measure reports, and request assistance from the SR AF 1321. In some applications, the obtaining of the SAI can be performed via control plane signaling on SR-5 (or similar RTC-5) from the SR AF 1321 to the MSH 1312, and then the MSH 1312 exposes the obtained SAI to the application 1311 via the SR-6 (or similar RTC-6) interface. In some embodiments, the SR-5 interface can include an EDGE-1 function and an EDGE-4 function, respectively. The EDGE-1 function helps the edge-enabled client instantiated by the MSH1312 to register with the EES at the SR AF 1321, and the EDGE-4 function enables the EEC to access the configuration of edge resources (e.g., EAS configuration information) from the edge configuration server (ECS). In other embodiments, the SAI (as obtained after an SR-1 provisioning call) can be directly exposed by the ASP 1331 via an application-specific SR-8 (or similar RTC-8) interface. This step 1302 is illustrated as "Service Access Information (SAI) Obtaining".
[0231] Next is the split management process. In some embodiments, as illustrated in step 1303a, the split management process can be triggered by the client, i.e., by the split rendering-aware application 1311, and in some other embodiments, the split management process can be triggered by the ASP 1331, as illustrated in step 1303b.
[0232] In the first embodiment where split management is triggered by the client, step 1303a continues. Application 1311 requests split rendering assistance from UE SRC 1313 via the SR-7 API. SRC 1313 communicates with MSH 1312 via the SR-6 API to determine the client media capabilities and functions (e.g., XR runtime API, XR runtime rendering capabilities, etc.) and available SRS instances according to the SAI configuration. SRC 1313 then negotiates the media center split configuration with SRS 1322b at the edge. In some embodiments, this may mean performing split management negotiation via the SDP process or dedicated split rendering management signaling (e.g., as an addition to WebRTC-specific signaling, etc.). For example, in the case of SDP-based signaling, the split is negotiated between SRC 1313 and SRS 1322b based on the SDP protocol, whereby the SDP offer / answer process is used to determine the media streams, media formats, and multiplexing supported by SRC 1315 and SRS 1322b during the split rendering session. After the negotiation of split management, SRS 1322b confirms the split configuration to SRC 1313 via the user plane interface SR-4 (or similar RTC-4), and the split rendering media delivery session is ready to start. As a result, SRC 1313 has completed the application request for split rendering via the SR-7 interface. In this first embodiment, step 1303a is illustrated as "client-driven split management".
[0233] In the second embodiment triggered by SRS1322b for split management, step 1303b continues. SRS1322b requests a split rendering session from SRC 1313 via the SR-4 (or similar RTC-4) interface. In some embodiments, this can be achieved in part through a signaling server function (i.e., implemented by a dedicated signaling server, etc.), while in other embodiments, this can be left to the SDP media session negotiation based on the offer / answer process. After receiving this request, SRC 1313 queries MSH 1312 regarding the UE media capabilities when given the supplied split rendering configuration and available SAI. MSH 1312 replies with the available media capabilities and information to SRC1313 via SR-6. SRC 1313 replies to the SRS1322b request via the SR-4 (or similar RTC-4) interface and negotiates media center split management. Once the negotiation is complete, the determined media configuration, media format, and multiplexing associated with the split configuration profile are returned to SRS1322b as the confirmed split rendering configuration. SRC 1313 then notifies the split rendering-aware application via the SR-7 interface (e.g., via WebSocket or a similar asynchronous mechanism) of the split rendering configuration and the fact that the split rendering session is ready to start. In this second embodiment, step 1303b is illustrated as "network-driven split management".
[0234] Another step 1304 includes media delivery. In some embodiments, this includes SRC 1313 establishing a media delivery split rendering session. In this case, SRC 1313 can request SRS1322b to create a split rendering session given the split rendering configuration profile determined in the previous step. In some embodiments, SRC 1313 is formed at least by the scene manager shim function, XR source manager, and media access function. Thus, such an SRC 1313 can trigger a request for SRS1322b to create a split rendering session by means of the scene manager function or alternatively by means of the MAF function. SRS1322b responds to confirm the request with a split rendering description of the split rendering configuration profile for this session. SRC1313 establishes a connection with SRS1322b and the split rendering session is initiated. In some embodiments, this can apply additional media center signaling given a signaling server on SR-4 (e.g., WebRTC on WebSockets or other signaling mechanisms) as a signaling sub-interface of SR-4. Once the session has been established, the media is considered to be configured in both the UL and DL directions between SRC 1313 and SRS1322b via SR-4. This step 1304 is illustrated as "media delivery establishment".
[0235] Next are additional steps 1305 that utilize the rendering loop. In some embodiments, a typical description of the rendering loop for split rendering includes SRC 1313 receiving or alternatively obtaining pose information and user actions from the XR runtime. In other embodiments, this latter information is supplemented by UL video and / or audio streams (e.g., for AR conferences or AR immersive and interactive applications). The pose, user actions, and any media streams are sent in the UL to SRS 1322b via SR-4m (i.e., the user plane media center sub-interface of SR-4). SRS 1322b processes the received information, information in its own service buffer, and at least partially renders the content of the next one or more frames for split-rendering-aware applications. In one embodiment, SRS 1322b may fuse one or more poses and user actions together to estimate the pose information as close as possible to the expected display time of the frame. In such an embodiment, this estimate is used for rendering, or alternatively, for pre-rendering the next frame. In other embodiments, SRS 1322b may select the latest pose information for the expected display time of the next frame for the rendering / pre-rendering operation. Next, SRS 1322b sends the rendered frame to SRC 1313 via SR-4m. In some embodiments, the rendered frame may additionally include interaction and immersion metadata used by SRS 1322b for rendering the frame (e.g., pose information, user actions / gestures). This data is sent as part of the video elementary stream associated with the split-rendering session together with the video-encoded rendered frame. In one embodiment, this interaction and immersion metadata may be embedded into the video elementary stream as user-specific metadata, e.g., as SEI messages for H.264, H.265, H.266, or alternatively as OBU metadata for AV1. The media stream is then sent by SRS 1322b to SRC 1313. The MAF function and the video codec decode the media stream and further expose the user data embedded in the elementary stream to the XR runtime 1314, or alternatively to the split-rendering-aware application 1311 via the SR-7 interface. In an embodiment, the XR runtime 1314 displays the rendered frame and uses the embedded interaction and immersion metadata for post reprojection and asynchronous time warping to correct any errors between the current pose information of the XR runtime 1314 at the display time and the pose and user action information of SRS 1322b estimated / used during the rendering operation.
[0236] In some embodiments, the payload of the interactive and immersive metadata being communicated can include at least two fields. In a first embodiment, the first field acts as a unique type identifier (i.e., UUID) that determines the syntax and semantics of the corresponding information payload located in the second field. In a second embodiment, the second field carries the interactive and immersive data that is embedded as metadata into the video coding elementary stream. In another embodiment, the syntax and semantics of the second field can be determined at least in part based on the first field. This can mean an index search (e.g., in a list of formats for interactive and immersive metadata supported by a particular communication endpoint such as a UE or alternatively an AS), a data repository or alternatively a registry search (e.g., a query of an Internet-based registry for one or more formats for interactive and immersive metadata), or a selection of a pre-configured resource (e.g., a selection of a format determined by an application for interactive and immersive metadata).
[0237] Figure 14 and Figure 15 provides a high-level overview of example implementations of interactive and immersive metadata for elementary streams of various video codecs (e.g., H.264 / H.265 / H.266).
[0238] Figure 14 Illustrates a representation 1400 of multimedia interactive and immersive user data as metadata within a video coding elementary stream for the MPEG H-26x series of video codecs. For a given NAL unit 1410, the NAL header 1411 and the NAL payload (raw bytes) 1412 are illustrated. The NAL payload 1412 includes a NAL SEI raw byte sequence payload 1413 which itself includes a first SEI message 1413a and a second SEI message 1413b. The first SEI message 1413a is illustrated as "UUID=76994094-c7bd-436b-ac8e-c5205da905cc" and "SEI message" and "interactive and immersive data". The second SEI message 1413b is illustrated as "UUID=8b6d5df6-be48-40f1-814d-20ae061d078d" and "SEI message" and "interactive and immersive data".
[0239] Figure 15Illustrated is the representation 1500 of multimedia interaction and immersive user data as metadata within the video coding elementary stream for an AV1 video codec. For a given OBU 1510 as a metadata OBU, the OBU header 1511 and the OBU payload 1512 are illustrated, where the OBU payload 1512 includes a UUID, which is illustrated as "UUID=76994094-c7bd-436b-ac8e-c5205da905cc" and is further illustrated as "Interactive and Immersive Data".
[0240] In some embodiments, a video decoder (e.g., a video decoder within the MAF in the SRC placed at the UE for DL split rendering service, or alternatively, a video decoder within the ASP in the SRS placed at the SRS for UL split rendering service) may deliver user data metadata (including interactive and immersive data) embedded within a video elementary stream (e.g., H.264, H.265, H.266, AV1, etc.). In another embodiment, the video decoder may expose the user data to other functional boxes (e.g., XR runtime, split rendering aware application, SRS). In one embodiment, this may be achieved based on a functional hook, handler, or callback, or alternatively based on other dedicated interfaces (e.g., raw buffer, event bus, etc.) that expose the interactive and immersive metadata payload based at least in part on a configured set of UUID filters that extract relevant payload information. In one example, the SRC is configured for a split rendering session with a media format on an H.264 video elementary stream that contains interactive and immersive data of the type UUID=76994094-c7bd-436b-ac8e-c5205da905cc. The SRC will filter and expose the corresponding user data SEI message passed by the MAF H.264 decoding instance. In the example, this may be done by the SRC via the SR-7 interface, or alternatively, by an API for any registered consumer, e.g., a split rendering aware application or alternatively the XR runtime. In some embodiments, this may occur after at least one UUID among one or more UUIDs is configured and the feature of interactive and immersive metadata transmission via the video coding elementary stream is enabled.
[0241] In order to enable filtering and output of interactive and immersive user data corresponding to metadata from the video elementary stream, it is thus necessary to signal the UUID used by split rendering and the associated metadata format carried on the media center user plane.
[0242] A UUID can generally be used as a reference for an identifier related to determining the types of interactive and immersive user data and the payload format. If not explicitly specified by an example, the UUID shall be considered a general identifier here.
[0243] Certain embodiments will now be described with reference to the signaling utilized. In particular, the following will be described: application-based signaling embodiments; network-supported signaling embodiments; and SDP-based signaling embodiments. The application-based signaling will be described first.
[0244] In an embodiment, the ASP may signal to a split-rendering-aware application the enabling of the transmission of interactive and immersive user data as metadata on one or more video elementary streams. The media stream is part of the user-plane media center service associated with the DL split-rendering video service (e.g., a split-rendering video stream including monocular or binocular buffering for an XR application), or alternatively, part of the user-plane media center service associated with the UL video service (e.g., one or more UL video streams associated with an AR interactive and immersive application).
[0245] In one example, the application may activate the feature based on a certain application-specific configuration (e.g., a JSON / YAML / XML media format description and configuration) shared between the server (i.e., the ASP application infrastructure) and the client (i.e., the split-rendering-aware application). In one example, such a configuration may contain an option equivalent to interaction-metadata-over-video-es = true to mark the enabling status of the transmission of interactive and immersive user data as metadata on the video-coded elementary stream. Additionally, in some examples, the interface for communicating this configuration may include a proprietary implementation or signaling protocol provided through a reliable communication channel (such as Figure 13a and Figure 13b the SR-8 interface outlined in). Some example implementations may use WebSockets, HTTP methods, SCTP, or other reliable and recognized messaging protocols to communicate this information.
[0246] In another embodiment, the configuration may also contain a list of one or more UUIDs supported by the application, each UUID corresponding to an interactive and immersive user data payload format. In some embodiments, these formats may depend on the XR runtime used by the UE, i.e., corresponding to the SRC in the split-rendering establishment. For example, the OpenXR hand-tracking extension format is associated with This is the case between specific hand-tracking formats for HoloLens 2. In other embodiments, the format can be co-encoded to correspond to one or more XR runtimes, e.g., a collection of OpenXR abstract formats valid for one or more devices. In other embodiments, the UUID can correspond to one or more types of interaction and immersion user data payloads. For example, in some examples, the UUID can correspond to common pose information, e.g., XrPosef of OpenXR, which corresponds to user head tracking, and another UUID can correspond to a collection of the following: composite pose, position, and velocity, each corresponding to a hand joint associated with user hand tracking, e.g., XR_EXT_hand_tracking of OpenXR.
[0247] In other embodiments, the application logic can be static with respect to the UUIDs supported by the application and the XR runtime. In such embodiments, the split-rendering-aware application is configured using a static configuration of the list of supported UUIDs.
[0248] Furthermore, in another embodiment, when at least one UUID is provided for filtering, the feature of the transmission of the interaction and immersion user data of the split-rendering application as metadata on the video encoding elementary stream is automatically enabled.
[0249] In some embodiments, the split-rendering-aware application enabled and configured with the interaction and immersion user data UUIDs can also share its configuration with the SRC. In turn, the SRC can apply the received configuration to filter the corresponding payloads of the interaction and immersion user data by the UUID after video decoding. These data payloads can also be exposed via other interfaces. In one example, the application can use the SR-7 API to configure the SRC with the appropriate configuration on the interaction and immersion user data UUIDs (or alternatively identifiers).
[0250] In another embodiment related to the UL AR split-rendering service, the SRS can request configuration information about the enabling of the transmission of the interaction and immersion user data as metadata on the video elementary stream, and the corresponding UUIDs of such metadata types and payloads, from the SR AF (which can be directly indicated by the ASP, e.g., via the SR-1 interface).
[0251] In one embodiment, an ASP that signals to an application configuration information enabling the transmission of interactive and immersive user data as metadata on a video elementary stream, and the corresponding UUIDs of such metadata types and payloads, may additionally include a mapping of UUIDs to the video elementary stream. In one example, this may be performed by means of an application-specific stream mapping (e.g., based on an application media stream identifier), while in another example, this may be achieved by means of the 5-tuple information of an available media stream (source address (src addr), destination address (dst addr), srd port, dst port, protocol identifier).
[0252] Embodiments related to network-supported signaling will now be described. In an embodiment, one or more application functions may signal to an application the enabling of the transmission of interactive and immersive user data as metadata on one or more video elementary streams. In another embodiment, the signaling of the configuration may additionally include the UUIDs determining the types and payload formats of the interactive and immersive user data to be embedded as metadata on the video elementary stream. In some embodiments, the configuration should be carried on a container that includes at least the following information: a boolean flag indicating the enabling of the transmission of interactive and immersive user data as metadata on one or more video elementary streams (e.g., interaction-metadata-over-video-es = true indicates that the feature is enabled, and interaction-metadata-over-video-es = false indicates that the feature is disabled); a list of one or more UUIDs or alternatively identifiers that determine the types and formats of the payloads of the interactive and immersive user data to be carried on the video-coded elementary stream (e.g., UUID = [ea17c3e6-9be4-4d8d-98bf-b5772b273307, afcd396b-9c96-4c8c-b453-85200febfc09,
[0253] c32dfb22-86fa-46b5-a584-6603a2d7a556], which corresponds to, for example, viewport pose tracking information (e.g., in radians relative to the horizontal and vertical axes as FoV), user pose information (e.g., floating-point representations of 3D vectors for orientation and floating-point representations of 4D quaternions), and user hand tracking information (e.g., a set of objects representing the hands and the corresponding hand joint positions and velocity vectors); where each piece of information also includes its corresponding XR space reference and timestamp to assist in synchronizing with the XR space and time of the XR runtime.
[0254] In an example implementation, such a container can be indicated by the SR AF to the MSH on the UE via the interface SR-5 from the split rendering edge DN. In this case, the information is used to partially indicate the media configuration (e.g., the video codec used, the interactive and immersive types, and the format of the metadata embedded in the video coding elementary stream). In another example, SR-5 can include functions specific to UL or DL media streams corresponding to the M5d interface or the M5u interface of the 5GMS architecture, while in other examples, SR-5 can include functions specific to the RTC-5 interface corresponding to the 5GS real-time communication (RTC) architecture.
[0255] In one embodiment, the MSH containing the configuration information enabling the transmission of interactive and immersive user data as metadata on one or more video elementary streams, and the corresponding UUIDs of such metadata, can expose this configuration to the SRC via an interface. The SRC can in turn request this information when starting or updating a split rendering media delivery session to appropriately process the interactive and immersive user data and further expose the interactive and immersive user data to other functional blocks. In one example, this communication can occur via the SR-6 API to expose the MSH configuration information regarding the transmission and format of interactive and immersive metadata to the SRC. In some examples, such an interface can be implemented via typical HTTP methods, e.g., GET, or any other request-response-based protocol.
[0256] In another embodiment related to UL AR split rendering services, the SRS can request from the SR AF, via an available API (e.g., SR-3), the configuration information enabling the transmission of interactive and immersive user data as metadata on video elementary streams, and the corresponding UUIDs of such metadata types and payloads.
[0257] In one embodiment, the SR AF signaling that signals to the MSH the configuration information enabling the transmission of interactive and immersive user data as metadata on video elementary streams, and the corresponding UUIDs of such metadata types and payloads, can additionally include a mapping of the UUIDs to the video elementary streams. In one example, this can be performed with the 5-tuple information (source address, destination address, srd port, dst port, protocol identifier) of the available media streams.
[0258] Embodiments related to SDP-based signaling will now be described.
[0259] In an embodiment, the split rendering server may signal to the application a list of one or more UUIDs corresponding to the enabling of transmission of one or more types and formats of interactive and immersive user data as metadata on one or more video elementary streams. To this end, the split rendering server signals this configuration information for each video elementary stream, thereby indicating which UUIDs each video-coded elementary stream supports. In addition, the split rendering server signals this configuration information to its corresponding split rendering client in the band via the media center user plane.
[0260] In one embodiment, the protocol for signaling this configuration is SDP, and the signaling of the UUIDs supported between the server and the client and the media negotiation are based on the SDP offer / answer process. The SDP offer / answer is performed before the split rendering content delivery media session is established, or alternatively before an update. In some embodiments, the SRS provides an SDP offer to the SRC. The SRC parses the offer, identifies the SDP attributes associated with the interactive and immersive user data, and determines whether to support the latter based on the parsed values. This may be based in part on information obtained from the MSH before the media session is established, which is about UE media capabilities, and / or split rendering-aware application media capabilities and configurations. To confirm support, the SRC copies the supported UUID configuration corresponding to the interactive and immersive user data types and formats in the SDP answer to the SRS. Thus, this configuration is applied to the DL service of the split rendering content delivery session.
[0261] In some other embodiments, the UL split rendering service may also include a video stream (e.g., an AR application), and the SDP offer / answer process is iterative. To this end, the SRC provides an SDP offer to the SRS, which has the supported UUIDs for video elementary stream embedded interactive and immersive metadata. The SRS processes the SDP offer and then provides its response in the SDP answer. If the SRS accepts the SDP UUID attribute offer, the SRS copies the SDP UUID attribute offer into the corresponding SDP answer, which is sent back to the SRC. The split rendering media delivery session starts from the determined SDP offer / answer negotiation result, and the supported UUID types and the corresponding payloads are embedded as interactive and immersive metadata in the video-coded elementary stream in the UL.
[0262] In one example, the SDP offer / answer procedure between the SRC and the SRS in the DL or alternatively in the UL is performed via the SR-4 user plane interface or alternatively via the RTC-4 user plane interface. In another example, the SDP offer / answer procedure between the SRC and the SRS in the DL or alternatively in the UL is performed via the SR-4m media center user plane interface over 5GS. In other examples, the SDP offer / answer procedure between the SRC and a dedicated signaling server interfacing with the SRS in the DL or alternatively in the UL is performed via the SR-4s signaling center user plane interface over 5GS.
[0263] In one embodiment, 3GPP-specific SDP attributes can be utilized on 5GS implementations to convey the UUID that determines the type and format of the interaction and immersive metadata payloads on the video elementary stream as SDP attributes. For this purpose, one or more SDP attributes can be used to list one or more UUIDs for the video stream. In an example, a UUID can be assigned to more than one video media stream. In one example, the SDP attribute tag for such a UUID is listed in the enhanced Backus-Naur form (ABNF) format as per RFC 4122 and is repeated below for completeness:
[0264]
[0265]
[0266] As described above, the SDP attribute tag for such a UUID can correspond to, for example, an SDP attribute of the form "a=3gpp-interaction-metadata-uuid:20354d7a-e4fe-47af-8ff6-187bca92f3f9". Thus, in some implementations, the UUID format can conform to ISO-IEC-11578 Annex A or to the ISO-IEC 9834-8 format. Thus, in an example, the fields of the UUID listed in the above ABNF description can be derived according to Version 4 UUID of ISO-IEC 9834-8, i.e., randomly derived.
[0267] In another example, a valid SDP list corresponding to either the SDP offer or the SDP answer is provided below:
[0268] m=video 49230 RTP / AVP 96
[0269] a=rtpmap:96 H.264 / 90000
[0270] a = 3gpp - interaction - metadata - uuid:ea17c3e6 - 9be4 - 4d8d - 98bf - b5772b273307 / / Viewport a = 3gpp - interaction - metadata - uuid:afcd396b - 9c96 - 4c8c - b453 - 85200febfc09 / / User Pose a = 3gpp - interaction - metadata - uuid:c32dfb22 - 86fa - 46b5 - a584 - 6603a2d7a556 / / Hand Tracking m = video 49231RTP / AVP 97
[0271] a = rtpmap:97 H.265 / 90000
[0272] a = 3gpp - interaction - metadata - uuid:ea17c3e6 - 9be4 - 4d8d - 98bf - b5772b273307 / / Viewport a = 3gpp - interaction - metadata - uuid:afcd396b - 9c96 - 4c8c - b453 - 85200febfc09 / / User Pose
[0273] In the above example, the indicated UUIDs are determined for each media stream. Thus, the H.264 media stream on port 49230 embeds the UUIDs for viewport pose tracking information, user pose information, and user hand tracking information in the video elementary stream, while the H.265 media stream on port 49231 embeds only the UUIDs for viewport pose tracking information and user pose information in the video elementary stream.
[0274] In one embodiment, once the SDP offer / answer configuration has determined the UUIDs for the media streams and the associated interaction and immersion metadata types and formats, the SRC or alternatively the SRS processes the user data after decoding and further exposes it to other functional blocks (e.g., XR runtime, split - rendering - aware applications, SRS media processing functions).
[0275] The disclosure herein presents the necessary configuration signaling to enable split rendering for XR applications, where interactive and immersive user data is transported over the video elementary stream, and the interactive and immersive user data is in the form of SEI messages (H.264, H.265, H.266) or alternatively OBU metadata (AV1). The solution supports multiple types of metadata based on UUIDs, which identify the type and format of each metadata payload to be carried over the video elementary stream. The proposed signaling mechanism for media configuration is based on three approaches: application layer signaling from the ASP to the application, followed by updating the media configuration to the split rendering client for a split rendering media delivery session; control plane signaling using the split rendering application function to indicate the media configuration to the media session handler, followed by updating the media configuration to the split rendering client for a split rendering media delivery session; user plane signaling by the SDP when establishing a split rendering media delivery session based on new SDP attributes; the new SDP attributes can configure a list of UUIDs identifying the metadata carried over the video elementary stream for each video media stream.
[0276] The problem solved by this disclosure is media configuration signaling for real-time transporting interactive and immersive multimedia data for split rendering of interactive and immersive XR applications. User interaction and immersive data (e.g., pose information, FoV tracking, user actions) serve as inputs for split rendering. However, for efficient display processing on the UE, the XR runtime requires metadata information about the user's pose and input, which is used by the split rendering server to render frames. These are necessary for post-reprojection. For this purpose, a real-time transport mechanism for rendering interactive and immersive metadata and media configuration signaling is necessary for efficiently displaying split-rendered frames.
[0277] This invention solves the problem by leveraging SEI messages and OBU metadata as the transport mechanism for interactive and immersive metadata for the split rendering architecture. The transport relies on the video coding elementary stream, thus grouping the rendered frames with their associated pose information for rendering at the edge. To assist the split rendering client, the media configuration needs to include signaling (e.g., 5-tuple) identifying the type of metadata carried over the video elementary stream and its mapping to the video media stream. Three signaling methods are proposed: i) the application signaling path, ii) the control plane signaling path of the real-time communication system AF, and iii) the user plane signaling based on the SDP offer / answer at the time of RTP session establishment.
[0278] The proposed solution is superior to the data channel solution using the WebRTC SCTP stack because it benefits from the advantages of RTP / SRTP, namely, FEC-based timing, synchronization, jitter management, and reliability. Additionally, the proposed method inherently synchronizes interaction and immersion metadata with one or more video streams for carriage. Any kind of signaling is outside the scope of WebRTC.
[0279] The proposed solution is superior to the RTP header extension solution because the RTP header extension solution does not limit the maximum payload size for interactive and immersive multimedia data types. Additionally, the proposed solution is transmitted as part of the RTP payload, thereby enabling RTP synchronization, jitter management, and reliability by means of FEC. RTP header extension signaling is entirely dependent on the SDP offer / answer process. SDP is a candidate signaling in the proposed disclosure, but a new SDP attribute is introduced, which is orthogonal to any RTP header extension-related signaling solution.
[0280] The proposed solution is a trade-off with respect to the solution of defining a new IETF RTP payload for interactive and immersive multimedia data. This trade-off is mainly to avoid the need to provide a complete RTP payload type specification.
[0281] In particular, the present disclosure provides an embodiment that utilizes application-based signaling. Signaling for renderer interaction and immersion metadata and subsequent mapping to video media streams is performed by an application service provider, which notifies the media configuration to the application client. The application client then correspondingly modifies the split renderer client and the media session handler media configuration via the 5GS exposed APIs of the latter two functional blocks.
[0282] Another embodiment utilizes network-supported signaling. Signaling for renderer interaction and immersion metadata and subsequent mapping to video media streams is performed by the split renderer AF, which notifies the media configuration to the media session handler. The media session handler exposes this configuration to the application client via its API and also to the split renderer client.
[0283] Another embodiment utilizes SDP-based signaling. Before media session establishment, signaling for renderer interaction and immersion metadata and subsequent mapping to video media streams is performed by means of the SDP offer / answer process. A new media description attribute a=3gpp-interaction-metadata-uuid is proposed: <uuid>, for mapping metadata types (identified by UUIDs) to one or more RTP media streams.
[0284] From the perspective of a split rendering server, the disclosure herein provides a method for configuring a multimedia split rendering content delivery session over a network, the method comprising: determining a media configuration that includes a mapping of one or more interactive and immersive multimedia data types to one or more video media streams, wherein the one or more interactive and immersive multimedia data types include non-video encoded data from one or more media sources; signaling the determined media configuration to a remote endpoint; establishing a multimedia split rendering content delivery session configuration, at least in part based on the signaled media configuration, the multimedia split rendering content delivery session configuration including at least one video media stream of at least one video encoded elementary stream, the at least one video encoded elementary stream including one or more mapped interactive and immersive multimedia data as non-video encoded metadata payloads; and utilizing the non-video encoded metadata payloads corresponding to the one or more interactive and immersive multimedia data included by the at least one video encoded elementary stream for split rendering associated with at least one video media stream of at least one video decoded elementary stream.
[0285] In some embodiments, the non-video encoded metadata payload includes at least two fields.
[0286] In some embodiments, the first field indicates a syntax and semantic identifier representation format of the interactive and immersive multimedia data type, and the second field encodes the interactive and immersive multimedia data according to the identifier representation format determined by the first field.
[0287] In some embodiments, the first field includes a Universal Unique Identifier (UUID) indication.
[0288] Some embodiments further include: matching one or more identifier representation formats corresponding to one or more interactive and immersive multimedia data types to a configuration mapping of one or more video media streams.
[0289] Some embodiments further include: performing signaling of the media configuration based on at least one of the following: signaling performed by an Application Service Provider (ASP) to a split rendering-aware application via an application interface; signaling between a Real-Time Communication Application Function (RTCAF) and a media session handler via a control plane interface (e.g., the RTC-5 reference interface between the RTC AF and the MSH in 5GS); and signaling between a split rendering server (SRS) and a split rendering client (SRC) via a user plane interface (e.g., the SR-4 reference interface between the SRS and the SRC in 5GS or alternatively the RTC-4 reference interface).
[0290] Some embodiments also include determining a media configuration based, at least in part, on at least one of the following: an application service provider (ASP) configuration that provisions at least one of a real-time communication application function (RTC AF), a provisioning function, a split rendering application function (SR AF), and a configuration function; and an application configuration of media capabilities that is semi-statically determined based, at least in part, on the hardware capabilities of a corresponding device that processes the logic of the application.
[0291] In some embodiments, at least one video coding elementary stream is determined based on at least one of the following: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
[0292] In some embodiments, the non-video coding metadata payload included in at least one video coding elementary stream includes at least one of the following: a supplementary enhancement information (SEI) message of the user data unregistered type; and a metadata open bitstream unit (OBU) of the unregistered user private data type.
[0293] In some embodiments, the interactive and immersive multimedia data types include one of the following: a user pose data representation (e.g., a timestamped 3D position vector and quaternion representation of an XR space that describes the orientation of a pose object with up to 6 DoF. Such a pose object can correspond to a user body component or part, such as a head, joint, hand, or a combination thereof); a user gesture tracking data representation (i.e., an array of one or more hands tracked according to the pose of one or more hands. For example, according to the OpenXR OpenXR_EXT_hand_tracking API specification, the tracking of each hand additionally includes an array of hand joint positions and velocities relative to the base XR space and XR runtime timestamp); a user body tracking data representation (e.g., a biovisual hierarchy (BVH) encoding of body and body part movements and associated pose objects); a user facial feature tracking data representation (e.g., an array of key points / feature positions, poses, or their encoding into a predefined facial expression class); a collection of one or more user actions (i.e., user actions and inputs to a physical or logical controller defined within the XR space. For example, according to an OpenXR XrAction handle, which captures different user inputs to a controller or HW command unit supported by an OpenXR-compliant AR / VR device); and a collection of one or more augmented reality (AR) object representations, each object including at least one of a graphical description and an associated position anchor (i.e., metadata that determines the position of an object or point in the XR space as an anchor for placing virtual 2D / 3D objects, such as text rendering, static video content, 2D / 3D video content, etc.).
[0294] Some embodiments also include split-rendering utilization of interaction and immersion data for at least one of the following: post-reprojection at a split-rendering client (SRC) for displaying one or more rendered frames that match the latest user pose and orientation information; partial or full pre-rendering of the next one or more frames at a split-rendering server (SRS); and estimating the most likely user pose and orientation at the split-rendering server (SRS), the most likely user pose and orientation being associated with the expected display time of the next one or more frames to be partially or fully pre-rendered.
[0295] In some embodiments, the mapping of interaction and immersion multimedia data types to media streams is based on at least one of the following: 5-tuple indication (e.g., as a 5-tuple describing an IP stream that includes (IP source address, IP destination address, source port, destination port, protocol identifier)), and media stream identifier (e.g., media description under SDP, unique media stream name provided in SDP media description attributes, etc.).
[0296] From the perspective of a split-rendering client, the disclosure herein also provides a method for configuring a multimedia split-rendering content delivery session over a network, the method including: receiving, over the network, a media configuration that includes a mapping of one or more interaction and immersion multimedia data types to one or more video media streams, where the one or more interaction and immersion multimedia data types include non-video-encoded data from one or more media sources; controlling a video decoder to generate a set of one or more information payloads corresponding to the non-video-encoded payloads, where each non-video-encoded payload includes at least two fields; and processing the one or more information payloads into one or more samples of interaction and immersion multimedia data generated by one or more media sources.
[0297] In some embodiments, the non-video-encoded metadata payload includes at least two fields.
[0298] In some embodiments, the first field indicates a syntax and semantic identifier representation format for the interaction and immersion multimedia data type, and the second field encodes the interaction and immersion multimedia data according to the identifier representation format determined by the first field.
[0299] In some embodiments, the first field includes a Universally Unique Identifier (UUID) indication.
[0300] Some embodiments also include: matching one or more identifier representation formats corresponding to one or more interaction and immersion multimedia data types to a configuration mapping of one or more video media streams.
[0301] Some embodiments further include: receiving media configuration based on at least one of the following: signaling performed by an application service provider (ASP) to a split rendering-aware application via an application interface; signaling between a real-time communication application function (RTCAF) and a media session handler via a control plane interface (e.g., the RTC-5 reference interface between RTC AF and MSH in 5GS); and signaling between a split rendering server (SRS) and a split rendering client (SRC) via a user plane interface (e.g., the SR-4 reference interface or alternatively the RTC-4 reference interface between SRS and SRC in 5GS).
[0302] Some embodiments further include: media configuration is based in part on at least one of the following: ASP configuration that provisions at least one of a real-time communication application function (RTC AF), a provisioning function, a split rendering application function (SR AF), and a configuration function; and application configuration of media capabilities that is semi-statically determined based in part on the hardware capabilities of a corresponding device that processes the logic of the application.
[0303] In some embodiments, at least one video coding elementary stream is determined based on at least one of the following: the H.264 video codec specification; the H.265 video codec specification; the H.266 video codec specification; and the AV1 video codec specification.
[0304] In some embodiments, the non-video coding metadata payload included in at least one video coding elementary stream includes at least one of the following: a supplementary enhancement information (SEI) message of the user data unregistered type; and a metadata open bitstream unit (OBU) of the unregistered user private data type.
[0305] In some embodiments, the interactive and immersive multimedia data types include one of the following: user pose data representation (e.g., a timestamped 3D position vector and quaternion representation of an XR space describing the orientation of a pose object with up to 6 DoF. Such a pose object can correspond to a user body component or part, such as the head, joints, hands, or a combination thereof); user gesture tracking data representation (i.e., an array of one or more hands tracked according to the pose of one or more hands, e.g., according to the OpenXR OpenXR_EXT_hand_tracking API specification, the tracking of each hand additionally includes an array of hand joint positions and velocities relative to the base XR space and XR runtime timestamp); user body tracking data representation (e.g., biovisual hierarchical (BVH) encoding of body and body part movements and associated pose objects); user facial feature tracking data representation (e.g., as an array of key points / feature positions, poses, or their encoding into predefined facial expression classes); a set of one or more user actions (i.e., user actions and inputs to physical or logical controllers defined within the XR space, e.g., according to an OpenXR XrAction handle, which captures different user inputs to controllers or HW command units supported by an OpenXR-compliant AR / VR device); and a set of one or more augmented reality (AR) object representations, each object including at least one of a graphic description and an associated location anchor (i.e., metadata that determines the position of an object or point in the XR space, serving as an anchor for placing virtual 2D / 3D objects, such as text rendering, static video content, 2D / 3D video content, etc.).
[0306] Some embodiments also include split rendering utilization of the interactive and immersive data for at least one of the following: post-reprojection at a split rendering client (SRC) for displaying one or more rendered frames that match the latest user pose and orientation information; partial or full pre-rendering of the next one or more frames at a split rendering server (SRS); and estimating the most likely user pose and orientation at the split rendering server (SRS), the most likely user pose and orientation being associated with the expected display time of the next one or more frames to be partially or fully pre-rendered.
[0307] In some embodiments, the mapping of the interactive and immersive multimedia data types to media streams is based on at least one of the following: 5-tuple indication (e.g., as a 5-tuple describing an IP stream that includes (IP source address, IP destination address, source port, destination port, protocol identifier)), and media stream identifier (e.g., media description under SDP, unique media stream name provided in SDP media description attributes, etc.).
[0308] The content of the present disclosure relates to using SEI messages / OBU metadata for transmitting interaction and immersion metadata associated with XR applications supported by split rendering, as well as the signaling required to enable such applications over 5GS.
[0309] It should be noted that the above methods and apparatuses illustrate rather than limit the invention, and those skilled in the art will be able to design many alternative arrangements without departing from the scope of the appended claims. The word "comprising" does not exclude the presence of other elements or steps than those listed in the claims, "a" or "an" does not exclude a plurality, and a single processor or other unit may implement the functions of several units recited in the claims. Any reference signs in the claims should not be construed as limiting their scope.
[0310] Furthermore, although examples are given in the context of a specific communication standard, these examples are not intended to limit the communication standards to which the disclosed methods and apparatuses can be applied. For example, while specific examples are given in the context of 3GPP, the principles disclosed herein can also be applied to another wireless communication system and, in fact, any communication system that uses routing rules.
[0311] The method can also be embodied in a set of instructions stored on a computer-readable medium, which, when loaded into a computer processor, a digital signal processor (DSP), etc., causes the processor to execute the above method.
[0312] Where mentioned, the OpenXR specification is described in the git reference version 1.0.26, the RTP payload format media type is available in the IANA standard RFC4855, and the SDP offer / answer model is described in the RFC standard 3264 titled "Offer / Answer Model with the Session Description Protocol (SDP)". In addition, the following 3GPP technical specifications Tdocs are relevant to the disclosure herein: S4aR220052 titled "Real-Time Metadata Transmission over RTP"; S4aR220053 titled "Real-Time Metadata Transmission over the Data Channel"; S4-221557 titled "Real-Time Metadata Transmission over the Data Channel"; S4-221555 titled "Real-Time Metadata Transmission over RTP"; and S4-230359 titled "Signaling Rendering Pose and Other Related Information".
[0313] The described methods and apparatuses may be practiced in other specific forms. The described methods and apparatuses are to be considered in all respects only illustrative and not restrictive. Thus, the scope of the invention is indicated by the appended claims rather than by the foregoing description. All changes within the meaning and equivalent scope of the claims should be included within their scope.
[0314] The following abbreviations are relevant to the field covered by this document: 3GPP, 3rd Generation Partnership Project; 5G, 5th Generation; 5GS, 5G System; 5QI, 5G QoS Identifier; AF, Application Function; AMF, Access and Mobility Management Function; AR, Augmented Reality; DL, Downlink; DTLS, Datagram Transport Layer Security; NAL, Network Abstraction Layer; NALU, NAL Unit; OBU, Open Bitstream Unit; PCF, Policy Control Function; PDU, Packet Data Unit; PPS, Picture Parameter Set; QoE, Quality of Experience; QoS, Quality of Service; RAN, Radio Access Network; RTCP, Real-Time Control Protocol; RTP, Real-Time Protocol; SDAP, Service Data Adaptation Protocol; SEI, Supplemental Enhancement Information; SMF, Session Management Function; SRTCP, Secure Real-Time Control Protocol; SRTP, Secure Real-Time Protocol; SR AF, Split Rendering AF; SRC, Split Rendering Client; SRS, Split Rendering Server; TLS, Transport Layer Security; UE, User Equipment; UL, Uplink; UPF, User Plane Function; VCL, Video Coding Layer; VMAF, Video Multimethod Assessment Fusion; VPS, Video Parameter Set; VR, Virtual Reality; WebRTC, Web Real-Time Communication; XR, Extended Reality; XR AS, XR Application Server; and XRM, XR Media.< / uuid>
Claims
1. An apparatus for wireless communication, comprising: a processor; and a memory coupled to the processor, the processor being configured to cause the apparatus to: determine a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate that the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata; signal the media configuration to a second apparatus; establish, at least in part based on the media configuration, a multimedia split rendering content delivery session including the video media stream with the second apparatus; and use the one or more data units of multimedia immersion and interaction data for split rendering of the video media stream.
2. The apparatus according to claim 1, wherein the non-video encoded metadata includes: a first field that includes: an identifier for a syntax and semantic representation format of the one or more data units of multimedia immersion and interaction data; and a second field that includes: the one or more data units of multimedia immersion and interaction data, the one or more data units of multimedia immersion and interaction data being encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
3. The apparatus according to claim 2, wherein the identifier is indicated by a Universally Unique Identifier (UUID).
4. The apparatus according to claim 3, wherein the UUID is unique for a specific application or session, or the UUID is globally unique.
5. The apparatus according to any one of claims 2 to 4, wherein the processor is configured to cause the apparatus to determine the media configuration by causing the apparatus to: determine the media configuration for a plurality of video media streams, the plurality of video media streams having respective encoded video streams and respective non-video encoded metadata, wherein the one or more parameters of the media configuration map the identifier of the first field of the respective non-video encoded metadata to the respective video media stream of the respective non-video encoded metadata.
6. The apparatus according to claim 5, wherein the one or more parameters of the media configuration map the first field of the respective non-video encoded metadata to the respective video media stream of the respective non-video encoded metadata based on at least one of: application-specific stream mapping; 5-tuple indication; and media stream description attributes.
7. The apparatus according to any one of the preceding claims, wherein the processor is configured to cause the apparatus to signal the media configuration by at least one of: an application interface; a control plane interface; and a user plane interface.
8. The apparatus according to any one of the preceding claims, wherein the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs, the list of video codecs including: H.264 Video Codec Specification; H.265 Video Codec Specification; H.266 Video Codec Specification; and AV1 Video Codec Specification.
9. The apparatus according to claim 8, wherein: the video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and the non-video-encoded metadata is encapsulated as the payload of user data in one or more Supplementary Enhancement Information (SEI) messages of the unregistered user data type; and / or the video codec includes an AV1 codec, and the non-video-encoded metadata is encapsulated as the payload of user data in one or more Metadata Open Bitstream Units (OBUs) of the unregistered user private data type.
10. The apparatus according to any one of the preceding claims, wherein the one or more data units of the multimedia immersion and interaction data comprise: Immersive and interactive data selected from the list comprising: user viewpoint data; user field of view data; user pose / orientation data; user gesture tracking data; user body tracking data; user facial feature tracking data; user action and / or user input data; split rendering pose and spatial information; and augmented reality object representation, the augmented reality object representation including at least one of the following: a graphical description of the object, and an object location anchor.
11. The apparatus according to any of the preceding claims, wherein the processor is further configured to cause the apparatus to determine the media configuration based on at least one of the following: an application configuration of an application hosted on the second device, wherein the application configuration includes one or more media capabilities of the second device; and / or an Application Service Provider (ASP) configuration, the ASP configuration being supplied to at least one of the following: a Real-Time Communication Application Function (RTC AF); a provisioning function; a Split Rendering Application Function (SR AF); and / or a configuration function.
12. A device for wireless communication, comprising: a processor; and a memory coupled to the processor, the processor being configured to cause the device to: receive, from a first device, a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters indicating that the encoded video stream of the video media stream includes one or more data units of multimedia immersive and interactive data as non-video-encoded metadata; configure the device to receive the video media stream using the media configuration; decode the encoded video stream of the video media stream, wherein the decoding includes: extracting the non-video-encoded metadata; and consume the one or more data units of multimedia immersive and interactive data from the non-video-encoded metadata.
13. The apparatus according to claim 12, wherein the non-video-encoded metadata includes: a first field, the first field including: an identifier of a syntax and semantic representation format for the one or more data units of multimedia immersive and interactive data; and A second field, the second field including: the one or more data units of multimedia immersion and interaction data, the one or more data units of multimedia immersion and interaction data being encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
14. The apparatus according to claim 13, wherein the identifier is indicated by a Universally Unique Identifier (UUID).
15. The apparatus according to claim 14, wherein the UUID is unique for a specific application or session, or the UUID is globally unique.
16. The apparatus according to any one of claims 13 to 15, wherein the media is configured for a plurality of video media streams, the plurality of video media streams including: A corresponding encoded video stream, and corresponding non-video encoded metadata, wherein the one or more parameters of the media configuration map the identifier of the first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata.
17. The apparatus according to claim 16, wherein the one or more parameters of the media configuration map the first field of the corresponding non-video encoded metadata to the corresponding video media stream of the corresponding non-video encoded metadata based on at least one of the following: Application-specific stream mapping; 5-tuple indication; and Media stream description attributes.
18. The apparatus according to any one of claims 12 to 17, wherein the processor is configured to cause the apparatus to receive the media configuration through at least one of the following: An application interface; A control plane interface; and A user plane interface.
19. The apparatus according to any one of claims 12 to 18, wherein the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs, the list of video codecs including: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; And AV1 video codec specification.
20. The apparatus according to claim 19, wherein: The video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and the non-video encoded metadata is encapsulated as the payload of user data in one or more Supplementary Enhancement Information (SEI) messages of the unregistered user data type; and / or The video codec includes an AV1 codec, and the non-video encoded metadata is encapsulated as the payload of user data in one or more Metadata Open Bitstream Units (OBUs) of the unregistered user private data type.
21. The apparatus according to any one of claims 12 to 20, wherein the one or more data units of the multimedia immersion and interaction data include immersion and interaction data selected from a list including the following: User viewpoint data; User field of view data; User pose / orientation data; User gesture tracking data; User body tracking data; User facial feature tracking data; User action and / or user input data; Split rendering pose and spatial information; And An augmented reality object representation, the augmented reality object representation including at least one of the following: a graphical description of the object, and an object location anchor.
22. The apparatus according to any one of claims 12 to 21, wherein the media configuration is based on at least one of the following: The application configuration of the application hosted on the second device, where the application configuration includes: One or more media capabilities of the second device; and / or An application service provider (ASP) configuration, the ASP configuration being supplied to at least one of the following: A real-time communication application function (RTC AF); A supply function; A split rendering application function (SR AF); and / or A configuration function.
23. A method for wireless communication, comprising: Determining a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate that the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata; Signaling the media configuration to a second device; Establishing, at least in part based on the media configuration, a multimedia split rendering content delivery session with the second device that includes the video media stream; And Using the one or more data units of multimedia immersion and interaction data for split rendering the video media stream.
24. The method according to claim 23, wherein the non-video encoded metadata includes: A first field that includes: an identifier for a syntax and semantic representation format of the one or more data units of multimedia immersion and interaction data; and A second field that includes: the one or more data units of multimedia immersion and interaction data, the one or more data units of multimedia immersion and interaction data being encoded according to the syntax and semantic representation format corresponding to the identifier of the first field.
25. The method according to claim 24, wherein the identifier is indicated by a universally unique identifier (UUID).
26. The method according to claim 25, wherein the UUID is unique for a specific application or session, or the UUID is globally unique.
27. The method according to any one of claims 24 to 26, wherein determining the media configuration includes: Determining the media configuration for a plurality of video media streams, the plurality of video media streams including: respective encoded video streams, and respective non-video encoded metadata, wherein the one or more parameters of the media configuration map the identifier of the first field of the respective non-video encoded metadata to the respective video media stream of the respective non-video encoded metadata.
28. The method according to claim 27, wherein the one or more parameters of the media configuration map the first field of the respective non-video encoded metadata to the respective video media stream of the respective non-video encoded metadata based on at least one of the following: An application-specific stream mapping; A 5-tuple indication; and A media stream description attribute.
29. The method according to any one of claims 23 to 28, wherein signaling the media configuration comprises: Signaling by at least one of the following: An application interface; A control plane interface; And User plane interface.
30. The method according to any one of claims 23 to 29, wherein the encoded video stream is encoded / decoded using at least one video codec selected from a list of video codecs, the list of video codecs including: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
31. The method according to claim 30, wherein: the video codec includes: an H.264 codec, an H.265 codec, or an H.266 codec, and the non-video encoded metadata is encapsulated as the payload of user data in one or more supplementary enhancement information (SEI) messages of the unregistered user data type; and / or the video codec includes an AV1 codec, and the non-video encoded metadata is encapsulated as the payload of user data in one or more metadata open bitstream units (OBUs) of the unregistered user private data type.
32. The method according to any one of claims 23 to 31, wherein the one or more data units of the multimedia immersion and interaction data comprise: Immersion and interaction data selected from a list including: User viewpoint data; User field of view data; User pose / orientation data; User gesture tracking data; User body tracking data; User facial feature tracking data; User action and / or user input data; Split rendering pose and spatial information; and Augmented reality object representation, the augmented reality object representation including at least one of the following: a graphical description of the object, and an object location anchor.
33. The method according to any one of claims 23 to 32, wherein determining the media configuration comprises: Determine the media configuration based on at least one of the following: The application configuration of the application hosted on the second device, wherein the application configuration includes: one or more media capabilities of the second device; and / or Application service provider (ASP) configuration, the ASP configuration being supplied to at least one of the following: Real-time communication application function (RTC AF); Provisioning function; Split rendering application function (SR AF); and / or Configuration function.
34. A method for wireless communication, comprising: Receiving, from a first device, a media configuration for split rendering of a video media stream, wherein the media configuration includes one or more parameters that indicate: the encoded video stream of the video media stream includes one or more data units of multimedia immersion and interaction data as non-video encoded metadata; Configuring the device to receive the video media stream using the media configuration; Decoding the encoded video stream of the video media stream, wherein the decoding includes: extracting the non-video encoded metadata; and Consuming the one or more data units of multimedia immersion and interaction data from the non-video encoded metadata.
35. The method according to claim 34, wherein the non-video encoded metadata includes: A first field, the first field including: an identifier for the syntax and semantic representation format of the one or more data units of multimedia immersion and interaction data; and Second field, the second field comprising: said one or more data units of multimedia immersion and interaction data, said one or more data units of multimedia immersion and interaction data being encoded according to said syntax and semantic representation format corresponding to the identifier of said first field.
Citation Information
Patent Citations
Centrifuge deposition device and continuous slab mold for processing polymeric-foam-generating liquid reactants
US4221555A