Method and apparatus for split rendering using multiple codecs
Patent Information
- Application Number
- GB2025002311
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2026-09-16
Smart Images

Figure 00000000_0000_ABST 
Figure 00000000_0001_ABST
Abstract
Description
[0001] This disclosure generally relates to communication networks that provide real-time communication sessions for split rendering. More particularly, this disclosure provides a method and apparatus for split rendering using multiple video / image codecs. Background
[0002] A mobile telecommunication network or cellular network (generally referred to herein as a communication network) enables communications between two or more communication devices, provides communication devices access to, delivers services provides by third-party applications to communication devices, or provides services offered by the communication network to communication devices.
[0003] A communication network and communication devices may operate in accordance with cellular technologies (otherwise referred to as radio access technologies), such as GSM, UTMS, LTE, LTE-A, and NR. Cellular technologies are standardized by various standards organization, such as the Third Generation Partnership Project (3GPP) or ETSI (European Telecommunications Standards Institute). 3GPP is currently developing standards for 5th generation cellular technologies (generally referred to a 5G or NR standards) and 6th generation cellular technologies (generally referred to a 6G standards). Communication networks that operate in accordance with 5G or NR standards are generally referred to as 5 G networks and communication networks that operate in accordance with 6G standards are generally referred to as 6G networks.
[0004] A communication network (e.g., a 5G network or a 6G network) includes access networks (e.g., radio access networks) that can communicate wirelessly with one or multiple communication devices by sharing available resources (e.g., bandwidth, transmit power, etc.) of the access network (e.g., radio access network). A communication network can also establish reliable, secure connectivity between communication devices and a data network. A communication network (e.g., a 5G network) may provide enhanced mobile broadband services (e.g., telephony, video, data, short message services messaging services), ultrareliable low-latency communication services (e.g., for extended reality XR services), or massive machine type communication services to communication devices.
[0005] A communication network that provides ultra-reliable low-latency communication services may facilitate real-time communications between a communication device and a provider of a XR service (generally referred to as an XR service provider). Such a communication network may enable a user equipment receiving XR data of an XR service to offload computationally intensive tasks of processing XR data of an XR service received from an XR service provider to a computing system deployed at an edge of the communication network (e.g., deployed in a data network located near the communication network or deployed near an access network of the communication network). For example, a user equipment receiving XR data from a XR service provider may offload rendering of a computer graphics scene to a computing system deployed at an edge of the computing system (generally referred to as an edge computing system) to reduce energy consumed by the user equipment when rendering an augmented reality or a virtual reality scene. Summary
[0006] This disclosure relates to a split rendering client of a device communicating with a split rendering server via a communication network to improve split rendering of a computer graphics scene of an XR service provided by an XR service provider to the user equipment.
[0007] In accordance with a first aspect of this disclosure, there is provided a device comprising at least one processor; and at least one memory storing instructions of a split rendering client which, when executed by the at least one processor, cause the device to perform: sending, to a split rendering server, information indicative of video / image codecs supported by the split rendering client; receiving, from the split rendering server, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames; sending, to the split rendering server, metadata comprising at least data indicative of a predicted pose of a user of the device; receiving, from the split rendering server, multiple encoded frames, wherein each respective encoded frame of the multiple encoded frames comprises a rendered composition layer of multiple rendered composition layers of a scene; for each respective encoded frame, decoding the respective encoded frame to obtain the respective rendered composition layer using one video / image codec of the one or more video / image codec, the one video / image codec corresponding to a video / image codec of the SRS that is used to encode the respective rendered composition layer; and forwarding each respective rendered composition layer to an extended reality (XR) runtime of the user equipment for composition of the scene and rendering of the scene.
[0008] The memory of the device may further store instructions of an extended reality runtime which, when executed by the at least one processor, cause the device to perform receiving each respective rendered composition layer; and composing the scene using each respective rendered composition layer; and rendering the scene on a display of the device.
[0009] The composing may comprise superimposing a first rendered composition layers on a second respective composition layer.
[0010] The metadata may further comprise at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device.
[0011] The instructions of the split rendering client, when executed by the at least one processor, may further cause the device to perform receiving, from the split rendering server, signaling comprising a relative drawing order of the multiple rendered composition layers.
[0012] The instructions of the split rendering client, when executed by the at least one processor, may further cause the device to perform receiving, from the split rendering server, signaling comprising a relative depth order of the multiple rendered composition layers.
[0013] The instructions of the split rendering client, when executed by the at least one processor, may further cause the device to perform sending, to the split rendering server, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
[0014] The device may be a user equipment or a head-mounted display.
[0015] In accordance with a second aspect of this disclosure, there is provided a split rendering server comprising at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the split rendering server to perform: receiving, from a split rendering client, information indicative of video / image codecs supported by the split rendering client; sending, to the split rendering client, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames; receiving, from the split rendering client, metadata comprising at least data indicative of a predicted pose of a user of device; rendering multiple composition layers of a scene for the predicted pose; for each respective rendered composition layer, encoding, using one video / image codec of the split rendering server, the respective rendered composition layer to generate an encoded frame comprising the respective the one video / image codec corresponding to one of the video / image codecs supported by the split rendering client; and sending each respective encoded frame to the split rendering client.
[0016] The metadata may further comprise at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device
[0017] The instructions, when executed by the at least one processor, may further cause the split rendering client to perform: sending to the split rendering client, signaling comprising a relative drawing order of the multiple rendered composition layers.
[0018] The instructions, when executed by the at least one processor, may further cause the split rendering client to perform: sending to the split rendering client, signaling comprising a relative depth order of the multiple rendered composition layers.
[0019] The instructions, when executed by the at least one processor, may further cause the split rendering client to perform: receiving, from the split rendering client, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
[0020] The instructions, when executed by the at least one processor, may further cause the split rendering client to perform: determining the split rendering configuration comprising the indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames based on the information indicative of video / image codecs supported by the split rendering client and the information indicative of a decoding latency of each respective video / image decoder supported by the split rendering client.
[0021] The above summary provides a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the present disclosure. Nor is it intended to be used to limit the scope of the present disclosure. Other aspects and features of the present disclosure will become apparent to those of ordinary skill in the art upon review of the following description of example implementations in conjunction with the accompanying figures. Brief Description of the Drawings
[0022] Reference will now be made, by way of example, to the accompanying drawings which show example implementations of the present application, and in which:
[0023] FIG. lisa schematic block diagram illustrating a communication network that connects a user equipment to a data network that includes a server hosting an application provider that provides extended reality services to the user equipment in accordance with an example implementation.
[0024] FIG. 2 is a schematic diagram of a MSE split rendering architecture in which a communication network connects a user equipment to a data network that includes one or more servers of an application provider that provides extended reality services to the user equipment;
[0025] FIG. 3 shows a diagram of a split rendering procedure that involves a split rendering client and a split rendering server according to an example implementation;
[0026] FIG. 4 is a flowchart of a method according to an example implementation;
[0027] FIG. 5 is a flowchart of a method according to an example implementation;
[0028] FIG. 6 is a schematic diagram illustrating physical and logical components of a split rendering server in accordance with an example implementation.
[0029] FIG. 7 is a schematic diagram illustrating physical and logical components of a user equipment comprising a split render client in accordance with an example implementation.
[0030] Similar reference numerals may have been used in different figures to denote similar components. Unless otherwise specifically noted, articles depicted in the drawings are not necessarily drawn to scale. Detailed Description
[0031] The subject matter is described herein with reference to the accompanying drawings, in which example implementations are shown. However, many different example implementations may be used, and thus the description should not be construed as limited to the implementations set forth herein. Rather, these example implementations are provided so that this application will be thorough and complete. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same elements, and prime notation is used to indicate similar elements, operations or steps in alternative example implementations. Separate boxes or illustrated separation of logical elements of illustrated systems and devices does not necessarily require physical separation of such logical elements, as communication between such logical elements may occur by way of messaging, function calls, shared memory space, and so on, without any such physical separation. As such, logical elements need not be implemented in physically or logically separated platforms, although such logical elements are illustrated separately for ease of explanation herein. Different devices may have different designs, such that although some devices implement some logical elements in hardware, other devices may implement such logical elements in a programmable processor with code obtained from a machine-readable medium. Lastly, elements referred to in the singular may be plural and vice versa, except wherein indicated otherwise either explicitly or inherently by context.
[0032] References in the present disclosure to “one implementation,” “an implementation,” “an example implementation,” and the like indicate that the implementation described may include a particular feature, structure, or characteristic, but it is not necessary that every implementation_includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same implementation. Further, when a particular feature, structure, or characteristic is described in connection with an implementation, it is submitted that it is within the knowledge of one skilled in the art to effect such feature, structure, or characteristic in connection with other implementations whether or not explicitly described.
[0033] Rendering videos for display on a user equipment, such as the user equipment 100, or on a head-mounted display (HMD) connected to the user equipment via a wireless connection or wirelessly, is computationally and power intensive. The computational and power requirements for rendering videos may be especially high for immersive applications which may collectively be known as extended reality (XR) applications. Examples of immersive applications (e.g., XR applications) include virtual reality (VR) applications, augmented reality (AR) applications, mixed reality (MR) applications. An immersive application (e.g., XR application) may require high quality graphics to ensure a reasonable quality of experience (QoE) to a user using the immersive application (e.g., XR application).
[0034] A user equipment and / or a display (such as a head-mounted display) connected to a user equipment may not have sufficient computation and / or energy resources to provide a reasonable QoE to user using an immersive application (e.g., XR application).
[0035] Before explaining in detail the examples, certain general principles of split rendering are described. Split rendering involves offload parts of processing related to rendering from a device to a remote server. Split rendering has multiple benefits, for instance, lowering the energy consumption of the device or lightening the form factor of the device. In split rendering scenarios, a device is continuously sending its pose information to the remote server and the remote server renders a scene for the requested pose and sends the rendered views back to the device after being encoded and packaged appropriately.
[0036] Conceptually, a split rendering system comprises a Split Rendering Server (SRS) and a Split Rendering Client (SRC). In split rendering a device (e.g., a SRC of the device is offloading rendering tasks to a remote server (e.g., an SRS). The device (e.g., the SRC) is continuously sending pose information (e.g., its position and orientation information) to the remote server (e.g., the SRS) and the remote server (e.g., SRS) renders a computer graphics scene (hereinafter referred to as scene) for a requested pose and sends a rendered frame of the scene back to the device (e.g., the SRC). The rendered frame of the scene may comprise depth and texture information to support better pose correction at the device (e.g., a head mounted display device or a user equipment connected to a head mounted display device).
[0037] The SRC may be deployed on a device (e.g., a head mounted display device or a user equipment connected to a head mounted display device) with low compute or power resources while SRS may be a computing device or computing system that has a large amount of compute and energy resources enabling the SRS to render computationally intensive scenes and encode rendered frames of the scene.
[0038] As noted above, a scene may be rendered by a virtual camera and the pose (e.g., position and orientation) of the virtual camera in the scene may be controlled by a user of the device. For non-immersive applications, a scene may be controlled by a user of the device via an input device of the device, such a mouse, a joystick, a game controller, and the like. For XR applications, a pose of the virtual camera may be controlled by a head mounted display (HMD) device worn by a user.
[0039] In split rendering scenarios, a SRC sends metadata which may include data indicative of a pose of the virtual camera and user input of a user of the device that includes the split rendering client to an SRS over a communication network and receives rendered frames for a scene over the communication network. This results in a latency between a moment the SRC sends the pose to SRS and the moment the SRC receives the corresponding frame of a scene from the SRS. The pose sent by the SRC to the SRC may be a predicted pose for an estimated display time for the corresponding frame of a scene. However, accuracy of a prediction of a pose for a future time point reduces as a length of the prediction window increases. A frame of a scene that is received by the SRC may be rendered for a pose (generally referred to as a render pose) which may be different from an actual pose at the time the frame of the computer graphics scene is displayed (generally referred to as a display pose). This mismatch may be due a latency between the time a pose is measured by the SRC until a time a corresponding frame of the computer graphics scene is displayed by the SRC. This is referred to as motion to render photon latency. Motion to render photon latency may be ameliorated by using pose correction techniques, such as space warp, asynchronous time warp, or other image-based rendering techniques. Pose correction techniques, however, may not be sufficient or may not be successful in ameliorating motion to render photon latency due to issues such as disocclusion and flicker. Disocclusion and flicker may be reduced by streaming depth maps for each frame of a computer graphics scene and using depth information when using pose correction techniques to ameliorate motion to render photon latency. However, streaming depth maps for each frame of a computer graphics scene and using depth information does not eliminate the all the issues related to pose correction. Objects in a scene may be at different depths from the virtual camera and may be occluded from the view of the virtual camera at a first pose and visible at a second pose. In an instance a first pose is a render pose, and a second pose is display pose, it may be impossible for a pose correction technique to provide a coherent frame of the computer graphics scene for the viewer. Further, different objects, or render frames thereof may need different amounts of pose correction.
[0040] Rendering, at an SRS, multiple composition layers of a scene and transmitting the rendered multiple composition layers to the SRC may be advantageous. Such an approach may ameliorate errors that are a result of pose correction techniques.
[0041] A communication network and a device that includes a split rendering client and receives XR services provided by a server hosting an XR application of a XR service provider via the communication network will now be explained with reference to FIG. 1 to FIG. 2 to assist in understanding the described examples.
[0042] Referring to FIG. 1 and FIG. 2 a schematic representation of a communication network 104 (hereinafter referred to as a network) that connects a device 100 (which may an end-user device or a user equipment) to a data network 102. The device 100 includes an immersive application (e.g., an XR application) running on the device 100, a split rendering client (SRC), a XR runtime and a media session handler (MSH). The data network 102 (which may be an edge data network) includes a Media Application Provider (AP), a Split Rendering Server (SRS), and a Media Application Function (RTC AF). The MSH is configured to communicate with the Media Application Function (RTC AF) via an RTC-5 interface. The Media Application Function (RTC-AF) is responsible for provisioning network resources, QoS allocation, and edge resource discovery. The SRC is responsible for acquiring the UE media capabilities (e.g., media capabilities of the device 100) and negotiating with the Media Application Function (RTC AF) to agree on the split rendering process at the Media Application Function (RTC AF). The SRS is also responsible for negotiation of SR session with the SRC, monitoring the SRS's resource usage, and managing and / or running a split rendering process.
[0043] The Media Application Provider (AP) of the data network 102 is configured to provide a XR service to the device 100 and to provision resources for split rendering via an RTC-1 interface. The Media Application Provider (AP) is also configured to deliver media to the SRS via an RTC-2 interface. Communications between the Media Application Function (RTC AF) and the SRC are via an RTC-3 interface. The RTC-3 interface may for example comprise an EDGE-3 interface (as described in clause 6.5.7 of 3GPP TS 23.558 V19.0). Signalling (where signalling refers to sending control plane messages, for example for WebRTC session setup or reconfiguration) and media delivery between the SRC and the SRS is through an RTC-4 interface. The Media Application Function (RTC AF) may provide split rendering information to the MSH via an RTC-5 interface. The SRC may discover the application via the RTC-6 interface. The SRC handles the XR runtime which creates XR experiences. The SRC discovers client media capabilities via an RTC-7 interface. The Application running on the device 100 and the Media Application Provider (AP) of the data network 102 interact via an RTC-8 interface.
[0044] The SRC delivers an immersive experience to a user using the device 100 or a device associated with the device 100, and the SRS supports delivery of the immersive experience by performing at least some of the rendering operations of the immersive experience offloaded by the SRC to the SRS.
[0045] Referring again to FIG. 1, the communication network 104 comprises an access network and a core network 106. The access network 106 provides wireless connectivity (e.g., radio connectivity) to the device 100 and connects the UE 100 to the core network 104 via a backhaul network. The access network 106 may comprise a Radio Access Network (RAN), a non-terrestrial network (e.g., a satellite network), a wireless local area network (WLAN) or any other type of network that provides wireless connectivity to the UE 100 and connects the UE 100 to the core network 108. The access network 106 provides wireless connectivity (e.g., radio connectivity) to the UE 100 via at least one access node of the access network 106. For example, an access node of the access network 106 may be a radio access node and may provide radio connectivity to the UE 102. A radio access node may comprise a gNodeB (gNB), an ng-eNodeB (ng-eNB)), an eNodeB (eNB). The access node of the access network 106 may be an access point, such as a Wireless Local Area Network (WLAN) access point and may provide wireless connectivity to the UE 100. In some implementations, the access node of the access network 106 may be an access point that provides wired connectivity to the UE 100 (generally referred to as fixed access). The backhaul network between the access network 106 and the core network 108 may comprises a fixed network, a satellite network, or a combination thereof. The core network 108 connects the access network 106 with a data network (DN) 102 via a N6 interface (not shown) of the core network 108. As noted above, the core network comprises Network Functions (NF) 110 (generally referred to a network function 110 and collectively referred to as network functions 110). Each respective network function 110 may be implemented as a software running on dedicated hardware (e.g., one or more physical computing devices, such as servers), as a virtualized network function (VNF) instantiated on a physical or virtual machine, or as a container provided by infrastructure of a cloud computing system. Data network 102 may be an external public network, a private data network, or an intra-operator data network (e.g., for providing IP Multimedia Subsystem (IMS) services). The UE 100 (otherwise referred to herein as a mobile terminal) may be configured to access the access network 106, register with the core network 108, establish one or more data sessions with the core network 108, and access services provided by the core network 108 and / or application functions (not shown) hosted on application servers (not shown) of the data network 102.
[0046] Some split rendering applications may benefit from rendering different composition layers of a scene to improve pose correction. Composition layers include various types of content, and each composition layer contributes to the overall appearance and functionality of a scene. Composition layers include a background layer, a geometry layer, a lighting layer, a texture layer, an effects layer, a user interface layer, an animation layer and a sound layer. A background layer is the base layer that sets the scene’s environment, such as sky, landscape, or any static background. The geometry layer contains the 3D models and objects within the scene. The 3D models and object include characters, buildings, vehicles, and other physical entities. The lighting layer defines the light sources and their properties, such as direction, intensity, color, and shadows and is crucial for creating realistic and dynamic lighting effects. The texture layer applies textures to the 3D models, adding details like colors, patterns, and surface properties (e.g., roughness, glossiness). The effect layer includes special effects like particle systems (smoke, fire, rain), post-processing effects (bloom, depth of field), and other visual enhancements. The user interface layer contains user interface elements, such as menus, buttons, and HUD (heads-up display) components, which are overlaid on the scene. The animation layer manages the animations of objects and characters, including their movements, transformations, and interactions. The sound layer, although not visual, includes the audio elements that are synchronized with the scene, such as background music, sound effects, and dialogue.
[0047] The composition layers may be managed separately to allow for flexibility and control during rendering. For example, in instances a distance-based level of detail rendering technique is used to render composition layers there may be different complexities at different depths from a virtual camera corresponding to a pose of a user of a end-user device. Similarly, in instances a pose and / or a gaze of a user is used during foveated rendering to increase quality where a viewer’s (e.g., user’s) eye is looking at, while lowering a quality of peripheral vision of the viewer without impacting a QoE experienced by the viewer. Those different areas may have different encoding complexity and may be sent through multiple composition layers as well. Different composition layers of a scene may be rendered and sent as different frames to alleviate pose correction. There may be other use cases where the SRC receiving multiple concurrent composition layers for a single display frame may be advantageous.
[0048] When multiple composition layers are used in split rendering scenarios, the multiple composition layers may be separately encoded at the SRS and streamed back to the end-user device (e.g., the SRC). The end-user device (e.g., SRC) decodes each separate encoded composition layer of the multiple composition layers and assembles the decoded composition layers to achieve the final rendering. Split rendering scenarios requirements, in terms of latency, are generally low, which means any process that delays the decoding in the end-user device (e.g., SRC) can have a negative impact on the quality of experience of the viewer.
[0049] If the multiple composition layers are encoded utilizing the same codec, multiple decoding instances will be instantiated utilizing the same codec (e.g., the same hardware chipset). Encoding of the multiple composition layers utilizing the same codec (e.g., the same hardware chipset) enables simultaneous decoding encoded multiple composition layers using the same codec chipset, however, introduces further latency because of the decoder instance pipelining. In such cases, however, they would require the SRC to use the same decoding chipset, potentially increasing the latency.
[0050] This disclosure alleviates this issue by utilizing multiple codecs to efficiently handle all composition layers when multiple composition layers are used in split rendering scenarios. Further, rendering a scene at the SRS in multiple composition layers and transmitting the multiple composition layers to the SRC may ameliorate errors that are a result of pose correction techniques.
[0051] This disclosure provides a method for split-rending that involves a split rendering system that includes a split rendering server, and a split rendering client. The method leverages a device’s (e.g., a user equipment or a head mounted display associated with the device) hardware decoding capabilities of the device (e.g., video / image codes supported by the device) to select a set of encoding profiles thereby enabling the lowest possible decoding latency in the split rendering client. The split rendering server receives information on the hardware decoding capabilities of a device (e.g., video-image codecs supported by the device) from a split rendering client of the device. The hardware decoding capabilities of the device include information indicative of decoding hardware (e.g., video / image codecs) that are supported by the device. The hardware decoding capabilities of the device may also include, for each respective video / image codec supported by the device, a decoding latency of the video / image codec (e.g., a time that video / image codec requires to decode an encoded composition layer). The split rendering server, based on the hardware decoding capabilities of a device (e.g., video / image codecs supported by the device) determines an optimal set of composition layers of a computer graphics scene and coding parameters for each composition layer in the set of optimal composition layers to minimize a decoding latency in the device Generally, the largest gains in compression are achieved on complex contents. The complexity of each composition layer can drive the selection of decoding hardware (e.g., a video / image codec) to be used for decoding rendered composition layers. For example, a H.264 codec can be used for encoding and decoding a rendered composition layer that includes static background, a H.265 codec can be used for encoding and decoding a rendered composition layer that includes animated textures with particles effects, and a JPEG-XS codec can be used for encoding and decoding a rendered a graphical user interface (GUI) layer. The split rendering server generates a split rendering configuration and sends the split rendering configuration to the split client. The split rendering configuration includes information that indicates which video / image codecs of the multiple video / image codecs supported by the split rendering client are be used by the split rendering client for decoding which ones of the rendered composition layers. Frames of a scene may be sent together with different priorities and in order to cope with decoding latencies of the video / image codecs of device.
[0052] A rendering loop includes the following. The split rendering server receives a pose prediction from the split rendering client. The split rendering server then renders multiple composition layers for the received pose. The split rendering server encodes the multiple composition layers with the coding parameters. The split rendering server sends the multiple encoded composition layers frames to the split rendering client.
[0053] A split rendering client of a device initiates a split rendering session and transmits the device’s) hardware decoding capabilities to the split rendering server. The device’s hardware decoding capabilities include information indicative of the video decoders that are supported by the device. Optionally, the split rendering client also transmits to the split rendering service the decoding latencies of hardware decoders supported by the device. The split rendering client receives a configuration from the split rendering server. The configuration includes decoding capabilities to be instantiated. The split rendering client instantiates all its required internal hardware and initialized everything properly. A rendering loop at the split rendering client then starts. The rendering loop includes the split rendering client sending a pose-prediction to the split rendering server. After sending a pose-prediction to the split rendering server, the split rendering client receives multiple encoded frames from the split rendering server. The split rendering client forwards the different frame to their respective decoders. The split rendering client forwards the decoded frames to the XR runtime for finale composition and rendering.
[0054] Referring now to FIG. 3, a diagram showing a procedure for multi-codec optimization is shown. The procedure for multi-codec optimization involves a SRC and an XR runtime of a device and an SRS. The procedure for multi-codec optimization begins at operation 1.
[0055] At operation 1, the SRC and the SRS negotiate a split rendering configuration for a split rendering session and setup the split rendering session. The negotiation of a split rendering configuration includes the SRS and the SRC exchanging capabilities including a rendering power of the SRS and the SRC respectively (e.g., a power available for split rendering at the SRC and the SRC) and video / image decoders supported by the SRC, and video / image encoders supported by the SRS. The split rendering configuration includes information indicative which video / image codec is to be used for encoding which type of rendered composition layer and to be used for decoding encoded frame the rendered composition layer. For example, the split rendering configuration may indicate that a H.264 decoder is to be used for
[0056] After the split rendering configuration is negotiated and the split rendering session is setup, a rendering loop starts. The rendering loop begins at operation 3.
[0057] At operation 3, the SRC sends metadata to the SRS. The metadata includes data indicative of a pose data (e.g., position and orientation data) of a user of the device (e.g., an end-user device or UE), timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, and data indicative of a gaze of the user of the device. In some implementations, the metadata includes tracking information that includes one or more of data from sensors of the device that track the movements of the user of the device, viewport information indicative of details above a portion of a scene that needs to be rendered, and / or information about objects in a scene, the properties of the objects and how the objects should be rendered (generally referred to as a scene description).
[0058] At operation 4, the SRS renders multiple composition layers of a scene for a predicted pose of the user of the device (e.g., a pose of the user predicted for a future time).
[0059] At operation 5, the SRS encodes the rendered composition layers as frames. The SRS encodes each respective rendered composition layer as a frame by feeding each respective rendered composition layer to one video / image codec that is indicated in the split rendering configuration to be used for encoding the respective rendered composition layer.
[0060] At 6, the SRS sends the encoded frames to the SRC. For example, a rendered background layer of the multiple composition layers may be encoded using H.264 codec and a rendered graphical user interface layer of the multiple composition layers may be encoded using a JPEG-XS encoder. In some implementations, the SRC may allocate a different level of importance to each encoded frame sent to the SRC. The order in which the encoded frames are sent by the SRC to the SRC may be adjusted by the SRS to account for decoding latencies of the different video / image codecs and transmission latencies. In some implementations, an encoded frame that is estimated to be slowest to decode at the SRC is the first encoded frame sent to the SRC. This enables the prioritization of decoding in the SRC based on the known decoding latencies of the video / image codecs at the SRC. In some implementations, a protocol packager which packages the encoded frames into protocol packets may insert a marker in the protocol packet indicating the relative importance of the encoded frame in the protocol packet. In some implementations, the SRS sends each respective encoded frame together with metadata associated with the respective encoded frame.
[0061] At 7, the SRC decodes the encoded frames by providing (e.g., feeding) each encoded frame into a video / image codec that corresponds to the video / image that was used by the SRS to encode the frame.
[0062] At 8, the SRC forwards the rendered composition layers to the XR runtime for composition and rendering.
[0063] At 9, the XR runtime receives the rendered composition layers, and composes a scene based on the rendered composition layers and renders the scene on a display of the device. In some implementations, the XR runtime composes the scene by superimposing a high quality rendered composition layer on a low quality rendered composition layer. In some implementations, prior to superimposing the high quality rendered composition layer on the low quality rendered composition layer, the high quality rendered composition layer, and the low quality rendered composition layer may be upscaled. In some implementation, prior to superimposing the high quality rendered composition layer on the low quality rendered composition layer, the high quality rendered composition layer, and the low quality rendered composition layer may be downscaled.
[0064] In some implementations, the SRC may send a signalling message that includes the encoded frames and a relative drawing order (e.g., a, depth or z-order) of the rendered composition layers may be sent from the SRS to the SRC. The signalling message may be sent over a real-time data channel, such as a WebRTC data channel that conforms with a split rendering application specific swap message format defined in 3GPP TS 26.565, or an IMS data channel. The signalling message may have the following format: CompositionParams Array 0..1 Array containing information about composition layers compLayerlD number / string l..n A unique identifier of the layer. z-order / depth number 1..1 z-order / depth of the layer. nearPlane / minDepth number l..n near plane of the virtual camera rendering the layer / or the lowest possible (absolute) depth value in the layer, with camera as the origin farPlane / maxDepth number l..n near plane of the virtual camera rendering the layer / or the highest possible (absolute) depth value in the layer, with camera as the origin
[0065] In some implementations, the signaling message sent from the SRS to the SRC may include: (i) an identifier of a rendered composition layer; (ii) an indication of the z-order of the rendered composition layer; (iii) an indication of the near plane of the virtual camera rendering the composition layer; (iv) an indication of the far plane of the virtual camera rendering the composition layer; (v) an indication of the total number of rendered composition layers; and (vi) an indication of post-processing to used on the rendered composition layer during composition of a scene.
[0066] In some implementations, after operation 2, the SRC sends a request for rendering to the SRS and the SRS sends an acknowledgement message to the SRC. The acknowledgement message indicates that the SRS accepted the request for rendering sent by the SCE.
[0067] In some implementations, during negotiation of the split rendering configuration, the SRC sends a signalling message to the SRC to request a split rendering configuration.
[0068] In some implementation, any of the signaling messages described herein may be an application specific RTCP message compliant with RFC4585. In some implementations, any of the signaling messages described herein may be an RTP header extension (RTP HE). In some implementations, the RTP HE may be a header extension to an RTP packet comprising an encoded frame that includes a rendered composition layer. In some implementations, a signalling message may be sent from the SRS to the SRC as an RTCP APP packet compliant with RFC 3550. In some implementations, a signalling message may be sent from the SRS to the SRC as an RTCP XR packet compliant with RFC 3611.
[0069] In some implementations, during setup of the split rendering session, a signalling messages indicating the rendered composition layers and the parameters are exchanged between the SRC and the SRS as parameters in SDP exchange.
[0070] In some implementations, the signalling message comprises a dedicated SEI message that can be carried alongside with every frame comprising a rendered composition layer. The dedicated SEI message may be as follows: composition_parameters( payloadSize) { Descriptor cp_composition_layer_id u(8) cp_number_layer u(8) for(i=0;i<cp_number_layer;i++){ cp_min_depth_range u(8) cp_max_depth_range u(8) cp_filtering_idx u(8) }
[0071] The SEI message above includes the following semantics: • cpcompositionlayerid indicates a position of the received rendered composition layer in the hierarchical rendering structure. • cp number layer indicates the number of layers in the hierarchical rendering structure. • cpmindepthrange indicates a lower boundary depth value of the received composition layer. • cpmaxdepthrange indicates the upper boundary depth value of the received composition layer • cpfilteringidx indicates which filter should be used at the boundaries of this composition layer. 0 means no filtering, 1 means bilinear filtering, 2 means bicubic filtering. • Note: other filtering, or post processing could be signalled here, including super-resolution, denoising, custom filters, and the like.
[0072] An example of a dedicated SEI message for when there are three (3) composition layers is as follows: • GUI Layer: a. cpcompositionlayerid = 0 b. cp min depth range = 0 c. cp max depth range = 1 d. cp filtering idx = 0 • Foreground layer: e. cp composition layer id = 1 f. cpmindepthrange = 1 g. cp max depth range = 5 h. cp filtering idx = 1 • Background layer: i. cpcompositionlayerid = 2 j. cpmindepthrange = 5 k. cp max depth range = 30 1. cpfilteringidx = 1
[0073] In some implementations, the signaling of the rendered composition layers is achieved with a lookup map that is being encoded as a separated video stream. This has the advantage of not requiring a dedicated video / image codec in the device, the lookup map being decoded with an existing video / image decoder of the device and fed to the XR runtime (e.g. the shader sampler function). The lookup map represents for each pixel location from which composition layer the pixel should be extracted. As a limited number of composition layers is 17 needed, this lookup can be efficiently encoded with screen-content coding tools assuming the range of pixel used in the input signal is known, and thus can be easily encoded, e.g. by leveraging palette encoding. In addition to encoding the layer identifier, some post-processing information can also be encoded. Hence, multiple technical realizations of this lookup map encoding are possible, for example: • Monochrome encoding for layer identifiers. • Monochrome encoding for layer identifiers and filter index. Layer identifier can be coded in 4-bits MSB and filter index in 4-bits LSB. • 4:4:4 coding with the Y channel encoding the lookup map, the U channel for the post-processing mode (e.g. filtering, up-sampling, denoising, etc ...) and the V channel to carry filter indexes or parameters.
[0074] Referring now to FIG. 4, a flowchart of an example method or process of an SRC is shown. As noted above, a device, such as device 100, comprises an SRC. The device also comprises at least one processor and at least one memory that stores instructions of the SRC which, when executed by the at least one processor, cause the device to perform the method or process shown in FIG. 4. The method or process begins at 402.
[0075] At 402, the SRC sends to a SRS, information indicative of video / image codecs supported by the split rendering client. In some implementations, the SRC also sends to the SRC, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
[0076] At 404, the SRC receives from the SRS, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames.
[0077] At 406, the SRC sends to the SRS, metadata comprising data indicative of a predicted pose of a user of the device. In some implementations, the metadata also includes at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device. At 408, the SRC receives from the SR, multiple encoded frames, wherein each respective encoded frame of the multiple encoded frames comprises a rendered composition layer of multiple rendered composition layers of a scene. In some implementations, the SRC also receives from the SRS metadata associated with each respective encoded frame. In some implementations, the SRC receives from the SRS signaling comprising a relative drawing order of the multiple rendered composition layers. In some implementations, the SRC receives from the SRS signaling comprising a relative depth order of the multiple rendered composition layers.
[0078] At 410, the SRC, for each respective encoded frame, decoding the respective encoded frame to obtain the respective rendered composition layer using one video / image codec of the one or more video / image codec, the one video / image codec corresponding to a video / image codec of the SRS that is used to encode the respective rendered composition layer.
[0079] At 412, the SRC forwards each respective rendered composition layer to an extended reality (XR) runtime of the user equipment for composition of the scene and rendering of the scene. The XR runtime may receive each respective rendered composition layer, compose the scene using each respective rendered composition layer; and render the scene on a display of the device.
[0080] Referring now to FIG. 5, a flowchart of an example method or process of an SRS is shown. The SRS comprises at least one processor and at least one memory that stores instructions, which when executed by the at least one processor, cause the SRS to perform the method or process shown in FIG. 5. The method or process begins at 502.
[0081] At 502, the SRS receives from the SRC, information indicative of video / image codecs supported by the split rendering client. In some implementations, the SRS also receives from the SRS, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
[0082] At 504, the SRS sends to the SRC a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames. In some implementations, the SRS determines the split rendering configuration comprising the indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames based on the information indicative of video / image codecs supported by the split rendering client and the information indicative of a decoding latency of each respective video / image decoder supported by the split rendering client.
[0083] At 506, the SRS receives from the SCR, metadata comprising at least data indicative of a predicted pose of a user of device. In some implementations, the metadata also includes at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device.
[0084] At 506, the SRS renders multiple composition layers of a scene for the predicted pose.
[0085] At 506, the SRS, for each respective rendered composition layer, encoding, using one video / image codec of the SRS, the respective rendered composition layer to generate an encoded frame comprising the respective the one video / image codec corresponding to one of the video / image codecs supported by the SRC. At 508, the SRS sends each respective encoded frame to the SRC. In some implementations, the SRC also sends metadata associated with each respective encoded frame. In some implementations, the SRS sends to the SRC signaling comprising a relative drawing order of the multiple rendered composition layers. In some implementations, the SRS sends to the SRC signaling comprising a relative depth order of the multiple rendered composition layers.
[0086] Reference is now made to FIG. 6 which shows physical and logical components of an example device 600 that includes a split rendering client in accordance with an implementation of this disclosure. Although FIG. 4 shows a single instance of each physical and / or logical component of the device 100, the user device 100 may include multiple instances of each physical and / or logical component shown in FIG. 6.
[0087] Device 100 may be a user equipment that is capable of sending and receiving radio signals or an HMD that communicates (e.g., via a wireless or wired connection) with a user equipment. Non-limiting examples of a user equipment include a mobile station (MS), a mobile device such as a mobile phone or what is known as a ’smart phone’, a computer provided with a wireless interface card or other wireless interface facility (e.g., USB dongle), a personal data assistant (PDA) or a tablet provided with wireless communication capabilities, a machine-type communications (MTC) device, an Internet of things (loT) type communication device or any combinations of these.
[0088] The device 100 also includes one or more processors 601, one or more memories 602 (collectively referred to as memory 602) and other components or circuitry 603 for use in software and hardware aided execution of operations of the split rendering client, the XR runtime and the media session handler described herein. In the implementation that the device 100 is a user equipment, the device is configured to perform, including control of access to and communications with radio access networks (e.g., the RAN illustrated in FIG. 1). The processor 601 is coupled to the memory 602. The one or more processors 601 may comprise a central processing unit (CPU), a microprocessor, multicore processor, a tensor processing unit (TPU), a graphics processing unit (GPU), a neural processing unit (NPU), a dedicated logic circuit, an application specific integrated circuit, a field programmable gate array (FPGA), a dedicated artificial intelligence processing unit, a hardware accelerator, a quantum processor, or any combinations thereof. The memory 602 may include a volatile or non-volatile memory (e.g., a flash memory', a random-access memory (RAM), and / or a read-only memory (ROM)).
[0089] The processor 601 may be configured to execute instructions 608 of the split rendering client described herein. The processor 601 may also be configured to execute instructions of the XR runtime and / or the media session handler described herein. The execution of the instructions 608 may for example cause the device 100 to perform one or more operations, including the operations of the split rendering client described herein with respect to FIG 4. The instructions of the split rendering client 608 may be stored in memory 602. Similarly, the instructions of the XR runtime and media session hander may be stored in memory 602.
[0090] In the implementation that the device 100 is a user equipment, the device also include an antenna array 604 and a transceiver 606 for transmitting wireless signals (e.g., radio signals) to access nodes of an access network (e.g., a radio access network nodes of a RAN) and / or receiving wireless signals from access nodes of an access network (e.g., a radio access network nodes of a RAN) over an air interface 607. The wireless signals (e.g., radio signals) may carry communications, such as voice, electronic mail (email), text messages, multimedia data, and / or machine data. The antenna array 606 may be arranged internally or externally to the user equipment 100. The antenna array 606 may comprise one or more antenna elements. The antenna array 606 may be a multi-input multi output (MIMO) antenna.
[0091] The processor 601, the at least one memory' 602, the transceiver 606 and other components or circuitry 603 of the user equipment 100 (e.g., a modem) can be provided on a circuit board, in chipsets, or in a system on chip (SOC). The circuit board, chipsets or SOC is denoted by reference 404 in Figure 3. Device 100 may optionally include display device 605, for example, a touch-sensitive display device. The device 100 also includes a battery (not shown). The device 100 may also include a speaker (not shown) and a microphone (not shown). The device 100 may also include a universal subscriber identity module (USIM) (not shown) or an embedded subscriber identity module (eSIM) (not shown).
[0092] Reference is now made to FIG. 7 which shows physical and logical components of an example of split rendering server 500 in accordance with an example implementation of this disclosure, although FIG. 7 shows a single instance of each logical and / or physical component of the split rendering server 700, there may be multiple instances of each logical and / or physical component shown in FIG. 7.
[0093] The split rendering server 700 includes one or more processors 702, such as a central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a dedicated logic circuitry, a graphics processing unit (GPU), a tensor processing unit, a neural processing unit, a dedicated artificial intelligence processing unit, a hardware accelerator, a quantum processor, or any combinations thereof. The one or more processors 702 may generally be referred to as processor 702 and collectively be referred to as processors 702.
[0094] The split rendering server 700 also includes one or more memories 704 (generally referred to a memory 504 and collectively referred to herein as “memories 704”), which may include volatile or non-volatile memory (e.g., a flash memory, a random-access memory (RAM), and / or a read-only memory (ROM)). Memory 704 may store computer code (e.g. instructions) for execution by at least one of the one or more processors 702. For example, the computer code (e.g. instructions) which is shown stored in the memory 704, may be executed by at least one of the one or more processors 702, to cause the split rendering server 700 to perform the operations of the method of FIG. x described herein.
[0095] In some examples, the split rendering server 700 may also include one or more electronic storage units (not shown), such as a solid-state drive, a hard disk drive, a magnetic disk drive and / or an optical disk drive. In some examples, one or more datasets and / or modules may be provided by an external memory or may be provided by a transitory or non-transitory computer-readable medium. Examples of non-transitory computer readable media include a RAM, a ROM, an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a CD-ROM, or other portable memory storage. The storage units and / or external memory may be used in conjunction with memory 704 to implement data storage, retrieval, and caching functions of the split rendering server 700.
[0096] The processor 702 and memory 704 of the split rendering server 700 may communicate with each other via a communication bus, for example.
[0097] It is noted that whilst some embodiments have been described in relation to 5G networks, similar principles can be applied in relation to other networks and communication systems such as 6G networks or 5G-Advanced networks. Therefore, although certain implementations were described above by way of example with reference to certain example communication networks, technologies and standards, implementations may be applied to any other suitable forms of communication systems than those illustrated and described herein.
[0098] It is also noted herein that while the above describes example implementations, there are several variations and modifications which may be made to this disclosure without departing from the scope of this disclosure.
[0099] In general, the various implementations may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. Some aspects of this disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of this disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[00100] As used in this disclosure, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (iii) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[00101] This definition of circuitry may apply to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[00102] The term “means for” may define at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause an apparatus at least to perform the steps listed after the term.
[00103] The implementations described in this disclosure may include computer software comprising instructions which when executed by a processor of a user equipment or a split rendering server, such as in the processor entity, or by hardware, or by a combination of software and hardware cause the communication device to carry out or perform the operations of the method shown in FIG. 8. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computer-executable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it.
[00104] Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non-transitory media. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
[00105] The scope of protection sought the disclosure is set out by the independent claims and the dependent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure.
[00106] The foregoing description has provided by way of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and 24 the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.
Claims
1. A device comprising:at least one processor; andat least one memory storing instructions of a split rendering client which, when executed by the at least one processor, cause the device to:sending, to a split rendering server, information indicative of video / image codecs supported by the split rendering client;receiving, from the split rendering server, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames;sending, to the split rendering server, metadata comprising at least data indicative of a predicted pose of a user of the device;receiving, from the split rendering server, multiple encoded frames, wherein each respective encoded frame of the multiple encoded frames comprises a rendered composition layer of multiple rendered composition layers of a scene;for each respective encoded frame, decoding the respective encoded frame to obtain the respective rendered composition layer using one video / image codec of the one or more video / image codec, the one video / image codec corresponding to a video / image codec of the SRS that is used to encode the respective rendered composition layer; andforwarding each respective rendered composition layer to an extended reality (XR) runtime of the user equipment for composition of the scene and rendering of the scene.
2. The device of claim 1, wherein the memory further stores instructions of an extended reality runtime which, when executed by the at least one processor, cause the device to perform:receiving each respective rendered composition layer;composing the scene using each respective rendered composition layer; and rendering the scene on a display of the device.
3. The device of claim 2, wherein the composing comprises superimposing a first rendered composition layers on a second respective composition layer.
4. The device of any of claims 1 to 3, wherein the metadata further comprises at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device5. The device of any of claims 1 to 4, wherein the instructions of the split rendering client when executed by the at least one processor further cause the device to perform:receiving, from the split rendering server, signaling comprising a relative drawing order of the multiple rendered composition layers.
6. The device of any of claims 1 to 5, wherein the instructions of the split rendering client when executed by the at least one processor further cause the device to perform:receiving, from the split rendering server, signaling comprising a relative depth order of the multiple rendered composition layers.
7. The device of any of claims 1 to 6, wherein the instructions of the split rendering client when executed by the at least one processor further cause the device to perform:sending, to the split rendering server, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
8. The device of any of claims 1 to 7, wherein the device is a user equipment or a headmounted display.
9. A split rendering server comprising:at least one processor; andat least one memory storing instructions which, when executed by the at least one processor, cause the split rendering server to perform:receiving, from a split rendering client, information indicative of video / image codecs supported by the split rendering client;sending, to the split rendering client, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames;receiving, from the split rendering client, metadata comprising at least data indicative of a predicted pose of a user of device;rendering multiple composition layers of a scene for the predicted pose;for each respective rendered composition layer, encoding, using one video / image codec of the split rendering server, the respective rendered composition layer to generate an encoded frame comprising the respective the one video / image codec corresponding to one of the video / image codecs supported by the split rendering client; andsending each respective encoded frame to the split rendering client.
10. The split rendering service of claim 9, wherein the metadata further comprises at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device11. The split rendering server of claim 9 or 10, wherein the instructions when executed by the at least one processor further cause the split rendering client to perform:sending to the split rendering client, signaling comprising a relative drawing order of the multiple rendered composition layers.
12. The split rendering server of any of claims 9 to 11, wherein the instructions when executed by the at least one processor further cause the split rendering client to perform: sending to the split rendering client, signaling comprising a relative depth order of the multiple rendered composition layers.
13. The split rendering server of any of claims 9 to 12, wherein the instructions when executed by the at least one processor further cause the split rendering client to perform: receiving, from the split rendering client, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
14. The split rendering server of any of claims 9 to 13, wherein the instructions when executed by the at least one processor further cause the split rendering client to perform: determining the split rendering configuration comprising the indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames based on the information indicative of video / image codecs supported by the split rendering client and the information indicative of a decoding latency of each respective video / image decoder supported by the split rendering client.
15. A method comprising:sending, by a split rendering client to a split rendering server, information indicative of video / image codecs supported by the split rendering client;receiving, by the split rendering client from a split rendering server, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames;sending, by the split rendering client to the split rendering server, metadata comprising at least data indicative of a predicted pose of a user of the device;receiving, by the split rendering client from the split rendering server, multiple encoded frames, wherein each respective encoded frame of the multiple encoded frames comprises a rendered composition layer of multiple rendered composition layers of a scene;for each respective encoded frame, decoding by the split rendering client, the respective encoded frame to obtain the respective rendered composition layer using one video / image codec of the one or more video / image codec, the one video / image codec corresponding to a video / image codec of the SRS that is used to encode the respective rendered composition layer; andforwarding, by the split rendering client, each respective rendered composition layer to an extended reality (XR) runtime of the user equipment for composition of the scene and rendering of the scene.
16. The method of claim 15, further comprising:receiving, by the XR runtime, each respective rendered composition layer;composing, by the XR runtime, the scene using each respective rendered composition layer; andrendering, by the XR runtime, the scene on a display of the device.
17. The method of claim 16, wherein the composing comprises superimposing a first rendered composition layers on a second respective composition layer.
18. The device of any of claims 15 to 17, wherein the metadata further comprises at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device19. The device of any of claims 15 to 18, further comprisingreceiving, by the split rendering client from the split rendering server, signaling comprising a relative drawing order of the multiple rendered composition layers.
20. The device of any of claims 15 to 19, further comprising:receiving, by the split rendering client from the split rendering server, signaling comprising a relative depth order of the multiple rendered composition layers.
21. The device of any of claims 15 to 20, further comprising:sending, to the split rendering server, for each respective video / image decoder supported by the split rendering client, information indicative of a decoding latency of the respective video / image decoder.
22. A method comprising:receiving, by a split rendering server from a split rendering client, information indicative of video / image codecs supported by the split rendering client;sending, by the split rendering server to the split rendering client, a split rendering configuration comprising an indication of one or more of the video / image codecs to be utilized by the split rendering client for decoding encoded frames;receiving, by the split rendering server from the split rendering client, metadata comprising at least data indicative of a predicted pose of a user of device;rendering by the split rendering server multiple composition layers of a scene for the predicted pose;for each respective rendered composition layer, encoding by the split rendering server, using one video / image codec of the split rendering server, the respective rendered composition layer to generate an encoded frame comprising the respective the one video / image codec corresponding to one of the video / image codecs supported by the split rendering client; andsending by the split rendering server each respective encoded frame to the split rendering client.
23. The method of claim 22, wherein the metadata further comprises at least one of timing information indicative of a first time at which a frame is expected to be displayed on a display of the device, or data indicative of a gaze of the user of the device24. The method of claim 22 or 23, further comprisingsending, by the split rendering server to the split rendering client, signaling comprising a relative drawing order of the multiple rendered composition layers.
25. The split rendering server of any of claims 22 to 24, further comprising:sending, by the split rendering server to the split rendering client, signaling comprising a relative depth order of the multiple rendered composition layers.T +44(0)30 0300 2000A
Citation Information
Patent Citations
ViewWO2025/014882A1onEspacenetopensinnewtab
ViewUS2024/0414415A1onEspacenetopensinnewtab