Apparatus, method and computer program for a pose error aware split-rendering

The proposed method enhances the prediction of the immersive applications.

WO2025257059A1PCT designated stage Publication Date: 2025-12-18NOKIA TECHNOLOGIES OY

Patent Information

Application Number
PCT/EP2025/065845
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-14
Filing Date
2025-06-06
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Communication devices, especially those with limited resources, struggle to provide high-quality immersive experiences due to computational and energy constraints, and existing technologies fail to optimize latency in pose prediction for split-rendering scenarios, which leads to poor performance in prediction-based rendering scenarios.

Method used

The proposed method utilizes a split rendering system that predicts pose prediction error for immersive applications.

Benefits of technology

The proposed method enhances the prediction of the immersive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025065845_18122025_PF_FP_ABST
    Figure EP2025065845_18122025_PF_FP_ABST
Patent Text Reader

Abstract

There is provided an apparatus comprising a split rendering client that is configured for obtaining a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display, providing rendering metadata to a split rendering server. The rendering metadata comprises the predicted pose and the time information. The split rendering client is further configured for receiving, from the split rendering server, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames, obtaining a second predicted pose at a second time different than the first time, decoding at least two rendered and encoded frames of the plurality of the rendered and encoded frames to generate at least two decoded frames, based on the render pose associated with each respective decoded frame and the second predicted pose, selecting a primary decoded frame and one or more secondary decoded frames from the at least two decoded frames, composing a video frame to be displayed based on the primary decoded frame and at least one or more secondary decoded frames and causing the video frame that is composed to be displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Apparatus, method and computer program for a pose error aware split-renderingRelated Applications

[0001] This patent application claims the benefit of priority of United Kingdom Patent Application No. 2408587.0 filed on June 14, 2024, which is hereby incorporated by reference as if reproduced in its entirety.Field

[0002] The present application relates to a method, apparatus, system and computer program and in particular but not exclusively to apparatus, methods, and computer programs for pose aware split-rendering.Background

[0003] A communication system can be seen as a facility that enables communication sessions between two or more communication devices, provides communication devices access to a data network, or provides services such as, extended reality services, to communication devices. A mobile or wireless communication network is one example of a communication network. A communication device may be provided with a service by an application server.

[0004] A mobile or wireless communication network may operate in accordance with standards such as those provided by 3GPP (Third Generation Partnership Project) or ETSI (European Telecommunications Standards Institute). Examples of mobile or wireless communication network that operate in accordance with 3 GPP standards are generally referred to as 4G (4th Generation) networks, 5G (5th Generation) network, 5G-Advanced networks and 6G networks.Summary

[0005] In a first aspect there is provided an apparatus comprising a split rendering client configured for obtaining a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display, providing rendering metadata to a split rendering server. The rendering metadata comprises the predicted pose and the time information. The split rendering client is further configured for receiving, from the split rendering server, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of theplurality of rendered and encoded frames, obtaining a- second predicted pose at a second time different than the first time, decoding at least two rendered and encoded frames of the plurality of the rendered and encoded frames to generate at least two decoded frames, based on the render pose associated with each respective decoded frame and the second predicted pose, selecting a primary decoded frame and one or more secondary decoded frames from the at least two decoded frames, composing a video frame to be displayed based on the primary decoded frame and at least one or more secondary decoded frames and causing the video frame that is composed to be displayed.

[0006] The split rendering client may be further configured for computing a pose prediction error for the predicted pose and providing the pose prediction error to the split rendering server.

[0007] The pose prediction error for the predicted pose may be computed using historical predicted poses and the predicted pose.

[0008] The pose prediction error for the predicted pose may comprise a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

[0009] Each respective rendered and encoded video frame of the plurality of rendered and encoded frames may be associated with at least one of a priority level or a quality of service profile.

[0010] Composing a frame may be further based on depth map determined based on the one or more secondary decoded frames.

[0011] In a second aspect there is provided an apparatus comprising a split rendering server configured for receiving rendering metadata from a split rendering client. The rendering metadata comprises a predicted pose of a virtual camera and a time information indicative of a time at which a video frame is expected to be displayed on a display. The split rendering server is further configured for obtaining a pose prediction error for the predicted pose, rendering and encoding a plurality of frames based on the rendering metadata, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error and providing, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0012] Obtaining the pose prediction error for the predicted pose may comprise receiving the pose prediction error from the split rendering client.

[0013] The pose prediction error for the predicted pose may comprise at least one of the following: a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

[0014] At least one of the at least associated render poses or the number thereof is based on at least one of the following: a magnitude of the pose prediction error, a magnitude of the predicted position error, a magnitude of the predicted orientation error, the confidence value associated with the predicted pose, an available channel capacity, a latency, a quality of service profile, and at least one historical predicted pose of the virtual camera or a time period between the rendering operation and the first time.

[0015] The plurality of frames may be encoded using multi-view encoding.

[0016] Each respective rendered and encoded frame of the plurality of rendered and encoded frames may be associated with at least one of a priority level or a quality of service profile.

[0017] In a third aspect there is provided a method of a split rendering client, the method comprising obtaining a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display, providing rendering metadata to a split rendering server, the rendering metadata comprising the predicted pose and the time information, receiving, from the split rendering server, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames, obtaining a- second predicted pose at a second time different than the first time, decoding at least two rendered and encoded frames of the plurality of the rendered and encoded frames to generate at least two decoded frames, based on the render pose associated with each respective decoded frame and the second predicted pose, selecting a primary decoded frame and one or more secondary decoded frames from the at least two decoded frames, composing a video frame to be displayed based on the primary decoded frame and at least one or more secondary decoded frames and causing the video frame that is composed to be displayed.

[0018] The method may further comprise computing a pose prediction error for the predicted pose and providing the pose prediction error to the split rendering server.

[0019] The pose prediction error for the predicted pose may be computed using historical predicted poses and the predicted pose.

[0020] The pose prediction error for the predicted pose may comprise a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

[0021] Each respective rendered and encoded video frame of the plurality of rendered and encoded frames may be associated with at least one of a priority level or a quality of service profile.

[0022] Composing a frame may be further based on depth map determined based on the one or more secondary decoded frames.

[0023] In a fourth aspect there is provided a method of a split rendering server, the method receiving rendering metadata from a split rendering client, the rendering metadata comprising a predicted pose of a virtual camera and a time information indicative of a time at which a video frame is expected to be displayed on a display, obtaining a pose prediction error for the predicted pose, rendering and encoding a plurality of frames based on the rendering metadata, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error and providing, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0024] Obtaining the pose prediction error for the predicted pose may comprise receiving the pose prediction error from the split rendering client.

[0025] The pose prediction error for the predicted pose may comprise at least one of the following: a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

[0026] At least one of the at least associated render poses or the number thereof is based on at least one of the following: a magnitude of the pose prediction error, a magnitude of the predicted position error, a magnitude of the predicted orientation error, the confidence value associated with the predicted pose, an available channel capacity, a latency, a quality of service profile, and at least one historical predicted pose of the virtual camera or a time period between the rendering operation and the first time.

[0027] The plurality of frames may be encoded using multi-view encoding.

[0028] Each respective rendered and encoded frame of the plurality of rendered and encoded frames may be associated with at least one of a priority level or a quality of service profile.

[0029] In a fifth aspect there is provided an apparatus comprising a split rendering client and comprising means for performing the method according to the third aspect.

[0030] The means may comprise at least one processor, and at least one memory storing instructions which, when executed by the at least one processor, cause the apparatus at least to perform the method according to the third aspect.

[0031] In a sixth aspect there is provided an apparatus comprising means for performing the third aspect.

[0032] The means may comprise at least one processor, and at least one memory storing instructions of a split render client, wherein the instructions when executed by the at least one processor, cause the apparatus to perform the method according to the third aspect.

[0033] In a seventh aspect there is provided an apparatus comprising at least one processor, and at least one memory storing instructions of a split render client which, when executed by the at least one processor, cause the apparatus at least to perform a method according to the third aspect.

[0034] In an eighth aspect there is provided an apparatus comprising at least one processor, and at least one memory storing instructions of a split rendering server, wherein the instructions, when executed by the at least one processor cause the apparatus at least to perform a method according to the fourth aspect.

[0035] In a ninth aspect there is provided a non-transitory computer readable medium comprising instructions of a spit rendering client wherein the instructions when executed by at least one processor of an apparatus cause the apparatus to perform the method according to the third aspect.

[0036] In a tenth aspect there is provided a non-transitory computer readable medium comprising program instructions of a spit rendering server wherein the instructions when executed by at least one processor of an apparatus cause the apparatus to perform the method according to the fourth aspect.

[0037] In the above, many different embodiments have been described. It should be appreciated that further embodiments may be provided by the combination of any two or more of the embodiments described above.Description of Figures

[0038] Embodiments will now be described, by way of example only, with reference to the accompanying Figures in which:

[0039] FIG. 1 is a schematic diagram of a communication network that connects a user equipment to a data network that includes one or more servers of an application provider that provides extended reality services to the user equipment;

[0040] FIG. 2 is a schematic diagram of an example communication device;

[0041] FIG. 3 is a schematic diagram of an example split-rendering server;

[0042] FIG. 4 shows a schematic diagram of an example communication network that has a MSE split rendering architecture that connects a user equipment to a data network that includes one or more servers of an application provider that provides extended reality services to the user equipment;

[0043] FIGs. 5A and 5G show a flowchart of a method according to an example embodiment;

[0044] FIG. 6 shows a flowchart of a method according to an example embodiment;

[0045] FIG. 7 shows a diagram of a split rendering procedure that involves a split rendering client and a split rendering server according to an example embodiment;

[0046] FIG. 8 shows a flowchart of a method according to an example embodiment;

[0047] FIG. 9 shows a flowchart of a method according to an example embodiment.Detailed description

[0001] Before explaining in detail the examples, certain general principles of a communication system of a communication network and communication devices that receive XR services provided by XR service providers are briefly explained with reference to FIG. 1, FIG. 2 and FIG. 3 to assist in understanding the technology underlying the described examples.

[0002] FIG. 1 shows a schematic representation of a 5G communication network 100 (hereinafter referred to as a 5G network) that connects a UE 102 (which may also be referred to as a communication device or a terminal) to a DN 110 that includes a Split-Rendering Server (SRS), a Real-Time Communications AF and an Application Provider (see FIG. 4). The communication network 100 may comprise a 5G radio access network (5G-RAN) 104 and a 5G core network (5GC) 106. One or more application functions (AF) 108 may be connected to the 5GC 100 (e.g., a network exposure function of the 5GC) to enable the one or more AFs to access services provided by network functions of the 5GC. Although the AF108 is shown in FIG. 1 as being external to the 5GC 106, in some embodiments, one or more of the AFs 108 may be a trusted AF that is part of 5GC (e.g., internal to the 5GC).

[0003] The 5GC 106 comprises one or more network functions (otherwise referred to a network entities). The one or more network functions of the 5GC 106 may comprise one or more Access and mobility Management Functions (AMF) 112, one or more session management functions (SMF) 114, an authentication server function (AUSF) 116, a Unified Data Management (UDM) 118, one or more user plane functions (UPF) 120, a Unified Data Repository (UDR) 122 and / or a Network Exposure Function (NEF) 124. The UPF is controlled by the SMF (Session Management Function) that receives policies from a PCF (Policy Control Function). Each respective network function of the 5G communicates with another network function of the 5G using a service-based interface.

[0004] The UPF of the 5GC 106 communicates with the Radio Access Network (RAN) via a N3 interface. The 5G-RAN 104 may comprise one or more radio access network nodes (RANs). A radio access network node may comprise a gNodeB (gNB) and a gNB may comprise one or more gNB Distributed Units (DUs) connected to a gNodeB (gNB) Centralized Unit (CU).

[0005] In FIG. 1, PDU Session Anchor (PSA) is the term given to a UPF (User Plane Function) which terminates at an N6 interface of a 5G network. A UPF that terminates at an N6 interface is said to support PDU Session Anchor functionality.

[0006] A communication device will now be described in more detail with reference to FIG. 2 showing a schematic, partially sectioned view of a communication device 200. Such a communication device is often referred to as user equipment (UE) or terminal. A communication device may be any device capable of sending and receiving wireless signals, including radio signals. Non-limiting examples of a communication device comprise a mobile station (MS) or mobile device such as a mobile phone or what is known as a ’smart phone’, a computer provided with a wireless interface card or other wireless interface facility (e.g., USB dongle), personal data assistant (PDA) or a tablet provided with wireless communication capabilities, voice over IP (VoIP) phones, portable computers, desktop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), smart devices, wireless customer-premises equipment (CPE), or any combinations of these or the like. A communication device may provide, for example, communication ofdata for carrying communications such as voice, electronic mail (email), text message, multimedia and so on. Users may thus be offered and provided numerous services via their communication devices. Non-limiting examples of these services comprise two-way or multi-way calls, data communication or multimedia services or simply an access to a data communications network system, such as the Internet. Users may also be provided broadcast or multicast data. Non-limiting examples of the content comprise downloads, television and radio programs, videos, advertisements, various alerts, and other information.

[0007] A communication device 200 is typically provided with or comprises at least one data processing entity 201, at least one memory 202 and other possible components 203 for use in software and hardware aided execution of tasks it is designed to perform, including control of access to and communications with access systems and other communication devices. The data processing, storage and other relevant components can be provided on an appropriate circuit board and / or in chipsets. This feature is denoted by reference 204. The user may control the operation of the mobile device by means of a suitable user interface such as key pad 205, voice commands, touch sensitive screen or pad, combinations thereof or the like. A display 208, a speaker and a microphone can be also provided. Furthermore, a mobile communication device may comprise appropriate connectors (either wired or wireless) to other devices and / or for connecting external accessories, for example hands-free equipment, thereto.

[0008] The communication device 200 may receive wireless signals (e.g., radio signals) over an air or radio interface 207 via appropriate apparatus for receiving and may transmit signals via appropriate apparatus for transmitting radio signals. In FIG. 2 transceiver apparatus is designated schematically by block 206. The transceiver apparatus 206 may be provided for example by means of a radio part and associated antenna arrangement. The antenna arrangement may comprise one or more antenna elements and may be arranged internally or externally to the mobile device.

[0009] FIG. 3 shows an example of split rendering server 300. The split rendering server 300 comprises at least one memory 301, at least one data processing unit 302, 303 and a network interface 304. The at least one data processing unit 302 and 303 may comprise a CPU 302 and a hardware accelerator 303. The memory 301 of the split rendering server stores instructions which when executed by the at least one data processing unit 302, 303, cause the split rendering server to perform operations, for example the operations described in FIG.6 or the operations associated with the steps 1, 2, 3b, 3c, 4, 5, and 6 of the procedure shown in FIG. 7.

[0010] Rendering videos for display on, e.g., a communication device such as the communication device 200 described with reference to FIG. 2 or a head-mounted display (HMD) connected to the communication device via a wireless connection or wirelessly, is computationally and power intensive. The computational and power requirements for rendering videos may be especially high for immersive applications which may collectively be known as extended reality (XR) applications. Examples of immersive applications (e.g., XR applications) include virtual reality (VR) applications, augmented reality (AR) applications, mixed reality (MR) applications. An immersive application (e.g., XR application) may require high quality graphics to ensure a reasonable quality of experience (QoE) to a user using the immersive application (e.g., XR application).

[0011] A communication device and / or a display (such as a head-mounted display) connected to a communication device may not have sufficient computation and / or energy resources to provide a reasonable QoE to user using an immersive application (e.g., XR application).

[0012] Split rendering may allow immersive applications (e.g., XR applications) that are resource intensive to be provided with a high quality of experience by resource constrained devices such as a communication device or a standalone Head Mounted Displays (HMDs) connected to a communication device.

[0013] Referring again to FIG. 1, a 5G network that has a service-based architecture may expose network resources and functionalities via standardized interfaces to entities that are authorized to provide a service (e.g., an XR service). A service of a provider (referred to as a service provider) is provided by one or more servers hosting or running an application. In some implementations, a 5G network hosts one or more servers of the service provider that provide a service (XR service) to communication devices via the 5G network. In some implementations, a 5G network may expose network resources and functionalities via standardized interfaces to a service provider (in which case the one or more servers of the service provider are located outside the 5G network) to enable the service provider to provide a service (e.g., XR service) to communication devices via the 5G network.

[0014] Referring now to FIG. 4, an example of a UE connected to a DN via 5G network (not shown but described with respect to FIG. 1). The 5G network has a Split Rendering MediaService Enabler (SR MSE) architecture. In FIG. 4, Split Rendering is performed by a split render client (SRC) of the UE and a split rendering server (SRS) of the DN. The SRS may reside inside (e.g., may be deployed or located in) the DN which may be an edge data network.

[0015] The UE includes an immersive application (e.g., an XR application) client and a media session handler (MSE) that is configured to communicate with the RTC-AF of the DN via a control plane of the 5G network. The RTC-AF is responsible for provisioning of network resources, QoS allocation, and edge resource discovery.

[0016] The UE also includes a split rendering client (SRC) that is responsible for acquiring the UE media capabilities (e.g., media capabilities of a UE) and negotiating with a RTC AF to agree on the split-rendering process at the RTC AF. The SRS is also responsible for negotiation of SR session with SRC, monitoring the SRS’s resource usage, and managing and / or running the split rendering process.

[0017] The 5G network that has a SR MSE architecture shown in FIG. 4, the DN has 5G media function residing at a 5G edge server that is an SRS. The 5G media function is responsible for negotiation of parameters and configuration of a split rendering session for establishment of the split rendering (SR) session with a SRC, monitoring usage of resources of the 5G edge server on which the 5G media function resides and managing and / or running a split rendering process. The Split-Rendering Server resides (e.g., is located in a DN). Although the SRS is illustrated as being located in an external DN, in some implementations, the SRS may reside (e.g., be deployed or located) in the 5G network.

[0018] In the example of a 5G network with a SR MSE architecture shown in FIG. 4, an Application Service Provider (AP) is the application provider that is configured to provides XR service to the UE which use split render via a communication network (e.g., the 5G network shown in FIG. 1). The AP is also configured to provision resources for splitrendering through RTC-1.

[0019] The AP (e.g., a 5G AP) may be authorised to use resources and functionalities provided by the 5G network. Provisioning may include the 5G AP requesting the 5G network to allocate appropriate resources and Quality of Service profiles for a split rendering session. The resources, may for example, include means to carry out compute and render operations for a split rendering session.

[0020] The AP (e.g., 5Gmay deliver (e.g., provide) media to an SRS through RTC-2. The communication between RTC AF and SRS is through RTC-3. RTC-3 is an interface that may for example include the EDGE-3 interface (as described in clause 6.5.7 of 3GPP TS 23.558 V19.0). Signalling (where signalling refers to sending control plane messages, for example for WebRTC session setup or reconfiguration) and media delivery between SRC and SRS is though RTC-4. The RTC AF may provide the split-rendering information to the Media Session Handler defined by RTC-5. SRC discovers the application through RTC-6. The SRC handles the XR runtime. The SRC discovers the client media capabilities through the RTC-7 interface. The 5G Application and AP interact through RTC-8-8.

[0021] In the example shown in FIG. 4, the SRC is the entity that delivers an immersive experience to a user using a UE or a device associated with a UE, while the SRS is the entity that supports the delivery of the immersive experience by performing at least some of the rendering operations of the immersive experience offloaded by the SRC.

[0022] There may be multiple different types of metadata sent from the SRC to the SRS. The SRC may send metadata comprising pose data which is used by SRS to render frames, metadata comprising state data (such as input action or object state) and other tracking data such as controller pose, trackable pose and / or gaze information.

[0023] Pose may be defined as a representation of a position and an orientation of an object in three dimensions. In the context of rendering, a pose may refer to a pose of a virtual camera that captures a video frame which may be displayed to a user. For some XR applications, a pose of a virtual camera may be associated with a position and / or an orientation of a Head Mounted Display (HMD). For other XR applications, the pose of the virtual camera may be associated with a position and / or an orientation of or an input from an input device such as a gaming controller, mouse, or keyboard. In both cases the user may be conceptually associated with a pose for which they expect a displayed frame.

[0024] For split-rendering, the SRC sends pose data to the SRS. The SRS renders a scene as a video frame for the requested pose and sends the rendered video frame back to the device (after the rendered frame is encoded and packaged appropriately). The rendered frame may comprise depth and texture information to support better pose correction on the device.

[0025] Split rendering may allow high quality and resource intensive immersive experiences to be delivered on resource constrained devices such as mobile devices, smart glasses and standalone Head Mounted Displays (HMDs).

[0026] Rendering for XR applications is latency critical. The pose for which a video frame has been generated should match the actual pose of the user of a device that is displaying the video frame. Latency due to, e.g., roundtrip transmission delay of a transmission though a 5G network to the UE and back to the AP, may mean that the actual pose of the user is not that for which pose data was sent.

[0027] To mitigate roundtrip transmission delay of a 5G network, prediction based rendering may be used. In prediction based rendering, a prediction of a pose is used to render a video frame on a SRS. For example, to account for latency between a time a pose is requested to a time a frame is displayed, XR runtimes may provide a pose which is predicted for the estimated time that the rendered frame is to be displayed.

[0028] In split-rendering scenarios where a predicted pose is used to remotely render a video frame, there may be a mismatch between the pose for which the video frame has been generated and the actual pose of the user when displaying the received video frame. This problem may be partially addressed by applying a pose correction at the receiver to reproject the received video frame to the actual pose.

[0029] For example, a pose prediction may vary in accuracy, depending on the pose prediction algorithm used to generate the pose prediction and the prediction horizon, i.e., the time window between a time of a pose prediction is requested and a time at which a pose prediction is generated. In the case generation of an incorrect pose prediction, a communication device or a device associated with a communication device will receive a video frame for display that is not aligned with a current pose of a user of the communication device or device associated with a communication device and a reprojection algorithm may need to be used to re-align the pose.

[0030] Using a reprojection algorithm to re-align a pose may be sufficient if a mismatch between the predicted pose and the current pose is low and if there is not much parallax and occlusions in a scene. However, reprojection using a reprojection algorithm to re-align a pose may lead to QoE drop for the user especially if the pose prediction error is large. QoE should be maintained when a pose prediction error cannot be sufficiently compensated by using a reprojection algorithm to re-align a pose in a communication device comprising the SRC.

[0031] The following shows a View Pose Prediction Error table that describes pose prediction parameters characterizing a pose prediction error that is used for evaluating QoE metrics for AR and / or MR services.Table 1: Viewer Pose Prediction Error

[0032] Clause 6.3.5.3 of TR 26.812 shows a call flow and describes a procedures for a SRC to measure a pose prediction error. The pose prediction error may be used to optimize rendering, coding and delivery of media data for a split-rendering session.

[0033] In other example embodiments, a pose prediction error may be computed by a SRS from pose information received from a SRC along with timing information.

[0034] Referring now to FIGs. 5A and 5B, a flowchart showing an example method or process of a SRC is shown. The example method of process is performed by an SRC. As noted above, the SRC is a logical entity of a communication device (e.g., UE). A communication device comprising or implementing an SRC may comprise at least one processor and at least one memory that stores instructions of the SRC, and when the instructions of the SRC are executed by the at least one processor, the communication device is caused to perform the method of the SRC shown in FIGs. 5A and 5B. The communication device may be a smart phone, smart glasses or an HMD.

[0035] At 501, the SRC obtains a predicted pose of a virtual camera at a first time and time information indicative of a time at which a video frame is expected to be displayed on a display.

[0036] At 502, the SRC provides rendering metadata to a split rendering server (SRS). The rendering metadata may comprise the predicted pose and the time information.

[0037] At 503, the SRC receives, from the SRS, a plurality of rendered and encoded video frames and a rendered pose associated with each respective rendered and encoded video frame of the plurality of rendered and encoded video frames.

[0038] At 504, the SRC obtains a second predicted pose at a second time different than the first time.

[0039] At 505, the SRC decodes at least two rendered and encoded video frames of the plurality of the rendered and encoded video frames to generate at least two decoded video frames.

[0040] At 506, the SRC, based on the render pose associated with each respective decoded video frame and the second predicted pose, selects a primary decoded video frame and one or more secondary decoded video frames from the at least two decoded video frames.

[0041] At 507, the SRC composes a video frame to be displayed based on the primary decoded video frame and at least one or more secondary decoded video frames.

[0042] At 508, the SRC causes the video frame that is composed to be displayed.

[0043] Referring now to FIG. 6, a flowchart of an example method of an SRS is shown. The method or process may be performed by a SRS. As noted above, in some implementations, an apparatus of a data network may comprise the SRS that is configured to perform the method shown in FIG. 6. In some implementations, an apparatus of a communication that comprises a SRS includes at least one processor and at least one memory that stores instructions of the SRS, and when the instructions of the SRS are executed by the at least one processor, cause the apparatus to perform the method of the SRS shown in FIG. 6.

[0044] At 601, the SRS receives rendering metadata from a split rendering client (SRC). The rendering metadata may comprise a predicted pose and time information indicative of a time at which a video frame is expected to be displayed on a display.

[0045] At 602, the SRS obtains a pose prediction error for the predicted pose.

[0046] At 603, the SRS renders and encodes a plurality of video frames based on the rendering metadata, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error.

[0047] At 604, the SRS provides, to the SRC, the plurality of rendered and encoded video frames and the render pose associated with each respective rendered and encoded video frame of the plurality of rendered and encoded video frames.

[0048] The methods of the SRC and SRS described above may provide a split-rendering system, comprising a split-rendering server and at least one split-rendering client, which can leverage a pose prediction error to optimize the end-user QoE (e.g., the QoE experienced by the end-user).

[0049] The following terms are used to explain example embodiments of the methods. The definitions provided for these terms are illustrative and should not be considered limiting.

[0050] The display pose of a video frame is the pose associated with a user at the point in time when the video frame is displayed.

[0051] The render pose of a frame is the pose associated with a virtual camera at the time point when the frame is rendered.

[0052] A view or scene may refer to the image rendered by a virtual camera at a given render pose. In addition to render pose, the view may depend on the camera configuration, e.g. field of view of the virtual camera, near plane, far plane distance, projection and such.

[0053] FIG. 7 shows a split rendering procedure that involves a split rendering client and a split render server according to an example embodiment.

[0054] At step 1, a split rendering session is established between an SRS and SRC. Establishing a split rendering session may comprise negotiating the division of rendering tasks between the SRS and the SRC, negotiating media formats or negotiating meta-data formats. For pose error aware rendering, the parameterization, and settings of pose aware rendering may also be negotiated, for example, using SDP negotiations.

[0055] After the split rendering session is established, a rendering loop starts. The rendering loop begins at steps 2. At step 2, the SRC sends rendering metadata to the SRS. The rendering metadata includes a predicted pose of a virtual camera and timing information indicative of a first time at which a video frame is expected to be displayed on a display. The SRS may make a new pose prediction for a particular display time based on the predicted pose received from the SRC.

[0056] At step 3, pose prediction error is computed, either at the SRS or SRC. Pose prediction error may be determined (e.g., computed) using historical pose predictions and the predicted pose. The pose prediction error may comprise position prediction error and orientation prediction errors as separate or separable quantities. Pose prediction error may comprise a confidence value associated with the predicted pose.

[0057] One example embodiment of computing a pose prediction error of step 3 is illustrated at step 3a. At step 3a, SRC calculates (e.g., determines or computes) a pose prediction error. The pose prediction error may be calculated (e.g., determined or computed) for each of the predicted poses for which a frame is rendered or measured periodically. In this example embodiment, SRC provides, or reports, the pose prediction error to the SRS at step 3b. The calculation (e.g., determination or computation) of the pose prediction error may be reported either continuously or periodically.

[0058] Another example embodiment, of computing a pose prediction error of step 3 is illustrated at step 3c. At 3c, the SRS calculates (e.g., determines or computes) the pose prediction error, either continuously or periodically.

[0059] In some example implementations, pose correction error may be calculated or reported in place of or in addition to pose prediction error. Pose correction error may refer to a measure of error between a ground truth frame for a pose corresponding to a particular display time (i.e., the video frame that would be rendered for the pose display time) and a frame displayed by the SRC at the display time. The displayed frame may be rendered by an SRS for a pose predicted for an estimated display time and then reprojected by the SRC to correspond to the pose at actual display time. In example embodiments, the steps below may be based on pose correction error in place of or in addition of pose prediction error.

[0060] At step 4, based on the predicted pose received in step 2 and the pose prediction error computed at step 3, the SRS selects a set of poses, comprising a primary pose and one or more secondary poses, to render frames for. The poses to render for may be called render poses.

[0061] The selected primary render pose may correspond to the received predicted pose.

[0062] Selecting one or more secondary render poses may be based on at least one of the following: the magnitude of the pose prediction error during a preceding time period of a pre-set length, a confidence value associated with the pose prediction error, an available channel capacity as estimated during a preceding time period of a pre-set length, a motion to render to photon latency during a preceding time period of a pre-set length, a QoS profile assigned to or target for the Split Rendering session, the position and orientation components of the pose prediction error, the confidence value associated with position and orientation components of the pose prediction error, a sequence of predicted poses received during a preceding time period of pre-set length or an estimated time duration between the render operation and the display time associated with the predicted pose.

[0063] For example, if the display time is far in the future (e.g. >100ms), more secondary poses may be selected.

[0064] Selecting a set of poses may include estimating the similarity between frames rendered from the secondary poses to the frame rendered from the primary pose.

[0065] The selection of the set of poses may be triggered if the pose prediction error is above a threshold.

[0066] The threshold which triggers selection of secondary poses may depend on position prediction error and orientation prediction errors individually. For example, if the position prediction error is above a threshold, regardless of the orientation prediction error, the threshold to trigger selection of secondary poses may be met or if the orientation predictionerror is above a threshold regardless of the position prediction error, the threshold to trigger selection of secondary poses may be met.

[0067] In an example implementation, if the position prediction error is below a first threshold and the orientation prediction error is below a second threshold, the number of secondary render poses selected may be 0.

[0068] Optionally, in the case where the number of secondary render poses selected is 0, the field of view (FoV) of the frame rendered for the primary pose may be increased. The increase in FoV may be based on the magnitude of the orientation prediction error.

[0069] In another example implementation, if the position prediction error is below a threshold, the number of secondary render poses selected may be 1, with the secondary render pose having the same position component as the primary render pose and the orientation component either being the same as the primary render pose, but with the FoV of the secondary frame is increased if the orientation prediction error is below a third threshold or different from the primary render pose, if the orientation prediction error is above the third threshold.

[0070] At step 5, based on the rendering metadata, the SRS renders the video frames for the selected render poses and encodes them for transmission to the SRC. Each respective rendered and encoded video frame of the plurality of rendered and encoded video frames has an associated render pose, wherein each associated render pose is determined based on the received predicted render pose and the pose prediction error.

[0071] Step 5 may include rendering a primary frame corresponding to the primary render pose and one or more secondary frame corresponding to the one or more secondary poses.

[0072] For rendering multiple frames, the SRS may instantiate virtual cameras corresponding to the selected secondary render poses and render the primary and secondary frames synchronously in parallel, transform existing virtual cameras to the selected secondary render poses and render the primary and secondary frames synchronously in parallel or use a single virtual camera to render the primary and secondary frames one after another.

[0073] At step 6, the rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames are provided, or transmitted, to the SRC.

[0074] A different priority may be allocated to the different frames, depending on error magnitude or similarity between frames. The priority may be used by the RAN scheduler to prioritize packets or to drop less important frames.

[0075] The frames rendered for the selected render poses may be encoded using a multiview codec such as MV-HEVC, each frame being encoded as a view. The primary view may correspond to the primary frame being encoded as the base-layer. The secondary view(s) may correspond to the secondary frame(s) being encoded as dependent enhancement layers.

[0076] Alternatively, the frames rendered for the selected render poses may be encoded using a single-layer codec, by leveraging frame packing. Coding of auxiliary depth or alpha channels may additionally be used. In some examples the encoded views may comprise color information or depth information or both. In some examples, only the encoded secondary view(s) may comprise depth information.

[0077] For rate control, a higher bitrate or resolution may be allocated for the primary frames and a lower bitrate or resolution may be allocated for the secondary frames. The bit budget or resolution of a secondary frame may be based on a difference between the corresponding secondary render pose and the primary render pose.

[0078] In some example embodiments, video streams corresponding to the primary and secondary frames are transmitted over RTP.

[0079] The encapsulation may further comprise packaging a render timestamps corresponding to each of the views being packaged.

[0080] For the packaging the frames may be encapsulated within an appropriate transport protocol packet, for example, RTP packets. An indicator of number of views being transmitted maybe encapsulated.

[0081] In some example embodiments, the render pose(s) of a frame may be indicated in the payload of a header extension of the RTP packet(s) comprising at least part of the frame.

[0082] In some example embodiments, only the render pose corresponding to the primary frame may be indicated, while the render poses of secondary frames may be indicated as relative pose differences.

[0083] In some example embodiments, only the render pose corresponding to the primary frame may be indicated, while the render poses of secondary poses are obtained by the SRC from the render pose and the currently used template of views.

[0084] Different importance may be allocated in an appropriate protocol header extension, depending on the pose used for rendering the encoded frame being encapsulated. For example, if RTP is used for media transport, in RTP PDU Importance (PSI) field of the PDU set header extension specified in 3GPP TS 26.522 may be set. Importance may be set for a frame, or a slice, or a tile, or a subpicture.

[0085] In an example implementation, the importance may be set to high when packaging primary frames and low when packaging secondary frames. The encapsulation may comprise packaging in the transport protocol packet or packet header, a corresponding render pose with each encapsulated and encoded view. For example, when RTP is used, the render pose may be packaged in an RTP header extension for pose as specified in 3GPP TS 26.522.

[0086] In some example embodiments, video streams corresponding to the primary and secondary frames are transmitted over SCTP.

[0087] At step 7, the SRC uses the received frames to compose a frame for the current display pose and displays the frame.

[0088] Step 7 may comprise decoding one or more of the encoded frames.

[0089] Step 7 may comprise obtaining a second predicted pose, for a second time different than the first time. Based on the predicted pose and the render pose associated with each respective decoded frame the method may comprise selecting a primary decoded frame and one or more secondary decoded frames from the at least two decoded frames. The primary frame may be selected such that the difference between the render pose of the selected primary frame and the predicted pose is the lowest or below a first threshold. The one or more secondary frames may be selected such that the difference between the render poses of the selected secondary frames and the predicted pose is the lower than a second threshold.

[0090] Composing a frame in step 7 may be further based on a depth map determined based on the one or more secondary decoded frames. Step 7 may comprise estimating a depth map for the primary frame based on the selected one or more secondary frames. Composing the frame may comprise reprojecting the primary frame using the primary frame and a depth map, using, for example, a depth image based rendering techniques.

[0091] In some examples, the reprojected video frame may be processed to correct reprojection artefacts or fill missing pixels in the reprojected video frame, for example using optical flow or in-painting techniques.

[0092] A video frame may be synthesised using the received video frames by selecting one primary video frame and using frame extrapolation techniques, such as optical flow, exploiting the secondary frames.

[0093] In one example embodiment, the SRC and the SRS may negotiate a set of possible configurations, or templates. A template may be defined as a pre-configured set of views coupled to a set of rendered poses to be generated by the SRC for the SRS.

[0094] Referring to FIG. 8, a flowchart of an example method or process of an SRC is shown. The SRC may be a logical entity. As noted above, a communication device (e.g., UE) may comprise a SRC. The communication device (e.g., UE) may comprise at least one processor and at least one memory that stores instructions of the SRC, and when the instructions of the SRC are executed by the at least one processor, the communication device is caused to perform the method of the SRC shown in FIG. 8. The communication device may comprise a smart phone, smart glasses or a HMD.

[0095] At 801, SRC performs a negotiation with a split rendering server to obtain at least one template for a split rendering session. In some examples, respective templates of the at least one template comprises a plurality of rendered image frames associated with a set of render poses. In some examples, each respective template of the at least one template comprises a plurality of render poses (e.g., a set of render poses), and for each respective render pose of the plurality of render poses (e.g., included in the set of render poses), positioning information of a virtual camera, orientation information of the virtual camera, and a field of view information of the virtual camera. In some examples, each respective template also comprises at least one of an indication of a number render poses (e.g., an indication of a number of render poses in the set of render poses), a depth channel or alpha channel (e.g., depth information or alpha information) for each respective render pose, or resolutions of frames rendered for each respective render pose.

[0096] At 802 the SRC obtains a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display.

[0097] At 803, the SRC provides rendering metadata to the split rendering server. The rendering metadata comprises the predicted pose and the time information.

[0098] At 804, the SRC, receives from the split rendering server, based on the negotiated at least one template, a plurality of rendered and encoded frames and a render poseassociated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0099] Referring to FIG. 9. a flowchart showing an example method or process of an SRS is shown. The example method or process is performed by an SRS. The SRS may be a logical entity. As noted above, in some implementations, an apparatus of a data network may comprise at least one processor and at least one memory that stores instructions of an SRS, and when the instructions of the SRS are executed by the at least one processor, the apparatus is caused to perform the method of the SRS shown in FIG. 9.

[0100] At 901, the SRS performs a negotiation with a split rendering client of a communication device (e.g., UE) to obtain at least one template for a split rendering session. In some examples, respective templates of the at least one template comprises a plurality of rendered image frames associated with a set of render poses. In some examples, each respective template of the at least one template comprises a plurality of render poses (e.g., a set of render poses), and for each respective render pose of the plurality of render poses (e.g., included in the set of render poses), positioning information of a virtual camera, orientation information of the virtual camera, and a field of view information of the virtual camera. In some examples, each respective template also comprises at least one of an indication of a number render poses (e.g., an indication of a number of render poses in the set of render poses), a depth channel or alpha channel (e.g., depth information or alpha information) for each respective render pose, or resolutions of frames rendered for each respective render pose.

[0101] At 902, the SRS receives rendering metadata from the split rendering client, the rendering metadata comprising a predicted pose of a virtual camera and a time information indicative of a time at which a video frame is expected to be displayed on a display.

[0102] At 903, the SRS obtains a pose prediction error for the predicted pose.

[0103] At 904, the SRS renders and encodes a plurality of frames based on the rendering metadata and the least one template, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error.

[0104] At 905, the SRS provides, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0105] In some embodiments, the SRS may determine to adapt the template for the split rendering session and provide a request to the split rendering server to adapt the template.

[0106] The SRS may modify the template or select a new template for the SRC to use.

[0107] In one example embodiment, the SRC re-evaluates the template and may request the SRC to switch to use a different template (e.g., change the template to another at least one template negotiated for the split rendering session), adjust or modify the parameters of the currently used template, adjust or modify the parameters of a different template and switch to that template and / or update the output accordingly. The re-evaluation of the template may be performed periodically.

[0108] This re-evaluation and the corresponding request may be based on user input detected at the apparatus associated with the SRC or XR experience, (e.g. logic, content, genre and “pace” of the XR experience). For example, a first-person shooter (FPS) game may be considered fast paced and may require fewer render poses.

[0109] Alternatively, or in addition, the re-evaluation and corresponding request may be based on at least one of the following: pose-error, user’s trajectory during a time window, controller speed, acceleration, trajectory, gaze pattern, pose pattern, user’s history on other games or XR applications or user’s environment.

[0110] Alternatively, or in addition, the re-evaluation and the corresponding request may be based on operating conditions of the apparatus or a change therein, network conditions detected by the apparatus, an estimated quality of experience of the split rendering session, an estimated quality of service of the split rendering session or a user input detected by the apparatus. The user input may indicate to the SRC to change the template and / or the user may be able to select a template to use.[OHl] As an example, operating conditions of the apparatus may include battery level, temperature or computational resources available for SR operation.

[0112] Network conditions may include, for example, channel capacity, network rtt, motion to render to photon latency, motion to high quality latency, jitter or packet interarrival time.

[0113] An estimate of user QoE or service QoS may be based on, for example, network conditions, operating conditions, and state of the SR media session. The state of the SRmedia session may comprise the state of the receiving or decoding buffer (capacity, fill rate, etc.), dropped frames (number, proportion, frequency etc.), corrupted frames (number, proportion, frequency etc.), video stalls (number, duration, etc.) or a no reference video quality metric.

[0114] In an example embodiment, a request to adjust parameters of the currently used template may comprise an indication to send only the primary frame of the plurality of rendered and encoded frames for a time period.

[0115] In an example embodiment, the SRC may request additional information about the rendering to be transmitted by the SRS.

[0116] Additional information about the rendering to be transmitted may include an intra coded reference picture for streams corresponding to the primary frame or the secondary frames, an IDR frame for streams corresponding to the primary frame or the secondary frames, a GDR request for streams corresponding to the primary frame or the secondary frames and / or a request for a particular bitrate for streams corresponding to the primary frame or the secondary frames.

[0117] In some example embodiments, the request to adapt the at least one template may be sent over RTCP packet or, RTCP app messages.

[0118] In some example embodiments, negotiation of the at least one template, and the requests for and acknowledgement of adaptation (e.g., adjustment or switching) of the at least one template may be sent over WebRTC data channel.

[0119] In some example embodiments, negotiation of the at least one template, and the requests for and acknowledgement of adaptation (e.g., adjustment or switching) of the at least one template may be sent over IMS data channel.

[0120] An apparatus may comprise means for obtaining a predicted pose of a virtual camera at a first time and time information indicative of a time at which a video frame is expected to be displayed on a display, providing rendering metadata to a split rendering server, the rendering metadata comprising the predicted pose and the time information, receiving, from the split rendering server, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames, obtaining a- second predicted pose at a second time different than the first time, decoding at least two rendered and encoded frames of the plurality of the rendered and encoded frames to generate at least two decoded frames, based on the render pose associated with each respective decoded frame and the second predicted pose, selecting a primarydecoded frame and one or more secondary decoded frames from the at least two decoded frames, composing a video frame to be displayed based on the primary decoded frame and at least one or more secondary decoded frames and causing the video frame that is composed to be displayed.

[0121] An apparatus may comprise means for performing a negotiation with a split rendering server of at least one template for a split rendering session, wherein the at least one template comprises a plurality of rendered image frames associated with a set of render poses, obtaining a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display, providing rendering metadata to the split rendering server, the rendering metadata comprising the predicted pose and the time information, receiving, from the split rendering server, based on the negotiated at least one template, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0122] The apparatus may comprise a user equipment, such as a mobile phone or HMD, be the user equipment, or be comprised in the user equipment, or a chipset for performing at least some actions of / for the user equipment. The user equipment may be the communication device as illustrated in FIG. 2.

[0123] An apparatus may comprise means for receiving rendering metadata from a split rendering client, the rendering metadata comprising a predicted pose of a virtual camera and a time information indicative of a time at which a video frame is expected to be displayed on a display, obtaining a pose prediction error for the predicted pose, rendering and encoding a plurality of frames based on the rendering metadata, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error and providing, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0124] An apparatus may comprise means for performing a negotiation with a split rendering client of at least one template for a split rendering session, wherein the at least one template comprises a plurality of rendered image frames associated with a set of render poses; receiving rendering metadata from the split rendering client, the rendering metadata comprising a predicted pose of a virtual camera and a time information indicative of atime at which a video frame is expected to be displayed on a display, obtaining a pose prediction error for the predicted pose, rendering and encoding a plurality of frames based on the rendering metadata and the a least one template, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error and providing, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

[0125] The apparatus may comprise a split rendering server as described with reference to FIG. 3, be the split rendering server, or be comprised in the split rendering server, or a chipset for performing at least some actions of / for the split rendering server.

[0126] It should be understood that the apparatuses may comprise or be coupled to other units or modules etc., such as radio parts or radio heads, used in or for transmission and / or reception. Although the apparatuses have been described as one entity, different modules and memory may be implemented in one or more physical or logical entities.

[0127] It is noted that whilst some embodiments have been described in relation to 5G networks, similar principles can be applied in relation to other networks and communication systems such as 6G networks or 5G-Advanced networks. Therefore, although certain embodiments were described above by way of example with reference to certain example architectures for wireless networks, technologies and standards, embodiments may be applied to any other suitable forms of communication systems than those illustrated and described herein.

[0128] It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention.

[0129] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.

[0130] In general, the various embodiments may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. Some aspects of the disclosuremay be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0131] As used in this application, the term “circuitry” may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(iii) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0132] This definition of circuitry may apply to all uses of this term “means” in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0133] The term “means for” may define at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause an apparatus at least to perform the steps listed after the term.

[0134] The embodiments of this disclosure may be implemented by computer software comprising instructions which when executed by a processor of a communication device, such as in the processor entity, or by hardware, or by a combination of software and hardware cause the communication device to carry out or perform the operations of the method shown in FIG. 8. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computer-executable components which, when the program is run, are configured to carry out embodiments. The one or more computerexecutable components may be at least one software code or portions of it.

[0135] Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non- transitory media. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).

[0136] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may comprise one or more of central processing units (CPU), graphics processing units (GPU), tensor processing units (TPUs), microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), FPGA, gate level circuits and processors based on multi core processor architecture, as non-limiting examples.

[0137] Embodiments of the disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logiclevel design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0138] The scope of protection sought the disclosure is set out by the independent claims and the dependent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure.

[0139] The foregoing description has provided by way of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.

Claims

CLAIMS1. An apparatus comprising: a split rendering client configured for: obtaining a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display; providing rendering metadata to a split rendering server, the rendering metadata comprising the predicted pose and the time information; receiving, from the split rendering server, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames; obtaining a- second predicted pose at a second time different than the first time; decoding at least two rendered and encoded frames of the plurality of the rendered and encoded frames to generate at least two decoded frames; based on the render pose associated with each respective decoded frame and the second predicted pose, selecting a primary decoded frame and one or more secondary decoded frames from the at least two decoded frames; composing a video frame to be displayed based on the primary decoded frame and at least one or more secondary decoded frames; and causing the video frame that is composed to be displayed.

2. The apparatus according to claim 1, wherein the split rendering client is further configured for computing a pose prediction error for the predicted pose and providing the pose prediction error to the split rendering server.

3. The apparatus according to claim 2, wherein the pose prediction error for the predicted pose is computed using historical predicted poses and the predicted pose.

4. The apparatus according to claim 2 or claim 3, wherein the pose prediction error for the predicted pose comprises a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

5. The apparatus according to any of claims 1 to 4, wherein each respective rendered and encoded video frame of the plurality of rendered and encoded frames is associated with at least one of a priority level or a quality of service profile.

6. The apparatus according to any of claims 1 to 5, wherein composing a frame is further based on depth map determined based on the one or more secondary decoded frames.

7. An apparatus comprising: a split rendering server configured for: receiving rendering metadata from a split rendering client, the rendering metadata comprising a predicted pose of a virtual camera and a time information indicative of a time at which a video frame is expected to be displayed on a display; obtaining a pose prediction error for the predicted pose; rendering and encoding a plurality of frames based on the rendering metadata, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error; and providing, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

8. The apparatus according to claim 7, wherein obtaining the pose prediction error for the predicted pose comprises receiving the pose prediction error from the split rendering client.

9. The apparatus according to claim 7 or claim 8, wherein the pose prediction error for the predicted pose comprises at least one of the following: a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

10. The apparatus according to claim 9, wherein at least one of the at least associated render poses or the number thereof is based on at least one of the following: a magnitude of the pose prediction error, a magnitude of the predicted position error, a magnitude of the predicted orientation error, the confidence value associated with the predicted pose, an available channel capacity, a latency, a quality of service profile, at least one historical predicted pose of the virtual camera or a time period between the rendering operation and the first time.

11. The apparatus according to any of claims 7 to 10, wherein the plurality of frames are encoded using multi-view encoding.

12. The apparatus according to any of claims 7 to 11, wherein each respective rendered and encoded frame of the plurality of rendered and encoded frames is associated with at least one of a priority level or a quality of service profile.

13. A method of a split rendering client, the method comprising: obtaining a predicted pose of a virtual camera at a first time and a time information indicative of a time at which a video frame is expected to be displayed on a display; providing rendering metadata to a split rendering server, the rendering metadata comprising the predicted pose and the time information; receiving, from the split rendering server, a plurality of rendered and encoded frames and a render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames; obtaining a second predicted pose at a second time different than the first time; decoding at least two rendered and encoded frames of the plurality of the rendered and encoded frames to generate at least two decoded frames; based on the render pose associated with each respective decoded frame and the second predicted pose, selecting a primary decoded frame and one or more secondary decoded frames from the at least two decoded frames; composing a video frame to be displayed based on the primary decoded frame and at least one or more secondary decoded frames; and causing the video frame that is composed to be displayed.

14. The method according to claim 13, further comprising: computing a pose prediction error for the predicted pose and providing the pose prediction error to the split rendering server.

15. The method according to claim 14, wherein the pose prediction error for the predicted pose is computed using historical predicted poses and the predicted pose.

16. The method according to claim 14 or claim 15, wherein the pose prediction error for the predicted pose comprises a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

17. The method according to any of claims 14 to 16, wherein each respective rendered and encoded video frame of the plurality of rendered and encoded frames is associated with at least one of a priority level or a quality of service profile.

18. The method according to any of claims 14 to 17, wherein composing a frame is further based on depth map determined based on the one or more secondary decoded frames.

19. A method of a split rendering server, the method comprising: receiving rendering metadata from a split rendering client, the rendering metadata comprising a predicted pose of a virtual camera and a time information indicative of a time at which a video frame is expected to be displayed on a display; obtaining a pose prediction error for the predicted pose; rendering and encoding a plurality of frames based on the rendering metadata, each respective rendered and encoded video frame of the plurality of rendered and encoded video frames having an associated render pose, wherein each associated render pose is determined based on the received predicted pose and the pose prediction error; and providing, to the split rendering client, the plurality of rendered and encoded frames and the render pose associated with each respective rendered and encoded frame of the plurality of rendered and encoded frames.

20. The method according to claim 19, wherein obtaining the pose prediction error for the predicted pose comprises receiving the pose prediction error from the split rendering client.

21. The method according to claim 19 or claim 20, wherein the pose prediction error for the predicted pose comprises at least one of the following: a predicted position error for the predicted pose, a predicted orientation error for the predicted pose or a confidence value associated with the predicted pose.

22. The method according to claim 21, wherein at least one of the at least associated render poses or the number thereof is based on at least one of the following: a magnitude of the pose prediction error, a magnitude of the predicted position error, a magnitude of the predicted orientation error, the confidence value associated with the predicted pose, an available channel capacity, a latency, a quality of service profile, at least one historical predicted pose of the virtual camera or a time period between the rendering operation and the first time.

23. The method according to any of claims 19 to 22, wherein the plurality of frames are encoded using multi-view encoding.

24. The method according to any of claims 19 to 23, wherein each respective rendered and encoded frame of the plurality of rendered and encoded frames is associated with at least one of a priority level or a quality of service profile.

25. A computer-readable medium comprising instructions which, when executed by at least one apparatus causes the apparatus to perform the method of any of claims 13 to 24.

Citation Information

Patent Citations

  • Deep learning based head motion prediction for extended reality

    US11144117B1

  • Method and device for reducing performance difference between contents and devices in communication system

    US20220066543A1

  • Method and device for performing rendering using latency compensatory pose prediction with respect to three-dimensional media data in communication system supporting mixed reality / augmented reality

    US20230316583A1

Cited By

  • Multi-model scheduling method and device based on error perception, equipment and storage medium

    CN121934985A

  • Error-aware multi-model scheduling method and device, equipment and storage medium

    CN121934985B