Apparatus, method and computer program for optimisation of a gaze aware rendering and encoding process
A gaze-based optimization profile in split-rendering architectures dynamically adjusts rendering and encoding to optimize resource allocation, addressing computational and power constraints and ensuring high-quality immersive experiences on communication devices.
Patent Information
- Application Number
- GB2024004753
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-10-15
AI Technical Summary
Existing technologies do not effectively address the computational and power resource constraints of communication devices in delivering high-quality immersive experiences, particularly in split-rendering architectures, failing to account for changes in operating conditions and network conditions.
Implementing a gaze-based optimization profile that dynamically adjusts rendering and encoding parameters based on gaze information, network conditions, and user input, using a split rendering client and server to optimize resource allocation in split-rendering sessions.
Enhances the quality of immersive experiences on resource-constrained devices by efficiently allocating rendering and encoding resources, adapting to changing conditions, and maintaining user experience quality.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
A communication system can be seen as a facility that enables communication sessions between two or more communication devices, or provides communication devices access to a data network. A mobile or wireless communication network is one example of a communication network. A communication device may be provided with a service by an application server. A mobile or wireless communication network may operate in accordance with standards such as those provided by 3GPP (Third Generation Partnership Project) or ETSI (European Telecommunications Standards Institute). Examples of mobile or wireless communication network that operate in accordance with 3GPP standards are generally referred to as 4G (4th Generation) networks, 5G (5th Generation) network, 5G-Advanced networks and 6G networks. Summary In a first aspect there is provided an apparatus comprising a split rendering client comprising means for sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client, sending, to the split rendering server, a request for adaptation of a gaze based optimization profile for a split rendering session, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile, receiving, from the split rendering server, an encoded video frame, decoding the encoded video frames to generate a decoded video frame and displaying the decoded video frame. The indication of an adaptation of a gaze based optimization profile may comprise an indication to change one or more parameters of the gaze optimization profile. The indication of an adaptation of a gaze based optimization profile may comprise an indication of a change to one or more parameters of the gaze optimization profile. The gaze based optimization profile may comprise a first spatial quality map used for performing rendering the video frame and a second spatial quality map used for encoding the video frame. The first spatial quality map for rendering the video frame may indicates a spatially varying quality of a rendered frame based on gaze information. The second spatial quality map for encoding the video frame may indicate a spatially varying quality of an encoded frame based on gaze information. The means may be for determining to request adaptation of the gaze based optimization profile based on at least one of the following: operating conditions of the apparatus, network conditions detected by the apparatus, an estimated quality of experience of the split rendering session, an estimated quality of service of the split rendering session or a user input detected by the apparatus. The indication to change one or more parameters of the gaze based optimization profile may comprise an indication to change a method of rendering or an indication to change a method of encoding. The indication of the adaptation to the gaze based optimization profile may comprise an indication to change the gaze based optimization profile to another gaze based optimization profile selected from a plurality of gaze based optimization profiles negotiated for the split rendering session. The means may be further for sending the request for adaptation of the gaze based optimization profile with metadata relating to the video frame for use in combination with the gaze based optimization profile for rendering and encoding the video frame. The means may be for receiving, from the split rendering server, information relating to the adapted gaze based optimization profile. The means may be for performing further processing of the received video frame based on the received information relating to the adapted gaze based optimization profile. In a second aspect there is provided an apparatus comprising a split rendering server comprising means for receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client, performing rendering using the metadata and a gaze based optimization profile for a split rendering session, encoding the video frame using the gaze based optimization profile to generate an encoded video frame and sending, to the split rendering client, the encoded video frame. The means may be further for performing a negotiation of the gaze based optimization profile for the split rendering session with the split rendering client. The negotiation may be performed based on at least one of the following: the capabilities of an apparatus comprising the split rendering client, the application being rendered in the split rendering session or user input detected by the apparatus. The gaze based optimization profile may comprise a first spatial quality map used for performing rendering and a second spatial quality map used for encoding the video frame. The first spatial quality map used for performing rendering the video frame may indicate a spatially varying quality of a rendered frame based on gaze information. The second spatial quality map used for encoding the video frame may indicate a spatially varying quality of an encoded frame based on gaze information. The means may further be for determining to adapt the gaze based optimization profile for the split rendering session based on the metadata that is received, adapting the gaze based optimization profile and performing rendering to generate a second video frame using the gaze based optimization profile that is adapted and encoding the second video frame using the gaze based optimization profile that is adapted. Determining to adapt the gaze based optimization profile may comprise receiving a request for adaptation of the gaze based optimisation profile for the split rendering session, wherein the request for adaptation of the gaze based optimisation profile comprises an indication of an adaptation to the gaze based optimisation profile and adapting the gaze based optimisation profile based on the request. The indication of an adaptation of a gaze based optimization profile may comprise an indication to change one or more parameters of the gaze optimization profile. The indication of an adaptation of a gaze based optimization profile may comprise an indication of a change to one or more parameters of the gaze optimization profile. The indication to change one or more parameters of the gaze based optimisation profile may comprise at least one of a request to change a method of rendering or an indication to change a method of encoding. The means may further be for receiving the request for adaptation of the gaze based optimisation profile with metadata relating to the video frame for use in combination with the gaze based optimisation profile for rendering and encoding the video frame. The indication of the adaptation to the gaze based optimisation profile may comprise an indication to change the gaze based optimization profile of another gaze based optimization profile selected from a plurality of gaze based optimization profiles negotiated for the split rendering session. The means may further be for sending, to the split rendering client, information relating to the adapted gaze based optimisation profile. In a third aspect there is provided a method comprising, at a split rendering client, sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client; sending, to the split rendering server, a request for adaptation of a gaze based optimization profile for a split rendering session, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile, receiving, from the split rendering server, an encoded video frame, decoding the encoded video frames to generate a decoded video frame and displaying the decoded video frame. The indication of an adaptation of a gaze based optimization profile may comprise an indication to change one or more parameters of the gaze optimization profile. The indication of an adaptation of a gaze based optimization profile may comprise an indication of a change to one or more parameters of the gaze optimization profile. The gaze based optimization profile may comprise a first spatial quality map used for performing rendering the video frame and a second spatial quality map used for encoding the video frame. The first spatial quality map for rendering the video frame may indicates a spatially varying quality of a rendered frame based on gaze information. The second spatial quality map for encoding the video frame may indicate a spatially varying quality of an encoded frame based on gaze information. The method may comprise determining to request adaptation of the gaze based optimization profile based on at least one of the following: operating conditions of the apparatus, network conditions detected by the apparatus, an estimated quality of experience of the split rendering session, an estimated quality of service of the split rendering session or a user input detected by the apparatus. The indication to change one or more parameters of the gaze based optimization profile may comprise an indication to change a method of rendering or an indication to change a method of encoding. The indication of the adaptation to the gaze based optimization profile may comprise an indication to change the gaze based optimization profile to another gaze based optimization profile selected from a plurality of gaze based optimization profiles negotiated for the split rendering session. The method may comprise sending the request for adaptation of the gaze based optimization profile with metadata relating to the video frame for use in combination with the gaze based optimization profile for rendering and encoding the video frame. The method may comprise receiving, from the split rendering server, information relating to the adapted gaze based optimization profile. The method may comprise performing further processing of the received video frame based on the received information relating to the adapted gaze based optimization profile. In a fourth aspect there is provided a method comprising, at a split rendering server, receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client, performing rendering using the metadata and a gaze based optimization profile for a split rendering session, encoding the video frame using the gaze based optimization profile to generate an encoded video frame and sending, to the split rendering client, the encoded video frame. The method may comprise performing a negotiation of the gaze based optimization profile for the split rendering session with the split rendering client The negotiation may be performed based on at least one of the following: the capabilities of an apparatus comprising the split rendering client, the application being rendered in the split rendering session or user input detected by the apparatus. The gaze based optimization profile may comprise a first spatial quality map used for performing rendering and a second spatial quality map used for encoding the video frame. The first spatial quality map used for performing rendering the video frame may indicate a spatially varying quality of a rendered frame based on gaze information. The second spatial quality map used for encoding the video frame may indicate a spatially varying quality of an encoded frame based on gaze information. The method may comprise determining to adapt the gaze based optimization profile for the split rendering session based on the metadata that is received, adapting the gaze based optimization profile and performing rendering to generate a second video frame using the gaze based optimization profile that is adapted and encoding the second video frame using the gaze based optimization profile that is adapted. Determining to adapt the gaze based optimization profile may comprise receiving a request for adaptation of the gaze based optimisation profile for the split rendering session, wherein the request for adaptation of the gaze based optimisation profile comprises an indication of an adaptation to the gaze based optimisation profile and adapting the gaze based optimisation profile based on the request. The indication of an adaptation of a gaze based optimization profile may comprise an indication to change one or more parameters of the gaze optimization profile. The indication of an adaptation of a gaze based optimization profile may comprise an indication of a change to one or more parameters of the gaze optimization profile. The indication to change one or more parameters of the gaze based optimisation profile may comprise at least one of a request to change a method of rendering or an indication to change a method of encoding. The method may comprise receiving the request for adaptation of the gaze based optimisation profile with metadata relating to the video frame for use in combination with the gaze based optimisation profile for rendering and encoding the video frame. The indication of the adaptation to the gaze based optimisation profile may comprise an indication to change the gaze based optimization profile of another gaze based optimization profile selected from a plurality of gaze based optimization profiles negotiated for the split rendering session. The method may comprise sending, to the split rendering client, information relating to the adapted gaze based optimisation profile. In a fifth aspect there is provided an apparatus comprising at least one processor, and at least one memory storing instructions which, when executed by the processor, cause the apparatus at least to perform a method according to the first or second aspects. In a sixth aspect there is provided a computer readable medium comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least the method according to the first or second aspects. In a seventh aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the method according to any of the fifth to eighth aspect. In the above, many different embodiments have been described. It should be appreciated that further embodiments may be provided by the combination of any two or more of the embodiments described above. Description of Figures Embodiments will now be described, by way of example only, with reference to the accompanying Figures in which: Figure 1 shows a schematic diagram of an example 5GS communication system; Figure 2 shows a schematic diagram of an example mobile communication device; Figure 3 shows a schematic diagram of an example control apparatus; Figure 4 shows a schematic diagram of an example MSE split rendering architecture; Figure 5 shows a schematic diagram of an example I BAGS split rendering architecture; Figure 6 shows a representation of relative acuity of human vision per degree of eccentricity from gaze vector; Figure 7 shows an illustration of a foveated gaze; Figure 8 shows a flowchart of a method according to an example embodiment; Figure 9 shows a flowchart of a method according to an example embodiment; Figure 10a shows a representation of an example quality map according to an example embodiment; Figure 10b shows a representation of an example quality map according to an example embodiment; Figure 11 shows a signalling flow chart according to an example embodiment. Detailed description Before explaining in detail the examples, certain general principles of a communication system of a communication network and communication devices that receive XR services provided by XR service providers are briefly explained with reference to Figure 1, Figure 2 and Figure 3 to assist in understanding the technology underlying the described examples. Figure 1 shows a schematic representation of a 5G network 100 that connects a UE 102 (which may also be referred to as a communication device or a terminal) to a DN 110 that includes a Split-Rendering Server (SRS), a Real-Time Communications AF and an Application Provider (see Figure 4). The 5G network 100 may comprise a 5G radio access network (5G-RAN) 104 and a 5G core network (5GC) 106. One or more application functions (AF) 108 may be connected to the 5GC 100 (e.g., a network exposure function of the 5GC) to enable the one or more AFs to access services provided by network functions of the 5GC.. Although the AF 108 is shown in Figure 1 as being external to the 5GC 106, in some embodiments, one or more of the AFs 108 may be a trusted AF that is part of 5GC (e.g., internal to the 5GC). The 5GC 106 comprises one or more network functions (otherwise referred to a network entities). The one or more network functions of the 5GC 106 may comprise one or more Access and mobility Management Functions (AMF) 112, one or more session management functions (SMF) 114, an authentication server function (AUSF) 116, a Unified Data Management (UDM) 118, one or more user plane functions (UPF) 120, a Unified Data Repository (UDR) 122 and / or a Network Exposure Function (NEF) 124. The UPF is controlled by the SMF (Session Management Function) that receives policies from a PCF (Policy Control Function). The UPF of the 5GC 106 is connected to the Radio Access Network (RAN) via a N3 interface. The 5G-RAN 104 may comprise one or more radio access network nodes (RANs). A radio access network node may comprise a gNodeB (gNB) and a gNB may comprise one or more gNB Distributed Units (DUs) connected to a gNodeB (gNB) Centralized Unit (CU). PDU Session Anchor (PSA) is the term given to the UPF (User Plane Function) which terminates the N6 interface of a PDU session within a 5G core network. In order to support selective traffic routing to the DN(data network), the SMF may control the data path of a PDU Session so that the PDU Session may simultaneously correspond to multiple N6 interfaces. The UPF that terminates each of these interfaces is said to support PDU Session Anchor functionality. Each PDU Session Anchor supporting a PDU Session provides a different access to the same data network DN. A communication device will now be described in more detail with reference to Figure 2 showing a schematic, partially sectioned view of a communication device 200. Such a communication device is often referred to as user equipment (UE) or terminal. A communication device may be any device capable of sending and receiving wireless signals, including radio signals. Non-limiting examples of a communication device comprise a mobile station (MS) or mobile device such as a mobile phone or what is known as a ’smart phone’, a computer provided with a wireless interface card or other wireless interface facility (e.g., USB dongle), personal data assistant (PDA) or a tablet provided with wireless communication capabilities, voice over IP (VoIP) phones, portable computers, desktop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), smart devices, wireless customer-premises equipment (CPE), or any combinations of these or the like. A mobile communication device may provide, for example, communication of data for carrying communications such as voice, electronic mail (email), text message, multimedia and so on. Users may thus be offered and provided numerous services via their communication devices. Non-limiting examples of these services comprise two-way or multi-way calls, data communication or multimedia services or simply an access to a data communications network system, such as the Internet. Users may also be provided broadcast or multicast data. Nonlimiting examples of the content comprise downloads, television and radio programs, videos, advertisements, various alerts, and other information. A mobile device is typically provided with at least one data processing entity 201, at least one memory 202 and other possible components 203 for use in software and hardware aided execution of tasks it is designed to perform, including control of access to and communications with access systems and other communication devices. The data processing, storage and other relevant components can be provided on an appropriate circuit board and / or in chipsets. This feature is denoted by reference 204. The user may control the operation of the mobile device by means of a suitable user interface such as key pad 205, voice commands, touch sensitive screen or pad, combinations thereof or the like. A display 208, a speaker and a microphone can be also provided. Furthermore, a mobile communication device may comprise appropriate connectors (either wired or wireless) to other devices and / or for connecting external accessories, for example hands-free equipment, thereto. The mobile device 200 may receive signals over an air or radio interface 207 via appropriate apparatus for receiving and may transmit signals via appropriate apparatus for transmitting radio signals. In Figure 2 transceiver apparatus is designated schematically by block 206. The transceiver apparatus 206 may be provided for example by means of a radio part and associated antenna arrangement. The antenna arrangement may be arranged internally or externally to the mobile device. Figure 3 shows an example of split rendering server 300. The split rendering server 300 comprises at least one memory 301, at least one data processing unit 302, 303 and an network interface 304. The at least one data processing unit 302 and 303 may comprise a CPU 302 and a hardware accelerator 303.The memory 301 of the split rendering server stores instructions which when executed by the at least one data processing unit 302, 303, cause the split rendering server to perform operations, for example the operations described in Figure 9 or the operations associated with steps 3, 4a, 4b, 5a, and 5b of the procedure shown in Figure 11. Rendering videos for display on, e.g., a communication device such as the communication device 200 described with reference to Figure 2 or a head-mounted display (HMD) connected to the communication device via a wireless connection or wirelessly, is computationally and power intensive. The computational and power requirements for rendering videos may be especially high for immersive applications which may collectively be known as extended reality (XR) applications. XR includes, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR) etc. Immersive applications may require high quality graphics for a reasonable QoE. A communication device and / or a display (such as a head-mounted display) connected to a communication device may not have sufficient computation and / or energy resources to provide a high quality XR experience to user of the XR application. Remote rendering provides an offloading paradigm of rendering computer generated graphics wherein at least part of the rendering operations are offloaded from a rendering client in a communication device, or display connected to the communication device, over a network to a remote rendering server. The rendering client receives rendered frames from the rendering server, possibly applies some correction, for example, pose correction, on the rendered frame and displays the frame. In split rendering is a realization of the offloading paradigm wherein a split rendering client performs some rendering operations in addition to simplistic pose correction, for example, the split rendering client may render some graphics objects locally. Split rendering (SR) may be a potential solution to the issue of providing high quality graphics for resource constrained devices since, in split rendering, the rendering workload is divided between two or more devices, for example, a split rendering client (SRC) and a split rendering server (SRS). In split-rendering scenarios, the SRC regularly sends metadata to the SRS which the SRS uses to render video frames of a scene. The SRS encodes the rendered video frames and then sends the rendered media to the SRC which decodes and displays it. The split rendering paradigm may allow high quality and resource intensive immersive experiences to be delivered on resource constrained devices such as mobile devices and standalone Head Mounted Displays (HMDs). Applications using split rendering services are envisioned to be available in 5G networks and beyond 5G networks. This may be achieved by leveraging the service-based architecture of 5G (and beyond 5G networks), by exposing network resources and functionality via standardized interfaces to authorized entities. This allows a network operator to either deploy an XR service itself (in which case the DN may be inside the 5G network) or expose resources to an authorized external application service provider (in which case the DN may be outside the 5G network, but still uses 5G network functionalities) The following functions are introduced in TS 26.565, a Split-Rendering Client (SRC), a Split-Rendering Sever (SRS), Application Function (AF), Application Service Provider, Application and Media Session Handler (MSH). The SRC is responsible for acquiring the UE media capabilities and negotiates with the RTC AS to agree on the split-rendering process at the RTC AS. The SRS is responsible for negotiation of SR session with SRC, monitoring the server’s edge resource usage, and managing / running the split rendering process. The AF is responsible for provisioning of network resources, QoS allocation, and edge resource discovery. The Application Service Provider (AP) is the application provider that offers the XR service. The Applications is the application running on a UE. The MSH is the entity on UE that is responsible for the control plane communication with the AF. Two examples of 5G networks that implement split rendering are referred to as Split Rendering Media Service Enabler (SR_MSE) and IMS based conversational services (IBAGS). An example a 5G network with a SR_MSE architecture is illustrated in Figure 4. In the 5G network with a SR_MSE architecture, 5G media function residing in a 5G edge server is a SRS . The 5G media function is responsible for negotiation of parameters and configuration of a split rendering session for establishment of the split rendering (SR) session with a SRC, monitoring usage of resources of the 5G edge server on which the 5G media function resides and managing and / running a split rendering process. The Split-Rendering Server resides in a DN. Although the DN is illustrated as an external DN, it may be external or part of the 5GS. In the example of a 5G network with a SR_MSE architecture shown in Figure 4, a 5G Application Service Provider or 5G Application Provider (AP) provisions resources for splitrendering through RTC-1. The 5G AP may be an entity providing applications to communication devices (e.g., UEs) which use split rendering, over a 5GS. The 5G AP may be authorised to use resources and functionalities provided by the 5G network. Provisioning may include the 5G AP requesting the 5G network to allocate appropriate resources and Quality of Service profiles for a split rendering session. The resources, may for example, include means to carry out compute and render operations for a split rendering session. The 5G AP may deliver (e.g., provide) media to a SRS through RTC-2. The communication between RTC AF and SRS is through RTC-3. RTC-3 is an interface that may for example include the EDGE-3 interface (as described in clause 6.5.7 of 3GPP TS 23.558 V19.0 ). Signaling (where signalling refers to user plane signalling, for example for WebRTC session setup or reconfiguration) and media delivery between SRC and SRS is though RTC-4. The RTC AF may provide the split-rendering information to the Media Session Handler defined by RTC-5. SRC discovers the application through RTC-6. The SRC handles the XR runtime. The SRC discovers the client media capabilities through the RTC-7 interface. The 5G Application and AP interact through RTC-8-8. An example communication network that has a SR_IBAS architecture is illustrated in Figure 5. In the communication network shown in Figure 5, the AR Media Function (ARMF) is an SRS and the ARMF is an IMS based Media Function (MF) or a Media Resource Function (MRF) that has AR media processing capabilities. In the example communication network shown in Figure 5, the functions of the MF / MRF (e.g., the entity comprising the MF / MRF) support AR conversational service by providing transcoding for terminals with limited capabilities. Additionally, the MF and / or MRF may collect spatial and media descriptions from communication devices (e.g., UEs) and generate scene descriptions for symmetrical AR call experiences, provide remote rendering for AR-MTSI clients in communication devices (e.g., terminals) with limited capabilities based on rendering negotiation. For remote rendering the AR-MTSI client provides AR metadata, e.g., pose data. This in practice defines the MF / MRF as a split rendering server and the AR-MTSI client as a split rendering client. In both communication networks shown in Figure 4 and Figure 5, Split Rendering is performed by an SRC of a UE and an SRS. The SRS may reside inside (e.g., may be deployed in) a Data Network (DN) which may be an edge data network as described in 3GPP TS 23.558. The DN may be part of the 5G network or it may be an external DN configured to access resources of a 5G network using appropriate means, for example those defined in clauses 4, 6 and 10 of TS 26.113 or clauses 5,8,9 and 10 of TS 26.510 . Conceptually, a SRC is an entity intended to deliver an immersive experience to a user of a UE or a device associated with a UE, while a SRS is an entity that supports the delivery of the immersive experience by performing at least some of the rendering operations of the immersive experience offloaded by the SRC. There may be multiple modalities of metadata sent from the SRC to the SRS. Modalities of metadata include pose data which is used by SRS to render frames, which may be reported for one or two views, state data (such as input action or object state) and other tracking data such as controller pose and trackable pose and gaze information. An XR head mounted display (HMD) may include an eye tracker. For example, an XR HMD may comprise eye tracking hardware which may provide an XR application of the XR HMD real time gaze information of a user of the XR HUD. The gaze information (or gaze data) of a user of an XR HUD may be used for rendering optimizations, for example, spatially vary rendering resource allocation based on gaze information and user interaction. Gaze information may be queried from XR run times using vendor and platform specific Application Programming Interfaces (APIs). These APIs may be software applications which exposes the functions and capabilities of hardware to other software applications . In split rendering examples, where gaze data is available at the SRC, it may be sent by the SRC to SRS as part of the rendering meta-data for rendering and encoding optimizations in addition to user interaction. Gaze data may comprise a gaze location and in some cases a confidence information indicating precision of the gaze location. If the use of eye gaze tracking is activated, the SRS may use this gaze location and, if available, confidence information to perform foveated rendering. In foveated rendering, areas of an image are rendered with a higher quality than other areas, to match the user’s current gaze. Alternatively, or in addition, if the use of eye gaze tracking is activated, the SRS may use this gaze location and, if available, confidence information to perform foveated encoding. In foveated encoding, areas of an image are encoded with a higher SNR quality than other areas, to match the user’s current gaze. Foveation refers to the non-uniform nature of human vision due to the physiology of the eye. Given a stimulus, such as a video frame, the human eye perceives the frame with the highest acuity where our gaze is fixated within the frame, while the acuity drops rapidly with distance from the gaze fixation. Figure 6 illustrates this phenomenon in terms of relative visual acuity as a function of angular eccentricity from the fovea-lens axis which in practical terms may be called a gaze vector. Foveated video encoding and foveated rendering take advantage of this non-uniformity of human vision to differentially allocate encoding and rendering resources to different areas of a frame being encoded or rendered, respectively. With gaze predictions, the SRS may produce or generate an importance map for the picture based on the confidence values associated with the gaze predictions. Additionally, the SRS may also use other information to produce or generate the importance map, such as content Regions of Interest. This importance map is passed to the encoder to properly allocate bits for the encoding of the picture. Foveated rendering and foveated encoding differentially allocate rendering and encoding resources to different areas in a frame based on where the viewer’s gaze is. The target being that resource allocation is such that on the final displayed frame, the area on and around the viewer’s gaze has high visual quality while the areas further away from the gaze have lower quality. Conceptually, a frame obtained by foveated rendering or from decoding a foveated encoded frame may comprise a high quality region centered around the user gaze, and a low quality region away from the user gaze, there may be other regions with intermediate qualities. Such regions may be located in between the high quality region and the low quality region to actuate a stepped quality drop off. The high quality region may be referred to as the foveal region. An example frame created by foveated rendering and / or foveated encoding and decoding may have the quality profile shown in Figure 7. In some embodiments, the intermediate quality may comprise of sub regions with quality difference between the sub-regions. In some embodiments, the high, intermediate and low quality regions may be further divided into sub regions which are not contiguous with each other on the frame. Systems using foveated rendering / encoding may be very sensitive to latency due to gaze tracking. High latency may lead to a mismatch between where a user is looking and the pose from which a frame is rendered. Several parameters may affect the latency and latency requirements of a foveated system, including, content characteristics of the media being rendered, logic of the underlying experience or application being executed, the degree of degradation / foveation applied to the image, foveation rendering / encoding technique, frame rate of the media application, sampling rate of the gaze tracker, reporting rate of the gaze and / or the size of the full-resolution image. Foveated rendering and encoding can be implemented in different fashions, with the underlying principle being spatially varying quality of a rendered frame or an encoded frame based on gaze location. One technique for foveated rendering is to vary the sampling rate during the pixel shading operation of rendering, so that samples at or around the gaze (i.e. the foveal region) are shaded normally in a one sample to one or more pixel fashion, while as the samples in the intermediate and low quality regions are shaded in a many sample to one pixel fashion, the number of samples contributing to a pixel being highest in the peripheral region. Another technique for foveated rendering is to vary the level of detail of a scene dynamically with gaze location. Yet another technique may be using more than one virtual camera: one high quality / resolution one for the high quality region desired around the gaze point and one or more progressively lower quality or resolution cameras for the intermediate and low quality regions desired for regions away from the gaze. A technique for foveated rendering as well as foveated encoding is to spatially vary the pixel resolution of a frame. For example, by reducing the image resolution towards the peripheries of a frame away from the user’s gaze (i.e. low and intermediate quality regions), while keeping the areas under and around the user’s (i.e. high quality region / foveal region) gaze at their normal resolution. When displaying the frame, the foveal region is displayed as it is and up- sampling is applied to the low-and intermediate quality parts. In case of foveated rendering, the spatial variation of pixel resolution may be achieved by rendering progressively fewer pixels in the periphery away from the gaze, while rendering all the pixels in the desired high quality region (in and around gaze) Another technique for foveated encoding is to spatially vary the Quantization Parameter (QP) applied to macroblocks in the frame during encoding, typically as a delta QP added on top of the QPs calculated by the encoder- State of the art work considers either gaze based rendering optimizations for local rendering or gaze based video encoding optimizations for remote rendering. While simultaneous gaze based optimizations for both rendering and video encoding for remote rendering have been studied in academia, these studies do not consider split-rendering architectures in wireless networks, and consequently do not account for changes in operating conditions of the SRS, SRC and variations in channel conditions. Figure 8 shows a flowchart of a method performed by a SRC according to an example embodiment. The SRC is a logical entity and an apparatus may comprise or implement the SRC. For example, an apparatus may comprise at least one processor and at least one memory that stores instructions of the SRC, and when the instructions of the SRC are executed by the at least one processor, cause the apparatus to perform the operations of the method shown in FIG. 8. The apparatus may comprise a communication device, such as a smart phone or a HMD. In 801, the method comprises sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client. In 802, the method comprises sending, to the split rendering server, a request adaptation of a gaze based optimization profile, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile. In 803, the method comprises receiving, from the split rendering server, an encoded video frame. In 804, the method comprises decoding the encoded video frames to generate a decoded video frame. In 805, the method comprises displaying the decoded video frame. Figure 9 shows a flowchart of a method performed by a SRS according to an example embodiment. The apparatus may comprise, implement or be a SRS. In 901, the method comprises receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client. In 902, the method comprises performing rendering using the metadata and a gaze based optimization profile for a split rendering session. In 903, the method comprises encoding the video frame using the gaze based optimization profile to generate an encoded video frame. In 902, the method comprises sending, to the split rendering client, the encoded video frame.. Methods as described with reference to Figures 8 and 9 use gaze data (or gaze information), which may be reported by the SRC to the SRS in the form of metadata to reduce the usage of rendering resources at the SRS and reduce the size of encoded video stream of the frames rendered at the SRS. This may maintain the QoE of an end user while adapting to changing conditions at the SRS, SRC and network. Metadata may include information needed to render and encode a frame with gaze based optimizations, for example pose information, gaze information and timing information. The metadata is for use with a virtual camera associated with the split rendering client. The gaze based optimization profile may comprise a first spatial quality map used for performing rendering the video frame and a second spatial quality map used for encoding the video frame. A first spatial quality map for rendering the video frame may indicate a spatially varying quality of a rendered frame based on gaze information. A second spatial quality map for encoding a video frame may indicate spatially varying the quality of an encoded video frame based on gaze information. A method as described with reference to Figure 9 may comprise performing a negotiation of the gaze based optimisation profile or of a set of gaze based optimization profiles for the split rendering session with the split rendering client. In an example embodiment, the SRC and SRS negotiate to agree on a gaze based optimization profile to be used during the split rendering session. In another example embodiment, the SRC and SRS negotiate to agree on a set of gaze based optimization profiles available for use during the split rendering session. A gaze based optimization profile may comprise at least one of the following: the foveated rendering method used, the initial parameters of a first spatial quality map (to be used with the foveated rendering method), the foveated encoding method used, the initial parameters of a second spatial quality map (to be used with the foveated encoding method), the initial parameters of a third spatial quality map or the parameters and factors to be used for adaptation of the first spatial quality map and the second spatial quality map, and if used, the third spatial quality map. The parameters of a spatial quality map comprise at least one of the following: size, shape and of the high quality region, size, shape and relative position of intermediate quality region(s), size, shape and relative position of the low quality region or quality desired for the high, intermediate and low quality regions. The desired quality may be indicated in terms relevant to the techniques of foveated rendering and encoding used. The negotiation of the gaze based optimization profile with the split rendering client may be performed based on at least one of the following: the capabilities of an apparatus comprising the split rendering client, the application being rendered in the split rendering session or user input detected by the apparatus. User input may comprise a user selecting an option from a list of profiles In an example embodiment, during the negotiation, the SRC may propose or agree to a gaze based optimization profile or a set thereof based on, for example, UE capabilities, application being rendered in the SR session, user choice and such. In an embodiment, the negotiated gaze based profile or set of gaze based optimization profiles may be indicated in the Split Rendering Configuration, for example, as defined in 3GPP standard TS 26.565, in a congruently named extra configuration parameter, for example, as shown in Table 1. gazeOptProfile Object 1..N A object corresponding to a gaze based profile. It may be only an identifier such as a URI / N or it may comprise all information needed to use the profile Table 1 In an example embodiment, gaze-based optimizations are enabled by adding new SR_CONFIG_FLAGs in the Split Rendering Configuration, as defined in TS 26.565. For example, the flags FU\G_FOVEATED_RENDERING and FLAG_FOVEATED_ENCODING may be defined forfoveated rendering and foveated encoding, respectively. These flags may be set depending on the result of the negotiation between SRC and SRS. Depending on the setting of these flags, SRS may generate the associated quality maps and encode / render the video accordingly. In an example embodiment, foveated rendering is achieved by variable rate shading, for example, the shading rate of the foveal, intermediate and low is set to 1 sample per pixel, 2x2 samples per pixel and 4x4 samples per pixel respectively. In an example embodiment, SRC and SRS negotiate a scaling factor for progressively reducing the shading rate of the intermediate and low quality parts of the image. In an example embodiment, foveated rendering is achieved by spatially varying the resolution of the frame by rendering fewer pixels away from the gaze point and more pixels closer to the gaze point. In an example of the above embodiment, SRC and SRS negotiate a scaling factor for progressively downscaling the resolution of the intermediate and low quality parts of the image. For example, with a scaling factor of 2, the intermediate and low quality areas can be down-sampled to 1 / 2x, 1 / 4x, 1 / 8x, etc. of the full the resolution image. In an embodiment, foveated encoding may be achieved by spatially varying the resolution of the frame during the encoding process. In an example of the above embodiment, SRC and SRS negotiate a scaling factor for progressively downscaling the resolution of the intermediate and low quality parts of the image. For example, with a scaling factor of 2, the intermediate and low quality areas can be down-sampled to 1 / 2x, 1 / 4x, 1 / 8x, etc. of the full the resolution image. In an example of the above embodiment, SRC and SRS may negotiate the scheme of spatially varying the resolution, i.e., whether the downscaling is performed using a linear, logarithmic or another function (e.g. based on a visual acuity fall-off model). In an example embodiment, foveated encoding is achieved by spatially varying QP applied during encoding a frame. In an example embodiment, SRC and SRS negotiate a QP increase rate for foveated encoding, progressively reducing the visual fidelity of the foveated part of the image. For example, with a QP increase rate of 5 and assuming that the high quality region (corresponding to a gaze of a user of a communication device (e.g., UE) comprising the SRC) is encoded with QP=22, SRS encodes two intermediate quality regions with the QP values 27, 32, and the low quality region with QP value 37, where QP values increase towards the boundaries of the image. SRS may have multiple foveation techniques or other gaze optimization techniques at its disposal, e.g., it may be able to apply different foveated rendering techniques as well as Level of Detail (LoD) rendering or use LoD rendering to achieve foveated rendering. In an embodiment, SRC and SRS negotiate whether one or more gaze optimization techniques will be used by SRS. In some embodiments Gaze optimization mode for foveated rendering or foveated encoding may be added as a parameter to the Split Rendering Configuration e.g. as defined in TS 26.565. An example shown is below, where the listed optimization modes are foveated rendering with linear and human visual system (HVS) optimized function, respectively, and LoD rendering. gazeOptMode enum 1..N The type indicates the gaze optimization mode configuration. Defined values are FOVEATED_RENDERING / ENCO DINGJJNEAR and FOVEATED_RENDERING / ENCO DING_HVS, FOVEATED_RENDERING / ENCO DING LOD. Other values may be added. A request for adaptation of the gaze based optimization profile may be sent with the metadata relating to the video frame for use in combination with the gaze based optimization profile for rendering and encoding the video frame. In an embodiment an adaptation message may comprises an indication to adapt a gaze based optimization profile, requests for, or acknowledgements of adaptation of a gaze based optimization profile. An example adaptation message from the SRC to SRS may be of three types, an indication to adapt a gaze based optimization profile, e.g., by reporting metric(s) to SRS such that the SRS determines to adapt the gaze based optimisation profile, a request to modify the parameters of a gaze optimization profile or it is a request to switch to a different gaze optimization profile. In some example embodiments, meta data messages comprising pose and timing information further comprise adaptation messages. In some example embodiments the requests for adaptation are sent asynchronously to the frame rendering loop. Alternatively, or in addition, a rendering server may determine to determine to adapt the profile for at least one of rendering or encoding the video frame based on gaze information of the viewer of the video frame. Determining to adapt the profile may comprise receiving a request to adapt the profile from a rendering client and adapting the profile based on the request. The indication of an adaptation of a gaze based optimization profile may comprise an indication to change one or more parameters of the gaze based optimization profile or an indication of a change to one or more parameters of the gaze optimization profile. A method as described with reference to Figure 8 may comprise determining to request an adaptation of the gaze based optimization profile based on at least one of the following: operating conditions of the apparatus, network conditions detected by the apparatus, an estimated quality of experience of the split rendering session, an estimated quality of service of the split rendering session or a user input detected by the apparatus. The user input may indicate to the SRC to change the profile and / or the user may be able to selected a gaze based optimization profile to use. For example, the requests for adaptation of a gaze based optimization profile may be based on operating conditions of the UE. As an example, operating conditions may include battery level, temperature or computational resources available for SR operation. Network conditions may include, for example, channel capacity, network rtt, motion to render to photon latency, motion to high quality latency, jitter or packet inter-arrival time. An estimate of user QoE or service QoS may be based on, for example, network conditions, operating conditions, and state of the SR media session. The state of the SR media session may comprise the state of the receiving or decoding buffer (capacity, fill rate, etc.), dropped frames (number, proportion, frequency etc.), corrupted frames (number, proportion, frequency etc.), video stalls (number, duration, etc.) or a no reference video quality metric. The requests for adaptation of a gaze based optimization profile may be based on user input, for example, a user may choose a larger high quality region, or a higher quality for the peripheral region or vice versa. The adaptation of a gaze based optimization profile may comprise at least one of the following: changing the method used for foveated rendering or foveated encoding or changing the parameters of the method used for foveated rendering or foveated encoding (e.g. size / quality / shape / relative shape of the different quality regions etc.). The indication to change one or more parameters of the gaze based optimization profile may comprise an indication to change a method of rendering or an indication to change a method of encoding. The indication of the adaptation to the gaze based optimization profile comprises an indication to change the gaze based optimization profile to another gaze based optimization profile selected from a plurality of gaze based optimization profiles negotiated for the split rendering session The computations related to the above tasks may be carried out by the SRC continuously, for example, considering a sliding time window or periodically with a defined frequency or these may be triggered by an event, such as a parameter related to UE operation, network conditions or media session comprised in the SR changing beyond a threshold. Figure 10a illustrates an example quality map 1000 for a video frame which divides the video frame into four concentric regions with different qualities based on gaze 1001 of a viewer of the frame. A high quality region 1002 is centred on the gaze point 1001, a first intermediate quality region 1003 covers an annular area around the high quality region 1002. A second intermediate quality region 1004 covers another annular area around the first intermediate quality region 1003, the rest of the frame is covered by a low quality region 1005. The area and quality of the high quality region 1002, the first intermediate quality region 1003, the second quality region 1004, and the low quality region 1005 may be configurable and may comprise parameters of a gaze based optimization profile. In some embodiments, between the four quality regions, the quality of the first intermediate quality region 1003 may be higher than the quality of the second intermediate quality region 1004, the high quality region 1002 may have the highest quality and the low quality region 1005 may have the lowest quality. Figure 10b illustrates an example quality map 1010 for a video frame which divides the video frame into three concentric regions with different qualities. The quality map 1010 includes a high quality region 1012 that is centred on a gaze point 1011, an intermediate quality concentric region 1013 that covers an annular area around the high quality region 1012 while the rest of the frame is covered by a low quality region 1014. The area and quality of the high quality region 1012, the intermediate quality region 1013, and the low quality region 1014 may be configurable and may comprise parameters of a gaze based optimization profile. In some embodiments, between the three quality regions, the high quality region 1012 may have the highest quality and the low quality region 1014 may have the lowest quality, while the intermediate quality 1013 region may have a quality in between that of the high quality region 1012 and the low quality region 1014. In an example embodiment, the first spatial quality map is generated by the SRS based on the metadata received from the SRC and on at least one of the following: o a latency metric e.g. network round trip time (rtt), motion to render to photon latency which may refer to the delay between the time pose of a user is registered by a UE to the time a corresponding frame is displayed to the user. Other examples of latency metrics on which the first spatial quality map may be based on include those defined in 3GPP TR 26.812. o a network condition metric. Network condition metrics may include, for example, a network capacity metric such estimated bandwidth, interarrival time of packets, packet jitter which may be based on variations in interarrival time of packets, network congestion status, packet loss percentage, etc. o a Quality of Experience (QoE) metric which may take into account various factors: ■ related to the content of an application being rendered by the SRS e.g. graphics quality, application logic (game play), dynamicity of the content. For XR applications, the nature of real and virtual content may also be accounted for. ■ factors related to the user of the UE in which the SRC is deployed, for example, their eye sight, their video quality preferences etc, as registered by the SR application or SRC in the UE (e.g. via user input or user behaviour history) o a Quality of Service (QoS) policy, for example, an application provider may aim to provide a particular level of quality of service for a split rendering session by provisioning appropriate resources in the 5G network, for example, using procedures defined in TS 26.510. o a predicted gaze or pose, an SRC may report metadata that includes a gaze value or pose value to the SRS for an estimated display time. At the time when meta-data is reported to the SRS, the SRC or the UE comprising the SRC may not be able to accurately specify a display time for a frame expected to be rendered on the basis of the meta-data, hence a display time may be estimated by the SRC or the UE comprising the SRC or another entity comprised in the SRC. The estimated display time may be in future at the time when the metadata is reported to the SRS, hence the SRC may predict gaze and pose values. The SRS may perform further predictions of gaze and pose values for use in generating the first quality map o a model of the human visual system (including possible visual impairments), for example, a visual acuity model shown in Figure 6. o optical characteristics of the viewing HMD, the optical characteristics, such as its lens shape, construction and power, field of view and such may affect how a video frame displayed on a display screen of the HMD appears to a viewer. In some cases some regions of the displayed video frame may be appear to be of lower quality while others may appear to be higher quality due to, for example, so called lens distortion. The optical characteristics of the viewing HMD may be advantageously used during rendering to save rendering, network or energy resources. o Rendering resource consumption, the shape and size of different quality regions of a first quality map may be adjusted to optimize consumption of resources used for rendering a frame. For example, if it is desired, a reduction in the consumption of rendering resources may be achieved by reducing the size of regions rendered with high quality in a frame, o a confidence value of the received gaze data, a high confidence value in the accuracy of received gaze data may allow the region of high quality cantered around gaze to be made smaller and vice versa. o content of the scene being rendered, the first quality map may take into account the characteristics of the content being rendered, for example, regions with high levels of visual saliency may be rendered with high quality, regardless of the gaze location in a frame. o application logic and other application information, e.g., controls, game physics, game logic, player actions, game state etc., or o a previously rendered and / or encoded frame, the first quality map may take into account the previous frames to ensure, for example, that the SRC is able to decode the current frame, In an example embodiment, the second spatial quality map is based on the received gaze data and on at least one of the following: o the first spatial quality map. To optimize usage of rendering resources and bit budget for encoding, the second quality map may be based, in part, on the first quality map, so that the encoding quality corresponds with the first quality map. In an example, the second quality map may be identical to the first quality map. In another example, the second quality map may be such that it ameliorates visual artefacts that may arise due rendering using the first quality map. For example, if in the first quality map, a frame is divided into concentric circles of different qualities, with the stepped or discrete quality variation between two adjacent regions, the second quality map may have a smooth transition of quality between different regions, which may ensure that artefacts introduced in the rendering stage are at least partially ameliorated by a filtering effect of encoding using the second quality map. o a latency metric e.g. network round trip time (rtt), motion to render to photon latency which may refer to the delay between the time pose of a user is registered by a UE to the time a corresponding frame is displayed to the user. Other examples of latency metrics on which the first spatial quality map may be based on include those defined in 3GPP TR 26.812. o a network condition metric. Network condition metrics may include, for example, a network capacity metric such estimated bandwidth, , interarrival time of packets, packet jitter which may be based on variations in interarrival time of packets, network congestion status, packet loss percentage, etc. o a Quality of Experience (QoE) metric which may take into account various factors: ■ related to the content of an application being rendered by the SRS e.g. graphics quality, application logic (game play), dynamicity of the content. For XR applications, the nature of real and virtual content may also be accounted for. ■ factors related to the user of the UE in which the SRC is deployed, for example, their eye sight, their video quality preferences etc, as registered by the SR application or SRC in the UE (e.g. via user input or user behaviour history) o a Quality of Service (QoS) policy, for example, an application provider may provide a particular level of quality of service for a split rendering session by provisioning appropriate resources in the 5G network, for example, using procedures defined in TS 26.510. o a predicted gaze or pose, an SRC may report metadata comprising a gaze value or pose value associated with a user of the UE comprising the SRC or a user of an HMD associated with the UE comprising the SRC to the SRS for an estimated display time. The estimated display time may be in future at the time when the meta-data is reported, hence the SRC may predict these gaze and pose values. The SRS may perform further predictions of gaze and pose values for use in generating the first quality map o a model of the human visual system (including possible visual impairments), for example, a visual acuity model shown in Figure 6. o optical characteristics of the viewing HMD, the optical characteristics, such as its lens shape, construction and power, field of view and such may affect how a video frame displayed on a display screen of the HMD appears to a viewer. In some cases some regions of the displayed video frame may be appear to be of lower quality while others may appear to be higher quality due to, for example, so called lens distortion These characteristics may be advantageously used during rendering to save resources. o Target bitrate, the shape and size of different quality regions of a first quality map may be adjusted to optimize consumption of bit budget allocated for encoding a rendered frame. For example, if it is desired, a reduction in the bitrate may be achieved by reducing the size of regions encoded with high quality in a frame, o a confidence value of the received gaze data, a high confidence value in the accuracy of received gaze data may allow the region of high quality centered around gaze to be made smaller and vice versa. o content of the scene being rendered, the first quality map may take into account the characteristics of the content being rendered, for example, regions with high levels of visual saliency may be encoded with high quality, regardless of the gaze location in a frame. o application logic and other application information, e.g., controls, game physics, game logic, player actions, game state etc., or o a previously rendered and / or encoded frame, the second quality map may take into account the previous frames to ensure, for example, that the SRC is able to decode the current frame, o encoding parameters, like codec, codec profile, rate control mode, GOP size, frame types etc These factors affect the quality of an encoded frame and may also affect the computational and time resources needed for encoding a video frame by the SRS and eventual decoding of the video frame by the SRC. The second quality map may for example define the quality levels of different regions of a frame based on these factors In some example embodiments, the first spatial quality map comprises a volumetric map comprising different levels of detail. In various example embodiments, the size of the different quality regions may be based on motion to render to photon latency, for example, the size of the high quality region may be increased if the motion to render to photon latency is high and vice versa. In an example embodiment, the first spatial quality map and / or the second spatial quality map are set to a uniform quality across the frame for a frame which is to be encoded as a reference frame. In another example embodiment, when the encoding uses the first spatial quality map and / or the second spatial quality map set areas corresponding to a prospective intra-coded region, e.g. a GDR slice, to high quality. In some example embodiments, the spatial quality maps may comprise a smooth / continuous quality fall based on distance from the gaze location. In an embodiment, the second spatial quality map includes delta quantizer information. This can be used by the rate control to locally increase, maintain, or decrease the quantization strength and directly impacts the quality. Alternatively, the second spatial quality map can signal penalties or bonuses to be applied to the Lagrangian multiplier lambda as it has more impact on the overall coding decision process. Alternatively, a weighting map can be used to locally modulate or to override the local bitrate allocation achieved by the rate control. In an embodiment, the second spatial quality map indicates what are the areas to be preserved, if not located in the foveated area. This can for example be useful for preserving head-up display information that people are likely to frequently look at (e.g. game minimap or interface). This can be also useful in monetization scenarios where advertisements are incorporated in virtual spaces, and where the advertising control wants a good and clear visibility. This can be done by transmitting similar maps are described in the previous embodiment, or a binary map can be used to highlight the important areas. In an example embodiment, the second spatial quality map may also indicate artifacts masking information. During foveated rendering and encoding, quality of peripheral vision is likely to be lowered, creating visual artifacts (e.g. blocking, ringing). Post-processing, including filtering and noise synthesis, can be used to mask those artifacts to the viewer, and enhance the quality of experience for the viewer. The second spatial quality map can indicate the required masking strength, e.g. in a range of 0 to 1, that would be use by the encoder to generate postprocessing SEI messages telling the receiver what processing to apply, and where. A film grain synthesis SEI message can be used. A method as described with reference to Figure 9 may comprise providing information relating to the profile or an adapted profile to the SRC from the SRS. Information relating to the profile may comprise post-processing information, for example, artefact masking information. In some example embodiments, information relating to the gaze based optimization profile may be sent from the SRS to the SRC. Information relating to the gaze based optimization profile may comprise rendered pose, timing information and information comprising one or more of the following: gaze location used for gaze based optimizations, gaze based optimization profile used for the frame, parameter values of the gaze based optimization profile used, indication of or information about artefact masking comprised in the second spatial quality map or other information corresponding to the first, second or third spatial quality map. The SRC may process the received stream of video frames based on the received information relating to the profile, adapted profile or rendered video frame. In some example embodiments, artefact masking information may be indicated to the SRC during session setup and / or during the session. In some embodiments, the first and the second spatial quality map are identical. In some embodiments, the second spatial quality map takes into account the visual effects of the first spatial quality map on the rendered frame. In another example embodiment, the third spatial quality map can indicate the preferred frame partitioning in terms of tiles / slices / subpictures, in order to increase the robustness of the transmission in some areas. In some example embodiments, the third spatial quality map indicates importance of the different tiles, subpictures, or / and slices comprised in the frame , e.g., tiles, subpictures or slices corresponding to high quality regions may be assigned the highest importance. This information may be shared with a network node for intelligent network traffic management. In an example of the above embodiment, the encoded frames are transmitted over RTP comprising header extensions wherein the header extension comprises the relative importance information indicated in the third map. In an example of the above embodiments can be then used by the RTP packager to set the value of the PSI field in the RTP header extension. The SRS may provide comprise adaptation messages to the SRS. The adaptation messages may comprise requests for, or acknowledgements of adaptation of a gaze based optimization profile. Adaptation messages from the SRS to SRC may comprise an indication or an acknowledgement of an adaptation or information related to the adaptation or information relating to the adapted gaze based optimization profile. If the SRC is performing postprocessing, as described below, postprocessing may be based on the information relating to the adapted gaze based optimization profile from the SRS to SRC. In some example embodiments, the adaptation messages may be sent together with the meta-data. In some embodiments, the adaptation messages may be sent asynchronously to the metadata. In some embodiments, the meta-data and / or adaptation messages are sent together (in-band, e.g. in an RTP header extension) with the rendered frame. In some example embodiments, the SRC may apply post processing operation(s) to the decoded frame. In example embodiments where foveated rendering or encoding is achieved by varying the pixel resolution of a frame, a post processing operation may comprise a spatially varying up-sampling operation (corresponding to the scaling factor and the scheme thereof used at the SRS and agreed to during the session negotiation). In some example embodiments, where the second spatial quality map indicates artefact masking information, the post processing step comprises applying artifact masking based on the artefact masking information comprised in the second quality map and indicated during session setup or in meta-data received from the SRS. In some example embodiments, the post processing step may comprise applying neural network based techniques, such as super-resolution. In some embodiments, the post-processing step may further take into account at least one of the following: visual content in the frame, characteristics of the application being rendered in the SR session, pose or gaze patterns of the user, a no-reference image or video quality metric, for example, to detect or estimate, blocking artefacts, temporal flicker etc., gaze (and / or pose error) between the gaze location used for gaze based optimization (and / or pose used rendering the frame) and a gaze location (and / or pose) estimated for a display time for the current frame. In some example embodiments, the adaptation messages are sent over a WebRTC data channel conformant to IETF RFC8831. In some example embodiments, the adaptation messages are sent as RTCP application layer feedback messages. In some example embodiments, the adaptation messages are sent in RTP header extension that conforms with RFC 8285 Figure 11 illustrates an example a split rendering procedure which uses aspects of gaze based rendering and encoding optimisations as described. At step 1 of the split rendering procedure, the SRS and SRC negotiate the parameters and configuration of a split rendering session and set up the split rendering session. At step 1, the SRC and SRS may agree on one or more gaze based optimization profiles to use for the split rendering session. After step 1, a rendering loop of the split rendering procedure is performed. A rendering loop may comprise a set of operations, some of which may repeat in a sequence, which result in new video frames being displayed or caused to be displayed by the SRC. In a split rendering session, at least a subset of the set of operations comprising a rendering loop may be executed by the SRS or the SRC. At step 2a, the SRC sends to the SRS metadata for a virtual camera, associated with an apparatus comprising the SRC (e.g., a UE or a HUD associated with the UE). The metadata may include pose information of a virtual camera associated with the UE, gaze information of a user, eye status of a user, and time information associated with a user of the apparatus that includes the SRC. In some embodiments, after step 2a, the split rendering procedure proceeds to step 2b. At step 2b, the SRC sends an adaptation message related to a gaze optimization profile. The adaptation message may be sent in synchronisation with the metadata sent as step 2a or the adaptation message may be sent asynchronously to the metadata sent at step 2a. The adaptation message is an example of an indication or request to adapt the profile for adaptation of a gaze based profile for a split rendering session. Step 2b may not be performed per video frame. At step 3, the SRS generates a plurality of quality maps. During step 3, the SRS generates a first spatial quality map at step 3a that is to be used for rendering a frame such that the rendering resources are differentially allocated to render different areas of the frame. Also during step 3,the SRS generates a second spatial quality map that is to be used for encoding a frame that is rendered such that the encoding bit budget is differentially allocated to different areas of the frame. An encoding bit budget may refer to storage size, for example, in bits or Megabits, for an output encoded frame that the video encoder targets when encoding an unencoded frame obtained, for example, as output of a rendering operation. After step 3, the split rendering procedure proceeds to step 4a. At step 4a, the SRS renders a frame according to the first quality map and encodes the frame that is rendered according to the second quality map. This is an example of performing rendering using the metadata and a gaze based optimization profile for a split rendering session and encoding the video frame using the gaze based optimization profile to generate an encoded video frame. At optional step 4b, the SRS generates a third quality map, that is to used to identify transmission priorities with which different tiles or slices or subpictures comprising the encoded video frame are sent. Transmission priorities may include information to the network indicating the importance of data packets comprising the encoded frame. The importance of a data packet may correspond to the transmission priority of the tile or slice or subpicture comprised in the data packet. Transmission priority may also indicate the order in which data packets corresponding to an encoded video frame are sent. At step 5a, the encoded frame and possible associated metadata is sent to the SRC. This is an example of sending the encoded video frame to the split rendering client. At step 5b, the SRS may send an adaption message related to the gaze optimization profile, either in sync with 5a or asynchronously to 5a. This is an example of sending information relating to the adapted gaze based optimization profile to a SRC. In some embodiments, step 5b may not be performed for each frame. At step 6, the SRC decodes the received encoded frame, and composes a frame for display using the decoded frame, and causes the composed frame to be displayed . The term “frame” and “video frame” are used interchangeably in the above description and have the same meaning. A video frame refers to an image from a sequence of images which when displayed successively creates the “moving picture”. Video frames can be raw, i.e., not encoded or compressed or they can be compressed in size by encoding. A video frame rendered by the SRS would be a raw video frame which is encoded by the SRS using an encoder before transmitting it to the SRC. In an example embodiment an apparatus is provided for carrying out the methods of Figures 8 and 9. In one example embodiment, an apparatus comprising a split rendering client may comprise means for sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client, sending, to the split rendering server, a request for adaptation of a gaze based optimization profile for a split rendering session, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile, receiving, from the split rendering server, an encoded video frame, decoding the encoded video frames to generate a decoded video frame and displaying the decoded video frame. The apparatus may comprise a user equipment, such as a mobile phone or HMD, be the user equipment, or be comprised in the user equipment, or a chipset for performing at least some actions of / for the user equipment. The user equipment may be the communication device as illustrated in Figure 2. In one example embodiment, an apparatus comprising a split rendering server may comprise means for receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client, performing rendering using the metadata and a gaze based optimization profile for a split rendering session, encoding the video frame using the gaze based optimization profile to generate an encoded video frame and sending, to the split rendering client, the encoded video frame. The apparatus may comprise a split rendering server as described with reference to Figure 3, be the split rendering server, or be comprised in the split rendering server, or a chipset for performing at least some actions of / for the split rendering server. It should be understood that the apparatuses may comprise or be coupled to other units or modules etc., such as radio parts or radio heads, used in or for transmission and / or reception. Although the apparatuses have been described as one entity, different modules and memory may be implemented in one or more physical or logical entities. It is noted that whilst some embodiments have been described in relation to 5G networks, similar principles can be applied in relation to other networks and communication systems such as 6G networks or 5G-Advanced networks. Therefore, although certain embodiments were described above by way of example with reference to certain example architectures for wireless networks, technologies and standards, embodiments may be applied to any other suitable forms of communication systems than those illustrated and described herein. It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. In general, the various embodiments may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. Some aspects of the disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry may apply to all uses of this term “means” in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “means for” may define at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause an apparatus at least to perform the steps listed after the term. The embodiments of this disclosure may be implemented by computer software comprising instructions which when executed by a processor of a communication device, such as in the processor entity, or by hardware, or by a combination of software and hardware cause the communication device to carry out or perform the operations of the method shown in FIG. 8. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computer-executable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non-transitory media. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may comprise one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), FPGA, gate level circuits and processors based on multi core processor architecture, as non-limiting examples. Embodiments of the disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. The scope of protection sought for various embodiments of the disclosure is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure. The foregoing description has provided by way of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.
Claims
1. An apparatus comprising a split rendering client comprising means for:sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client;sending, to the split rendering server, a request for adaptation of a gaze based optimization profile for a split rendering session, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile;receiving, from the split rendering server, an encoded video frame;decoding the encoded video frames to generate a decoded video frame; and displaying the decoded video frame.
2. The apparatus of claim 1, wherein the indication of an adaptation of a gaze based optimization profile comprises an indication to change one or more parameters of the gaze optimization profile.
3. The apparatus of claim 1, wherein the indication of an adaptation of a gaze based optimization profile comprises an indication of a change to one or more parameters of the gaze optimization profile.
4. The apparatus according to any of claims 1 to 3, wherein the gaze based optimization profile comprises a first spatial quality map used for performing rendering the video frame and a second spatial quality map used for encoding the video frame.
5. The apparatus according to claim 4, wherein the first spatial quality map for rendering the video frame indicates a spatially varying quality of a rendered frame based on gaze information and the second spatial quality map for encoding the video frame indicates a spatially varying quality of an encoded frame based on gaze information.
6. The apparatus according to any of claims 1 to 5, wherein the means is for determining to request adaptation of the gaze based optimization profile based on at least one of the following: operating conditions of the apparatus, network conditions detected by the apparatus, an estimated quality of experience of the split rendering session, an estimated quality of service of the split rendering session or a user input detected by the apparatus.
7. The apparatus according to claim 2, wherein the indication to change one or more parameters of the gaze based optimization profile comprises an indication to change a method of rendering or an indication to change a method of encoding.
8. The apparatus according to any of claims 1 to 7, wherein the indication of the adaptation to the gaze based optimization profile comprises an indication to change the gaze based optimization profile to another gaze based optimization profile selected from a plurality of gaze based optimization profiles negotiated for the split rendering session.
9. The apparatus according to any of claims 1 to 8, wherein the means is further for sending the request for adaptation of the gaze based optimization profile with metadata relating to the video frame for use in combination with the gaze based optimization profile for rendering and encoding the video frame.
10. The apparatus according to any of claims 1 to 9, wherein the means is for receiving, from the split rendering server, information relating to the adapted gaze based optimization profile.
11. The apparatus according to claim 10, wherein the means is for performing further processing of the received video frame based on the received information relating to the adapted gaze based optimization profile.
12. A split rendering server comprising means for:receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client;performing rendering using the metadata and a gaze based optimization profile for a split rendering session;encoding the video frame using the gaze based optimization profile to generate an encoded video frame; andsending, to the split rendering client, the encoded video frame.
13. The apparatus according to claim 12, wherein the means is further for performing a negotiation of the gaze based optimization profile for the split rendering session with the split rendering client.
14. The apparatus according to claim 13, wherein the negotiation is performed based on at least one of the following: the capabilities of an apparatus comprising the split rendering client, the application being rendered in the split rendering session or user input detected by the apparatus.
15. The apparatus according to any of claims 12 to 14, wherein the gaze based optimization profile comprises a first spatial quality map used for performing rendering and a second spatial quality map used for encoding the video frame.
16. The apparatus according to claim 15, wherein the first spatial quality map used for performing rendering the video frame indicates a spatially varying quality of a rendered frame based on gaze information and the second spatial quality map used for encoding the video frame indicates a spatially varying quality of an encoded frame based on gaze information.
17. The apparatus according to any of claims 12 to 16, wherein the means is further for: determining to adapt the gaze based optimization profile for the split rendering session based on the metadata that is received;adapting the gaze based optimization profile;performing rendering to generate a second video frame using the gaze based optimization profile that is adapted; andencoding the second video frame using the gaze based optimization profile that is adapted.
18. The apparatus according to claim 17, wherein determining to adapt the gaze based optimization profile comprises:receiving a request for adaptation of the gaze based optimisation profile for the split rendering session, wherein the request for adaptation of the gaze based optimisation profile comprises an indication of an adaptation to the gaze based optimisation profile; and adapting the gaze based optimisation profile based on the request.
19. The apparatus of claim 18, wherein the indication of an adaptation of a gazebased optimization profile comprises an indication to change one or more parameters of the gaze optimization profile.
20. The apparatus of claim 18, wherein the indication of an adaptation of a gazebased optimization profile comprises an indication of a change to one or more parameters of the gaze optimization profile.
21. The apparatus according to claim 19, wherein the indication to change one or more parameters of the gaze based optimisation profile comprises at least one of a request to change a method of rendering or an indication to change a method of encoding.
22. A method comprising, at a split rendering client:sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client;sending, to the split rendering server, a request for adaptation of a gaze based optimization profile for a split rendering session, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile;receiving, from the split rendering server, an encoded video frame;decoding the encoded video frames to generate a decoded video frame; and displaying the decoded video frame.
23. A method comprising, at a split rendering server:receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client;performing rendering using the metadata and a gaze based optimization profile for a split rendering session;encoding the video frame using the gaze based optimization profile to generate an encoded video frame; andsending, to the split rendering client, the encoded video frame.
24. A computer readable medium comprising instructions which, when executed by an apparatus comprising a split rendering client, cause the apparatus to perform at least the following:sending, to a split rendering server, metadata for a virtual camera associated with the split rendering client;sending, to the split rendering server, a request for adaptation of a gaze based optimization profile for a split rendering session, wherein the request for adaptation of the gaze based optimization profile comprises an indication of an adaptation to the gaze based optimization profile;receiving, from the split rendering server, an encoded video frame;decoding the encoded video frames to generate a decoded video frame; and displaying the decoded video frame.5 25. A computer readable medium comprising instructions which, when executed by anapparatus comprising a split rendering server, cause the apparatus to perform at least the following:receiving, from a split rendering client, metadata for a virtual camera associated with the split rendering client;10 performing rendering using the metadata and a gaze based optimization profile for asplit rendering session;encoding the video frame using the gaze based optimization profile to generate an encoded video frame; andsending, to the split rendering client, the encoded video frame.
Citation Information
Patent Citations
Single-stream foveal display transport
US20210063741A1
Foviation and hdr
US20210329316A1