Apparatus, method and computer program
The multimodal split-rendering technique in XR systems addresses the issue of visual artifacts by generating and encoding multiple layers with layering indications, enhancing the Quality of Experience by minimizing distortion and improving pose correction accuracy.
Patent Information
- Application Number
- GB2023017097
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-14
AI Technical Summary
Existing split-rendering techniques in extended reality (XR) systems, such as those used in 5G networks, suffer from visual artifacts due to pose correction, particularly when static elements are warped, leading to a decrease in Quality of Experience (QoE).
Implementing a multimodal split-rendering approach that generates multiple layers of images based on positional information, where each layer is associated with a layering indication, and encodes these layers as separate video streams, allowing for selective pose correction and reducing the impact of visual artifacts.
Enhances the Quality of Experience (QoE) by minimizing distortion of static elements and improving the accuracy of pose correction in XR environments, thereby maintaining a smoother and more immersive user experience.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field The present application relates to a method, apparatus, system and computer program and in particular but not exclusively to multimodal split-rendering for extended reality. Background A communication system can be seen as a facility that enables communication sessions between two or more entities such as user terminals, base stations and / or other nodes by providing carriers between the various entities involved in the communications path. A communication system can be provided for example by means of a communication network and one or more compatible communication devices. The communication sessions may comprise, for example, communication of data for carrying communications such as voice, video, electronic mail (email), text message, multimedia and / or content data and so on. Nonlimiting examples of services provided comprise two-way or multi-way calls, data communication or multimedia services and access to a data network system, such as the Internet. In a wireless communication system at least a part of a communication session between at least two stations occurs over a wireless link. Examples of wireless systems comprise public land mobile networks (PLMN), satellite based communication systems and different wireless local networks, for example wireless local area networks (WLAN). Some wireless systems can be divided into cells, and are therefore often referred to as cellular systems. A user can access the communication system by means of an appropriate communication device or terminal. A communication device of a user may be referred to as user equipment (UE) or user device. A communication device is provided with an appropriate signal receiving and transmitting apparatus for enabling communications, for example enabling access to a communication network or communications directly with other users. The communication device may access a carrier provided by a station, for example a base station of a cell, and transmit and / or receive communications on the carrier. The communication system and associated devices typically operate in accordance with a given standard or specification which sets out what the various entities associated with the system are permitted to do and how that should be achieved. Communication protocols and / or parameters which shall be used for the connection are also typically defined. One example of a communications system is Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (UTRAN) (3G radio). Other examples of communication systems are the long-term evolution (LTE) of the Universal Mobile Telecommunications System (UMTS) radio-access technology and so-called 5G or New Radio (NR) networks. NR is being standardized by the 3rd Generation Partnership Project (3GPP). Other examples of communication systems include 5G-Advanced (NR Rel-18 and beyond) and 6G. Summary In a first aspect there is provided an apparatus for a server comprising means for receiving positional information relating to a three-dimensional spatial environment at the server from a user equipment, means for rendering an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication, means for encoding the at least one layer as at least one video stream and means for providing the at least one video stream and the at least one layering indication to the user equipment. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three-dimensional spatial environment. The apparatus may comprise means for determining the at least one layering indication based on at least one of the following: render layers in a rendering pipeline, a depth stencil of the rendered image, or elements of a user interface. The apparatus may comprise means for providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream. The apparatus may comprise means for providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bitstream. The apparatus may comprise means for providing to the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The apparatus may comprise means for providing to the user equipment texture information associated with the image in a first bitstream and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The apparatus may comprise means for determining a number of the at least one layer based on a negotiation during a rendering session setup. The apparatus may comprise means for providing or means for receiving a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a second aspect there is provided an apparatus for a user equipment comprising means for providing positional information relating to a three dimensional spatial environment to a server from the user equipment, means for receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, means for decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment and means for applying correction based on the layering indication to the texture information associated with the image. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three dimensional spatial environment. The apparatus may comprise means for receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bit stream. The apparatus may comprise means for receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bit stream. The apparatus may comprise means for receiving at the user equipment texture information associated with the image in a first bit stream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The apparatus may comprise means for receiving at the user equipment, texture information associated with the image in a first bitstream, and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The apparatus may comprise means for determining a number of the at least one layer based on a negotiation during a rendering session setup. The apparatus may comprise means for providing or means for receiving a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a third aspect there is provided a method comprising receiving positional information relating to a three-dimensional spatial environment at a server from a user equipment, rendering an image from the three dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication, encoding the at least one layer as at least one video stream and providing the at least one video stream and the at least one layering indication to the user equipment. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three-dimensional spatial environment. The method may comprise determining the at least one layering indication based on at least one of the following: render layers in a rendering pipeline, a depth stencil of the rendered image, or elements of a user interface. The method may comprise providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream. The method may comprise providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bitstream. The method may comprise providing to the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The method may comprise providing to the user equipment texture information associated with the image in a first bitstream and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The method may comprise determining a number of the at least one layer based on a negotiation during a rendering session setup. The method may comprise providing or receiving a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a fourth aspect there is provided a method comprising providing positional information relating to a three-dimensional spatial environment to a server from a user equipment, receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment and applying correction based on the layering indication to the texture information associated with the image. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three dimensional spatial environment. The method may comprise receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bit stream. The method may comprise receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bit stream. The method may comprise receiving at the user equipment texture information associated with the image in a first bit stream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The method may comprise receiving at the user equipment, texture information associated with the image in a first bitstream, and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The method may comprise determining a number of the at least one layer based on a negotiation during a rendering session setup. The method may comprise providing or receiving a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a fifth aspect there is provided an apparatus for a server comprising at least one processor, and at least one memory storing instructions which, when executed by the processor, cause the apparatus at least to receive positional information relating to a three-dimensional spatial environment at the server from a user equipment, render an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication;, encode the at least one layer as at least one video stream and provide the at least one video stream and the at least one layering indication to the user equipment. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three-dimensional spatial environment. The apparatus may be caused to determine the at least one layering indication based on at least one of the following: render layers in a rendering pipeline, a depth stencil of the rendered image, or elements of a user interface. The apparatus may be caused to provide texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream. The apparatus may be caused to provide texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bitstream. The apparatus may be caused to provide to the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The apparatus may be caused to provide to the user equipment texture information associated with the image in a first bitstream and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The apparatus may be caused to determine a number of the at least one layer based on a negotiation during a rendering session setup. The apparatus may be caused to provide or receive a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a sixth aspect there is provided an apparatus for a user equipment comprising at least one processor, and at least one memory storing instructions which, when executed by the processor, cause the apparatus at least to provide positional information relating to a three-dimensional spatial environment to a server from the user equipment, receive at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, decode the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment and apply correction based on the layering indication to the texture information associated with the image. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three dimensional spatial environment. The apparatus may be caused to receive the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bit stream. The apparatus may be caused to receive the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bit stream. The apparatus may be caused to receive at the user equipment texture information associated with the image in a first bit stream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The apparatus may be caused to receive at the user equipment, texture information associated with the image in a first bitstream, and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The apparatus may be caused to determine a number of the at least one layer based on a negotiation during a rendering session setup. The apparatus may be caused to provide or receive a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a seventh aspect there is provided a computer readable medium comprising instructions which, when executed by an apparatus for a server, cause the apparatus to perform at least the following receiving positional information relating to a three-dimensional spatial environment at the server from a user equipment, rendering an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication, encoding the at least one layer as at least one video stream and providing the at least one video stream and the at least one layering indication to the user equipment. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three-dimensional spatial environment. The apparatus may be caused to perform determining the at least one layering indication based on at least one of the following: render layers in a rendering pipeline, a depth stencil of the rendered image, or elements of a user interface. The apparatus may be caused to perform providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream. The apparatus may be caused to perform providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bitstream. The apparatus may be caused to perform providing to the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The apparatus may be caused to perform providing to the user equipment texture information associated with the image in a first bitstream and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The apparatus may be caused to perform determining a number of the at least one layer based on a negotiation during a rendering session setup. The apparatus may be caused to perform providing or receiving a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In an eighth aspect there is provided a computer readable medium comprising instructions which, when executed by an apparatus for a user equipment, cause the apparatus to perform at least the following providing positional information relating to a three-dimensional spatial environment to a server from the user equipment, receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment and applying correction based on the layering indication to the texture information associated with the image. The positional information may comprise at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three dimensional spatial environment. The apparatus may be caused to perform receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bit stream. The apparatus may be caused to perform receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bit stream. The apparatus may be caused to perform receiving at the user equipment texture information associated with the image in a first bit stream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. The apparatus may be caused to perform receiving at the user equipment, texture information associated with the image in a first bitstream, and the at least one layering indication as at least one range of texture values in the first bit stream. The indication of the at least one range may be carried as one or more SEI messages in one or more bitstreams. The indication of the at least one range may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. The apparatus may be caused to perform determining a number of the at least one layer based on a negotiation during a rendering session setup. The apparatus may be caused to perform providing or receiving a request to modify the number of the at least one layer. The at least one layering indication may comprise at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping. In a ninth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus to perform at least the method according to the third or fourth aspect. In the above, many different embodiments have been described. It should be appreciated that further embodiments may be provided by the combination of any two or more of the embodiments described above. Description of Figures Embodiments will now be described, by way of example only, with reference to the accompanying Figures in which: Figure 1 shows a schematic diagram of an example 5GS communication system; Figure 2 shows a schematic diagram of an example mobile communication device; Figure 3 shows a schematic diagram of an example control apparatus; Figure 4 shows an example architecture for split-rendering; Figure 5 shows a flowchart of a method according to an example embodiment; Figure 6 shows a flowchart of a method according to an example embodiment; Figure 7 shows a schematic diagram of layers being encoded in a single layer video; Figure 8 shows a schematic diagram of encoding layer information in the depth image; Figure 9 shows a schematic diagram of Illustration of layered rendering and encoding of each layer into its own video stream. Detailed description Before explaining in detail the examples, certain general principles of a wireless communication system and mobile communication devices are briefly explained with reference to Figure 1, Figure 2 and Figure 3 to assist in understanding the technology underlying the described examples. An example of a suitable communications system is the 5G or NR concept. Network architecture in NR may be similar to that of LTE-advanced. Base stations of NR systems may be known as next generation NodeBs (gNBs). Changes to the network architecture may depend on the need to support various radio technologies and finer Quality of Service (QoS) support, and some on-demand requirements for e.g., QoS levels to support Quality of Experience (QoE) for a user. Also network aware services and applications, and service and application aware networks may bring changes to the architecture. Those are related to Information Centric Network (ICN) and User-Centric Content Delivery Network (UC-CDN) approaches. NR may use Multiple Input - Multiple Output (MIMO) antennas, many more base stations or nodes than the LTE (a so-called small cell concept), including macro sites operating in co-operation with smaller stations and perhaps also employing a variety of radio technologies for better coverage and enhanced data rates. Future networks may utilise network functions virtualization (NFV) which is a network architecture concept that proposes virtualizing network node functions into “building blocks” or entities that may be operationally connected or linked together to provide services. A virtualized network function (VNF) may comprise one or more virtual machines running computer program codes using standard or general type servers instead of customized hardware. Cloud computing or data storage may also be utilized. In radio communications, this may mean node operations to be carried out, at least partly, in a server, host or node operationally coupled to a remote radio head. It is also possible that node operations will be distributed among a plurality of servers, nodes or hosts. It should also be understood that the distribution of labour between core network operations and base station operations may differ from that of the LTE or even be non-existent. Figure 1 shows a schematic representation of a 5G system (5GS) 100. The 5GS may comprise a user equipment (UE) 102 (which may also be referred to as a communication device or a terminal), a 5G radio access network (5GRAN) 104, a 5G core network (5GCN) 106, one or more internal or external application functions (AF) 108 and one or more data networks (DN) 110. An example 5G core network (CN) comprises functional entities. The 5GCN 106 may comprise one or more Access and mobility Management Functions (AMF) 112, one or more session management functions (SMF) 114, an authentication server function (ALISF) 116, a Unified Data Management (UDM) 118, one or more user plane functions (UPF) 120, a Unified Data Repository (UDR) 122 and / or a Network Exposure Function (NEF) 124. The UPF is controlled by the SMF (Session Management Function) that receives policies from a PCF (Policy Control Function). The CN is connected to a UE via the Radio Access Network (RAN). The 5GRAN may comprise one or more gNodeB (gNB) Distributed Unit (DU) functions connected to one or more gNodeB (gNB) Centralized Unit (CU) functions. The RAN may comprise one or more access nodes. A User Plane Function (UPF) referred to as PDU Session Anchor (PSA) may be responsible for forwarding frames back and forth between the DN and the tunnels established over the 5G towards the UE(s) exchanging traffic with the DN. An example mobile communication device will now be described in more detail with reference to Figure 2 showing a schematic, partially sectioned view of a communication device 200. Such a communication device is often referred to as user equipment (UE) or terminal. An appropriate mobile communication device may be provided by any device capable of sending and receiving radio signals. Non-limiting examples comprise a mobile station (MS) or mobile device such as a mobile phone or what is known as a ’smart phone’, a computer provided with a wireless interface card or other wireless interface facility (e.g., USB dongle), personal data assistant (PDA) or a tablet provided with wireless communication capabilities, voice over IP (VoIP) phones, portable computers, desktop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehiclemounted wireless terminal devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), smart devices, wireless customerpremises equipment (CPE), or any combinations of these or the like. A mobile communication device may provide, for example, communication of data for carrying communications such as voice, electronic mail (email), text message, multimedia and so on. Users may thus be offered and provided numerous services via their communication devices. Non-limiting examples of these services comprise two-way or multi-way calls, data communication or multimedia services or simply an access to a data communications network system, such as the Internet. Users may also be provided broadcast or multicast data. Non-limiting examples of the content comprise downloads, television and radio programs, videos, advertisements, various alerts, and other information. A mobile device is typically provided with at least one data processing entity 201, at least one memory 202 and other possible components 203 for use in software and hardware aided execution of tasks it is designed to perform, including control of access to and communications with access systems and other communication devices. The data processing, storage and other relevant components can be provided on an appropriate circuit board and / or in chipsets. This feature is denoted by reference 204. The user may control the operation of the mobile device by means of a suitable user interface such as key pad 205, voice commands, touch sensitive screen or pad, combinations thereof or the like. A display 208, a speaker and a microphone can be also provided. Furthermore, a mobile communication device may comprise appropriate connectors (either wired or wireless) to other devices and / or for connecting external accessories, for example hands-free equipment, thereto. The mobile device 200 may receive signals over an air or radio interface 207 via appropriate apparatus for receiving and may transmit signals via appropriate apparatus for transmitting radio signals. In Figure 2 transceiver apparatus is designated schematically by block 206. The transceiver apparatus 206 may be provided for example by means of a radio part and associated antenna arrangement. The antenna arrangement may be arranged internally or externally to the mobile device. Figure 3 shows an example of a control apparatus 300 for a communication system, for example to be coupled to and / or for controlling a station of an access system, such as a RAN node, e.g. a base station, eNB or gNB, a relay node or a core network node such as an MME or Serving Gateway (S-GW) or Packet Data Network Gateway (P-GW), or a core network function such as AMF / SMF, or a server or host. The method may be implemented in a single control apparatus or across more than one control apparatus. The control apparatus may be integrated with or external to a node or module of a core network or RAN. In some embodiments, base stations comprise a separate control apparatus unit or module. In other embodiments, the control apparatus can be another network element such as a radio network controller or a spectrum controller. In some embodiments, each base station may have such a control apparatus as well as a control apparatus being provided in a radio network controller. The control apparatus 300 can be arranged to provide control on communications in the service area of the system. The control apparatus 300 comprises at least one memory 301, at least one data processing unit 302, 303 and an input / output interface 304. Via the interface, the control apparatus can be coupled to a receiver and a transmitter of the base station. The receiver and / or the transmitter may be implemented as a radio front end or a remote radio head. XR (Extended Reality) is a term that encompasses various immersive technologies, including Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). These technologies aim to create experiences that blend the physical and digital worlds, providing users with interactive and immersive environments (e.g., three-dimensional spatial environments). Rendering in XR refers to the process of generating and displaying the digital content (e.g., images) in these immersive environments (also referred to as scenes). The following relates to split-rendering scenarios, as defined, for example in 3GPP standards. Split rendering in 5G refers to a technique used to offload some of the processing tasks related to rendering and graphics from a user's device (such as a smartphone or AR / VR headset) to remote servers or the edge cloud. In this scenario, a device delegates rendering tasks to a remote server by continuously sending its pose information for scene rendering and receiving rendered views in return. These views may include depth and texture data for improved pose correction. The pose correction adapts the view to the current pose, accounting for position and orientation changes in 3D space. Asynchronous time warping (ATW) is one method for achieving this correction. Depth data is a term used to describe a map of per-pixel data containing depth-related information. It is a container for per-pixel distance or disparity information captured by compatible camera devices. Texture data (also referred to as texture information) may include the information that defines the color, brightness, contrast, and surface details of a 3D model. In a split-rendering delivery model, it is likely that the pose of the viewer may change between the instant it sends the view to the network for rendering, and the instant it receives the rendered view for display. If the pose changed, the device would perform a pose correction to adjust the view to the current pose. Pose may comprise rotational and positional parameters that describe viewpoint position and orientation in 3D space. Additional arguments describing the viewport such as near and far plane as well as projection parameters may exist. Figure 4 shows an example management architecture for split-rendering scenarios. In this case, a device (e.g., a UE) offloads rendering tasks to a remote server. In one embodiment, the device continuously sends its pose (or positional) information to the server and the server renders images of the scene (which may be defined as a three dimensional spatial environment) for the requested pose and sends the rendered images as a view back to the device. The view may include depth and texture information to support pose correction at the receiver. The correction may involve both rotational and positional adjustments, and may possibly utilize reprojection. The split rendering model may include SR-4s for user-plane signaling (WebRTC and ICE) and SR-4m for media and metadata exchange. However, pose correction, while reducing motion-to-photon latency, may introduce visual artifacts, impacting the Quality of Experience (QoE). Specifically, static or flat elements in the rendered video, like a user interface, may become warped and distorted when pose correction is applied, negatively affecting the QoE. The pose correction can be achieved for instance, using asynchronous time warping (ATW). Other techniques that allow to warp the received view based on the differences in the renderpose and the current accurate pose may be utilized. Terminology-wise warping could be replaced with reprojection, and it may include both positional and rotational corrections that are applied to the received rendered view that may comprise texture information and optionally also depth information. In the context of split rendering, the SR-4 interface is further classified as SR-4s and SR-4m sub-interfaces. The SR-4s interface covers all user-plane signaling, including WebRTC and ICE signaling. The SR-4m serves for media and metadata exchange between the split rendering client and the split rendering server. The SWAP protocol allows for the definition of application-specific messages. For Split Rendering, the following application-specific messages are supported the configuration message carries the split rendering configuration information from the SRC to the SRS. It shall be identified by the type “urn:3gpp:sr-mse:sr-configuration” and the object shall be formatted according to clause 8.4.2.2of TS 25.565. The rendering description message carries the description of the split rendered media from the SRS to SRC. It shall be identified by the type “urn:3gpp:sr-mse:sr-description”. The rendering description message provides the semantics of the media that is delivered over WebRTC from the SRS to SRC. While the pose correction is a good way to computationally compensate for motion to photon latency, it may introduce visual artifacts reducing the Quality of Experience (QoE). If the rendered video embeds static or flat elements that do not need to be time-warped (e.g., a static user interface), then a negative impact on the QoE would be observed once the pose correction is applied as the static elements are warped and distorted too. A practical example is cloud gaming when game control and rendering are achieved in the cloud. In that case, a rendered frame is generally composed by the scene itself and by a userinterface, anchored in the user field of view (FoV). While a changed pose justifies the need to warp a portion of the received image representing the scene, pose correction should not be done for the user-interface, which is expected to remain anchored at the same position. Figure 5 shows a flowchart according to an example method. The method may be performed at a server. In 501, the method comprises receiving positional information relating to a three-dimensional spatial environment at a server from a user equipment. In 502, the method comprises rendering an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication. In 503, the method comprises encoding the at least one layer as at least one video stream. In 504, the method comprises providing the at least one video stream and the at least one layering indication to the user equipment. Figure 6 shows a flowchart of a method according to an example embodiment. The method may be performed at a UE. In 601, the method comprises providing positional information relating to a three-dimensional spatial environment to a server from a user equipment. In 602, the method comprises receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment. In 603, the method comprises decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment. In 604, the method comprises applying correction based on the layering indication to the texture information associated with the image. The server may comprise a composition engine for performing rendering. The server may comprise a demultiplexer (demux). 5 The positional information (also referred to as pose information) may comprise at least one of the following: one or more orientation parameters or one or more positional parameters, wherein the one or more orientation parameters and one or more positional parameters relate to a viewing position of the image relative to the three-dimensional spatial environment. 10 Table 1 provides specific examples of parameters which may be comprised in pose information. pose Object 1..1 An object that carries the pose information for a particular view. orientation Object 1..1 Represents the orientation of the view pose as a quaternion based on the reference XR space. x number 1..1 Provides the x coordinate of the quaternion. y number 1..1 Provides the y coordinate of the quaternion. z number 1..1 Provides the z coordinate of the quaternion. w number 1..1 Provides the w coordinate of the quaternion. position Object 0..1 Represents the location in 3D space of the pose based on the reference XR space. For eye gaze poses, the position is not required. X number 1..1 Provides the x coordinate of the position vector. y number 1..1 Provides the y coordinate of the position vector. z number 1..1 Provides the z coordinate of the position vector. Table 1 Methods as described with reference to Figure 5 and Figure 6 may leverage a multimodal delivery of rendered video based on warping criterion, reducing QoE drops when pose correction is performed on the UE. An example embodiment of a method performed by the UE and a server comprises the following steps. In step 1, a user sends its pose information to the server. This is an example of providing positional information relating to a three-dimensional spatial environment to a server from a user equipment. In step 2, the server receives the pose information and optionally applies a pose prediction to anticipate network latency. In step 3, the server renders the scene for a requested pose and generates several independent layers, depending on their sensitivity to receiver pose correction. This is an example of rendering an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication. In step 4, the different layers are encoded as one or more video streams. This is an example of encoding the at least one layer as at least one video stream. A video stream may comprise one or more bitstreams. Step 4 may generate additional metadata that can be carried through the system, either as a supplemental enhancement information (SEI) message attached to the video bitstreams, or through the 5G interface, or service-level signaling. The parameters of the additional metadata that may be transmitted include but are not limited to indications of hierarchical position of the layers, indications of sensitivity to time warping or sensitivity to space warping. The parameters may be transmitted as ranges of depth, or alpha values, or dynamic range codewords on which the pose correction should be applied. In step 5, the one or more video streams are sent from the server to the user equipment with the pose the one or more video streams were rendered with. This is an example of providing the at least one video stream and the at least one layering indication to the user equipment. Step 5 may be implemented by leveraging multi-track signaling, for example through the Realtime communications (RTC) Service Description Protocol (SDP), or any other type of service signaling (including Dynamic Adaptive Streaming over HTTP (DASH)ZHTTP Live Streaming (HLS) manifest, or proprietary template). In an example embodiment, the layer meta-data is transmitted with the corresponding layer media in a Real-time Transport Protocol (RTP) header extension compliant with RF8285. An example one byte-header extension could be of the form: 0 12 3 01234567890123456789012345678901 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | OxBE | OxDE | length | +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | ID | L=1 | LayerlD | +-+-+-+-+-+-+.+.+.+-+-+-+-+-+-+-+ And an example two-byte header extension could be of the form: 0 12 3 01234567890123456789012345678901 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | 0x100 |appbits| length | +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | ID | L=1 | LayerlD | +-+-+-+-+-+-+-+-+.+-+-+-+-+-+-+-+ Where layerlD may indicate the layer identifier of the frame or NAL unit in the current RTP header. LayerlD maybe an 8bit integer or a bitmask. In one example embodiment, texture data belonging to all layers is encoded in a single video, and layer identification is encoded as another video, where each pixel value corresponds to layer indication for the same pixel coordinate in the texture video. This is an example of providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bitstream. Alternatively, or in addition, a method as described with reference to Figure 5 may comprise providing to the user equipment, texture information associated with the image in a first bitstream, the at least one layering indication as at least one range of texture values in the first bitstream. In this example scenario, layer information may be combined into a single layer video stream, where the pixel value indicates the layer information for the respective pixels in the texture and depth videos. This is illustrated in Figure 7. Depth information may be already streamed to the user equipment to accommodate translational pose correction, which is why leveraging depth video for layer carriage may be particularly useful. Leveraging depth video does not increase the amount of video data sent towards the user equipment, although it may reduce the accuracy of the depth information. In another embodiment an alpha channel, side metadata is generated to indicate the layer indication. This is an example of for providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream. In an example scenario, the produced video stream(s) embed(s) an alpha channel, and / or depth channel to carry the layering information. This can be done by adding auxiliary channels to a simple High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC) channel. In another embodiment layer indication is encoded as range of depth values in the video stream representing depth information of the rendered view. The range of values allocated for layer indications in the depth video is negotiated before starting split rendering session. Each pixel of the depth frame falling in the layer indication range describes the layer indication for the matching pixel in the texture frame. This is an example of providing to the user equipment; texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream. In an example scenario illustrated in Figure 8, layer information may be encoded in the depth video by allocating a subset of depth values to signal the layer of depth to which the corresponding pixels in the texture video belong. The indication of the at least one range (i.e., the at least one range of texture values or the at least one range of depth values) may be carried as one or more SEI messages in one or more bitstreams. 5 Texture information associated with the image comprises texture values or texture data of the image. For instance, an SEI message could be defined with the following syntax: pose_correction_parameters( payloadSize) { Descriptor pcp_scene_layer_id u(16) pcp_scene_correction_mode u(2) if( pcp_scene_correction_mode == 1){ pcp_variable_correction u(1) if( pcp_variable_correction == 0){ pcp_correction_sensitivity u(16) } else{ pcp_number_custom_areas u(16) for(i=1;i<pcp_number_custom_areas;p++){ pcp_area_x[i] u(16) pcp_area_y[i] u(16) pcp_area_w[i] u(16) pcp_area_h[i] u(16) pcp_area_correction_sensitivity[i] u(16) } } else if( pcp_scene_correction_mode == 2){ pcp_layering_metric u(2) pcp_n u m be rj nterval s u(16) for (i=0; i<pcp_number_intervals; i++){ pcp_interval_upper_bound[i] u
[16] pcp_interval_correction_sensitivity[i] u
[16] } } } Table 2 With the following semantic: • pcp_scene_layer_id indicates the position of the received layer in the hierarchical rendering structure. This is used by a composition engine of the server to decide at which hierarchical level the layer should be rendered. • pcp_scene_correction_mode indicates the correction mode to be applied for the received frame. 0 means no correction should be applied, 1 means a pose correction should be applied, 2 means that a range-based adaptive correction is used. • pcp_variable_correction indicates if a uniform correction should be applied to the received frame, or if a dynamic per area correction is expected. 0 means a uniform correction is applied, 1 means a dynamic correction is applied. • pcp_correction_sensitivity indicates the intensity of the pose correction that should be applied to the received frame. • pcp_number_custom_areas indicates, when pcp_variable_correction is equal to 1, how many custom areas are present in the received frame. • pcp_area_x[i], pcp_area_y[i], pcp_area_h[i] and pcp_area_w[i], respectively indicates the (x,y) position and the width and height of the i-th area in the received frame. • pcp_area_correction_sensitivity[i] indicates the intensity of pose correction that should be applied on the i-th area in the received frame. • pcp_layering_metric indicates, when range-based adaptive correction is used, what metric is used to derive the different layers within the texture. 0 means depth channel information is used, 1 means alpha channel information is used, 2 means dynamic range codewords are used. • pcp_number_intervals indicates in how many intervals the layering is described for the selected metric. • pcp_interval_upper_bound[i] indicates the upper bound value of the i-th interval. • pcp_interval_correction_sensitivity[i] indicates the intensity of pose correction that should be applied on the i-th interval in the received frame. Table 3 shows an alternative example SEI message. pose_correction_parameters( payloadSize ) { Descriptor pcp_metric u(1) pcp_n_intervals u(16) for (i=0; i<pcp_n_intervals; i++){ pcp_interval_upper_bound[i] u(16) pcp_interval_correction_sensitivity[i] u(16) } u(1) } With the following semantic: • pcp_metric indicates what metric is used to extract different layers from video texture. 0 means alpha ranges are used, 1 means depth ranges are used. • pcp_n_intervals indicates in how many intervals the layering is described for the selected metric. • pcp_interval_upper_bound[i] indicates the upper bound value of the i-th interval. • pcp_interval_correction_sensitivity[i] indicates the intensity of pose correction that should be applied on the i-th interval in the received frame. 0 means no pose correction should be applied, other values describes different degrees of pose-correction sensitivity. The indication of the at least one range (i.e., the at least one range of texture values or the at least one range of depth values) may be carried as one or more messages and sent via an out of band communication channel. The out of band communication channel may comprise one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel. In step 6, the UE decodes the one or more video streams, extracts layering indications for the decoded texture information, and forwards the decoded frames and the layering information to the XR runtime. This is an example of receiving at least one layering indication at the user equipment and decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment. Step 6 may require a standard video decoding to be used. If an SEI message approach is used to deliver side-metadata, the demux of the server shall be able to forward the SEI message to the composition engine to be interpreted. In step 7, pose correction is applied to the texture information as described by the layering indications. This is an example of applying correction based on the layering indication to the texture information associated with the image. The at least one layer indication may comprise parameters indicating at least one of the following: a hierarchical position of the at least one layer, a sensitivity of the at least one layer to time warping or a sensitivity of the at least one layer to space warping. Step 7 may be implemented using any pose correction method but for example, may consider the sensitivity to time warping as an input parameter. There are a number of different alternatives for pose corrections, such as the ATW, but pose correction may be defined as processing the received composed image or its components / layers to reflect the latest pose of the user equipment. In step 8, the final pose-corrected texture image is displayed to the viewer. Composing the different layers may be based on the received metadata, for example, via SEI messages. In an example embodiment, OpenXR runtime is used for composition and layer 3 of figure 2 is assigned an XrCompositionLayerQuad. The capability of a SRC and SRS to use multimodal split rendering can be indicated in the split rendering configuration. For example, the following configuration parameter may be used. In addition, the renderingFlag is set to FLAG_ALPHA_BLENDING. Name Type Description poseLayers Object Provides an array of layers that will be created by the SRS, and can be decoded by the SRC. If multimodal split rendering is not supported the parameter is omitted or has only one element (layer). The array consists of index and pose sensitivity for each layer. index number Index number of the layer. Pose_sensitivity Number / string The sensitivity of the pose correction to be applied to that layer. In an example embodiment, one or more fields defined for the pose correction parameters above are transported using an RTP header extension. In another example embodiment, the one or more fields are transported over the data channel. In this case the pcp_scene_layer_id is appropriately associated with the media id or media description for each layer. The association may be defined in SDP with a new parameter, e.g., a=pcp_scene_layer_id <id> under the media description of each layer. Alternatively, the information may be shared as part of application-specific split-rendering description message. Determining the at least one layering indication may be based on render layers in a rendering pipeline, a depth stencil of the rendered image or elements of a user interface (UI). Elements of a UI may comprise, e.g., static or flat elements in the three-dimensional spatial environments such as but not limited to the menu, map or instructions as shown in Figures 7, 8 and 9. In one example embodiment, the layer indications can be derived from the renderlayers in the rendering pipeline, by analyzing the depth stencil of the rendered view, or by utilizing the awareness that the rendering engine has of visual elements of a UI in the view. Figure 9 show a rendered scene can be split in multiple layers having their own depths and sensitivity to time warping. The number of layers may be a rendering-side decision depending on the scene composition and its overall sensitivity to pose correction. Typically, two to three layers are expected to contain sufficient information for the pose-correction. There are applications however, where the number of layers may be increased. A method as described with reference to Figure 5 may comprise determining a number of the at least one layers based on a negotiation during a rendering session setup. The method may comprise providing from the server to the UE or receiving a request at the UE from the server to modify the number of the at least one layers. In an example embodiment, the number of layers is negotiated by the UE and the SRS, at the split rendering session setup. In another embodiment, during a session, the UE or SRS may request a change in the number of layers based, for example, on change in operating conditions of the UE or the SRS or changes in scene content. In an example embodiment, the change in number of layers is requested and acknowledged by the UE via application specific RTCP feed-back messages and by the SRS via RTP header extension messages or over a webrtc data channel protocol with message type set to “urn:3gpp:split-rendering:vN:xxxx”, where N may be an appropriate protocol version identifier and xxxx may be an appropriate message type identifier, for example, “renderLayers” An apparatus for a server may comprise may comprise means for receiving positional information relating to a three dimensional spatial environment at the server from a user equipment, means for rendering an image from the three dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication, means for encoding the at least one layer as at least one video stream and means for providing the at least one video stream and the at least one layering indication to the user equipment. The apparatus may comprise a server, be the user server or be comprised in the server or a chipset for performing at least some actions of / for the server. An apparatus for a user equipment may comprise means for providing positional information relating to a three dimensional spatial environment to a server from the user equipment, means for receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, means for decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment and means for applying correction based on the layering indication to the texture information associated with the image. The apparatus may comprise a user equipment, such as a mobile phone, be the user equipment or be comprised in the user equipment or a chipset for performing at least some actions of / for the user equipment. It should be understood that the apparatuses may comprise or be coupled to other units or modules etc., such as radio parts or radio heads, used in or for transmission and / or reception. Although the apparatuses have been described as one entity, different modules and memory may be implemented in one or more physical or logical entities. It is noted that whilst some embodiments have been described in relation to 5G networks, similar principles can be applied in relation to other networks and communication systems such as 6G networks or 5G-Advanced networks. Therefore, although certain embodiments were described above by way of example with reference to certain example architectures for wireless networks, technologies and standards, embodiments may be applied to any other suitable forms of communication systems than those illustrated and described herein. It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. In general, the various embodiments may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. Some aspects of the disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The embodiments of this disclosure may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computerexecutable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non-transitory media. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may comprise one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), FPGA, gate level circuits and processors based on multi core processor architecture, as non-limiting examples. Embodiments of the disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. The scope of protection sought for various embodiments of the disclosure is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure. The foregoing description has provided byway of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.
Claims
1. An apparatus for a server comprising:means for receiving positional information relating to a three-dimensional spatial environment at the server from a user equipment;means for rendering an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication;means for encoding the at least one layer as at least one video stream; and means for providing the at least one video stream and the at least one layering indication to the user equipment.
2. The apparatus according to claim 1, wherein the positional information comprises at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three-dimensional spatial environment.
3. The apparatus according to any of claims 1 to 2, further comprising means for determining the at least one layering indication based on at least one of the following: render layers in a rendering pipeline, a depth stencil of the rendered image, or elements of a user interface.
4. The apparatus according to any of claims 1 to 3, comprising means for providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream.
5. The apparatus according to any of claims 1 to 3, comprising means for providing texture information associated with the image to the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bitstream.
6. The apparatus according to any of claims 1 to 3, comprising means for providing to the user equipment:texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, andthe at least one layering indication as at least one range of depth values in the second bitstream.
7. The apparatus according to any of claims 1 to 3, comprising means for providing to the user equipment:texture information associated with the image in a first bitstream and the at least one layering indication as at least one range of texture values in the first bitstream.
8. The apparatus according to any of claims 6 or 7, wherein the indication of the at least one range is carried as one or more SEI messages in one or more bitstreams.
9. The apparatus according to any of claims 1 to 8, wherein the indication of the at least one range is carried as one or more messages and sent via an out of band communication channel.
10. The apparatus according to claim 9, wherein the out of band communication channel comprises one or more of: an user plane signalling channel, a control plane signalling channel, and an application layer negotiation channel.
11. The apparatus according to any of claims 1 to 10, comprising means for determining a number of the at least one layer based on a negotiation during a rendering session setup.
12. The apparatus according to claim 11, comprising means for providing or means for receiving a request to modify the number of the at least one layer.
13. The apparatus according to any of claims 1 to 12, wherein the at least one layering indication comprises at least one parameter indicating at least one of the following: a hierarchical position of the at least one layer, sensitivity of the at least one layer to time warping or sensitivity of the at least one layer to space warping.
14. An apparatus for a user equipment comprising:means for providing positional information relating to a three dimensional spatial environment to a server from the user equipment;means for receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment;means for decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment; andmeans for applying correction based on the layering indication to the texture information associated with the image.
15. The apparatus according to claim 14, wherein the positional information comprises at least one of the following: one or more orientation parameters; or one or more positional parameters, wherein the one or more orientation parameters and the one or more positional parameters relate to a viewing position of the image relative to the three dimensional spatial environment.
16. The apparatus according to any of claims 14 to 15, comprising means for receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication in an alpha channel of the first bitstream.
17. The apparatus according to any of claims 14 to 15, comprising means for receiving the texture information associated with the image at the user equipment in a first bitstream and the at least one layering indication to the user equipment in a second bit stream.
18. The apparatus according to any of claims 14 to 15, comprising means for receiving at the user equipment;texture information associated with the image in a first bit stream, depth information associated with the image in a second bitstream, and the at least one layering indication as at least one range of depth values in the second bitstream.
19. The apparatus according to any of claims 14 to 15, comprising means for receiving at the user equipment;texture information associated with the image in a first bitstream, andthe at least one layering indication as at least one range of texture values in the first bit stream.
20. A method comprising:receiving positional information relating to a three-dimensional spatial environment at a server from a user equipment;rendering an image from the three dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication;encoding the at least one layer as at least one video stream; andproviding the at least one video stream and the at least one layering indication to the user equipment.
21. A method comprising:providing positional information relating to a three-dimensional spatial environment to a server from a user equipment;receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment;decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment; andapplying correction based on the layering indication to the texture information associated with the image.
22. An apparatus for a server comprising:at least one processor, and at least one memory storing instructions which, when executed by the processor, cause the apparatus at least to:receive positional information relating to a three-dimensional spatial environment at the server from a user equipment;render an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication;encode the at least one layer as at least one video stream; andprovide the at least one video stream and the at least one layering indication to the user equipment.
23. An apparatus for a user equipment comprising: at least one processor, and at least one memory storing instructions which, when executed by the processor, cause the apparatus at least to: provide positional information relating to a three-dimensional spatial environment to a server from the user equipment; receive at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment; decode the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment; and apply correction based on the layering indication to the texture information associated with the image.
24. A computer readable medium comprising instructions which, when executed by an apparatus for a server, cause the apparatus to perform at least the following: receiving positional information relating to a three-dimensional spatial environment at the server from a user equipment; rendering an image from the three-dimensional spatial environment based on the positional information, wherein rendering the image based on the positional information comprises generating at least one layer, each at least one layer associated with a corresponding at least one layering indication; encoding the at least one layer as at least one video stream; and providing the at least one video stream and the at least one layering indication to the user equipment.
25. A computer readable medium comprising instructions which, when executed by an apparatus for a user equipment, cause the apparatus to perform at least the following: providing positional information relating to a three-dimensional spatial environment to a server from the user equipment; receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment;decoding the at least one video stream, wherein decoding the at least one video stream comprises determining texture information associated with an image from the three-dimensional spatial environment; andapplying correction based on the layering indication to the texture information5 associated with the image.40
Citation Information
Patent Citations
Method, an apparatus and a computer program product for volumetric video encoding and decoding
US20220217400A1