Apparatus, method and computer program for split rendering of augmented reality

By rendering and layering images on the server side and transmitting layering instructions through out-of-band communication channels, the problem of visual artifacts introduced by attitude correction in wireless communication systems is solved, thus improving the user experience quality of extended reality.

CN121794992APending Publication Date: 2026-04-03NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-01
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing wireless communication systems suffer from visual artifacts introduced by pose correction in extended reality, especially during segmentation rendering when pose correction of static or planar elements leads to a decline in the quality of experience.

Method used

By rendering images on the server side based on the location information provided by the user device, and encoding the images into video streams in layers, and transmitting layer indication information using out-of-band communication channels, the user device performs texture and depth information correction to achieve pose correction.

Benefits of technology

It improves the quality of extended reality experiences, reduces visual artifacts caused by posture correction, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121794992A_ABST
    Figure CN121794992A_ABST
Patent Text Reader

Abstract

There is provided an apparatus for a server, comprising means for receiving, at the server, location information relating to a three-dimensional space environment from a user device, means for rendering an image from the three-dimensional space environment based on the location information, where rendering the image based on the location information comprises generating at least one layer, each at least one layer is associated with a respective at least one hierarchy indication, means for encoding the at least one layer into at least one video stream, and means for providing the at least one video stream and the at least one hierarchy indication to the user device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a method, apparatus, system, and computer program, and particularly, but not exclusively, to multimodal segmentation rendering for extended reality. Background Technology

[0002] A communication system can be viewed as a facility that enables a communication session between two or more entities (e.g., user terminals, base stations, and / or other nodes) by providing carrier waves between various entities involved in the communication path. The communication system can be provided, for example, through a communication network and one or more compatible communication devices. The communication session may include, for example, data communication for carrying communications such as voice, video, email, text messages, multimedia, and / or content data. Non-limiting examples of the services provided include two-way or multi-way calling, data communication or multimedia services, and access to data network systems (e.g., the Internet).

[0003] In wireless communication systems, at least a portion of a communication session between at least two stations occurs via a wireless link. Examples of wireless systems include Public Land Mobile Networks (PLMNs), satellite-based communication systems, and various wireless local area networks (WLANs). Some wireless systems can be divided into cells and are therefore often referred to as cellular systems.

[0004] Users can access the communication system through appropriate communication equipment or terminals. The user's communication equipment may be referred to as User Equipment (UE) or User Device. The communication equipment is equipped with appropriate signal receiving and transmitting means for enabling communication, such as access to a communication network or direct communication with other users. The communication equipment can access a carrier provided by a station (e.g., a base station in a cell) and transmit and / or receive communication on that carrier.

[0005] Communication systems and related equipment typically operate according to a given standard or specification that defines what the various entities associated with the system are allowed to do and how they should be implemented. Communication protocols and / or parameters applied to the connection are also typically defined. One example of a communication system is the Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (UTRAN) (3G radio). Other examples of communication systems are the Long Term Evolution (LTE) of UMTS Radio Access Technology and so-called 5G or New Radio (NR) networks. NR is being standardized by the 3rd Generation Partnership Project (3GPP). Other examples of communication systems include 5G-Advanced (NR Rel-18 and later versions) and 6G. Summary of the Invention

[0006] In a first aspect, an apparatus for a server is provided, comprising components for receiving location information relating to a three-dimensional spatial environment from a user device at the server, components for rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with a corresponding at least one layer indication, components for encoding the at least one layer into at least one video stream, and components for providing the at least one video stream and the at least one layer indication to the user device.

[0007] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0008] The apparatus may include components for determining the at least one layer indication based on at least one of the following: a rendering layer in the rendering pipeline, a depth template of the rendered image, or an element of the user interface.

[0009] The apparatus may include components for providing texture information associated with the image to the user equipment in a first bitstream and for providing the at least one layer indication in the alpha channel of the first bitstream.

[0010] The apparatus may include components for providing texture information associated with the image to the user equipment in a first bitstream and for providing the at least one layer indication to the user equipment in a second bitstream.

[0011] The apparatus may include components for providing the user equipment with texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0012] The apparatus may include components for providing the user equipment with texture information associated with the image in a first bitstream and the at least one layer indication in the first bitstream as at least one range of texture values.

[0013] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0014] The indication of at least one range can be carried as one or more messages and sent via an out-of-band communication channel.

[0015] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0016] The apparatus may include components for determining the number of the at least one layer based on negotiation during the rendering session establishment.

[0017] The apparatus may include components for providing or receiving requests to modify the number of the at least one layer.

[0018] The at least one layer indication may include at least one parameter indicating at least one of the following: the layer position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0019] In a second aspect, an apparatus for a user equipment is provided, comprising components for providing location information related to a three-dimensional spatial environment from the user equipment to a server, components for receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, components for decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment, and components for applying a correction to the texture information associated with the image based on the layering indication.

[0020] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0021] The apparatus may include components for receiving texture information associated with the image in a first bitstream at the user equipment and for receiving the at least one layering indication in an alpha channel of the first bitstream.

[0022] The apparatus may include components for receiving texture information associated with the image in a first bitstream at the user equipment and for receiving the at least one layering indication sent to the user equipment in a second bitstream.

[0023] The apparatus may include components for receiving at the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and the at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0024] The apparatus may include components for receiving, at the user equipment, texture information associated with the image in a first bitstream, and at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0025] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0026] The indication of at least one range can be carried as one or more messages and sent via an out-of-band communication channel.

[0027] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0028] The apparatus may include components for determining the number of the at least one layer based on negotiation during the rendering session establishment.

[0029] The apparatus may include components for providing or for receiving requests to modify the number of the at least one layer.

[0030] The at least one layer indication may include at least one parameter indicating at least one of the following: the layer position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0031] In a third aspect, a method is provided, comprising receiving location information related to a three-dimensional spatial environment from a user device at a server, rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information comprises generating at least one layer, each at least one layer being associated with a corresponding at least one layer indication, encoding the at least one layer into at least one video stream, and providing the at least one video stream and the at least one layer indication to the user device.

[0032] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0033] The method may include determining the at least one layer indication based on at least one of the following: a rendering layer in the rendering pipeline, a depth template of the rendered image, or an element of the user interface.

[0034] The method may include providing texture information associated with the image to the user equipment in a first bitstream, and providing the at least one layering indication in the alpha channel of the first bitstream.

[0035] The method may include providing the user equipment with texture information associated with the image in a first bitstream, and providing the user equipment with the at least one layering indication in a second bitstream.

[0036] The method may include providing the user equipment with texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0037] The method may include providing the user equipment with texture information associated with the image in a first bitstream, and the at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0038] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0039] The indication of at least one range can be carried as one or more messages and sent via an out-of-band communication channel.

[0040] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0041] The method may include determining the number of the at least one layer based on negotiation during the rendering session establishment.

[0042] The method may include providing or receiving a request to modify the number of the at least one layer.

[0043] The at least one layer indication may include at least one parameter indicating at least one of the following: the layer position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0044] In a fourth aspect, a method is provided, comprising providing location information related to a three-dimensional spatial environment from a user equipment to a server, receiving at the user equipment at at least one video stream and at least one layering indication associated with the at least one video stream from the server, decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment, and applying a correction to the texture information associated with the image based on the layering indication.

[0045] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0046] The method may include receiving texture information associated with the image in a first bitstream at the user equipment, and receiving the at least one layering indication in the alpha channel of the first bitstream.

[0047] The method may include receiving texture information associated with the image in a first bitstream at the user equipment, and receiving the at least one layering indication sent to the user equipment in a second bitstream.

[0048] The method may include receiving at the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0049] The method may include receiving, at the user equipment, texture information associated with the image in a first bitstream, and the at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0050] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0051] The indication of at least one range can be carried as one or more messages and sent via an out-of-band communication channel.

[0052] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0053] The method may include determining the number of the at least one layer based on negotiation during the rendering session establishment.

[0054] The method may include providing or receiving a request to modify the number of the at least one layer.

[0055] The at least one layer indication may include at least one parameter indicating at least one of the following: the layer position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0056] In a fifth aspect, an apparatus for a server is provided, comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the processor, causing the apparatus to receive, at least at the server, location information relating to a three-dimensional spatial environment, and to render an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with a corresponding at least one layer indication, encoding the at least one layer into at least one video stream, and providing the at least one video stream and the at least one layer indication to the user equipment.

[0057] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0058] The apparatus can be configured to determine the at least one layer indication based on at least one of the following: a rendering layer in the rendering pipeline, a depth template of the rendered image, or an element of the user interface.

[0059] The apparatus may be configured to provide texture information associated with the image to the user equipment in a first bitstream, and to provide the at least one layering indication in the alpha channel of the first bitstream.

[0060] The apparatus can be configured to provide the user equipment with texture information associated with the image in a first bitstream, and to provide the user equipment with the at least one layering indication in a second bitstream.

[0061] The apparatus may be configured to provide the user equipment with texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0062] The apparatus may be configured to provide the user equipment with texture information associated with the image in a first bitstream, and the at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0063] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0064] The indication of at least one range can be carried as one or more messages and transmitted via an out-of-band communication channel.

[0065] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0066] The apparatus can be configured to determine the number of the at least one layer based on negotiation during the rendering session establishment.

[0067] The device can be configured to provide or receive requests to modify the number of the at least one layer.

[0068] The at least one layer indication may include at least one parameter, which indicates at least one of the following: the hierarchical position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0069] In a sixth aspect, an apparatus for a user equipment is provided, comprising at least one processor and at least one memory storing instructions, which, when executed by the processor, cause the apparatus to perform at least the following operations: providing location information related to a three-dimensional spatial environment from the user equipment to a server; receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment; and decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment and applying correction to the texture information associated with the image based on the layering indication.

[0070] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0071] The apparatus may be configured to receive texture information associated with the image in a first bitstream at the user equipment, and to receive the at least one layering indication in the alpha channel of the first bitstream.

[0072] The apparatus may be configured to receive texture information associated with the image in a first bitstream at the user equipment, and to receive the at least one layering indication sent to the user equipment in a second bitstream.

[0073] The apparatus may be configured to receive at the user equipment texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0074] The apparatus may be configured to receive, at the user equipment, texture information associated with the image in a first bitstream, and at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0075] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0076] The indication of at least one range can be carried as one or more messages and transmitted via an out-of-band communication channel.

[0077] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0078] The apparatus can be configured to determine the number of the at least one layer based on negotiation during the rendering session establishment.

[0079] The device can be configured to provide or receive requests to modify the number of the at least one layer.

[0080] The at least one layer indication may include at least one parameter, which indicates at least one of the following: the hierarchical position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0081] In a seventh aspect, a computer-readable medium is provided, including instructions that, when executed by an apparatus for a server, cause the apparatus to perform at least the following operations: receiving location information related to a three-dimensional spatial environment from a user device at the server; rendering an image from the three-dimensional spatial environment based on the location information; wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with a corresponding at least one layer indication; encoding the at least one layer into at least one video stream; and providing the at least one video stream and the at least one layer indication to the user device.

[0082] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0083] The apparatus may be configured to perform a determination of the at least one layering instruction based on at least one of the following: a rendering layer in the rendering pipeline, a depth template of the rendered image, or an element of the user interface.

[0084] The apparatus may be configured to provide texture information associated with the image to the user equipment in a first bitstream and to provide the at least one layering indication in the alpha channel of the first bitstream.

[0085] The apparatus may be configured to provide the user equipment with texture information associated with the image in a first bitstream and with the at least one layering indication in a second bitstream.

[0086] The apparatus may be configured to perform the following: providing the user equipment with texture information associated with the image in a first bitstream, depth information associated with the image in a second bitstream, and at least one layering indication as at least one range of depth values ​​in the second bitstream.

[0087] The apparatus may be configured to perform the function of providing the user equipment with texture information associated with the image in a first bitstream and at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0088] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0089] The indication of at least one range can be carried as one or more messages and transmitted via an out-of-band communication channel.

[0090] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0091] The apparatus can be configured to perform negotiation during the establishment of a rendering session to determine the number of the at least one layer.

[0092] The device can be configured to perform a request to provide or receive a modification of the number of the at least one layer.

[0093] The at least one layer indication may include at least one parameter indicating at least one of the following: the hierarchical position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0094] In an eighth aspect, a computer-readable medium is provided, including instructions that, when executed by a device of a user equipment, cause the device to perform at least the following operations: providing location information related to a three-dimensional spatial environment from the user equipment to a server; receiving at the user equipment at the server at at least one video stream and at least one layering indication associated with the at least one video stream; and decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment and applying correction to the texture information associated with the image based on the layering indication.

[0095] The location information may include at least one of the following: one or more orientation parameters; or one or more location parameters, wherein the one or more orientation parameters and the one or more location parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

[0096] The apparatus can be configured to receive texture information associated with the image in a first bitstream at the user equipment, and to receive the at least one layering indication in the alpha channel of the first bitstream.

[0097] The apparatus can be configured to receive texture information associated with the image in a first bitstream at the user equipment, and receive the at least one layering indication directed to the user equipment in a second bitstream.

[0098] The apparatus can be configured to receive texture information associated with the image in a first bitstream, receive depth information associated with the image in a second bitstream, and receive the at least one layer indication as at least one range of depth values ​​in the second bitstream.

[0099] The apparatus can be configured to receive texture information associated with the image in a first bitstream at the user equipment, and to receive the at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0100] The indication of at least one range can be carried as one or more SEI messages in one or more bit streams.

[0101] The indication of at least one range can be carried as one or more messages and sent via an out-of-band communication channel.

[0102] The out-of-band communication channel may include one or more of the following: user plane signaling channel, control plane signaling channel, and application layer negotiation channel.

[0103] The apparatus can be configured to perform a negotiation based on the establishment of the rendering session to determine the number of the at least one layer.

[0104] The apparatus can be configured to perform a request to provide or receive a modification of the number of the at least one layer.

[0105] The at least one layer indication may include at least one parameter indicating at least one of the following: the hierarchical position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

[0106] In a ninth aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing a device to execute at least the method described in accordance with a third or fourth aspect.

[0107] Many different embodiments have been described above. It should be understood that further embodiments can be provided by combining any two or more of the above embodiments. Attached Figure Description

[0108] The embodiments will now be described by way of example only, with reference to the accompanying drawings: Figure 1 A schematic diagram of an example 5GS communication system is shown; Figure 2 A schematic diagram of an example mobile communication device is shown; Figure 3 A schematic diagram of an example control device is shown; Figure 4 An example architecture for segmented rendering is shown; Figure 5 A flowchart of a method according to an example embodiment is shown; Figure 6 A flowchart of a method according to an example embodiment is shown; Figure 7 This diagram illustrates the coding layer in a single-layer video. Figure 8 This diagram illustrates the encoding of layer information in a depth image. Figure 9 A schematic diagram illustrates layered rendering and how each layer is encoded into its own video stream. Detailed Implementation

[0109] Before explaining the examples in detail, refer to Figure 1 , Figure 2 and Figure 3 A brief explanation of some general principles of wireless communication systems and mobile communication devices is provided to aid in understanding the technology upon which the examples are based.

[0110] Examples of suitable communication systems are the 5G or NR concepts. The network architecture in NR may resemble that of LTE-advanced. Base stations in NR systems may be referred to as next-generation NodeBs (gNBs). Changes in network architecture may depend on the need to support various wireless technologies and more granular Quality of Service (QoS) support, as well as some on-demand requirements, such as QoS levels supporting Quality of User Experience (QoE). Furthermore, network-aware services and applications, and service- and application-aware networks, may bring about architectural changes. These relate to Information Center Network (ICN) and User-Centric Content Delivery Network (UC-CDN) approaches. NR may use multiple-input multiple-output (MIMO) antennas, significantly more base stations or nodes than LTE (the so-called small cell concept), including macro sites cooperating with smaller sites, and may also employ various wireless technologies to achieve better coverage and enhanced data rates.

[0111] Future networks may leverage Network Functions Virtualization (NFV), a network architecture concept that proposes virtualizing network node functions as "building blocks" or entities that can operate connections or be linked together to provide services. Virtualized network functions (VNFs) may include one or more virtual machines running computer program code using standard or general-purpose servers instead of custom hardware. Cloud computing or data storage may also be utilized. In wireless communications, this could mean that node operations are performed, at least partially, within servers, hosts, or nodes coupled to remote wireless head operations. Node operations may also be distributed across multiple servers, nodes, or hosts. It should also be understood that the workload distribution between core network operations and base station operations may differ from, or even not exist, in LTE.

[0112] Figure 1 A schematic representation of a 5G system (5GS) 100 is shown. The 5GS may include a user equipment (UE) 102 (also referred to as a communication device or terminal), a 5G radio access network (5GRAN) 104, a 5G core network (5GCN) 106, one or more internal or external application functions (AF) 108, and one or more data networks (DN) 110.

[0113] The example 5G core network (CN) includes functional entities. 5GCN 106 may include one or more Access and Mobility Management Functions (AMF) 112, one or more Session Management Functions (SMF) 114, Authentication Server Function (AUSF) 116, Unified Data Management (UDM) 118, one or more User Plane Functions (UPF) 120, Unified Data Repository (UDR) 122, and / or Network Openness Function (NEF) 124. The UPF is controlled by the SMF (Session Management Function), which receives policies from the PCF (Policy Control Function).

[0114] The CN connects to the UE via a radio access network (RAN). The 5G RAN may include one or more gNodeB (gNB) distributed unit (DU) functions connected to one or more gNodeB (gNB) centralized unit (CU) functions. The RAN may include one or more access nodes.

[0115] The User Plane Function (UPF), known as the PDU Session Anchor (PSA), is responsible for forwarding frames back and forth between the DN and the tunnel established by the UE that exchanges traffic with the DN via 5G.

[0116] Now refer to Figure 2 A more detailed description of the example mobile communication device, Figure 2 A schematic partial cross-sectional view of a communication device 200 is shown. Such a communication device is generally referred to as a user equipment (UE) or terminal. Suitable mobile communication devices can be provided by any device capable of transmitting and receiving radio signals. Non-limiting examples include mobile stations (MS) or mobile devices, such as mobile phones or so-called "smartphones," computers equipped with wireless interface cards or other wireless interface facilities (e.g., USB dongles), personal digital assistants (PDAs) or tablets equipped with wireless communication capabilities, VoIP phones, portable computers, desktop computers, image capture terminal devices (e.g., digital cameras), gaming terminal devices, music storage and playback devices, in-vehicle wireless terminal devices, wireless endpoints, mobile stations, laptop embedded devices (LEEs), laptop mounted devices (LMEs), smart devices, wireless client devices (CPEs), or any combination of these devices or similar devices. Mobile communication devices can, for example, provide data communication for carrying communications such as voice, email, text messaging, multimedia, etc. Therefore, numerous services can be provided and supplied to users through their communication devices. Non-limiting examples of these services include two-way or multi-way calling, data communications or multimedia services, or simply access to data communications network systems (such as the Internet). Broadcast or multicast data may also be provided to users. Non-limiting examples of content include downloads, television and radio programs, videos, advertisements, various alarms, and other information.

[0117] Mobile devices typically include at least one data processing entity 201, at least one memory 202, and other possible components 203 for software and hardware assistance in performing tasks designed to be performed, including controlling access to and communication with access systems and other communication devices. The data processing, storage, and other related components may be housed in suitable circuit boards and / or chipsets. This feature is indicated by reference numeral 204. Users can control the operation of the mobile device through a suitable user interface, such as a keyboard 205, voice commands, a touchscreen or touchpad, combinations thereof, or the like. A display 208, a speaker, and a microphone may also be provided. Furthermore, mobile communication devices may include suitable connectors (wired or wireless) for connecting to other devices and / or for connecting external accessories, such as hands-free devices.

[0118] Mobile device 200 can receive signals via air or radio interface 207 through appropriate means for receiving, and can transmit signals via appropriate means for transmitting radio signals. Figure 2 In the diagram, the transceiver device is schematically represented by block 206. The transceiver device 206 can be provided, for example, via a radio section and an associated antenna arrangement. The antenna arrangement can be located inside or outside the mobile device.

[0119] Figure 3 An example of a control device 300 for a communication system is shown, such as coupled to and / or used to control access to the system, such as RAN nodes, like base stations, eNBs or gNBs, relay nodes, or core network nodes such as MMEs or Serving Gateways (S-GWs) or Packet Data Network Gateways (P-GWs), or core network functions such as AMFs / SMFs, or servers or hosts. The method can be implemented in a single control device or across multiple control devices. The control device can be integrated with or external to nodes or modules of the core network or RAN. In some embodiments, the base station includes a separate control device unit or module. In other embodiments, the control device can be another network element, such as a radio network controller or a spectrum controller. In some embodiments, each base station can have such a control device as well as control devices provided in the radio network controller. The control device 300 can be arranged to provide control over communications within the service area of ​​the system. The control device 300 includes at least one memory 301, at least one data processing unit 302, 303, and an input / output interface 304. The control device can be coupled to a receiver and transmitter of the base station via the interface. The receiver and / or the transmitter can be implemented as a radio front end or a remote radio head end.

[0120] XR (Extended Reality) is a term that encompasses a variety of immersive technologies, including Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). These technologies aim to create experiences that blend the physical and digital worlds, providing users with interactive and immersive environments (e.g., three-dimensional spatial environments). Rendering in XR refers to the process of generating and displaying digital content (e.g., images) within these immersive environments (also known as scenes).

[0121] The following content relates to segmented rendering scenes, such as those defined in the 3GPP standard.

[0122] Split rendering in 5G refers to a technique used to offload some of the processing tasks related to rendering and graphics from user devices (such as smartphones or AR / VR headsets) to remote servers or the edge cloud. In this scenario, the device delegates rendering tasks to a remote server by continuously sending its pose information for scene rendering and receiving the returned rendered views. These views may include depth and texture data to improve pose correction. Pose correction adapts the view to the current pose, taking into account changes in position and orientation in 3D space. Asynchronous Time Warp (ATW) is one method to achieve this correction. Depth data is a term used to describe a pixel-by-pixel data mapping that contains depth-related information. It is a container for pixel-by-pixel distance or parallax information captured by compatible camera devices. Texture data (also known as texture information) can include information that defines the color, brightness, contrast, and surface details of a 3D model.

[0123] In a segmented rendering transport model, the viewer's pose is likely to change between the moment the viewer sends a view to the network for rendering and the moment the rendered view is received for display. If the pose changes, the device will perform pose correction to adjust the view to the current pose. The pose can include rotational and positional parameters describing the viewpoint's position and orientation in 3D space. Additional parameters describing the viewport may also exist, such as near and far planes and projection parameters.

[0124] Figure 4 An example management architecture for segmenting a rendering scene is illustrated. In this scenario, the device (e.g., the UE) offloads the rendering task to a remote server. In one embodiment, the device continuously sends its pose (or position) information to the server, and the server renders an image of the scene (which can be defined as a three-dimensional spatial environment) for the requested pose and sends the rendered image back to the device as a view. The view may include depth and texture information to support pose correction at the receiver.

[0125] The corrections may involve rotation and position adjustments and may utilize reprojection. The segmented rendering model may include SR-4s for user plane signaling (WebRTC and ICE) and SR-4m for media and metadata exchange. However, while pose correction reduces motion-to-photon latency, it can introduce visual artifacts, impacting Quality of Experience (QoE). Specifically, static or planar elements in rendered video (such as user interfaces) may become distorted and warped when pose correction is applied, negatively affecting QoE.

[0126] Pose correction can be achieved, for example, using asynchronous time warp (ATW). Other techniques can be utilized that allow warping of the received view based on the difference between the rendered pose and the current accurate pose. In technical terms, warping can be replaced by reprojection, and it can include position and rotation corrections applied to the received rendered view, which may include texture information and optionally depth information.

[0127] In the context of segmented rendering, the SR-4 interface is further subdivided into SR-4s and SR-4m sub-interfaces. The SR-4s interface covers all user plane signaling, including WebRTC and ICE signaling. SR-4m is used for media and metadata exchange between segmented rendering clients and segmented rendering servers.

[0128] The SWAP protocol allows for the definition of application-specific messages. For split-rendering, the following application-specific messages are supported: Configuration messages transmit split-rendering configuration information from the SRC to the SRS. It should be identified by the type "urn:3gpp:sr-mse:sr-configuration", and the object should be formatted according to clause 8.4.2.2 of TS 25.565. Render description messages transmit a description of the split-rendering media from the SRS to the SRC. It should be identified by the type "urn:3gpp:sr-mse:sr-description". Render description messages provide semantics for the media transmitted from the SRS to the SRC via WebRTC.

[0129] While attitude correction is a good way to computationally compensate for motion-to-photon delay, it can introduce visual artifacts, thereby reducing the quality of experience (QoE).

[0130] If the rendered video embeds static or planar elements that do not require time warping (e.g., static user interfaces), then once pose correction is applied, a negative impact on QoE will be observed because the static elements will also be warped and distorted.

[0131] A practical example is cloud gaming, where game control and rendering are performed in the cloud. In this case, a rendered frame typically consists of the scene itself and the user interface anchored in the user's field of view (FoV). While pose changes justify distorting the portion of the received image representing the scene, pose correction should not be applied to the user interface, as it is expected to remain anchored in the same position.

[0132] Figure 5 A flowchart illustrating an example method is shown. This method can be executed on a server.

[0133] In 501, the method includes receiving location information related to a three-dimensional spatial environment from a user device at a server.

[0134] In 502, the method includes rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with a corresponding at least one layer indication.

[0135] In 503, the method includes encoding the at least one layer into at least one video stream.

[0136] In 504, the method includes providing the user equipment with the at least one video stream and the at least one layering indication.

[0137] Figure 6 A flowchart of a method according to an example embodiment is shown. The method can be executed at the UE.

[0138] In 601, the method includes providing location information related to a three-dimensional spatial environment from a user device to a server.

[0139] In 602, the method includes receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment.

[0140] In 603, the method includes decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment.

[0141] In 604, the method includes applying correction to the texture information associated with the image based on the layering indication.

[0142] The server may include a compositing engine for performing rendering. The server may include a demultiplexer.

[0143] Position information (also known as pose information) may include at least one of the following: one or more orientation parameters or one or more position parameters, wherein the one or more orientation parameters and one or more position parameters relate to the viewing position of the image relative to the three-dimensional spatial environment.

[0144] Table 1 provides specific examples of parameters that can be included in the attitude information.

[0145]

[0146] Table 1 refer to Figure 5 and Figure 6 The described method can reduce QoE drop when performing attitude correction on the UE by leveraging multimodal delivery of rendered video based on a distortion criterion.

[0147] An example embodiment of the method performed by the UE and the server includes the following steps.

[0148] In step 1, the user sends their pose information to the server. This is an example of providing location information related to the three-dimensional spatial environment from the user device to the server.

[0149] In step 2, the server receives attitude information and optionally applies attitude prediction to predict network latency.

[0150] In step 3, the server renders the scene for the requested pose and generates multiple independent layers, depending on their sensitivity to receiver pose correction. This is an example of rendering an image from a 3D spatial environment based on location information, where rendering the image based on location information includes generating at least one layer, each of which is associated with at least one corresponding layering indicator.

[0151] In step 4, the different layers are encoded into one or more video streams. This is an example of encoding at least one layer into at least one video stream. A video stream may include one or more bitstreams.

[0152] Step 4 can generate additional metadata that can be transmitted through the system, either as a Supplemental Enhancement Information (SEI) message appended to the video bitstream, or via the 5G interface or service-level signaling. Parameters of the additional metadata that can be transmitted include, but are not limited to, indications of the layer's layer location, sensitivity to temporal warp, or sensitivity to spatial warp. Parameters can be transmitted as depth ranges, alpha values, or dynamic range codewords to which attitude correction should be applied.

[0153] In step 5, one or more video streams, along with the poses used to render them, are sent from the server to the user device. This is an example of providing the user device with at least one video stream and at least one layering instruction.

[0154] Step 5 can be achieved by utilizing multitrack signaling, such as through the Real-Time Communication (RTC) Service Description Protocol (SDP), or any other type of service signaling, including HTTP-based Dynamic Adaptive Streaming (DASH) / HTTP Live Streaming (HLS) manifests, or proprietary templates.

[0155] In the example embodiment, the layer metadata, along with the corresponding layer media, is transmitted in a Real-Time Transport Protocol (RTP) header extension compliant with RF8285.

[0156] Example single-byte header extensions can take the following form:

[0157] Example double-byte header extensions can take the following form:

[0158] The layerID can indicate the layer identifier of the frame or NAL unit in the current RTP header. The layerID can be an 8-bit integer or a bitmask.

[0159] In one example embodiment, texture data belonging to all layers is encoded in a single video, and layer identifiers are encoded in another video, where each pixel value corresponds to a layer indicator at the same pixel coordinates in the texture video. This is an example of providing texture information associated with an image to a user device in a first bitstream and providing the at least one layer indicator to the user device in a second bitstream.

[0160] Alternatively or otherwise, refer to Figure 5 The described method may include providing a user device with texture information associated with an image in a first bitstream, and the at least one layer indication as at least one range of texture values ​​in the first bitstream.

[0161] In this example scenario, layer information can be combined into a single-layer video stream, where pixel values ​​indicate the layer information of corresponding pixels in the texture and depth video. This is in Figure 7 As shown in the diagram, depth information may have already been streamed to the user equipment to accommodate translational attitude corrections, which is why using depth video for layer carrying can be particularly useful. Using depth video does not increase the amount of video data sent to the user equipment, although it may reduce the accuracy of the depth information.

[0162] In another embodiment, an alpha channel and side metadata are generated to indicate layer indication. This is an example of providing texture information associated with the image to the user device in the first bitstream and providing the at least one layer indication in the alpha channel of the first bitstream.

[0163] In the example scenario, the resulting video stream embeds an alpha channel and / or a depth channel to carry layered information. This can be done by adding auxiliary channels to the Simple Efficient Video Coding (HEVC) or Universal Video Coding (VVC) channels.

[0164] In another embodiment, layer indicators are encoded as a range of depth values ​​in a video stream representing depth information of the rendered view. The range of values ​​assigned to the layer indicators in the depth video is negotiated before the rendering session begins. Each pixel in a depth frame falling within the layer indicator range describes the layer indicator of the matching pixel in the texture frame. This is an example of providing a user device with texture information associated with an image in a first bitstream, depth information associated with an image in a second bitstream, and said at least one layer indicator as at least one range of depth values ​​in the second bitstream.

[0165] exist Figure 8 In the example scenario shown, layer information can be encoded in the depth video by assigning a subset of depth values ​​to indicate the depth layer to which the corresponding pixel in the signal texture video belongs.

[0166] The indication of at least one range (i.e., at least one range of texture values ​​or at least one range of depth values) can be carried as one or more SEI messages in one or more bitstreams.

[0167] Texture information associated with an image includes the image's texture values ​​or texture data.

[0168] For example, SEI messages can be defined using the following syntax:

[0169] Table 2 It has the following semantics: The pcp_scene_layer_id indicates the position of the received layer in the layered rendering architecture. The server's compositing engine uses this information to determine at which layer level the layer should be rendered.

[0170] pcp_scene_correction_mode indicates the correction mode to be applied to the received frame. 0 indicates that no correction should be applied, 1 indicates that attitude correction should be applied, and 2 indicates that range-based adaptive correction should be used.

[0171] pcp_variable_correction indicates whether uniform correction should be applied to the received frame or dynamic per-region correction is expected. 0 indicates uniform correction should be applied, and 1 indicates dynamic correction should be applied.

[0172] pcp_correction_sensitivity indicates the strength of attitude correction applied to received frames.

[0173] pcp_number_custom_areas indicates how many custom areas exist in the received frame when pcp_variable_correction equals 1.

[0174] pcp_area_x[i], pcp_area_y[i], pcp_area_h[i], and pcp_area_w[i] indicate the (x, y) position, width, and height of the i-th region in the received frame, respectively.

[0175] pcp_area_correction_sensitivity[i] indicates the attitude correction intensity that should be applied to the i-th region in the received frame.

[0176] `pcp_layering_metric` indicates what metric is used to derive different layers within a texture when using range-based adaptive correction. 0 indicates the use of depth channel information, 1 indicates the use of alpha channel information, and 2 indicates the use of dynamic range codewords.

[0177] pcp_number_intervals indicates the number of intervals that describe the layer for the selected metric.

[0178] pcp_interval_upper_bound[i] indicates the upper bound of the i-th interval.

[0179] pcp_interval_correction_sensitivity[i] indicates the attitude correction intensity that should be applied to the i-th interval in the received frame.

[0180] Table 3 shows alternative example SEI messages.

[0181]

[0182] It has the following semantics: pcp_metric indicates what metric is used to extract different layers from the video texture. 0 indicates the use of the alpha range, and 1 indicates the use of the depth range.

[0183] pcp_n_intervals indicates the number of intervals that describe the stratification for the selected metric.

[0184] pcp_interval_upper_bound[i] indicates the upper bound of the i-th interval.

[0185] pcp_interval_correction_sensitivity[i] indicates the attitude correction intensity that should be applied to the i-th interval in the received frame. 0 indicates that attitude correction should not be applied, and other values ​​describe different degrees of attitude correction sensitivity.

[0186] The indication of the at least one range (i.e., at least one range of the texture values ​​or at least one range of the depth values) can be carried as one or more messages and transmitted via an out-of-band communication channel. The out-of-band communication channel may include one or more of the following: a user plane signaling channel, a control plane signaling channel, and an application layer negotiation channel.

[0187] In step 6, the UE decodes the one or more video streams, extracts a layering indication for the decoded texture information, and forwards the decoded frames and layering information to the XR runtime. This is an example of receiving at least one layering indication and decoding the at least one video stream at the user equipment, wherein decoding the at least one video stream includes determining texture information associated with an image from a three-dimensional spatial environment.

[0188] Step 6 may require the use of standard video decoding. If the SEI message method is used to transmit side metadata, the server's demultiplexer should be able to forward the SEI messages to the compositing engine for interpretation.

[0189] In step 7, pose correction is applied to the texture information as described by the layering indicator. This is an example of applying correction to texture information associated with an image based on the layering indicator. The at least one layering indicator may include parameters indicating at least one of the following: the layering location of the at least one layer, the sensitivity of the at least one layer to temporal distortion, or the sensitivity of the at least one layer to spatial distortion.

[0190] Step 7 can be implemented using any pose correction method, but for example, the sensitivity to time warp can be considered as an input parameter. Several different alternatives exist for pose correction, such as ATW, but pose correction can be defined as processing the received synthetic image or its components / layers to reflect the latest pose of the user equipment.

[0191] In step 8, the final pose-corrected texture image is displayed to the viewer.

[0192] Different layers can be combined based on received metadata, for example, via SEI messages.

[0193] In one example embodiment, the OpenXR runtime is used for composition, and Figure 2 The third layer is assigned an XrCompositionLayerQuad.

[0194] The ability of SRC and SRS to use multi-mode segmented rendering can be indicated in the segmented rendering configuration. For example, the following configuration parameters can be used. Additionally, renderingFlag is set to FLAG_ALPHA_BLENDING.

[0195]

[0196] In one example embodiment, one or more fields defined for the attitude correction parameters described above are transmitted using RTP header extension. In another example embodiment, the one or more fields are transmitted over a data channel. In this case, pcp_scene_layer_id is appropriately associated with the media ID or media description for each layer. This association can be defined using new parameters in the SDP, for example, using a=pcp_scene_layer_id under the media description for each layer. <id>Alternatively, this information can be shared as part of an application-specific segmented rendering description message.

[0197] Determining the at least one layer indication can be based on a rendering layer in the rendering pipeline, a depth template of the rendered image, or elements of the user interface (UI). UI elements can include, for example, static or planar elements in a three-dimensional spatial environment, such as, but not limited to, elements such as… Figure 7 , 8 And the menu, map, or instructions shown in Figure 9. In one example embodiment, layer indications can be derived from the rendering layer in the rendering pipeline by analyzing the depth template of the rendered view or by leveraging the rendering engine's perception of the visual elements of the UI in the view.

[0198] Figure 9 The scene being rendered can be divided into multiple layers, each with its own depth and sensitivity to time warping.

[0199] The number of layers can be a render-side decision depending on the scene composition and its overall sensitivity to pose correction. Typically, two to three layers are expected to contain enough information for pose correction. However, there are applications where the number of layers may be increased.

[0200] For reference Figure 5 The described method may include determining the number of the at least one layer based on negotiation during the rendering session setup.

[0201] The method may include providing a request from the server to the UE or receiving a request at the UE from the server to modify the number of the at least one layer.

[0202] In one example embodiment, the number of layers is negotiated by the UE and SRS when the segmented rendering session is established.

[0203] In another embodiment, during a session, the UE or SRS may request a change in the layer number based on, for example, changes in the UE or SRS's operating conditions or changes in the scene content.

[0204] In one example embodiment, the change in the number of layers is requested and acknowledged by the UE via an application-specific RTCP feedback message, and by the SRS via an RTP header extension message or via the WebRTC data channel protocol (message type set to "urn:3gpp:split-rendering:vN:xxxx"), where N can be an appropriate protocol version identifier and xxxx can be an appropriate message type identifier, such as "renderLayers".

[0205] The apparatus for a server may include components for receiving location information related to a three-dimensional spatial environment from a user device at the server, components for rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with at least one corresponding layer indication, components for encoding the at least one layer into at least one video stream, and components for providing the at least one video stream and the at least one layer indication to the user device.

[0206] The device may include a server, which may be a user server, or may be contained in a server or in a chipset for performing at least some of the server's actions.

[0207] The apparatus for a user equipment may include components for providing location information related to a three-dimensional spatial environment from the user equipment to a server, components for receiving at least one video stream and at least one layering indication associated with the at least one video stream from the server at the user equipment, components for decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment, and components for applying correction to the texture information associated with the image based on the layering indication.

[0208] The device may include user equipment, such as a mobile phone, which is user equipment or is included in a user equipment or a chipset for performing at least some of the actions of the user equipment.

[0209] It should be understood that the device may include or be coupled to other units or modules, such as radio components or radio heads used for transmission and / or reception. Although the device has been described as a single entity, different modules and memories may be implemented in one or more physical or logical entities.

[0210] Note that while some embodiments have been described in conjunction with 5G networks, similar principles can be applied to other networks and communication systems, such as 6G networks or 5G-Advanced networks. Therefore, although some embodiments have been described above by way of example with reference to certain example architectures of wireless networks, technologies, and standards, these embodiments can be applied to any other suitable form of communication system besides those described herein.

[0211] It should also be noted that although exemplary embodiments have been described above, several changes and modifications can be made to the disclosed solutions without departing from the scope of the invention.

[0212] As used herein, "at least one of the following: " and "at least one of " and similar wording, where the list of two or more elements is connected by "and" or "or", means at least any one element, or at least any two or more elements, or at least all elements.

[0213] Generally, various embodiments can be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of this disclosure can be implemented in hardware, while others can be implemented using firmware or software, which can be executed by a controller, microprocessor, or other computing device, although this disclosure is not limited thereto. While various aspects of this disclosure may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be fully understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0214] As used in this application, the term "circuit" may refer to one or more or all of the following: (a) Hardware circuit implementation only (e.g., implementation only in analog and / or digital circuits) and (b) A combination of hardware circuitry and software, such as (if applicable): (i) a combination of analog and / or digital hardware circuitry and software / firmware, and (ii) Any part of the hardware processor and software (including digital signal processors), software and memory, which work together to enable a device (e.g., a mobile phone or server) to perform various functions. (iii) Hardware circuitry and / or processors, such as microprocessors or a portion thereof, which require software (e.g., firmware) to operate, but which may be absent when not in use.

[0215] This definition of circuit system applies to all uses of the term in this application, including in any claim. As another example, as used herein, the term circuit system also covers only hardware circuitry or a processor (or processors) or a portion thereof and its accompanying software and / or firmware implementation. The term circuit system also covers, for example and if applicable to a particular claim element, baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices.

[0216] Embodiments of this disclosure can be implemented via computer software executable by a mobile device's data processor, such as in a processor entity, or via hardware, or via a combination of software and hardware. Computer software or programs, also known as program products, include software routines, applets, and / or macros, and can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer-executable components that, when the program runs, are configured to execute the embodiments. The one or more computer-executable components may be at least one piece of software code or a portion thereof.

[0217] Furthermore, it should be noted that any block of the logical flow in the accompanying drawings may represent a program step, or interconnected logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. Software may be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc. These physical media are non-transitory media. As used herein, the term "non-transitory" is a limitation on the medium itself (i.e., tangible, not signaling), not a limitation on the persistence of data storage (e.g., RAM and ROM).

[0218] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include one or more of general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), FPGAs, gate-level circuits, and processors based on multi-core processor architectures, as non-limiting examples.

[0219] Embodiments of this disclosure can be implemented in various components such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Complex and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready for etching and formation on semiconductor substrates.

[0220] The scope of protection sought by the various embodiments of this disclosure is set forth in the independent claims. Embodiments and features described in this specification that are not within the scope of the independent claims, if any, should be interpreted as examples that aid in understanding the various embodiments of this disclosure.

[0221] The foregoing description provides a complete and informative description of exemplary embodiments of the present disclosure by way of non-limiting example. However, various modifications and adjustments will become apparent to those skilled in the art when the foregoing description is read in conjunction with the accompanying drawings and appended claims. Nevertheless, all such and similar modifications to the teachings of this disclosure will still fall within the scope of the invention as defined in the appended claims. Indeed, there are other embodiments that include combinations of one or more embodiments with any other embodiments discussed above.< / id>

Claims

1. An apparatus for a server, comprising: A component for receiving location information related to the three-dimensional spatial environment from a user device at the server; A component for rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with at least one corresponding layering indicator; Components for encoding the at least one layer into at least one video stream; as well as A component for providing the user equipment with the at least one video stream and the at least one layer indication.

2. The apparatus of claim 1, wherein the position information includes at least one of the following: one or more orientation parameters; or one or more position parameters, wherein the one or more orientation parameters and the one or more position parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

3. The apparatus according to any one of claims 1 to 2, further comprising a component for determining the at least one layer indication based on at least one of the following: a rendering layer in the rendering pipeline, a depth template of the rendered image, or an element of the user interface.

4. The apparatus according to any one of claims 1 to 3, wherein the component for providing includes components for providing texture information associated with the image to the user equipment in a first bitstream and for providing the at least one layer indication in the alpha channel of the first bitstream.

5. The apparatus according to any one of claims 1 to 3, wherein the component for providing includes a component for providing texture information associated with the image to the user equipment in a first bitstream, and wherein the component for providing the at least one video stream and the at least one layer indication includes a component for providing the at least one layer indication to the user equipment in a second bitstream.

6. The apparatus according to any one of claims 1 to 3, wherein the component for providing includes a component for providing the following to the user equipment: Texture information associated with the image in the first bitstream, The depth information associated with the image in the second bitstream, and The at least one layer indication in the second bitstream as at least one range of depth values.

7. The apparatus according to any one of claims 1 to 3, wherein the component for providing includes a component for providing the following to the user equipment: Texture information associated with the image in the first bitstream and at least one layer indication as at least one range of texture values ​​in the first bitstream.

8. The apparatus according to any one of claims 6 or 7, wherein the indication of the at least one range is carried as one or more SEI messages in one or more bit streams.

9. The apparatus according to any one of claims 1 to 8, wherein the indication of the at least one range is carried as one or more messages and transmitted via an out-of-band communication channel.

10. The apparatus of claim 9, wherein the out-of-band communication channel comprises one or more of the following: a user plane signaling channel, a control plane signaling channel, or an application layer negotiation channel.

11. The apparatus according to any one of claims 1 to 10, further comprising a component for determining the number of the at least one layer based on negotiation during the rendering session establishment.

12. The apparatus of claim 11, further comprising components for providing a request to modify the number of the at least one layer or for receiving a request to modify the number of the at least one layer.

13. The apparatus according to any one of claims 1 to 12, wherein the at least one layer indication includes at least one parameter indicating at least one of the following: the hierarchical position of the at least one layer, the sensitivity of the at least one layer to time distortion, or the sensitivity of the at least one layer to spatial distortion.

14. An apparatus for a user equipment, comprising: Components for providing location information related to the three-dimensional spatial environment from the user equipment to the server; A component for receiving at the user equipment at the server at least one video stream and at least one layer indication associated with the at least one video stream; A component for decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment; as well as A component for applying correction to the texture information associated with the image based on the layering indication.

15. The apparatus of claim 14, wherein the position information includes at least one of: one or more orientation parameters; or one or more position parameters, wherein the one or more orientation parameters and the one or more position parameters are related to the viewing position of the image relative to the three-dimensional spatial environment.

16. The apparatus according to any one of claims 14 to 15, wherein the receiving component includes a component for receiving the texture information associated with the image in the first bitstream at the user equipment and receiving the at least one layer indication in the alpha channel of the first bitstream.

17. The apparatus of any one of claims 14 to 15, wherein the receiving component includes a component for receiving the texture information associated with the image in a first bitstream at the user equipment and for receiving the at least one layer indication directed to the user equipment in a second bitstream.

18. The apparatus according to any one of claims 14 to 15, wherein the receiving component includes a component for receiving the following at the user equipment: Texture information associated with the image in the first bitstream, The depth information associated with the image in the second bitstream, and The at least one layer indication is at least one range of depth values ​​in the second bitstream.

19. The apparatus according to any one of claims 14 to 15, wherein the receiving component includes a component for receiving the following at the user equipment: Texture information associated with the image in the first bitstream, and The at least one layer indication serves as at least one range of texture values ​​in the first bitstream.

20. A method for using a server, the method comprising: The server receives location information related to the three-dimensional spatial environment from the user equipment. Rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with at least one corresponding layering indicator; Encode the at least one layer into at least one video stream; as well as Provide the user equipment with the at least one video stream and the at least one layering indication.

21. A method for a user equipment, the method comprising: The user equipment provides the server with location information related to the three-dimensional spatial environment; At the user equipment, at least one video stream and at least one layering indication associated with the at least one video stream are received from the server. Decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment; and Correction is applied to the texture information associated with the image based on the layering indication.

22. An apparatus for a server, comprising: At least one processor, and at least one memory storing instructions, which, when executed by the processor, cause the device to at least: The server receives location information related to the three-dimensional spatial environment from the user equipment. Rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with at least one corresponding layering indicator; Encode the at least one layer into at least one video stream; as well as Provide the user equipment with the at least one video stream and the at least one layering indication.

23. An apparatus for a user equipment, comprising: At least one processor, and at least one memory storing instructions, wherein when the instructions are executed by the processor, the means causes the device to at least: The user equipment provides the server with location information related to the three-dimensional spatial environment; At the user equipment, at least one video stream and at least one layering indication associated with the at least one video stream are received from the server. Decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment; and Correction is applied to the texture information associated with the image based on the layering indication.

24. A computer-readable medium comprising instructions that, when executed by a means for a server, cause the means to perform at least the following operations: The server receives location information related to the three-dimensional spatial environment from the user equipment. Rendering an image from the three-dimensional spatial environment based on the location information, wherein rendering the image based on the location information includes generating at least one layer, each at least one layer being associated with at least one corresponding layering indicator; Encode the at least one layer into at least one video stream; as well as Provide the user equipment with the at least one video stream and the at least one layering indication.

25. A computer-readable medium comprising instructions, when executed by a means for a user equipment, causing the means to perform at least the following operations: The user equipment provides the server with location information related to the three-dimensional spatial environment; At the user equipment, at least one video stream and at least one layering indication associated with the at least one video stream are received from the server. Decoding the at least one video stream, wherein decoding the at least one video stream includes determining texture information associated with an image from the three-dimensional spatial environment; and Correction is applied to the texture information associated with the image based on the layering indication.