Segmentation rendering method and device based on UE estimation
The UE estimates and sends future pose information, the server performs rendering, and the UE corrects it, ultimately solving the problem of inaccurate pose estimation in segmentation rendering and improving the accuracy of rendering results and user experience.
Patent Information
- Application Number
- CN202480011861.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-10
- Filing Date
- 2024-01-24
- Publication Date
- 2025-09-05
AI Technical Summary
During the segmentation rendering process, it is difficult for the user equipment (UE) to accurately estimate the pose, resulting in a mismatch between the rendering results and the actual display time, affecting the user experience.
The UE estimates its pose information at a future time point and sends it to the server. The server renders based on this information and returns the result. The UE then performs the final pose correction to ensure the match.
Improves the accuracy of rendering results and user experience, reducing interruption time when screen content does not match expected content.
Smart Images

Figure CN120604502A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and apparatus for providing segmentation rendering based on User Equipment (UE) estimation in a communication system. Background Art
[0002] Fifth-generation (5G) mobile communication technology defines wide frequency bands, enabling high transmission rates and new services. 5G mobile communication technology can be implemented not only in frequency bands "below 6 GHz," such as 3.5 GHz, but also in frequency bands "above 6 GHz," known as millimeter waves (mmWave), including 28 GHz and 39 GHz. Furthermore, to achieve transmission rates fifty times faster than 5G mobile communication technology and ultra-low latency one-tenth that of 5G mobile communication technology, 6G mobile communication technology (referred to as a "beyond 5G system") is being considered for implementation in the terahertz band (e.g., the 95 GHz to 3 THz band).
[0003] In the initial stage of 5G mobile communication technology, in order to support services associated with enhanced Mobile Broadband (eMBB), Ultra-Reliable & Low Latency Communications (URLLC) and massive Machine-Type Communications (mMTC) and meet the performance requirements associated therewith, standardization is underway on the following items: beamforming and massive MIMO for mitigating radio wave path loss and increasing radio wave transmission range in millimeter waves, parameter sets (numerology) for dynamic operation of time slot formats for efficient utilization of millimeter wave resources and timeslot formats (for example, operation of multiple subcarrier spacings), initial access technology supporting multi-beam transmission and broadband, definition and operation of BWP (bandwidth part), new channel coding methods such as LDPC (low-density parity check) codes for large-capacity data transmission and polar codes for highly reliable transmission of control information, L2 preprocessing, and network slicing for providing specialized networks tailored to specific services.
[0004] Currently, in view of the services to be supported by 5G mobile communication technologies, discussions are underway on improvements and performance enhancements to initial 5G mobile communication technologies, and there is already physical layer standardization on technologies such as Vehicle-to-Everything (V2X) for assisting driving determination of autonomous vehicles based on information sent by vehicles about the location and status of vehicles and for enhancing user convenience, New Radio Unlicensed (NR-U) for system operation in unlicensed frequency bands that complies with various regulatory requirements, NR UE power saving, Non-Terrestrial Network (NTN) as UE-satellite direct communication for ensuring coverage in areas where communication with terrestrial networks is unavailable, and positioning.
[0005] Furthermore, in the area of radio interface architecture / protocols, standardization is underway on technologies such as the Industrial Internet of Things (IIoT), which supports new services through interworking and integration with other industries; Integrated Access and Backhaul (IAB), which provides nodes for expanding network service areas by integrating wireless backhaul and access links; mobility enhancements including conditional handover and Dual Active Protocol Stack (DAPS) handover; and two-step random access (NR two-step RACH) for simplifying the random access procedure. In the area of system architecture / services, standardization is also underway on a 5G baseline architecture (e.g., a service-based architecture or service-based interface) for combining Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies; and Mobile Edge Computing (MEC), which allows for receiving services based on the location of the UE.
[0006] If such a 5G mobile communication system is commercialized, the already exponentially growing number of connected devices will be connected to the communication network, and it is therefore expected that enhanced functionality and performance of the 5G mobile communication system and the integrated operation of connected devices will be necessary. To this end, new research is being planned related to the following: efficient support for extended reality (XR) (XR = AR + VR + MR) for augmented reality (AR), virtual reality (VR), mixed reality (MR), etc., 5G performance improvement and complexity reduction through the use of artificial intelligence (AI) and machine learning (ML), AI service support, metaverse service support, and drone communication.
[0007] Furthermore, such advancements in 5G mobile communication systems will serve not only as a foundation for the development of new waveforms, Full Dimensional MIMO (FD-MIMO), multi-antenna transmission technologies (such as array antennas and massive antennas) for ensuring coverage in the terahertz band for 6G mobile communication technology, metamaterial-based lenses and antennas for improving coverage of terahertz band signals, high-dimensional spatial multiplexing technologies using orbital angular momentum (OAM), and reconfigurable intelligent surfaces (RIS), but will also serve as a foundation for the development of full-duplex technologies for improving the frequency efficiency and system networks of 6G mobile communication technology, AI-based communication technologies for achieving system optimization by leveraging satellites and AI (artificial intelligence) from the design stage and internalizing end-to-end AI support functions, and next-generation distributed computing technologies for implementing services with a complexity level that exceeds the operational capabilities of UEs by utilizing ultra-high-performance communication and computing resources.
[0008] Split rendering may include an operation in which a device, such as a server, performs a rendering process on behalf of a user equipment (UE) and transmits a rendering result of the rendering process to the UE. In split rendering, the server may perform the rendering process using information about the content to be rendered and viewpoint information (e.g., pose and / or field of view) provided by the UE.
[0009] The above information is presented as background information only to assist with an understanding of the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above might be applicable as prior art with respect to the present disclosure. Summary of the Invention
[0010] Aspects of the present disclosure are to address at least the above-mentioned problems and / or disadvantages and to provide at least the advantages described below. Therefore, one aspect of the present disclosure is to provide a method and apparatus for providing split rendering for computing power allocation between a user equipment (UE) and a server in a communication system.
[0011] Another aspect of the present disclosure is to provide a method and apparatus for performing calculation (eg, rendering process) on a server based on an estimated pose for segmented rendering provided by a UE.
[0012] Another aspect of the present disclosure is to provide a method and apparatus for correcting a final pose to be applied to a result to be displayed based on a calculation result (eg, a rendering result) received from a server.
[0013] Another aspect of the present disclosure is to define components to be used for segmented rendering in a communication system, and the time required for the components can be measured or estimated.
[0014] Another aspect of the present disclosure is to provide a method and apparatus for transmitting information about a time required to split rendered components between a UE and a server.
[0015] Another aspect of the present disclosure is to provide a method and apparatus for indicating or recommending a segmentation rendering operation (eg, pose selection) to a server.
[0016] Another aspect of the present disclosure is to provide a method and apparatus for notifying a UE of information related to a split rendering operation (eg, pose selection) of a server.
[0017] Additional aspects will be set forth in part in the description which follows and, in part, will be apparent from the description, or may be learned by practice of the presented embodiments.
[0018] Technical Solution
[0019] According to one aspect of the present disclosure, a method performed by a user equipment (UE) supporting split rendering in a communication system is provided. The method includes estimating, at a first time point, a pose of the UE at a second time point, the pose indicating a target display time of a first media frame, sending the first media frame and first pose information related to the estimated pose to a server, receiving, from the server, a second media frame generated by rendering based on the first media frame and the estimated pose and second pose information related to the second media frame, generating a third media frame by correcting the second media frame based on metadata and an actual pose of the UE, and displaying the third media frame at a third time point.
[0020] According to another aspect of the present disclosure, a method performed by a server supporting split rendering in a communication system is provided. The method includes receiving a first media frame and first pose information related to an estimated pose of a user equipment (UE) from the UE, generating a second media frame by rendering based on the first media data and the estimated pose, and sending the second media frame and second pose information related to the second media frame to the UE.
[0021] According to another aspect of the present disclosure, a UE for supporting split rendering in a communication system is provided. The UE includes a transceiver, a memory, and a processor coupled to the transceiver and the memory, wherein the memory stores one or more computer programs including computer-executable instructions, which, when executed by the processor, cause the UE to estimate a pose of the UE at a second time point indicating a target display time of a first media frame at a first time point, send the first media frame and first pose information related to the estimated pose to a server, receive from the server a second media frame generated by rendering based on the first media frame and the estimated pose, and second pose information related to the second media frame, generate a third media frame by correcting the second media frame based on metadata and an actual pose of the UE, and display the third media frame at a third time.
[0022] According to another aspect of the present disclosure, a server for supporting split rendering in a communication system is provided. The server includes a network interface, a memory, and a processor coupled to the network interface and the memory, wherein the memory stores one or more computer programs including computer-executable instructions, which, when executed by the processor, cause the server to receive a first media frame and first pose information associated with an estimated pose of the UE from a UE, generate a second media frame by rendering based on the first media data and the estimated pose, and send the second media frame and second pose information associated with the second media frame to the UE.
[0023] According to another aspect of the present disclosure, one or more non-transitory computer-readable storage media storing computer-executable instructions are provided, which, when executed by one or more processors of a user equipment (UE), cause the UE to perform operations. The operations include: estimating a pose of the UE at a second time point indicating a target display time of a first media frame at a first time point, sending the first media frame and first pose information associated with the estimated pose to a server, receiving from the server a second media frame generated by rendering based on the first media frame and the estimated pose and second pose information associated with the second media frame, generating a third media frame by correcting the second media frame based on metadata and an actual pose of the UE, and displaying the third media frame at a third time point.
[0024] Other aspects, advantages, and salient features of the present disclosure will become apparent to those skilled in the art from the following detailed description, which, taken in conjunction with the accompanying drawings, discloses various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0026] Figure 1 shows the structure of a communication system according to an embodiment of the present disclosure;
[0027] Figure 2 shows pose correction of segmentation calculation according to an embodiment of the present disclosure;
[0028] Figure 3a shows the result of incorrect estimation based on the server according to an embodiment of the present disclosure;
[0029] Figure 3b shows the result of an incorrect estimation based on a server according to an embodiment of the present disclosure;
[0030] Figure 4 shows the results of incorrect estimation based on the server according to an embodiment of the present disclosure;
[0031] Figure 5 The structure of a segmentation rendering system based on user equipment (UE) estimation according to an embodiment of the present disclosure is shown;
[0032] Figure 6 Components of each stage of segmentation rendering according to an embodiment of the present disclosure are shown;
[0033] Figure 7 is a sequence diagram illustrating components of each stage of segmented rendering according to an embodiment of the present disclosure;
[0034] Figure 8 shows the information structure of a pose set according to an embodiment of the present disclosure;
[0035] Figure 9 shows the information structure of a pose pair according to an embodiment of the present disclosure;
[0036] Figure 10 shows the information structure of SR metadata according to an embodiment of the present disclosure;
[0037] Figure 11 is a sequence diagram illustrating the updating of a pose set according to an embodiment of the present disclosure;
[0038] Figure 12 is a sequence diagram illustrating correction of an estimated pose according to an embodiment of the present disclosure;
[0039] Figure 13 is a sequence diagram illustrating the reversal of the order of estimated poses according to an embodiment of the present disclosure;
[0040] Figure 14 is a sequence diagram illustrating a process for split rendering between a UE and a server according to an embodiment of the present disclosure;
[0041] Figure 15a and Figure 15b is a sequence diagram illustrating a process for split rendering between a UE and a server according to various embodiments of the present disclosure;
[0042] Figure 16 is a conceptual diagram illustrating an Internet Protocol (IP) packet structure including a media frame in a wireless communication system according to an embodiment of the present disclosure;
[0043] Figure 17 A real-time transport protocol (RTP) header extension structure including a pose pair according to an embodiment of the present disclosure is shown;
[0044] Figure 18 shows a PoseData field included in an RTP header extension according to an embodiment of the present disclosure;
[0045] Figure 19 shows an SRMetaData field included in an RTP header extension according to an embodiment of the present disclosure;
[0046] Figure 20 is a block diagram showing a configuration of a UE in a communication system according to an embodiment of the present disclosure; and
[0047] Figure 21 is a block diagram illustrating a configuration of a server in a wireless communication system according to an embodiment of the present disclosure.
[0048] Throughout the drawings, it should be noted that like reference numbers are used to depict the same or similar elements, features, and structures. DETAILED DESCRIPTION
[0049] The following description, with reference to the accompanying drawings, is provided to facilitate a fuller understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. It includes various specific details to aid understanding, but these details are to be considered merely as exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. Furthermore, descriptions of well-known functions and structures may be omitted for clarity and conciseness.
[0050] The terms and words used in the following description and claims are not limited to the bibliographical meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.
[0051] It should be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more of such surfaces.
[0052] When describing the embodiments of the present disclosure, descriptions related to technical contents that are well known in the art and not directly related to the present disclosure will be omitted. The omission of such unnecessary descriptions is intended to prevent the main idea of the present disclosure from being obscured and to shift the main idea more clearly. In addition, the terms described below are terms defined based on the functions in the present disclosure and may vary depending on the user, the user's intention, or habits. Therefore, the definition of terms should be based on the content throughout the specification.
[0053] For the same reason, in the accompanying drawings, some elements may be exaggerated, omitted or schematically shown. In addition, the size of each element does not fully reflect the actual size. In the accompanying drawings, the same or corresponding elements are provided with the same reference numerals.
[0054] By reference to the embodiments described below in conjunction with the accompanying drawings, the advantages and features of the present disclosure and the ways of achieving them will be apparent. However, the present disclosure is not limited to the embodiments set forth below, but can be implemented in various different forms. The following embodiments are provided only to fully disclose the present disclosure and to inform those skilled in the art of the scope of the present disclosure, and the present disclosure is limited only by the scope of the appended claims. Throughout the specification, the same or similar figure marks represent the same or similar elements. In addition, in the description of the present disclosure, when it is determined that the description may make the subject matter of the present disclosure unnecessarily unclear, the detailed description of the known functions or configurations incorporated herein will be omitted. The terms to be described below are terms based on functional definitions in the present disclosure and may be different according to the user, the user's intention or custom. Therefore, the definition of terms should be based on the content in the entire specification.
[0055] In the following description, a base station (BS) is an entity that allocates resources to a terminal and may be at least one of a gNode B, an eNode B, a Node B (or an xNode B, where x is an alphabetic alphabet consisting of "g" and "e"), a wireless access unit, a base station controller, a satellite, an onboard device, and a node on a network. User equipment (UE) may include a mobile station (MS), a cellular phone, a smartphone, a computer, or a multimedia system capable of performing communication functions. In this disclosure, a "downlink (DL)" refers to a radio link via which a base station transmits signals to a terminal, and an "uplink (UL)" refers to a radio link via which a terminal transmits signals to a base station. Additionally, a "sidelink (SL)" may exist, which refers to a radio link via which a UE transmits signals to another UE.
[0056] In addition, in the following description, LTE, LTE-A, or 5G systems may be described by way of example, but the embodiments of the present disclosure may also be applied to other communication systems with similar technical backgrounds or channel types. Examples of such communication systems may include advanced 5G, advanced NR, or 6th generation (5G) mobile communication technology developed beyond 5G mobile communication technology (or new radio, NR), and in the following description, "5G" may be a concept covering existing LTE, LTE-A, or other similar services. In addition, based on the determination of those skilled in the art, the embodiments of the present disclosure may also be applied to other communication systems with some modifications without significantly departing from the scope of the present disclosure.
[0057] In this document, it will be understood that each block of the flowchart diagram and the combination of blocks in the flowchart diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions specified in one or more flowchart blocks. These computer program instructions can also be stored in a computer-usable or computer-readable memory, which can instruct the computer or other programmable data processing device to act in a specific manner so that the instructions stored in the computer-usable or computer-readable memory produce an article of manufacture including an instruction device that implements the functions specified in the flowchart block or multiple blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are performed on the computer or other programmable device to produce a computer-implemented process so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flowchart blocks.
[0058] In addition, each block of the flowchart diagram may represent a module, segment or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative embodiments, the functions marked in the blocks may not occur in order. For example, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functions involved.
[0059] As used in the embodiments of the present disclosure, the term "unit" refers to a software element or hardware element that performs a predetermined function, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). However, the term "unit" is not always limited to software or hardware. A "unit" can be configured to be stored in an addressable storage medium or executed by one or more processors. Therefore, a "unit" includes, for example, software elements, object-oriented software elements, class elements or task elements, processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and parameters. The elements and functions provided by a "unit" can be combined into a smaller number of elements or "units" or divided into a larger number of elements or "units." Furthermore, elements and "units" can be implemented as one or more CPUs within a playback device or secure multimedia card. Furthermore, a "unit" in the embodiments may include one or more processors.
[0060] It should be understood that the blocks in each flowchart and the combination of the flowcharts can be executed by one or more computer programs comprising computer executable instructions. The entirety of one or more computer programs can be stored in a single memory, or one or more computer programs can be divided into different parts stored in different multiple memories.
[0061] Any functions or operations described herein may be processed by a processor or a combination of processors. A processor or a combination of processors is a circuit that performs processing and includes circuits such as an application processor (AP, e.g., a central processing unit (CPU)), a communication processor (CP, e.g., a modem), a graphics processing unit (GPU), a neural processing unit (NPU) (e.g., an artificial intelligence (AI) chip), a wireless fidelity (Wi-Fi) chip, a Bluetooth™ chip, a global positioning system (GPS) chip, a near-field communication (NFC) chip, a connectivity chip, a sensor controller, a touch controller, a fingerprint sensor controller, a display driver integrated circuit (IC), an audio codec chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on a chip (SoC), an IC, and the like.
[0062] Figure 1The structure of a communication system according to an embodiment of the present disclosure is shown.
[0063] refer to Figure 1 , a communication system may include at least one terminal (UE) (e.g., UE 120) capable of accessing network 100. Network 100 may include one or more nodes (e.g., at least one base station and at least one core network (CN) node) that support wireless access for UE 120. A base station may be referred to as an "access point (AP)", "eNodeB (eNB)", "gNodeB (gNB)", "fifth generation node (5G node)", "radio point", "transmission / reception point (TRP)", or any other term having equivalent technical meaning.
[0064] UE 120 is a device that can be used by a user to communicate over a wireless channel. According to an embodiment of the present disclosure, UE 120 is a device that performs machine-type communication (MTC) and may not be carried by a user. UE 120 may be referred to as a "user equipment (UE)," "mobile station," "subscriber station," "remote UE," "wireless terminal," "user device," or any other term with an equivalent technical meaning.
[0065] The UE 120 may include a split rendering client and may be connected to a server 110 providing split rendering (eg, a split rendering (SR) edge application service (EAS) server) through a network 100 .
[0066] Split rendering is a technology for allocating computing power between UE 120 and server 110. When the required performance of an application or content to be executed on UE 120 (e.g., content reproduction) is relatively high compared to the performance of UE 120, UE 120 may generate factors / parameters for executing the content or application and deliver the factors / parameters to server 110. Server 110 may execute the content or application based on the received factors / parameters and then deliver the result (e.g., rendering result) to UE 120.
[0067] Depending on the usage scenario, a time difference may occur between the computation time required for server 110 to execute content or an application, the time required to transmit factors (e.g., pose information), and / or the time required to transmit computation results (e.g., media frames). For example, in augmented reality (AR), UE 120 may receive a result obtained by executing on server 110 at a second time point using the spatial "position and orientation (e.g., direction)" (hereinafter referred to as pose) of UE 120 at a first time point as a factor. The greater the gap between the first time point and the second time point, the greater the degree of nausea felt by the user.
[0068] Figure 2Shown is the pose correction of the segmentation calculation according to an embodiment of the present disclosure.
[0069] refer to Figure 2 , the UE 120 may perform an operation of estimating a second time point (t2) 204 at a first time point (t1) 202, an operation of estimating a pose of the UE 120 at the estimated second time point (t2) 204, and an operation of sending the estimated pose to the server 110. After the server 110 performs server calculation (e.g., rendering), the UE 120 may perform an operation of receiving a calculation result (e.g., a calculation result or a rendering result) obtained by factoring the estimated pose from the server 110, and an operation of performing UE calculation (e.g., final pose correction) on the received result based on a difference between the pose of the UE at the second time point 204, which has been estimated at the first time point 202, and the actual pose measured at the second time point 204.
[0070] The time it takes for server 110 to perform rendering based on the user's pose for a rendering result to be displayed on the display of UE 120 and perceived by the user's eyes is referred to as pose-to-rendering-to-photon (P2R2P) latency. Although "photon" primarily describes the phenomenon of a rendering result being displayed on a display and perceived by the user's eyes, the point in time when the user perceives the rendering result using a device other than a display (e.g., an audio device or a haptic device) can also be expressed as a "photon time point."
[0071] Estimating the pose of UE 120 may include the aforementioned operations of estimating the second time point and estimating the pose of UE 120 at the second time point. The period between the first time point and the second time point may include at least one of the following: the time UE 120 transmits the estimated pose; the computation time required for server 110 to execute the content or application using the received pose as a factor; the transmission time required for the result generated as a result of the computation to be transmitted from server 110 to UE 120; or the UE computation time required for UE 120 to correct the result using the final pose at the second time point. The second time point may be determined based on one or more combinations of the performance of UE 120, the transmission performance of the wireless communication network (e.g., network 100) to which UE 120 has accessed, the number of other UEs simultaneously accessing network 100 or server 110, the complexity of the content and / or application (hereinafter referred to as content / application) to be executed by UE 120, or the resource performance of a server computing instance allocated by server 110 to provide the split rendering service. Furthermore, for the same combination, the second time point may be determined based on at least one of the movement of the user and UE 120, changes in the complexity of the content / application selected by the user, or changes in computing time due to changes in complexity. Therefore, the second time point estimated by UE 120 at the first time point and the second time point that actually occurs or is actually displayed are very likely different. In other words, accurately estimating the second time point may be difficult.
[0072] Figure 3a Results based on incorrect estimation by a server according to an embodiment of the present disclosure are shown.
[0073] Figure 3b Results based on incorrect estimation by a server according to an embodiment of the present disclosure are shown.
[0074] Figure 4 Results based on incorrect estimation by a server according to an embodiment of the present disclosure are shown.
[0075] refer to Figure 3a , the second time point 304a estimated by the UE 120 may be earlier than the actually occurring second time point 304. For example, although the UE 120 has estimated a time 80 ms after the first time point 302 as the second time point 304a, when the actually occurring second time point 304 after performing the split rendering operations (e.g., transmission, server calculation, reception, and UE calculation) occurs 100 ms after the first time point 302, the content that the UE 120 may display to the user at the actually occurring second time point 304 may be a result regarding the time that has already passed (e.g., the estimated second time point 304a).
[0076] refer to Figure 3b , the UE 120 may need another 100ms 306a to send the pose to the server 110 again and receive the result, and the UE 120 may only reproduce the usable result at a time point 306 that is 200ms later from the time point 302 at which the second time point 304a has been estimated for the first time.
[0077] refer to Figure 4 , the second time point 406 estimated by the UE 120 may be later than the actually occurring second time point 404. For example, although the UE 120 has estimated a time 120 ms after the first time point 402 as the second time point 406, when the actually occurring second time point 404 after performing the split rendering operations (e.g., transmission, server calculation, reception, and UE calculation) occurs 100 ms after the first time point 402, the content that the UE 120 can display to the user at the actually occurring second time point 404 may not yet be available or may correspond to future content.
[0078] Although UE 120 may wait for an additional 20 ms from the actual second time point 404 for reproducing the result provided by server 110, the later the estimated second time point 406 is than the actual second time point 404, the larger the pose estimation error may be, and therefore, the result may not match the actual pose at the second time point 404, resulting in degradation of content quality based on the final pose correction.
[0079] In the prior art, although the UE 120 can estimate the second point in time based on past statistical records, the above-mentioned problem caused by the difference between the estimated second point in time and the second point in time that actually occurs may result in an interruption duration (e.g., downtime) during which the content on the screen does not match the expected content.
[0080] Since segmented rendering may operate based on statistical records managed by the UE 120 itself, it is necessary to define the time components that exist between the UE 120 and the server 110, as well as the exchange of estimates of the time components and actual calculation times.
[0081] Although the number of poses estimated by UE 120, the number of poses actually processed by server 110, and the number of frames per second at which UE 120 displays the results of server 110 may be different, server 110 does not know which pose among the poses sent by UE 120 will be used to execute the content or application, and the UE does not know which pose has been used to generate the results of server 110. For example, UE 120 may acquire 1000 to 4000 poses per second via at least one sensor, and the number of poses estimated by UE 120 and the number of poses sent from UE 120 to server 110 may be 1000 or more. However, the number of frames per second of the video that can be reproduced by UE 120 may be, for example, 60 frames, and therefore the number of poses processed by server 110 may also be 60. In the prior art, UE 120 cannot request server 110 to select a pose from 1000 poses per second to process 60 frames per second. In the prior art, when the server 110 has selected and processed 60 frames among 1000 poses per second, the server 110 has been unable to inform the UE 120 of the pose for each frame.
[0082] In the following embodiments of the present disclosure, segmented rendering may include a UE estimating a pose, a server performing calculations (e.g., rendering) based on the estimated pose sent from the UE, the UE receiving the server's calculation results (e.g., rendering results), and the UE correcting the calculation results based on the final pose and outputting the corrected calculation results to a display. The following embodiments may define components for configuring each stage of segmented rendering and measure the time required for the components of each stage. The following embodiments may also transmit information about the time required for the components between the UE and the server. The following embodiments may also allow the UE to instruct or recommend server operations (e.g., pose selection).
[0083] Figure 5 The structure of a segmentation rendering system based on UE estimation according to an embodiment of the present disclosure is shown.
[0084] refer to Figure 5 , the UE 500 may generate information (e.g., pose) about the absolute or relative position of the UE 500 in the space in which the UE 500 is located and the orientation (e.g., direction) that the UE 500 is facing from at least one sensor 502 (e.g., a motion sensor and / or a camera) included internally or externally.
[0085] The pose generated at each time instant (e.g., input pose) may be applied as input to a segmentation rendering (SR) estimator 504. The SR estimator 504 may include a time estimator and a pose estimator, wherein the time estimator estimates a second time point (T2), where T2 is a target display time, and the pose estimator generates an estimated pose at which the UE 500 will be positioned at time T2 estimated by the time estimator based on the input pose.
[0086] The estimated pose generated by the SR estimator 504 can be stored in a pose buffer 506 before being sent to a server 530 (eg, a segmented rendering server). A target display time and an estimated pose for the target display time are referred to as a pose pair (eg, a pose and metadata pair).
[0087] The UE pose manager 508 (e.g., a client pose manager) may bundle poses (e.g., pose pairs) already input and stored in the pose buffer 506 into at least one pose set and transmit the pose set to the server 530. The pose sets may be transmitted periodically or aperiodically, as determined by the UE pose manager 508. At least one pose set in a transmission may include estimated poses for a duration that does not overlap with at least one pose set included in a previous transmission, or may include estimated poses for a duration that at least partially overlaps with at least one pose set included in a previous transmission. For example, if an aspect of the user's movement changes from a previously estimated aspect, a new estimated pose relative to an expected pose in the already transmitted poses that has not yet been processed by the server 530 may be included in the pose set and retransmitted to the server 530.
[0088] The TA server pose manager 534 of the server 530 may store the poses included in the pose set received from the UE 500 in the server pose buffer 536, and may select at least one pose required at the time point when the renderer 542 starts new rendering from the poses stored in the server pose buffer 536. Regardless of the order in which the pose sets are received from the UE 500, the server pose manager 534 may recognize the order in which the UE 500 has sent the pose sets, and may compare the pose pairs already received and stored in the server pose buffer 536 with the pose pairs of the pose sets sent and received later, to update the pose pairs having the same target display time.
[0089] The SR manager of server 530 can perform at least one of the following through negotiation between an application (e.g., application 522) of UE 500 that wishes to use the split rendering service and service provider 550: determining UE data (e.g., pose sets and UE capability information) to be provided by UE 500 to server 530; determining the subject matter of computation to be performed by server 530 (e.g., rendering of content or application); determining the type and form of results (e.g., server results) to be provided (e.g., 2D video / audio streams and pose pairs); establishing a transport session for transmitting UE data and server results; or allocating server computation resources. SR manager 532 can determine to operate multiple renderers (e.g., renderer 542) and multiple encoders (e.g., encoder 538) that produce results of varying quality by comparing target presentation times received from UE 500 with rendering and encoding times of server 530; can operate the renderers and encoders; and can determine to transmit results to UE 500 that meet requirements in the renderer and encoder outputs. At least one renderer 542 and at least one encoder 538 can be configured with a split rendering function (SRF).
[0090] The renderer 542 is one of the server instances allocated and initialized by the SR manager 532. It can execute (e.g., render) content or applications specified by the SR manager 532 on behalf of the UE 500 and deliver the results to the encoder 538. The renderer 542 can receive at least one pose pair from the server pose manager 534 and execute the specified content or application, assuming that the UE 500 is in an estimated pose at the target display time of the pose pair. The renderer 542 can receive configurations regarding the type and form of one or more results from the SR manager 532 and output one or more results based on the configurations. When configurations regarding the type and form of two or more results are received, the renderer 542 can include two or more logical renderers corresponding to the results. Different renderers can generate different results for a pair of poses received from the server pose manager 534, and each result and the rendering start time or rendering time can be delivered to the encoder 538 (e.g., one or more logical encoders).
[0091] The encoder 538 is one of the server instances assigned and initialized by the SR manager 532 and can encode the results of the renderer 542 into a form that can be delivered to the UE 500 (e.g., at least one media frame). The encoder 538 can receive from the renderer 542 information regarding the pose pairs used by the renderer 542 to generate each audio or video frame, as well as time information related to the start time of each rendering and the time each rendering takes. Depending on the configuration of the SR manager 532, the encoder 538 can insert the pose pairs and time information as additional information in the media frame or pass it to the packetizer 540. The SR manager 532 can operate one or more encoders 538 that support one or more different encoding qualities for each rendering result. When executing two or more encoders 538, the computation time required on the server 530 can include both the time required for rendering and encoding until the point in time when the generation of the media frame is complete.
[0092] The session and service established by the SR manager 532 for the split rendering service may suggest a target time for end-to-end operation, and in an embodiment of the present disclosure, the target time may be suggested by the UE 500 or the service provider 550. In order to meet the target computing time on the server 530 in the total time of the end-to-end operation, the SR manager 532 may operate multiple renderers 542 and multiple encoders 538, and may select one or more results completed within the target computing time or one result completed in the fastest time, and deliver the selected results to the packetizer 540.
[0093] The packetizer 540 may generate the encoding results (e.g., media frames) received from the encoder 538 into a form (e.g., one or more packets) for transmission to the UE 500 via the network. In addition to the encoding results, the packetizer 540 may also receive input of pose information (e.g., a pose set or pose pair) selected by the server pose manager 534 and SR metadata (e.g., time information regarding the start time and the time taken for rendering by the renderer 542, configuration information regarding the type and form of rendering, the time taken for encoding and rendering, and / or encoding quality information). The packetizer 540 may associate and manage the pose information and SR metadata for each media frame.
[0094] In an embodiment of the present disclosure, when the packetizer 540 transmits media frames and SR metadata through one stream, the packetizer 540 may insert the SR metadata as a header extension into the headers of a first packet and subsequent packets including the media frame (or at least a portion of the media frame) within the payload, and deliver the packets to the UE 500. In an embodiment of the present disclosure, when the SR metadata is transmitted through a second stream separate from the first stream through which the media frames are transmitted, at least one packet of the second stream may include information capable of identifying the media frame associated with the SR metadata transmitted through the first stream (for example, the sequence number of the packet including the media frame associated with the SR metadata within the payload) and the SR metadata.
[0095] UE 500 can receive media frames and their associated pose and SR metadata (e.g., pose and metadata pairs) through one or two streams, can understand the content of the media frames based on the pose and SR metadata, and can perform post-processing on UE 500 (e.g., UE calculations such as scene synthesis and / or pose correction).
[0096] The media access function (MAF) 510 of the UE 500 may extract media frames and SR metadata from one or more streams received from the server 530 and associate them with each other. The SR metadata may be stored in an SR meta buffer 524, the media frames may be decoded and stored in a media frame buffer 526, and the SR metadata and media frames may be managed in pairs.
[0097] The scene manager 512 may position the media frames read from the media frame buffer 526 at a logical position within the space where the UE 500 is located based on the SR metadata read from the SR meta buffer 524 .
[0098] The pose corrector 514 may correct the media frame by performing a spatial transformation (eg, warping) to represent the media frame based on the final actual pose of the UE 500 .
[0099] The media frames corrected by the pose corrector 514 and the SR metadata for each media frame may be stored in a display frame buffer 516 before display. When the media frame needs to be displayed, the SR metadata and the target display time may be delivered to a metric collector 518 (eg, a metric analyzer).
[0100] The corrected media frames and the SR metadata for each media frame can be stored in a display frame buffer 516 before display. When the media frames are output to a display 520, the time components taken to render, encode, receive, and display the media frames starting from the moment when the display of the media frames is estimated based on the SR metadata (e.g., the first time point) can be delivered to a metric collector 518.
[0101] Metrics collector 518 can receive SR metadata and generate statistics for the SR metadata, where the amount of SR metadata corresponds to the most recent few slices or seconds. Metrics collector 518 can analyze the SR metadata and, as a result of the analysis, generate time components taken to render, encode, receive, and display the media frame from the moment the media frame's display is estimated (e.g., the first time point). Based on these time components, metrics collector 518 can derive the end-to-end time taken for rendering and display from the pose estimate and identify the most time-consuming processes. The analysis results (e.g., time components) from metrics collector 518 can be applied to SR estimator 504 and used to correct the T2 time estimate in SR estimator 504.
[0102] Figure 6 Components of each stage of segmentation rendering according to an embodiment of the present disclosure are shown.
[0103] refer to Figure 6 Pose acquisition 602 may include the operation of acquiring a pose from sensor 502. The final pose 604 generated by pose acquisition 602 may be used in pose estimation 606 to estimate the UE's future position (e.g., estimated pose 608) within a target display time (e.g., estimated T2), which is later than the current time (T1) by the pose-to-render-to-photon (P2R2P) delay. The P2R2P time can be determined as the sum of time components, which will be described later. The time components can be obtained during the negotiation phase and actual operation of the split rendering service. Pose buffer 610 may include the operation of storing SR metadata and pose pairs in the UE pose buffer 506. The SR metadata includes the estimated T1 and the estimated T2, and the pose pair includes the estimated pose. Network transmission 612 (e.g., uplink transmission) may include the operation of transmitting a pose set including one or more pose pairs to the server 530. Each pose set may include a version, which is an identifier capable of identifying the order in which the corresponding pose set has been generated, or a version identifier indicating a pose set transmission time (T1').
[0104] The server 530 can determine whether it has received the same pose set or pose pair as the previously received pose set or pose pair by using the version or version identifier. Based on the pose set sending time, the server 530 can determine the sending order of the corresponding pose sets.
[0105] The server pose caching 614 may include storing the pose pairs received by the server pose manager 534 in the server pose buffer 536, deriving the D_up time, which is the UE uplink delay time obtained by subtracting T1' from the current time when the pose pair was stored, and adding the D_up time to the SR metadata associated with the pose pair. When the renderer 542 completes rendering of the previous media frame (not shown) and enters a waiting state, the server pose manager 534 may deliver a pose pair read from the server pose buffer 536 to the renderer 542. At this time, the server pose manager 534 may calculate the render-to-photon (R2P) as a value obtained by subtracting T3 (which is the rendering start time) from T2 based on the time information (e.g., T2) received from the metric collector 518 of the UE 500. The renderer 542 may add T3, which is the rendering start time, and T4, which is the rendering completion time, to the SR metadata and perform rendering based on the pose pair 616, and may deliver the media frame obtained as a result of performing the rendering together with the SR metadata to the encoder 538. Upon completing encoding 618 of the media frame, the encoder 538 may add T5, which is the encoding completion time, to the SR metadata and deliver the media data obtained as a result of performing the encoding together with the SR metadata to the packetizer 540.
[0106] The packetizer 620 may include an operation of generating one or more packets including media data and SR metadata, the media data being a result of performing encoding. Network reception 622 (e.g., downlink reception) may include an operation of the MAF 510 of the UE 500 receiving one or more packets from the server 530 via one or more streams.
[0107] Network buffering 624 may include an operation in which the MAF 510 associates media frames and SR metadata obtained from packets received from the server 530 with each other and stores them in the media frame buffer 526 and the SR metadata buffer 524, respectively. The MAF 510 may add a time T6 indicating when a decodable frame has been received from the packet to the SR metadata. Media frame buffering 628 may include an operation in which the media frames generated by decoding 626 are stored in the media frame buffer 526.
[0108] Scene synthesis 630 may be performed by the scene manager 512. Immediately after performing scene synthesis 630, pose correction 636 may be performed at time T7 based on the final pose 634 obtained by the UE 500 via pose acquisition 632, and the final correction result of pose correction 636 may be added to the display frame buffer 516 via the display frame buffer 638. A display 640 of the final correction result may be output at the actual display time (actual T2). The metric collector 518 may add the actual T2 to the SR metadata. The metric collector 518 may derive at least one of the time point at which pose estimation was performed (T1), the estimated P2R2P (estimated T2 - T1), the time point at which the pose was sent (T1'), the pose sending time (D_up), the rendering start time (T3), the rendering time (T4 - T3), the encoding time (T5 - T4), the frame sending time (T6 - T5), or the UE processing time (T2 - T6). The metric collector 518 may compare the estimated P2R2P (estimated T2 - T1) with the actual P2R2P (actual T2 - T1) to derive the P2R2P to be used for the next estimation.
[0109] In an embodiment of the present disclosure, SR metadata may be generated in a state including only T1 and T2 for a pose pair, and may include additional values as the process (e.g., operations 614, 616, 618, 622, and 640) proceeds. In an embodiment of the present disclosure, SR metadata may be transmitted from the server 530 to the UE 500 by including statistical values generated from previous poses in fields T1 to T7.
[0110] The SR manager 532 may obtain the value of the P2R2P time from a report by the metric collector 518 of the UE 500, or identify the value of the P2R2P time based on statistics of the SR metadata described in the pose set, and may determine whether the P2R2P time is suitable for the needs of the UE 500 or the service provider 550. If the P2R2P time is determined to be unsuitable for the needs of the UE 500 or the service provider 550, the SR manager 532 may change the configuration of the renderer 542 and the encoder 538, or may use the results of a faster renderer and encoder generated by using two or more renderers and / or two or more encoders with two or more different configurations. The SR manager 532 may change the configuration of the renderer 542 and the encoder 538 based on the time taken by the UE 500 to enable faster processing of media data on the UE 500 (e.g., scene management and pose correction). For example, the SR manager 532 may configure the renderer 542 and the encoder 538 to generate media frames at a lower resolution to reduce the time spent on rendering and encoding, and the renderer 542 may generate processing hint information including depth and occupancy images and media data including 2D images to reduce processing time on the UE 500, and provide the media data and the processing hint information to the UE 500. The UE 500 can reduce the time it takes to detect images from the received media data for processing based on the processing hint information.
[0111] Figure 7 is a sequence diagram illustrating components of each stage of segmented rendering according to an embodiment of the present disclosure.
[0112] refer to Figure 7 , operations 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, and 714 may correspond to Figure 6 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624, 626, 628, 630, 632, 634, 636, 638, and 640. Operation 715 includes the actual pose-to-render-to-photons T1 to T2 between the XR source manager 504 and the display 520. Operation 716 includes the rendering-to-photons between the renderer 542 and the display 520. Operation 717 includes the motion-to-photons between the XR runtime and the display 520.
[0113] According to an embodiment of the present disclosure, UE 500 can send an estimated pose within an estimated target presentation time to server 530 to perform a segmented rendering service. UE 500 also receives rendered media frames from server 530, along with metadata (e.g., SR metadata) including the pose used to generate the media frames. The current MeCAR PD v0.4 describes user poses as pose and time, but it is unclear whether the time can be interpreted as the pose acquisition time or the target presentation time.
[0114] In order to make an accurate estimate of the target display time, the UE 500 may refer to the statistical record of the delays that have been previously estimated and actually occurred. The estimated delay is the time that has been estimated ( Figure 7 The actual delay that actually occurs is the gap between the time the estimate was made (T1) and the time the photon was displayed (T2.actual).
[0115] When UE 500 transmits two or more poses (e.g., a pose group) at once, the pose group can be considered a sequence of pairs of poses and metadata including multiple pieces of time information. When UE 500 intends to overwrite some of the already transmitted poses that have been updated using the most recently estimated parameters, server 530 can identify versions between poses or pose groups received from UE 500. If a pose has not yet been rendered, server 530 can replace a pose with the same target display time with a pose or pose group with a closer T1′.
[0116] When the frequency of poses stored in the server pose buffer 536 is more frequent (dense) than the frame reproduction frequency of the UE 500, split rendering can select the pose closest to the photon-to-rendering time. The photon-to-rendering time of the most recent frame information can help the server 530 select an appropriate pose.
[0117] The UE 500 may send the set of poses to a segmented rendering function (SRF) of the server 530 (e.g., the renderer 542 and the encoder 538), and the server 530 (e.g., the SRF) may generate rendered media frames based on the poses in the set of poses. Each pose may be associated with temporal metadata, such as the time at which the pose estimation was performed (T1), the estimated target display time of the content (T2.estimated), and the time at which the pose set was sent (T1').
[0118] The gap between the actual display time (T2.actual) and the estimated time (T1) is the pose-to-render-to-photon (P2RTP) delay, which allows the UE 500 to know the amount of processing time and connection latency to split the rendering loop. The next pose estimate can refer to the pose-to-render-to-photon delay used to estimate the new T2.estimated.
[0119] The split rendering of the server 530 may be referred to as T1'. T1' is the time when a set of poses is sent from the UE 500 when more than one pair of poses and metadata for the same target display time is received from the UE 500. T1' may be used by the server 530 to manage poses, for example, allowing the UE to update previously estimated information (e.g., estimated poses) by resubmitting new poses with the same target display time.
[0120] The server 530 may send the rendered media frames and associated metadata via the split rendering function to the UE 500. The metadata may include time information associated with the pose used for rendering (e.g., estimated T1 and / or T2) and the time when the renderer 542 has started rendering (T3), and may be used by the UE 500 to measure photon rendering (R2P) latency.
[0121] In an embodiment of the present disclosure, the pose and metadata information sent from the UE 500 may include at least one of the following parameters:
[0122] -Pose;
[0123] - the time to make the estimate (T1);
[0124] - estimated target display time (t2-estimated);
[0125] - the time at which the transmission took place (T1'); or
[0126] - Render to photon time (T2.actual - T3).
[0127] In an embodiment of the present disclosure, the pose and metadata information associated with the media frame sent from the server 530 may include at least one of the following parameters:
[0128] -Pose;
[0129] - the time at which the estimate was made (T1);
[0130] - estimated target display time (t2-estimated); or
[0131] - Time to start rendering (T3).
[0132] Figure 8 FIG. 4 shows an information structure of a pose set according to an embodiment of the present disclosure.
[0133] refer to Figure 8 , the pose set 800 may include at least one of timePoseSetSent, numberOfPosePair, earliestDeadline (estimated T2earlist), latestDeadline (estimated T2latest), or posePair[].
[0134] timePoseSetSent is the time (T1′) immediately before the pose set is sent from the UE 500 to the server 530 .
[0135] numberOfPosePair indicates the number of pose pairs described in the pose set.
[0136] earliestDeadline indicates the estimated pose time of the first pose pair to arrive.
[0137] latestDeadline indicates the estimated pose time of the latest pose pair to arrive.
[0138] The pose pair indicates one or more pose pairs including the estimated pose.
[0139] In an embodiment of the present disclosure, UE 500 may send multiple pose pairs instead of pose set 800. Each pose pair may be sent along with a pose pair identifier. UE 500 may store all sent pose pairs until a media frame is received, and then, when a media frame is received, compare the pose pairs using the pose pair identifiers to determine whether the estimation was successful (compare the estimated T2 with the actual T2) and refine the factors used in the estimation.
[0140] Figure 9 The information structure of the pose pair according to an embodiment of the present disclosure is shown.
[0141] refer to Figure 9 , the pose pair 900 may include at least one of estimatedPose or srMetadata.
[0142] estimatedPose may refer to a pose estimated by the UE 500 and may include poseType and pose[].
[0143] The pose type is a representation of the estimated pose (estimatedPose) and can indicate a quaternion or indicate position (xyz) and orientation (rpy: roll, pitch, and yaw).
[0144] srMetadata refers to the metadata associated with estimatedPose.
[0145] Figure 10 FIG. 4 shows an information structure of SR metadata according to an embodiment of the present disclosure.
[0146] refer to Figure 10 , the SR metadata (e.g., SR metadata 1000) may include at least one of timePoseEstimated (T1), timePoseSetSent (T1'), timeEstimatedDisplayTarget (T2 estimate), timeActualDisplay (T2 actual), timeRenderStarted (T3), timeRenderFinished (T4), timeEncodeFinished (T5), timeFrameReceive (T6), timeLateStageReprojection (T7), or timeRenderToPhotonPercentile[].
[0147] timePoseEstimated refers to the time point (T1) at which pose estimation has been performed.
[0148] The timePoseSetSent(T1') of a posePair has the same value as the timePoseSetSent of the poseSet to which the posePair belongs. The server 530 can use the value of timePoseSetSent to determine the order in which pose sets are transmitted and store the pose sets in the server pose buffer 536. In addition, when there are two different posePairs with the same value of timeEstimatedDisplayTarget, the server 530 can use timePoseSetSent to identify the poseSet from which each posePair comes. Of the two posePairs, the one with the larger value of timePoseSetSent is the newer estimate and therefore has a higher priority. For example, a posePair included in a previously transmitted poseSet can be deleted.
[0149] timeEstimatedDisplayTarget refers to the estimated target display time (estimated T2) used for the pose estimation above.
[0150] timeActualDisplay refers to the actual display time (actual T2) of the pose estimated at time point T1.
[0151] timeRenderStarted refers to the time point (T3) at which rendering of the media frame has started. The metric collector 518 of the UE 500 can use the value of timeRenderStarted to identify the actual R2P time.
[0152] timeRenderFinished refers to the time point when the rendering of the media frame has been completed (T4).
[0153] timeEncodeFinished refers to the time point (T5) when the encoding of the media frame has been completed.
[0154] timeFrameReceived refers to the time point at which the media frame is received in a decodable unit (T6).
[0155] timeLateStageReprojection refers to the time point (T7) at which the media frame is corrected according to the actual final pose of the UE 500 .
[0156] timeRenderToPhotonPercentile refers to the statistic of render-to-photon (R2P) time at previous points in time (for example, just before P50, P90, P95, and P99).
[0157] Figure 11 is a sequence diagram illustrating the updating of a pose set according to an embodiment of the present disclosure.
[0158] refer to Figure 11 , the UE 500 including the segmentation rendering client may continuously obtain the current pose of the UE 500 from the sensor 502 , and estimate the time point T2 and the pose at the time point T2 by using the SR estimator 504 .
[0159] In operation 1101 , the estimated pose may be stored in the pose buffer 506 .
[0160] In operations 1102 , 1103 , and 1104 , the UE pose manager 508 may transmit the estimated pose stored in the pose buffer 506 to the server 530 for each specific time point.
[0161] In an embodiment of the present disclosure, in operation 1104 , the UE pose manager 508 may send a pose set including timePoseSetSent(T1) to the server 530 .
[0162] In operation 1105 , the server pose manager 534 may add T1 ′ to the SR metadata of the pose pairs in the received pose set (eg, pose set #1) and store pose set #1 in the server pose buffer 536 .
[0163] In operation 1106 , the server pose buffer 536 may sort and store the pose pairs according to timeEstimatedDisplayTarget ( T2 ).
[0164] In an embodiment of the present disclosure, estimated poses for the same target display time that have been delivered by UE 500 to server 530 can be corrected. In an embodiment of the present disclosure, server 530 can determine the order in which poses sent by UE 500 are sent, regardless of whether the communication path between UE 500 and server 530 ensures the order of transmission, as described below.
[0165] In order to enable the server 530 to determine the order in which the poses sent by the UE 500 are to be sent, the UE 500 may bundle multiple estimated poses to be sent and send them in the form of a repository called a pose set ("poseSet"). Figure 8 As shown, poseSet 800 includes timePoseSetSent, and timePoseSetSent may be listed at time T1′, which is immediately before poseSet 800 is generated and sent by UE pose manager 508. In an embodiment of the present disclosure, when timePoseSetSent is not listed within poseSet 800, server pose manager 534 may regard the value of the sending timestamp of the packet containing poseSet 800 as timePoseSetSent.
[0166] The server pose manager 534 can extract pose pairs (e.g., pose pair 900) from the received poseSet 800 and store them in the server pose buffer 536. In an embodiment of the present disclosure, the timePoseSetSent field of the srMetadata (e.g., srMetadata 1000) within posePair 900 is configured with the value of timePoseSetSent corresponding to the poseSet 800 to which posePair 900 belongs. At the point in time when the server pose manager 534 has received poseSet 800, the non-null fields in srMetadata 1000 are timePoseEstimated and timeEstimatedDisplayTarget, and the other fields may be null because they have not yet reached or replicated timePoseSetSent. The pose pairs stored in the server pose buffer 536 can be processed so that the value of timePoseSetSent is filled and the pose pairs are sorted in the order of timeEstimatedDisplayTarget.
[0167] Figure 12 is a sequence diagram illustrating correction of an estimated pose according to an embodiment of the present disclosure.
[0168] refer to Figure 12 , the UE 500 (eg, the SR estimator 504 ) may estimate a pose in operation 1201 , and may generate a pose set (eg, pose set # 1 ) including the estimated pose in operation 1203 .
[0169] In operation 1204 , the UE pose manager 508 of the UE 500 may transmit the pose set # 1 including timePoseSetSent(T 1 ′) to the server 530 .
[0170] In operation 1205 , the server pose manager 534 may add T1 ′ to the SR metadata of the pose pairs of pose set # 1 .
[0171] In operation 1206 , the server pose manager 534 may sort pose set # 1 according to timeEstimatedDisplayTarget ( T2 ) and store it in the server pose buffer 536 .
[0172] In operation 1207, UE 500 (eg, UE pose manager 508) may send pose set #2 including a new estimated pose to server 530 using new estimated parameters to correct the already sent estimated pose (eg, pose set #1) to the new estimated pose.
[0173] In operation 1208 , the server pose manager 534 may replace an existing pose pair (eg, pose set # 1 ) having the same T2 as the newly received pose pair (eg, pose set # 2 ) with the newly received pose pair.
[0174] In operation 1209 , the server pose manager 534 may sort pose set # 2 according to timeEstimatedDisplayTarget ( T2 ) and store it in the server pose buffer 536 .
[0175] A second pose set (e.g., pose set #2) that is retransmitted to correct an already transmitted first pose set (e.g., pose set #1) may include a second pose pair having the same timeEstimatedDisplayTarget as the first pose pair included in the first pose set. After storing the pose pair of the first pose set in the server pose buffer 536, the server pose manager 534 may treat the second pose pair as a copy of the first pose pair having the same timeEstimatedDisplayTarget associated with storing the pose pair of the second pose set in the server pose buffer 536. The server pose manager 534 may compare the timePoseSetSent information of the first pose pair and the second pose pair to determine that the UE 500 has recently transmitted a pose pair with a larger value (e.g., the second pose pair), and may determine that the UE 500 sent the second pose pair with the intention of correcting the previous estimate (e.g., the first pose pair). The server pose manager 534 may replace the first pose pair of the first pose set, which has been first sent and stored in the buffer, with the second pose pair of the second pose set.
[0176] Figure 13 is a sequence diagram illustrating the reversal of the order of estimating poses according to an embodiment of the present disclosure.
[0177] refer to Figure 13 , the UE 500 (eg, the SR estimator 504 ) may estimate a pose in operation 1301 , and may generate pose sets (eg, pose set # 1 and pose set # 2 ) including the estimated pose in operations 1303 and 1304 , respectively.
[0178] In operation 1304 a , the pose set # 2 generated later may first be sent by the UE pose manager 508 of the UE 500 to the server 530 .
[0179] In operation 1305 , the server pose manager 534 may add T1 ′ to the SR metadata of the pose pairs of pose set # 2 .
[0180] In operation 1306 , the server pose manager 534 may sort pose set # 2 according to timeEstimatedDisplayTarget ( T2 ) and store it in the server pose buffer 536 .
[0181] When a communication path in which the order in which packets are sent is not guaranteed is used between the UE 500 and the server 530, in operation 1304b, the first pose set (e.g., pose set #1) first sent by the UE 500 may arrive at the server 530 later than the second pose set (e.g., pose set #2).
[0182] In operation 1307 , the server pose manager 534 may store the pose pairs of the first pose set in the server pose buffer 536 .
[0183] When the server pose manager 534 processes the pose pairs of the first pose set after the pose pairs of the second pose set are stored in the server pose buffer 536, the server pose manager 534 can identify the transmission order of the pose pair (e.g., the first pose pair) having the same timeEstimatedDisplayTarget as the previously stored pose pair (e.g., the second pose pair). The server pose manager 534 can compare the timePoseSetSent of the first pose pair and the second pose pair with each other to determine that the pose pair (e.g., the second pose pair) with the larger value was sent later by the UE 500, and can determine that the arrival order of the first pose pair and the second pose pair has been reversed during the transmission process.
[0184] In operation 1308 , the server pose manager 534 may determine that the second pose pair of the second pose set received earlier and stored in the server pose buffer 536 is not to be replaced by the first pose pair of the first pose set received later.
[0185] According to an embodiment of the present disclosure, the 5G Media Service Enabler (MSE) (e.g., SR_MSE) of 3GPP SA4 is an enhanced MSE for supporting multimedia services over a 5G network. The 5G MSE builds on the existing MSE to provide advanced multimedia capabilities, such as improved video and audio quality, low latency, and high stability, and can support the requirements of 5G services and applications. The 5G MSE can provide end-to-end multimedia services over a 5G network, including supporting the 5G-defined Internet Protocol (IP) Multimedia Subsystem (IMS) by leveraging features of the 5G network, such as network slicing. The relevant 3GPP specifications for the 5G MSE are part of 3GPP Release 16 and subsequent releases.
[0186] The Split Rendering Media Service Enabler (SR_MSE) supports split rendering based on 5G MSE. The architecture for UE and server to provide split rendering functionality and the requirements for application APIs to provide the above functionality are discussed.
[0187] Figure 14 is a sequence diagram illustrating a process for split rendering between a UE and a server according to an embodiment of the present disclosure. The illustrated process may be defined in SR_MSE (TS 26.565 V0.2.0 202 2-11).
[0188] refer to Figure 14 , in operation 1401 , the UE 500 (eg, the scene manager 512 ) may establish a split rendering session with the server 530 (eg, SREAS).
[0189] In operation 1402 , the server 530 may transmit a description of the output of the segmented rendering (eg, rendering results) to the UE 500 (eg, the scene manager 512 ).
[0190] In operation 1403 , the UE 500 (eg, the scene manager 512 ) may establish a connection with the server 530 .
[0191] In operation 1404 , the UE 500 (eg, the XR runtime module) may deliver the user input and pose information including the estimated pose to the XR source management module (eg, the SR estimator 504 ) of the UE 500 .
[0192] In operation 1405 , the UE 500 (eg, an XR source management module) may transmit pose information and user input to the server 530 .
[0193] In operation 1406 , the server 530 may perform rendering of the requested pose based on the pose information and the user input.
[0194] In operation 1407 , the server 530 may transmit the next buffered frame (eg, media frame) to the UE 500 (eg, the MAF 510 ).
[0195] In operation 1408 , the UE 500 (eg, the MAF 1408 ) may perform decoding and processing of media data of the buffered frame.
[0196] In operation 1409 , the UE 500 (eg, the MAF 510 ) may transmit a media frame (eg, a raw buffered frame) generated as a result of decoding and processing to the XR runtime module via the scene manager 512 .
[0197] In operation 1410 , the UE 500 (eg, an XR runtime module) may synthesize, render, correct, and display a raw buffered frame.
[0198] Figure 15a and 15b is a sequence diagram illustrating a process for split rendering between a UE and a server according to various embodiments of the present disclosure.
[0199] refer to Figure 15a and 15b In operation 1501 , a session (eg, a split rendering session) may be established between a scene manager 512 (eg, a scene rendering engine) of a UE 500 (eg, a split rendering client) and an SR manager 532 of a server 530 .
[0200] In operation 1502 , the server 530 (eg, the SR manager 532 ) may provide the UE 500 (eg, the scene manager 512 ) with a description including properties and associated information of a media frame as an output of rendering (eg, a rendering result).
[0201] In operation 1503, the UE 500 (e.g., the scene manager 512 and the MAF 510) may establish a connection with the server 530 (e.g., the server pose manager 534 and the packetizer 540) for network transmission (e.g., uplink transmission) and for reception of media frames (e.g., downlink reception).
[0202] In operation 1504, the UE 500 (e.g., the metric collector 518) may collect and analyze metrics related to server and client performance (e.g., at least one of network transmission speed, central processing unit (CPU) / graphics processing unit (GPU) processing speed, media processing speed, or media frame attribute information) from the server 530 and / or the UE 500.
[0203] In operation 1505 , performance metric values associated with segmentation rendering (eg, SR statistics) may be delivered from the metric collector 518 to the SR estimator 504 of the UE 500 .
[0204] In operation 1506 , pose information obtained from the UE 500 (eg, the XR runtime module) and user input may be delivered to the SR estimator 504 .
[0205] In operation 1507 , the SR estimator 504 may estimate a target display time ( T2 ) based on the pose according to the pose information and the user input and the performance metric.
[0206] In operation 1508 , the SR estimator 504 may estimate a pose at time T2.
[0207] In operation 1509 , the estimated pose may be delivered to the UE pose manager 508 and may be stored in the pose buffer 506 .
[0208] In operation 1510 , the stored poses (eg, a set of estimated poses) may be sent to the server pose manager 534 in the form of a pose set.
[0209] In operation 1511, when there is a pose pair (e.g., a new pose pair) in the received pose set that has the same T2 as a pose pair (e.g., a previous pose pair) already stored in the server pose buffer 536, the server pose manager 534 may overwrite the previous pose pair with information of the new pose pair.
[0210] In operation 1512 , the server pose manager 534 may select at least one pose pair from among the pose pairs stored in the server pose buffer 153 based on the performance metric values of the server 530 and the UE 500 , and transmit the pose in the selected pose pair to the renderer 542 .
[0211] In operation 1513 , the server 530 (eg, the renderer 542 , the encoder 538 , and the packetizer 540 ) may perform rendering, encoding, and packetization based on the selected pose.
[0212] In operation 1514 , the server 530 (eg, the packetizer 540 ) may transmit the rendered media frame and SR metadata related thereto to the UE 500 (eg, the MAF 510 ) (eg, downlink transmission).
[0213] In operation 1515 , the UE 500 (eg, the MAF 510 ) may decode the media frame and the SR metadata received from the server 530 and store the media frame and the SR metadata in the media frame buffer 526 and the SR metadata buffer 524 .
[0214] In operation 1516 , the metric collector 518 may calculate the components of the plurality of time metrics and update statistics regarding the time spent on each component.
[0215] In operation 1517 , the buffer data read from the media frame buffer 526 and the SR metabuffer 524 may be delivered to the metric collector 518 , the scene manager 512 , and the XR runtime modules (eg, the pose corrector 514 ).
[0216] In operation 1518 , the XR runtime module (eg, the pose corrector 514 ) may synthesize a scene based on the buffer data and perform rendering, and may perform pose correction to correct a difference between the estimated pose and the actual pose.
[0217] In operation 1519 , the UE 500 (eg, the metric collector 518 ) may measure actual P2R2P delay and R2P delay based on the displayed results.
[0218] In a wireless communication system according to an embodiment of the present disclosure, the server 530 may transmit an IP packet including an encapsulated media frame to the UE 500 by using Real-time Transport Protocol (RTP) or Secure RTP (SRTP).
[0219] Figure 16 is a conceptual diagram illustrating an IP packet structure including a media frame in a wireless communication system according to an embodiment of the present disclosure.
[0220] refer to Figure 16 , the media frame may be included in the RTP payload 1610 of the IP packet 1600. The IP packet 1600 may also include an RTP header 1620, a User Datagram Protocol (UDP) header 1630, and an IP header 1640 before the RTP payload 1610. The RTP header (X120) may include an RTP header extension (e.g., header extension 1602). The first 12 octets or 12 bytes of the RTP header 1620 may be included in all RTP packets, and an identifier of a contributing source (CSRC) may be added by the mixer. Each field of the RTP header 1620 has the following meaning:
[0221] -version (V): A 2-bit field indicating the version of RTP. RTP according to IETF RFC 3550 has a value of 2;
[0222] -padding (P): a 1-bit field with a value of 1 when the RTP packet includes padding octets;
[0223] - Extension (X): a 1-bit field having a value of 1 when the RTP packet includes an RTP header extension 1602;
[0224] -CSRC Count (CC): a 4-bit field indicating the number of CSRC identifiers following the 12-octet fixed header;
[0225] - Flag (M): 1-bit field, and its use is determined by the RTP profile. For example, when a video frame is divided into multiple RTP packets and transmitted, only the value of the M field of the last RTP packet in the RTP packets can be configured to be 1;
[0226] - Payload Type (PT): A 7-bit field used to identify the format of the RTP payload. The value of the field can be determined using a static mapping determined by the RTP profile or a dynamic mapping determined by an out-of-band method using the Session Description Protocol (SDP);
[0227] - Sequence number: 16-bit field that increases by 1 each time each RTP packet is sent. It can be used for loss detection and packet reordering at the receiver;
[0228] -Timestamp: A 32-bit field that may indicate an acquisition time point or a reproduction time point of a data sample included in a corresponding RTP packet;
[0229] -SSRC: a 32-bit field indicating the identifier of the synchronization source; and
[0230] -CSRS: A 32-bit field indicating the identifier of the contributing source.
[0231] exist Figure 16 In the embodiment, although a media frame is shown as being included in a single RTP payload 1610, the media frame may be segmented according to its data size, and the RTP payload 1610 may include at least a portion of the media frame (e.g., a segmented portion). For example, a single media frame may be transmitted via multiple IP packets. The RTP headers (e.g., RTP header 1620) of the IP packets transmitting a single media frame may have the same value in the timestamp field. Consecutive media frames may be transmitted via an RTP stream, which is a stream of consecutive RTP packets. The RTP stream may be identified by the SSRC field in an RTP session, which is defined as an association between entities participating in RTP-based communication (e.g., UE 500 and server 530).
[0232] In a wireless communication system according to an embodiment of the present disclosure, a segmentation rendering server (e.g., server 530) may perform rendering based on a first pose pair to provide media frames associated with the first pose pair and a second pose pair obtained by updating the first pose pair to a segmentation rendering client (e.g., UE 500). The second pose pair may be a pose pair in which only the value of the SR metadata in the first pose pair is updated.
[0233] In a wireless communication system according to an embodiment of the present disclosure, a segmented rendering server (e.g., server 530) may include a pose pair associated with a media frame in an RTP header extension (e.g., RTP header extension 1602) to transmit to a segmented rendering client (e.g., UE 500). The RTP header extension including the pose pair may be included in an RTP packet carrying the associated media frame. In an embodiment of the present disclosure, the RTP header extension including the pose pair may be added to another RTP header of an RTP stream carrying RTP packets including the media frame, or may be added to the RTP header of a separate RTP stream.
[0234] In an embodiment of the present disclosure, an RTP header extension can be identified by a Uniform Resource Name (URN) as a global identifier and an ID field as a local identifier. The mapping between the global identifier URN and the local identifier ID field can be negotiated using a separate protocol (e.g., SDP) during the RTP session establishment process. In an embodiment of the present disclosure, an RTP packet can include one or more RTP header extensions, and the format and usage of the RTP header extension included in the RTP packet can be identified by the ID field included in the RTP header extension.
[0235] According to an embodiment of the present disclosure, the pose pair may be sent via one RTP header extension or a combination of at least two RTP header extensions.
[0236] The RTP header extension including the pose pair according to an embodiment of the present disclosure may include at least one of the following information:
[0237] - associated media frame identification information: as an example, the value of the timestamp field or the sequence number of the first RTP header of the associated media frame;
[0238] - Pose pair identification information: as an example, the time when the UE has sent the pose set including the pose pair to the server;
[0239] - pose information: as an example, an identifier indicating the format of a pose, including at least one of a quaternion or a position (xyz) and an orientation (e.g., roll, pitch, yaw), and a pose value according to the pose format; and
[0240] -SR metadata: As an example, Figure 10 At least one of the values shown in .
[0241] In an embodiment of the present disclosure, when a pose pair associated with a media frame is transmitted in the same RTP packet header extension as the RTP packet including the media frame, the associated media frame identification information may be omitted. In an embodiment of the present disclosure, when the RTP packet transmitting the media frame and the RTP packet with the RTP header extension including the pose pair associated with the media frame are transmitted via different RTP streams, the associated media frame identification information may also include RTP stream identification information, which includes the RTP packet used to transmit the media frame. The RTP stream identification information may include, for example, the value of the SSRC field in the RTP header used to transmit the media frame and at least one of the sending and receiving addresses and port numbers of the IP packet and the UDP packet. The RTP stream identification information may be exchanged between the UE 500 and the server 530 using out-of-band signaling (e.g., SDP).
[0242] According to an embodiment of the present disclosure, SR metadata may include one or more elements representing time, and may also include an identifier indicating the time representation format of the element. The time representation format may be, for example, the same as the representation format of the timestamp field of the RTP header containing the SR metadata within the header extension, or the same as the time representation format of the Network Time Protocol (NTP), or a format that only takes some bits from the NTP format, or a representation format separately defined by a service provider. The method of transmitting SR metadata using an RTP header extension according to an embodiment of the present disclosure may use one RTP header extension that supports multiple time representation formats, or may use at least one RTP header extension with different URNs according to different time representation formats.
[0243] Figure 17 An RTP header extension structure including a pose pair according to an embodiment of the present disclosure is shown.
[0244] refer to Figure 17 , the RTP header extension 1700 containing the pose pair (e.g., RTP header extension 1602) may include at least one of a local identifier (ID) field, a length (L) field, or a flag field 1702, wherein the flag field 1702 may determine the format and purpose of the data structure of the following pose data field 1704 and the data structure of the following SRMetaData field 1706.
[0245] In an embodiment of the present disclosure, the bits (e.g., F0, F1, F2, F3, F4, F5, F6, and F7) of the flag field 1702 may have the following meanings:
[0246] -F0: an identifier indicating whether the PoseData field 1704 contains a value of a pose used in rendering the associated media frame. For example, a value of 0 in F0 may indicate that the PoseData field 1704 contains only information for identifying an estimated pose stored by the UE 500, while a value of 1 in F0 may indicate that the PoseData field 1704 contains an estimated pose;
[0247] - F1: an identifier that indicates the representation format of the pose when the PoseData field 1704 contains a pose used in rendering the associated media frame (e.g., when the value of F0 is 1). For example, an F1 value of 0 may indicate that the pose value is represented in quaternion format, and an F1 value of 1 may indicate that the pose value is represented in the format of position (e.g., xyz) and orientation (e.g., roll, pitch, and yaw);
[0248] - F2F3: a 2-bit identifier indicating the time format of the PoseData field 1704 and the SRMetaData field 1706. For example, a value of 00 for F2F3 may indicate the use of a 64-bit length NTP timestamp format, a value of 01 for F2F3 may indicate the use of only the middle 32 bits of the 64-bit length NTP timestamp, and a value of 10 for F2F3 may indicate the use of the same format as the timestamp in the RTP header;
[0249] - F4: A 1-bit identifier that indicates whether the parameter contained in the SRMetaData field 1706 is an absolute or relative time. For example, with respect to the value of the parameter contained in the SRMetaData field 1706, a value of 0 for F4 may be a relative time to be measured with reference to timePoseSetSent, and a value of 1 for F4 may be an absolute time.
[0250] - F5: a 1-bit identifier indicating whether the SRMetaData field 1706 includes the D_up parameter (UE uplink time); and
[0251] - F6F7: A 2-bit identifier that, when the SRMetaData field 1706 has the same format as the timestamp of the RTP header, indicates a parameter of the SRMetaData field 1706 that matches the timestamp of the RTP header including the SRMetaData field 1706. For example, the timestamp of the RTP header including the SRMetaData field 1706 may match a value of T3 (timeRenderStarted) when the value of F6F7 is 00, a value of T4 (timeRenderFinished) when the value of F6F7 is 01, or a value of T5 (timeEncoderFinished) when the value of F6F7 is 10. In this case, the parameter that matches the timestamp value of the RTP header may be omitted from the SRMetaData field 1706.
[0252] The above embodiment may assume that the RTP header extension identified by a single URN (e.g., URN:3gpp:SRMeta) supports all formats of the PoseData field 1704 and the SRMetaData field 1706 as determined by the flag field. The RTP header extension 1602 according to an embodiment of the present disclosure may include a combination of RTP header extensions identified by two or more URNs that only support a portion of the PoseData field 1704 and the SRMetaData field 1706 as determined by the flag field. For example, urn:3gpp:SRMeta-NTP, urn:3gpp:SRMeta-NTP-compact, and urn:3gpp:SRMeta-RTP-timestamp may each refer to an RTP header extension that includes a 64-bit NTP timestamp, the middle 32 bits of the NTP timestamp, and a timestamp that indicates time information as an RTP packet header.
[0253] Figure 18 FIG. 2 shows a PoseData field included in an RTP header extension according to an embodiment of the present disclosure.
[0254] refer to Figure 18, the pose data field 1800 (e.g., the pose data field 1704) may include at least one of the following: timePoseSetSent(T1') 1802, which indicates the time when the pose set for use in rendering the associated media frame has been sent from the UE, timePoseEstimated(T1) 1804, which indicates the time point at which the pose estimation has been performed, timeEstimatedDisplayTaget(T2) 1806, which indicates the estimated target display time used in the pose estimation, or pose[] 1808, which indicates the pose used in rendering the associated media frame.
[0255] The presence or value of each of the parameters 1802, 1804, 1806, and 1808 configuring the PoseData field 1800 may be determined by the format (e.g., a format recognizable by URN) of the RTP header extension (e.g., RTP header extension 1700) including the PoseData field 1800 or by Figure 17 The flag field 1702 shown in FIG.
[0256] Figure 19 An SRMetaData field included in an RTP header extension according to an embodiment of the present disclosure is illustrated.
[0257] refer to Figure 19 , the SRMetaData field 1900 (e.g., the SRMetaData field 1706) may include at least one of: TimeRenderStart (T3) 1902, which indicates a time point at which rendering of the associated media frame has started for a pose used to render the associated media frame, timeRenderFinished (T4) 1904, which indicates a time point at which rendering of the associated media frame has been completed, timeEncoderFinished (T5) 1906, which indicates a time point at which encoding of the associated media frame has been completed, or uplinkDelay (D_up) 1908, which indicates a sending delay time of a pose set containing a pose used when rendering the associated media frame.
[0258] The presence or value of each of the parameters 1902, 1904, 1906, and 1908 configuring the SRMetaData field 1900 may be determined by the format of the RTP header extension (e.g., RTP header extension 1700) (e.g., recognizable as a URN) that includes the SRMetaData field 1900 or by Figure 17 The flag field 1702 described in is controlled.
[0259] In a wireless communication system according to an embodiment of the present disclosure, a split rendering server (e.g., server 530) may transmit pose pairs associated with media frames through a WebRTC data channel established between a split rendering client (e.g., UE 500) and the server 530.
[0260] The WebRTC data channel can use the UDP / Datagram Transport Layer Security (DTLS) / Stream Control Transmission Protocol (SCTP) protocol and can be configured by an SCTP stream pair with the same SCTP stream identifier. The data unit of the WebRTC data channel can be an SCTP block. According to an embodiment of the present disclosure, the SCTP block including the pose pair may include at least one of the following information:
[0261] - Associated media frame identification information: as an example, including the sequence number or timestamp value of the first RTP header of the associated media frame;
[0262] - pose pair identification information: as an example, an identifier indicating a format of a pose including at least one of a quaternion or a position (e.g., xyz) and an orientation (e.g., roll, pitch, and yaw), and a pose value based on the pose format; and
[0263] -SR metadata: For example, Figure 10 At least one of the parameters of the SR metadata shown.
[0264] In an embodiment of the present disclosure, the above-mentioned information element may be encoded into a continuous bit string and included in an SCTP block, and its specific format and usage may be the same as when the above-mentioned RTP header extension is used.
[0265] Figure 20 is a block diagram illustrating a configuration of a UE in a communication system according to an embodiment of the present disclosure.
[0266] refer to Figure 20 , UE 500 may include a processor 2010, a transceiver 2020, and a memory 2030. The processor 2010, the transceiver 2020, and the memory 2030 of UE 500 may be configured according to Figure 1 、 Figure 2 、 Figure 3a 、 Figure 3b 、 Figures 4 to 14 、 Figure 15a 、 Figure 15b and Figures 16 to 19 The UE 500 may be operated according to the method described in the aforementioned embodiment. However, the components of the UE 500 are not limited to the aforementioned examples. For example, the UE 500 may include more or fewer components than those described above. In addition, at least one of the processor 2010, the transceiver 2020, and the memory 2030 may be implemented in the form of a single chip.
[0267] Transceiver 2020 is a term that collectively refers to both a receiver and a transmitter. UE 500 can transmit or receive signals to or from a base station or another network entity (e.g., server 530) via transceiver 2020. In this case, the transmitted or received signal may include at least one of control information and data. To this end, transceiver 2020 may include an RF transmitter that up-converts and amplifies the frequency of a transmitted signal, and an RF receiver that performs low-noise amplification and down-converts the frequency of a received signal. This is merely one example configuration of transceiver 2020, and the components of transceiver 2020 are not limited to an RF transmitter and an RF receiver.
[0268] The transceiver 2020 may receive an RF signal via a communication method defined by the 3GPP standard and output the RF signal to the processor 2010, and may transmit control information or data output from the processor 2010 to the server 530 via the network 100 (e.g., a base station) via the RF signal. The transceiver 2020 may receive a signal transmitted by the server 530 via the network 100 and provide the signal to the processor 2010.
[0269] The memory 2030 can store Figure 1 、 Figure 2 、 Figure 3a 、 Figure 3b 、 Figures 4 to 14 、 Figure 15a 、 Figure 15b and Figures 16 to 19 The memory 2030 may store programs and data required for the operation of the UE 500 according to at least one embodiment of the present invention. In addition, the memory 2030 may store control information and / or data included in the signal acquired by the UE 500. The memory 2030 may include a storage medium such as a read-only memory (ROM), a random access memory (RAM), a hard disk, a CD-ROM, and a (digital versatile disk) DVD, or a combination of storage media.
[0270] The processor 2010 can control a series of operations to enable the UE 500 to Figure 1 、 2, 3a, 3b, 4 to 14, 15a, 15b, and 16 to 19. The processor 2010 may include at least one processing circuit (e.g., an application processor (AP) and / or a communication processor (CP)). The processor 2010 may include (e.g., execute) the previously described components of the UE 500 (e.g., at least one of the SR estimator 504, the application 522, the UE pose manager 508, the metric collector 518, the MAF 510, the scene manager 512, or the pose corrector 514). At least one of the pose buffer 506, the SR meta buffer 524, the media frame buffer 526, or the display frame buffer 516 may be included in the processor 2010 or the memory 2030. Although not shown, the UE 500 may also include at least one sensor 502 and a display 520.
[0271] Figure 21 is a block diagram illustrating a configuration of a server in a wireless communication system according to an embodiment of the present disclosure.
[0272] refer to Figure 21 , the server 530 may include a processor 2110, a network interface (NW IF) 2120, and a memory 2130. The processor 2110, the network interface 2120, and the memory 2130 of the server 530 may be configured according to Figure 1 、 Figure 2 、 Figure 3a 、 Figure 3b 、 Figures 4 to 14 、 Figure 15a 、 Figure 15b and Figures 16 to 19 The server 530 may be operated according to the method described in the above embodiment. However, the components of the server 530 are not limited to the above examples. For example, the server 530 may include more or fewer components than those described above. In addition, at least one of the processor 2110, the network interface 2120, and the memory 2130 may be implemented in the form of a single chip.
[0273] The network interface 2120 may include a receiver and a transmitter, and the server 530 may transmit or receive a signal to or from a UE (e.g., UE 500) or another network entity through the network interface 2120. The transmitted or received signal may include at least one of control information and data.
[0274] The memory 2130 can store Figure 1 、 Figure 2 、 Figure 3a 、 Figure 3b 、 Figures 4 to 14 、 Figure 15a 、 Figure 15b and Figures 16 to 19The memory 2130 may store programs and data necessary for the operation of at least one of the server 530 in the embodiments of the present invention. In addition, the memory 2130 may store control information and / or data included in the signal obtained from the server 530. The memory 2130 is a storage medium such as a read-only memory (ROM), a random access memory (RAM), a hard disk, a compact disc ROM (CD-ROM), and a digital versatile disk (DVD), or a combination of storage media.
[0275] The processor 2110 can control a series of operations to enable the server 530 to Figure 1 、 Figure 2 、 Figure 3a 、 Figure 3b 、 Figures 4 to 14 、 Figure 15a 、 Figure 15b and Figures 16 to 19 The processor 2110 may include at least one processing circuit (e.g., an AP and / or a CP). The processor 2110 may include (e.g., execute) the components of the server 530 described above (e.g., at least one of the SR manager 532, the server pose manager 534, the renderer 542, the encoder 538, or the packetizer 540). The server pose buffer 536 may be included in the processor 2110 or the memory 2130.
[0276] According to an embodiment of the present disclosure, one or more non-transitory computer-readable storage media store computer-executable instructions, which, when executed by one or more processors of a user equipment (UE), cause the UE to perform operations, including: estimating the posture of the UE at a second time point at a first time point, the second time point indicating a target display time of a first media frame, sending the first media frame and first posture information related to the estimated posture to a server, receiving from the server a second media frame generated by rendering based on the first media frame and the estimated posture and second posture information associated with the second media frame, generating a third media frame by correcting the second media frame based on metadata and the actual posture of the UE, and displaying the third media frame at a third time point.
[0277] In an embodiment of the present disclosure, the operation may further include measuring a pose-to-rendering-to-photon (P2R2P) delay from the first time point to a third time point and a rendering-to-photon (R2P) delay from the time when rendering of the second media frame has started to the third time point.
[0278] In the above detailed embodiments of the present disclosure, the elements included in the present disclosure are expressed in the singular or plural, depending on the detailed embodiment presented. However, for ease of description, the singular form or plural form is appropriately selected for the situation presented, and the present disclosure is not limited to elements expressed in the singular or plural. Therefore, an element expressed in the plural may also include a single element, or an element expressed in the singular may also include multiple elements.
[0279] While the present disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. A method performed by a user equipment (UE) supporting split rendering in a communication system, the method comprising: estimating, at the first time point, a pose of the UE at a second time point indicating a target display time of the first media frame; sending a first media frame and first pose information associated with the estimated pose to a server; receiving, from the server, a second media frame generated by rendering based on the first media frame and the estimated pose, and second pose information associated with the second media frame; as well as Based on the second media frame and the second pose information, a third media frame is displayed at a third time point.
2. The method of claim 1 , further comprising measuring at least one of a pose-to-render-to-photon (P2R2P) latency from the first point in time to a third point in time and a render-to-photon (R2P) latency from a time when rendering of the second media frame has begun to the third point in time.
3. The method according to claim 1, wherein The first pose information includes predicted pose information representing a position and an orientation of an estimated pose and first temporal metadata associated with the estimated pose.
4. The method according to claim 3, wherein: The first temporal metadata includes at least one of the following: T1, indicates the first time point; T2.estimated, indicating the second time point; T1', indicating the time when the first pose information is sent from the UE to the server; as well as Render-to-photon latency, indicated as T2.actual - T3, where T2.actual indicates the time the photon has been displayed, and T3 indicates the time the rendering started.
5. The method according to claim 1, wherein The second posture information includes: Pose information, indicating the pose used for rendering; and Second temporal metadata associated with the second media frame.
6. The method according to claim 5, wherein: The second temporal metadata includes at least one of the following: T1, indicates the first time point; T2.estimated, indicating the second time point; T3, the actual time when the server starts rendering the first media frame; as well as T5 indicates the time when the second media frame is output from the server.
7. A method performed by a server supporting split rendering in a communication system, the method comprising: Receiving a first media frame and first pose information associated with an estimated pose from a user equipment (UE), wherein the estimated pose corresponds to a pose of the UE at a target display time; generating a second media frame by rendering based on the first media frame and the estimated pose, and generating second pose information associated with the second media frame; and Send the second media frame and the second posture information to the UE.
8. The method according to claim 7, wherein: The first pose information includes predicted pose information representing a position and an orientation of an estimated pose and first temporal metadata associated with the estimated pose.
9. The method according to claim 8, wherein The first temporal metadata includes at least one of the following: T1 indicates the first time point at which the estimated pose is made in the UE; T2.estimated, indicating the target display time estimated by the UE for the first media frame; T1', indicating the time when the first pose information is sent from the UE to the server; as well as Render-to-photon latency, indicated as T2.actual - T3, where T2.actual indicates the time the photon has been displayed, and T3 indicates the time the rendering started.
10. The method according to claim 7, wherein: The second posture information includes: Pose information, indicating the pose used for rendering; and Second temporal metadata associated with the second media frame.
11. The method according to claim 10, wherein: The second temporal metadata includes at least one of the following: T1 indicates the first time point at which the estimated pose is made in the UE; T2.estimated, indicating the target display time estimated by the UE for the first media frame; T3, the actual time when the server starts rendering the first media frame; as well as T5 indicates the time when the second media frame is output from the server.
12. A user equipment (UE) for supporting split rendering in a communication system, the UE comprising: transceiver; Memory; and a processor, coupled with a transceiver and a memory, The memory stores one or more computer programs including computer-executable instructions, which, when executed by the processor, cause the UE to: estimating, at the first time point, a pose of the UE at a second time point indicating a target display time of the first media frame, sending a first media frame and first pose information associated with the estimated pose to a server, receiving, from the server, a second media frame generated by rendering based on the first media frame and the estimated pose, and second pose information associated with the second media frame, and Based on the second media frame and the second pose information, a third media frame is displayed at a third time point.
13. The UE according to claim 12, wherein: The first pose information includes predicted pose information representing a position and an orientation of an estimated pose and first temporal metadata associated with the estimated pose, The first time metadata includes at least one of the following: T1, indicates the first time point; T2.estimated, indicating the second time point; T1', indicating the time when the first pose information is sent from the UE to the server; and Render-to-photon latency, indicated as T2.actual - T3, where T2.actual indicates the time the photon has been displayed, and T3 indicates the time the rendering started.
14. The UE according to claim 12, wherein: The second posture information includes: Pose information, indicating the pose used for rendering; and second temporal metadata associated with the second media frame, The second time metadata includes at least one of the following: T1, indicates the first time point; T2.estimated, indicating the second time point; T3, indicating the actual time when the server starts rendering the first media frame; and T5 indicates the time when the second media frame is output from the server.
15. A server for supporting split rendering in a communication system, the server comprising: Network interface; Memory; as well as a processor coupled with a network interface and a memory, The memory stores one or more computer programs including computer-executable instructions that, when executed by the processor, cause the server to: Receiving a first media frame and first pose information associated with an estimated pose from a user equipment (UE), wherein the estimated pose corresponds to a pose of the UE at a target display time; generating a second media frame by rendering based on the first media frame and the estimated pose, and generating second pose information associated with the second media frame; and Send the second media frame and the second posture information to the UE.