Method for managing scene rendering based on confidence of attitude prediction

By acquiring and utilizing the predicted information of user pose and its confidence level, and selecting appropriate frames for display or reprojection, the latency and quality instability of scene rendering in pose prediction management in existing technologies are solved, and a more efficient and stable rendering process is achieved.

CN121986286APending Publication Date: 2026-05-05INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-08-07
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In augmented reality, extended reality, virtual reality, and mixed reality experiences, existing technologies struggle to effectively utilize the confidence level of pose prediction to manage scene rendering, resulting in rendering latency and unstable image quality.

Method used

By obtaining the predicted information of the user's pose and its confidence level, appropriate frames are selected for display or reprojection. Based on the pose confidence information, new frames are selected for rendering or previously rendered frames are displayed. The pose error is calculated and its effectiveness is determined, thus optimizing the rendering process.

Benefits of technology

It improves the efficiency of the rendering process and image quality, reduces latency, and enhances the stability and consistency of the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121986286A_ABST
    Figure CN121986286A_ABST
Patent Text Reader

Abstract

Some embodiments of a method may include obtaining a first predicted frame display time; obtaining first prediction attitude information, wherein the first prediction attitude information represents the prediction of the user attitude at the display time of the first prediction frame; determining first attitude confidence information, the first attitude confidence information indicating a confidence level of the first predicted attitude information; selecting a frame for display based at least in part on the first attitude confidence information; and causing display of the selected frame. When the attitude confidence information indicates a sufficiently high confidence, a newly rendered frame is obtained for display. When the attitude confidence information indicates a low confidence, a previously rendered frame may be used for display after re-projection.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims the benefit of European Patent Application No. EP23306698, filed on October 3, 2023, entitled “METHOD TO MANAGE SCENERENDERING BASED ON THE CONFIDENCE OF THE POSE PREDICTION”; and European Patent Application No. EP23306353, filed on August 9, 2023, entitled “METHOD TO MANAGE SCENERENDERING BASED ON THE CONFIDENCE OF THE POSEPREDICTION”, the entire contents of which are hereby incorporated by reference. Background Technology

[0002] In augmented reality (AR), extended reality (XR), virtual reality (VR), and / or mixed reality (MR) experiences, computer-generated virtual elements are displayed to the user, either in the user's real environment or in a virtual environment using various devices such as VR headsets, optical see-through glasses, or video see-through devices such as smartphones, tablets, and headsets.

[0003] Pose prediction is used to compensate for the round-trip time required to render a virtual scene. For each view, pose prediction allows the final render to be adjusted by taking into account the movement of the AR device during rendering computation.

[0004] For each view, pose information can be predicted at the start of the rendering loop based on previous / past pose values ​​at a given time. For example, embedded device sensors (such as depth or color cameras) and inertial measurement units (IMUs) can be used to collect pose information for user pose estimation. Summary of the Invention

[0005] The embodiments described herein include methods used in video encoding and decoding (collectively, “code processing”).

[0006] A first example method according to some embodiments may include: obtaining a first predicted frame display time; obtaining first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; obtaining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; and causing the selected frame to be displayed.

[0007] Some embodiments of the first example method may further include: obtaining a second predicted frame display time; and performing a reprojection of the selected frame based on the second predicted frame display time before causing the selected frame to be displayed.

[0008] In some embodiments of the first example method, in response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

[0009] Some embodiments of the first example method may also include rendering the selected frame based on the first predicted pose.

[0010] In some embodiments of the first example method, in response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

[0011] In some embodiments of the first example method, the first attitude confidence information includes multiple flags.

[0012] In some embodiments of the first example method, the first attitude confidence information includes at least a position validity flag and an orientation validity flag.

[0013] In some embodiments of the first example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame may be selected as the selected frame for display.

[0014] In some embodiments of the first example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame may be selected as the selected frame for display.

[0015] In some embodiments of the first example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0016] Some embodiments of the first example method may further include: obtaining second predicted pose information, the second predicted pose information representing a prediction of the user pose at the display time of the selected frame; and obtaining second pose confidence information, the second pose confidence information indicating the confidence level of the second predicted pose information.

[0017] Some embodiments of the first example method may further include: calculating the attitude error based on the difference between the first predicted attitude information and the second predicted attitude information.

[0018] Some embodiments of the first example method may further include: determining whether to calculate the attitude error based on the first attitude confidence information and the second attitude confidence information, wherein the attitude error is calculated in response to determining to calculate the attitude error.

[0019] Some embodiments of the first example method may further include: determining whether the attitude error is valid based on the first attitude confidence information and the second attitude confidence information.

[0020] In some embodiments of the first example method, obtaining the first attitude confidence information includes determining the first attitude confidence information.

[0021] A first example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0022] A second example method according to some embodiments may include: obtaining a first predicted frame display time; obtaining first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; determining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; transmitting the determined first pose confidence information to an edge application server, the determined first pose confidence information indicating the confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; and causing the selected frame to be displayed.

[0023] Some embodiments of the second example method may further include: obtaining a second predicted frame display time; and performing a reprojection of the selected frame based on the second predicted frame display time before causing the selected frame to be displayed.

[0024] In some embodiments of the second example method, the first attitude confidence information may include multiple flags.

[0025] In some embodiments of the second example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0026] In some embodiments of the second example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0027] In some embodiments of the second example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0028] In some embodiments of the second example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0029] Some embodiments of the second example method may further include: obtaining second predicted pose information, the second predicted pose information representing a prediction of the user pose at the display time of the selected frame; and obtaining second pose confidence information, the second pose confidence information indicating the confidence level of the second predicted pose information.

[0030] Some embodiments of the second example method may further include: calculating the attitude error based on the difference between the first predicted attitude information and the second predicted attitude information.

[0031] Some embodiments of the second example method may further include: determining whether to calculate the attitude error based on the first attitude confidence information and the second attitude confidence information, wherein the attitude error is calculated in response to determining to calculate the attitude error.

[0032] Some embodiments of the second example method may further include: determining whether the attitude error is valid based on the first attitude confidence information and the second attitude confidence information.

[0033] A second example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0034] A third example method according to some embodiments may include: receiving a first predicted frame display time; receiving first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; receiving first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; rendering the selected frame; and sending the rendered frame to an extended reality (XR) application device.

[0035] In some embodiments of the third example method, in response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

[0036] In some embodiments of the third example method, the selected frame is rendered based on the first predicted pose.

[0037] In some embodiments of the third example method, in response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

[0038] In some embodiments of the third example method, the first attitude confidence information may include multiple flags.

[0039] In some embodiments of the third example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0040] In some embodiments of the third example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0041] In some embodiments of the third example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0042] In some embodiments of the third example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0043] Some embodiments of the third example method may also include sending the first predicted frame display time to the extended reality (XR) application device.

[0044] A third example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0045] A fourth example method according to some embodiments may include: obtaining a first predicted frame display time; obtaining first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; transmitting the first predicted frame display time, the first predicted pose information, and information indicating the state of an extended reality (XR) view to an edge application server; obtaining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; and causing the selected frame to be displayed.

[0046] Some embodiments of the fourth example method may further include: obtaining a second predicted frame display time; and performing a reprojection of the selected frame based on the second predicted frame display time before causing the selected frame to be displayed.

[0047] In some embodiments of the fourth example method, the first attitude confidence information may include multiple flags.

[0048] In some embodiments of the fourth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0049] In some embodiments of the fourth example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0050] In some embodiments of the fourth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0051] In some embodiments of the fourth example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0052] Some embodiments of the fourth example method may further include: obtaining second predicted pose information, the second predicted pose information representing a prediction of the user pose at the display time of the selected frame; and obtaining second pose confidence information, the second pose confidence information indicating the confidence level of the second predicted pose information.

[0053] Some embodiments of the fourth example method may further include: calculating the attitude error based on the difference between the first predicted attitude information and the second predicted attitude information.

[0054] Some embodiments of the fourth example method may further include: determining whether to calculate the attitude error based on the first attitude confidence information and the second attitude confidence information, wherein the attitude error is calculated in response to determining to calculate the attitude error.

[0055] Some embodiments of the fourth example method may further include: determining whether the attitude error is valid based on the first attitude confidence information and the second attitude confidence information.

[0056] A fourth example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0057] A fifth example method according to some embodiments may include: receiving a first predicted frame display time; receiving the first predicted frame display time; receiving first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; receiving information indicating the state of an extended reality (XR) view; determining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; rendering the selected frame; and sending the rendered frame and the determined first pose confidence information to an extended reality (XR) application device.

[0058] In some embodiments of the fifth example method, in response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

[0059] In some embodiments of the fifth example method, the selected frame is rendered based on the first predicted pose.

[0060] In some embodiments of the fifth example method, in response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

[0061] In some embodiments of the fifth example method, the first attitude confidence information may include multiple flags.

[0062] In some embodiments of the fifth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0063] In some embodiments of the fifth example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0064] In some embodiments of the fifth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0065] In some embodiments of the fifth example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0066] Some embodiments of the fifth example method may also include sending the first predicted frame display time to the extended reality (XR) application device.

[0067] A fifth example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0068] A sixth example method / apparatus according to some embodiments may include: one or more processors configured to perform any of the methods listed above.

[0069] A seventh example method / apparatus according to some embodiments may include a computer-readable storage medium storing instructions for causing one or more processors to perform any of the methods listed above.

[0070] An eighth example method / apparatus according to some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any of the methods listed above.

[0071] An example computer-readable medium according to some embodiments may include instructions for causing one or more processors to perform any of the methods listed above.

[0072] In some embodiments of the example computer-readable medium, the computer-readable medium is a non-transitory storage medium.

[0073] An example computer program according to some embodiments may include instructions that, when the program is executed by one or more processors, cause the one or more processors to perform any of the methods listed above.

[0074] An example signal according to some embodiments may include a scene description file generated according to any of the methods listed above.

[0075] In another embodiment, an encoder device and a decoder device are provided to perform the methods described herein. The encoder device or decoder device may include a processor configured to perform the methods described herein. The device may include a computer-readable medium (e.g., a non-transitory medium) storing instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., a non-transitory medium) stores video encoded using any of the methods described herein.

[0076] One or more embodiments of this invention also provide a computer-readable storage medium having instructions stored thereon for performing bidirectional optical flow, encoding or decoding video data according to any of the methods described above. This embodiment also provides a computer-readable storage medium having a bitstream generated according to the methods described above stored thereon. This embodiment also provides a method and apparatus for transmitting a bitstream generated according to the methods described above. This embodiment also provides a computer program product including instructions for performing any of the described methods. Attached Figure Description

[0077] Figure 1A This is a schematic side view illustrating an example waveguide display that can be used with extended reality (XR) applications according to some embodiments.

[0078] Figure 1B This is a schematic side view illustrating example alternative display types that can be used with extended reality applications according to some embodiments.

[0079] Figure 1C This is a schematic side view illustrating example alternative display types that can be used with extended reality applications according to some embodiments.

[0080] Figure 1D This is a system diagram illustrating a set of example interfaces for a system according to some embodiments.

[0081] Figure 1E This is a system diagram illustrating an example communication system according to some embodiments.

[0082] Figure 2 is a message sequence diagram illustrating an example process for measuring attitude error and time error in attitude prediction.

[0083] Figure 3 This is a schematic timing diagram illustrating an example process of using a second prediction (T2.predicted2) based on the display time to improve prediction accuracy.

[0084] Figure 4 is a message sequence diagram illustrating an example process of using attitude confidence information according to some embodiments, wherein the confidence level is calculated and checked by an XR application.

[0085] Figure 5 is a message sequence diagram illustrating an example process of using attitude confidence information according to some embodiments, wherein the confidence level is checked by an edge application server.

[0086] Figure 6 is a message sequence diagram illustrating an example process of using attitude confidence information according to some embodiments, wherein the confidence level is calculated and checked by the edge application server.

[0087] Figure 7 This is a flowchart illustrating an example process, according to some embodiments, including frame selection using attitude confidence information.

[0088] Figure 8 This is a flowchart illustrating an example process, according to some embodiments, including frame selection using attitude confidence information.

[0089] Figure 9 This is a flowchart illustrating an example process, according to some embodiments, including attitude error determination using attitude confidence information.

[0090] Figure 10 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of predicted pose, according to some embodiments.

[0091] Figure 11 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of predicted pose, according to some embodiments.

[0092] Figure 12 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of predicted pose, according to some embodiments.

[0093] Figure 13 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of predicted pose, according to some embodiments.

[0094] Figure 14 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of predicted pose, according to some embodiments.

[0095] The entities, connections, arrangements, etc., depicted and described in conjunction with the various figures are presented by way of example rather than limitation. Therefore, any and all statements or other indications regarding what a particular figure “depicts,” what a particular element or entity in a particular figure “is” or “has,” and any and all similar statements that might be interpreted in isolation and out of context as absolute and therefore limiting, should only be correctly interpreted as constructively beginning with a clause such as “In at least one embodiment, …”. For the sake of simplicity and clarity, this implicit leading clause will not be repeated in the detailed description. Annoying . Detailed Implementation

[0096] Figure 1A This is a schematic side view illustrating an example waveguide display that can be used with extended reality (XR) applications according to some embodiments. An image is projected by an image generator 102. The image generator 102 can use one or more of a variety of techniques for projecting images. For example, the image generator 102 can be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light-emitting diode (LED) display (including organic LED (OLED) or micro LED (µLED) displays), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other types of image generators or light engines.

[0097] The light representing the image 112 generated by the image generator 102 is coupled into the waveguide 104 by a diffraction input coupler 188. The input coupler 188 diffracts the light representing the image 112 into one or more diffraction orders. For example, one of the rays 108 representing part of the bottom of the image is diffracted by the input coupler 188, and one of the diffraction orders 110 (e.g., second order) is at an angle capable of propagating through the waveguide 104 by total internal reflection. The image generator 102 displays the image according to the instructions of the control module 124, which operates to render image data, video data, point cloud data, or other displayable data.

[0098] At least a portion of the light 110, already coupled into waveguide 104 by diffraction input coupler 188, is coupled out of the waveguide by diffraction output coupler 114. At least some of the light coupled out of waveguide 104 replicates the incident angle of the light coupled into the waveguide. For example, in the illustration, output coupled rays 116a, 116b, and 116c replicate the angle of input coupled ray 108. Because the light leaving the output coupler replicates the direction of the light entering the input coupler, the waveguide essentially replicates the original image 112. The user's eye 118 can focus on the replicated image.

[0099] exist Figure 1A In the example, output coupler 114 outputs only a portion of the coupled light on each reflection, allowing a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). In this way, even if the eye is not perfectly aligned with the center of the output coupler, at least some of the light from each part of the image may reach the user's eye. For example, if eye 118 moves downwards, beam 116c may enter the eye even if beams 116a and 116b do not, so the user can still perceive the bottom of image 112 despite the positional shift. Therefore, output coupler 114 partially functions as an outgoing pupil dilator in the vertical direction. The waveguide may also include one or more additional outgoing pupil dilators (…). Figure 1A (Not shown in the image) to expand the exit pupil in the horizontal direction.

[0100] In some embodiments, waveguide 104 is at least partially transparent relative to light originating from outside the waveguide display. For example, at least some of the light 120 from a real-world object (such as object 122) passes through waveguide 104, allowing the user to see the real-world object when using the waveguide display. Multiple diffraction orders and thus multiple images will exist as the light 120 from the real-world object also passes through diffraction grating 114. To minimize the visibility of multiple images, it is desirable that the zeroth-order diffraction (not deflected by 114) has high diffraction efficiency for both light 120 and the zeroth order, while higher diffraction orders have lower energy. Therefore, in addition to expanding and output coupling the virtual image, output coupler 114 is preferably configured to allow the zeroth order of the real image to pass through. In such embodiments, the image displayed by the waveguide display may appear to be superimposed on the real world.

[0101] Figure 1BThis is a schematic side view illustrating example alternative display types that can be used with extended reality applications according to some embodiments. In the XR head-mounted display device 130, a control module 132 controls a display 134 (which may be an LCD) to display images. The head-mounted display includes a partially reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) the image displayed on the LCD to make the image visible to the user. The partially reflective surface 136 also allows at least some external light to pass through, thereby allowing the user to see their surroundings.

[0102] Figure 1C This is a schematic side view illustrating example alternative display types that can be used with extended reality applications according to some embodiments. In the XR head-mounted display device 140, a control module 142 controls a display 144 (which may be an LCD) to display an image. The image is focused by one or more lenses of a display optics 146 to make the image visible to the user. Figure 1C In this example, external light does not reach the user's eyes directly. However, in some such embodiments, an external camera 148 may be used to capture images of the external environment and display such images on a display 144 along with any virtual content that may also be displayed.

[0103] The embodiments described herein are not limited to any particular type or structure of XR display device.

[0104] Figure 1D This is a system diagram illustrating a set of example interfaces for a system according to some embodiments. Interfaces such as... Figure 1D The system shown implements an extended reality display device and its control electronics. System 150 may be embodied as an apparatus including the various components described below and configured to perform one or more aspects described in this document. Examples of such apparatus include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 150 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 150 are distributed across multiple ICs and / or discrete components. In various embodiments, system 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 150 is configured to implement one or more aspects described in this document.

[0105] System 150 includes at least one processor 152 configured to execute instructions loaded thereon to implement various aspects described herein, such as those described. Processor 152 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 150 includes at least one memory 154 (e.g., a volatile memory device and / or a non-volatile memory device). System 150 may include a storage device 158, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 158 may include internal storage, attached storage (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0106] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 156 may include its own processor and memory. The encoder / decoder module 156 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device can include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 156 may be implemented as a separate element of system 150, or it may be incorporated within processor 152 as a combination of hardware and software known to those skilled in the art.

[0107] Program code to be loaded onto processor 152 or encoder / decoder 156 to execute the various aspects described herein may be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. According to various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder / decoder module 156 may store one or more of various items during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0108] In some embodiments, the memory within processor 152 and / or encoder / decoder module 156 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 152 or encoder / decoder module 152) is used for one or more of these functions. External memory may be memory 154 and / or storage device 158, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding, i.e., a new standard developed by JVET (Joint Video Experts Group)).

[0109] Inputs to the components of system 150 can be provided through various input devices, as indicated in box 172. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1C Other examples not shown include composite video.

[0110] In various embodiments, the input device of block 172 has corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select desired data packets. The RF section in various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various embodiments rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as insert amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0111] Additionally, USB and / or HDMI terminals may include corresponding interface processors for connecting system 150 to other electronic devices via USB and / or HDMI connections. It should be understood that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 152, as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 152. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 152 and an encoder / decoder 156 operating in conjunction with memory and storage elements, to process the data stream as needed for presentation on an output device.

[0112] Various components of system 150 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data between them using suitable connection means 174 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).

[0113] System 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 162. The communication interface 160 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 162 may be implemented, for example, within a wired and / or wireless medium.

[0114] In various embodiments, data is streamed to or otherwise provided to system 150 using a wireless network, such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these embodiments, the Wi-Fi signal is received via a communication channel 162 and a communication interface 160 adapted for Wi-Fi communication. The communication channel 162 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other top-level communications. Other embodiments use a set-top box to provide streaming data to system 150, delivering data via an HDMI connection to input box 172. Still other embodiments use an RF connection to input box 172 to provide streaming data to system 150. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0115] System 150 can provide output signals to various output devices, including display 176, speaker 178, and other peripheral devices 180. Display 176 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 176 can be used in televisions, tablets, laptops, mobile phones, or other devices. Display 176 can also be integrated with other components (e.g., in a smartphone) or standalone (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 180 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 180 based on the output of system 150 to provide functionality. For example, an optical disc player performs the function of playing the output of system 150.

[0116] In various embodiments, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention are used to transmit control signals between system 150 and display 176, speaker 178, or other peripheral devices 180. Output devices may be communicatively coupled to system 150 via dedicated connections through corresponding interfaces 164, 166, and 168. Alternatively, output devices may be connected to system 150 via communication interface 160 using communication channel 162. Display 176 and speaker 178 may be integrated into a single unit along with other components of system 150 in an electronic device, such as a television. In various embodiments, display interface 164 includes a display driver, such as a timing controller (TCon) chip.

[0117] For example, if the RF input section 172 is part of a separate set-top box, the display 176 and speaker 178 can alternatively be separate from one or more of the other components. In various embodiments where the display 176 and speaker 178 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0118] System 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscope sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors can be used to determine information such as the user's position and orientation. Where system 150 is used as a control module (such as control modules 124, 132) for an extended reality display, the user's position and orientation can be used to determine how to render image data so that the user perceives the correct portion of a virtual object or scene from the correct perspective. In the case of a head-mounted display device, the position and orientation of the device itself can be used to determine the user's position and orientation for rendering virtual content. In the case of other display devices (such as telephones, tablets, computer monitors, or televisions), other inputs can be used to determine the user's position and orientation for rendering content. For example, a user can select and / or adjust the desired viewpoint and / or viewing direction by using a touchscreen, keypad or keyboard, trackball, joystick, or other inputs. Where the display device has sensors such as accelerometers and / or gyroscopes, the viewpoint and orientation can be selected and / or adjusted based on the movement of the display device for rendering content.

[0119] The embodiments can be executed by processor 152 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, memory 154 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 152 can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0120] Figure 1E This diagram illustrates an example communication system 182 that can implement one or more of the disclosed embodiments. Communication system 182 can be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. Communication system 182 enables multiple wireless users to access such content through shared system resources including wireless broadband. For example, communication system 182 can employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0121] like Figure 1EAs shown, the communication system 182 may include wireless transmit / receive units (WTRUs) 184a, 184b, 184c, 184d, RAN 186, CN 188, public switched telephone network (PSTN) 190, Internet 192, and other networks 194. However, it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 184a, 184b, 184c, and 184d can be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 184a, 184b, 184c, and 184d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 184a, 184b, 184c, and 184d may be interchangeably referred to as WTRUs.

[0122] The communication system 182 may also include base station 196a and / or base station 196b. Each of base stations 196a and 196b may be any type of device configured to wirelessly interface with at least one of WTRUs 184a, 184b, 184c, and 184d to facilitate access to one or more communication networks such as CN 188, the Internet 192, and / or other networks 194. For example, base stations 196a and 196b may be base transceiver stations (BTS), Node-B, eNode B, home Node BB, home eNode B, gNB, NR Node B, site controllers, access points (APs), wireless routers, etc. Although base stations 196a and 196b are each depicted as a single element, it will be understood that base stations 196a and 196b may include any number of interconnected base stations and / or network elements.

[0123] Base station 196a may be part of RAN 186, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 196a and / or base station 196b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a specific geographic area that may be relatively fixed or may change over time. A cell may also be divided into cell sectors. For example, the cell associated with base station 196a may be divided into three sectors. Therefore, in one embodiment, base station 196a may include three transceivers, i.e., one transceiver per sector of the cell. In embodiments, base station 196a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0124] Base stations 196a and 196b can communicate with one or more of WTRUs 184a, 184b, 184c, and 184d via air interface 198, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 198 can be established using any suitable radio access technology (RAT).

[0125] More specifically, as described above, communication system 182 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 196a in RAN 186 and WTRUs 184a, 184b, 184c can implement radio technologies, such as using Wideband CDMA (WCDMA) to establish Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) for air interface 198. WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0126] In the embodiments, base station 196a and WTRUs 184a, 184b, 184c may implement radio technologies, such as using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) to establish Evolved UMTS Terrestrial Radio Access (E-UTRA) for air interface 198.

[0127] In the embodiments, base station 196a and WTRUs 184a, 184b, 184c may implement radio technologies, such as using New Radio (NR) to establish NR radio access for air interface 198.

[0128] In the embodiments, base station 196a and WTRUs 184a, 184b, and 184c can implement various radio access technologies. For example, base station 196a and WTRUs 184a, 184b, and 184c can, for instance, use the dual connectivity (DC) principle to jointly implement LTE radio access and NR radio access. Therefore, the air interface used by WTRUs 184a, 184b, and 184c can be characterized by various types of radio access technologies and / or by transmissions sent to / from various types of base stations (e.g., eNBs and gNBs).

[0129] In other embodiments, base station 196a and WTRUs 184a, 184b, 184c may implement radio technologies such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., Global System for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.

[0130] Figure 1EBase station 196b can be, for example, a wireless router, a home Node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for use by drones), roads, etc. In one embodiment, base station 196b and WTRUs 184c, 184d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 196b and WTRUs 184c, 184d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 196b and WTRUs 184c, 184d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1E As shown, base station 196b can be directly connected to Internet 192. Therefore, base station 196b does not need to access Internet 192 via CN 188.

[0131] RAN 186 can communicate with CN 188, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 184a, 184b, 184c, and 184d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 188 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although Figure 1E As not shown, but will be understood, RAN 186 and / or CN 188 can communicate directly or indirectly with other RANs that use the same RAT as or a different RAT than RAN 186. For example, in addition to connecting to RAN 186, which may be utilizing NR radio technology, CN 188 can also communicate with another RAN (not shown) using GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0132] CN 188 can also serve as a gateway for WTRUs 184a, 184b, 184c, and 184d to access PSTN 190, the Internet 192, and / or other networks 194. PSTN 190 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 192 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 194 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 194 may include another CN connected to one or more RANs, which may use the same RAT as RAN 186 or a different RAT.

[0133] Some or all of the WTRUs 184a, 184b, 184c, and 184d in communication system 182 may include multi-mode capabilities (e.g., WTRUs 184a, 184b, 184c, and 184d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1E The WTRU 184c shown can be configured to communicate with a base station 196a that can use cellular-based radio technology and with a base station 196b that can use IEEE 802 radio technology.

[0134] An overview of pose measurement and rendering This article uses the OpenXR framework to describe example implementations, but it is envisioned that the principles described herein can be applied to rendering using other frameworks.

[0135] In the example embodiment, pose prediction and tracking of the AR display device are managed by a module called the XR runtime. Within the runtime module using the OpenXR framework, the function `xrLocateViews` is called to obtain pose values ​​at a specific time. This function returns the associated state. For example, `xrLocateViews` can accept the display time of the returned result data as input. Applications use `xrLocateViews` to retrieve viewer pose and projection parameters for rendering each view, to be used in the compositing projection layer. In addition to the predicted pose information, the `xrLocateViews` function can also set the values ​​of four flags. In some embodiments, the values ​​of the four flags or the flags themselves can be read to obtain specific information about the pose estimation state. Different implementations may use different techniques to provide the state information.

[0136] In the example embodiment, the flag is used to provide information indicating the confidence level of the pose estimation. In other embodiments, different formats may be used to provide such information. In some embodiments, the confidence information is used to at least partially determine the procedural flow, call flow, and / or pose error measurement for rendering frames.

[0137] In an example embodiment (e.g., see Figure 4), the confidence level is calculated and checked locally by the XR application (which may be hosted by the UE (User Equipment)). In another example embodiment (e.g., see Figure 5), the confidence level is calculated locally by the XR application (e.g., hosted by the UE) and sent to an edge application server for checking. In yet another example embodiment (e.g., see Figure 6), view status flags for information validity and tracking of location and orientation (e.g., with OpenXR) will be given. XrViewStateFlags The flag is sent to the edge application server, which calculates and checks the confidence level (instead of the XR application doing so).

[0138] Figure 2 is a message sequence diagram illustrating an example process for measuring pose error and time error in pose prediction, where the confidence level is calculated by the XR application. Figure 2 shows a measurement process 200 for a cloud-based rendered scene as an example. The example in Figure 2 does not change the call flow based on the confidence level of the pose estimation, and is described here to illustrate embodiments that do utilize information about the confidence level of the pose estimation. Not all steps described below need to be performed in all embodiments. The XR runtime and XR application can be on the same device, such as a UE (User Equipment), or on different devices, such as AR glasses (which host the XR runtime) and a UE (which host the XR application).

[0139] In step 1 (208), XR application 204 estimates the round-trip time (RTT) between XR application 204 and edge application server (EAS) 206.

[0140] In step 2 (210), the XR application 204 queries the next display time. In embodiments using OpenXR, this (and step 3) can be achieved by calling the xrWaitFrame function.

[0141] In step 3 (212), XR runtime 202 responds with information indicating the next display time.

[0142] In step 4 (214), XR applies 204 to predict the display time, i.e., the initial prediction, and uses "initial" because a second prediction / estimation will be made later. This predicted display time is referred to as T2.predicted1.

[0143] In step 5 (216), the XR application 204 queries the predicted pose at the initial prediction display time T2.predicted1. This step and step 7 can be achieved by calling the function xrLocateViews in OpenXR.

[0144] In step 6 (218), XR runtime 202 predicts the pose, and this prediction occurs at time T1.

[0145] In step 7 (220), XR runtime 202 returns the predicted pose (P.predicted1).

[0146] In step 8 (222), the XR application 204 sends the predicted pose (P.predicted1) and the associated initial predicted display time (T2.predicted1) to the edge application server (EAS) 206 for performing rendering. In some embodiments, the user device or other system is used to perform rendering instead of using EAS 206.

[0147] In step 9 (224), EAS 206 renders the predicted pose (P.predicted1) and compresses the rendered frames.

[0148] In step 10 (226), EAS 206 returns the rendered frame along with the initial predicted display time (T2.predicted1) to the XR application.

[0149] In step 11 (228), the XR application 204 sends the rendered frame to the XR runtime, for example, via a swapchain. In some embodiments, this is done by calling the xrReleaseSwapchainImage function in OpenXR. The XR application passes the display time for rendering the frame, and this can be done by calling the xrEndFrame function in OpenXR.

[0150] In step 12 (230), the XR application 204 queries to predict the display time. This is to obtain a more accurate prediction of the display time than the prediction in step 4, because the future time to be predicted is shorter at this point.

[0151] In step 13 (232), XR runtime 202 returns an updated prediction of the display time (T2.predicted2).

[0152] In step 14 (234), XR runtime 202 performs a reprojection for attitude correction. The actual display playback time is referred to as T2.actual.

[0153] In step 15 (236), the XR application queries the pose associated with the update prediction (T2.predicted2) for the display time. This can be achieved by calling the xrLocateViews function in OpenXR.

[0154] In step 16 (238), XR runtime 202 performs attitude estimation.

[0155] In step 17 (240), XR runtime 202 returns the pose estimate (P.predicted2).

[0156] In step 18 (242), XR application 204 calculates the attitude error estimate (P.predicted1 - P.predicted2) and the time error estimate (T2.predicted1 - T2.predicted2).

[0157] According to some embodiments, the user equipment may be, for example, a WTRU, such as... Figure 1E 184a, 184b, 184c, and 184d. According to some embodiments, the user equipment may be implemented, for example, as... Figure 1D The example system 150 is an example or a part thereof. It should be understood that these are merely examples, and the user equipment is not limited to such example implementations.

[0158] According to some embodiments, the edge application server can be, for example, as... Figure 1E The shown network is part of or communicates with a server. According to some embodiments, the edge application server can communicate with... Figure 1E The edge application server communicates with one or more of the WTRUs (e.g., user equipment, such as head-mounted displays or smartphones) 184a, 184b, 184c, and 184d. According to some embodiments, the edge application server can be implemented, for example, as... Figure 1D The example system 150 is an example of or part of that example. It should be understood that these are merely examples, and the edge application server is not limited to such example implementations.

[0159] Figure 3 This is a schematic timing diagram illustrating an example process of using a second prediction (T2.predicted2) of the display time to improve prediction accuracy. In this example timing diagram 300, two queries are used to predict the display time of the same frame.

[0160] The first query occurs at the first "now" time 312 in step 2, and the query result is used to determine the target display time for the rendering process in step 4. In step 4, for some embodiments, the XR application predicts the initial display time based on RTT 314 and frame rate (T2.predicted1(306)). Figure 3 As shown, the next predicted display time 304 is equal to the (previous) display time 302 plus 1 / FrameRate (308). In addition, T2.predicted1 (306) is equal to the next predicted display time 304 plus 1 / FrameRate (310).

[0161] The second query occurs at a second “now” time 320, which is closer to the actual display time, as shown in steps 12 to 13, and thus provides higher accuracy. Figure 3 The relationship between the actual display times 316 and 318 and the prediction T2.predicted2(322) is shown in the figure.

[0162] Within the OpenXR framework, the xrLocateViews function is defined as follows:

[0163] The parameters used by the xrLocateViews function can be described as follows: · session It is provided XrSession The handle.

[0164] · viewLocateInfo It refers to validity XrViewLocateInfo A pointer to the structure, which includes information indicating the estimated display time.

[0165] · viewState It is an output structure that contains viewer state information, and in some embodiments, it may include flags for conveying confidence information.

[0166] · viewCapacityInput This is an input parameter that specifies the capacity of the view array.

[0167] · viewCountOutput It is the output parameter for the valid count of the identifier view.

[0168] · views yes XrView An array that can be used to provide attitude information.

[0169] The estimated display time is provided to the xrLocateViews function via the viewLocateInfo parameter. In the "view" parameter of type XRView, the xrLocateViews function provides pose information for the estimated display time. This estimated display time can be the target display time for a given frame.

[0170] The function xrLocateViews returns an array of XrView elements (one XrView element for each view of the specified view configuration type) and an XrViewState containing additional state data shared across all views. XrView elements can be configured to include pose information (e.g., information of type XrPosef).

[0171]

[0172] The XrViewState structure contains additional state data, as follows:

[0173] structure XrViewState Includes the field XrViewStateFlags XrViewStateFlags . XrViewStateFlags The field contains information indicating the state of all views. XrViewStateFlagBits The bitmask is as follows:

[0174] Typically, these status flags are used to indicate the following information: • XR_VIEW_STATE_ORIENTATION_VALID_BIT indicates whether all XrView orientations contain valid data.

[0175] • XR_VIEW_STATE_POSITION_VALID_BIT indicates whether all XrView locations contain valid data.

[0176] • XR_VIEW_STATE_ORIENTATION_TRACKED_BIT indicates whether all XrView orientations represent actively tracked orientations.

[0177] • XR_VIEW_STATE_POSITION_TRACKED_BIT indicates whether all XrView positions represent actively tracked positions.

[0178] In the example embodiment, these status flags are also used to indicate the confidence level of the attitude information.

[0179] Signaling and use of confidence information In the example call flow shown in Figure 2, the function xrLocateViews It is called twice (step 5 and step 15) to obtain the predicted pose.

[0180] At steps 7 and 17, attitude estimation and information regarding the validity and tracking of position and orientation information are obtained. This information regarding validity and tracking can be used... XrViewStateFlags Provided when marked.

[0181] In example embodiments, the XR runtime provides information indicating the confidence level of the pose estimation. This information can be used to manage quality of experience (QOE) and rendering. In some embodiments using the OpenXR framework, existing... XrViewStateFlags The flags provide information indicating the confidence level of the attitude estimation, namely: ·XR_VIEW_STATE_ORIENTATION_VALID_BIT ·XR_VIEW_STATE_POSITION_VALID_BIT ·XR_VIEW_STATE_POSITION_TRACKED_BIT ·XR_VIEW_STATE_ORIENTATION_TRACKED_BIT In OpenXR, these flags are provided as bitmasks. To obtain the value, the mask is combined with XrViewStateFlagBits (e.g., XrViewStateFlagBits&XR_VIEW_STATE_ORIENTATION_VALID_BIT).

[0182] For simplicity, XR_VIEW_STATE_POSITION_VALID_BIT should be understood as the result of a masking operation in the following text. The same applies to the other three flags. For descriptive purposes, 0 is used for "invalid" and 1 is used to indicate "valid".

[0183] The four state flags together can distinguish sixteen (24) different states. However, so many different states are not needed to provide the necessary information about the validity of pose information and tracking. Therefore, the general use of the four XrViewStateFlags includes redundant information. For example, VALID information can be considered more important than TRACKED information. That is, if VALID equals 0, TRACKED is irrelevant; when VALID equals 1, TRACKED can equal 0. In addition to the general information about validity and tracking, other conventions can be adopted to allow the four XrViewStateFlags to provide confidence level information.

[0184] In the example embodiment, the XrViewStateFlags provided by the xrLocateViews function (e.g., in step 7) are used to signal the confidence level of the predicted pose information. As an example, flag values ​​in XrViewStateFlags can be used to signal the confidence level, as indicated in Table 1. However, in different embodiments, different confidence levels can be assigned to combinations of flags. In the example in Table 1, “X” indicates that, given other flag values ​​in the same row, the flag value marked “X” does not affect the determined confidence level.

[0185]

[0186] Table 1 As described above, in the example embodiments (e.g., see Figure 4), the confidence level is calculated and checked locally by the XR application (which may be hosted by the UE (User Equipment)). In another example embodiment (e.g., see Figure 5), the confidence level is calculated locally by the XR application (e.g., hosted by the UE) and sent to an edge application server for checking. In yet another example embodiment (e.g., see Figure 6), view state flags (e.g., with OpenXR's XrViewStateFlags flag) indicating the validity of the information and the tracking of position and orientation are sent to the edge application server, which calculates and checks the confidence level (instead of the XR application doing so). The example embodiments and the categories of embodiments are merely examples of how some example processing can be transferred to different degrees between the XR application (e.g., hosted by the UE) and the edge application server, and the embodiments described herein are not limited to these implementations.

[0187] Figure 4 is a message sequence diagram illustrating an example process of using pose confidence information according to some embodiments, wherein the confidence level is calculated and checked by an XR application. In some embodiments, by xrLocateViews The confidence level indicated by the returned flag (e.g., step 7) is used to modify the rendering process of the call flow shown in Figure 2. For example, some embodiments include step 7bis (422) of calculating and checking the confidence level as shown in Figure 4. In some embodiments, the frame to be displayed depends on the confidence level. For example, according to the call flow shown in Figure 2, in response to a confidence level at or above a threshold (which may be a predetermined threshold), a predicted pose can be used to render the scene. However, in response to a confidence level below the threshold, in some embodiments, a predicted pose is not used to render the scene. Examples of rendering processes that depend on the confidence level are provided in Table 2. Regardless of the confidence level, steps 12 (432), 13 (434), and 14 (436) are still performed in the example embodiment of message flow 400 shown in Figure 4.

[0188] In step 1 (408), XR application 404 estimates the round-trip time (RTT) between XR application 404 and edge application server (EAS) 406.

[0189] In step 2 (410), the XR application returns a 404 query for the next display time. In embodiments using OpenXR, this (and step 3) can be achieved by calling the xrWaitFrame function.

[0190] In step 3 (412), XR runtime 402 responds with information indicating the next display time.

[0191] In step 4 (414), XR applies 404 to predict the display time, i.e., the initial prediction, and uses "initial" because a second prediction / estimation will be performed later. This predicted display time is referred to as T2.predicted1.

[0192] In step 5 (416), the XR application performs a 404 query to determine the predicted pose at the initial prediction display time T2.predicted1. This step, along with step 7, can be achieved by calling the function xrLocateViews in OpenXR.

[0193] In step 6 (418), XR runtime 402 predicts the pose, and the prediction occurs at time T1.

[0194] In step 7 (420), XR runtime 402 returns the predicted pose (P.predicted1).

[0195] In step 7bis (422), the XR application 404 calculates and checks the confidence level. If the confidence level is below a pre-configured threshold (e.g., 5 / 8), the XR application 404 may use the previously rendered frame, depending on the strategy. Therefore, steps 8 (424), 9 (426), and 10 (428) are skipped. In this case, step 11 (430) is performed using the previously rendered frame.

[0196] In step 8 (424), the XR application 404 sends the predicted pose (P.predicted1) and the associated initial predicted display time (T2.predicted1) to the edge application server (EAS) 406 for performing rendering. In some embodiments, the user device or other system is used to perform rendering instead of using EAS 406.

[0197] In step 9 (426), EAS 406 renders the predicted pose (P.predicted1) and compresses the rendered frames.

[0198] In step 10 (428), EAS 406 returns the rendered frame along with the initial predicted display time (T2.predicted1) to the XR application.

[0199] In step 11 (430), the XR application 404 sends the rendered frame to the XR runtime, for example, via a swap chain. In some embodiments, this is done by calling the xrReleaseSwapchainImage function in OpenXR. The XR application passes the display time for rendering the frame, and this can be done by calling the xrEndFrame function in OpenXR.

[0200] In step 12 (432), XR applies a 404 query to predict the display time. This is to obtain a more accurate prediction of the display time than the prediction in step 4, because the future time to be predicted is shorter at this point.

[0201] In step 13 (434), XR runtime 402 returns an updated prediction of the display time (T2.predicted2).

[0202] In step 14 (436), XR runtime 402 performs a reprojection for attitude correction. The actual display playback time is referred to as T2.actual (438).

[0203] In step 15 (440), the XR application performs a 404 query to determine the pose associated with the update prediction (T2.predicted2) for the display time. This can be achieved by calling the xrLocateViews function in OpenXR.

[0204] In step 16 (442), XR runtime 402 performs attitude estimation.

[0205] In step 17 (444), XR runtime 402 returns the pose estimate (P.predicted2).

[0206] In step 18 (446), XR application 404 calculates the attitude error estimate (P.predicted1 - P.predicted2) and the time error estimate (T2.predicted1 - T2.predicted2).

[0207]

[0208] Table 2. Rendering based on confidence level In the example in Table 2, the confidence level threshold is 5 / 8, but in other embodiments, other thresholds may be used.

[0209] In some embodiments, information about the confidence level of the predicted pose can be inferred directly from the values ​​of the flags; for example, it is not necessary to calculate a specific numerical confidence level. For instance, in response to determining that both XR_VIEW_STATE_POSITION_VALID_BIT and XR_VIEW_STATE_ORIENTATION_VALID_BIT are equal to "1", it can be determined that the confidence level of the predicted pose is sufficiently high (e.g., at or above a threshold), while if either or both of these bits are equal to "0", it can be determined that the confidence level of the predicted pose is low (e.g., below a threshold). Alternatively, other logical combinations of these bits can be used to determine whether the confidence level of the predicted pose is high enough to render a scene using that predicted pose.

[0210] Other example embodiments may offload various processing features to the edge application server, as shown in the following examples.

[0211] Figure 5 is a message sequence diagram illustrating an example process of using attitude confidence information according to some embodiments, wherein the confidence level is checked by an edge application server. Figure 5 shows an example message sequence 500.

[0212] In step 1 (508), XR application 504 estimates the round-trip time (RTT) between XR application 504 and Edge Application Server (EAS) 506. In step 2 (510), XR application 504 queries the next display time. In embodiments using OpenXR, this (and step 3) can be achieved by calling the xrWaitFrame function. In step 3 (512), XR runtime 502 responds with information indicating the next display time. In step 4 (514), XR application 504 predicts the display time, i.e., the initial prediction, and uses "initial" because a second prediction / estimation will be performed later. This predicted display time is referred to as T2.predicted1. In step 5 (516), XR application 504 queries the predicted pose at the initial predicted display time T2.predicted1. This step and step 7 can be achieved by calling the xrLocateViews function in OpenXR. In step 6 (518), XR runtime 502 predicts the pose, and the prediction occurs at time T1. In step 7 (520), XR runtime 502 returns the predicted pose (P.predicted1).

[0213] In step 7bis (522), the XR application 504 calculates the confidence level. In step 8 (524), the confidence level, along with the predicted pose (P.predicted1) and the predicted initial display time (T2.predicted1), is sent to the edge application server 506. In step 8bis (526), ​​the edge application server 506 checks the received confidence level. If the confidence level is below a pre-configured threshold (e.g., 5 / 8), the edge application server 506 may use the previously rendered frame, according to a policy. Therefore, step 9 (528) is skipped, and step 10 (530) is performed using the previously rendered frame. Steps 11 (532), 12 (534), and 13 (536) are retained.

[0214] In step 11 (532), the XR application 504 sends the rendered frame to the XR runtime, for example, via a swapchain. In some embodiments, this is done by calling the xrReleaseSwapchainImage function in OpenXR. The XR application passes the display time for rendering the frame, and this can be done by calling the xrEndFrame function in OpenXR. In step 12 (534), the XR application 504 queries the predicted display time. This is to obtain a more accurate prediction of the display time than the prediction in step 4, since the future time to be predicted is shorter at this point. In step 13 (536), the XR runtime 502 returns an updated prediction of the display time (T2.predicted2). In step 14 (538), the XR runtime 502 performs a reprojection for pose correction. The actual display playback time is referred to as T2.actual (540). In step 15 (542), the XR application 504 queries the pose associated with the updated prediction of the display time (T2.predicted2). This can be achieved by calling the xrLocateViews function in OpenXR. In step 16 (544), XR runtime 502 performs attitude estimation. In step 17 (546), XR runtime 502 returns the attitude estimate (P.predicted2). In step 18 (548), XR application 504 calculates the attitude error estimate (P.predicted1 - P.predicted2) and the time error estimate (T2.predicted1 - T2.predicted2).

[0215] Figure 6 is a message sequence diagram illustrating an example process of using attitude confidence information according to some embodiments, wherein the confidence level is calculated and checked by an edge application server. Figure 6 shows an example message sequence 600.

[0216] In step 1 (608), XR application 604 estimates the round-trip time (RTT) between XR application 404 and Edge Application Server (EAS) 606. In step 2 (610), XR application 604 queries the next display time. In embodiments using OpenXR, this (and step 3) can be achieved by calling the xrWaitFrame function. In step 3 (612), XR runtime 602 responds with information indicating the next display time. In step 4 (614), XR application 604 predicts the display time, i.e., the initial prediction, and uses "initial" because a second prediction / estimation will be performed later. This predicted display time is referred to as T2.predicted1. In step 5 (616), XR application 604 queries the predicted pose at the initial predicted display time T2.predicted1. This step and step 7 can be achieved by calling the xrLocateViews function in OpenXR. In step 6 (618), XR runtime 602 predicts the pose, and the prediction occurs at time T1. In step 7 (620), XR runtime 602 returns the predicted pose (P.predicted1).

[0217] In step 8 (622), XR application 604 sends a predicted pose (P.predicted1) with view state flags (e.g., XrViewStateFlags flag) and a predicted initial display time (T2.predicted1). In step 8bis, edge application server 606 uses the received view state flags to calculate a confidence level. Edge application server 606 checks the confidence level. If the confidence level is below a pre-configured threshold (e.g., 5 / 8), then, according to the policy, edge application server 606 may use the previously rendered frame. Therefore, step 9 (624) is skipped. In this case, step 10 (626) is performed with the previously rendered frame. Edge application server 606 may send the calculated confidence level along with the rendered frame, which can be used in step 18 (644) to calculate the overall confidence level. Steps 11 (628), 12 (630), and 13 (632) are retained.

[0218] In step 11 (628), the XR application 604 sends the rendered frame to the XR runtime, for example, via a swapchain. In some embodiments, this is done by calling the xrReleaseSwapchainImage function in OpenXR. The XR application passes the display time for rendering the frame, and this can be done by calling the xrEndFrame function in OpenXR. In step 12 (630), the XR application 604 queries the predicted display time. This is to obtain a more accurate prediction of the display time than the prediction in step 4, since the future time to be predicted is shorter at this point. In step 13 (632), the XR runtime 602 returns an updated prediction of the display time (T2.predicted2). In step 14 (634), the XR runtime 602 performs a reprojection for pose correction. The actual display playback time is referred to as T2.actual (636). In step 15 (638), the XR application 604 queries the pose associated with the updated prediction of the display time (T2.predicted2). This can be achieved by calling the xrLocateViews function in OpenXR. In step 16 (640), the XR runtime 602 performs attitude estimation. In step 17 (642), the XR runtime 602 returns the attitude estimate (P.predicted2). In step 18 (844), the XR application 604 calculates the attitude error estimate (P.predicted1 - P.predicted2) and the time error estimate (T2.predicted1 - T2.predicted2).

[0219] In other words, the differences between Figures 4, 5, and 6 can be seen in steps 7bis, 8, 8bis, and 10 (for some embodiments, step 7bis applies only to Figures 4 and 5, and step 8bis applies only to Figures 5 and 6). Steps 1 to 7 are the same in Figures 4, 5, and 6. Additionally, step 9 is the same in Figures 4, 5, and 6. Furthermore, steps 11 to 18 are the same in Figures 4, 5, and 6. In some embodiments, previously rendered frames can be used for display, and steps 8, 9, and 10 of Figures 4, 5, and 6, as well as step 8bis of Figure 6 (if applicable), can be skipped, and step 11 of Figures 4, 5, and 6 can be performed using the previously rendered frames.

[0220] In some embodiments, as shown in step 7bis of Figure 4, the XR application calculates and checks the confidence level. In some embodiments, as shown in step 7bis of Figure 5, the XR application calculates the confidence level. In some embodiments, as shown in Figure 6, there is no step 7bis.

[0221] In some embodiments, as shown in step 8 of Figure 4, the XR application sends the predicted pose P.predicted1 and the predicted initial display time (T2.predicted1) to the edge application server. In some embodiments, as shown in step 8 of Figure 5, the XR application sends the predicted pose P.predicted1 with a confidence level and the predicted initial display time (T2.predicted1) to the edge application server. In some embodiments, as shown in step 8 of Figure 6, the XR application sends the predicted pose P.predicted1 with the view state XrViewStateFlags flag and the predicted initial display time (T2.predicted1) to the edge application server.

[0222] In some embodiments, as shown in Figure 4, step 8bis is omitted. In some embodiments, as shown in step 8bis in Figure 5, the edge application server checks the confidence level. In some embodiments, as shown in step 8bis in Figure 6, the edge application server calculates and checks the confidence level.

[0223] In some embodiments, as shown in step 10 of Figure 4, the edge application server returns the rendered frame and the predicted initial display time (T2.predicted1) to the XR application. In some embodiments, as shown in step 10 of Figure 5, the edge application server returns the rendered frame and the predicted initial display time (T2.predicted1) to the XR application. In some embodiments, as shown in step 10 of Figure 6, the edge application server returns the rendered frame, the predicted initial display time (T2.predicted1), and the confidence level to the XR application.

[0224] Figure 7 This is a flowchart illustrating an example process, according to some embodiments, including frame selection using attitude confidence information. Figure 7 The method performed in some embodiments is illustrated. At 702, a first predicted display time is obtained, which represents a first prediction of when a particular frame will be displayed. This prediction may be based on the frame rate and an estimate of the round-trip time (RTT) used to obtain the rendered frame.

[0225] At position 704, the first predicted pose is obtained. The first predicted pose represents the prediction of what the user's pose will be at the first predicted display time. First pose confidence information is also obtained, representing the confidence level of the first predicted pose. In some embodiments, the first predicted pose and first pose confidence information can be obtained by calling the function xrLocateViews() in the OpenXR system or by using other techniques from other XR runtime systems.

[0226] At 706, it is determined whether the first pose confidence information indicates a sufficiently high confidence level. In some embodiments, this can be performed by converting the value of the flag XrViewStateFlags to a numerical confidence level and determining whether that numerical confidence level is at or above a threshold. In some embodiments, this can be performed by performing logical operations on some or all of the values ​​in XrViewStateFlags. For example, a high pose confidence can only be determined if both XR_VIEW_STATE_POSITION_VALID_BIT and XR_VIEW_STATE_ORIENTATION_VALID_BIT are equal to "1".

[0227] If the pose confidence level is high enough, at 708, a new frame is obtained based on the first predicted pose. In some embodiments, this can be done by sending the first predicted pose and the first predicted display time to an edge application server and receiving the result from the server. At 710, a second predicted display time is obtained. This can be obtained by calling xrWaitFrame() in the OpenXR system or by using other techniques from other XR runtime systems. At 712, the newly rendered frame is reprojected based on the second predicted display time. At 718, the reprojected frame is displayed, for example, by providing it to the user device and / or by actually displaying the reprojected frame (if execution is performed). Figure 7 If the system has built-in display capabilities, the method allows the reprojected frame to be displayed to the user.

[0228] On the other hand, if the pose confidence level is not high enough, the newly rendered frame is not obtained based on the first predicted pose. At 714, the second predicted display time is obtained. This can be obtained by calling xrWaitFrame() in the OpenXR system or by using other techniques from other XR runtime systems. At 716, the previously rendered frame is reprojected based on the second predicted display time. At 718, the reprojected frame is displayed to the user.

[0229] Figure 8 This is a flowchart illustrating an example process, according to some embodiments, including frame selection using attitude confidence information. (The remaining text appears to be unrelated and possibly machine-generated.) Figure 8The example embodiment is described further below. At 802, a first predicted display time is obtained, for example, in steps 4 (214, 414, 514, 614) and 702. At 804, a first predicted pose is obtained. The first predicted pose represents a prediction of what the user's pose will be at the first predicted display time. First pose confidence information is also obtained, representing the confidence level of the first predicted pose. At 806, a frame to be displayed is selected based on the first pose confidence level information. For example, in response to determining that the first pose confidence level information is below a threshold, a previously rendered frame can be selected. In response to determining that the first pose confidence level information is greater than or equal to a threshold, a new frame can be rendered based on the first predicted pose information (e.g., via EAS). At 808, a second predicted display time is obtained. The second predicted display time can be obtained by calling xrWaitFrame() in the OpenXR system or by using other techniques of other XR runtime systems. At 810, the selected frame is reprojected based on the second predicted display time (depending on the selection, the frame can be a newly rendered frame or a previously rendered frame). At 812, the reprojected frame is displayed to the user.

[0230] Use of confidence information in attitude error measurement Figure 9 This is a flowchart illustrating an example process, according to some embodiments, including attitude error determination using attitude confidence information. In some embodiments, the process for attitude error measurement is performed based on the attitude confidence level obtained as described above. Regarding... Figure 9 Describe the example process.

[0231] Figure 9 The method shown is the same as Figure 8 The method shown is performed similarly. At 902, the first predicted display time is obtained. At 904, the first predicted pose is obtained. At 906, the frame to be displayed is selected based on the first pose confidence level information. At 908, the second predicted display time is obtained. At 910, the selected frame is reprojected based on the second predicted display time (depending on the selection, this frame can be a newly rendered frame or a previously rendered frame). At 912, the reprojected frame is displayed to the user.

[0232] Additionally, at position 914, a second predicted pose is obtained. The second predicted pose represents the prediction of the user's pose at the second prediction display time. Second pose confidence information is also obtained, representing the confidence level of the second predicted pose. In some embodiments, the second predicted pose and second pose confidence information can be obtained by calling the function xrLocateViews() in the OpenXR system or by using other techniques from other XR runtime systems.

[0233] At position 916, the attitude error can be determined. The attitude error can be determined by calculating the difference between the first predicted attitude and the second predicted attitude, for example, using (P.predicted1 - P.predicted2).

[0234] In some embodiments, at step 18 (446, 548, 644), it is determined whether the attitude error is calculated based on first attitude confidence information (e.g., as returned at step 7 (420, 520, 620)) and / or second attitude confidence information (e.g., as returned at step 17 (444, 546, 642)). In some embodiments, at step 18 (446, 548, 644), the attitude error is calculated, and it is determined whether the attitude error is determined to be valid based on the first attitude confidence information and / or the second attitude confidence information.

[0235] In some embodiments, the bitwise product of the first attitude confidence information and the second attitude confidence information is calculated. For example, the final flag value can be determined as follows:

[0236]

[0237]

[0238]

[0239] In some embodiments, the confidence level is determined from the final flag values ​​using the confidence levels indicated in Table 1; however, in other embodiments, other confidence values ​​may be assigned. In some embodiments, the validity (or invalidity) of the attitude error measurement is determined based on the numerical confidence level, as shown in Table 3. In some embodiments, whether to compute the attitude error measurement is determined based on the numerical confidence level, as shown in Table 3.

[0240]

[0241] Table 3. Attitude Error Measurement Based on Confidence Level Therefore, the measurement is valid if both the confidence levels from steps 7 and 17 are equal to 1. In some embodiments, a low confidence level is associated with the attitude error measurement if the product of the confidence levels from steps 7 and 17 is greater than or equal to a threshold (e.g., 5 / 8 in Table 3). In some embodiments, the attitude error measurement is considered invalid if the result is below this threshold.

[0242] In some embodiments, if the first attitude confidence information (e.g., obtained in step 7) indicates a confidence level below a threshold (e.g., below 5 / 8), one or more of steps 15, 16, 17, and 18 are omitted. In some embodiments, if the confidence level is higher than or equal to a threshold (such as 5 / 8), the attitude measurement error is calculated (e.g., at step 18).

[0243] The calculated confidence level and attitude error can be provided to the application. Based on this, the application can decide whether or not to adjust the process.

[0244] In some embodiments, the time error estimation remains unchanged and is performed as shown in Figure 2.

[0245] Figure 10 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of a predicted pose, according to some embodiments. In some embodiments, example process 1000 may include obtaining 1002 a first predicted frame display time. In some embodiments, example process 1000 may also include obtaining 1004 first predicted pose information, which represents a prediction of the user's pose at the first predicted frame display time. In some embodiments, example process 1000 may also include determining 1006 first pose confidence information, which indicates a confidence level of the first predicted pose information. In some embodiments, example process 1000 may also include selecting 1008 a frame for display, at least in part based on the first pose confidence information. In some embodiments, example process 1000 may also include causing 1010 to display the selected frame.

[0246] Figure 11 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of a predicted pose, according to some embodiments. For some embodiments, example process 1100 may include obtaining 1102 a first predicted frame display time. For some embodiments, example process 1100 may also include obtaining 1104 first predicted pose information, which represents a prediction of the user pose at the first predicted frame display time. For some embodiments, example process 1100 may also include determining 1106 first pose confidence information, which indicates a confidence level of the first predicted pose information. For some embodiments, example process 1100 may also include transmitting 1108 the determined first pose confidence information, which indicates a confidence level of the first predicted pose information, to an edge application server. For some embodiments, example process 1100 may also include selecting 1110 a frame for display, at least in part based on the first pose confidence information. For some embodiments, example process 1100 may also include causing 1112 to display the selected frame.

[0247] Figure 12 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of a predicted pose, according to some embodiments. In some embodiments, example process 1200 may include receiving 1202 a first predicted frame display time. In some embodiments, example process 1200 may also include receiving 1204 first predicted pose information, which represents a prediction of a user pose at the first predicted frame display time. In some embodiments, example process 1200 may also include receiving 1206 first pose confidence information, which indicates a confidence level of the first predicted pose information. In some embodiments, example process 1200 may also include selecting 1208 a frame for display, at least in part based on the first pose confidence information. In some embodiments, the example process may also include rendering 1210 the selected frame. In some embodiments, example process 1200 may also include sending 1212 the rendered frame to an extended reality (XR) application device. In some embodiments, example process 1200 may also include sending the first predicted frame display time to an XR application device. In some embodiments, the XR application device already knows the display time of the first predicted frame and keeps tracking that value for later use.

[0248] Figure 13 This is a flowchart illustrating an example process for rendering a scene based on a confidence level of a predicted pose, according to some embodiments. In some embodiments, example process 1300 may include obtaining 1302 a first predicted frame display time. In some embodiments, example process 1300 may also include obtaining 1304 first predicted pose information, which represents a prediction of the user's pose at the first predicted frame display time. In some embodiments, example process 1300 may also include transmitting 1306 the first predicted frame display time, the first predicted pose information, and information indicating the extended reality (XR) view state to an edge application server. In some embodiments, example process 1300 may also include obtaining 1308 first pose confidence information, which indicates the confidence level of the first predicted pose information. In some embodiments, example process 1300 may also include selecting 1310 a frame for display, at least in part based on the first pose confidence information. In some embodiments, example process 1300 may also include causing 1312 to display the selected frame.

[0249] Figure 14This is a flowchart illustrating an example process for rendering a scene based on a confidence level of a predicted pose, according to some embodiments. In some embodiments, example process 1400 may include receiving 1402 a first predicted frame display time. In some embodiments, example process 1400 may also include receiving 1404 first predicted pose information representing a prediction of a user pose at the first predicted frame display time. In some embodiments, example process 1400 may also include receiving 1406 information indicating the extended reality (XR) view state. In some embodiments, example process 1400 may also include determining 1408 first pose confidence information indicating a confidence level of the first predicted pose information. In some embodiments, example process 1400 may also include selecting 1410 a frame for display, at least in part based on the first pose confidence information. In some embodiments, example process 1400 may also include rendering 1412 the selected frame. In some embodiments, example process 1400 may also include sending 1414 the rendered frame and the determined first pose confidence information to an extended reality (XR) application device. In some embodiments, example process 1400 may further include sending the first predicted frame display time to the XR application device. In some embodiments, the XR application device already knows the first predicted frame display time and keeps tracking this value for later use.

[0250] While methods and systems according to some embodiments are generally discussed in the context of extended reality (XR), some embodiments can be applied to any XR context, such as virtual reality (VR) / mixed reality (MR) / augmented reality (AR) contexts. Additionally, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments can be applied to wearable devices (which may or may not be attached to the head) capable of supporting, for example, XR, VR, AR, and / or MR.

[0251] A first example method according to some embodiments may include: obtaining a first predicted frame display time; obtaining first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; obtaining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; and causing the selected frame to be displayed.

[0252] Some embodiments of the first example method may further include: obtaining a second predicted frame display time; and performing a reprojection of the selected frame based on the second predicted frame display time before causing the selected frame to be displayed.

[0253] In some embodiments of the first example method, in response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

[0254] Some embodiments of the first example method may also include rendering the selected frame based on the first predicted pose.

[0255] In some embodiments of the first example method, in response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

[0256] In some embodiments of the first example method, the first attitude confidence information includes multiple flags.

[0257] In some embodiments of the first example method, the first attitude confidence information includes at least a position validity flag and an orientation validity flag.

[0258] In some embodiments of the first example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame may be selected as the selected frame for display.

[0259] In some embodiments of the first example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame may be selected as the selected frame for display.

[0260] In some embodiments of the first example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0261] Some embodiments of the first example method may further include: obtaining second predicted pose information, the second predicted pose information representing a prediction of the user pose at the display time of the selected frame; and obtaining second pose confidence information, the second pose confidence information indicating the confidence level of the second predicted pose information.

[0262] Some embodiments of the first example method may further include: calculating the attitude error based on the difference between the first predicted attitude information and the second predicted attitude information.

[0263] Some embodiments of the first example method may further include: determining whether to calculate the attitude error based on the first attitude confidence information and the second attitude confidence information, wherein the attitude error is calculated in response to determining to calculate the attitude error.

[0264] Some embodiments of the first example method may further include: determining whether the attitude error is valid based on the first attitude confidence information and the second attitude confidence information.

[0265] In some embodiments of the first example method, obtaining the first attitude confidence information includes determining the first attitude confidence information.

[0266] A first example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0267] A second example method according to some embodiments may include: obtaining a first predicted frame display time; obtaining first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; determining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; transmitting the determined first pose confidence information to an edge application server, the determined first pose confidence information indicating the confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; and causing the selected frame to be displayed.

[0268] Some embodiments of the second example method may further include: obtaining a second predicted frame display time; and performing a reprojection of the selected frame based on the second predicted frame display time before causing the selected frame to be displayed.

[0269] In some embodiments of the second example method, the first attitude confidence information may include multiple flags.

[0270] In some embodiments of the second example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0271] In some embodiments of the second example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0272] In some embodiments of the second example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0273] In some embodiments of the second example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0274] Some embodiments of the second example method may further include: obtaining second predicted pose information, the second predicted pose information representing a prediction of the user pose at the display time of the selected frame; and obtaining second pose confidence information, the second pose confidence information indicating the confidence level of the second predicted pose information.

[0275] Some embodiments of the second example method may further include: calculating the attitude error based on the difference between the first predicted attitude information and the second predicted attitude information.

[0276] Some embodiments of the second example method may further include: determining whether to calculate the attitude error based on the first attitude confidence information and the second attitude confidence information, wherein the attitude error is calculated in response to determining to calculate the attitude error.

[0277] Some embodiments of the second example method may further include: determining whether the attitude error is valid based on the first attitude confidence information and the second attitude confidence information.

[0278] A second example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0279] A third example method according to some embodiments may include: receiving a first predicted frame display time; receiving first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; receiving first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; rendering the selected frame; and sending the rendered frame to an extended reality (XR) application device.

[0280] In some embodiments of the third example method, in response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

[0281] In some embodiments of the third example method, the selected frame is rendered based on the first predicted pose.

[0282] In some embodiments of the third example method, in response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

[0283] In some embodiments of the third example method, the first attitude confidence information may include multiple flags.

[0284] In some embodiments of the third example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0285] In some embodiments of the third example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0286] In some embodiments of the third example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0287] In some embodiments of the third example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0288] Some embodiments of the third example method may also include sending the first predicted frame display time to the extended reality (XR) application device.

[0289] A third example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0290] A fourth example method according to some embodiments may include: obtaining a first predicted frame display time; obtaining first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; transmitting the first predicted frame display time, the first predicted pose information, and information indicating the state of an extended reality (XR) view to an edge application server; obtaining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; and causing the selected frame to be displayed.

[0291] Some embodiments of the fourth example method may further include: obtaining a second predicted frame display time; and performing a reprojection of the selected frame based on the second predicted frame display time before causing the selected frame to be displayed.

[0292] In some embodiments of the fourth example method, the first attitude confidence information may include multiple flags.

[0293] In some embodiments of the fourth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0294] In some embodiments of the fourth example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0295] In some embodiments of the fourth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0296] In some embodiments of the fourth example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0297] Some embodiments of the fourth example method may further include: obtaining second predicted pose information, the second predicted pose information representing a prediction of the user pose at the display time of the selected frame; and obtaining second pose confidence information, the second pose confidence information indicating the confidence level of the second predicted pose information.

[0298] Some embodiments of the fourth example method may further include: calculating the attitude error based on the difference between the first predicted attitude information and the second predicted attitude information.

[0299] Some embodiments of the fourth example method may further include: determining whether to calculate the attitude error based on the first attitude confidence information and the second attitude confidence information, wherein the attitude error is calculated in response to determining to calculate the attitude error.

[0300] Some embodiments of the fourth example method may further include: determining whether the attitude error is valid based on the first attitude confidence information and the second attitude confidence information.

[0301] A fourth example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0302] A fifth example method according to some embodiments may include: receiving a first predicted frame display time; receiving the first predicted frame display time; receiving first predicted pose information, the first predicted pose information representing a prediction of a user pose at the first predicted frame display time; receiving information indicating the state of an extended reality (XR) view; determining first pose confidence information, the first pose confidence information indicating a confidence level of the first predicted pose information; selecting a frame for display based at least in part on the first pose confidence information; rendering the selected frame; and sending the rendered frame and the determined first pose confidence information to an extended reality (XR) application device.

[0303] In some embodiments of the fifth example method, in response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

[0304] In some embodiments of the fifth example method, the selected frame is rendered based on the first predicted pose.

[0305] In some embodiments of the fifth example method, in response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

[0306] In some embodiments of the fifth example method, the first attitude confidence information may include multiple flags.

[0307] In some embodiments of the fifth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag.

[0308] In some embodiments of the fifth example method, the first pose confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that both the position validity flag and the orientation validity flag are set, a newly rendered frame is selected as the selected frame for display.

[0309] In some embodiments of the fifth example method, the first attitude confidence information may include at least a position validity flag and an orientation validity flag, and in response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

[0310] In some embodiments of the fifth example method, the first attitude confidence information may include at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

[0311] Some embodiments of the fifth example method may also include sending the first predicted frame display time to the extended reality (XR) application device.

[0312] A fifth example device according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0313] A sixth example method / apparatus according to some embodiments may include: one or more processors configured to perform any of the methods listed above.

[0314] A seventh example method / apparatus according to some embodiments may include a computer-readable storage medium storing instructions for causing one or more processors to perform any of the methods listed above.

[0315] An eighth example method / apparatus according to some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any of the methods listed above.

[0316] An example computer-readable medium according to some embodiments may include instructions for causing one or more processors to perform any of the methods listed above.

[0317] In some embodiments of the example computer-readable medium, the computer-readable medium is a non-transitory storage medium.

[0318] An example computer program according to some embodiments may include instructions that, when the program is executed by one or more processors, cause the one or more processors to perform any of the methods listed above.

[0319] An example signal according to some embodiments may include a scene description file generated according to any of the methods listed above.

[0320] This disclosure describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in a specific manner and are generally described in a way that may sound limiting, at least for the purpose of illustrating the various features. However, this is for the purpose of clarity of description and does not limit the disclosure or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in previous documents.

[0321] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are specifically shown, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of implementation. At least one aspect relates generally to video encoding and decoding, and at least one other aspect relates generally to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions thereon stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having bitstreams generated according to any of the methods stored thereon.

[0322] This document describes various methods, each of which includes one or more steps or actions for implementing the described method. The order and / or use of specific steps and / or actions may be modified or combined unless the correct operation of the method requires a specific order. Additionally, in various embodiments, terms such as "first," "second," etc., may be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a sequence of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding, but may occur, for example, before, during, or in a time period overlapping with the second decoding.

[0323] For example, various numerical values ​​may be used in this disclosure. Specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0324] The embodiments described herein can be implemented by computer software, either by a processor or other hardware or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the processor can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0325] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / apparatus.

[0326] The embodiments and aspects described herein may be implemented in, for example, methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single embodiment (e.g., discussed only as a method), embodiments of the discussed features may be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, with suitable hardware, software, and firmware. Methods may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0327] The references to "an embodiment" or "an embodiment" or "an implementation" or "implementation", as well as other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in a manifestation" or "in an implementation" or "in an implementation", as well as any other variations, appearing throughout this disclosure, do not necessarily refer to the same embodiment.

[0328] Additionally, this disclosure may relate to “determining” various types of information. Determining information may include one or more of the following: for example, estimated information, calculated information, predicted information, or information retrieved from memory.

[0329] Furthermore, this disclosure may relate to “accessing” various types of information. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., retrieving from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0330] Additionally, this disclosure may relate to "receiving" various types of information. Like "access," the intent to receive is a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., retrieving from memory). Furthermore, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0331] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As yet another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). This can be extended to as many as the items listed.

[0332] Furthermore, as used herein, the term “signaling” specifically refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a particular parameter among several parameters for region-based filter parameter selection for artifact removal filtering. In this way, in embodiments, the same parameter is used at both the encoder and decoder sides. Thus, for example, the encoder can send (explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without sending (implicitly signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntactic elements, flags, etc., are used to send information to the corresponding decoder. Although the verb form of the word “signal” has been referred to above, the word “signal” can also be used as a noun herein.

[0333] Implementation methods can generate various signals that are formatted to carry information, such as information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementation methods. For example, the signal may be formatted to carry a bit stream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0334] Several embodiments are described. Features of these embodiments may be provided individually or in any combination across various claim classes and types. Furthermore, embodiments may include one or more of the following features, means, or aspects individually or in any combination across various claim classes and types: • A bitstream or signal that includes one or more of the described syntactic elements or their variants.

[0335] • A bitstream or signal, which includes syntax for conveying information generated according to any embodiment of the described embodiments.

[0336] • Create and / or send and / or receive and / or decode bit streams or signals, which include one or more of the described syntactic elements or their variants.

[0337] • Create and / or send and / or receive and / or decode according to any of the described embodiments.

[0338] • A method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the described embodiments.

[0339] It should be noted that the various hardware elements, one or more of which are described in the embodiments, are referred to as “modules”, which implement (i.e., perform, execute, etc.) the various functions described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that a person skilled in the art would consider suitable for a given implementation. Each described module may also include executable instructions for performing one or more functions described as being performed by the respective module, and it should be noted that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and may be stored in any suitable non-transitory computer-readable medium such as RAM, ROM, etc., commonly referred to as RAM, ROM, etc.

[0340] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multifunction disks (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A method comprising: Obtain the display time of the first predicted frame; Obtain first predicted pose information, which represents the prediction of the user's pose at the display time of the first predicted frame; Obtain first attitude confidence information, which indicates the confidence level of the first predicted attitude information; The frame to be displayed is selected based at least in part on the first attitude confidence information; as well as This causes the selected frame to be displayed.

2. The method according to claim 1, further comprising: Obtain the display time of the second predicted frame; as well as Before the selected frame is displayed, the selected frame is reprojected based on the second predicted frame display time.

3. The method according to any one of claims 1 to 2, wherein: In response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

4. The method according to claim 3, further comprising: The selected frame is rendered based on the first predicted pose.

5. The method according to any one of claims 1 to 2, wherein: In response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

6. The method according to any one of claims 1 to 5, wherein, The first attitude confidence information includes multiple flags.

7. The method according to any one of claims 1 to 6, wherein, The first attitude confidence information includes at least a position validity flag and an orientation validity flag.

8. The method according to any one of claims 1 to 7, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to the determination that both the position validity flag and the orientation validity flag have been set, a newly rendered frame is selected as the selected frame for display.

9. The method according to any one of claims 1 to 8, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

10. The method according to any one of claims 1 to 9, wherein, The first attitude confidence information includes at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

11. The method according to any one of claims 1 to 10, further comprising: Obtain second predicted pose information, which represents a prediction of the user's pose at the display time of the selected frame; as well as Obtain second attitude confidence information, which indicates the confidence level of the second predicted attitude information.

12. The method of claim 11, further comprising: The attitude error is calculated based on the difference between the first predicted attitude information and the second predicted attitude information.

13. The method of claim 12, further comprising: The attitude error is calculated in response to the determination to calculate the attitude error, based on the first attitude confidence information and the second attitude confidence information.

14. The method of claim 12, further comprising: The validity of the attitude error is determined based on the first attitude confidence information and the second attitude confidence information.

15. The method according to any one of claims 1 to 14, wherein, Obtaining the first attitude confidence information includes determining the first attitude confidence information.

16. An apparatus comprising: processor; as well as A non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 1 to 15.

17. A method comprising: Obtain the display time of the first predicted frame; Obtain first predicted pose information, which represents the prediction of the user's pose at the display time of the first predicted frame; Determine first attitude confidence information, which indicates the confidence level of the first predicted attitude information; The determined first pose confidence information is transmitted to the edge application server, the determined first pose confidence information indicating the confidence level of the first predicted pose information; The frame to be displayed is selected based at least in part on the first attitude confidence information; as well as This causes the selected frame to be displayed.

18. The method of claim 17, further comprising: Obtain the display time of the second predicted frame; as well as Before the selected frame is displayed, the selected frame is reprojected based on the second predicted frame display time.

19. The method according to any one of claims 17 to 18, wherein, The first attitude confidence information includes multiple flags.

20. The method according to any one of claims 17 to 19, wherein, The first attitude confidence information includes at least a position validity flag and an orientation validity flag.

21. The method according to any one of claims 17 to 20, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to the determination that both the position validity flag and the orientation validity flag have been set, a newly rendered frame is selected as the selected frame for display.

22. The method according to any one of claims 17 to 21, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

23. The method according to any one of claims 17 to 22, wherein, The first attitude confidence information includes at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

24. The method according to any one of claims 17 to 23, further comprising: Obtain second predicted pose information, which represents a prediction of the user's pose at the display time of the selected frame; as well as Obtain second attitude confidence information, which indicates the confidence level of the second predicted attitude information.

25. The method of claim 24, further comprising: The attitude error is calculated based on the difference between the first predicted attitude information and the second predicted attitude information.

26. The method of claim 25, further comprising: The attitude error is calculated in response to the determination to calculate the attitude error, based on the first attitude confidence information and the second attitude confidence information.

27. The method of claim 26, further comprising: The validity of the attitude error is determined based on the first attitude confidence information and the second attitude confidence information.

28. An apparatus comprising: processor; as well as A non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 17 to 27.

29. A method comprising: Receive the first predicted frame and display the time; Receive first predicted pose information, which represents the prediction of the user's pose at the display time of the first predicted frame; Receive first attitude confidence information, which indicates the confidence level of the first predicted attitude information; The frame to be displayed is selected based at least in part on the first attitude confidence information; Render the selected frame; as well as Send rendered frames to extended reality (XR) application devices.

30. The method according to claim 29, wherein: In response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

31. The method according to claim 30, wherein, The selected frame is rendered based on the first predicted pose.

32. The method according to any one of claims 30 to 31, wherein: In response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

33. The method according to any one of claims 30 to 32, wherein, The first attitude confidence information includes multiple flags.

34. The method according to any one of claims 30 to 33, wherein, The first attitude confidence information includes at least a position validity flag and an orientation validity flag.

35. The method according to any one of claims 30 to 34, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to the determination that both the position validity flag and the orientation validity flag have been set, a newly rendered frame is selected as the selected frame for display.

36. The method according to any one of claims 30 to 35, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

37. The method according to any one of claims 30 to 36, wherein, The first attitude confidence information includes at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

38. The method according to any one of claims 30 to 37, further comprising sending the first predicted frame display time to the extended reality (XR) application device.

39. An apparatus comprising: processor; as well as A non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 30 to 38.

40. A method comprising: Obtain the display time of the first predicted frame; Obtain first predicted pose information, which represents the prediction of the user's pose at the display time of the first predicted frame; The first predicted frame display time, the first predicted pose information, and information indicating the extended reality (XR) view status are transmitted to the edge application server. Obtain first attitude confidence information, which indicates the confidence level of the first predicted attitude information; The frame to be displayed is selected based at least in part on the first attitude confidence information; as well as This causes the selected frame to be displayed.

41. The method of claim 40, further comprising: Obtain the display time of the second predicted frame; as well as Before the selected frame is displayed, the selected frame is reprojected based on the second predicted frame display time.

42. The method according to any one of claims 40 to 41, wherein, The first attitude confidence information includes multiple flags.

43. The method according to any one of claims 40 to 42, wherein, The first attitude confidence information includes at least a position validity flag and an orientation validity flag.

44. The method according to any one of claims 40 to 43, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to the determination that both the position validity flag and the orientation validity flag have been set, a newly rendered frame is selected as the selected frame for display.

45. The method according to any one of claims 40 to 44, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

46. ​​The method according to any one of claims 40 to 45, wherein, The first attitude confidence information includes at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

47. The method according to any one of claims 40 to 46, further comprising: Obtain second predicted pose information, which represents a prediction of the user's pose at the display time of the selected frame; as well as Obtain second attitude confidence information, which indicates the confidence level of the second predicted attitude information.

48. The method of claim 47, further comprising: The attitude error is calculated based on the difference between the first predicted attitude information and the second predicted attitude information.

49. The method of claim 48, further comprising: The attitude error is calculated in response to the determination to calculate the attitude error, based on the first attitude confidence information and the second attitude confidence information.

50. The method of claim 48, further comprising: The validity of the attitude error is determined based on the first attitude confidence information and the second attitude confidence information.

51. An apparatus, the apparatus comprising: processor; as well as A non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 40 to 50.

52. A method comprising: Receive the first predicted frame and display the time; Receive first predicted pose information, which represents the prediction of the user's pose at the display time of the first predicted frame; Receive information indicating the status of the extended reality (XR) view; Determine first attitude confidence information, which indicates the confidence level of the first predicted attitude information; The frame to be displayed is selected based at least in part on the first attitude confidence information; Render the selected frame; as well as Send rendered frames and determined first pose confidence information to the extended reality (XR) application device.

53. The method according to claim 52, wherein: In response to determining that the first pose confidence information indicates a confidence level at least as large as a threshold, a newly rendered frame is selected as the selected frame for display.

54. The method according to claim 53, wherein, The selected frame is rendered based on the first predicted pose.

55. The method according to claim 52, wherein: In response to determining that the first pose confidence information indicates a confidence level below a threshold, a previously rendered frame is selected as the selected frame for display.

56. The method according to any one of claims 52 to 55, wherein, The first attitude confidence information includes multiple flags.

57. The method according to any one of claims 52 to 56, wherein, The first attitude confidence information includes at least a position validity flag and an orientation validity flag.

58. The method according to any one of claims 52 to 57, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to the determination that both the position validity flag and the orientation validity flag have been set, a newly rendered frame is selected as the selected frame for display.

59. The method according to any one of claims 52 to 58, in, The first attitude confidence information includes at least a position validity flag and an orientation validity flag, and In response to determining that at least one of the position validity flag and the orientation validity flag is not set, a previously rendered frame is selected as the selected frame for display.

60. The method according to any one of claims 52 to 59, wherein, The first attitude confidence information includes at least a position validity flag, an orientation validity flag, a position tracking flag, and an orientation tracking flag.

61. The method according to any one of claims 52 to 60, further comprising: The first predicted frame display time is sent to the extended reality (XR) application device.

62. An apparatus comprising: processor; as well as A non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 52 to 60.

63. An apparatus comprising one or more processors configured to perform the method as described in any one of claims 1 to 15, 17 to 27, 29 to 38, 40 to 50, and 52 to 61.

64. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method as described in any one of claims 1 to 15, 17 to 27, 29 to 38, 40 to 50, and 52 to 61.

65. An apparatus comprising at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform the method as described in any one of claims 1 to 15, 17 to 27, 29 to 38, 40 to 50, and 52 to 61.

66. A computer-readable medium comprising instructions for causing one or more processors to perform the method as described in any one of claims 1 to 15, 17 to 27, 29 to 38, 40 to 50, and 52 to 61.

67. The computer-readable medium of claim 66, wherein the computer-readable medium is a non-transitory storage medium.

68. A computer program product comprising instructions that, when the program is executed by one or more processors, cause the one or more processors to perform the method as described in any one of claims 1 to 15, 17 to 27, 29 to 38, 40 to 50, and 52 to 61.

69. A signal comprising a scene description file generated by the method according to any one of claims 1 to 15, 17 to 27, 29 to 38, 40 to 50, and 52 to 61.