Eye gaze prediction for split XR
By utilizing fully connected networks and gated recurrent units in conjunction with a predictor model in extended reality (XR) rendering, and predicting eye gaze based on head pose data, the problem of inaccurate eye gaze prediction is solved, thereby improving rendering accuracy and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies suffer from inaccurate eye gaze prediction in extended reality (XR) rendering, which affects the user experience.
By acquiring eye gaze and head pose data from user devices, and utilizing a fully connected network (FCN) and gated recurrent unit (GRU) blocks, combined with a predictor model, eye gaze is predicted based on high-sampling-rate head pose data, generating more accurate predicted eye gaze, and rendering frames based on this.
It improves the accuracy of extended reality rendering and user experience, and enhances frame display quality by predicting eye gaze, especially in split XR systems.
Smart Images

Figure CN121752980A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit and priority of Indian Provisional Patent Application No. 202341060303, filed on September 7, 2023, entitled “EYE GAZE PREDICTION FOR SPLIT XR”, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0002] This disclosure relates generally to processing systems, and more specifically, to one or more techniques for graphics processing. Background Technology
[0003] Computing devices typically perform graphics and / or display processing (e.g., utilizing a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices can include, for example, computer workstations, mobile phones (such as smartphones), embedded systems, personal computers, tablet computers, and video game consoles. A GPU is configured to execute a graphics processing pipeline that includes one or more processing stages that operate together to execute graphics processing commands and output frames. A CPU controls the operation of a GPU by issuing one or more graphics processing commands to it. Modern CPUs are typically capable of executing multiple applications concurrently, each of which may require the use of a GPU during execution. A display processor can be configured to convert digital information received from the CPU into analog values and can issue commands to a display panel to display visual content. Devices that provide content for visual presentation on a display can utilize a CPU, GPU, and / or display processor.
[0004] Current techniques for extended reality (XR) rendering may be based on inaccurate eye gaze prediction. Improved techniques for XR rendering are needed. Summary of the Invention
[0005] The following is a simplified summary of one or more aspects of the invention to provide a basic understanding of these aspects. This summary is not a broad overview of all anticipated aspects, nor is it intended to identify key or essential elements of all aspects, nor to depict the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.
[0006] In one aspect of this disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus includes: a memory; and a processor coupled to the memory, configured to obtain, based on information stored in the memory, an indication of an eye gaze associated with a user device. The processor is configured to obtain an indication of a head posture associated with the user device. The processor is configured to determine a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture. The processor is configured to output an indication of the determined set of predicted eye gazes.
[0007] In one aspect of this disclosure, a method, computer-readable medium, and apparatus are provided. The apparatus includes: a memory; and a processor coupled to the memory, configured, based on information stored in the memory, to obtain a first indication of a first set of eye gazes associated with a user device. The processor is configured to obtain a second indication of a second set of head postures associated with the user device. The processor is configured to determine a third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of the obtained second set of head postures associated with a first subgroup of the obtained first set of eye gazes. The processor is configured to output a third indication of the determined third set of predicted eye gazes.
[0008] In some aspects, the technology described herein relates to a graphics processing method comprising: obtaining a first indication of a first set of eye gazes associated with a user device; obtaining a second indication of a second set of head poses associated with the user device; determining a third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of the obtained second set of head poses associated with a first subgroup of the obtained first set of eye gazes; and outputting a third indication of the determined third set of predicted eye gazes.
[0009] In some respects, the technology described herein relates to a method that further includes: associating the second subgroup of the acquired second set of head poses with the first subgroup of the acquired first set of eye gazes.
[0010] In some respects, the technology described herein relates to a method in which associating the second subgroup of a second set of head poses with the first subgroup of a first set of eye gazes comprises: associating the second subgroup of a second set of head poses with the first subgroup of a first set of eye gazes based on a first set of time indicators associated with the first set of eye gazes and a second set of time indicators associated with the second set of head poses.
[0011] In some respects, the techniques described herein relate to a method in which the first set of eye gazes includes eye gaze time series data, and the second set of head poses includes head pose time series data.
[0012] In some aspects, the technology described herein relates to a method that further includes: obtaining a fourth indication of a fourth set of blinks associated with the user device; and selecting a third subgroup of a first set of eye gazes based on the fourth set of blinks associated with the user device, wherein determining the third set of predicted eye gazes based on the first set of eye gazes and a second subgroup of a second set of head postures associated with the first subgroup of the first set of eye gazes comprises: determining the third set of predicted eye gazes based on the selected third subgroup of the first set of eye gazes and a fourth subgroup of the second set of head postures associated with the selected third subgroup of the first set of eye gazes.
[0013] In some aspects, the technology described herein relates to a method that further includes: determining a third subgroup of a first set of eye gazes associated with a saccade; and selecting a fourth subgroup of the first set of eye gazes based on the determined third subgroup of the first set of eye gazes associated with the saccade, wherein determining the third set of predicted eye gazes based on the first set of eye gazes and a second subgroup of a second set of head postures associated with the first subgroup of the first set of eye gazes includes: determining the third set of predicted eye gazes based on the selected fourth subgroup of the first set of eye gazes and a fifth subgroup of the selected fourth subgroup of the second set of head postures associated with the selected fourth subgroup of the first set of eye gazes.
[0014] In some aspects, the technology described herein relates to a method that further includes: converting a third subgroup of the first set of eye gazes into a fourth set of six-dimensional (6D) representations of the first set of eye gazes, wherein determining the third set of predicted eye gazes based on the obtained first set of eye gazes and the second subgroup of the obtained second set of head poses associated with the first subgroup of the obtained first set of eye gazes includes: determining the third set of predicted eye gazes based on the converted fourth set of 6D representations of the first set of eye gazes.
[0015] In some respects, the techniques described herein relate to a method in which a third set of predicted eye gazes is determined based on a first set of obtained eye gazes and a second set of obtained head poses associated with the first subset of the obtained first set of eye gazes, comprising: determining the third set of predicted eye gazes based on a first set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks.
[0016] In some respects, the techniques described herein relate to a method in which determining a third set of predicted eye gazes based on a first set of obtained eye gazes and a second set of obtained head poses associated with that first set of obtained eye gazes comprises: determining the third set of predicted eye gazes based on a predictor model that uses the first set of obtained eye gazes and the second set of obtained head poses associated with that first set of obtained eye gazes as input.
[0017] In some respects, the techniques described herein relate to a method in which the predictor model includes at least one of a regression model, a time series predictor model, or a neural network.
[0018] In some aspects, the technology described herein relates to a method that further includes: determining a third subgroup of the first set of eye gazes associated with a user tracking a virtual object; determining a fourth set of predicted eye gazes based on the determined third subgroup of the first set of eye gazes relative to the user's head; and outputting a fourth indication of the determined fourth set of predicted eye gazes.
[0019] In some aspects, the technology described herein relates to a method that further includes: determining that a fourth subgroup of the first set of eye gazes is not associated with the user tracking the virtual object, wherein determining the third set of predicted eye gazes based on the obtained first set of eye gazes and the second subgroup of the obtained second set of head poses associated with the first subgroup of the obtained first set of eye gazes includes: determining the third set of predicted eye gazes based on the determined fourth subgroup of the first set of eye gazes and the fifth subgroup of the obtained second set of head poses associated with the determined fourth subgroup of the first set of eye gazes.
[0020] In some aspects, the technology described herein relates to a method in which obtaining the first set of eye gazes includes generating the first set of eye gazes at a first sampling rate via an eye gaze camera associated with the user equipment, and obtaining the second set of head poses includes generating the second set of head poses at a second sampling rate via an inertial measurement unit (IMU) associated with the user equipment.
[0021] In some respects, the techniques described herein relate to a method in which the first sampling rate (e.g., 90 Hz) is less than the second sampling rate (e.g., 1000 Hz).
[0022] In some respects, the technology described herein relates to a method in which the first set of eye gazes includes a third set of eye gaze time series data, and the second set of head poses includes a fourth set of head pose time series data.
[0023] In some aspects, the technology described herein relates to a method in which the first set of eye gazes includes first translation information and first orientation information of at least one eye of a user associated with the user device, and the second set of head poses includes second translation information and second orientation information of the user's head associated with the user device.
[0024] In some aspects, the technology described herein relates to a method in which outputting the third indication for a determined third set of predicted eye gazes comprises: sending the third indication for the determined third set of predicted eye gazes; receiving a fourth indication for a fourth set of rendered frames based on sending the third indication for the determined third set of predicted eye gazes; and outputting a fifth indication for the fourth set of rendered frames.
[0025] In some respects, the techniques described herein relate to a method in which the fifth instruction for the fourth set of rendered frames is output including storing the fourth set of rendered frames in at least one of a memory, a buffer, or a cache.
[0026] In some aspects, the technology described herein relates to a method that further includes: decoding the fourth set of rendered frames; and warping the decoded fourth set of rendered frames based on the latest available head pose of the user associated with the user device, wherein the fifth indication of the fourth set of rendered frames includes: outputting a sixth indication of the fifth set of warped decoded rendered frames.
[0027] In some respects, the techniques described herein relate to a method in which each of the fourth set of render frames includes a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the third set of predicted eye gazes.
[0028] In some aspects, the technology described herein relates to a method in which obtaining the first indication of eye gaze associated with the user device comprises: receiving the first indication of eye gaze associated with the first set of eyes from the user device, wherein obtaining the second indication of head posture associated with the user device comprises: receiving the second indication of head posture associated with the second set of eyes from the user device.
[0029] In some respects, the technology described herein relates to a method in which outputting the third indication of a determined third set of predicted eye gazes comprises: outputting the third indication to a renderer, the method further comprising: rendering a fourth set of frames via the renderer based on the determined third set of predicted eye gazes; and sending the rendered fourth set of frames to the user device.
[0030] In some aspects, the techniques described herein relate to a method in which rendering the fourth set of frames based on a determined third set of predicted eye gazes includes: rendering a first region and a second region of each frame in the fourth set of frames, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the determined third set of predicted eye gazes.
[0031] In some respects, the techniques described herein relate to a method in which sending a rendered fourth group of frames to the user equipment includes: encoding the rendered fourth group of frames; and sending the encoded rendered fourth group of frames.
[0032] In one aspect of this disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus includes: a memory; and a processor coupled to the memory, and configured, based on information stored in the memory, to: provide at a first time instance a first indication of a set of head poses of a user's head and a second indication of a set of eye gazes of the user's eyes; obtain a rendered frame based on the first and second indications, the rendered frame being based on a predicted eye gaze of the user at a second time instance occurring after the first time instance; and output an indication of the rendered frame.
[0033] In one aspect of this disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus includes: a memory; and a processor coupled to the memory, and configured, based on information stored in the memory, to: obtain a first indication of a set of head poses of a user's head and a second indication of a set of eye gazes of the user's eyes, wherein the first indication and the second indication correspond to a first time instance; generate a predicted eye gaze of the user based on the first indication and the second indication, wherein the predicted eye gaze of the user corresponds to a second time instance occurring after the first time instance; render a frame based on the predicted eye gaze of the user; and output the rendered frame.
[0034] In some aspects, the technology described herein relates to a method for graphics processing, the method comprising: obtaining an indication of an eye gaze associated with a user device; obtaining an indication of a head pose associated with the user device; determining a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose; and outputting an indication of the determined set of predicted eye gazes.
[0035] In some respects, the technology described herein relates to a method that further includes: associating the head posture with the eye gaze.
[0036] In some respects, the technology described herein relates to a method in which associating a head pose with an eye gaze includes: associating the head pose with the eye gaze based on a first time indicator associated with the eye gaze and a second time indicator associated with the head pose.
[0037] In some respects, the techniques described herein relate to a method in which the indication of eye gaze includes eye gaze time-series data, and the indication of head posture includes head posture time-series data.
[0038] In some aspects, the technology described herein relates to a method that further includes: obtaining an indication of a second eye gaze associated with the user device; obtaining an indication of a set of blinks associated with the user device; and determining that the second eye gaze is associated with at least one blink in the set of blinks, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture includes: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one blink in the set of blinks.
[0039] In some aspects, the technology described herein relates to a method that further includes: obtaining an indication of a second eye gaze associated with the user device; obtaining an indication of a set of saccades associated with the user device; and determining that the second eye gaze is associated with at least one saccade in the set of saccades, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture includes: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one saccade in the set of saccades.
[0040] In some aspects, the technology described herein relates to a method that further includes: converting the eye gaze into a six-dimensional (6D) representation of the eye gaze, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture includes: determining the set of predicted eye gazes based on the converted 6D representation of the eye gaze.
[0041] In some aspects, the techniques described herein relate to a method in which determining the set of predicted eye gazes based on an obtained indication of eye gaze and an obtained indication of head posture includes: determining the set of predicted eye gazes based on a first set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks.
[0042] In some respects, the techniques described herein relate to a method in which determining a set of predicted eye gazes based on an obtained indication of eye gaze and an obtained indication of head posture comprises: determining the set of predicted eye gazes based on a predictor model using the eye gaze and the head posture as input.
[0043] In some respects, the techniques described herein relate to a method in which the predictor model includes at least one of a regression model, a time series predictor model, or a neural network.
[0044] In some aspects, the technology described herein relates to a method that further includes: obtaining an indication of a second eye gaze associated with the user device; determining that the second eye gaze is associated with a user tracking a virtual object; in response to determining that the eye gaze is associated with the user tracking the virtual object, determining a second set of predicted eye gazes based on the obtained indication of the second eye gaze relative to the user's head; and outputting an indication of the determined second set of predicted eye gazes.
[0045] In some aspects, the technology described herein relates to a method that further includes: determining that the eye gaze is not associated with the user tracking the virtual object, wherein determining the set of predicted eye gazes based on obtained indications of the eye gaze and obtained indications of the head posture includes: determining the set of predicted eye gazes based on obtained indications of the eye gaze and obtained indications of the head posture in response to determining that the eye gaze is not associated with the user tracking the virtual object.
[0046] In some aspects, the technology described herein relates to a method in which obtaining the indication of an eye gaze associated with the user equipment includes: generating a set of eye gazes including the eye gaze at a first sampling rate via an eye gaze camera associated with the user equipment, wherein obtaining the head pose includes: generating a set of head poses including the head pose via an inertial measurement unit (IMU) associated with the user equipment at a second sampling rate.
[0047] In some respects, the techniques described herein relate to a method in which the first sampling rate is less than the second sampling rate.
[0048] In some aspects, the technology described herein relates to a method in which the indication of eye gaze includes first translation information and first orientation information of at least one eye of a user associated with the user device, and the indication of head posture includes second translation information and second orientation information of the user's head associated with the user device.
[0049] In some aspects, the techniques described herein relate to a method in which outputting an indication of a determined set of predicted eye gazes comprises: sending the indication of the determined set of predicted eye gazes; receiving an indication of a set of rendered frames after sending the indication of the determined set of predicted eye gazes; and outputting the indication of the set of rendered frames.
[0050] In some respects, the techniques described herein relate to a method in which the output of the instruction for the set of rendered frames includes storing the set of rendered frames in at least one of a memory, a buffer, or a cache.
[0051] In some aspects, the technology described herein relates to a method that further includes: decoding the set of rendered frames; and warping the decoded set of rendered frames based on the latest available head pose of a user associated with the user device, wherein outputting an indication of the set of rendered frames includes: outputting an indication of the set of warped decoded rendered frames.
[0052] In some respects, the techniques described herein relate to a method in which each of the set of rendered frames includes a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the set of predicted eye gazes.
[0053] In some aspects, the technology described herein relates to a method in which obtaining the indication of eye gaze associated with the user device includes receiving the indication of eye gaze from the user device, and obtaining the indication of head posture associated with the user device includes receiving the indication of head posture from the user device.
[0054] In some respects, the techniques described herein relate to a method in which outputting an indication of a determined set of predicted eye gazes includes: outputting the indication to a renderer, wherein the method further includes: rendering a set of frames via the renderer based on the determined set of predicted eye gazes; and sending the rendered set of frames to the user device.
[0055] In some aspects, the techniques described herein relate to a method in which rendering a set of frames based on a determined set of predicted eye gazes includes: rendering a first region and a second region of each frame in the set of frames, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the determined set of predicted eye gazes.
[0056] In some respects, the techniques described herein relate to a method in which sending a rendered set of frames to the user equipment includes: encoding the rendered set of frames; and sending the encoded rendered set of frames.
[0057] To achieve the foregoing and related objectives, one or more aspects include the features fully described below and specifically pointed out in the claims. The following description and drawings set forth some exemplary features of one or more aspects in detail. However, these features indicate only some of the various ways in which the principles of the various aspects may be employed, and this description is intended to include all such aspects and their equivalents. Attached Figure Description
[0058] Figure 1 This is a block diagram illustrating an example of a system for generating content based on one or more techniques of this disclosure.
[0059] Figure 2 Example graphics processors (e.g., graphics processing units (GPUs)) according to one or more technologies according to this disclosure are illustrated.
[0060] Figure 3 Example images or surfaces are illustrated according to one or more techniques of this disclosure.
[0061] Figure 4 This is an illustration of an example of a split extended reality (XR) system according to one or more technologies of this disclosure.
[0062] Figure 5 The diagram illustrates an example aspect of a predictive model for eye fixation prediction based on one or more techniques according to this disclosure.
[0063] Figure 6 This is a diagram illustrating example simulation results related to a prediction model according to one or more techniques of this disclosure.
[0064] Figure 7A The illustrations are examples of aspects related to using XR metadata to enhance eye gaze prediction according to one or more techniques of this disclosure.
[0065] Figure 7B The illustrations are examples of aspects related to using XR metadata to enhance eye gaze prediction according to one or more techniques of this disclosure.
[0066] Figure 8 This is a flowchart illustrating an example workflow for predicting a user's gaze according to one or more techniques of this disclosure.
[0067] Figure 9This is an illustration of an example aspect of a set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks configured to predict a user's gaze according to one or more techniques of this disclosure.
[0068] Figure 10 These are illustrations illustrating example aspects of gaze predictor design according to one or more techniques of this disclosure.
[0069] Figure 11 This is a connection flowchart illustrating an example of a server configured to provide predictive gaze to a client according to one or more techniques of this disclosure.
[0070] Figure 12 This is a connection flowchart illustrating an example of a server configured to provide predictive gaze to a client according to one or more techniques of this disclosure.
[0071] Figure 13 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure.
[0072] Figure 14 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure.
[0073] Figure 15 This is a flowchart of an example method for graphical processing according to one or more techniques of this disclosure. Detailed Implementation
[0074] Various aspects of the systems, apparatuses, computer program products, and methods will be described more fully below with reference to the accompanying drawings. However, this disclosure may be embodied in many different forms and should not be construed as limited to any particular structure or function presented throughout this disclosure. Rather, these aspects are provided to make this disclosure comprehensive and complete, and to fully convey the scope of this disclosure to those skilled in the art. Based on the teachings herein, those skilled in the art will understand that the scope of this disclosure is intended to cover any aspect of the systems, apparatuses, computer program products, and methods disclosed herein, whether implemented independently of or in combination with other aspects of this disclosure. For example, any number of aspects set forth herein may be used to implement an apparatus or practice. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods practiced using structures, functionalities, or structures and functionalities other than or different from the various aspects of the disclosure set forth herein. Any aspect disclosed herein may be embodied by one or more elements of the claims.
[0075] Although various aspects are described herein, many variations and substitutions of these aspects fall within the scope of this disclosure. While some potential benefits and advantages of the aspects of this disclosure are mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or objective. Rather, the aspects of this disclosure are intended to be broadly applicable to different wireless technologies, system configurations, processing systems, networks, and transmission protocols, some of which are illustrated by way of example in the accompanying drawings and the description below. The detailed description and drawings are merely illustrative and not limiting of this disclosure, and the scope of this disclosure is defined by the appended claims and their equivalents.
[0076] Several aspects are presented with reference to various apparatuses and methods. These apparatuses and methods are described in detail and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether these elements are implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system.
[0077] For example, an element, any part of an element, or any combination of elements can be implemented as a “processing system” including one or more processors (which may also be referred to as processing units). Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, system-on-a-chip (SoCs), baseband processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic units, discrete hardware circuits, and other suitable hardware configured to perform the various functionalities described throughout this disclosure. One or more processors in the processing system can execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other names, software is broadly understood to mean instructions, instruction sets, code, code segments, program code, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc.
[0078] The term "application" can refer to software. As described herein, one or more technologies can refer to an application (e.g., software) configured to perform one or more functions. In such examples, the application may be stored in memory (e.g., on-chip memory of a processor, system memory, or any other memory). Hardware described herein, such as a processor, may be configured to execute the application. For example, an application may be described as including code that, when executed by the hardware, causes the hardware to perform one or more technologies described herein. As an example, the hardware may access and execute code accessed from memory to perform one or more technologies described herein. In some examples, components are identified in this disclosure. In such examples, a component may be hardware, software, or a combination thereof. Each component may be a separate component or a subcomponent of a single component.
[0079] In one or more examples described herein, the described functionality can be implemented in hardware, software, or any combination thereof. If implemented in software, the functionality can be stored or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media can be any available medium accessible by a computer. By way of example, and not limitation, such computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disc storage devices, magnetic disk storage devices, other magnetic storage devices, combinations of computer-readable media of the types described above, or any other medium that can be used to store computer-executable code in the form of instructions or data structures accessible by a computer.
[0080] As used herein, instances of the term "content" may refer to "graphic content," "image," etc., regardless of whether the term is used as an adjective, noun, or other part of speech. In some examples, as used herein, the term "graphic content" may refer to content produced by one or more processes in a graphics processing pipeline. In other examples, as used herein, the term "graphic content" may refer to content produced by a processing unit configured to perform graphics processing. In yet another example, as used herein, the term "graphic content" may refer to content produced by a graphics processing unit.
[0081] Split rendering (i.e., split rendering process) refers to an example where a client device (e.g., a head-mounted unit (HMU) or head-mounted display (HMD)) collaborates with a server to facilitate the display of graphical content on the client device's screen, where the client device and the server can communicate with each other via wired and / or wireless communication. For example, a portion of a rendering task (or other task) can be offloaded to the server. The server can perform the rendering task (or other task). The server can send the output of the rendering task (or other task) to the client device. The client device can perform additional processing based on the received output to display the graphical content. Split rendering (e.g., split XR rendering) allows the client device to render relatively high-quality graphical content on its screen while conserving the client device's battery life.
[0082] In a split-rendering system, a client device can generate the user's eye gaze and send it to a server. The server can then use the eye gaze to render and encode a frame. For example, the server can render and encode the frame at a higher quality in the concave region and at a lower quality in the non-concave region. However, the user's eye gaze may change between the time the eye gaze is generated and the time the rendered frame is displayed. Therefore, the user may focus on the non-concave region of the frame (e.g., the lower-quality region of the frame), which could affect the user experience.
[0083] This paper describes various techniques related to eye gaze prediction for split extended reality (XR). In an example, a device (e.g., a client device) provides a first indication of a set of head poses for the user's head and a second indication of a set of eye gazes for the user's eyes at a first time instance. The device (e.g., the client device) obtains a rendered frame based on the first and second indications, which is based on the predicted eye gazes of the user at a second time instance occurring after the first time instance. The device (e.g., the client device) outputs an indication of this rendered frame. Instead of providing the first indication of the head poses and the second indication of the eye gazes, the device (e.g., the client device) can obtain a rendered frame based on the predicted eye gazes of the user at the time the frame is to be displayed, which improves the user experience.
[0084] In another example, a device (e.g., a server) obtains a first indication of a set of head poses for a user's head and a second indication of a set of eye gazes for the user's eyes, wherein the first and second indications correspond to a first time instance. The device (e.g., the server) generates a predicted eye gaze for the user based on the first and second indications, wherein the predicted eye gaze for the user corresponds to a second time instance occurring after the first time instance. The device (e.g., the server) renders a frame based on the user's predicted eye gaze. The device (e.g., the server) outputs the rendered frame. Compared to generating the user's predicted eye gaze based on the first indication of the set of head poses and the second indication of the set of eye gazes, this device (e.g., the server) can generate a more accurate predicted eye gaze compared to a device that predicts eye gazes based on eye gazes (without head poses). Furthermore, since the frame is rendered based on the user's predicted eye gaze, the user experience is improved.
[0085] In another example, the device may be configured to obtain an indication of eye gaze associated with a user device. The device may be configured to obtain an indication of head posture associated with the user device. The device may be configured to determine a set of predicted eye gazes based on the obtained indication of eye gaze and the obtained indication of head posture. The device may be configured to output an indication of the determined set of predicted eye gazes.
[0086] In some respects, eye gaze predictors can leverage head pose prediction to assist eye gaze prediction in split XR systems. Since head pose data can be obtained at a higher sampling rate than eye gaze data, eye gaze predictors can use high-sample-rate head pose data to predict eye gazes. An autoregressive model can be used on past eye gazes (e.g., the past 3 frames) to predict future gazes (e.g., the next 3 frames). Head pose rotation can be additionally used as input to this autoregressive model. XR metadata, such as object depth and motion vectors from the encoder, can also be used as input to this autoregressive model to enhance the prediction.
[0087] The examples described herein may relate to the use and functionality of a graphics processing unit (GPU). As used herein, a GPU can be any type of graphics processor, and a graphics processor can be any type of processor designed or configured to process graphical content. For example, a graphics processor or GPU can be a dedicated circuit designed to process graphical content. As an additional example, a graphics processor or GPU can be a general-purpose processor configured to process graphical content.
[0088] Figure 1This is a block diagram illustrating an example content generation system 100 configured to implement one or more technologies of this disclosure. The content generation system 100 includes a device 104. Device 104 may include one or more components or circuitry for performing the various functions described herein. In some examples, one or more components of device 104 may be components of a System-on-a-Chip (SOC). Device 104 may include one or more components configured to perform one or more technologies of this disclosure. In the illustrated example, device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, device 104 may include multiple components (e.g., a communication interface 126, a transceiver 132 or antenna, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). Display 131 may refer to one or more displays 131. For example, display 131 may include a single display or multiple displays, which may include a first display and a second display. The first display may be a left-eye display, and the second display may be a right-eye display. In some examples, the first and second displays may receive different frames for presentation on the first and second displays. In other examples, the first and second displays may receive the same frames used for rendering on both displays. In yet another example, the results of graphics processing may not be displayed on the device; for example, the first and second displays may not receive any frames used for rendering on either display. Instead, the frames or graphics processing results may be transferred to another device. In some respects, this can be referred to as split rendering.
[0089] Processing unit 120 may include internal memory 121. Processing unit 120 may be configured to perform graphics processing using graphics processing pipeline 107. Content encoder / decoder 122 may include internal memory 123. In some examples, device 104 may include a processor configured to perform one or more display processing techniques on one or more frames generated by processing unit 120, and then display those frames through one or more displays 131. Although the processor in example content generation system 100 is configured as display processor 127, it should be understood that display processor 127 is one example of a processor and other types of processors, controllers, etc., may be used instead of display processor 127. Display processor 127 may be configured to perform display processing. For example, display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by processing unit 120. One or more displays 131 may be configured to display or otherwise present the frames processed by display processor 127. In some examples, one or more displays 131 may include one or more of the following: liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, projection display device, augmented reality display device, virtual reality display device, head-mounted display, or any other type of display device.
[0090] Memory (such as system memory 124) external to processing unit 120 and content encoder / decoder 122 may be accessible to processing unit 120 and content encoder / decoder 122. For example, processing unit 120 and content encoder / decoder 122 may be configured to read from and / or write to external memory (such as system memory 124). Processing unit 120 may be communicatively coupled to system memory 124 via a bus. In some examples, processing unit 120 and content encoder / decoder 122 may be communicatively coupled to internal memory 121 via the bus or via a different connection.
[0091] Content encoder / decoder 122 can be configured to receive graphic content from any source, such as system memory 124 and / or communication interface 126. System memory 124 can be configured to store received encoded or decoded graphic content. Content encoder / decoder 122 can be configured to receive encoded or decoded graphic content from system memory 124 and / or communication interface 126, for example, in the form of encoded pixel data. Content encoder / decoder 122 can be configured to encode or decode any graphic content.
[0092] Internal memory 121 or system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, internal memory 121 or system memory 124 may include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, magnetic data media or optical storage media, or any other type of memory. According to some examples, internal memory 121 or system memory 124 may be a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that internal memory 121 or system memory 124 is not removable or that its contents are static. For example, system memory 124 may be removed from device 104 and moved to another device. Alternatively, system memory 124 may not be removable from device 104.
[0093] Processing unit 120 may be a CPU, GPU, GPGPU, or any other processing unit configured to perform graphics processing. In some examples, processing unit 120 may be integrated into the motherboard of device 104. In other examples, processing unit 120 may reside on a graphics card mounted in a port on the motherboard of device 104, or may otherwise be incorporated into a peripheral device configured to interoperate with device 104. Processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic components, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, processing unit 120 may store instructions for software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 121) and may use one or more processors to execute instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) may be considered as one or more processors.
[0094] The content encoder / decoder 122 can be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 may be integrated into the motherboard of device 104. The content encoder / decoder 122 may include one or more processors, such as one or more microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic components, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, the content encoder / decoder 122 may store instructions for software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 123) and may use one or more processors to execute instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors.
[0095] In some aspects, the content generation system 100 may include a communication interface 126. The communication interface 126 may include a receiver 128 and a transmitter 130. The receiver 128 may be configured to perform any of the receiving functions described herein with respect to device 104. Additionally, the receiver 128 may be configured to receive information from another device, such as eye or head positioning information, rendering commands, and / or location information. The transmitter 130 may be configured to perform any of the transmitting functions described herein with respect to device 104. For example, the transmitter 130 may be configured to transmit information to another device, which may include a request for content. The receiver 128 and the transmitter 130 may be combined to form a transceiver 132. In such an example, the transceiver 132 may be configured to perform any of the receiving and / or transmitting functions described herein with respect to device 104.
[0096] Refer again Figure 1In some aspects, processing unit 120 may include eye gaze predictor 198, configured to: provide at a first time instance a first indication of a set of head poses for a user's head and a second indication of a set of eye gazes for the user's eyes; obtain a rendered frame based on the first and second indications, the rendered frame being based on the user's predicted eye gaze at a second time instance occurring after the first time instance; and output an indication of the rendered frame. In some aspects, processing unit 120 may include eye gaze predictor 199, configured to: obtain a first indication of a set of head poses for a user's head and a second indication of a set of eye gazes for the user's eyes, wherein the first and second indications correspond to a first time instance; generate a predicted eye gaze for the user based on the first and second indications, wherein the user's predicted eye gaze corresponds to a second time instance occurring after the first time instance; render a frame based on the user's predicted eye gaze; and output the rendered frame.
[0097] In other respects, the eye gaze predictor 198 may be configured to obtain an indication of an eye gaze associated with a user device. The eye gaze predictor 198 may be configured to obtain an indication of a head posture associated with the user device. The eye gaze predictor 198 may be configured to determine a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture. The eye gaze predictor 198 may be configured to output an indication of the determined set of predicted eye gazes.
[0098] While the following description may focus on graphics processing, the concepts described herein are applicable to other similar processing techniques. Furthermore, although the following description may emphasize split rendering, the concepts described herein are also applicable to non-split rendering.
[0099] Devices such as device 104 can refer to any device, apparatus, or system configured to perform one or more of the technologies described herein. For example, a device can be a server, base station, user equipment, client device, station, access point, computer (such as a personal computer, desktop computer, laptop computer, tablet computer, computer workstation, or mainframe computer), end product, apparatus, telephone, smartphone, server, video game platform or console, handheld device (such as a portable video game device or personal digital assistant (PDA)), wearable computing device (such as a smartwatch, augmented reality device, or virtual reality device), non-wearable device, display or display device, television, set-top box, intermediate network device, digital media player, video streaming device, content streaming device, in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more of the technologies described herein. The processes described herein may be described as being performed by a specific component (e.g., GPU), but in other embodiments, other components (e.g., CPU) consistent with the disclosed embodiments may be used to perform them.
[0100] A GPU can process various types of data or data packets within its pipeline. For example, in some aspects, a GPU can process two types of data or data packets, such as context register packets and draw call data. Context register packets can be a set of global state information, such as information about global registers, shaders, or constant data, which can adjust how the graphics context will be processed. For example, a context register packet may include information about the color format. In some aspects of a context register packet, there may be one or more bits indicating which workload belongs to the context register. Additionally, multiple functions or programs can run simultaneously and / or in parallel. For example, a function or program may describe an operation, such as a color mode or color format. Therefore, context registers can define various states of the GPU.
[0101] Context states can be used to determine how individual processing units (e.g., vertex extractors (VFDs), vertex shaders (VSs), shader processors, or geometry processors) operate and / or in which mode they operate. To do this, the GPU uses context registers and programming data. In some aspects, the GPU can generate workloads in the pipeline based on the context register definitions of modes or states, such as vertex or pixel workloads. Certain processing units (e.g., VFDs) can use these states to determine certain functions, such as how to aggregate vertices. Because these modes or states can change, the GPU may need to modify the corresponding context. Additionally, the workload corresponding to a mode or state may follow the changed mode or state.
[0102] Figure 2 Example GPU 200 is illustrated according to one or more technologies according to this disclosure. For example... Figure 2 As shown, GPU 200 includes a command processor (CP) 210, a draw call group 212, a VFD 220, a VS 222, a vertex cache (VPC) 224, a triangle setup engine (TSE) 226, a rasterizer (RAS) 228, a Z-process engine (ZPE) 230, a pixel interpolator (PI) 232, a fragment shader (FS) 234, a rendering backend (RB) 236, an L2 cache (UCHE) 238, and system memory 240. Although Figure 2 The GPU 200 includes processing units 220 to 238, but the GPU 200 may include multiple additional processing units. Additionally, processing units 220 to 238 are merely examples, and the GPU may use any combination or order of processing units in accordance with this disclosure. The GPU 200 also includes a command buffer 250, a context register group 260, and a context state 261.
[0103] like Figure 2 As shown, the GPU can use a CP (e.g., CP 210) or a hardware accelerator to resolve the command buffer into context register groups (e.g., context register group 260) and / or draw call data groups (e.g., draw call group 212). Subsequently, CP 210 can transfer the context register group 260 or the draw call group 212 to a processing unit or block in the GPU via a separate path. Furthermore, the command buffer 250 can alternate between different states of the context registers and draw calls. For example, the command buffer can simultaneously store the following information: the context register of context N, the draw call of context N, the context register of context N+1, and the draw call of context N+1.
[0104] GPUs can render images in a variety of different ways. In some cases, GPUs can render images using direct rendering and / or tiled rendering. In a tiled rendering GPU, an image can be divided or separated into different parts or tiles. After the image is divided, each part or tile can be rendered individually. A tiled rendering GPU can divide a computer graphics image into a grid format, so that each part of the grid (i.e., a tile) is rendered individually. In some aspects of tiled rendering, the image can be divided into different bins or tiles during binning passes. In some aspects, a visibility stream can be constructed during binning passes, where visible primitives or draw calls can be identified. A rendering pass can be performed after a binning pass. In contrast to tiled rendering, direct rendering does not divide a frame into smaller bins or tiles. Instead, in direct rendering, the entire frame is rendered at once (i.e., without binning passes). Additionally, some types of GPUs allow both tiled rendering and direct rendering (e.g., flex rendering).
[0105] In some respects, a GPU can apply the drawing or rendering process to different bins or tiles. For example, a GPU can render a bin and perform all drawing for the primitives or pixels within that bin. During the bin-based rendering process, the rendering target can be located in GPU Internal Memory (GMEM). In some instances, after rendering a bin, the contents of the rendering target can be moved to system memory, and GMEM can be freed to render the next bin. Additionally, a GPU can render another bin and perform drawing for the primitives or pixels within that bin. Thus, in some respects, there may be a small number of bins covering all the drawing on a surface, for example, four bins. Furthermore, a GPU can loop through all the drawing in a bin but perform drawing only for visible drawing calls, i.e., drawing calls that include visible geometry. In some respects, a visibility stream can be generated, for example, in binning passes, to determine the visibility information of each primitive in an image or scene. For example, such a visibility stream can identify whether a primitive is visible. In some respects, this information can be used to remove invisible primitives, such that, for example, invisible primitives are not rendered in a rendering pass. Additionally, at least some primitives that are marked as visible can be rendered in the rendering pass.
[0106] In some aspects of tile rendering, there can be multiple processing stages or passes. For example, rendering can be performed in two passes, such as a binning, visibility, or box visibility pass and a rendering or box rendering pass. During a visibility pass, the GPU can input a rendering workload, record the positions of primitives or triangles, and then determine which primitives or triangles fall into which bins or regions. In some aspects of a visibility pass, the GPU can also identify or mark the visibility of each primitive or triangle in the visibility stream. During a rendering pass, the GPU can input a visibility stream and process one bin or region at a time. In some aspects, the visibility stream can be analyzed to determine which primitives or primitive vertices are visible or invisible. Thus, visible primitives or primitive vertices can be processed. By doing so, the GPU can reduce the unnecessary workload of processing or rendering invisible primitives or triangles.
[0107] In some aspects, certain types of primitive geometry, such as localized geometry, can be processed during visibility passes. Additionally, primitives can be categorized into different bins or regions based on their localization or position. In some instances, categorizing primitives or triangles into different bins can be performed by determining visibility information for those primitives or triangles. For example, the GPU can determine the visibility information for each primitive in each bin or region or write it to, for example, system memory. This visibility information can be used to determine or generate a visibility stream. In a rendering pass, the primitives in each bin can be rendered individually. In these cases, the visibility stream can be retrieved from memory and used to remove primitives that are not visible to that bin.
[0108] Some aspects of the GPU or GPU architecture can provide multiple different options for rendering (e.g., software rendering and hardware rendering). In software rendering, the driver or CPU can process each view... Figure 1 The entire frame geometry is copied each time. Additionally, some different states can change depending on the viewpoint. Therefore, in software rendering, the software can copy the entire workload by changing some states that can be used for rendering for each viewpoint in the image. In some respects, this can lead to increased overhead because the GPU may submit the same workload multiple times for each viewpoint in the image. In hardware rendering, the hardware or GPU may be responsible for copying or processing the geometry for each viewpoint in the image. Therefore, the hardware can manage the copying or processing of primitives or triangles for each viewpoint in the image.
[0109] Figure 3 An image or surface 300 according to one or more techniques of this disclosure is illustrated, including multiple elements divided into multiple boxes. For example... Figure 3As shown, the image or surface 300 includes a region 302, which includes primitives 321, 322, 323, and 324. Primitives 321, 322, 323, and 324 are divided or placed into different bins, such as bins 310, 311, 312, 313, 314, and 315. Figure 3 This example illustrates tile rendering using multiple viewpoints for primitives 321-324. For instance, primitives 321-324 are in a first viewpoint 350 and a second viewpoint 351. Therefore, GPU processing or rendering of an image or surface 300 including region 302 can utilize multi-view or multi-view rendering.
[0110] As indicated in this article, GPUs or graphics processors can use tile rendering architectures to reduce power consumption or save memory bandwidth. As further stated above, this rendering method divides the scene into multiple bins, along with visibility paths that identify the visible triangles within each bin. Therefore, in tile rendering, the entire screen can be divided into multiple bins or tiles. The scene can then be rendered multiple times, for example, once or multiple times for each bin.
[0111] In various aspects of graphics rendering, some graphics applications may render a single target (i.e., the rendering target) once or multiple times. For example, in graphics rendering, the frame buffer on system memory can be updated multiple times. The frame buffer can be part of memory or random access memory (RAM) (e.g., containing bitmaps or storage devices) to help store display data for the GPU. The frame buffer can also be a memory buffer containing a complete frame of data. Additionally, the frame buffer can be a logical buffer. In some aspects, updating the frame buffer can be performed in bin or tile rendering, where, as discussed above, the surface is divided into multiple bins or tiles, and each bin or tile can then be rendered individually. Furthermore, in tile rendering, the frame buffer can be divided into multiple bins or tiles.
[0112] As this article points out, in some respects, such as in boxed or tiled rendering architectures, frame buffers allow data to be repeatedly stored or written to them, for example, when rendering from different types of memory. This can be referred to as unresolving the frame buffers or system memory. For example, when storing or writing to one frame buffer and then switching to another, the data or information on the frame buffer can be resolved from the GMEM at the GPU to system memory, i.e., memory in dual data rate (DDR) RAM or dynamic RAM (DRAM).
[0113] In some respects, system memory can also be system-on-chip (SoC) memory or another chip-based memory, such as on a device or smartphone, used for storing data or information. System memory can also be a physical data storage device shared by the CPU and / or GPU. In some respects, system memory can be, for example, a DRAM chip on a device or smartphone. Therefore, SoC memory can be a chip-based method for storing data.
[0114] In some respects, GMEM can be on-chip memory at the GPU, which can be implemented using static RAM (SRAM). Alternatively, GMEM can be stored on the device (e.g., a smartphone). As indicated herein, data or information can be transferred between system memory or DRAM and GMEM, for example, at the device. In some respects, system memory or DRAM can reside at the CPU or GPU. Furthermore, data can be stored in DDR or DRAM. In some respects, such as in bin or tiled rendering, a small portion of the memory can be stored at the GPU, for example, in GMEM. In some cases, storing data at GMEM may utilize a larger processing workload and / or consume more power compared to storing data at the frame buffer or system memory.
[0115] Users can wear display devices to experience extended reality (XR) content. XR can refer to technologies that blend digital experiences with aspects of the real world (i.e., XR content). XR can include augmented reality (AR), mixed reality (MR), and / or virtual reality (VR). In AR, AR objects can be overlaid on the real-world environment perceived through a display device. In an example, AR content can be experienced through AR glasses that include transparent or translucent surfaces. When a user views the environment through the glasses, AR objects can be projected onto the transparent or translucent surface of the glasses. Generally, AR objects may not exist in the real world, and the user may not interact with them. In MR, MR objects can be overlaid on the real-world environment perceived through a display device, and the user can interact with them. In some aspects, MR objects can include "video perspective" with added virtual content. In an example, a user can "touch" the MR object being displayed to them (i.e., the user can place their hand in the real world where the MR object appears to be positioned from the user's perspective), and the MR object can "move" based on being touched (i.e., the position of the MR object on the display can change). Typically, MR content can be experienced through MR glasses (similar to AR glasses) worn by the user or through a head-mounted display (HMD) worn by the user. An HMD may include a camera and one or more display panels. The HMD captures images of the environment perceived through the camera and displays an image of the environment to the user, overlaid with MR objects. Unlike the transparent or translucent surfaces of AR / MR glasses, one or more display panels of an HMD may not be transparent or translucent. In VR, users can experience a fully immersive digital environment where the real world is obscured. VR content can be experienced through an HMD. XR devices can refer to devices capable of displaying XR content.
[0116] Figure 4 Figure 400 illustrates an example of a split extended reality (XR) system according to one or more techniques of this disclosure. In a split XR system, a client 404 (e.g., an HMD) may send head-mounted display (HMD) pose data (i.e., orientation information, eye gaze data, etc.) as data 406 to a server 402. The server 402 may render XR content based on the HMD pose data and eye gaze data, compress the XR content, and send the XR content, along with the rendered pose, as an encoded bitstream 408 to the client 404. The client 404 may decompress and render the received content after distorting the content using the rendered pose and the latest display pose information. The round-trip time between the client 404 and the server 402 may be represented by a round-trip time interval 410.
[0117] In a split XR system, an eye gaze can be generated at the client HMD at time "t" milliseconds (ms) and transmitted to the server. The server uses the head pose and eye gaze to render and encode a frame. The server can then send the frame back to the client HMD, which can decode and display the frame at time "t+T" milliseconds (ms). The server can use the eye gaze to render / encode the frame at high quality within the fovea region associated with the frame and at lower quality outside the fovea region; however, in either case, the user's eyes can move continuously. If the eye moves significantly from its position at time "t" at time "t+T", the user may focus their eyes on a portion of the frame that is likely to be of lower quality, thus degrading the user experience. The aspects presented in this paper relate to algorithms that can accurately predict eye gazes at time "t+T".
[0118] Figure 5 Figure 500 illustrates example aspects of a prediction model for eye gaze prediction related to one or more techniques according to this disclosure. The aspects presented herein can utilize the fact that head pose and eye gaze can be generated together. The aspects presented herein can utilize head pose prediction to assist eye gaze prediction. In the examples, head pose can be obtained at a high sampling rate (e.g., an inertial measurement unit (IMU) can sample at 1000 Hz), while eye gaze can be obtained at eye-tracking camera frames per second (fps) (e.g., such as 30 Hz, or at intervals of 33 ms).
[0119] In one respect, the predictor model can be a time series predictor. In another respect, the predictor model can be an autoregressive model. In yet another respect, the predictor model can be a neural network. In yet another respect, the predictor model can be a transformer.
[0120] The aspects presented in this article can be linked to the advantages of using historical eye fixation samples to predict eye fixation.
[0121] In one scenario, a user may be looking at a stationary object in a virtual world while moving their head. In such cases, the user's eye gaze may move continuously in the opposite direction of the head. Using the aspects described herein, head pose information can be converted into world coordinates for eye gazes at different time instances (corresponding to eye gazes of stationary objects projected onto the image). World coordinates may be located at the same position (because the object is not moving). Using this information, along with a predicted head pose at time "T+t", the aspects presented herein determine the eye gaze at time "T+t".
[0122] In another scenario, the user may be gazing at a moving object in the virtual world, and the user may move his / her head to track the moving object. In such cases, the algorithm described in this paper can take into account gaze movement relative to the movement of the head and the object. For example, if the object is moving in the x-direction, the actual eye gaze may not move much, and eye gaze prediction can be performed.
[0123] In another scenario, if head movement is largely static and eye gaze is continuous, the eyes may be following a moving object. In such cases, the algorithm described in this paper can learn the object's motion. If the object's motion is linear, an autoregressive model may be appropriate. If the object's motion is non-linear, a neural network can be used to predict the object's motion. In these cases, head pose may not be able to assist in eye gaze prediction.
[0124] According to the first method described herein, eye gazes at times t-66ms, t-33ms, and t ms can be used to predict eye gazes at times t+33ms, t+66ms, and t+99ms using an autoregressive model. For example, predictor model 516 can use head pose data 502, head pose data 504, head pose data 506, eye gaze data 510, eye gaze data 512, eye gaze data 514, and predicted head pose data 508 at time “T+t” as input. Although predictor model 516 is shown to receive three eye gaze and head pose data measures in addition to one predicted head pose measure, predictor model 516 can have any number of historical eye gaze and head pose measures and / or any number of predicted head pose measures as input to determine the predicted eye gaze 518 at time “T+t”. According to the second method, eye gaze and head pose rotation at times t-66ms, t-33ms, and t ms can be used to predict eye gaze at times t+33ms, t+66ms, and t+99ms using an autoregressive model. According to the baseline method, there may be no prediction of eye gaze. For example, an eye gaze at time t can be used as a prediction of eye gaze at times t+33ms, t+66ms, and t+99ms. For the first, second, and baseline methods, rotational errors between the baseline true values and the predicted eye gazes can be calculated at times t+33ms, t+66ms, and t+99ms. Eye gazes are available at 30fps, and therefore eye gaze samples can be spaced 33ms apart.
[0125] Figure 6 Figure 600 illustrates example simulation results 602 and 604 related to a prediction model according to one or more techniques of this disclosure.
[0126] Simulation results demonstrate that using head pose information and eye gaze can help estimate future eye gaze. When using head pose information, rotational errors can be reduced by 99%. The dataset can simulate most stationary objects. The head can move around in the scene gazing at the object without using the object's depth information. The aspects presented in this paper can lead to improved algorithms. To generate improved algorithms, more data can be generated that simulates the other cases discussed above. Improved algorithms can involve nonlinear predictors (i.e., nonlinear prediction models) that use neural networks to model object motion.
[0127] Figure 7A Figure 700 illustrates an example aspect of using XR metadata to enhance eye gaze prediction according to one or more techniques of this disclosure. XR metadata may include the depth of an object. Utilizing the depth of the object (the one the eye is gazing at) can assist in tracking the object in three-dimensional (3D) space. Utilizing the object's depth can help utilize the translational component of head pose. For example, rotational and translational components from head pose can be used to more accurately predict eye gaze. Between frames 702, 704, and 706, a sensor can track a user's eye gaze 708 displayed in the frame relative to an object 710. An eye gaze predictor can calculate a movement 714 of the eye gaze 708 and a motion vector 716 relative to the object 710. The increment of the eye gaze 708 between frames 702, 704, and 706 can be similar to the motion vector 716 in both magnitude and direction. In other words, the user can track the same object between frames 702, 704, and 706.
[0128] XR metadata can include motion vectors from the encoder. The aspects presented in this paper can utilize motion vectors to detect whether the eye has started to gaze at a different object. If the eye has started tracking a different object, the trajectory of the previous object may lead to inaccuracies (e.g., due to attempts to predict a new trajectory based on a sample history of previous trajectories).
[0129] One aspect of this paper relates to a method for predicting eye gaze in XR, which uses past eye gazes from an eye-tracking system and head poses from an on-device head-tracking system. Obtaining the eye gaze at a future time instance can be based on the head pose at the predicted time instance as a first step. Obtaining the eye gaze at a future time instance can be based on predicting the motion of objects in the space surrounding the user that the user's eyes may be tracking. Object motion can be defined based on rotation around the user (e.g., a change in orientation). Additional cues for determining object location or object motion can be obtained from metadata / data associated with the XR, such as depth vectors and motion vectors.
[0130] Figure 7BFigure 750 illustrates an example aspect of using XR metadata to enhance eye gaze prediction according to one or more techniques of this disclosure. For example, between frames 752, 754, and 756, a sensor may track a user's eye gaze 708 relative to objects 710 and 712 displayed in the frames. An eye gaze predictor may calculate a movement 714 of the eye gaze 708 and a motion vector 716 relative to object 710. The increment of the eye gaze 708 between frames 702, 704, and 706 may have different directions when compared to the motion vector 716. In other words, the user may track different objects between frames 702, 704, and 706 (e.g., tracking object 710 first, then object 712). The eye gaze predictor can then make predictions using input samples corresponding to the new objects.
[0131] Figure 8 This is a flowchart 800 illustrating an example workflow for predicting a user's gaze according to one or more techniques of this disclosure. An eye tracker 802 can capture combined eye gaze directions (e.g., x-vector, y-vector, z-vector) for each of the user's eyes over time. At 804, the eye gaze predictor can remove samples taken at and near the time of the user's blink, as such samples may not be as accurate as samples captured within at least a threshold time period from the time the user blinks. For example, a blink can last for 200 ms. The eye gaze predictor can remove samples before and after the blink (e.g., samples within 50 ms before and 50 ms after the blink), or it can remove the first four samples before the blink and the last four samples after the blink. In some aspects, the eye gaze predictor can be configured to remove different numbers of samples before and after the blink, e.g., four samples before the blink and two samples after the blink. A camera can detect a user's blink by tracking eyelid patterns around one or both of the user's eyes.
[0132] At 806, the eye gaze predictor can convert a captured eye gaze into a 6D representation. For example, the eye gaze predictor can convert a gaze direction vector of [x, y, z] into a 6D rotation. A 6D rotation can represent a continuous representation of the rotation. In some aspects, the eye gaze predictor can use a z-vector to obtain the optical axis, where the optical axis = [0, 0, 1]. In some aspects, the eye gaze predictor can use a quaternion head pose [qw, qx, qy, qz] (relative to the world). In other aspects, the eye gaze predictor can use only the eye gaze vector, for example, by setting the head pose to an identifier [0, 0, 0, 1] (relative to the head). The eye gaze predictor can convert the head pose (or, if relative to the head, an identifier of the head pose) into a rotation matrix R. The eye gaze predictor can multiply the rotation matrix with the gaze to obtain a vector product V = inv(R). Gaze prediction. This eye gaze predictor takes the cross product of vector products to obtain the axis, where axis = optical axis × V. This eye gaze predictor calculates the angle, where angle = acos(optical axis V). This eye gaze predictor converts the axis angle into a rotation matrix with 9 elements. This eye gaze predictor discards the last column of the 9-element rotation matrix to obtain a 6D representation.
[0133] At 808, the eye gaze predictor can apply filters to improve the accuracy of the 6D representation input. For example, a causal filter or a low-pass filter. In some respects, the filter can be an exponential filter, such as α x current_sample + (1-α) x previous_sample.
[0134] At 810, the GRU predictor can receive a 6D filtered input and determine the 6D prediction output based on historical input data. For example, Figure 9 The GRU predictor is illustrated in Figure 900. When the GRU predictor receives a filtered 6D input, it updates its hidden state by extrapolating 6D predicted gazes to multiple future timestamps. At 812, the eye gaze predictor converts the 6D prediction output into gaze directions (e.g., x, y, and z vectors) to determine the user's gaze point. At 814, the eye gaze predictor outputs each predicted gaze to the renderer to render a frame based on the predicted gaze and its associated timestamp.
[0135] Figure 9 This is a diagram 900 illustrating an example aspect of a set of fully connected network (FCN) blocks 902 and a set of gated recurrent unit (GRU) blocks 904 configured to predict a user's gaze according to one or more techniques of this disclosure. The set of FCN blocks 902 may be the final layer of an artificial neural network for predicting the gaze of a set of eyes. The set of FCN blocks 902 may be connected to each GRU block of the final layer of the set of GRU blocks 904. The set of GRU blocks may include a gating mechanism for the recurrent neural network. The GRU blocks may accept a 6D continuous representation of the user's gaze, based on an input X. t The number of 6D prediction outputs used to predict user gaze, where X represents information about the gaze at time t. Information about the gaze may include data from filters (such as...). Figure 8The 6D filtered input is the filter at position 808 in the image. Information about gaze may include at least one of the following: a set of predicted head pose, field of view (FOV), estimated depth of the object the user may be focusing on, and / or motion vectors from the encoder for a moving object (e.g., a virtual moving object, a physical moving object) the user may be focusing on. The encoder may receive the trajectory of a virtual moving object from an application that generates virtual content for a display. The encoder may determine the trajectory of a physical moving object by analyzing information from a camera that detects objects moving within the user's FOV. The GRU may be configured to be based on the input X... t To predict the next gaze. In some respects, a set of GRU blocks 904 can be arranged in a set of layers to improve prediction accuracy. Although four blocks are shown in a set of GRU blocks 904, more or fewer GRU layers can be used to increase / decrease prediction accuracy. In other words, a set of four GRU blocks allows each timestamp prediction to be based on at least four previously captured eye gazes. A set of GRU blocks 904 can be fed into a set of FCN blocks 902 to provide a set of gaze predictions for future timestamps. The predicted eye gaze of the FCN block can be fed into the next set of GRU layers. In other words, if the FCN block predicts eye gaze Y... t+1 Then predict the eye's gaze at Y t+1 This can be fed into the input of the bottommost GRU layer in this group of GRU blocks 904. This group of FCN blocks 902 can predict 9 future time periods (i.e., Y). t+1 To Y t+9 () or eye gaze at the next 9 timestamps (TFP). For example, if each input X t Captured by a camera that captures eye gaze every 11 ms, then Figure 9 The illustrated GRU predictor can predict eye gazes in the future 99ms. In some aspects, the GRU predictor can be designed to predict the number of eye gazes with a delay greater than or equal to the round-trip time between sending an eye gaze to the server and receiving a rendered frame from the server. For example, if the round-trip time delay is 90ms, the GRU predictor can be designed to predict the number of eye gazes with a delay greater than or equal to that delay (e.g., 9 predicted eye gazes with an 11ms time interval can be used to predict eye gazes in the 99ms period greater than 90ms).
[0136] Figure 10Figure 1000 illustrates an example aspect of a gaze predictor design according to one or more techniques of this disclosure. Eye tracker input 1002 may include a set of captured eye gazes over time, e.g., eye gazes every 11 ms. Eye gazes may be captured as gaze vectors, e.g., in [x, y, z] directions. Eye tracker input 1002 may include a validity flag. This validity flag may indicate the presence of a blink during a time period. In other words, an eye gaze vector sent with the validity flag set may indicate an eye gaze captured when there was no blink during that time period, and an eye gaze vector sent without the validity flag set may indicate an eye gaze captured when there was a blink during that time period. Eye tracker input 1002 may include head pose, e.g., as a quaternion head pose [qw, qx, qy, qz].
[0137] Eye tracker input 1002 may be received by blink detector 1006, which may be configured to remove samples taken during and around the time of a user blink, as such samples may not be as accurate as samples captured within at least a threshold time period from the user's blink. In other words, blink detector 1006 may output a null value when a captured eye gaze is received during a blink. Blink detector 1006 may send a reset flag to gaze predictor 1018, which may reset the state of any GRU within a threshold time period of the detected blink. For example, gaze predictor 1018 may erase GRU state information associated with the previous four eye gaze postures, or may revert the state of the GRU to four eye gaze postures. In other words, blink detector 1006 may be configured to erase data associated with a threshold number of eye gaze postures captured before the detected blink. Blink detector 1006 may also be configured to ignore captured eye gaze postures captured shortly after a blink. For example, the blink detector 1006 can output a null value for the captured eye gaze posture (e.g., the next 2 eye gaze postures or the next 4 eye gazes) within a threshold time after a blink is detected.
[0138] The blink detector 1006 outputs the filtered gaze direction to the converter 1008 to convert the filtered gaze direction into a 6D representation of the head pose. In other words, the converter 1008 can set the head pose to the identifier [0, 0, 0, 1] when calculating the rotation matrix R. The converter 1008 can filter the 6D representation using a low-pass filter 1010, and the filtered 6D representation can then be fed to the sequence detector 1016 to determine the sequence of gaze directions.
[0139] Blink detector 1006 outputs the filtered gaze direction to converter 1012 to convert the filtered gaze direction into 6D relative to the user's absolute world. In other words, converter 1008 can use quaternion head pose [qw, qx, qy, qz] when calculating the rotation matrix R. Converter 1012 can filter the 6D representation using low-pass filter 1014, and the filtered 6D representation can then be fed to sequence detector 1016 to determine the sequence of gaze directions.
[0140] Sequence detector 1016 can detect the state of a user's eye gaze based on a pattern detected by a 6D representation. For example, sequence detector 1016 can determine that the user's eye gaze is in a tracking state, such as tracking a moving object. In some aspects, the user may be tracking a virtual moving object. Sequence detector 1016 can determine that the user's movement vector matches (within an error margin) the movement vector of a virtual object moving within the user's FOV. If sequence detector 1016 determines that the user's eye gaze is in a tracking state, sequence detector 1016 can indicate this situation to XOR block 1028, which can then allow the 6D representation of the user's head from low-pass filter 1010 to be output to gaze predictor 1018. In another example, sequence detector 1016 can determine that the user's eye gaze is in a vestibular-ocular retraction (VOR) state, such as tracking a stationary object (e.g., a virtual stationary object or a physically stationary object) or tracking a moving physical object. Sequence detector 1016 can determine that the user's movement vector matches (within an error margin) the movement vector of a physical object moving within the user's FOV. In other words, the user can move their head while their eyes are focused on a stationary virtual or physical object, or the user is tracking a moving physical object. If the sequence detector 1016 determines that the user's eye gaze is in a VOR state, the sequence detector 1016 can indicate this to the XOR block 1030, which can then allow a 6D representation of the user's world from the low-pass filter 1014 to be output to the gaze predictor 1018.
[0141] In some respects, the sequence detector 1016 can determine that the user's gaze is in a saccade state. In other words, the sequence detector 1016 can determine that the user's gaze is moving so quickly that the user is not focusing on anything. In other words, the measured speed of movement of the eye gaze may be greater than a threshold. The sequence detector 1016 can output an indication of the saccade to the gaze predictor 1018, which can then reset the state of the GRU. In some respects, the gaze predictor 1018 can treat the saccade indication as the same as a blink, for example by erasing the previous four eye gazes, or by reversing the state of a set of GRUs back to four eye gazes. In other respects, the gaze predictor 1018 can treat the saccade indication as different from a blink, for example by erasing the previous two eye gazes instead of four eye gazes. In either case, the gaze predictor 1018 can disable the prediction of eye gazes for a period of time in response to the detected saccade. If sequence detector 1016 determines that the user's eye gaze is in a saccade state, sequence detector 1016 may indicate such a situation to XOR block 1032, which may then simply pass the raw eye gaze data from low-pass filter 1010 to a set of output gazes 1026, essentially disabling gaze prediction during the saccade. In other words, sequence detector 1016 may be unable to determine the sequence based on eye gaze data, which may result in uncertain predictions being output as a set of output gazes 1026 within the time period associated with the detected saccade (e.g., during the saccade, for a first threshold number of eye gazes before the detected saccade, and / or for a second threshold number of eye gazes after the detected saccade).
[0142] The gaze predictor 1018 may receive filtered gaze direction from blink detector 1006, determined sequence information from sequence detector 1016, and / or filtered 6D representation from low-pass filters 1010 and 1014 to predict the user's gaze. In some aspects, the gaze predictor 1018 may reset its predictor state in response to blink detector 1006 detecting a blink. In some aspects, the gaze predictor 1018 may reset its predictor state in response to sequence detector 1016 detecting a saccade. The gaze predictor 1018 may also receive additional input 1004, such as a set of predicted head poses, the user's field of view (FOV), and / or depth / motion vectors from the encoder. In some aspects, the gaze predictor 1018 may ignore / disable data associated with vectors or objects outside the user's FOV (e.g., movement vectors of virtual objects that the user cannot see because they are outside the user's FOV). In some aspects, the gaze predictor 1018 can analyze the combined eye gaze of a user's two eyes. In other aspects, the gaze predictor 1018 can analyze the eye gaze of each of the user's eyes individually. For example, the gaze predictor 1018 can analyze the eye gaze of each of the user's eyes, where the gaze predictor 1018 predicts that the user is following an object with a depth less than or equal to a threshold. In other words, the (physical or virtual) object can be so close to the user's eyes that the user can be looking at each other, and the user's left and right eye gazes can be opposite to each other. The gaze predictor 1018 can output predicted gazes as a set of output gazes 1026, which may include x-gaze vectors, y-gaze vectors, z-gaze vectors (e.g., the next 9 predicted eye gazes) at different timestamps.
[0143] Figure 11 This is a connection flowchart 1100 illustrating an example of a server configured to provide predictive gaze to a client according to one or more technologies of this disclosure. Client 1102 may include an HMU. The server may include a server that renders XR content for display by client 1102. Client 1102 and server 1104 may jointly constitute a split XR system, wherein client 1102 may send HMD pose data to server 1104, and server 1104 renders XR content along with rendering pose information to client 1102. HMD pose data may include head pose data and / or eye gaze data. Server 1104 may compress the rendered data before sending it to client 1102. Client 1102 may decompress and render a frame after warping the rendered frame based on the latest rendered pose information of the user obtained by client 1102.
[0144] At 1108, client 1102 may collect head pose and eye gaze data from a set of sensors monitoring the user (e.g., sensors of an HMU). These sensors may include, for example, a camera monitoring one or both eyes of the user wearing the HMU, or an IMU monitoring the head pose of the user wearing the HMU. Head pose data may include head pose time-series data. Eye gaze data may include eye gaze time-series data. The sensors monitoring the user's head pose and eye gaze information may generate data at different sampling rates. For example, the IMU may generate head pose data at a first sampling rate, and the camera may generate eye gaze data at a second sampling rate, and the first sampling rate may be greater than the second sampling rate. In some aspects, client 1102 may correlate eye gaze data with head pose data by matching timestamps. In other aspects, client 1102 may send raw eye gaze data and head pose data to server 1104 for server 1104 to correlate the eye gaze data with the head pose data by matching timestamps. In other words, a subgroup of head pose data may be associated with a set of eye gaze data, or vice versa. Head pose data and / or eye gaze data may include translation and orientation information. Eye gaze data may include eye gaze data collected individually from each of the user's eyes, or a combination of gazes from both eyes (e.g., the average of estimated x, y, and z values).
[0145] Client 1102 may send instructions to server 1104 regarding the collected data 1110 for processing. Server 1104 may receive instructions regarding the collected data 1110 from client 1102. In some aspects, client 1102 and server 1104 may have wireless transceivers that respectively transmit and receive such data. In some aspects, the collected data 1110 may include a set of predicted head poses of the user. In other aspects, server 1104 may generate a set of predicted head poses of the user.
[0146] At 1112, server 1104 can generate a set of predicted eye gazes based on the collected data 1110 received from client 1102. Server 1104 can generate this set of predicted eye gazes based on historical eye gaze data about the head. Server 1104 can generate this set of predicted eye gazes based on the user's historical eye gaze data and historical head pose data. Server 1104 can generate this set of predicted eye gazes based on XR metadata (such as the depth of physical / virtual objects, the object's motion vector, and / or the user's FOV). Server 1104 can generate this set of predicted eye gazes based on other data (such as predicted head pose data). Server 1104 can generate this set of predicted eye gazes based on predictor models (such as regression models, time series predictor models, or neural networks). In some aspects, server 1104 can use... Figure 8 Workflow Figure 9 FCN blocks and GRU blocks and / or Figure 10 A gaze predictor is used to generate the set of predicted eye gazes.
[0147] At 1114, server 1104 may render a set of frames based on the set of predicted eye gazes. Server 1104 may render the first region of the frame at a higher quality level than the second region of the frame based on the prediction that the user's predicted eye gaze will be in the first region during the time period corresponding to the predicted eye gaze.
[0148] Server 1104 may send an indication to client 1102 of a set of rendered frames 1116. Client 1102 may receive the indication of the set of rendered frames 1116 from server 1104. In some aspects, server 1104 may encode the set of rendered frames before sending, and client 1102 may decode the encoded set of rendered frames after receiving. At 1118, client 1102 may output the rendered frames to, for example, a user's display. In some aspects, client 1102 may warp the rendered frames based on the user's current head pose and current eye gaze. The user's current head pose and current eye gaze may be matched with the user's predicted eye gaze and / or predicted head pose matched with the current timestamp.
[0149] Figure 12This is a connection flowchart 1200 illustrating an example of a server configured to provide predictive gaze to a client according to one or more technologies of this disclosure. Client 1202 may include an HMU. The server may include a server that renders XR content for display by client 1202. Client 1202 and server 1204 may jointly constitute a split XR system, wherein client 1202 may send HMD pose data to server 1204, and server 1204 renders XR content along with rendering pose information to client 1202. HMD pose data may include head pose data and / or eye gaze data. Server 1204 may compress the rendered data before sending it to client 1202. Client 1202 may decompress and render a frame after warping the rendered frame based on the user's latest rendered pose information obtained by client 1202.
[0150] At 1208, client 1202 may collect head pose and eye gaze data from a set of sensors (e.g., sensors of an HMU) monitoring the user. These sensors may include, for example, a camera monitoring one or both eyes of a user wearing an HMU, or an IMU monitoring the head pose of a user wearing an HMU. Head pose data may include head pose time-series data. Eye gaze data may include eye gaze time-series data. The sensors monitoring the user's head pose and eye gaze information may generate data at different sampling rates. For example, the IMU may generate head pose data at a first sampling rate, and the camera may generate eye gaze data at a second sampling rate, and the first sampling rate may be greater than the second sampling rate. In some aspects, client 1202 may correlate eye gaze data with head pose data by matching timestamps. In other words, a subgroup of head pose data may be correlated with a set of eye gaze data, a subgroup of eye gaze data may be correlated with a set of head pose data, or a subgroup of eye gaze data may be correlated with a subgroup of head pose data. Head pose data and / or eye gaze data may include translation information and orientation information. Eye gaze data may include eye gaze data collected individually from each of the user's eyes, or a combination of gazes from both eyes (e.g., the average of estimated x, y, and z values). In some respects, client 1202 may predict a set of predicted head poses of the user.
[0151] At 1212, client 1202 can generate the set of predicted eye gazes based on the collected data. Client 1202 can generate the set of predicted eye gazes based on historical eye gaze data about the head. Client 1202 can generate the set of predicted eye gazes based on the user's historical eye gaze data and historical head pose data. Client 1202 can generate the set of predicted eye gazes based on XR metadata (such as the depth of physical / virtual objects, the motion vectors of objects, and / or the user's FOV). Server 1104 can generate the set of predicted eye gazes based on other data (such as predicted head pose data). Client 1202 can generate the set of predicted eye gazes based on predictor models (such as regression models, time series predictor models, or neural networks). In some aspects, client 1202 can use... Figure 8 Workflow Figure 9 FCN blocks and GRU blocks and / or Figure 10 A gaze predictor is used to generate the set of predicted eye gazes.
[0152] Client 1202 may send instructions to server 1204 regarding the generated prediction data 1213 for processing. Server 1204 may receive instructions regarding the generated prediction data 1213 from client 1202. In some aspects, client 1202 and server 1204 may have wireless transceivers that respectively transmit and receive such data. In some aspects, the generated prediction data 1213 may include a set of predicted eye gazes of the user, the user's predicted head pose, and / or XR metadata.
[0153] At 1214, server 1204 may render a set of frames based on the set of predicted eye gazes. Server 1204 may render the first region of the frame at a higher quality level than the second region of the frame, based on the prediction that the user's predicted eye gaze will be in the first region during the time period corresponding to the predicted eye gaze.
[0154] Server 1204 may send an indication to client 1202 of a set of rendered frames 1216. Client 1202 may receive the indication of the set of rendered frames 1216 from server 1204. In some aspects, server 1204 may encode the set of rendered frames before sending, and client 1202 may decode the encoded set of rendered frames after receiving. At 1218, client 1202 may output the rendered frames to, for example, a user's display. In some aspects, client 1202 may warp the rendered frames based on the user's current head pose and current eye gaze. The user's current head pose and current eye gaze may be matched with the user's predicted eye gaze and / or predicted head pose matched with the current timestamp.
[0155] Figure 13This is a flowchart 1300 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be executed by means such as a graphics processing device, a graphics processor (e.g., a GPU), a CPU, device 104, a client device, a wireless communication device, etc., as in combination with... Figures 1 to 6 , Figure 7A , Figure 7B and Figures 8 to 11 The method is used in various aspects. In the example, the method (including the various aspects detailed below) can be performed by the eye gaze predictor 198.
[0156] At 1302, the device (e.g., a client device) may provide, at a first instance, a first indication of a set of head postures of the user's head and a second indication of a set of eye gazes of the user's eyes. In one example, 1302 may be performed by an eye gaze predictor 198. In another example, 1302 may be performed by... Figure 11 The client 1102 executes this, and can provide instructions to the server 1104 at a first-time instance regarding the collected data 1110. The collected data 1110 may include a first indication of a set of head poses of the user's head and a second indication of a set of eye poses of the user's eyes, collected at 1108. In another example, 1302 may be executed by... Figure 12 The client 1202 executes a function that can provide the eye gaze predictor 198 with an indication of the data collected at 1203 at a first-time instance. The collected data may include information about a set of head postures of the user's head collected at 1208 and a second indication of a set of eye postures of the user's eyes.
[0157] At 1304, the device (e.g., a client device) obtains a rendered frame based on the first and second instructions, the rendered frame being based on the user's predicted eye gaze at a second time instance occurring after the first time instance. In one example, 1304 may be performed by an eye gaze predictor 198. In another example, 1304 may be performed by... Figure 11 The client 1102 executes the process, and can receive instructions from the server 1104 regarding a set of rendered frames 1116. The client 1102 can receive the instructions regarding the set of rendered frames 1116 based on the first and second instructions. At least one rendered frame in the set of rendered frames 1116 can be based on the user's predicted eye pose at a second time instance occurring after the first time instance. In another example, 1304 can be... Figure 12The client 1202 executes, and the client may receive instructions for a set of rendered frames 1216 based on the generated prediction data 1213. The generated prediction data 1213 may be based on the first instruction and the second instruction at 1212. At least one rendered frame in the set of rendered frames 1116 may be based on the user's predicted eye pose at a second time instance occurring after the first time instance (e.g., at least one prediction data in the generated prediction data 1213).
[0158] At 1306, the device (e.g., a client device) outputs an instruction for the rendered frame. In this example, 1306 may be performed by an eye gaze predictor 198. In another example, 1306 may be performed by... Figure 11 The client 1102 in the example can execute instructions for the rendered frame. In another example, 1306 can be executed by... Figure 12 The client 1202 in the middle executes, and the client can output instructions for the rendered frame.
[0159] In one aspect, providing the first instruction and the second instruction may include sending the first instruction and the second instruction to a server at the first time instance, and obtaining the render frame may include receiving the render frame from the server.
[0160] In one aspect, the instruction to output the rendered frame may include: (1) storing the rendered frame in at least one of a memory, a buffer, or a cache, or (2) sending the rendered frame for display.
[0161] In one aspect, the set of head poses may include head pose time series data, and the set of eye gazes may include eye gaze time series data.
[0162] In one aspect, the device (e.g., a client device) can generate the set of head poses via an inertial measurement unit (IMU).
[0163] In one aspect, the device (e.g., a client device) can generate the set of eye gazes via a camera.
[0164] In one aspect, generating the set of head poses may include generating the set of head poses at a first sampling rate, wherein generating the set of eye gazes may include generating the set of eye gazes at a second sampling rate, wherein the first sampling rate may be greater than the second sampling rate.
[0165] In one respect, the rendering frame can be encoded, and the device (e.g., a client device) can decode the rendering frame.
[0166] In one aspect, the device (e.g., a client device) may warp the decoded rendered frame based on the user’s latest available head pose, wherein the indication of the rendered frame output may include the indication of the decoded warped rendered frame output.
[0167] In one aspect, the set of head poses may include first translation information and first orientation information of the user's head, and wherein the set of eye gazes may include second translation information and second orientation information of the user's eyes.
[0168] In one aspect, the rendered frame may include a first region and a second region, wherein the first region may be associated with a first quality level and the second region may be associated with a second quality level, wherein the first quality level may be greater than the second quality level, and wherein the first region may correspond to the user's predicted eye gaze at the second time instance.
[0169] In one aspect, the set of head poses may correspond to a first time instance that occurs before or concurrently with the first time instance, and wherein the set of eye gazes may correspond to a second time instance that occurs before or concurrently with the first time instance.
[0170] In one aspect, the user's set of eye gazes may correspond to the user's eye gaze on the display, and wherein the user's predicted eye gaze may correspond to the user's predicted eye gaze on the display.
[0171] In one aspect, the rendered frame may be further based on extended reality (XR) metadata, which includes at least one of depth information associated with an object or motion vectors associated with the object, wherein the object may be a virtual object or a real-world object.
[0172] Figure 14 This is a flowchart 1400 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be performed by means such as a graphics processing device, a graphics processor (e.g., a GPU), a CPU, device 104, a server, a wireless communication device, etc., as combined with... Figures 1 to 6 , Figure 7A , Figure 7B and Figures 8 to 11 The method is used in various aspects. In the example, the method (including the various aspects detailed below) can be performed by the eye gaze predictor 199.
[0173] At 1402, the device (e.g., a server) obtains a first indication of a set of head postures for the user's head and a second indication of a set of eye gazes for the user's eyes, wherein the first indication and the second indication correspond to a first temporal instance. In one example, 1402 may be performed by an eye gaze predictor 199. In another example, 1402 may be performed by... Figure 11 Server 1104 executes the process, and this server can receive collected data 1110 from client 1102. The collected data 1110 may include a first indication of a set of head poses for the user's head and a second indication of a set of eye poses for the user's eyes. The first and second indications may correspond to a first-time instance. In another example, 1402 may be... Figure 12 The client 1202 executes, and the client can obtain a first indication of a set of head poses of the user's head and a second indication of a set of eye poses of the user's eyes collected at 1208. The first indication and the second indication may correspond to a first time instance.
[0174] At 1404, the device (e.g., a server) generates a predicted eye gaze for the user based on the first instruction and the second instruction, wherein the predicted eye gaze for the user corresponds to a second time instance occurring after the first time instance. In one example, 1404 may be performed by an eye gaze predictor 199. In another example, 1404 may be performed by... Figure 11 Server 1104 in the process executes a function that can generate a predicted eye pose for the user at 1112 based on the first and second instructions. This predicted eye pose for the user may correspond to a second time instance occurring after the first time instance. In another example, 1404 may be executed by... Figure 12 The client 1202 executes, and at 1212, the client can generate a predicted eye pose for the user based on the first instruction and the second instruction. The user's predicted eye pose may correspond to a second time instance that occurs after the first time instance.
[0175] At 1406, the device (e.g., a server) renders a frame based on the user's predicted eye gaze. In this example, 1406 may be performed by an eye gaze predictor 199. In another example, 1406 may be performed by... Figure 11 Server 1104 in the example can render a frame at 1114 based on the user's predicted eye pose. In another example, 1406 can be... Figure 12 The client 1202 executes, which can render frames based on the user's predicted eye pose.
[0176] At 1408, the device (e.g., a server) outputs the rendered frame. In this example, 1408 may be performed by an eye gaze predictor 199. In another example, 1408 may be performed by... Figure 11 Server 1104 in the example executes this, and this server can send a set of rendered frames 1216 to client 1202. In another example, 1408 can be... Figure 12 The client 1202 executes, and this client can output the rendered frame at 1218.
[0177] Figure 15 This is a flowchart 1500 of an example method for graphics processing according to one or more techniques of this disclosure. The method can be performed by means such as a graphics processing device, a graphics processor (e.g., a GPU), a CPU, device 104, client 1102, server 1104, a server, a wireless communication device, etc., as combined with... Figures 1 to 6 , Figure 7A , Figure 7B and Figures 8 to 11 The method is used in various aspects. In the example, the method (including the various aspects detailed below) can be performed by the eye gaze predictor 199.
[0178] At 1502, the device can obtain a first indication of the gaze of a first set of eyes associated with the user equipment. For example, 1502 can be generated by... Figure 11 The server 1104 in the process executes the process and can receive collected data 1110 from the client 1102. The collected data 1110 may include a first indication of eye gaze associated with a first set of eyes on the user device. In another example, 1502 may be... Figure 12 In one example, client 1202 executes the function, which may receive at 1208 a first indication of eye gaze associated with a first set of eyes on the user device. In another example, 1502 may be executed by eye gaze predictor 198. In yet another example, 1502 may be executed by eye gaze predictor 199.
[0179] At 1504, the device can obtain a second indication of a second set of head postures associated with the user equipment. For example, 1504 can be generated by... Figure 11 The server 1104 in the process executes the process and can receive collected data 1110 from the client 1102. The collected data 1110 may include a second indication of a second set of head postures associated with the user equipment. In another example, 1504 may be executed by... Figure 12 In one example, client 1202 executes a second instruction at 1208 for a second set of head poses associated with the user device. In another example, 1504 may be executed by eye gaze predictor 198. In yet another example, 1504 may be executed by eye gaze predictor 199.
[0180] At 1506, the device can determine a third set of predicted eye gazes based on a first set of obtained eye gazes and a second subset of obtained second set of head postures associated with a first subset of the first set of obtained eye gazes. For example, 1506 can be... Figure 11 Server 1104 in the middle executes, and the server can determine a third set of predicted eye gazes at 1112 based on a first set of obtained eye gazes and a second subgroup of obtained second set of head poses associated with a first subgroup of the obtained first set of eye gazes. In another example, 1506 can be performed by Figure 12 The client 1202 executes, and at 1212, the client can determine a third set of predicted eye gazes based on a first set of obtained eye gazes and a second subgroup of obtained head poses associated with a first subgroup of the first set of obtained eye gazes. In another example, 1506 can be executed by eye gaze predictor 199. In another example, 1506 can be executed by eye gaze predictor 198. In another example, 1506 can be executed by eye gaze predictor 199.
[0181] At point 1508, the device can output a third indication of the determined third set of predicted eye fixations. For example, 1508 can be generated by... Figure 11 Server 1104 in the middle executes, and the server can output a third indication of the determined third set of predicted eye gazes to the renderer at 1114 to render a set of frames based on the predicted eye gazes. In another example, 1508 can be executed by Figure 12 The client 1202 executes this, and can send instructions to the server 1204 regarding the generated prediction data 1213. The generated prediction data 1213 may include a third instruction regarding the determined third set of predicted eye gazes. In another example, 1508 may be executed by the eye gaze predictor 199.
[0182] In one aspect, obtaining the first instruction and the second instruction may include receiving the first instruction and the second instruction from a client device, and wherein outputting the rendered frame may include sending the rendered frame to the client device.
[0183] In one aspect, the set of head poses may include head pose time series data, and the set of eye gazes may include eye gaze time series data.
[0184] In one aspect, the set of head poses may be associated with a first sampling rate, and the set of eye gazes may be associated with a second sampling rate, wherein the first sampling rate may be greater than the second sampling rate.
[0185] In one aspect, the apparatus can encode the rendered frame, wherein outputting the rendered frame may include sending the encoded rendered frame.
[0186] In one aspect, the set of head poses may include first translation information and first orientation information of the user's head, and wherein the set of eye gazes may include second translation information and second orientation information of the user's eyes.
[0187] In one aspect, the rendered frame may include a first region and a second region, wherein the first region may be associated with a first quality level and the second region may be associated with a second quality level, wherein the first quality level may be greater than the second quality level, and wherein the first region may correspond to the user's predicted eye gaze at the second time instance.
[0188] In one aspect, the set of head poses may correspond to a first time instance that occurs before or concurrently with the first time instance, and wherein the set of eye gazes may correspond to a second time instance that occurs before or concurrently with the first time instance.
[0189] In one aspect, the user's set of eye gazes may correspond to the user's eye gaze on the display, and wherein the user's predicted eye gaze may correspond to the user's predicted eye gaze on the display.
[0190] In one aspect, the device may obtain extended reality (XR) metadata, wherein generating the user's predicted eye gaze may include further generating the user's predicted eye gaze based on the XR metadata.
[0191] In one aspect, the XR metadata may include at least one of depth information associated with an object or motion vectors associated with the object.
[0192] In one respect, the object can be a virtual object or a real-world object.
[0193] In one aspect, generating the user's predicted eye gaze may include generating the user's predicted eye gaze via a predictor model.
[0194] In one aspect, the predictor model may include at least one of a regression model, a time series predictor model, or a neural network.
[0195] The configuration provides a method or apparatus for graphics processing. The apparatus may be a GPU, a CPU, or some other processor capable of performing graphics processing. In various aspects, the apparatus may be a processing unit 120 within device 104, or it may be other hardware within device 104 or another device. The apparatus may include components for obtaining an indication of an eye gaze associated with a user device. The apparatus may include components for obtaining an indication of a head posture associated with the user device. The apparatus may include components for determining a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture. The apparatus may include components for outputting the determined set of predicted eye gazes. The apparatus may include components for associating the head posture with the eye gaze. The apparatus may include components for associating the head posture with the eye gaze by associating the head posture with the eye gaze based on a first time indicator associated with the eye gaze and a second time indicator associated with the head posture. The indication of the eye gaze may include eye gaze time-series data. The indication of the head posture may include head posture time-series data. The device may include components for obtaining an indication of a second eye gaze associated with the user device. The device may include components for obtaining an indication of a set of blinks associated with the user device. The device may include components for determining that the second eye gaze is associated with at least one blink in the set of blinks. The device may include components for: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one blink in the set of blinks, and determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture. The device may include components for obtaining an indication of a second eye gaze associated with the user device. The device may include components for obtaining an indication of a set of saccades associated with the user device. The device may include components for determining that the second eye gaze is associated with at least one saccade in the set of saccades. The device may include components for: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one saccade in the set of saccades, and determining the set of predicted eye gazes based on an indication of the eye gaze and an indication of the head posture obtained. The device may include components for converting the eye gaze into a 6D representation of the eye gaze. The device may include components for: determining the set of predicted eye gazes based on the converted 6D representation of the eye gaze, and determining the set of predicted eye gazes based on an indication of the eye gaze and an indication of the head posture obtained.The apparatus may include components for: determining a set of predicted eye gazes based on a first set of FCN blocks and a set of GRU blocks, and determining a set of predicted eye gazes based on an indication of the eye gaze and an indication of the head posture obtained from PBCH blocks. The apparatus may include components for: determining a set of predicted eye gazes based on a predictor model using the eye gaze and the head posture as input, and determining a set of predicted eye gazes based on an indication of the eye gaze and an indication of the head posture obtained from PBCH blocks. The predictor model may include at least one of a regression model, a time series predictor model, or a neural network. The apparatus may include components for obtaining an indication of a second eye gaze associated with the user device. The apparatus may include components for determining that the second eye gaze is associated with a user tracking a virtual object. The apparatus may include components for: determining a second set of predicted eye gazes based on an indication of the second eye gaze relative to the user's head in response to determining that the eye gaze is associated with the user tracking the virtual object. The apparatus may include components for outputting an indication of the determined second set of predicted eye gazes. The apparatus may include components for determining that the eye gaze is not associated with the user tracking the virtual object. The apparatus may include components for: determining a set of predicted eye gazes based on an obtained indication of the eye gaze and an obtained indication of the head posture in response to determining that the eye gaze is not associated with a user tracking the virtual object; and determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture. The apparatus may include components for: obtaining the indication of the eye gaze associated with the user device by generating a set of eye gazes including the eye gaze at a first sampling rate via an eye gaze camera associated with the user device. The apparatus may include components for obtaining the head posture by generating a set of head postures including the head posture via an IMU associated with the user device at a second sampling rate. The first sampling rate may be greater than the second sampling rate. The indication of the eye gaze may include first translation information and first orientation information of at least one eye of the user associated with the user device. The indication of the head posture may include second translation information and second orientation information of the user's head associated with the user device. The apparatus may include components for outputting an indication of a determined set of predicted eye gazes by sending the indication of the determined set of predicted eye gazes. The apparatus may include components for receiving an indication of a set of rendered frames after sending the indication of the determined set of predicted eye gazes. The apparatus may include components for outputting an indication of the set of rendered frames.The apparatus may include components for outputting the indication of a determined set of predicted eye gazes by: (a) sending the indication of the determined set of predicted eye gazes, (b) receiving an indication of a set of rendered frames after sending the indication of the determined set of predicted eye gazes, and (c) outputting the indication of the set of rendered frames. The apparatus may include components for outputting the indication of the set of rendered frames by storing the set of rendered frames in at least one of a memory, a buffer, or a cache. The apparatus may include components for decoding the set of rendered frames. The apparatus may include components for warping the decoded set of rendered frames based on the latest available head pose of a user associated with the user equipment. The apparatus may include components for outputting the indication of the set of rendered frames by outputting an indication of the warped, decoded set of rendered frames. Each rendered frame in the set of rendered frames may include a first region and a second region. The first region may be associated with a first quality level, and the second region may be associated with a second quality level. The first quality level may be greater than the second quality level. The first region may correspond to at least one predicted eye gaze in the set of predicted eye gazes. The apparatus may include components for obtaining an indication of an eye gaze associated with the user equipment by receiving an indication of that eye gaze from the user equipment. The apparatus may include components for obtaining an indication of a head posture associated with the user equipment by receiving an indication of that head posture from the user equipment. The apparatus may include components for outputting the indication of an indication for a determined set of predicted eye gazes by (a) outputting the indication to a renderer. The apparatus may include components for rendering a set of frames based on the determined set of predicted eye gazes via the renderer. The apparatus may include components for sending the rendered set of frames to the user equipment. The apparatus may include components for rendering the set of frames based on the determined set of predicted eye gazes by rendering a first region and a second region of each frame in the set of frames. The first region may be associated with a first quality level, and the second region may be associated with a second quality level. The first quality level may be greater than the second quality level. The first region may correspond to at least one predicted eye gaze in the determined set of predicted eye gazes. The apparatus may include components for transmitting a rendered set of frames to the user equipment by: (a) encoding the rendered set of frames, and (b) transmitting the encoded rendered set of frames. In some aspects, the components may include an eye gaze predictor 198. In some aspects, the components may include an eye gaze predictor 199.
[0196] The apparatus may include components for providing a first indication of a set of head poses for a user's head and a second indication of a set of eye gazes for the user's eyes at a first time instance. The apparatus may further include components for obtaining a rendered frame based on the first and second indications, the rendered frame being based on a predicted eye gaze of the user at a second time instance occurring after the first time instance. The apparatus may further include components for outputting an indication of the rendered frame. The apparatus may further include components for generating the set of head poses via an inertial measurement unit (IMU). The apparatus may further include components for generating the set of eye gazes via a camera. The apparatus may further include components for decoding the rendered frame. The apparatus may further include components for warping the decoded rendered frame based on the user's latest available head pose, wherein outputting the indication of the rendered frame includes outputting an indication of the decoded warped rendered frame.
[0197] The configuration provides a method or apparatus for graphics processing. This apparatus may be a GPU, a CPU, or some other processor capable of performing graphics processing. In various aspects, the apparatus may be a processing unit 120 within device 104, or it may be other hardware within device 104 or another device. The apparatus may include components for obtaining a first indication of a set of head poses for a user's head and a second indication of a set of eye gazes for the user's eyes, wherein the first indication and the second indication correspond to a first time instance. The apparatus may further include components for generating a predicted eye gaze for the user based on the first indication and the second indication, wherein the predicted eye gaze for the user corresponds to a second time instance occurring after the first time instance. The apparatus may further include components for rendering a frame based on the user's predicted eye gaze. The apparatus may further include components for outputting the rendered frame. The apparatus may further include components for encoding the rendered frame, wherein outputting the rendered frame includes sending the encoded rendered frame. The apparatus may further include components for obtaining extended reality (XR) metadata, wherein generating the user's predicted eye gaze includes further generating the user's predicted eye gaze based on the XR metadata.
[0198] The configuration provides a method or apparatus for graphics processing. This apparatus may be a GPU, a CPU, or some other processor capable of performing graphics processing. In various aspects, the apparatus may be a processing unit 120 within device 104, or it may be other hardware within device 104 or another device. The apparatus may include components for obtaining a first indication of a first set of eye gazes associated with a user device. The apparatus may include components for obtaining a second indication of a second set of head postures associated with the user device. The apparatus may include components for determining a third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of the obtained second set of head postures associated with a first subgroup of the obtained first set of eye gazes. The apparatus may include components for outputting a third indication of the determined third set of predicted eye gazes. The apparatus may include components for associating the second subgroup of the obtained second set of head postures with the first subgroup of the obtained first set of eye gazes. The device may include components for associating a second subgroup of a obtained second set of head postures with the first subgroup of a obtained first set of eye gazes based on a first set of time indicators associated with the first set of eye gazes and a second set of time indicators associated with the second set of head postures. The first set of eye gazes may include eye gaze time-series data. The second set of head postures may include head posture time-series data. The device may include components for obtaining a fourth indication of a fourth set of blinks associated with the user device. The device may include components for selecting a third subgroup of the obtained first set of eye gazes based on the fourth set of blinks associated with the user device. The device may include components for: determining a third group of predicted eye gazes based on a selected third subgroup of obtained first group of eye gazes and a fourth subgroup of obtained second group of head postures associated with the selected third subgroup of obtained first group of eye gazes; and determining the third group of predicted eye gazes based on the obtained first group of eye gazes and the second subgroup of obtained second group of head postures associated with the first subgroup of obtained first group of eye gazes. The device may include components for determining the third subgroup of obtained first group of eye gazes associated with saccades. The device may include components for: selecting a fourth subgroup of obtained first group of eye gazes based on the determined third subgroup of obtained first group of eye gazes associated with the saccades.The device may include components for: determining a third set of predicted eye gazes based on a selected fourth subgroup of a first set of eye gazes and a fifth subgroup of a second set of head postures associated with the selected fourth subgroup of the first set of eye gazes; and determining the third set of predicted eye gazes based on the first set of eye gazes and a second subgroup of a second set of head postures associated with the first subgroup of the first set of eye gazes. The device may include components for converting the third subgroup of the first set of eye gazes into a fourth set of six-dimensional (6D) representations of the first set of eye gazes. The device may include components for: determining the third set of predicted eye gazes based on the converted fourth set of 6D representations of the first set of eye gazes; and determining the third set of predicted eye gazes based on the first set of eye gazes and a second subgroup of a second set of head postures associated with the first subgroup of the first set of eye gazes. The device may include components for: determining the third set of predicted eye gazes based on a first set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks, and determining the third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of obtained second set of head poses associated with the first subgroup of the obtained first set of eye gazes. The device may also include components for: determining the third set of predicted eye gazes based on a predictor model, and determining the third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of obtained second set of head poses associated with the first subgroup of the obtained first set of eye gazes, the predictor model using the obtained first set of eye gazes and the second subgroup of obtained second set of head poses associated with the first subgroup of the obtained first set of eye gazes as input. The predictor model may include at least one of a regression model, a time series predictor model, or a neural network. The device may also include components for determining the association of the third subgroup of the first set of eye gazes with a user tracking a virtual object. The device may include components for determining a fourth set of predicted eye gazes based on a determined third subgroup of the first set of eye gazes relative to the user's head. The device may include components for outputting a fourth indication of the determined fourth set of predicted eye gazes. The device may include components for determining that the fourth subgroup of the first set of eye gazes is not associated with the user tracking the virtual object. The device may include components for determining the third set of predicted eye gazes based on the determined fourth subgroup of the first set of eye gazes and a fifth subgroup of a second set of head postures obtained in relation to the determined fourth subgroup of the first set of eye gazes, and determining the third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of the obtained second set of head postures obtained in relation to the first subgroup of the obtained first set of eye gazes.The apparatus may include components for obtaining the first set of eye gazes by generating the first set of eye gazes at a first sampling rate via an eye gaze camera associated with the user equipment. The apparatus may include components for obtaining the second set of head poses by generating the second set of head poses at a second sampling rate via an inertial measurement unit (IMU) associated with the user equipment. The first sampling rate may be greater than the second sampling rate. The first set of eye gazes may include a third set of eye gaze time-series data. The second set of head poses may include a fourth set of head pose time-series data. The first set of eye gazes may include first translation information and first orientation information of at least one eye of the user associated with the user equipment. The second set of head poses may include second translation information and second orientation information of the user's head associated with the user equipment. The apparatus may include components for outputting the third indication of the determined third set of predicted eye gazes by: (a) sending the third indication of the determined third set of predicted eye gazes, (b) receiving a fourth indication of a fourth set of rendered frames based on sending the third indication of the determined third set of predicted eye gazes, and (c) outputting a fifth indication of the fourth set of rendered frames. The apparatus may include components for outputting the fifth indication of the fourth set of rendered frames by storing the fourth set of rendered frames in at least one of a memory, a buffer, or a cache. The apparatus may include components for: decoding the fourth set of rendered frames; and distorting the decoded fourth set of rendered frames based on the latest available head pose of the user associated with the user equipment. The apparatus may include components for outputting the fifth indication of the fourth set of rendered frames by outputting a sixth indication of the distorted, decoded fifth set of rendered frames. Each rendered frame in the fourth set of rendered frames may include a first region and a second region. The first region may be associated with a first quality level, and the second region may be associated with a second quality level. The first quality level may be greater than the second quality level. The first region may correspond to at least one predicted eye gaze in the third set of predicted eye gazes. The apparatus may include components for obtaining the first indication of the first set of eye gazes associated with the user equipment by receiving the first indication of the first set of eye gazes from the user equipment. The apparatus may include components for obtaining a second indication of the second set of head poses associated with the user equipment by receiving the second indication of the second set of head poses from the user equipment. The apparatus may include components for outputting a third indication of a determined third set of predicted eye gazes by outputting the third indication to a renderer. The apparatus may include components for rendering a fourth set of frames via the renderer based on the determined third set of predicted eye gazes; and for sending the rendered fourth set of frames to the user equipment. The apparatus may include components for rendering the fourth set of frames based on the determined third set of predicted eye gazes by rendering a first region and a second region of each frame in the fourth set of frames.The first region may be associated with a first quality level, and the second region may be associated with a second quality level. The first quality level may be greater than the second quality level. The first region may correspond to at least one predicted eye gaze in the determined third set of predicted eye gazes. The apparatus may include components for transmitting the rendered fourth set of frames to the user equipment by: (a) encoding the rendered fourth set of frames, and (b) transmitting the encoded rendered fourth set of frames.
[0199] It should be understood that the specific order or hierarchy of boxes / steps in the processes, flowcharts, and / or call flowcharts disclosed herein are merely illustrative of example methods. It should be understood that the specific order or hierarchy of boxes / steps in these processes, flowcharts, and / or call flowcharts may be rearranged based on design preferences. Furthermore, some boxes / steps may be combined and / or omitted. Other boxes / steps may also be added. The appended method claims provide the elements of various boxes / steps in an exemplary order, but are not intended to limit one to the given specific order or hierarchy.
[0200] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but should be given the full scope consistent with the language of the claims, wherein, unless specifically stated otherwise, references to elements in the singular are not intended to mean “one and only one,” but rather “one or more.” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.
[0201] Unless otherwise specified, the term "some" refers to one or more, and unless otherwise specified in the context, the term "or" may be interpreted as "and / or". Combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" include any combination of A, B, and / or C, and may include multiple A, multiple B, or multiple C. Specifically, combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" may be only A, only B, only C, A and B, A and C, B and C, or A and B and C, wherein any such combination may include one or more members of A, B, or C. The various aspects described throughout this disclosure are all structural and functional equivalents known now or hereafter to those skilled in the art, and are expressly incorporated herein by reference and intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be offered to the public, whether or not such disclosure is explicitly recited in the claims. The terms “module,” “mechanism,” “element,” “device,” etc., cannot replace the word “component.” Therefore, no claim element will be interpreted as a functional component unless the element is explicitly described using the phrase “component for…”. Unless otherwise stated, the phrase “processor” may refer to “any processor in one or more processors” (e.g., one processor in one or more processors, more than one processor in one or more processors, or all processors in one or more processors), and the phrase “memory” may refer to “any memory in one or more memories” (e.g., one memory in one or more memories, more than one memory in one or more memories, or all memories in one or more memories).
[0202] In one or more examples, the functionality described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term "processing unit" is used throughout this disclosure, such a processing unit may be implemented in hardware, software, firmware, or any combination thereof. If any functionality, processing unit, technique, or other module described herein is implemented in software, then such functionality, processing unit, technique, or other module may be stored on or transmitted on a computer-readable medium as one or more instructions or code.
[0203] Computer-readable media may include computer data storage media and communication media, including any media that facilitates the transfer of computer programs from one place to another. In this way, computer-readable media may generally correspond to: (1) a tangible computer-readable storage medium that is non-transitory; or (2) a communication medium, such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to extract instructions, code, and / or data structures for implementing the techniques described in this disclosure. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, compressed optical disc read-only memory (CD-ROM) or other optical disc storage devices, magnetic disk storage devices, or other magnetic storage devices. As used herein, magnetic disks and optical discs include compressed optical discs (CD), laser optical discs, optical discs, digital versatile optical discs (DVD), floppy disks, and Blu-ray discs, wherein magnetic disks typically magnetically copy data, while optical discs optically copy data using lasers. Combinations of the above should also be included within the scope of computer-readable media. Computer program products may include computer-readable media.
[0204] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in any hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware. Therefore, the term "processor" as used herein can refer to any of the above-described structures or any other structure suitable for implementing the techniques described herein. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0205] The following aspects are merely illustrative and may be combined with other aspects or teachings described herein without limitation.
[0206] Aspect 1 is a graphics processing method, the method comprising: providing at a first time instance a first indication of a set of head poses of a user’s head and a second indication of a set of eye gazes of the user’s eyes; obtaining a render frame based on the first indication and the second indication, the render frame being based on a predicted eye gaze of the user at a second time instance occurring after the first time instance; and; and outputting an indication of the render frame.
[0207] Aspect 2 is the method according to aspect 1, wherein providing the first indication and the second indication includes sending the first indication and the second indication to a server at the first time instance, and wherein obtaining the rendering frame includes receiving the rendering frame from the server.
[0208] Aspect 3 is the method according to aspect 2, wherein the graphics processing is performed on a wireless communication device including at least one of a transceiver or an antenna, wherein receiving the rendering frame includes receiving the rendering frame via at least one of the transceiver or the antenna.
[0209] Aspect 4 is the method according to any one of aspects 1 to 3, wherein outputting the instruction for the rendered frame comprises: (1) storing the rendered frame in at least one of a memory, a buffer, or a cache, or (2) sending the rendered frame for display.
[0210] Aspect 5 is the method according to any one of Aspects 1 to 4, wherein the set of head poses includes head pose time series data, and wherein the set of eye gazes includes eye gaze time series data.
[0211] Aspect 6 is a method according to any one of aspects 1 to 5, the method further comprising: generating the set of head poses via an inertial measurement unit (IMU); and generating the set of eye gazes via a camera.
[0212] Aspect 7 is the method according to aspect 6, wherein generating the set of head poses includes generating the set of head poses at a first sampling rate, wherein generating the set of eye gazes includes generating the set of eye gazes at a second sampling rate, wherein the first sampling rate is less than the second sampling rate.
[0213] Aspect 8 is a method according to any one of Aspects 1 to 7, wherein the rendering frame is encoded, the method further comprising: decoding the rendering frame; and warping the decoded rendering frame based on the user's latest available head pose, wherein outputting the indication to the rendering frame includes outputting an indication to the decoded warped rendering frame.
[0214] Aspect 9 is a method according to any one of aspects 1 to 8, wherein the set of head postures includes first translation information and first orientation information of the user's head, and wherein the set of eye gazes includes second translation information and second orientation information of the user's eyes.
[0215] Aspect 10 is a method according to any one of aspects 1 to 9, wherein the rendered frame includes a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to the predicted eye gaze of the user at the second time instance.
[0216] Aspect 11 is a method according to any one of aspects 1 to 10, wherein the set of head postures corresponds to a first time instance that occurs before or concurrently with the first time instance, and wherein the set of eye gazes corresponds to a second time instance that occurs before or concurrently with the first time instance.
[0217] Aspect 12 is a method according to any one of aspects 1 to 11, wherein the user's set of eye gazes corresponds to the user's eye gaze on the display, and wherein the user's predicted eye gaze corresponds to the user's predicted eye gaze on the display.
[0218] Aspect 13 is a method according to any one of Aspects 1 to 12, wherein the rendering frame is further based on Extended Reality (XR) metadata, the Extended Reality (XR) metadata including at least one of depth information associated with an object or motion vectors associated with the object, and wherein the object is a virtual object or a real-world object.
[0219] Aspect 14 is a method for graphics processing, the method comprising: obtaining a first indication of a set of head poses of a user's head and a second indication of a set of eye gazes of the user's eyes, wherein the first indication and the second indication correspond to a first time instance; generating a predicted eye gaze of the user based on the first indication and the second indication, wherein the predicted eye gaze of the user corresponds to a second time instance occurring after the first time instance; rendering a frame based on the predicted eye gaze of the user; and outputting the rendered frame.
[0220] Aspect 15 is the method according to aspect 14, wherein obtaining the first indication and the second indication includes: receiving the first indication and the second indication from a client device, and wherein outputting the rendered frame includes sending the rendered frame to the client device.
[0221] Aspect 16 is the method according to aspect 15, wherein the method is performed by a wireless communication device, the wireless communication device including at least one of a transceiver or an antenna, wherein transmitting the rendered frame includes transmitting the rendered frame via at least one of the transceiver or the antenna.
[0222] Aspect 17 is a method according to any one of aspects 14 to 16, wherein the set of head poses includes head pose time series data, and wherein the set of eye gazes includes eye gaze time series data.
[0223] Aspect 18 is a method according to any one of aspects 14 to 17, wherein a set of head poses is associated with a first sampling rate and a set of eye gazes is associated with a second sampling rate, and wherein the first sampling rate is less than the second sampling rate.
[0224] Aspect 19 is a method according to any one of aspects 14 to 18, the method further comprising: encoding a rendered frame, wherein outputting the rendered frame includes sending the encoded rendered frame.
[0225] Aspect 20 is a method according to any one of aspects 14 to 19, wherein the set of head postures includes first translation information and first orientation information of the user's head, and wherein the set of eye gazes includes second translation information and second orientation information of the user's eyes.
[0226] Aspect 21 is a method according to any one of aspects 14 to 20, wherein the rendered frame includes a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to the predicted eye gaze of the user at the second time instance.
[0227] Aspect 22 is a method according to any one of aspects 14 to 21, wherein the set of head postures corresponds to a first time instance that occurs before or concurrently with the first time instance, and wherein the set of eye gazes corresponds to a second time instance that occurs before or concurrently with the first time instance.
[0228] Aspect 23 is a method according to any one of aspects 14 to 22, wherein the user's set of eye gazes corresponds to the user's eye gazes on the display, and wherein the user's predicted eye gazes correspond to the user's predicted eye gazes on the display.
[0229] Aspect 24 is a method according to any one of aspects 14 to 23, the method further comprising: obtaining extended reality (XR) metadata, wherein generating the user's predicted eye gaze includes further generating the user's predicted eye gaze based on the XR metadata.
[0230] Aspect 25 is the method according to aspect 24, wherein the XR metadata includes at least one of depth information associated with an object or motion vectors associated with the object.
[0231] Aspect 26 is the method according to aspect 25, wherein the object is a virtual object or a real-world object.
[0232] Aspect 27 is a method according to any one of aspects 14 to 26, wherein generating the predicted eye gaze of the user includes generating the predicted eye gaze of the user via a predictor model.
[0233] Aspect 28 is the method according to aspect 27, wherein the predictor model includes at least one of a regression model, a time series predictor model, or a neural network.
[0234] Aspect 29 is a method for graphics processing, the method comprising: obtaining a first indication of a first set of eye gazes associated with a user device; obtaining a second indication of a second set of head poses associated with the user device; determining a third set of predicted eye gazes based on the obtained first set of eye gazes and a second subgroup of the obtained second set of head poses associated with a first subgroup of the obtained first set of eye gazes; and outputting a third indication of the determined third set of predicted eye gazes.
[0235] Aspect 30 is the method according to aspect 29, the method further comprising: associating a second subgroup of the obtained second group of head postures with a first subgroup of the obtained first group of eye gazes.
[0236] Aspect 31 is the method according to aspect 30, wherein associating the second subgroup of the obtained second group of head poses with the first subgroup of the obtained first group of eye gazes includes: associating the second subgroup of the obtained second group of head poses with the first subgroup of the obtained first group of eye gazes based on a first set of time indicators associated with the first group of eye gazes and a second set of time indicators associated with the second group of head poses.
[0237] Aspect 32 is the method according to any one of aspects 29 to 31, wherein the first set of eye gazes includes eye gaze time series data, and wherein the second set of head postures includes head posture time series data.
[0238] Aspect 33 is a method according to any one of aspects 29 to 32, the method further comprising: obtaining a fourth indication of a fourth group of blinks associated with the user device; and selecting a third subgroup of a first group of eye gazes based on the fourth group of blinks associated with the user device, wherein determining the third group of predicted eye gazes based on the first group of eye gazes and a second subgroup of a second group of head postures associated with the first subgroup of the first group of eye gazes comprises: determining the third group of predicted eye gazes based on the selected third subgroup of the first group of eye gazes and a fourth subgroup of the second group of head postures associated with the selected third subgroup of the first group of eye gazes.
[0239] Aspect 34 is a method according to any one of aspects 29 to 33, the method further comprising: determining a third subgroup of the obtained first group of eye gazes associated with saccades; and selecting a fourth subgroup of the obtained first group of eye gazes based on the determined third subgroup of the obtained first group of eye gazes associated with the saccades, wherein determining the third group of predicted eye gazes based on the obtained first group of eye gazes and a second subgroup of the obtained second group of head postures associated with the first subgroup of the obtained first group of eye gazes comprises: determining the third group of predicted eye gazes based on the selected fourth subgroup of the obtained first group of eye gazes and a fifth subgroup of the obtained second group of head postures associated with the selected fourth subgroup of the obtained first group of eye gazes.
[0240] Aspect 35 is a method according to any one of aspects 29 to 34, the method further comprising: converting a third subgroup of the first group of eye gazes into a fourth group of six-dimensional (6D) representations of the first group of eye gazes, wherein determining the third group of predicted eye gazes based on the obtained first group of eye gazes and a second subgroup of the obtained second group of head poses associated with the first subgroup of the obtained first group of eye gazes comprises: determining the third group of predicted eye gazes based on the converted fourth group of 6D representations of the first group of eye gazes.
[0241] Aspect 36 is a method according to any one of aspects 29 to 35, wherein determining the third set of predicted eye gazes based on a first set of obtained eye gazes and a second subgroup of obtained second set of head poses associated with the first subgroup of the first set of obtained eye gazes comprises: determining the third set of predicted eye gazes based on a first set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks.
[0242] Aspect 37 is a method according to any one of aspects 29 to 36, wherein determining the third set of predicted eye gazes based on a first set of obtained eye gazes and a second subgroup of obtained second set of head poses associated with the first subgroup of obtained first set of eye gazes comprises: determining the third set of predicted eye gazes based on a predictor model, the predictor model using the first set of obtained eye gazes and the second subgroup of obtained second set of head poses associated with the first subgroup of obtained first set of eye gazes as input.
[0243] Aspect 38 is the method according to aspect 37, wherein the predictor model includes at least one of a regression model, a time series predictor model, or a neural network.
[0244] Aspect 39 is a method according to any one of aspects 29 to 38, the method further comprising: determining a third subgroup of the first group of eye gazes associated with a user tracking a virtual object; determining a fourth group of predicted eye gazes based on the determined third subgroup of the first group of eye gazes relative to the user's head; and outputting a fourth indication of the determined fourth group of predicted eye gazes.
[0245] Aspect 40 is the method according to aspect 39, the method further comprising: determining that a fourth subgroup of the first group of eye gazes is not associated with the user tracking the virtual object, wherein determining the third group of predicted eye gazes based on the obtained first group of eye gazes and a second subgroup of the obtained second group of head poses associated with the first subgroup of the obtained first group of eye gazes comprises: determining the third group of predicted eye gazes based on the determined fourth subgroup of the first group of eye gazes and a fifth subgroup of the obtained second group of head poses associated with the determined fourth subgroup of the first group of eye gazes.
[0246] Aspect 41 is a method according to any one of aspects 29 to 40, wherein obtaining the first set of eye gazes comprises: generating the first set of eye gazes at a first sampling rate via an eye gaze camera associated with the user equipment, wherein obtaining the second set of head poses comprises: generating the second set of head poses at a second sampling rate via an inertial measurement unit (IMU) associated with the user equipment.
[0247] Aspect 42 is the method according to aspect 41, wherein the first sampling rate is less than the second sampling rate.
[0248] Aspect 43 is the method according to any one of aspects 29 to 42, wherein the first set of eye gazes includes a third set of eye gaze time series data, and wherein the second set of head postures includes a fourth set of head posture time series data.
[0249] Aspect 44 is a method according to any one of aspects 29 to 43, wherein the first set of eye gazes includes first translation information and first orientation information of at least one eye of a user associated with the user device, and wherein the second set of head postures includes second translation information and second orientation information of the user's head associated with the user device.
[0250] Aspect 45 is a method according to any one of aspects 29 to 44, wherein outputting the third indication for the determined third set of predicted eye gazes comprises: sending the third indication for the determined third set of predicted eye gazes; receiving a fourth indication for a fourth set of rendered frames based on sending the third indication for the determined third set of predicted eye gazes; and outputting a fifth indication for the fourth set of rendered frames.
[0251] Aspect 46 is the method according to aspect 45, wherein outputting the fifth instruction for the fourth set of rendering frames includes storing the fourth set of rendering frames in at least one of a memory, a buffer, or a cache.
[0252] Aspect 47 is the method according to aspect 45, the method further comprising: decoding the fourth set of rendered frames; and warping the decoded fourth set of rendered frames based on the latest available head pose of the user associated with the user equipment, wherein outputting the fifth indication for the fourth set of rendered frames comprises: outputting a sixth indication for the fifth set of warped decoded rendered frames.
[0253] Aspect 48 is the method according to aspect 45, wherein each of the fourth set of rendering frames includes a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the third set of predicted eye gazes.
[0254] Aspect 49 is a method according to any one of aspects 29 to 48, wherein obtaining the first indication of eye gaze associated with the first set of eyes associated with the user equipment comprises: receiving the first indication of eye gaze associated with the first set of eyes from the user equipment, wherein obtaining the second indication of head posture associated with the second set of head postures associated with the user equipment comprises: receiving the second indication of head posture associated with the second set of head postures from the user equipment.
[0255] Aspect 50 is the method according to aspect 49, wherein outputting the third indication for the determined third set of predicted eye gazes comprises: outputting the third indication to a renderer, the method further comprising: rendering a fourth set of frames via the renderer based on the determined third set of predicted eye gazes; and sending the rendered fourth set of frames to the user equipment.
[0256] Aspect 51 is the method according to aspect 50, wherein rendering the fourth set of frames based on the determined third set of predicted eye gazes comprises: rendering a first region and a second region of each frame in the fourth set of frames, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the determined third set of predicted eye gazes.
[0257] Aspect 52 is the method according to aspect 50, wherein sending the rendered fourth group of frames to the user equipment includes: encoding the rendered fourth group of frames; and sending the encoded rendered fourth group of frames.
[0258] Aspect 53 is a method for graphics processing, the method comprising: obtaining an indication of an eye gaze associated with a user device; obtaining an indication of a head pose associated with the user device; determining a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose; and outputting an indication of the determined set of predicted eye gazes.
[0259] Aspect 54 is the method according to aspect 53, the method further comprising: associating the head posture with the eye gaze.
[0260] Aspect 55 is the method according to aspect 54, wherein associating the head posture with the eye gaze includes: associating the head posture with the eye gaze based on a first time indicator associated with the eye gaze and a second time indicator associated with the head posture.
[0261] Aspect 56 is the method according to any one of aspects 53 to 55, wherein the indication of eye gaze includes eye gaze time series data, and wherein the indication of head posture includes head posture time series data.
[0262] Aspect 57 is a method according to any one of aspects 53 to 56, the method further comprising: obtaining an indication of a second eye gaze associated with the user device; obtaining an indication of a set of blinks associated with the user device; and determining that the second eye gaze is associated with at least one blink in the set of blinks, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture comprises: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one blink in the set of blinks.
[0263] Aspect 58 is a method according to any one of aspects 53 to 57, the method further comprising: obtaining an indication of a second eye gaze associated with the user equipment; obtaining an indication of a set of saccades associated with the user equipment; and determining that the second eye gaze is associated with at least one saccade in the set of saccades, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture comprises: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one saccade in the set of saccades.
[0264] Aspect 59 is a method according to any one of aspects 53 to 58, the method further comprising: converting the eye gaze into a six-dimensional (6D) representation of the eye gaze, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture comprises: determining the set of predicted eye gazes based on the converted 6D representation of the eye gaze.
[0265] Aspect 60 is a method according to any one of aspects 53 to 59, wherein determining the set of predicted eye gazes based on the obtained indications of the eye gaze and the obtained indications of the head posture comprises: determining the set of predicted eye gazes based on a first set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks.
[0266] Aspect 61 is a method according to any one of aspects 53 to 59, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head posture comprises: determining the set of predicted eye gazes based on a predictor model using the eye gaze and the head posture as input.
[0267] Aspect 62 is the method according to aspect 61, wherein the predictor model includes at least one of a regression model, a time series predictor model, or a neural network.
[0268] Aspect 63 is a method according to any one of aspects 53 to 62, the method further comprising: obtaining an indication of a second eye gaze associated with the user device; determining that the second eye gaze is associated with a user tracking a virtual object; determining a second set of predicted eye gazes based on the obtained indication of the second eye gaze relative to the user's head in response to determining that the eye gaze is associated with the user tracking the virtual object; and outputting an indication of the determined second set of predicted eye gazes.
[0269] Aspect 64 is the method according to aspect 63, the method further comprising: determining that the eye gaze is not associated with the user tracking the virtual object, wherein determining the set of predicted eye gazes based on obtained indications of the eye gaze and obtained indications of the head posture comprises: determining the set of predicted eye gazes based on obtained indications of the eye gaze and obtained indications of the head posture in response to determining that the eye gaze is not associated with the user tracking the virtual object.
[0270] Aspect 65 is a method according to any one of aspects 53 to 64, wherein obtaining the indication of the eye gaze associated with the user equipment comprises: generating a set of eye gazes including the eye gaze at a first sampling rate via an eye gaze camera associated with the user equipment, wherein obtaining the head pose comprises: generating a set of head poses including the head pose via an inertial measurement unit (IMU) associated with the user equipment at a second sampling rate.
[0271] Aspect 66 is the method according to aspect 65, wherein the first sampling rate is less than the second sampling rate.
[0272] Aspect 67 is a method according to any one of aspects 53 to 66, wherein the indication of eye gaze includes first translation information and first orientation information of at least one eye of the user associated with the user device, and wherein the indication of head posture includes second translation information and second orientation information of the user's head associated with the user device.
[0273] Aspect 68 is a method according to any one of aspects 53 to 67, wherein outputting the indication for a determined set of predicted eye gazes comprises: sending the indication for the determined set of predicted eye gazes; receiving an indication for a set of rendered frames after sending the indication for the determined set of predicted eye gazes; and outputting the indication for the set of rendered frames.
[0274] Aspect 69 is the method according to aspect 68, wherein outputting the instruction for the set of rendering frames includes storing the set of rendering frames in at least one of a memory, a buffer, or a cache.
[0275] Aspect 70 is the method according to aspect 68, the method further comprising: decoding the set of rendered frames; and warping the decoded set of rendered frames based on the latest available head pose of a user associated with the user device, wherein outputting the indication for the set of rendered frames comprises: outputting an indication for the set of warped decoded rendered frames.
[0276] Aspect 71 is the method according to aspect 68, wherein each of the set of rendered frames includes a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the set of predicted eye gazes.
[0277] Aspect 72 is a method according to any one of aspects 53 to 71, wherein obtaining the indication of eye gaze associated with the user equipment comprises: receiving the indication of eye gaze from the user equipment, wherein obtaining the indication of head posture associated with the user equipment comprises: receiving the indication of head posture from the user equipment.
[0278] Aspect 73 is the method according to aspect 72, wherein outputting the indication for the determined set of predicted eye gazes comprises: outputting the indication to a renderer, wherein the method further comprises: rendering a set of frames via the renderer based on the determined set of predicted eye gazes; and sending the rendered set of frames to the user device.
[0279] Aspect 74 is the method according to aspect 73, wherein rendering the set of frames based on a determined set of predicted eye gazes comprises: rendering a first region and a second region of each frame in the set of frames, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze in the determined set of predicted eye gazes.
[0280] Aspect 75 is the method according to aspect 74, wherein sending a rendered set of frames to the user equipment includes: encoding the rendered set of frames; and sending the encoded rendered set of frames.
[0281] Aspect 76 is an apparatus for graphics processing, the apparatus including at least one processor coupled to a memory and configured to implement the method according to any one of aspects 1 to 75.
[0282] Aspect 77 can be combined with aspect 76, and includes the device as a wireless communication device.
[0283] Aspect 78 is an apparatus for graphics processing, the apparatus comprising components for implementing the method according to any one of aspects 1 to 75.
[0284] Aspect 79 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer-executable code that, when executed by at least one processor, causes the at least one processor to implement the method according to any one of aspects 1 to 75.
[0285] Various aspects have been described herein. These and other aspects are within the scope of the following claims.
Claims
1. An apparatus for graphics processing, the apparatus comprising: a memory; and a processor coupled to the memory and configured to, based at least in part on information stored in the memory: obtain an indication of eye gaze associated with a user device; obtain an indication of head pose associated with the user device; determine a set of predicted eye gazes based on the obtained indication of eye gaze and the obtained indication of head pose; and output an indication of the determined set of predicted eye gazes.
2. The apparatus of claim 1, wherein the processor is further configured to: correlate the head pose with the eye gaze.
3. The apparatus of claim 2, wherein to correlate the head pose with the eye gaze, the processor is configured to: correlate the head pose with the eye gaze based on a first time indicator associated with the eye gaze and a second time indicator associated with the head pose.
4. The apparatus of claim 1, wherein the indication of eye gaze comprises eye gaze time series data, wherein the indication of head pose comprises head pose time series data.
5. The apparatus of claim 1, wherein the processor is further configured to: obtain an indication of a second eye gaze associated with the user device; obtain an indication of a set of blinks associated with the user device; and to determine the set of predicted eye gazes based on the obtained indication of eye gaze and the obtained indication of head pose, the processor is configured to: determining that the second eye gaze is associated with at least one blink of the set of blinks, wherein, avoid determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one blink in the set of blinks.
6. The apparatus of claim 1, wherein the processor is further configured to: obtain an indication of a second eye gaze associated with the user device; obtain an indication of a set of saccades associated with the user device; and to determine the set of predicted eye gazes based on the obtained indication of eye gaze and the obtained indication of head pose, the processor is configured to: determining that the second eye gaze is associated with at least one saccade in the set of saccades, wherein, avoid determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one saccade in the set of saccades.
7. The apparatus of claim 1, wherein the processor is further configured to: to determine the set of predicted eye gazes based on the obtained indication of eye gaze and the obtained indication of head pose, the processor is configured to: converting the eye gaze into a six-dimensional (6D) representation of the eye gaze, wherein, determine the set of predicted eye gazes based on a transformed 6D representation of the eye gaze. to determine the set of predicted eye gazes based on the obtained indication of eye gaze and the obtained indication of head pose, the processor is configured to:
8. The apparatus of claim 1, wherein, determining the set of predicted eye gazes based on a first set of fully connected network (FCN) blocks and a set of gated recurrent unit (GRU) blocks.
9. The apparatus of claim 1, wherein, To determine the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose, the processor is configured to: determine the set of predicted eye gazes based on a predictor model that uses the eye gaze and the head pose as inputs.
10. The apparatus of claim 9, wherein the predictor model comprises at least one of a regression model, a time series predictor model, or a neural network.
11. The apparatus of claim 1, wherein the processor is further configured to: obtain an indication of a second eye gaze associated with the user device; determine that the second eye gaze is associated with a user tracking a virtual object; determine a second set of predicted eye gazes based on the obtained indication of the second eye gaze relative to a head of the user in response to determining that the eye gaze is associated with the user tracking the virtual object; and output an indication of the determined second set of predicted eye gazes.
12. The apparatus of claim 11, wherein the processor is further configured to: To determine the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose, the processor is configured to: determining that the eye gaze is not associated with the user tracking the virtual object, wherein, determine the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose in response to determining that the eye gaze is not associated with the user tracking the virtual object. To obtain the indication of the eye gaze associated with the user device, the processor is configured to:
13. The apparatus of claim 1, wherein, generate a set of eye gazes including the eye gaze at a first sampling rate via an eye gaze camera associated with the user device, wherein, to obtain the head pose, the processor is configured to: generate a set of head poses including the head pose at a second sampling rate via an inertial measurement unit (IMU) associated with the user device.
14. The apparatus of claim 13, wherein the first sampling rate is less than the second sampling rate.
15. The apparatus of claim 1, wherein the indication of the eye gaze comprises first translational information and first orientation information of at least one eye of a user associated with the user device, wherein the indication of the head pose comprises second translational information and second orientation information of a head of the user associated with the user device. To output the indication of the determined set of predicted eye gazes, the processor is configured to:
16. The apparatus of claim 1, wherein, send the indication of the determined set of predicted eye gazes; receive an indication of a set of rendered frames after sending the indication of the determined set of predicted eye gazes; and output an indication of the set of rendered frames. To send the indication of the determined set of predicted eye gazes, the processor is configured to:
17. The apparatus of claim 16, wherein the apparatus is a wireless communication device including at least one of a transceiver or an antenna coupled to the processor, wherein, transmit, via at least one of the transceiver or the antenna, the indication of the determined set of predicted eye gazes, wherein, to receive the indication of the set of rendered frames after transmitting the indication of the determined set of predicted eye gazes, the processor is configured to: receive, via at least one of the transceiver or the antenna, the indication of the set of rendered frames.
18. The apparatus of claim 16, wherein, to output the indication of the set of rendered frames, the processor is configured to: store the set of rendered frames in at least one of a second memory, buffer, or cache.
19. The apparatus of claim 16, wherein the processor is further configured to: decode the set of rendered frames; and warping the decoded set of rendered frames based on a latest available head pose of a user associated with the user device, wherein, to output the indication of the set of rendered frames, the processor is configured to: output an indication of a set of warped decoded rendered frames.
20. The apparatus of claim 16, wherein each rendered frame of the set of rendered frames comprises a first region and a second region, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze of the set of predicted eye gazes.
21. The apparatus of claim 1, wherein, to obtain the indication of the eye gaze associated with the user device, the processor is configured to: receive, from the user device, the indication of the eye gaze, wherein, to obtain the indication of the head pose associated with the user device, the processor is configured to: receive, from the user device, the indication of the head pose.
22. The apparatus of claim 21, wherein, to output the indication of the determined set of predicted eye gazes, the processor is configured to: output the indication to a Tenderer, wherein the processor is further configured to: render, via the Tenderer, a set of frames based on the determined set of predicted eye gazes; and transmit, to the user device, the rendered set of frames.
23. The apparatus of claim 22, wherein the apparatus is a wireless communication device comprising at least one of a transceiver or an antenna coupled to the processor, wherein, to transmit, to the user device, the rendered set of frames, the processor is configured to: transmit, via at least one of the transceiver or the antenna, the rendered set of frames to the user device.
24. The apparatus of claim 22, wherein, to render the set of frames based on the determined set of predicted eye gazes, the processor is configured to: render a first region and a second region of each frame of the set of frames, wherein the first region is associated with a first quality level and the second region is associated with a second quality level, wherein the first quality level is greater than the second quality level, and wherein the first region corresponds to at least one predicted eye gaze of the determined set of predicted eye gazes.
25. The apparatus of claim 24, wherein, to transmit, to the user device, the rendered set of frames, the processor is configured to: encode the rendered set of frames; and transmit the encoded rendered set of frames.
26. The apparatus of claim 1, wherein the apparatus comprises a wireless communication device.
27. A method for graphics processing, the method comprising: obtaining an indication of an eye gaze associated with the user device; obtaining an indication of a head pose associated with the user device; determining a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose; and outputting an indication of the determined set of predicted eye gazes.
28. The method of claim 27, the method further comprising: obtaining an indication of a second eye gaze associated with the user device; obtaining an indication of a set of blinks associated with the user device; and determining that the second eye gaze is associated with at least one blink of the set of blinks, wherein determining the set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose comprises: avoiding determining the set of predicted eye gazes based on the second eye gaze in response to determining that the second eye gaze is associated with at least one blink of the set of blinks.
29. The method of claim 27, the method further comprising: obtaining an indication of a second eye gaze associated with the user device; determining that the second eye gaze is associated with a user tracking a virtual object; determining a second set of predicted eye gazes based on the obtained indication of the second eye gaze relative to a head of the user in response to determining that the eye gaze is associated with the user tracking the virtual object; and outputting an indication of the determined second set of predicted eye gazes.
30. A computer-readable medium storing computer executable code that, when executed by a processor, causes the processor to: obtain an indication of an eye gaze associated with the user device; obtain an indication of a head pose associated with the user device; determine a set of predicted eye gazes based on the obtained indication of the eye gaze and the obtained indication of the head pose; and output an indication of the determined set of predicted eye gazes.