Method for rendering video frames, computer-readable storage medium, and computer system
By predicting the user's scandal landing point and performing post-update updates in the GPU buffer, the rendering pipeline is optimized, and the problem of image blur in virtual reality scenes is solved, improving image clarity and user experience, while reducing power consumption.
Patent Information
- Application Number
- CN202210589739.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-17
- Filing Date
- 2019-05-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2039-05-07
AI Technical Summary
In virtual reality scenarios, the user's eye movement speed is faster than the update speed of traditional rendering pipelines, resulting in blurred images, especially when the user's gaze moves from one area to another, the prior art cannot update the high-resolution area in time, affecting the user experience.
By predicting the user's sacrificial landing point and performing post-update updates in the GPU-accessible buffer, the rendering pipeline is optimized to ensure that the fovea area is consistent with the user's line of sight direction, achieving high-resolution rendering.
It improves the image clarity of virtual reality scenes, reduces the risk of motion sickness for users, improves user experience, and reduces power consumption.
Smart Images

Figure CN115131483B_ABST
Abstract
Description
[0001] Inventors: Andrew Young and Javier Fernandez Rico
[0002] This divisional application is a divisional application of the application with the application date of May 7, 2019, the application number of 201980046421.9, and the invention title of "Eye Tracking for Fast Foveal Rendering in an HMD Environment Using Prediction and Late Updates to the GPU". Technical Field
[0003] The present disclosure relates to computer-generated images, and more particularly to the real-time rendering of computer-generated graphics. Background Art
[0004] Computer rendering of virtual reality (VR) scenes in a rendering pipeline requires central processing unit (CPU) and graphics processing unit (GPU) resources. Although only a small portion of the wide viewing range is displayed, the VR scene can be rendered within that viewing range. Additionally, VR scenes may be more complex than traditional scenes and may also require a higher frame rate for image processing to avoid motion sickness, all of which result in a high power consumption rate.
[0005] To save power, portions of the display can be rendered at a higher resolution than other portions. For example, the portion of the screen that a user is fixating on can be rendered at a higher resolution than other portions that the user is not fixating on (such as those in the periphery). Rendering at a lower resolution portion of the display in the periphery can save processing resources, and since the user does not focus on the periphery, the low resolution does not degrade the user's viewing experience. However, the movement of the eyes of a user viewing a VR scene may be faster than the update of frames through the rendering pipeline. Thus, because the eyes are faster than the computer rendering pipeline, when the user moves to a portion that may have been in the periphery of the scene before, that portion may still be rendered at a low resolution until the update catches up with the eye movement. This results in a blurry image for the user.
[0006] It is in this context that the embodiments of the present disclosure have emerged. Summary of the Invention
[0007] Embodiments of the present disclosure relate to predicting the landing point of a saccade associated with a user viewing a display of a head-mounted display (HMD), and updating information in a rendering pipeline including a central processing unit (CPU) and a graphics processing unit (GPU) by late-updating the predicted landing point to a buffer accessible by the GPU for immediate use. Several inventive embodiments of the present disclosure are described below.
[0008] In one embodiment, a method for predicting eye movements in an HMD is disclosed. The method includes tracking the movement of a user's eyes at a plurality of sample points using a gaze tracking system disposed in the HMD. The method includes determining a speed of the movement based on the movement of the eyes. The method includes determining that the user's eyes are in a saccade state when the speed reaches a threshold speed. The method includes predicting a landing point corresponding to the saccade eye direction on a display of the HMD.
[0009] In one embodiment, a method for updating information of a rendering pipeline including a CPU and a GPU is disclosed. The method includes executing an application on the CPU in a first frame period to generate primitives of a scene of a first video frame. A frame period corresponds to an operating frequency of the rendering pipeline, and the rendering pipeline is configured to perform sequential operations by the CPU and the GPU in consecutive frame periods before scanning and outputting a corresponding video frame to a display. The method includes receiving, at the CPU in a second frame period, gaze tracking information of a user's eyes experiencing a saccade. The method includes predicting, at the CPU in the second frame period, a landing point corresponding to the saccade eye gaze direction on a display of a head-mounted display (HMD) at least based on the gaze tracking information. The method includes performing a late update operation by the CPU in the second frame period by transmitting the predicted landing point to a buffer accessible to the GPU. The method includes performing one or more shader operations in the GPU in the second frame period to generate pixel data of pixels of the HMD based on the primitives of the scene of the first video frame and based on the predicted landing point, wherein the pixel data includes at least color and texture information, and wherein the pixel data is stored in a frame buffer. The method includes scanning and outputting the pixel data from the frame buffer to the HMD in a third frame period.
[0010] In another embodiment, a non - transitory computer - readable medium is disclosed that stores a computer program for updating information for a rendering pipeline including a CPU and a GPU. The computer - readable medium includes program instructions for executing an application on the CPU in a first frame period to generate primitives of a scene of a first video frame. The frame period corresponds to an operating frequency of the rendering pipeline, which is configured to perform sequential operations by the CPU and the GPU in successive frame periods before scanning out a corresponding video frame to a display. The computer - readable medium includes program instructions for receiving, at the CPU in a second frame period, eye gaze - tracking information of a user's eyes undergoing a saccade. The computer - readable medium includes program instructions for predicting, at the CPU in the second frame period, a landing point corresponding to the saccade eye gaze direction on a display of a head - mounted display (HMD) at least based on the gaze - tracking information. The computer - readable medium includes program instructions for performing a late - update operation at the CPU in the second frame period by transmitting the predicted landing point to a buffer accessible by the GPU. The computer - readable medium includes program instructions for performing one or more shader operations in the GPU in the second frame period to generate pixel data of pixels of the HMD based on the primitives of the scene of the first video frame and based on the predicted landing point, where the pixel data includes at least color and texture information, and where the pixel data is stored in a frame buffer. The computer - readable medium includes program instructions for scanning out the pixel data from the frame buffer to the HMD in a third frame period.
[0011] In yet another embodiment, a computer system is disclosed having a processor and a memory coupled to the processor, the memory storing instructions which, if executed by the computer system, cause the computer system to perform a method for updating information of a rendering pipeline including a CPU and a GPU. The method includes executing an application on the CPU in a first frame period to generate primitives of a scene of a first video frame. A frame period corresponds to an operating frequency of the rendering pipeline, which is configured to perform sequential operations by the CPU and the GPU in consecutive frame periods before scanning out a corresponding video frame to a display. The method includes receiving, at the CPU in a second frame period, eye gaze tracking information of a user's eyes undergoing a saccade. The method includes predicting, at the CPU in the second frame period, a landing point corresponding to the saccade eye gaze direction on a display of a head-mounted display (HMD) based at least on the eye gaze tracking information. The method includes performing a late update operation by the CPU in the second frame period by transmitting the predicted landing point to a buffer accessible by the GPU. The method includes performing one or more shader operations in the GPU in the second frame period to generate pixel data of pixels of the HMD based on the primitives of the scene of the first video frame and based on the predicted landing point, wherein the pixel data includes at least color and texture information, and wherein the pixel data is stored in a frame buffer. The method includes scanning out the pixel data from the frame buffer to the HMD in a third frame period.
[0012] Other aspects of the present disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The present disclosure may best be understood by reference to the following description in conjunction with the accompanying drawings, in which:
[0014] Figure 1A A system configured to provide an interactive experience with VR content and to predict a landing point of a saccade associated with a user viewing a display of an HMD is shown, where some input controls may be provided by a hand-held controller and some input controls may be managed by tracking body parts (such as by a camera).
[0015] Figure 1BA system is shown that is configured to provide an interactive experience with VR content and to predict the landing points of saccades associated with a user viewing the display of an HMD, where some input controls for editing can be provided via a handheld controller and some input controls can be managed by tracking body parts (such as implemented via a camera), where the camera also tracks the movement of the HMD for the purpose of beam tracking an RF transmitter that transmits data to the HMD.
[0016] Figure 1C A system is shown that is configured to provide an interactive experience with VR content and to predict the landing points of saccades associated with a user viewing the display of an HMD, where some input controls for editing can be provided via a handheld controller and some input controls can be managed by magnetically tracking body parts (such as partially implemented via a magnetic source).
[0017] Figure 2 Conceptually shown is the functionality of an HMD in accordance with an embodiment of the present disclosure that combines with the execution of a video game and provides a 3D editing space for editing 3D digital content.
[0018] Figures 3A to 3C Shown is a view of an example display housing when looking towards the inner surface, where the face is designed to contact the display housing at the inner surface, and the view shows the interior of an HMD including an eye tracking sensor.
[0019] Figure 4A Shown is a prediction engine configured to predict the landing points of saccades associated with a user viewing the display of an HMD in accordance with an embodiment of the present disclosure.
[0020] Figure 4B Shown is a recurrent neural network for predicting the landing points of saccades associated with a user viewing the display of an HMD in accordance with an embodiment of the present disclosure.
[0021] Figure 4C Shown is an example neural network for constructing a saccade motion model for one or more users viewing a VR scene in an HMD in accordance with an embodiment of the present disclosure.
[0022] Figure 5A Shown is a rendering pipeline without saccade prediction, which shows how frame updates are slower than eye movements, such that the image is blurry to the user after the eye movement is complete.
[0023] Figure 5BShows the resulting effect of a rendering pipeline according to an embodiment of the present disclosure, the rendering pipeline being configured with saccade prediction of the eye movements of a user viewing the display of an HMD such that after the eye movement is completed, the image is made clear to the user by advancing the update of the high-resolution foveal region in the rendering pipeline.
[0024] Figure 6A Shows the eye displacement and velocity of a saccade of a user viewing the display of an HMD according to an embodiment of the present disclosure.
[0025] Figure 6B Shows the sampling of eye orientation data at various sample points in the velocity profile of a saccade of a user viewing the display of an HMD according to an embodiment of the present disclosure.
[0026] Figure 6C Shows the collection of eye orientation data for one or more sets of sample points for predicting the landing point of a saccade associated with a user viewing the display of an HMD according to an embodiment of the present disclosure.
[0027] Figure 6D Shows a table according to an embodiment of the present disclosure, the table listing the eye orientation data for a set of sample points for predicting the landing point of a saccade associated with a user viewing the display of an HMD.
[0028] Figure 6E Shows the line-of-sight direction vector for determining the velocity of a user's eye according to an embodiment of the present disclosure.
[0029] Figure 7 Is a flowchart showing the steps in a method for predicting the landing point of a saccade associated with a user viewing the display of an HMD according to an embodiment of the present disclosure, and includes the convergence of multiple predictions of the landing point of a saccade associated with a user viewing the display of an HMD using eye orientation data from a set of sample points collected during the saccade.
[0030] Figure 8 Shows a computer system implementing a rendering pipeline according to an embodiment of the present disclosure, the rendering pipeline being configured for foveal rendering, including: predicting the landing point of a saccade associated with a user viewing the display of an HMD, and providing the landing point as a late update to a buffer accessible to the GPU of the computer system for immediate use in rendering a high-resolution foveal region centered on the landing point in a corresponding video frame.
[0031] Figure 9AShows a rendering pipeline that receives and uses eye-tracking information when generating video frames during application execution according to an embodiment of the present disclosure, where the rendering pipeline does not implement late updating of eye-tracking information.
[0032] Figure 9B Is a flowchart showing steps in a method for updating information of a rendering pipeline by predicting a landing point on a display of an HMD according to an embodiment of the present disclosure, where the landing point corresponds to the orientation of the eyes of a user viewing the display during or at the end of a saccade, and where the predicted landing point is used by the GPU to render a high-resolution foveal region centered on the landing point in a corresponding video frame.
[0033] Figure 9C Shows a rendering pipeline that receives and uses eye-tracking information when generating video frames during application execution according to an embodiment of the present disclosure, where a landing point corresponding to the orientation of the eyes of a user viewing the HMD during or at the end of a saccade is predicted on the HMD, and where the predicted landing point is used by the GPU to render a high-resolution foveal region centered on the landing point in a corresponding video frame.
[0034] Figure 10A Is a flowchart showing steps in a method for updating information of a rendering pipeline by late-updating eye-tracking information to a buffer accessible by the GPU for immediate use according to an embodiment of the present disclosure.
[0035] Figure 10B Shows a rendering pipeline that receives and uses eye-tracking information when generating video frames during application execution according to an embodiment of the present disclosure, where the rendering pipeline implements late updating of eye-tracking information.
[0036] Figure 11A Is a flowchart showing steps in a method for updating information of a rendering pipeline by late-updating a predicted landing point on a display of an HMD to a buffer accessible by the GPU for immediate use according to an embodiment of the present disclosure, where the landing point corresponds to the orientation of the eyes of a user viewing the display during or at the end of a saccade.
[0037] Figure 11B Shows a rendering pipeline that receives and uses eye-tracking information when generating video frames during application execution according to an embodiment of the present disclosure, where the rendering pipeline late-updates a predicted landing point on a display of the HMD to a buffer accessible by the GPU for immediate use, where the landing point corresponds to the orientation of the eyes of a user viewing the display during or at the end of a saccade.
[0038] Figure 12FIG. 9 is a component diagram of an example device that can be used to implement aspects of various embodiments of the present disclosure. FIG. 9 is a diagram showing components of a head-mounted display according to an embodiment of the present disclosure.
[0039] Figure 13 is a diagram showing components of a head-mounted display according to an embodiment of the present disclosure.
[0040] Figure 14 is a block diagram of a gaming system according to various embodiments of the present disclosure. DETAILED DESCRIPTION
[0041] Although the following detailed description includes many specific details for purposes of illustration, those of ordinary skill in the art will appreciate that many variations and changes of the following details are within the scope of the present disclosure. Accordingly, in setting forth various aspects of the present disclosure as described below, there is no loss of generality in the claims appended hereto and no limitation is imposed on these claims.
[0042] Generally, various embodiments of the present disclosure describe systems and methods for predicting a landing point associated with the line-of-sight direction of a user's eyes during and / or at the end of a saccade defined in relation to a user viewing a display of an HMD on a display. Specifically, when the user's line of sight moves from one fixation point to another in a normal manner, a velocity map of the measured portion of the saccade that defines the user's eye movement can be used to predict the characteristics of the entire saccade. In this way, one or more eye directions can be predicted based on velocity analysis, where the eye directions correspond to one or more landing points on the display. Once the target landing points on the display are known, the rendered frames can be updated in consideration of the target landing points for display on the HMD. For example, the foveal region corresponding to the area at or around the target landing point on the display can be updated such that the movement of the eyes is consistent with the display at the target landing point. The foveal region (e.g., the position where the eyes are focused and directed) is rendered at a high resolution, and the extrafoveal region (e.g., the periphery) can be rendered at a lower resolution.
[0043] Based on the above general understanding of the various embodiments, example details of the embodiments will now be described with reference to the various figures.
[0044] Throughout the specification, references to "gaming applications" are intended to represent any type of interactive application that is directed by the execution of input commands. For purposes of illustration only, interactive applications include applications for gaming, word processing, video processing, video game processing, etc. Additionally, the terms video game and gaming application are interchangeable.
[0045] Throughout the specification, reference is made to the user's pan. Typically, pan refers to the rapid and simultaneous movement of the user's eyes when moving from a fixation point on a display to another fixation point. The panning motion of the eyes is typically performed in a specific direction, and need not be performed in a rotational manner. Panning motion may reach a peak angular velocity of more than 900 degrees per second, and lasts any time between 20 milliseconds (ms) to 200ms. The angular displacement (degrees) of the eyes during the panning may change upwards to approximately 90 degrees, but displacements exceeding 20 degrees to 50 degrees may be accompanied by head movement.
[0046] Predicting the landing point of the user's eye during a saccade
[0047] Figure 1A A system for interactive game play of a gaming application according to an embodiment of the present disclosure is shown. A user 100 is shown wearing an HMD 102, wherein the HMD 102 is worn in a manner similar to glasses, goggles, or a helmet and is configured to display a video game from an interactive gaming application or other content from an interactive application to the user 100. The HMD 102 provides a highly immersive experience to the user because it provides a display mechanism in close proximity to the user's eyes. Thus, the HMD 102 can provide a display area to each eye of the user that occupies a large portion of, or even the entire field of view of the user.
[0048] Figure 1A The system in is configured to update the target landing point on the display of HMD 102 so that the movement of the user's eyes is consistent with the presentation of the foveal area on the display at the updated target landing point. In particular, the saccade prediction of the landing point can be performed at one or more of HMD 102, computer 106 and cloud game server 114, either alone or in combination. The prediction is performed by a saccade prediction engine 300 including a deep learning engine 190, which is configured to perform one or both of the following two items: generating a saccade model based on a saccade measured for a test subject (e.g., collection of eye orientation data or parameters) by training; and comparing the eye orientation data of the user's current saccade with the trained saccade model to predict on the display a landing point associated with the user's line of sight direction during and / or at the end of the saccade.
[0049] In one embodiment, the HMD 102 can be connected to a computer or game console 106. The connection to the computer 106 can be wired or wireless. In some implementations, the HMD 102 can also communicate with the computer via an alternative mechanism or channel (such as via a network 112 to which both the HMD 102 and the computer 106 are connected). The computer 106 can be any general-purpose or special-purpose computer known in the art, including but not limited to game consoles, personal computers, laptops, tablet computers, mobile devices, cellular phones, tablets, thin clients, set-top boxes, streaming devices, etc. In one embodiment, the computer 106 can be configured to execute a game application and output video and audio from the game application for rendering by the HMD 102. The computer 106 is not limited to executing game applications, but can also be configured to execute an interactive application that outputs VR content 191 for rendering by the HMD 102. In one embodiment, the computer 106 is configured to predict a landing point associated with the line-of-sight direction of the user's eyes during and / or at the end of a saccade defined in relation to the user viewing the display on the HMD's display. In other embodiments, the prediction of the landing point can be performed by one or more of the HMD 102, the computer 106, and the cloud gaming server 114, either alone or in combination.
[0050] The user 100 can operate the controller 104 to provide input for the game application. The connection to the computer 106 can be wired or wireless. Additionally, the camera 108 can be configured to capture one or more images of the interactive environment in which the user 100 is located. The captured images can be analyzed to determine the position and movement of the user 100, various parts of the user (e.g., tracking gestures for input commands), the HMD 102, and the controller 104. In one embodiment, the controller 104 includes lights or other marker elements that can be tracked to determine its position and orientation. Additionally, the HMD 102 can include one or more lights that can be tracked to determine the position and orientation of the HMD 102. The tracking function, partially implemented by the camera 108, provides input commands generated by the movement of the controller 104 and / or the body parts (e.g., hands) of the user 100. The camera 108 can include one or more microphones to capture sound from the interactive environment. The sound captured by the microphone array can be processed to identify the location of the sound source. The sound from the identified location can be selectively utilized or processed to exclude other sounds that do not originate from the identified location. Additionally, the camera 108 can be defined to include multiple image capture devices (e.g., a stereo camera pair), IR cameras, depth cameras, and combinations thereof.
[0051] In another embodiment, computer 106 functions as a thin client in its communication with cloud gaming provider 112 over a network. Cloud gaming provider 112 maintains and executes the game application that user 102 is playing. Computer 106 transmits inputs from HMD 102, controller 104, and camera 108 to the cloud gaming provider, which processes these inputs to affect the game state of the executing game application. Outputs from the executing game application, such as video data, audio data, and haptic feedback data, are transmitted to computer 106. Computer 106 may also process the data before transmission, or may transmit the data directly to the relevant device. For example, video and audio streams are provided to HMD 102, while haptic feedback data is used to generate vibration feedback commands provided to controller 104.
[0052] In one embodiment, HMD 102, controller 104, and camera 108 themselves may be networked devices connected to network 110 to communicate with cloud gaming provider 112. For example, computer 106 may be a local network device such as a router that does not otherwise perform video game processing but facilitates the transfer of network traffic. The connections of HMD 102, controller 104, and camera (e.g., image capture device) 108 to the network may be wired or wireless.
[0053] In yet another embodiment, computer 106 may execute a portion of the game application while the remainder of the game application is executed on cloud gaming provider 112. In other embodiments, portions of the game application may also be executed on HMD 102. For example, cloud gaming provider 112 may service a request to download the game application from computer 106. When the request is serviced, cloud gaming provider 112 may execute a portion of the game application and provide game content to computer 106 for rendering on HMD 102. Computer 106 may communicate with cloud gaming provider 112 over network 110. Inputs received from HMD 102, controller 104, and camera 108 are transmitted to cloud gaming provider 112 while the game application is being downloaded to computer 106. Cloud gaming provider 112 processes the inputs to affect the game state of the executing game application. Outputs from the executing game application, such as video data, audio data, and haptic feedback data, are transmitted to computer 106 for delivery to the corresponding device.
[0054] Once the game application has been fully downloaded to the computer 106, the computer 106 can execute the game application and resume the gameplay of the game application from where it was interrupted on the cloud game provider 112. Inputs from the HMD 102, the controller 104, and the camera 108 are processed by the computer 106, and the game state of the game application is adjusted in response to the inputs received from the HMD 102, the controller 104, and the camera 108. In such an embodiment, the game state of the game application at the computer 106 is synchronized with the game state at the cloud game provider 112. Synchronization can be performed periodically to keep the state of the game application up-to-date at both the computer 106 and the cloud game provider 112. The computer 106 can transmit output data directly to the relevant devices. For example, video and audio streams are provided to the HMD 102, while haptic feedback data is used to generate vibration feedback commands provided to the controller 104.
[0055] Figure 1B Shown is a system configured to provide an interactive experience with VR content and to provide a 3D editing space for editing 3D digital content according to an embodiment of the present disclosure. Additionally, the system (e.g., the HMD 102, the computer 106, and / or the cloud 114) is configured to update a target landing point on the display of the HMD 102 such that the movement of the user's eyes is consistent with the rendering of the foveal region (high-resolution region) on the display at the updated target landing point. Figure 1B Similar to Figure 1A the system described in, where a transmitter / receiver (transceiver) 110 configured to pass data to the HMD 102, for example via an RF signal, is added. The transceiver 110 is configured to transmit video and audio from the game application to the HMD 102 for rendering thereon (either via a wired connection or a wireless connection). Additionally, the transceiver 110 is configured to transmit images, video, and audio of 3D digital content within the 3D editing space for editing purposes. In this implementation, according to an embodiment of the present disclosure, the camera 108 can be configured to track the movement of the HMD 102 such that the transceiver 110 can beam most of its RF power (as transmitted through the RF radiation pattern) towards the HMD 102 (e.g., for the purpose of transmitting data).
[0056] Figure 1C Shown is a system configured to provide an interactive experience with VR content according to an embodiment of the present disclosure. Additionally, the system (e.g., the HMD 102, the computer 106, and / or the cloud 114) is configured to update a target landing point on the display of the HMD 102 such that the movement of the user's eyes is consistent with the rendering of the foveal region (high-resolution region) on the display at the updated target landing point. Figure 1C Similar to Figure 1AThe system described is similar, with the addition of a magnetic source 116 configured to emit a magnetic field to enable magnetic tracking of the HMD 102, the controller 104 (e.g., configured as an interface controller), or any object configured with a magnetic sensor (e.g., a glove, strip located on a body part such as a finger). For example, the magnetic sensor can be an inductive element. In particular, the magnetic sensor can be configured to detect the magnetic field (e.g., intensity, orientation) emitted by the magnetic source 116. Information collected from the magnetic sensor can be used to determine and track the position and / or orientation of the HMD 102, the controller 104, and other interface objects, etc., in order to provide input commands executed within the 3D editing space. In an embodiment, the magnetic tracking is combined with tracking performed by cameras 108 and / or inertial sensors within the HMD 102, the controller 104, and / or other interface objects.
[0057] In some implementations, an interface object (e.g., the controller 104) is tracked relative to the HMD 102. For example, the HMD 102 can include an outward-facing camera that captures an image including the interface object. In other embodiments, the HMD 102 can include an IR emitter for tracking external objects such as interface objects. The captured image can be analyzed to determine the position / orientation of the interface object relative to the HMD 102, and the known position / orientation of the HMD 102 can be used to determine the position / orientation and / or movement of the interface object in the local environment.
[0058] The way the user 100 docks with the game application or the virtual reality scene of the 3D editing space displayed in the HMD 102 can vary, and other interface devices can be used in addition to the interface object (e.g., the controller 104). For example, various types of single-handed and two-handed controllers 104 can be used. In some implementations, the controller 104 itself can be tracked by tracking the lights included in the controller or by tracking the shape, sensors, and inertial data associated with the controller 104. Using these various types of controllers 104 or even just gestures made and captured by one or more cameras and magnetic sensors, one can dock, control, manipulate, interact with, and participate in the virtual reality game environment presented on the HMD 102.
[0059] Figure 2Conceptually shows the functionality of the HMD 102 in combination with the generation of VR content 291 (e.g., executing an application and / or a video game, etc.) according to an embodiment of the present disclosure, which includes updating a target landing point on the display of the HMD 102 such that the movement of the user's eyes is consistent with the rendering of the foveal region (e.g., high-resolution region) on the display at the updated target landing point. The saccade prediction of the landing point can be performed by one or more of the HMD 102, the computer 106, and the cloud gaming server 114, either alone or in combination. In an embodiment, the VR content engine 220 is executing on the HMD 102. In other embodiments, the VR content engine 220 is executing on a computer 106 (not shown) communicatively coupled to the HMD 102 and / or combined with the HMD 102. The computer can be local to the HMD (e.g., part of a local area network), or can be remote (e.g., part of a wide area network, a cloud network, etc.) and accessed via a network. The communication between the HMD 102 and the computer 106 can follow a wired or wireless connection protocol. In an example, the VR content engine 220 that executes the application can be a video game engine that executes a game application and is configured to receive inputs to update the game state of the game application. Figure 2 The following description is described in the context of the VR content engine 220 that executes the game application for the sake of brevity and clarity, and is intended to represent the execution of any application capable of generating VR content 291. The game state of the game application can be defined at least in part by the values of various parameters of the video game, and the parameters define various aspects of the current game-playing process, such as the presence and location of objects, the condition of the virtual environment, the triggering of events, the user profile, the view perspective, etc.
[0060] In the illustrated embodiment, by way of example, the VR content engine 220 receives controller input 261, audio input 262, and motion input 263. It can be from a game controller separate from the HMD 102 (such as the handheld game controller 104 (e.g., Sony 4 wireless controller, Sony Define the controller input 261 in the operation of a motion controller (e.g., a wearable controller such as a wearable glove interface controller). For example, the controller input 261 may include direction input, button presses, trigger activation, motion, gestures, or other types of input processed from the operation of a game controller. The audio input 262 may be processed from the microphone 251 of the HMD 102 or from a microphone included in the image capture device 208 or other locations within the local system environment. The motion input 263 may be processed from the motion sensors 259 included in the HMD 102 or from the image capture device when the image capture device 108 captures an image of the HMD 102. For example, in the case of executing a game application, the VR content engine 220 receives inputs processed according to the configuration of the content engine 220 used as a game engine to update the game state of a video game. The engine 220 outputs the game state data to various rendering modules, and the rendering modules process the game state data to define the content to be presented to the user.
[0061] In the illustrated embodiment, the video rendering module 283 is defined to render a video stream for presentation on the HMD 102.
[0062] The lenses of the optics 270 in the HMD 102 are configured for viewing VR content 291. The display screen 1304 is disposed behind the lenses of the optics 270 such that when the user wears the HMD 102, the lenses of the optics 270 are between the display screen 1304 and the user's eyes 260. In this way, the video stream can be presented by the display screen / projector mechanism 1304 and viewed by the user's eyes 260 through the optics 270. For example, an HMD user can choose to interact with interactive VR content 291 (e.g., a VR video source, video game content, etc.) by wearing the HMD for the purpose of editing 3D digital content in a 3D editing space. The interactive virtual reality (VR) scene from a video game can be rendered on the display screen 1304 of the HMD. In this way, during game development, the HMD 102 allows the user to edit and view the interactive VR scene. Similarly, during the game play process (including review and editing), the HMD allows the user to be fully immersed in the game play by providing the display mechanism of the HMD near the user's eyes. The display area defined in the display screen of the HMD for rendering content may occupy most or even the entire field of view of the user. Typically, each eye is supported by the associated lenses of the optics 270 that are viewing one or more display screens.
[0063] The audio rendering module 282 is configured to render an audio stream for a user to listen to. In one embodiment, the audio stream is output through a speaker 152 associated with the HMD 102. It should be understood that the speaker 152 may take the form of an open-air speaker, headphones, or any other type of speaker capable of presenting audio.
[0064] In one embodiment, a gaze tracking sensor 265 is included in the HMD 102 to enable tracking of a user's gaze. Although only one gaze tracking sensor 265 is included, it should be noted that more than one gaze tracking sensor may be employed to track the user's gaze, as will be described with respect to Figures 3A to 3C what is described. For example, in some embodiments, only one eye is tracked (e.g., using one sensor), while in other embodiments, both eyes are tracked using multiple sensors. The gaze tracking sensor 265 may be one or more of a camera, an optical sensor, an infrared sensor, an EMG (electromyogram) sensor, an optical reflector sensor, a distance sensor, and an optic flow sensor, a Doppler sensor, a microphone, etc. Generally, the sensor 265 may be configured to detect rapid eye movements, such as changes in eye movement direction, acceleration, and velocity. For example, a gaze tracking camera captures an image of the user's eye, and the image is analyzed to determine the direction of the user's gaze. In one embodiment, information regarding the direction of the user's gaze may be used to influence video rendering. For example, if it is determined that the user's eyes are viewing a particular direction, video rendering for that direction may be prioritized or emphasized. In embodiments of the present disclosure, the gaze direction and / or other eye orientation data may be used to predict a landing point associated with the corresponding gaze direction of the user's eyes during and / or at the end of a saccade defined in relation to a user viewing a display of the HMD. Saccade prediction may be performed by a saccade prediction engine 400, which will be described with respect to Figures 4A to 4C what is further described. The saccade prediction engine 400 may also be used in combination with a deep learning engine 190, which is configured to perform repetitive and computationally intensive operations. Specifically, the deep learning engine 190 may include and execute functions for saccade modeling and saccade prediction to update a target landing point on the display of the HMD 102 such that the movement of the user's eyes is consistent with the presentation of the foveal region (high-resolution region) on the display at the updated target landing point. It should be understood that the user's gaze direction may be defined with respect to the head-mounted display, with respect to the real environment in which the user is located, and / or with respect to a virtual environment rendered on the head-mounted display. Since the gaze direction may be defined with respect to the screen of the HMD, the gaze direction may be converted to a position on the screen. This position may be the center of the foveal region rendered at a high resolution.
[0065] Broadly speaking, the analysis of the images captured by the eye tracking sensor 265 provides the direction of the user's line of sight relative to the HMD 102 when considered alone. However, when combined with the tracked position and orientation of the HMD 102, the user's real-world line of sight direction can also be determined, since the position and orientation of the HMD 102 are synonymous with the position and orientation of the user's head. That is, the user's real-world line of sight direction can be determined from tracking the position movement of the user's eyes as well as tracking the position and orientation of the HMD 102. When rendering a view of the virtual environment on the HMD 102, the user's real-world line of sight direction can be applied to determine the user's virtual-world line of sight direction within the virtual environment.
[0066] Additionally, the haptic feedback module 281 is configured to provide signals to haptic feedback hardware included in the HMD 102 or another device (such as the controller 104) operated by the HMD user. The haptic feedback can take various forms of touch, such as vibration feedback, temperature feedback, pressure feedback, etc.
[0067] Figures 3A to 3C A view of an example display housing when looking towards the inner surface is shown, with the face designed to contact the display housing at the inner surface, and the view shows the interior of the HMD including the eye tracking sensor.
[0068] In particular, Figure 3A A view of an example display housing 102a when looking towards the inner surface is shown, with the face designed to contact the display housing 102a at the inner surface. As shown, the interface surface 102e surrounds the display housing 102a such that when worn, the display housing 102a substantially covers the user's eyes and facial features surrounding the eyes. This reduces the light entering the area that the user is viewing through the optics 102b, thus providing a more realistic viewing of the virtual reality scene provided by the HMD 102. When placing the display housing 102a on the user's head, the user's nose can slide into or be placed within the nose insertion area 102d. The nose insertion area 102d is the area between the optics 102b and at the lower part of the display housing 102a.
[0069] The flap 102c is designed to move or bend when the user's nose is at least partially placed in the nose insertion area 102d. As shown, the proximity sensor 206 is integrated within the display housing 102a and is directed towards an area within the nose insertion area 102d so as to capture information when the user's nose is at least partially placed within the nose insertion area 102d. The flap 102c is designed to fit near the user's nose, and when the display housing 102a is placed on the user's face, the flap helps to prevent light transmission towards the optics 102b and the user's eyes.
[0070] Also shown in Figure 3A is that the proximity sensor 302 is integrated into the inner surface of the display housing 102a and is located between the optics 102b. Thus, the placement of the proximity sensor 302 will be separated from the user's forehead, which may be closer to the interface surface 102e. However, the presence of the user's face within the HMD 102 can be sensed by the proximity sensor 302. Additionally, when the HMD 102 is worn, the proximity sensor 302 can also sense information regarding the distance, texture, image, and / or general characteristics of the user's face. As described above, the proximity sensor 302 can be defined by multiple sensors, which can be integrated at the same or different locations within the display housing 102a.
[0071] Also shown is a gaze detection sensor 265, which can be integrated at a location between the optics 102b of the display housing 102a. The gaze detection sensor 265 is configured to monitor the movement of the user's eyes when viewing through the optics 102b. The gaze detection sensor can be used to identify where the user is looking within the VR space. In an alternative embodiment, if the gaze detection sensor 265 is used to monitor the user's eyes, this information can be used for the user's avatar face such that the avatar face has eyes that move similarly to the movement of the user's eyes. The gaze detection sensor 265 can also be used to monitor when the user may be experiencing motion sickness.
[0072] The line-of-sight detector sensor 265 is configured to capture one or more parameters related to eye orientation. Information from the line-of-sight detector sensor 265 can be used to determine the line-of-sight direction (e.g., angle θ) of the user's eyes based on the orientation of the eye pupils, where the pupil is an opening in the center of the eye that allows light to enter and illuminate the retina. The line-of-sight detector sensor 265 can work in conjunction with one or more light sources (not shown) that emit energy of one or more wavelengths of invisible light (e.g., infrared) for illuminating the eyes. For example, the light source can be a light-emitting diode (LED) that directs light energy towards the eyes. The line-of-sight detector sensor 265 can be used to capture reflections of the eye's pupils, corneas, and / or irises, where the reflections are then analyzed (e.g., by a processor in the HMD 102, computer 106, etc.) to determine the line-of-sight direction and / or orientation of the pupils, which can be converted into the line-of-sight direction of the eyes. The line-of-sight direction can be referenced relative to the HMD 102 and / or the real-world space (e.g., angle θ). Various known techniques can be implemented to determine the line-of-sight orientation and / or direction, such as bright pupil tracking, dark pupil tracking, etc. Figure 4A A line-of-sight tracking system 820 including one or more light sources 401 and the one or more line-of-sight detection sensors 265 shown is configured to capture eye orientation data for determining the direction and / or orientation of the user's pupils and / or eyes.
[0073] Additionally, additional information can be determined based on the line-of-sight direction. For example, eye movement data such as the speed and acceleration of the eyes can be determined. The tracked movement of the eyes can be used to determine the user's saccades. Information from the sensors can also be used to track the user's head. For example, the information can respond to the position, movement, orientation, and change in orientation of the head. This information can be used to determine the line-of-sight direction within the real-world environment.
[0074] Figures 3B to 3C Different perspective views of the HMD 102 are also shown, which illustrate various placement positions of the line-of-sight direction sensor 265. For example, Figure 3B Examples of line-of-sight detection sensors 265a and 265b placed outside the optics 102b to capture the eye line of sight are shown. Figure 3C Include line-of-sight detection sensors 265x and 265y located between the optics 102b to capture the eye line of sight. The position of the line-of-sight detection sensors can vary within the display housing 102a and is generally positioned to provide a view towards the user's eyes. These illustrations are provided to show that the line-of-sight detection sensors can be flexibly positioned in different locations within the HMD 102.
[0075] Figure 4AShown is a prediction engine 400 configured to predict a landing point of a saccade associated with a user viewing a display of an HMD, according to one embodiment of the present disclosure. As previously described, the prediction engine 400 may be located at one or more of the HMD 102, the computer 106, and the cloud gaming server 114.
[0076] As shown, the eye tracking system 1220 is configured to determine the line of sight direction and / or orientation of the user's pupil and / or eyes. The line of sight direction may be relative to a display, such as the display of the HMD 102. As previously described, the eye tracking system 1220 includes one or more light sources 401 and one or more line of sight detection sensors 265. In particular, information from the eye tracking system 1220 is collected at one or more sample points. For example, the information may be collected periodically, with a period sufficient to sample the eye one or more times during a saccade. For example, the information may include the line of sight direction of the eye at a particular moment. The information for one or more sample points is retained in the storage device 1206 for later access, including the information for the current sample point.
[0077] In addition, the information for the current sample point is passed as an input to the prediction engine 400. More specifically, in one embodiment, the Δθ velocity generator 410 analyzes the information from the current sample point 402 and the information from the previous sample point 403 (passed from the storage device 1206 or retained in a buffer 405 accessible to the generator 410) to determine the velocity of the eye movement. Thus, the velocity generator 410 is configured to determine the velocity of the eye movement at a particular sample point based on the information from the current sample point 402 and the information from the previous sample point 403. For example, the information may be the line of sight direction at a particular time. In another embodiment, central difference velocity estimation is performed instead of backward difference. In this way, detection can be delayed and the previous position and the next position can be used to obtain a smoother velocity estimate. This may help reduce false positives when performing saccade detection.
[0078] The velocity information (e.g., dθ / dt) is provided as an input to the saccade recognizer 420. The velocity generator 410 may employ various techniques to determine when the user's eye movement is in a saccade state. In one embodiment, when the velocity reaches and / or exceeds a threshold, the eye and / or the eye movement is in a saccade state. The threshold is selected to avoid noisy information that may not necessarily indicate that the eye is making a saccade. For example, the threshold is higher than the velocity typically found when the eye is performing smooth pursuit (such as when tracking an object). For illustration only, saccade detection may be performed within 10 ms.
[0079] As described above, a saccade defines the rapid and simultaneous movement of a user's eyes when moving from one fixation point to another on a display. Saccadic movements can reach peak angular velocities of over 900 degrees per second and last for any time between 20 milliseconds (ms) and 200 ms. At a frame rate of 120 Hertz (Hz), a saccade can last for any time between 2 and 25 frames. For example, an HMD refreshes at a rate of 90 Hz or 120 Hz to minimize user discomfort (e.g., due to motion sickness).
[0080] Once it is determined that the eyes and / or eye movements are in a saccade state, the prediction engine 400 is configured to determine a landing point on the display at which the user's line of sight is directed. That is, at a particular point during the saccade (e.g., midpoint, end point, etc.), the landing point can be determined by the prediction engine 400 and more specifically by the deep learning engine 190, as Figure 4B shown. In particular, the sample set collector 430 collects information from a set of sample points to include information from the current sample point 402. Velocity information determined from the set of sample points can also be determined such that at least a portion of a full velocity map can be generated for the saccade experienced by the user. Information including a portion of the velocity map is provided as an input to the deep learning engine 190 to determine the landing point.
[0081] For example, Figure 4B illustrates a recurrent neural network as the deep learning engine 190 according to one embodiment of the present disclosure, the recurrent neural network being used to predict the landing point of a saccade associated with a user viewing a display of an HMD. The recurrent neural network includes a long short-term memory (LSTM) module 440 and a fully connected multi-layer network 450 (e.g., a multi-layer perceptron). In particular, the deep learning engine 190 is configured to compare the input information 451 (e.g., a portion of the velocity map, etc.) with a saccade model generated and / or known by the deep learning engine 190. For example, a portion of the saccade to be analyzed is compared with a velocity map constructed from multiple saccades of a test subject. In other embodiments, the input to the neural network can include information other than velocity, such as the velocity at each sample point, the line of sight direction at each sample point, and the time at each sample point. In this way, based on the saccade model constructed and / or known by the deep learning engine 190, the landing point corresponding to the direction of the user's eyes on the display can be determined for any point during the saccade. As shown, the output 452 of the deep learning engine 190 includes a vector (X F-n ), the vector indicating the line of sight direction of the user pointing to the determined landing point. Optionally, a time (t n ) parameter predicts when the user's eyes point to the landing point. The time (t n)The parameter can refer to one or more points, such as the start of a saccade, the point that determines the saccade, the latest sample point in the sample set 451 of sample points, etc.
[0082] Figure 4C An example neural network according to an embodiment of the present disclosure is shown. The neural network is, for example, used to construct a saccade model and / or a velocity map of those saccade models based on the measured saccades of a test subject, and to perform prediction of a landing point on the display of an HMD, where the landing point is associated with the line-of-sight direction of either eye of the user during and / or at the end of a saccade defined in association with the user viewing (e.g., the) display of the HMD. Specifically, the deep learning or machine learning engine 190 in the saccade prediction engine 400 is configured to receive information related to the eye orientation data of the user (e.g., line-of-sight direction, time, a portion of the velocity map of the saccade, etc.) as input. As described above, the deep learning engine 190 utilizes artificial intelligence (including deep learning algorithms, reinforcement learning, or other artificial intelligence-based algorithms) to construct saccade models, such as velocity maps of those saccade models, to identify the saccades the user is currently experiencing and predict the position where the line-of-sight direction points at any point during the saccade.
[0083] That is, during the learning and / or modeling phase, the deep learning engine 190 uses input data (e.g., measurements of the saccades of a test subject) to create saccade models (including velocity maps of those saccade models), which can be used to predict the landing point on the display where the user's eye points. For example, the input data can include multiple measurements of the saccades of a test subject, which, when fed into the deep learning engine 190, are configured to create one or more saccade models, and for each saccade model, a saccade recognition algorithm can be used to identify when the current saccade matches that saccade model.
[0084] In particular, the neural network 190 represents an example of an automated analysis tool for analyzing a data set to determine the corresponding user's responses, actions, behaviors, needs, and / or requirements. Different types of neural networks 190 are possible. In an example, the neural network 190 supports deep learning. Thus, a deep neural network, a convolutional deep neural network, and / or a recurrent neural network using supervised training or unsupervised training can be implemented. In another example, the neural network 190 includes a deep learning network that supports reinforcement learning. For example, the neural network 190 is set up to support a Markov decision process (MDP) of a reinforcement learning algorithm.
[0085] Generally, neural network 190 represents a network of interconnected nodes, such as an artificial neural network. Each node learns some information from the data. Knowledge can be exchanged between nodes through the interconnections. The input to the neural network 190 activates a set of nodes. In turn, this set of nodes activates other nodes, thus propagating knowledge about the input. This activation process is repeated across other nodes until an output is provided.
[0086] As shown, neural network 190 includes a hierarchy of nodes. At the lowest hierarchical level, there is an input layer 191. The input layer 191 includes a set of input nodes. For example, during monitoring of a test user / object undergoing a corresponding saccade (e.g., eye orientation data), each of these input nodes is mapped to local data 115 that is actively collected by an actuator or passively collected by a sensor.
[0087] At the highest hierarchical level, there is an output layer 193. The output layer 193 includes a set of output nodes. The output nodes represent decisions (e.g., predictions) related to information about the currently experienced saccade. As previously described, the output nodes can match the saccade experienced by the user with previously modeled saccades and also identify the predicted landing point on the display (e.g., of the HMD) that the user's line of sight is directed towards during and / or at the end of the saccade.
[0088] These results can be compared with predetermined and true results obtained from previous interactions and monitoring of the test object to refine and / or modify the parameters used by the deep learning engine 190 to iteratively determine a suitable saccade model for a given set of inputs and the predicted landing point on the display corresponding to the user's line of sight during and / or at the end of the saccade. That is, the nodes in the neural network 190 learn the parameters of the saccade model, and such decisions can be made using the saccade model when refining the parameters.
[0089] Specifically, there is a hidden layer 192 between the input layer 191 and the output layer 193. The hidden layer 192 includes "N" hidden layers, where "N" is an integer greater than or equal to 1. In turn, each hidden layer also includes a set of hidden nodes. The input nodes are interconnected to the hidden nodes. Similarly, the hidden nodes are interconnected to the output nodes such that the input nodes are not directly interconnected to the output nodes. If there are multiple hidden layers, the input nodes are interconnected to the hidden nodes of the lowest hidden layer. In turn, these hidden nodes are interconnected to the hidden nodes of the next hidden layer, and so on. The hidden nodes of the next highest hidden layer are interconnected to the output nodes. The interconnection connects two nodes. The interconnection has a numerical weight that can be learned, such that the neural network 190 adapts to the input and is able to learn.
[0090] Typically, the hidden layer 192 allows knowledge about the input nodes to be shared among all tasks corresponding to the output nodes. To this end, in one implementation, the transformation f is applied to the input nodes through the hidden layer 192. In one example, the transformation f is non-linear. Different non-linear transformations f can be used, including for example the rectified linear function f(x) = max(0, x).
[0091] The neural network 190 also uses a cost function c to find an optimal solution. The cost function measures the deviation between the prediction output by the neural network 190 (defined as f(x) for a given input x) and the ground truth or target value y (e.g., the expected result). The optimal solution represents a situation where the cost of no solution is lower than the cost of the optimal solution. For data for which such ground truth labels are available, an example of the cost function is the mean squared error between the prediction and the ground truth. During the learning process, the neural network 190 can use the backpropagation algorithm to adopt different optimization methods to learn the model parameters (e.g., the weights of the interconnections between the nodes in the hidden layer 192) that minimize the cost function. An example of such an optimization method is stochastic gradient descent.
[0092] In one example, the training dataset for the neural network 190 can be from the same data domain. For example, the neural network 190 is trained to learn the patterns and / or characteristics of similar saccades of a test subject based on a given set of inputs or input data. For example, the data domain includes eye orientation data. In another example, the training dataset comes from different data domains to include input data other than the baseline. In this way, the neural network 190 can use eye orientation data to identify saccades, or can be configured to generate a saccade model for a given saccade based on eye orientation data.
[0093] Figure 5A A rendering pipeline 501 without saccade prediction according to an embodiment of the present disclosure is shown, which shows how frame updates are slower than eye movements, such that the image is blurred to the user after the eye movement is completed. As previously described, the rendering pipeline 501 can be implemented within the HMD 102, the computer 106, and the cloud gaming server 114 alone or in combination.
[0094] Although the rendering pipeline 501 without landing point prediction enabled is shown in Figure 5A , it should be understood that in the embodiments of the present disclosure, the rendering pipeline 501 can be optimized to analyze the eye tracking information in order to identify saccades and eye movements, and predict the landing point (e.g., open) where the line of sight of the user's eyes 260 points during and / or at the end of a saccade on a display (e.g., of the HMD 102), as Figure 5B shown. That is, in Figure 5BIn which, the rendering pipeline 501 can be configured to perform foveal rendering based on the prediction of the landing point, as will be further described below with respect to Figure 5B Further described.
[0095] Specifically, the rendering pipeline includes a central processing unit (CPU) 1202, a graphics processing unit (GPU) 1216, and a memory accessible to both (e.g., vertex buffer, index buffer, depth or Z buffer, frame buffer for storing the rendered frame to be passed to the display, etc.). The rendering pipeline (or graphics pipeline) illustrates the general process of rendering an image, such as when using a 3D (three-dimensional) polygon rendering process. For example, the rendering pipeline 501 for the rendered image outputs corresponding color information for each pixel in the display, where the color information can represent texture and shading (e.g., color, shadow, etc.).
[0096] The CPU 1202 can generally be configured to perform object animation. The CPU 1202 receives input geometry corresponding to an object within the 3D virtual environment. The input geometry can be represented as vertices within the 3D virtual environment, along with information corresponding to each vertex. For example, an object within the 3D virtual environment can be represented as a polygon (e.g., triangle) defined by vertices, where the surface of the corresponding polygon is then processed by the rendering pipeline 501 to achieve the final effect (e.g., color, texture, etc.). The operation of the CPU 1202 is well-known and is generally described herein. Generally, the CPU 1202 implements one or more shaders (e.g., compute, vertex, etc.) to perform object animation frame by frame based on the forces applied to and / or exerted by the object (e.g., external forces such as gravity, and internal forces of the object causing movement). For example, the CPU 1202 performs physical simulation and / or other functions of the object within the 3D virtual environment. Then, the CPU 1202 issues draw commands for the polygon vertices to be executed by the GPU 1216.
[0097] In particular, the animation results generated by the CPU 1202 can be stored in a vertex buffer and then accessed by the GPU 1216, which is configured to perform the projection of polygon vertices onto a display (e.g., of an HMD) and the tessellation of the projected polygons for the purpose of rendering the polygon vertices. That is, the GPU 1216 can be configured to further construct the polygons and / or primitives that make up the objects within the 3D virtual environment, including performing lighting, shadow, and shading calculations for the polygons, depending on the lighting of the scene. Additional operations can be performed, such as culling to identify and ignore primitives outside the frustum, and rasterization to project the objects in the scene onto the display (e.g., projecting the objects onto an image plane associated with the user's viewpoint). At a simplified level, rasterization involves looking at each primitive and determining which pixels are affected by that primitive. Fragmentation of the primitive can be used to break the primitive into pixel-sized fragments, where each fragment corresponds to a pixel in the display and / or a reference plane associated with the rendering viewpoint. When rendering a frame on the display, one or more fragments of one or more primitives may contribute to the color of a pixel. For example, for a given pixel, the fragments of all primitives in the 3D virtual environment are combined into the pixel of the display. That is, the overall texture and shading information for the corresponding pixel is combined to output the final color value of the pixel. These color values can be stored in a frame buffer and, when the corresponding image of the scene is displayed frame by frame, these color values are scanned into the corresponding pixels.
[0098] The rendering pipeline 501 can include a gaze tracking system 1220, which is configured to provide gaze direction and / or orientation information to the CPU 1202. This gaze direction information can be used for the purpose of performing foveated rendering, where the foveal region is rendered at a high resolution and corresponds to the direction of the user's gaze. Figure 5A The rendering pipeline 501 configured for foveated rendering but without saccade prediction (i.e., saccade prediction is turned off) is shown. That is, no landing point prediction is performed, and as a result, the foveal region displayed on the HMD has a foveal region that is inconsistent with the user's eye movements because each calculated foveal region is out of date when it is displayed, especially during eye movements. Additionally, Figure 5A A timeline 520 is shown, which indicates the time at which frames (e.g., F1 to F8) in the scan output sequence from the rendering pipeline 501 are scanned. The sequence of frames F1 to F8 is also part of the saccades of the user viewing the display.
[0099] As Figure 5AAs shown, the rendering pipeline is shown as including operations sequentially performed by a gaze tracking system 1220, a CPU 1202, a GPU 1216, and a raster engine that scans the rendered frame out to a display 1210. For illustration, rendering pipeline sequences 591 through 595 are shown. Due to space limitations, other pipeline sequences, such as the sequence for frames F3 through F22, are not shown. In the example as Figure 5A shown, each component of the rendering pipeline 501 operates at the same frequency. For example, the gaze tracking system 1220 can output gaze direction and / or orientation information at 120 Hz, which can be the same frequency used by the rendering pipelines of the CPU 1202 and the GPU 1216. In this way, the gaze direction of the user's eye 260 can be updated for each frame scanned out in the rendering pipeline. In other embodiments, the gaze tracking system 1220 does not operate at the same frequency, such that the gaze direction information may not be aligned with the rendered frame being scanned out. In that case, if the frequency of the gaze tracking system 1220 is slower than the frequency used by the CPU 1202 and the GPU 1216, the gaze direction information can add further latency.
[0100] Gaze tracking information can be used to determine the foveal region rendered at high resolution. Regions outside the foveal region are displayed at a lower resolution. However, as Figure 5AAs shown, without saccade prediction, by the time the eye-tracking information is used to determine the frame to be scanned out, at least two frame cycles and at most three frame cycles have passed, after which the corresponding frame using the eye-tracking information is displayed. For example, in the rendering pipeline sequence 591, the eye-tracking information is determined in the first frame cycle at time t-20 (the midpoint of the saccade) and is passed to the CPU 1202. In the second frame cycle at time t-21, the CPU 1202 performs physical simulation on the object and passes the polygon primitives together with the drawing instructions to the GPU 1216. In the third frame cycle at time t-23, the GPU performs primitive assembly to generate the rendered frame (F23). Additionally, the GPU may render the foveal region corresponding to the line-of-sight direction passed in the first frame cycle at time t-20, and the foveal region was determined at least two frame cycles earlier. In the fourth frame cycle at time t-23, the frame F23 including the foveal region is scanned out. It is worth noting that in the rendering pipeline sequence 591, the eye-tracking information determined at time t-20 is outdated by at least the frame cycles at t-21 and t-22 (two frame cycles), and possibly a part of the third frame cycle. Similarly, the pipeline sequence 592 scans out the frame F24 at time t-24, which has the foveal region defined in the first frame cycle at time t-21 previously. Likewise, the pipeline sequence 593 scans out the frame F25 at time t-25, which has the foveal region defined in the first frame cycle at time t-22 previously. Furthermore, the pipeline sequence 594 scans out the frame F26 at time t-26, which has the foveal region defined in the first frame cycle at time t-23 previously. And the pipeline sequence 595 scans out the frame F27 at time t-27, which has the foveal region defined in the first frame cycle at time t-24 previously.
[0101] Because the eye 260 continuously moves past the points (e.g., time) detected for each rendering pipeline (e.g., at the start of rendering pipeline sequence 591, or 592 or 593, etc.), the foveal region in the frame (e.g., frame F27) of the corresponding rendering pipeline sequence (e.g., sequence 595) being scanned may be out of date by at least 2 to 3 frame periods. For example, the rendered frame F27 at the time of scan output will have a foveal region that does not match the direction of the user's line of sight. Specifically, the shown display 1210 shows the frame F27 at time t - 27, where the saccade path 510 (between frame F0 and F27) is superimposed on the display 1210 and shows the fixation point A (e.g., direction 506 and vector XF - 0) corresponding to the start of the saccade. For illustration, at the start of the saccade path 510 at time t - 0, the frame F1 was scanned and output, where the foveal region is centered on the fixation point A. The saccade path 510 includes a fixation point B corresponding to the end of the saccade, or at least a second point of the saccade. For illustration, at the end of the saccade path 510 at time t - 27, the frame F27 is scanned and output.
[0102] Furthermore, because there is no prediction of the saccade path performed by the rendering pipeline 501, the line of sight direction information provided by the eye tracking system 1220 is out of date by at least two or three frame periods. Thus, when the frame F27 is scanned and output for the rendering pipeline sequence 595 at time t - 27, although the eye 260 is fixated on the fixation point B (eye direction 507 and vector X F-27 )), the rendering pipeline sequence 595 uses the out-of-date line of sight direction information provided at time t - 24. That is, the line of sight direction information determined at time t - 24 propagates through the rendering pipeline sequence 595 for scan output at time t - 27. Specifically, the eye tracking system 1220 notes the line of sight direction pointing to the point 591 on the display 1210 at time t - 24. Thus, at time t - 24, when the frame F24 is scanned and output, the user's eye 260 points to the point 591 on the display. Whether the foveal region of the rendered frame F24 is correctly located on the display 1210 may not matter because during a saccade, the image received by the eye 260 is not fully processed and may appear blurry to the viewer. However, when the frame F27 is scanned and output on the display 1210 at time t - 27, even if the rendered foveal region 549 is calculated to be near the point 591 at time t - 24, the user's eye points to the fixation point B (as shown by the dashed region 592). Thus, for a user whose eye 260 points to and focuses on the region 592 at time t - 27, the frame F27 appears blurry because the region 592 is calculated to be in the periphery and can be rendered at a lower resolution, while the out-of-date foveal region 549 (where the eye is not pointing) is rendered at a high resolution, as previously described.
[0103] Figure 5B shows the resulting effect of a rendering pipeline according to an embodiment of the present disclosure, the rendering pipeline being configured with saccade prediction of the eye movements of a user viewing a display of an HMD such that, after the eye movement is completed, the image is made clear to the user by advancing the update of the high-resolution foveal region in the rendering pipeline. For example, Figure 5A the rendering pipeline 501 shown in now has saccade prediction enabled, more specifically, landing point prediction enabled. That is, the rendering pipeline 501 is now optimized to analyze eye-tracking information in order to identify saccades and eye movements and to predict the landing point (e.g., turn on) at which the line of sight of the user's eyes 260 points during and / or at the end of a saccade on a display (e.g., of the HMD 102). In this way, in Figure 5B , the rendering pipeline 501 is now configured to perform foveal rendering based on the prediction of the landing point. For illustrative purposes only, Figure 5B shows the prediction of the landing point at the end of a saccade 510 in one embodiment, but in other embodiments, it is possible to predict the landing point during the saccade (e.g., predicting the landing point 3 to 5 frame cycles beyond the current sample point).
[0104] Specifically, it shows that the display 1210 presents frame F27 at time t-27. The saccade path 510 is superimposed on the display 1210, and the fixation point A corresponding to the start of the saccade is shown (e.g., direction 506 and vector X F-0 ). For illustration, at the start of the saccade path 510 at time t-0, frame F1 is scanned out, with the foveal region centered on the fixation point A. The saccade path 510 includes a fixation point B corresponding to the end of the saccade, or at least a second point of the saccade. For illustration, frame F27 is scanned out at the end of the saccade path 510 at time t-27.
[0105] When each frame is scanned out, saccade prediction is performed within the rendering pipeline 501. In one embodiment, saccade prediction and / or landing point prediction can be performed within the CPU 1202, GPU 1216, or a combination of both. In another embodiment, saccade prediction is performed remotely and passed as an input into the rendering pipeline 501. Once the prediction is performed, the GPU 1216 can render a frame with a foveal region based on the landing point prediction. Specifically, the GPU 1216 can modify the foveal rendering such that instead of relying on outdated line-of-sight direction information as previously described in Figure 5A , the predicted landing point is used to determine the position of the foveal region.
[0106] Specifically, the predicted landing points of the fixation point B are superimposed on the display 1210. These landing points are determined in the previous rendering pipeline sequence. Specifically, by the time frame F8 and subsequent frames are scanned out, the predicted landing points of the saccade have converged to, for example, the fixation point B. As shown in the figure, at a certain moment after the saccade is detected (e.g., during the scan out of frame F5 at time t - 5), the prediction is performed. For example, the prediction can be performed starting from the rendering of frame F5 and subsequent frames. The saccade 510 was previously introduced in Figure 5A and includes the fixation point A as the starting point, and the fixation point B (e.g., as the end point or a predefined point within the saccade, such as the future 3 to 5 frame periods).
[0107] When frame F5 is scanned out, the predicted landing point (e.g., centered on vector X F-5 ) is shown as the predicted foveal region 541, which deviates from the fixation point B. In the next rendering pipeline sequence, when frame F6 is scanned out, the predicted landing point (e.g., centered on vector X F-6 ) is shown as the predicted foveal region 542, which is close to but still deviates from the fixation point B. Due to the prediction convergence, in the next rendering pipeline sequence, when frame F7 is scanned out, the predicted landing point (centered on vector X F-7 ) is shown as the predicted foveal region 543, which is very close to the fixation point B. The convergence may occur in the next rendering pipeline sequence when frame F8 is scanned out, where the predicted landing point (e.g., centered on vector X F-8 ) is shown as the predicted foveal region 592 (bold), which is centered on the fixation point B. For any subsequent rendering pipeline sequence, the foveal region 592 is used for rendering and is centered on the fixation point B, such as when rendering and scanning out frames F9 to F27 of the saccade 510. In this way, when frame F27 is rendered, due to the landing point prediction to the fixation point B, the foveal region 592 is consistent with the movement of the user's eye 260, such as at the end of the saccade. Moreover, since the prediction of the landing point converges together with the rendering and scan out of frame F8, all frames F9 to F27 may already have the foveal region 592 to prepare for the user's eye movement. Thus, instead of rendering at the foveal region 549 without prediction (as described in Figure 5A ), the predicted foveal region 592 is used to update the target landing point (e.g., a defined number of future frame periods, the end of the saccade, etc.), such that when the eye reaches the predicted landing point, the frame is rendered using the foveal region centered on the predicted landing point.
[0108] Figure 6AFIG. 600A shows the eye displacement and velocity of saccades of a user viewing a display of an HMD in accordance with an embodiment of the present disclosure. FIG. 600A includes a vertical axis 610A showing the angular velocity (dθ / dt) of eye movement during a saccade. Additionally, FIG. 600A includes another vertical axis 610B showing the angular displacement (θ). FIG. 600A includes a horizontal axis 615 showing time and includes a time series of saccades between time t-0 and approximately t-27 and / or t-28.
[0109] For illustration only, FIG. 600A shows the angular displacement of a saccade in line 630. As previously introduced, a saccade defines the rapid and simultaneous movement of a user's eyes when moving from one fixation point on a display to another. As shown, the angular movement as shown by the displacement line 630 of the eye is in a particular direction (e.g., from left to right). That is, in the example of FIG. 600A, during the saccade, the line of sight direction of the eye moves between 0 degrees and 30 degrees.
[0110] Accordingly, for illustration only, FIG. 600A shows the eye velocity during a saccade in line 620. The velocity profiles of different saccades generally follow the same shape as shown in line 620. For example, at the start of a saccade, the saccade velocity follows a linear progression (e.g., between time t-0 and t-8). After the linear progression, the velocity may tend to level off, such as between time t-8 and t-17. The velocity profile in line 620 shows a sharp decline in velocity after leveling off until the end of the saccade (such as between time t-17 and t-27).
[0111] Embodiments of the present disclosure match a portion of the velocity profile of a current saccade (e.g., the linear portion of line 620) with a modeled saccade (e.g., constructed during training of the deep learning engine 190). The landing point of the current saccade can approximate the landing point of the modeled saccade at any point during the saccade and can be predicted for the current saccade.
[0112] Figure 6B FIG. 600B shows the sampling of eye orientation / tracking data at various sample points in the velocity profile of saccades of a user viewing a display of an HMD in accordance with an embodiment of the present disclosure. FIG. 600B follows Figure 6A that of FIG. 600A, which includes a vertical axis 610A showing the angular velocity (dθ / dt) of eye movement during a saccade, and a horizontal axis 615, but is cropped to show only the velocity of the saccade in line 620.
[0113] Specifically, at each sample point during a saccade, eye orientation / tracking data is collected from the eye tracking system 1220. For illustrative purposes only, sample points may occur at least at times t-0, t-1, t-2, ......, t-27, ......, t-n. For example, the sample point S1 shown on line 620 is associated with the eye tracking data (e.g., line of sight direction, velocity, etc.) at time t-1, the sample point S2 is associated with the eye tracking data at time t-2, the sample point S3 is associated with the eye orientation / tracking data at time t-4, the sample point S5 is associated with the eye tracking data at time t-5, the sample point S6 is associated with the eye tracking data at time t-6, the sample point S7 is associated with the eye tracking data at time t-7, the sample point S8 is associated with the eye tracking data at time t-8, and so on. For example, the data collected at each sample point may include line of sight direction, time, and other information, as previously described. Based on the data, velocity information of the user's eyes can be determined. In some embodiments, velocity data can be collected directly from the eye tracking system 1220.
[0114] Thus, during the saccade indicated by the velocity line 620, eye tracking data is collected and / or determined for at least sample points S1 to approximately S 27 . The sample points S1 to S8 are highlighted in FIG. 600B to show the predicted convergence of the landing point of the saccade (e.g., the end point of saccade 510), as previously indicated in Figure 5B , which shows convergence at approximately time t-8 corresponding to sample point S8.
[0115] Figure 6C Shows the collection of eye orientation / tracking data for one or more sets of sample points for predicting the landing point of a saccade associated with the eyes 260 of a user viewing a display of an HMD according to one embodiment of the present disclosure. Figure 6C Shows the use of information at the sample points introduced in Figure 6B FIG. 600B of to predict the landing point of a saccade of a user's eyes.
[0116] As shown, the eye orientation / tracking data is collected by the eye tracking system 1220 on the eyes 260 at a plurality of sample points 650 (e.g., S1 to S at each of times t-1 to t-27). For illustrative purposes only, the saccade is shown to move between 0 and 30 degrees between fixation point A and fixation point B. 27 )
[0117] Specifically, eye orientation / tracking data is collected and / or velocity data is determined from each sample point 650. The circles 640 highlighting the sample points are shown enlarged and include at least sample points S1 to S8. For example, velocity data V1 is associated with sample point S1, velocity data V2 is associated with sample point S2, velocity data V3 is associated with sample point S3, velocity data V4 is associated with sample point S4, velocity data V5 is associated with sample point S5, velocity data V6 is associated with sample point S6, velocity data V7 is associated with sample point S7, and at least velocity data V8 is associated with sample point S8. Additional data for the remaining sample points is collected and / or determined but not shown in circle 640.
[0118] For prediction purposes, once the user's eye is identified as being in a saccade state, information is collected from the set of sample points. For illustrative purposes, saccade identification may occur at time t-5 associated with sample point S5, which also coincides with the start of the rendering of future frame F8. In one embodiment, saccade identification is confirmed once the velocity of the eye reaches and / or exceeds a threshold velocity.
[0119] After saccade identification, a prediction of a predefined landing point is performed. Specifically, information from the set of sample points is identified. At a minimum, the information includes the measured and / or calculated angular velocity. The set may contain a predefined number of sample points, including the current sample point. For example, the set may contain 1 to 10 sample points. In one embodiment, the set may contain 3 to 5 sample points to reduce error.
[0120] For illustrative purposes, the set may contain 4 sample points, including the current sample point, as Figure 6C described. A sliding window for collecting information from the set of sample points is shown. For example, at the frame period or time corresponding to the current sample point S5, window (w1) includes sample points S2 to S5, where the corresponding information (e.g., velocity) from those sample points is used to predict the landing point. At the next frame period or time corresponding to the next current sample point S6, window (w2) includes sample points S3 to S6, where the corresponding information is used to predict the updated landing point. Further, at the next frame period or time corresponding to the next current sample point S7, window (w3) includes sample points S4 to S7, where the corresponding information is used to predict the updated landing point. Convergence may occur at the next frame period or time corresponding to the next current sample point S8, where window (w4) includes sample points S5 to S8.
[0121] Previously with respect to Figure 5B convergence was described. Confirmation of convergence may occur along with subsequent predictions of the landing point, such as for windows w5,......, w27. In one embodiment, once convergence is confirmed, prediction may be disabled.
[0122] Figure 6D Shows table 600D according to an embodiment of the present disclosure, which lists eye orientation data of a set of sample points for predicting the landing point (e.g., saccade end point) associated with a user viewing a display of an HMD. Figure 6D Aligned with Figure 5B shows the prediction and convergence of the predicted landing point at fixation point B.
[0123] In particular, column 661 shows the window name (e.g., w1 to w5); column 662 shows the set of sample points; column 663 shows the predicted landing point consistent with the end of the saccade, where the angular displacement is referenced to the start point of the saccade at fixation point A; and column 664 shows the predicted saccade end time (e.g., represented by a frame or frame period), where the predicted end time is referenced to the start time of the saccade at fixation point A.
[0124] For example, window w1 uses information (e.g., velocity) from a set of sample points including sample points S2 to S5 to predict the landing point (saccade end point). The predicted line of sight direction of the user's eyes for the predicted landing point is vector X F-5 , which has an angle of 42 degrees. Figure 5B shows the predicted landing point centered on fixation region 541. Additionally, the predicted end time or duration of the saccade is predicted to be approximately the time t-38 associated with frame and / or frame period F38.
[0125] Furthermore, window w2 uses information (e.g., velocity) from a set of sample points including sample points S3 to S6 to predict an updated landing point (saccade end point). The predicted line of sight direction of the user's eyes for the predicted landing point is vector X F-6 , which has an angle of 18 degrees. Figure 5B shows the predicted landing point centered on fixation region 542. Additionally, the predicted end time or duration of the saccade is predicted to be approximately the time t-20 associated with frame and / or frame period F20.
[0126] Window w3 uses information (e.g., velocity) from a set of sample points including sample points S4 to S7 to predict an updated landing point (saccade end point). The predicted line of sight direction of the user's eyes for the predicted landing point is vector X F-7 , which has an angle of 28 degrees, close to fixation point B at a 30-degree angle. Figure 5B shows the predicted landing point centered on fixation region 543 close to fixation point B. Additionally, the predicted end time or duration of the saccade is predicted to be approximately the time t-25 associated with frame and / or frame period F25.
[0127] Windows w4 and w5 show the convergence of the predicted landing points (e.g., saccade end points). That is, the predictions associated with these windows show a landing point of 30 degrees (e.g., from fixation point A). For example, window (w4) predicts the landing point using sample points S5 to S8. The predicted line-of-sight direction of the user's eye and the predicted landing point is vector X F-8 , which has an angle of 30 degrees, which is also the angle with respect to fixation point B. The predicted end time or duration of the saccade is predicted to be approximately the time t-27 associated with frame and / or frame period F27. Also, window (w5) uses sample points S6 to S9 to predict the same landing point of 30 degrees at fixation point B, where the predicted end time or duration of the saccade is also the time t-27 associated with frame and / or frame period F27. Thus, convergence occurs at window (w4), and confirmation of the convergence occurs at window (w5). Subsequent predictions should show the converged landing point.
[0128] With the detailed description of the various modules of the game console, HMD, and cloud game server, now according to one embodiment of the present disclosure, relative to Figure 7 Flowchart 700 describes a method for predicting a landing point associated with the line-of-sight direction of the eyes of a user experiencing a saccade on a display (e.g., of an HMD), where the landing point can occur at any point during or at the end of the saccade. As previously mentioned, flowchart 700 shows the processes and data flows of the operations involved in predicting the landing point at one or more of the HMD, game console, and cloud game server. In particular, the method of flowchart 300 can be at least partially performed by Figures 1A to 1C , Figure 2 and Figures 4A to 4C 's saccade prediction engine 400.
[0129] At 710, the method includes tracking the movement of at least one eye of the user at a plurality of sample points using a line-of-sight tracking system disposed in the HMD. For example, eye orientation / tracking data can be collected to include at least the line-of-sight direction. For example, according to one embodiment of the present disclosure, the line-of-sight direction can be shown at various times from t0 to t5 in Figure 6E . In Figure 6E , at time t0, the line-of-sight direction is defined by vector X t0 ; the line-of-sight direction defined by vector X t1 is at time t1; the line-of-sight direction defined by vector X t2 is at time t2; the line-of-sight direction defined by vector X t3 is at time t3; the line-of-sight direction defined by vector X t4 is at time t4; and the line-of-sight direction defined by vector X t5 is at time t5.
[0130] At 720, the method includes determining the velocity of the motion based on the tracking. As Figure 6E shown, the line-of-sight direction vectors can be used to determine the velocity of the user's eyes. That is, the velocity of the eyes can be determined based on a first eye or line-of-sight direction and a second eye or line-of-sight direction from two sample points. In particular, the line-of-sight direction between two sample points, the angle between the two line-of-sight directions, and the time between the two sample points can be used to determine the velocity between the two sample points. For example, one of a variety of techniques (including the trigonometric functions defined in the following equations) can be used to determine the angle (θ) between two sample points. By way of illustration, in one embodiment, the following equation (1) is used to determine the angle (θ) between two sample points obtained at time t n and time t n-1 . Referring to Figure 6E , the angle θ2 can be determined from vectors X t1 and X t2 , the angle θ3 can be determined from vectors X t2 and X t3 , θ4 can be determined from vectors X t3 and X t4 , and θ5 can be determined from vectors X t4 and X t5 .
[0131]
[0132] Equation 1 gives the angle between the line-of-sight directions at two sample points obtained at time t n and time t n-1 . To calculate the velocity between two sample points, the angle is divided by Δt, the duration between the two sample points, as shown in the following equation (2).
[0133] Velocity (degrees per second) = θ / (t n – t n-1 ) (2)
[0134] Thus, the velocity (v2) between sample points obtained at times t1 and t2 can be determined using vectors X t1 and X t2 , the velocity (v3) between sample points obtained at times t2 and t3 can be determined using vectors X t2 and X t3 , the velocity (v4) between sample points obtained at times t3 and t4 can be determined using vectors X t3 and X t4 , and the velocity (v5) between sample points obtained at times t4 and t5 can be determined using vectors X t4 and X t5 .
[0135] In one embodiment, at 730, the method includes determining that the user's eyes are in a saccade state when a speed reaches a threshold speed. In other embodiments, other methods may be used to determine that the user's eyes are in a saccade state. As previously mentioned, the threshold speed is predefined to avoid identifying saccades when the eyes may be undergoing another type of movement (e.g., smooth pursuit) or when the data is noisy.
[0136] At 740, the method includes predicting a landing point corresponding to the saccade eye direction on a display of the HMD. In one embodiment, the direction corresponds to the line of sight direction of the eyes. Since the line of sight direction can be defined relative to the screen of the HMD, the line of sight direction can be converted to a position on the screen, where the position is the landing point. The landing point can be used as the center of the foveal region rendered at high resolution for the frame. In one embodiment, the landing point can occur at any point during the saccade to include the midpoint of the saccade corresponding to the intermediate direction of the eyes. For example, in one embodiment, the landing point can occur a predefined number of frame periods beyond the current frame period. In another embodiment, the landing point can occur at the end of the saccade and correspond to the fixation direction of the eyes.
[0137] The prediction of the landing point can include: collecting eye orientation / tracking data when tracking the movement of the eyes for a set of sample points. That is, the landing point is predicted using information from the set of sample points. The eye orientation / tracking data includes at least the eyes and / or the line of sight direction relative to the HMD, where at least one sample point of the set occurs during the saccade. As previously mentioned, speed information can be determined from the eye orientation / tracking data, and the speed data can also be used to predict the landing point. Additionally, the eye orientation / tracking data of the set of sample points is provided as an input to a recurrent neural network (e.g., a deep learning engine). For example, the neural network is trained based on previously measured eye orientation data of multiple saccades of a test subject. In one embodiment, the recurrent neural network includes a long short-term memory neural network and a fully connected multi-layer perceptron network. The recurrent neural network can be configured to compare a part of the eye speed map constructed from the eye orientation data of the set of sample points with the eye speed map constructed from multiple saccades of the test subject. The trained saccades in the recurrent neural network can be utilized to perform a match between parts of the eye speed map of the user's saccades. Once the match is made, one or more predicted landing points of the user's saccades can approximate one or more landing points of the trained saccades. In this way, the recurrent neural network can be used to predict the landing point of the saccade (e.g., the end point of the saccade, or the intermediate point during the saccade) using information from the set of sample points.
[0138] Additionally, different sets of sample point data can be used to update the prediction of the landing point using subsequent predictions. For example, a first landing point is predicted at 741 associated with a first current sample point. The prediction of the first landing point is based on eye orientation data of a first set of sample points including the first sample point and at least one previous sample point. The eye orientation data includes the eye and / or line of sight direction relative to the HMD. At 742, an updated prediction is performed associated with a second current sample point after the first sample point in the saccade. The update of the landing point includes predicting a second landing point based on eye orientation data of a second set of sample points including the second sample point and at least one previous sample point (e.g., the first sample point).
[0139] At decision step 743, the method determines whether there is convergence of the predicted landing points. For example, convergence may occur when the two predicted landing points are within a threshold measurement (e.g., the incremental distance between the two predicted landing points on the display). In one embodiment, convergence occurs when the two predicted landing points are the same.
[0140] If there is no convergence, the method proceeds to 744, where another prediction is performed. In particular, at the next sample point after the previous sample point, the next landing point is predicted based on the eye orientation / tracking data from the next sample point and at least one previous sample point. The method returns to decision step 743 to determine whether there is convergence.
[0141] On the other hand, if there is convergence, the method proceeds to 745, where the last predicted landing point is selected as the landing point predicted for the saccade. That is, due to the convergence, the last calculated landing point is used as the predicted landing point.
[0142] In one embodiment, foveal rendering can be performed based on the predicted landing point. For example, a first video frame can be rendered for display, where the first video frame includes a foveal region centered at the predicted landing point on the display. The foveal region can be rendered at a high resolution. Additionally, the non-foveal region of the display includes the rest of the display and is rendered at a lower resolution. Further, the first video frame having the foveal region is presented on the display of the HMD, where when the first video frame is displayed, the predicted eye orientation is towards the landing point (i.e., corresponding to the foveal region).
[0143] In another embodiment, when rendering a frame for display on an HMD, additional measures can be taken to reduce power consumption. In particular, during a saccade, the user may not be able to view intermediate frames that are rendered and displayed because eye movements may be too fast. As such, rendering of at least one of the intermediate video frames can be terminated to save computational resources that would otherwise be used for rendering. That is, the method includes terminating rendering of at least one video frame to be rendered before a first video frame during a saccade.
[0144] In yet another embodiment, when rendering a frame for display on an HMD, another measure can be taken to reduce power consumption. In particular, since the user may not be able to view intermediate frames that are rendered and displayed during a saccade, the entire video frame may be rendered at a lower resolution or a low resolution. That is, the foveal region is not rendered for the frame. In other words, the method includes rendering at least one video frame to be rendered before a first video frame during a saccade at a low resolution.
[0145] Late updating the predicted landing point to a buffer accessible by the GPU
[0146] Figure 8 A computer system implementing a rendering pipeline 800 according to an embodiment of the present disclosure is shown, the rendering pipeline being configured for foveal rendering and including: predicting a landing point of a saccade associated with a user viewing a display of an HMD, and providing the landing point as a late update to a buffer accessible to a GPU of the computer system for immediate use in rendering a high-resolution foveal region centered on the landing point in a corresponding video frame.
[0147] The rendering pipeline 800 illustrates a general process of rendering an image using a 3D (three-dimensional) polygon rendering process, but is modified to perform additional programmable elements within the pipeline to perform foveal rendering, such as predicting the landing point and late-updating the landing point for immediate rendering of the foveal region of a corresponding video frame. The rendering pipeline 800 outputs corresponding color information for each pixel in the display of the rendered image, where the color information can represent textures and shadings (e.g., colors, shadows, etc.). The rendering pipeline 800 can be implemented at least in Figures 1A to 1C computer system 106, system 1200, Figure 2 and Figure 13 HMD 102 of Figure 14 and within client device 1410 of
[0148] The rendering pipeline 800 includes a CPU 1202 configured to execute an application, and a GPU configured to execute one or more programmable shader operations for foveal rendering. The programmable shader operations include processing vertex data, assembling vertices into primitives (e.g., polygons), performing rasterization relative to a display to generate fragments from the primitives, then calculating the color and depth values for each fragment, and blending the fragments pixel by pixel for storage in a frame buffer for display.
[0149] As shown, the application is configured to create geometric primitives 805, such as vertices within a 3D virtual environment, and information corresponding to each vertex. For example, an application executed on the CPU 1202 makes changes to the scene of the 3D virtual environment, where these changes represent the application of physical properties, animation, deformation, acceleration techniques, etc. Scene changes (e.g., object movement) are calculated frame by frame based on forces (e.g., external forces such as gravity and internal forces causing movement) applied and / or exerted on the object. Objects in the scene can be represented using polygons (e.g., triangles) defined by the geometric primitives. Then, the surface of the corresponding polygon is processed by the GPU 1216 in the rendering pipeline 800 to achieve the final effect (e.g., color, texture, etc.). Vertex attributes can include normals (e.g., which direction is light relative to the vertex), colors (e.g., RGB - red, green, and blue triplets, etc.), and texture coordinate / mapping information. The vertices are stored in the system memory 820 and then accessed by the GPU and / or transferred to the GPU's memory 840.
[0150] In addition, the CPU 1202 is configured to perform saccade prediction. In particular, the CPU 1202 includes a prediction engine 400 configured to predict the landing point of a saccade associated with a user viewing the display (e.g., an HMD). That is, the prediction engine 400 predicts the movement of the user's eyes during the identified saccade in order to predict the landing point on the display corresponding to the predicted direction of the eyes at any point during the saccade. A prediction time corresponding to the predicted direction of the eyes is also generated. As previously described, the eye tracking system 1220 provides eye orientation information to the prediction engine 400 for use during the prediction phase. For example, the eye tracking system 1220 is configured to determine the pupil and / or the line of sight direction and / or orientation of the user's eyes at discrete time points. The line of sight direction can be relative to the display. In particular, information from the eye tracking system 1220 is collected at one or more sample points 801 and provided to the prediction engine 400.
[0151] The CPU 1202 includes a late update module 830 that is configured to perform late updates of the eye tracking information and / or prediction of the landing point of the eye movement during a saccade, as will be described below with respect to FIGS. 9-11. For example, the predicted landing point and / or other eye tracking information 802 can be provided as a late update to a buffer accessible to the GPU (e.g., GPU memory 840) for immediate processing. In this way, when rendering a video frame for display, the prediction information and / or eye tracking information 802 can be used within the same frame cycle or almost immediately, rather than waiting to be processed by the GPU in the next frame cycle.
[0152] The GPU 1216 includes one or more shader operations, such as a rasterizer, a fragment shader, and a renderer, including an output combiner and a frame buffer for rendering images and / or video frames of a 3D virtual environment, but not all are shown in Figure 8 it.
[0153] In particular, the vertex shader 850 can also construct the primitives that make up the objects within the 3D scene. For example, the vertex shader 410 can be configured to perform lighting and shading calculations for polygons, depending on the lighting of the scene. The vertex processor 410 can also perform additional operations, such as clipping (e.g., identifying and ignoring primitives outside the frustum defined by the viewing position in the game world).
[0154] The primitives output by the vertex processor 410 are fed to a rasterizer (not shown) that is configured to project the objects in the scene onto the display according to the viewpoint within the 3D virtual environment. At a simplified level, the rasterizer looks at each primitive and determines which pixels are affected by the corresponding primitive. In particular, the rasterizer divides the primitive into pixel-sized fragments, where each fragment corresponds to a pixel in the display and / or a reference plane associated with the rendering viewpoint (e.g., the camera view).
[0155] The output from the rasterizer is provided as input to the foveal fragment processor 430, which performs a shading operation on the fragment at its core to determine how the color and brightness of the primitive vary with the available lighting. For example, the fragment processor 430 can determine the Z-depth, color, alpha value for transparency, normal, and texture coordinates (e.g., texture details) of the distance from the viewing position for each fragment, and can also determine an appropriate lighting level based on the available lighting, darkness, and color of the fragment. In addition, the fragment processor 430 can apply a shadow effect to each fragment.
[0156] In an embodiment of the present invention, fragments of a particle system are rendered differently depending on whether the fragment is inside or outside the foveal region (e.g., contributing to pixels inside or outside the foveal region). In one embodiment, the foveal fragment shader 860 is configured to determine which pixels are located in the high-resolution foveal region. For example, the foveal fragment shader 860 can use the predicted landing points and / or gaze tracking information provided to the GPU memory 840 during a late update operation to determine the pixels of the foveal region and accordingly determine the fragments corresponding to those pixels. In this way, the foveal fragment shader 430 performs the shading operations as described above based on whether the fragment is within the foveal region or the peripheral region. The shading operations are used to process fragments located within the foveal region of the displayed image at high resolution without considering processing efficiency in order to obtain detailed texture and color values for the fragments within the foveal region. On the other hand, the foveal fragment processor 430 performs the shading operations on fragments located within the peripheral region while considering processing efficiency in order to process the fragments with sufficient detail with a minimum number of operations (such as providing sufficient contrast). The output of the fragment processor 430 includes the processed fragments (e.g., texture and shading information including shadows) and is passed to the next stage of the rendering pipeline 800.
[0157] The output merger component 870 calculates the characteristics of each pixel based on the fragments that compose and / or affect each corresponding pixel. That is, the fragments of all primitives in the 3D game world are combined into 2D color pixels for display. For example, the fragments contributing to the texture and shading information of the corresponding pixel are combined to output the final color value of the pixel that is passed to the next stage in the rendering pipeline 800. The output merger component 870 can perform an optional blending of values between the fragments and / or pixels determined by the foveal fragment shader 860. The output merger component 870 can also perform additional operations such as clipping (identifying and ignoring fragments outside the viewing frustum) and culling (ignoring fragments occluded by closer objects) to the viewing position.
[0158] The pixel data (e.g., color value) of each pixel in the display 1210 is stored in the frame buffer 880. When displaying the corresponding image of the scene, these values are scanned into the corresponding pixels. In particular, the display reads the color values for each pixel row by row, from left to right or from right to left, from top to bottom or from bottom to top, or in any other pattern from the frame buffer and illuminates the pixels with these pixel values when displaying the image.
[0159] Figure 9A A rendering pipeline 900A that receives and uses gaze tracking information when generating video frames during the execution of an application according to an embodiment of the present disclosure is shown, where the rendering pipeline does not implement a late update of the gaze tracking information and does not provide saccade prediction. According to an embodiment of the present disclosure, Figure 9ASimilar to Figure 5A rendering pipeline 501, both show how frame updates are slower than eye movements, such that the images shown during and after a saccade are blurry to the user.
[0160] Specifically, rendering pipeline 900A illustrates the general process for rendering an image (e.g., a 3D polygon rendering process), and is configured to perform foveated rendering based on current eye-tracking information. Rendering pipeline 900A includes CPU 1202 and GPU 1216, as well as a memory (e.g., vertex, index, depth, and frame buffers) accessible by both. Rendering pipeline 900A performs functions similar to rendering pipeline 800, including outputting corresponding pixel data (e.g., color information) for each pixel in a display (e.g., an HMD), where the color information may represent texture and shading (e.g., color, shadows, etc.).
[0161] The rendering pipeline operates at a specific frequency, where each period corresponding to the frequency can be defined as a frame period. For example, for an operating frequency of 120 Hz, the frame period is 8.3 ms. Thus, in the rendering pipeline, the pipeline sequence includes sequential operations performed by the CPU and GPU in consecutive frame periods before scanning a video frame out of the frame buffer to the display. Figure 9A Two pipeline sequences 901 and 902 are shown. For illustrative purposes, only two pipeline sequences are shown.
[0162] The rendering pipeline receives eye-tracking information from, for example, eye-tracking system 1220. As shown, eye-tracking system 1220 presents to the CPU at the start of each frame period of each rendering sequence 901 and 902. In one embodiment, eye-tracking system 1220 operates at the same frequency as that used by rendering pipeline 900A. Thus, the line-of-sight direction of the user's eye can be updated in each frame period. In other embodiments, eye-tracking system 1220 operates at a frequency different from that used by rendering pipeline 900A.
[0163] The gaze tracking information can be used to determine the foveal region to be rendered at high resolution. For example, in pipeline sequence 901, the gaze tracking information is presented to CPU 1202 in frame period 1. The gaze tracking information can include a vector X1 corresponding to the direction of the line of sight of the eye relative to the display. The gaze tracking information may have been collected in the previous frame period. Also, in frame period 1, CPU 1202 can perform a physical simulation on the object and pass the polygon primitives along with the drawing instructions to GPU 1216. Thus, as described above, in the second frame period, GPU 1216 typically performs primitive assembly to generate the rendered frame. In addition, GPU 1216 is capable of providing foveal rendering based on the gaze tracking information. That is, the GPU can render a video frame having a foveal region corresponding to the line of sight direction passed in frame period 1. The non-foveal region is rendered at a low resolution. In frame period 3, the video frame (F3) is scanned out 910 to the display.
[0164] Also, in pipeline sequence 902, the gaze tracking information is presented to CPU 1202 in frame period 2. The gaze tracking information can include a vector X2 corresponding to the direction of the line of sight of the eye relative to the display. Also, in frame period 2, CPU 1202 can perform a physical simulation on the object and pass the polygon primitives along with the drawing instructions to GPU 1216. Thus, as described above, in frame period 3, GPU 1216 typically performs primitive assembly to generate the rendered frame. In addition, GPU 1216 is capable of providing foveal rendering based on the gaze tracking information. That is, the GPU can render a video frame having a foveal region corresponding to the line of sight direction passed in frame period 2. The non-foveal region is rendered at a low resolution. In frame period 4, the video frame (F4) is scanned out 910 to the display.
[0165] Thus, in Figure 9A without saccade prediction and without late update, by the time the corresponding video frame is scanned out, the gaze tracking information for rendering the foveal region of the video frame may be outdated by 2 to 4 frame periods (or 16 ms to 32 ms). That is, eye movements may be faster than the time to render the video frame, and thus when the corresponding video frame is displayed, the foveal region is not aligned with the line of sight direction. This problem is more prominent at even lower operating frequencies. For example, a rendering pipeline operating at 60 Hz (a frame period of 16 ms) will still have gaze tracking information that is outdated by 2 to 4 frame periods, but the time for these periods is doubled, resulting in a latency range of 32 ms to 64 ms, after which the foveal region is aligned with the line of sight direction.
[0166] Figures 9B to 9CShows a prediction of a landing point corresponding to the orientation of the eyes of a user viewing a display (e.g., an HMD) during or at the end of a saccade. In particular, Figure 9B Is a flowchart showing steps in a method according to an embodiment of the present disclosure for updating a rendering pipeline by predicting a landing point on a display of an HMD, where the landing point corresponds to the orientation of the eyes of a user viewing the display during or at the end of a saccade, and where the predicted landing point is used by a GPU to render a high-resolution foveal region centered on the landing point in a corresponding video frame. Figure 9C Shows a rendering pipeline 900C that receives and uses eye-tracking information during the generation of video frames during the execution of an application according to an embodiment of the present disclosure, where a landing point corresponding to the predicted line-of-sight direction and / or orientation of the eyes of a user viewing the HMD during or at the end of a saccade is predicted on the HMD, and where the predicted landing point is used by the GPU to render a high-resolution foveal region centered on the landing point in a corresponding video frame.
[0167] The rendering pipeline 900C illustrates the general process for rendering an image and is configured to perform foveal rendering based on saccade prediction. That is, the rendering pipeline 900C provides saccade prediction without late updates. The rendering pipeline 900C includes a CPU 1202 and a GPU 1216, as well as a memory (e.g., vertex, index, depth, and frame buffers) accessible to both. The rendering pipeline 900C performs functions similar to those of the rendering pipeline 800, including outputting corresponding pixel data (e.g., color information) for each pixel in a display (e.g., an HMD), where the color information may represent texture and shading (e.g., color, shadows, etc.).
[0168] A rendering pipeline receives eye-tracking information from, for example, a game tracking system 1220. As shown, the eye-tracking information is presented to the CPU at the start of each frame period, as previously described. The eye-tracking information is used to predict a landing point on a display (e.g., an HMD) that the user's eyes are pointing to during and / or at the end of a saccade. The prediction of the landing point has been previously described at least in part in Figure 7 which is described.
[0169] At 915, the method includes executing an application on a CPU during a first frame period (e.g., frame period 1) to generate scene primitives of a first video frame. For example, the rendering pipeline 900C may perform physical emulation on objects and pass polygon primitives along with draw instructions to the GPU 1216. The frame period corresponds to the operating frequency of the rendering pipeline, which is configured to perform sequential operations by the CPU 1202 and the GPU 1216 in successive frame periods before scanning out the corresponding video frame to a display. In one embodiment, the frequency of the gaze tracking system 1220 is the same as that of the rendering pipeline 900C, but in other embodiments the frequencies are different.
[0170] At 920, the method includes receiving, at the CPU during the first frame period, gaze tracking information for the eyes of a user who is experiencing a saccade. For example, the information may be collected by the gaze tracking system 1220 in a previous frame period. As shown, the gaze tracking information is presented to the CPU at the start of frame period 1 and may include a vector X1 corresponding to the direction of the user's gaze relative to the display.
[0171] At 930, the method includes predicting, at the CPU during frame period 1 based on the gaze tracking information, a landing point on the display (e.g., HMD) corresponding to the direction of the eyes of the user viewing the display (e.g., vector X F-1 ). For greater accuracy, a history of the gaze tracking information is used to predict a landing point on the display corresponding to the predicted gaze direction of the user's eyes during or at the end of the saccade (e.g., vector X F-1 ), as previously described in Figure 7 . The prediction may include a predicted time when the user's gaze is directed at the predicted landing point.
[0172] At 940, the method includes transmitting the predicted landing point corresponding to the predicted gaze direction of the user's eyes (e.g., vector X F-1 ) to a buffer accessible by the GPU. In this way, the predicted landing point (corresponding to vector X F-1 ) can be used in the pipeline sequence 911. In particular, the method includes, at 950, performing one or more shader operations in the GPU 1216 during frame period 2 to generate pixel data for the pixels of the display based on the scene primitives of the first video frame and based on the predicted landing point (corresponding to vector X F-1 ). The pixel data includes at least color and texture information, where the pixel data is stored in a frame buffer. Additionally, during frame period 2, the GPU 1216 may render the first video frame with a foveal region corresponding to the predicted gaze direction of the user's eyes (vector X F-1)The corresponding predicted landing point. The foveal region is rendered at a high resolution, and the non-foveal region is rendered at a low resolution.
[0173] At 960, the method includes scanning pixel data from the frame buffer to the display in a third frame period. As shown, the video frame (F3) is scanned out in frame period 3.
[0174] Similarly, the pipeline sequence 912 is configured for foveal rendering based on the predicted landing point. As shown, in Figure 9C frame period 2 of, the CPU 1202 executes an application to generate the scene primitives of the second video frame. For example, the CPU 1202 can perform a physical simulation on an object and pass polygon primitives along with draw instructions to the GPU 1216. As shown in frame period 2, the rendering pipeline 900C receives the eye tracking information (vector X2), and predicts at least based on the current eye tracking information at the CPU the landing point corresponding to the direction of the eyes of the user viewing the display (e.g., vector X F-2 ) of the saccade. That is, there is a mapping between vector X F-2 and the predicted landing point. For more accuracy, the history of the eye tracking information (e.g., the history collected during the saccade) is used to predict on the display the landing point corresponding to the predicted line-of-sight direction of the user's eyes during and / or at the end of the saccade. In frame period 2 or 3, the predicted landing point (corresponding to vector X F-2 ) is transferred to a buffer accessible by the GPU. In this way, the predicted landing point (corresponding to vector X F-2 ) is available to the pipeline sequence 912 in frame period 3. In particular, in frame period 3, the GPU 1216 can render the second video frame with a foveal region corresponding to the predicted landing point corresponding to the line-of-sight direction of the eyes of the user viewing the display during or at the end of the saccade (e.g., vector X F-2 ). The foveal region is rendered at a high resolution, and the non-foveal region is rendered at a low resolution. In frame period 4, the pixel data of the second video frame (e.g., F4) is scanned out.
[0175] In this way, in Figures 9B to 9CIn the case of saccade prediction without late update, even if it may take multiple cycles to generate an accurate prediction, by the time the scan outputs the corresponding video frame, the predicted landing point of the foveal region for rendering the video frame may at least keep up with the eye movement, and in some cases may be faster than the eye movement (e.g., the foveal region is waiting for the eye movement to catch up). That is, with saccade prediction, the time to render the video frame may be faster than the eye movement, and thus when the corresponding video frame is displayed, the foveal region may be aligned with the line of sight direction, or may be ahead, such that the foveal region in the corresponding video frame is ready and waiting for the eye movement to reach the predicted line of sight direction (during and / or at the end of the saccade).
[0176] By the detailed description of various modules of a computer system, a game console, an HMD, and a cloud gaming server, now according to one embodiment of the present disclosure regarding Figure 10A flowchart 1000A of Figure 10B and the rendering pipeline 1000B shown in Figures 1A to 1C a computer system 106 of Figure 2 and Figure 13 an HMD 102 of Figure 14 and a client device 1410 of
[0177] Before scanning and outputting the corresponding video frame to the display, the rendering pipeline 1000B performs sequential operations by the CPU and GPU in consecutive frame cycles. The rendering pipeline 1000B illustrates the general process for rendering an image and is configured to perform foveal rendering based on the late update of the eye tracking information. That is, the rendering pipeline 1000B provides eye tracking through late update. The rendering pipeline 1000B includes a CPU 1202 and a GPU 1216, as well as a memory (e.g., vertex, index, depth, and frame buffers) accessible by both. The rendering pipeline 1000B performs functions similar to those of the rendering pipeline 800, including outputting corresponding pixel data (e.g., color information) for each pixel in a display (e.g., an HMD), where the color information may represent textures and shadings (e.g., colors, shadows, etc.).
[0178] At 1010, the method includes executing an application on the CPU in a first frame cycle to generate the scene primitives of a first video frame. As shown in the first frame cycle, in the rendering pipeline 1000B, the CPU 1202 may perform a physical simulation on the object and pass the polygon primitives together with the drawing instructions to the GPU 1216. As described above, the frame cycle corresponds to the operating frequency of the rendering pipeline.
[0179] Moreover, as previously described, the eye-tracking information can be provided to the CPU 1302 on a per-frame basis. As shown, the eye-tracking information is presented to the CPU 1202 at the start of each frame period. For example, in frame period 1, eye-tracking information including vector X1 is provided, in frame period 2, eye-tracking information vector X2 is provided, in frame period 3, eye-tracking information vector X3 is provided, and so on. The eye-tracking information corresponds to, for example, the direction of the user's gaze relative to the display (e.g., HMD). The eye-tracking information may have been generated in a frame period that occurred prior to being passed to the CPU. Additionally, the eye-tracking information is used to determine the high-resolution foveal region. In one embodiment, the eye-tracking system 1220 provides the eye-tracking information at the same frequency as the rendering pipeline 1000B, but in other embodiments the frequencies are different. As will be described below, the eye-tracking information can be provided to the GPU for immediate use through a late update.
[0180] At 1020, the method includes receiving eye-tracking information at the CPU during a second frame period. As Figure 10B shown in frame period 2 of, the rendering pipeline 1000B receives the eye-tracking information (vector X2). Note that without any late update operations (e.g., as Figure 5A shown), the previously presented eye-tracking information (vector X1) at frame period 1 is not used in the pipeline sequence 1001 because the information may become stale during the execution of the pipeline sequence 1001. For illustrative purposes, the eye-tracking information (vector X1) can be used with a previous pipeline sequence not shown, or it can be retained without being executed.
[0181] On the other hand, the pipeline sequence 1001 having CPU operations executed in frame period 1 can utilize the late update feature of the rendering pipeline 1000B to utilize the latest eye-tracking information. In particular, at 1030, the method includes performing a late update operation by the CPU in frame period 2 by transmitting the eye-tracking information to a buffer accessible by the GPU.
[0182] Thus, instead of using the out-of-date by one frame period eye-tracking information (vector X1) received in frame period 1, pipeline sequence 1001 can perform GPU operations in frame period 2 using the latest eye-tracking information, which is vector X2 received at the CPU in the same frame period 2. In particular, at 1040, the method includes performing one or more shader operations in the GPU in frame period 2 to generate pixel data for the pixels of the display based on the scene primitives of the first video frame and based on the eye-tracking information (vector X2). The pixel data includes at least color and texture information, where the pixel data is stored in the frame buffer. In particular, in frame period 2, GPU 1216 can render the first video frame with a foveal region corresponding to the eye-tracking information (vector X1) corresponding to the measured and latest line-of-sight direction of the user. The foveal region is rendered at a high resolution, and the non-foveal region is rendered at a low resolution.
[0183] At 1050, the method includes scan-outputting the pixel data from the frame buffer to the display in a third frame period. As shown, video frame (F3) is scan-output in frame period 3.
[0184] Similarly, pipeline sequence 1002 is configured for foveal rendering based on a late update of the eye-tracking information. As shown, in frame period 2, CPU 1202 executes an application to generate the scene primitives of the second video frame. For example, CPU 1202 can perform a physical simulation on an object and pass polygon primitives along with draw instructions to GPU 1216. As Figure 10B shown in frame period 3, rendering pipeline 1000B receives the eye-tracking information (vector X3). Also, in frame period 3, the CPU performs a late update operation by transmitting the eye-tracking (vector X3) information to a buffer accessible by the GPU. In frame period 3, one or more shader operations are performed in GPU 1216 to generate pixel data for the pixels of the display based on the scene primitives of the second video frame and based on the eye-tracking information (vector X3). The pixel data includes at least color and texture information, where the pixel data is stored in the frame buffer. In particular, in frame period 3, GPU 1216 can render the second video frame with a foveal region corresponding to the eye-tracking information (vector X2) corresponding to the measured and latest line-of-sight direction of the user. The foveal region is rendered at a high resolution, and the non-foveal region is rendered at a low resolution. In frame period 4, the pixel data of the second video frame (e.g., F4) is scan-output.
[0185] Thus, in Figures 10A to 10BIn this case, through post-update, the corresponding video frame with the foveal region is scanned and output, where the foveal region corresponds to the eye-tracking information received in the previous frame period. That is, through post-update, the video frame is rendered almost as fast as the eye movement. In this way, it takes less time for the foveal region to align with the line-of-sight direction during and at the end of a saccade.
[0186] By the detailed description of various modules of a computer system, a game console, an HMD, and a cloud gaming server, now according to an embodiment of the present disclosure regarding Figure 11A flowchart 1100A and Figure 11B rendering pipeline 1100B shown in describe a method for updating information in a rendering pipeline including a CPU and a GPU by: predicting a landing point associated with the predicted line-of-sight direction of a user viewing a display (e.g., an HMD) during a saccade, and post-updating the predicted landing point to a buffer accessible by the GPU for immediate use. Flowchart 1100A and rendering pipeline 1100B can be implemented by at least Figures 1A to 1C computer system 106, system 1200, Figure 2 and Figure 13 HMD 102 of and Figure 14 client device 1410 of.
[0187] Before scanning and outputting the corresponding video frame to the display, rendering pipeline 1100B performs sequential operations by the CPU and the GPU in consecutive frame periods. Rendering pipeline 1100B illustrates the general process for rendering an image and is configured to perform foveal rendering based on post-update of saccade prediction. That is, rendering pipeline 1100B provides eye tracking through post-update of the predicted landing point, which corresponds to the predicted line-of-sight direction and / or orientation of the eyes of a user viewing a display (e.g., an HMD) during or at the end of a saccade. The prediction may include the predicted time when the user's line of sight points to the predicted landing point. The GPU uses the predicted landing point to render a high-resolution foveal region centered on the landing point in the corresponding video frame. Rendering pipeline 1000B includes CPU 1202 and GPU 1216, and a memory (e.g., vertex, index, depth, and frame buffers) accessible by both. Rendering pipeline 1000B performs functions similar to those of rendering pipeline 800, including outputting corresponding pixel data (e.g., color information) for each pixel in a display (e.g., an HMD), where the color information may represent texture and shading (e.g., color, shadow, etc.).
[0188] At 1110, the method includes executing an application on a CPU in a first frame period (e.g., frame period 1) to generate scene primitives of a first video frame. As shown in frame period 1, in rendering pipeline 1000B, CPU 1202 may perform a physical simulation on an object and pass polygon primitives along with draw instructions to GPU 1216. A frame period corresponds to the operating frequency of the rendering pipeline, which is configured for sequential operations to be performed by CPU 1202 and GPU 1216 in successive frame periods before the corresponding video frame is scanned out to a display. In one embodiment, the gaze tracking system 1220 provides gaze tracking information at the same frequency as the rendering pipeline 1100B, but in other embodiments the frequencies are different.
[0189] Gaze tracking information may be received at the CPU in the first frame period (e.g., frame period 1). The gaze tracking information may have been generated by the gaze tracking system 1220 in the previous frame period. As shown, the gaze tracking information is presented to the CPU at the start of frame period 1. For example, the gaze tracking information may include a vector X1 corresponding to the direction of the user's gaze relative to the display. Additionally, in frame period 1, the CPU predicts, at least based on the current gaze tracking information, a landing point on the display (e.g., HMD) corresponding to the direction of the user's eyes (vector X F-1 ) viewing the display. Because of late updates, the predicted landing point (corresponding to vector X F-1 ) is not used in pipeline 1101 when updated current gaze tracking information becomes available, as will be described below.
[0190] At 1120, the method includes receiving, at the CPU in a second frame period, gaze tracking information of the eyes of a user who has experienced a saccade. The gaze tracking information may have been generated by the gaze tracking system 1220 in the previous frame period. As shown, the gaze tracking information is presented to the CPU at the start of frame period 2. For example, the gaze tracking information may include a vector X2 corresponding to the direction of the user's gaze relative to the display.
[0191] At 1130, the method includes predicting, at least based on the current gaze tracking information (vector X2) at the CPU in frame period 2, a landing point on the display (e.g., HMD) corresponding to the direction of the user's eyes (vector X F-2 ) viewing the display. For greater accuracy, a history of the gaze tracking information (e.g., a history collected during the saccade) is used to predict a landing point on the display corresponding to the predicted gaze direction of the user's eyes during and / or at the end of the saccade, as previously described at least in part in Figure 7 . The predicted landing point may correspond to the predicted gaze direction of the user's eyes (vector X) during and / or at the end of the saccadeF-2 )。
[0192] At 1140, the CPU 1202 performs a late update operation to transfer the predicted landing point (corresponding to vector X F-2 ) to a buffer accessible by the GPU. The transfer is completed during frame cycle 2 and is immediately available for the GPU to use. That is, the transfer occurs before the GPU 1216 begins its operation in frame cycle 2 of pipeline 1101. In this way, the updated current gaze tracking information (e.g., X2 collected in the middle of pipeline sequence 1101) is used instead of the gaze tracking information (e.g., X1) collected at the start of pipeline sequence 1101 to generate the predicted landing point.
[0193] In this way, the predicted landing point (corresponding to vector X F-2 ) can be immediately used in pipeline sequence 1101. In particular, at 1150, the method includes performing one or more shader operations in the GPU 1216 during frame cycle 2 to generate pixel data for the pixels of the display based on the scene primitives of the first video frame and based on the predicted landing point (corresponding to vector X F-2 ). The pixel data includes at least color and texture information, where the pixel data is stored in the frame buffer. In particular, during frame cycle 2, the GPU 1216 can render the first video frame with a foveal region corresponding to the predicted landing point corresponding to the predicted gaze direction (vector X F-2 ) of the user's eye during or at the end of a saccade. The foveal region is rendered at a high resolution, and the non-foveal region is rendered at a low resolution.
[0194] At 1160, the method includes scan-outputting the pixel data from the frame buffer to the display during the third frame cycle. As shown, the video frame (F3) is scan-output during frame cycle 3.
[0195] Similarly, pipeline sequence 1102 is configured for foveal rendering based on a late update of the predicted landing point. As shown, in Figure 11B 's frame cycle 2, the CPU 1202 executes an application to generate the scene primitives of the second video frame. For example, the CPU 1202 can perform a physical simulation on an object and pass polygon primitives along with draw instructions to the GPU 1216. As shown in frame cycle 2, the rendering pipeline 1100B receives the gaze tracking information (vector X2).
[0196] During frame cycle 3, at the CPU, predict the landing point corresponding to the direction of the saccade of the user's eye viewing the display (e.g., vector X F-3 ) based at least on the current gaze tracking information. That is, vector X F-3There is a mapping to the predicted landing point. For greater accuracy, a history of the eye-tracking information (e.g., the history collected during a saccade) is used to predict on the display a landing point corresponding to the predicted line-of-sight direction of the user's eyes during and / or at the end of the saccade.
[0197] Also, in frame cycle 3, the CPU performs a late update operation by transferring the predicted landing point corresponding to vector X F-3 to a buffer accessible by the GPU. In frame cycle 3, one or more shader operations are performed in GPU 1216 to generate pixel data for the pixels of the display based on the scene primitives of the second video frame and based on the predicted landing point corresponding to vector X F-3 The pixel data includes at least color and texture information, where the pixel data is stored in a frame buffer. In particular, in frame cycle 3, GPU 1216 may render the second video frame with a foveal region corresponding to the predicted landing point corresponding to the line-of-sight direction of the eyes of the user viewing the display during or at the end of the saccade (e.g., vector X F-3 ). The foveal region is rendered at a high resolution, and the non-foveal region is rendered at a low resolution. In frame cycle 4, the pixel data of the second video frame (e.g., F4) is scanned out.
[0198] Thus, in Figure 11B , through the late update of saccade prediction (e.g., the landing point), the landing point prediction can utilize the latest current eye-tracking information collected during the pipeline sequence and not necessarily at the start of the pipeline sequence. Thus, by the time the corresponding video frame is scanned out, the predicted landing point for the foveal region used to render the video frame may at least keep up with the eye movement, and in most cases will be faster than the eye movement (e.g., the displayed foveal region is waiting for the eye movement to catch up). That is, through the late update of saccade prediction, the time to render the video frame may be faster than the eye movement, and thus when the corresponding video frame is displayed, the foveal region may be aligned with the line-of-sight direction, or may be ahead, such that the foveal region in the corresponding video frame is ready and waiting for the eye movement to reach the predicted line-of-sight direction (during and / or at the end of the saccade).
[0199] In one embodiment, when rendering a frame for display on an HMD, additional measures may be taken to reduce power consumption. In particular, during a saccade, the user may not be able to view the intermediate frames being rendered and displayed because the eye movement may be too fast. Thus, based on the predicted time to the landing point (e.g., during or at the end of the saccade), the rendering of at least one intermediate video frame that occurs before the video frame corresponding to the predicted landing point is displayed may be terminated to save the computational resources that would otherwise be used for rendering.
[0200] In yet another embodiment, when rendering a frame for display on an HMD, another measure can be taken to reduce power consumption. In particular, since the user may not be able to view the intermediate frames rendered and displayed during a saccade, the entire video frame may be rendered at a lower resolution or a low resolution. That is, the foveal region is not rendered for those intermediate frames. In other words, based on the predicted time to the landing point (e.g., during or at the end of a saccade), at least one intermediate video frame that occurs before displaying the video frame corresponding to the predicted landing point is rendered at a low resolution.
[0201] Figure 12 The components of an example apparatus 1200 that can be used to implement aspects of various embodiments of the present disclosure are shown. For example, Figure 12 An exemplary hardware system suitable for implementing an apparatus according to one embodiment is shown, the apparatus being configured to predict and post-update a target landing point on a display such that the movement of the user's eyes is consistent with the rendering of the foveal region at the updated target landing point on the display. The example apparatus 1200 is generally described because the prediction of the landing point can be performed in the context of an HMD as well as a more traditional display. The block diagram shows an apparatus 1200 that can be incorporated or can be a personal computer, a video game console, a personal digital assistant, or other digital device suitable for practicing the embodiments of the present disclosure. The apparatus 1200 includes a central processing unit (CPU) 1202 for running software applications and optionally an operating system. The CPU 1202 can include one or more homogeneous or heterogeneous processing cores. For example, the CPU 1202 is one or more general-purpose microprocessors having one or more processing cores. Other embodiments can be implemented using one or more CPUs having a microprocessor architecture that is particularly suitable for highly parallel and computationally intensive applications, such as media and interactive entertainment applications, or applications configured to provide a predicted landing point on a display that is associated with the line-of-sight direction of a user's eyes during and / or at the end of a saccade defined in association with a user viewing the display, as previously described.
[0202] Memory 1204 stores applications and data for use by CPU 1202. Storage device 1206 provides non-volatile storage and other computer-readable media for the applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray Disc, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 1208 conveys user input from one or more users to device 1200. Examples of user input devices may include keyboards, mice, joysticks, touchpads, touchscreens, static or video recorders / cameras, tracking devices for recognizing gestures, and / or microphones. Network interface 1214 allows device 1200 to communicate with other computer systems via an electronic communication network, and may include wired or wireless communication on local area networks and wide area networks such as the Internet. Audio processor 1212 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 1202, memory 1204, and / or storage device 1206. Components of device 1200, including CPU 1202, memory 1204, data storage device 1206, user input device 1208, network interface 1210, and audio processor 1212, are connected via one or more data buses 1222.
[0203] Graphics subsystem 1214 is also connected to data bus 1222 and components of device 1200. Graphics subsystem 1214 includes graphics processing unit (GPU) 1216 and graphics memory 1218. Graphics memory 1218 includes display memory (e.g., frame buffer) for storing pixel data of each pixel of an output image. Graphics memory 1218 may be integrated with GPU 1216 in the same device, connected to GPU 1216 as a separate device, and / or implemented within memory 1204. Pixel data may be provided directly from CPU 1202 to graphics memory 1218. Alternatively, CPU 1202 provides data and / or instructions that define a desired output image to GPU 1216, and GPU 1216 generates pixel data of one or more output images from the data and / or instructions. Data and / or instructions that define a desired output image may be stored in memory 1204 and / or graphics memory 1218. In an embodiment, GPU 1216 includes 3D rendering capabilities for generating pixel data of an output image from instructions and data that define the geometry, lighting, shading, texture, motion, and / or camera parameters of a scene. GPU 1216 may also include one or more programmable execution units capable of executing shader programs.
[0204] The graphics subsystem 1214 periodically outputs pixel data of an image from the graphics memory 1218 for display on the display device 1210 or projection by the projection system 1240. The display device 1210 can be any device capable of displaying visual information in response to a signal from the device 1200, including CRT, LCD, plasma, and OLED displays. For example, the device 1200 can provide an analog or digital signal to the display device 1210.
[0205] Additionally, the device 1200 includes a gaze tracking system 1220, which includes a gaze tracking sensor 265 and a light source (e.g., emitting invisible infrared light), as described above.
[0206] It should be understood that the embodiments described herein can be executed on any type of client device. In some embodiments, the client device is a head-mounted display (HMD) or a projection system. Figure 13 is a diagram showing components of a head-mounted display 102 according to an embodiment of the present disclosure. The HMD 102 can be configured to predict a landing point associated with the line-of-sight direction of a user's eyes during and / or at the end of a saccade defined in relation to a user viewing the display on the HMD, and provide the predicted landing point to the GPU during a later update operation.
[0207] The head-mounted display 102 includes a processor 1300 for executing program instructions. A memory 1302 is provided for storage purposes and can include both volatile and non-volatile memory. A display 1304 is included, which provides a visual interface that the user can view. A battery 1306 is provided as a power source for the head-mounted display 102. The motion detection module 1308 can include any one of various types of motion-sensitive hardware, such as a magnetometer 1310A, an accelerometer 1312, and a gyroscope 1314.
[0208] An accelerometer is a device for measuring acceleration and the reaction force of gravity sensing. Uniaxial and multi-axial models can be used to detect the magnitude and direction of acceleration in different directions. Accelerometers are used to sense tilt, vibration, and shock. In one embodiment, three accelerometers 1312 are used to provide the direction of gravity, which gives an absolute reference for two angles (world space pitch angle and world space roll angle).
[0209] A magnetometer measures the strength and direction of the magnetic field near the head-mounted display. In one embodiment, three magnetometers 1310A are used within the head-mounted display to ensure an absolute reference for the world space yaw angle. In one embodiment, the magnetometer is designed to span the Earth's magnetic field of ±80 microteslas. The magnetometer is affected by metals and provides a yaw angle measurement that varies monotonically with the actual yaw angle. The magnetic field may be distorted due to metals in the environment, which results in a distorted yaw angle measurement. If necessary, information from other sensors such as gyroscopes or cameras can be used to calibrate this distortion. In one embodiment, an accelerometer 1312 is used with the magnetometer 1310A to obtain the tilt angle and azimuth angle of the head-mounted display 102.
[0210] A gyroscope is a device for measuring or maintaining orientation based on the principle of angular momentum. In one embodiment, three gyroscopes 1314 provide information about motion across the respective axes (x, y, and z) based on inertial sensing. The gyroscope helps detect rapid rotation. However, the gyroscope can drift over time without an absolute reference. This requires periodically resetting the gyroscope, which can be done using other available information such as position / orientation determination based on visual tracking of an object, accelerometers, magnetometers, etc.
[0211] A camera 1316 is provided for capturing images and image streams of the real environment. More than one camera may be included in the head-mounted display 102, including a rear-facing camera (a camera that points away from the user when the user is viewing the display of the head-mounted display 102), and a front-facing camera (a camera that points towards the user when the user is viewing the display of the head-mounted display 102). Additionally, a depth camera 1318 may be included in the head-mounted display 102 for sensing depth information of objects in the real environment.
[0212] In one embodiment, a camera integrated on the front of the HMD can be used to provide warnings about safety. For example, if the user is approaching a wall or an object, the user can be warned. In one embodiment, a contour map of the physical objects in the room can be provided to the user to warn them of their presence. The contour can be, for example, an overlay in the virtual environment. In some embodiments, a view of a reference marker superimposed on, for example, the floor can be provided to the HMD user. For example, the marker can provide a reference to the user as to where the center of the room is where the user is playing a game. For example, this can provide the user with visual information as to where the user should move to avoid hitting the walls or other objects in the room. Haptic warnings and / or audio warnings can also be provided to the user to provide increased safety when the user is wearing the HMD and playing a game or navigating content with the HMD.
[0213] The head-mounted display 102 includes a speaker 252 for providing audio output. Moreover, a microphone 251 may be included for capturing audio from the real environment, including sounds from the surrounding environment, speech uttered by the user, and the like. The head-mounted display 102 includes a haptic feedback module 281 for providing haptic feedback to the user. In one embodiment, the haptic feedback module 281 is capable of causing movement and / or vibration of the head-mounted display 102 to provide haptic feedback to the user.
[0214] An LED 1326 is provided as a visual indicator of the status of the head-mounted display 102. For example, the LED may indicate battery level, power-on, etc. A card reader 1328 is provided to enable the head-mounted display 102 to read information from and write information to a memory card. A USB interface 1330 is included as an example of an interface for enabling connection to a peripheral device or connection to other devices such as other portable devices, computers, etc. In various embodiments of the head-mounted display 102, any one of various types of interfaces may be included to achieve greater connectivity of the head-mounted display 102.
[0215] A Wi-Fi module 1332 is included to enable connection to the Internet via wireless networking technology. Moreover, the head-mounted display 102 includes a Bluetooth module 1334 for enabling wireless connection to other devices. A communication link 1336 may also be included for connecting to other devices. In one embodiment, the communication link 1336 uses infrared transmission for wireless communication. In other embodiments, the communication link 1336 may use any one of various wireless or wired transmission protocols to communicate with other devices.
[0216] An input button / sensor 1338 is included to provide an input interface for the user. Any one of various types of input interfaces may be included, such as buttons, touch pads, joysticks, trackballs, etc. An ultrasonic communication module 1340 may be included in the head-mounted display 102 to facilitate communication with other devices via ultrasonic technology.
[0217] A biosensor 1342 is included to enable detection of physiological data from the user. In one embodiment, the biosensor 1342 includes one or more dry electrodes for detecting the user's bioelectrical signals through the user's skin.
[0218] An optoelectronic sensor 1344 is included to respond to signals from a transmitter (e.g., an infrared base station) placed in a three-dimensional physical environment. The game console analyzes information from the optoelectronic sensor 1344 and the transmitter to determine position and orientation information related to the head-mounted display 102.
[0219] In addition, a gaze tracking system 1320 is included and configured to enable tracking of a user's gaze. For example, the system 1320 can include a gaze tracking camera (e.g., a sensor) to capture an image of the user's eyes, and then analyze the image to determine the direction of the user's gaze. In one embodiment, information about the direction of the user's gaze can be used to influence video rendering and / or predict a landing point on the display where the user's gaze will point during or at the end of a saccade. Similarly, video rendering in the direction of the gaze can be prioritized or emphasized, such as by providing more detail and higher resolution through foveated rendering, higher resolution particle system effects displayed in the foveal region, lower resolution particle system effects displayed outside the foveal region, or faster updates in the area the user is viewing.
[0220] The foregoing components of the head-mounted display 102 are described only as exemplary components that may be included in the head-mounted display 102. In various embodiments of the present disclosure, the head-mounted display 102 may include or may not include some of the various above-described components. Embodiments of the head-mounted display 102 may additionally include other components that are currently not described but are known in the art for purposes of facilitating aspects of the present disclosure as described herein.
[0221] Those skilled in the art will appreciate that in various embodiments of the present disclosure, the above-described head-mounted device can be used in combination with interactive applications displayed on a display to provide various interactive functions. The exemplary embodiments described herein are provided by way of example only and not by way of limitation.
[0222] It should be noted that access services delivered over a wide geographic area, such as providing access to games of the current embodiments, often use cloud computing. Cloud computing is a style of computing in which dynamically scalable and usually virtualized resources are provided as a service over the Internet. Users do not need to be experts in the technical infrastructure of the "cloud" that supports them. Cloud computing can be divided into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services typically provide general applications (such as video games) accessible online from a web browser, while the software and data are stored on servers in the cloud. Based on the way the Internet is depicted in computer network diagrams, the term cloud is used as a metaphor for the Internet and is an abstract concept of the complex infrastructure it hides.
[0223] The game client uses a game processing server (GPS) (or simply referred to as the "game server") to play single-player and multi-player video games. Most video games conducted over the Internet are run via a connection to a game server. Typically, the game uses a dedicated server application that collects data from players and distributes it to other players. This is more effective and efficient than a peer-to-peer arrangement, but it requires a separate server to host the server application. In another embodiment, the GPS establishes communication between players and their respective gaming devices to exchange information without relying on a centralized GPS.
[0224] A dedicated GPS is a server that runs independently of the client. Such a server typically runs on dedicated hardware located within a data center, thus providing more bandwidth and dedicated processing capabilities. For most PC-based multi-player games, a dedicated server is the preferred method for hosting the game server. Large multi-player online games run on dedicated servers, which are typically hosted by the software company that owns the game title, thus allowing them to control and update the content.
[0225] The user accesses the remote service using a client device that includes at least a CPU, a display, and I / O. The client device can be a PC, a mobile phone, a netbook, a PDA, etc. In one embodiment, the network identification executed on the game server determines the type of device used by the client and adjusts the communication method employed. In other cases, the client device uses a standard communication method (such as html) to access the application on the game server via the Internet.
[0226] Embodiments of the present disclosure can be practiced with a variety of computer system configurations including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. The present disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a wired or wireless network.
[0227] It should be understood that a given video game or game application can be developed for a specific platform and a specific associated controller device. However, when such a game is available through the game cloud system presented herein, the user can use a different controller device to access the video game. For example, the game may be developed for a game console and its associated controller, while the user may be accessing the cloud-based version of the game from a personal computer using a keyboard and mouse. In such a case, the input parameter configuration can define the mapping of the input generated from the controller device available to the user (in this case, the keyboard and mouse) to the input acceptable for executing the video game.
[0228] In another example, a user may access a cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen-driven device. In such a case, the client device and the controller device are integrated together in the same device, where input is provided via detected touchscreen inputs / gestures. For such a device, the input parameter configuration may define specific touchscreen inputs corresponding to game inputs for a video game. For example, during the operation of a video game, buttons, directional keys, or other types of input elements may be displayed or overlaid to indicate locations on the touchscreen where the user may touch to generate game inputs. Gestures such as a swipe in a particular direction or a particular touch movement may also be detected as game inputs. In one implementation, a guide indicating how to provide input via the touchscreen in order to play the game may be provided to the user (e.g., before starting the gameplay of a video game) so that the user can get used to operating the controls on the touchscreen.
[0229] In some implementations, the client device serves as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection to transfer input from the controller device to the client device. Subsequently, the client device may process these inputs and then transmit the input data to the cloud gaming server via a network (e.g., a network accessed via a local networking device such as a router). However, in other implementations, the controller itself may be a networked device, having the ability to directly transmit inputs to the cloud gaming server via the network without first transmitting such inputs through the client device. For example, the controller may be connected to a local networking device (such as the aforementioned router) to send data to and receive data from the cloud gaming server. Thus, although it may still be required for the client device to receive the video output from a cloud-based video game and render it on a local display, input latency can be reduced by allowing the controller to bypass the client device and directly send inputs to the cloud gaming server via the network.
[0230] In one embodiment, the networked controller and the client device can be configured to send certain types of input directly from the controller to the cloud gaming server and other types of input via the client device. For example, inputs whose detection does not rely on any additional hardware or processing outside of the controller itself can bypass the client device and be sent directly from the controller to the cloud gaming server via the network. Such inputs can include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), etc. However, inputs that utilize additional hardware or require processing by the client device can be sent to the cloud gaming server by the client device. These may include video or audio captured from the gaming environment, which may be processed by the client device before being sent to the cloud gaming server. Additionally, inputs from the motion detection hardware of the controller can be processed by the client device in conjunction with the captured video to detect the position and motion of the controller, which will then be transmitted by the client device to the cloud gaming server. It should be understood that the controller device according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.
[0231] In particular, Figure 14 is a block diagram of a gaming system 1400 according to various embodiments of the present disclosure. The gaming system 1400 is configured to provide a video stream to one or more clients 1410 via a network 1415, such as in a single-player mode or a multi-player mode. The gaming system 1400 generally includes a video server system 1420 and an optional game server 1425. The video server system 1420 is configured to provide a video stream to one or more clients 1410 with a minimum quality of service. For example, the video server system 1420 can receive game commands that change the state of a video game or the perspective within a video game and provide an updated video stream reflecting such state changes to the clients 1410 with a minimum latency. The video server system 1420 can be configured to provide video streams in a variety of alternative video formats, including formats that have not yet been defined. Additionally, the video stream can include video frames configured to be presented to the user at a variety of frame rates. Typical frame rates are 30 frames per second, 80 frames per second, and 820 frames per second. However, higher or lower frame rates are included in alternative embodiments of the present disclosure.
[0232] The client 1410 (referred to herein individually as 1410A, 1410B, etc.) can include a head-mounted display, a terminal, a personal computer, a game console, a tablet computer, a telephone, a set-top box, a kiosk, a wireless device, a numeric keypad, a stand-alone device, a handheld gaming device, etc. Generally, the client 1410 is configured to receive an encoded (e.g., compressed) video stream, decode the video stream, and present the resulting video to a user, such as a gamer. The process of receiving the encoded video stream and / or decoding the video stream typically includes storing the individual video frames in a receive buffer of the client. The video stream can be presented to the user on a display integrated with the client 1410 or on a separate device such as a monitor or a television. The client 1410 is optionally configured to support more than one gamer. For example, a game console can be configured to support two, three, four, or more simultaneous players. Each of these players can receive a separate video stream, or a single video stream can include frame regions generated specifically for each player, such as frame regions generated based on each player's perspective. The client 1410 is optionally geographically dispersed. The number of clients included in the game system 1400 can vary widely, ranging from one or two to thousands, tens of thousands, or more. As used herein, the term "gamer" is used to refer to a person who plays a game, and the term "gaming device" is used to refer to a device used for playing a game. In some embodiments, a gaming device can refer to multiple computing devices that cooperate to deliver a gaming experience to a user. For example, a game console and an HMD can cooperate with a video server system 1420 to deliver a game viewed through the HMD. In one embodiment, the game console receives a video stream from the video server system 1420, and the game console forwards the video stream or an update to the video stream to the HMD for rendering.
[0233] The client 1410 is configured to receive a video stream via the network 1415. The network 1415 can be any type of communication network, including a telephone network, the Internet, a wireless network, a powerline network, a local area network, a wide area network, a private network, etc. In a typical embodiment, the video stream is transported via a standard protocol such as TCP / IP or UDP / IP. Alternatively, the video stream is transported via a proprietary standard.
[0234] A typical example of the client 1410 is a personal computer that includes a processor, non-volatile memory, a display, decoding logic, network communication capabilities, and an input device. The decoding logic can include hardware, firmware, and / or software stored on a computer-readable medium. Systems for decoding (and encoding) video streams are well known in the art and vary according to the particular encoding scheme used.
[0235] Client 1410 may also (but is not required to) include a system configured to modify received video. For example, the client may be configured to perform further rendering, overlay one video image on another, crop video images, etc. For example, client 1410 may be configured to receive various types of video frames, such as I-frames, P-frames, and B-frames, and process these frames as images for display to the user. In some embodiments, members of client 1410 are configured to perform further rendering, shading, conversion to 3D, or similar operations on the video stream. Members of client 1410 are optionally configured to receive more than one audio or video stream. The input devices of client 1410 may include, for example, a single-handed game controller, a two-handed game controller, a gesture recognition system, a gaze recognition system, a voice recognition system, a keyboard, a joystick, a pointing device, a force feedback device, a motion and / or position sensing device, a mouse, a touch screen, a neural interface, a camera, an input device yet to be developed, etc.
[0236] The video stream (and optionally the audio stream) received by client 1410 is generated and provided by video server system 1420. As further described elsewhere herein, the video stream includes video frames (while the audio stream includes audio frames). The video frames are configured (e.g., they include pixel information in an appropriate data structure) to meaningfully contribute to the images shown to the user. As used herein, the term "video frame" is used to refer to a frame that mainly includes information configured to contribute to (e.g., implement) the images shown to the user. Most of the teachings herein regarding "video frames" can also be applied to "audio frames".
[0237] Client 1410 is generally configured to receive input from the user. These inputs may include game commands configured to change the state of a video game or otherwise affect the gaming process. Game commands may be received using the input devices and / or may be automatically generated by computational instructions executed on client 1410. The received game commands are transmitted from client 1410 to video server system 1420 and / or game server 1425 via network 1415. For example, in some embodiments, the game commands are transmitted to game server 1425 via video server system 1420. In some embodiments, separate copies of the game commands are transmitted from client 1410 to game server 1425 and video server system 1420. The transmission of the game commands optionally depends on the identity of the command. Optionally, the game commands are transmitted from client 1410A via a different route or communication channel used to provide the audio or video stream to client 1410A.
[0238] The game server 1425 is optionally operated by an entity different from the video server system 1420. For example, the game server 1425 can be operated by the publisher of a multi-player game. In this example, the video server system 1420 is optionally regarded as a client by the game server 1425 and is optionally configured to be a prior art client that executes a prior art game engine from the perspective of the game server 1425. The communication between the video server system 1420 and the game server 1425 optionally occurs via the network 1415. Thus, the game server 1425 can be a prior art multi-player game server that sends game state information to multiple clients, and one of the multiple clients is the video server system 1420. The video server system 1420 can be configured to communicate with multiple instances of the game server 1425 simultaneously. For example, the video server system 1420 can be configured to provide multiple different video games to different users. Each of these different video games can be supported by a different game server 1425 and / or published by a different entity. In some embodiments, several geographically dispersed instances of the video server system 1420 are configured to provide game videos to multiple different users. Each of these instances of the video server system 1420 can communicate with the same instance of the game server 1425. The communication between the video server system 1420 and one or more game servers 1425 optionally occurs via a dedicated communication channel. For example, the video server system 1420 can be connected to the game server 1425 via a high-bandwidth channel dedicated to the communication between the two systems.
[0239] The video server system 1420 includes at least a video source 1430, an I / O device 1445, a processor 1450, and a non-transitory storage device 1455. The video server system 1420 can include one computing device or be distributed among multiple computing devices. These computing devices are optionally connected via a communication system such as a local area network.
[0240] The video source 1430 is configured to provide a video stream, such as a streaming video or a series of video frames that form a moving picture. In some embodiments, the video source 1430 includes a video game engine and rendering logic. The video game engine is configured to receive game commands from a player and maintain a copy of the video game state based on the received commands. The game state includes the position of objects in the game environment and typically includes the perspective. The game state can also include the properties, images, colors, and / or textures of the objects.
[0241] The game state is typically maintained based on game rules and game commands such as move, turn, attack, set focus, interact, use, etc. Optionally, a part of the game engine is set within the game server 1425. The game server 1425 can maintain a copy of the game state based on game commands received from multiple players using geographically dispersed clients. In these cases, the game state is provided by the game server 1425 to the video source 1430, where a copy of the game state is stored and rendering is performed. The game server 1425 can receive game commands directly from the clients 1410 via the network 1415, and / or can receive game commands via the video server system 1420.
[0242] The video source 1430 typically includes rendering logic, such as hardware, firmware, and / or software stored on a computer-readable medium such as the storage device 1455. The rendering logic is configured to create video frames of a video stream based on the game state. All or part of the rendering logic is optionally set within a graphics processing unit (GPU). The rendering logic typically includes processing stages that are configured to determine three-dimensional spatial relationships between objects and / or apply appropriate textures, etc., based on the game state and the perspective. The rendering logic produces raw video, which is then typically encoded before being transmitted to the clients 1410. For example, the raw video can be encoded according to: Adobe standards,.wav, H.264, H.263, On2, VP6, VC-1, WMA, Huffyuv, Lagarith, MPG-x, Xvid, FFmpeg, x264, VP6-8, realvideo, mp3, etc. The encoding process produces a video stream, which is optionally encapsulated for delivery to a decoder on a remote device. The video stream is characterized by a frame size and a frame rate. Typical frame sizes include 800×600, 1280×720 (e.g., 720p), 1024×768, but any other frame size can be used. The frame rate is the number of video frames per second. The video stream can include different types of video frames. For example, the H.264 standard includes "P" frames and "I" frames. An I frame includes information for refreshing all macroblocks / pixels on a display device, while a P frame includes information for refreshing a subset thereof. The data size of a P frame is typically smaller than that of an I frame. As used herein, the term "frame size" refers to the number of pixels within a frame. The term "frame data size" is used to refer to the number of bytes required to store a frame.
[0243] In an alternative embodiment, video source 1430 includes a video recording device, such as a camera. The camera can be used to generate delayed video or live video, which can be included in the video stream of a computer game. The resulting video stream optionally includes both rendered images and images recorded using a still camera or video camera. Video source 1430 may also include a storage device configured to store previously recorded video to be included in the video stream. Video source 1430 may also include a motion or positioning sensing device configured to detect the motion or position of an object (e.g., a person), and logic configured to determine a game state or generate video based on the detected motion and / or position.
[0244] Video source 1430 is optionally configured to provide overlays configured to be placed on top of other video. For example, these overlays can include command interfaces, login instructions, messages to game players, images of other game players, video feeds of other game players (e.g., webcam video). In embodiments of client 1410A that include a touchscreen interface or a sight detection interface, the overlays can include virtual keyboards, joysticks, touchpads, etc. In one example of an overlay, a player's voice is overlaid on the audio stream. Video source 1430 optionally also includes one or more audio sources.
[0245] In embodiments where video server system 1420 is configured to maintain a game state based on input from more than one player, each player can have a different perspective including a viewing position and orientation. Video source 1430 is optionally configured to provide a separate video stream for each player based on each player's perspective. Additionally, video source 1430 can be configured to provide different frame sizes, frame data sizes, and / or encodings to each of clients 1410. Video source 1430 is optionally configured to provide 3D video.
[0246] I / O device 1445 is configured to enable video server system 1420 to send and / or receive information, such as video, commands, information requests, game states, sight information, device motion, device position, user motion, client identification, player identification, game commands, security information, audio, etc. I / O device 1445 generally includes communication hardware, such as a network card or modem. I / O device 1445 is configured to communicate with game server 1425, network 1415, and / or client 1410.
[0247] The processor 1450 is configured to execute logic, such as software, included within various components of the video server system 1420 discussed herein. For example, the processor 1450 may be programmed with software instructions to perform the functions of the video source 1430, the game server 1425, and / or the client qualifier 1460. The video server system 1420 optionally includes more than one instance of the processor 1450. The processor 1450 may also be programmed with software instructions to execute commands received by the video server system 1420 or to coordinate the operation of the various elements of the game system 1400 discussed herein. The processor 1450 may include one or more hardware devices. The processor 1450 is an electronic processor.
[0248] The storage device 1455 includes non-transitory analog and / or digital storage devices. For example, the storage device 1455 may include an analog storage device configured to store video frames. The storage device 1455 may include a computer-readable digital storage device, such as a hard disk drive, an optical disc drive, or a solid-state storage device. The storage device 1455 is configured to store (e.g., via an appropriate data structure or file system) video frames, artificial frames, video streams including both video frames and artificial frames, audio frames, audio streams, etc. The storage device 1455 is optionally distributed among multiple devices. In some embodiments, the storage device 1455 is configured to store software components of the video source 1430 discussed elsewhere herein. These components may be stored in a format ready to be provided as needed.
[0249] The video server system 1420 optionally further includes a client qualifier 1460. The client qualifier 1460 is configured to remotely determine the capabilities of a client such as client 1410A or 1410B. These capabilities may include the capabilities of the client 1410A itself and the capabilities of one or more communication channels between the client 1410A and the video server system 1420. For example, the client qualifier 1460 may be configured to test the communication channel via the network 1415.
[0250] The client qualifier 1460 can manually or automatically determine (e.g., discover) the capabilities of the client 1410A. Manual determination includes communicating with the user of the client 1410A and asking the user to provide capabilities. For example, in some embodiments, the client qualifier 1460 is configured to display an image, text, etc. within the browser of the client 1410A. In one embodiment, the client 1410A is an HMD that includes a browser. In another embodiment, the client 1410A is a game console with a browser that can be displayed on the HMD. The displayed object requests the user to input information such as the operating system, processor, video decoder type, network connection type, display resolution, etc. of the client 1410A. The information input by the user is transmitted back to the client qualifier 1460.
[0251] Automatic determination can occur, for example, by executing an agent on the client 1410A and / or by sending a test video to the client 1410A. The agent can include computing instructions embedded in a web page or installed as an add-on, such as java script. The agent is optionally provided by the client qualifier 1460. In various embodiments, the agent can find out the processing capabilities of the client 1410A, the decoding and display capabilities of the client 1410A, the latency reliability and bandwidth of the communication channel between the client 1410A and the video server system 1420, the display type of the client 1410A, the firewall present on the client 1410A, the hardware of the client 1410A, the software executed on the client 1410A, the registry entries within the client 1410A, etc.
[0252] The client qualifier 1460 includes hardware, firmware, and / or software stored on a computer-readable medium. The client qualifier 1460 is optionally provided on a computing device separate from one or more other elements of the video server system 1420. For example, in some embodiments, the client qualifier 1460 is configured to determine the characteristics of the communication channel between the client 1410 and more than one instance of the video server system 1420. In these embodiments, the information discovered by the client qualifier can be used to determine which instance of the video server system 1420 is most suitable for delivering streaming video to one of the clients 1410.
[0253] Although specific embodiments have been provided to demonstrate predicting and post-updating a target landing point on a display such that the movement of the user's eyes is consistent with the presentation of the foveal region on the display at the updated target landing point, these are described by way of example and not limitation. Those skilled in the art who have read this disclosure will recognize other embodiments that fall within the spirit and scope of this disclosure.
[0254] It should be understood that the various features disclosed herein can be combined or assembled into specific implementations of the various embodiments defined herein. Accordingly, the provided examples are merely some possible examples and are not limited to the various implementations that may be possible by combining various elements to define more implementations. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.
[0255] Embodiments of the present disclosure can be practiced with a variety of computer system configurations including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. Embodiments of the present disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a wired or wireless network.
[0256] In view of the above embodiments, it should be understood that embodiments of the present disclosure can employ various computer-implemented operations involving data stored in a computer system. These operations are those that require physically manipulating physical quantities. Any operation described herein that forms part of an embodiment of the present disclosure is a useful machine operation. Embodiments of the present disclosure also relate to apparatus or devices for performing these operations. The apparatus may be specially constructed for the required purpose or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in the computer. In particular, various general-purpose machines may be used with a computer program written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus to perform the required operations.
[0257] The present disclosure can also be implemented as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data that can thereafter be read by a computer system. Examples of computer-readable media include hard disk drives, network-attached storage (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. The computer-readable medium can include computer-readable tangible media distributed on network-coupled computer systems such that the computer-readable code is stored and executed in a distributed fashion.
[0258] Although method operations are described in a particular order, it should be understood that other housekeeping operations may be performed between the operations, or the operations may be adjusted so that they occur at slightly different times, or the operations may be distributed across the system as long as the processing of the overlapping operations is performed in the desired manner and the system allows the processing operations to occur at various intervals associated with the processing.
[0259] Although the foregoing disclosure has been described in some detail for purposes of clarity of understanding, it will be apparent that some changes and modifications may be practiced within the scope of the appended claims. Accordingly, the present embodiment is to be considered illustrative and not restrictive, and the embodiments of the present disclosure are not limited to the details given herein but may be modified within the scope and equivalents of the appended claims.
Claims
1. A method for rendering a video frame, comprising: Generating primitives of a scene of the video frame; Receiving gaze tracking information of at least one eye of a user during a frame period; Predicting a landing point on a display corresponding to a gaze direction of the at least one eye during the frame period based on the gaze tracking information; And Rendering the video frame during the frame period based on the primitives and the predicted landing point; Wherein receiving the gaze tracking information includes tracking a first eye and a second eye of the user to increase the number of sample points for determining the gaze tracking information.
2. The method according to claim 1, wherein Rendering the video frame includes: Performing one or more shader operations to generate pixel data for a plurality of pixels of the display based on the primitives and the predicted landing point.
3. The method according to claim 2, further comprising: Storing the pixel data in a frame buffer; And Scanning and outputting the pixel data from the frame buffer to the display in a subsequent frame period.
4. The method according to claim 3, Wherein the subsequent frame period is the next frame period.
5. The method according to claim 1, wherein rendering the video frame includes: Rendering a foveal region of high resolution in the video frame, wherein the foveal region is centered on the landing point.
6. The method according to claim 1, Detecting that at least one eye of the user is undergoing a saccade; and Predicting a landing point based on the saccade.
7. The method according to claim 6, further comprising: Tracking the movement of at least one eye of the user; Determining a speed of the movement based on the tracking; Determining that the at least one eye is in a saccade when the speed reaches a threshold speed; And Predicting a landing point corresponding to a direction of the at least one eye of the saccade.
8. The method according to claim 1, further comprising: Executing an application in a previous frame period that occurred before the frame period to generate scene primitives of the video frame.
9. The method according to claim 1, Wherein the display is disposed within a head-mounted display.
10. A non-transitory computer-readable medium storing a computer program for executing a method, the computer-readable medium comprising: Program instructions for generating primitives of a scene of a video frame; Program instructions for receiving gaze tracking information of at least one eye of a user during a frame period; Program instructions for predicting a landing point on a display corresponding to a gaze direction of the at least one eye during the frame period based on the gaze tracking information; And Program instructions for rendering the video frame during the frame period based on the primitives and the predicted landing point; Wherein receiving the gaze tracking information includes tracking a first eye and a second eye of the user to increase the number of sample points for determining the gaze tracking information.
11. The non-transitory computer-readable medium according to claim 10, wherein the program instructions for rendering the video frame include: Program instructions for performing one or more shader operations to generate pixel data for a plurality of pixels of the display based on the primitives and the predicted landing point.
12. The non-transitory computer-readable medium according to claim 10, wherein the program instructions for rendering the video frame include: Program instructions for rendering a high-resolution foveal region in the video frame, wherein the foveal region is centered on the landing point.
13. The non-transitory computer-readable medium according to claim 10, further comprising: Program instructions for detecting that at least one eye of the user is undergoing a saccade; And Program instructions for predicting a landing point based on the saccade.
14. The non-transitory computer-readable medium according to claim 13, further comprising: Program instructions for tracking the movement of at least one eye of the user; Program instructions for determining the speed of the movement based on the tracking; Program instructions for determining that the at least one eye is in a saccade when the speed reaches a threshold speed; And Program instructions for predicting a landing point corresponding to the direction of at least one eye of the saccade.
15. The non-transitory computer-readable medium according to claim 10, further comprising: Program instructions for executing an application in a previous frame period that occurred before the frame period to generate scene primitives of the video frame.
16. A computer system, comprising: A processor; A memory coupled to the processor and storing instructions therein, which, if executed by the computer system, cause the computer system to perform a method, the method comprising: Generating primitives of a scene of a video frame; Receiving eye-tracking information of at least one eye of a user during a frame period; Based on the eye-tracking information, predicting a landing point on a display corresponding to the gaze direction of the at least one eye during the frame period; and Rendering the video frame during the frame period based on the primitives and the predicted landing point; Wherein receiving the eye-tracking information includes tracking a first eye and a second eye of the user to increase the number of sample points for determining the eye-tracking information.
17. The computer system according to claim 16, wherein in the method, rendering the video frame includes: Performing one or more shader operations to generate pixel data for a plurality of pixels of the display based on the primitives and the predicted landing point.
18. The computer system according to claim 16, wherein in the method, rendering the video frame includes: Rendering a high-resolution foveal region in the video frame, wherein the foveal region is centered on the landing point.
19. The computer system according to claim 16, the method further comprising: Detecting that at least one eye of the user is undergoing a saccade; And Predicting a landing point based on the saccade.
20. The computer system according to claim 16, the method further comprising: Executing an application in a previous frame period that occurred before the frame period to generate scene primitives of the video frame.
Citation Information
Patent Citations
Gaze and saccade based graphical manipulation
US20180129280A1
Image stabilization for color-sequential displays
US8970495B1