Input prediction for preloading rendering data
By fetching and loading graphics data based on a user's gaze and gestures in VR gaming, the method addresses rendering delays, providing a smoother and more immersive experience.
Patent Information
- Application Number
- JP2024525147
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-29
- Filing Date
- 2022-11-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-11-18
AI Technical Summary
Players in VR gaming environments often experience delays in rendering due to the high detail and computational resources required, which can detract from the immersive experience.
A method and system that fetches and loads graphics data into system memory based on a user's gaze and gestures, predicting interactions with content items and pre-fetching associated graphics data to reduce rendering delays.
This approach eliminates delays in rendering content items by pre-loading graphics data into system memory, enhancing the user's VR gaming experience with smoother interactions and improved image quality.
Smart Images

Figure 0007689633000001 
Figure 0007689633000002 
Figure 0007689633000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to fetching graphics data used to render game scenes, and more particularly to a method and system for fetching and loading graphics data into system memory based on a user's gaze and gestures. [Background technology]
[0002] 2. Description of Related Art The video game industry has undergone many changes over the years. In particular, the virtual reality (VR) gaming industry has seen tremendous growth over the years and is expected to continue growing at the same compound annual growth rate in the future. VR games can provide players with an immersive experience in which the player is immersed in a three-dimensional (3D) artificial environment while interacting with the VR game scene that is introduced to the player. There is a growing trend in the VR gaming industry to improve and develop unique ways to enhance the experience of VR games. Summary of the Invention [Problem to be solved by the invention]
[0003] For example, during a player's gameplay, when the player is immersed in the VR environment, the player may explore and interact with various virtual objects in the VR environment. In some cases, as the player moves through the VR scene and interacts with the virtual objects in the VR scene, the player may experience delays in rendering the VR scene because the graphics are highly detailed and require a large amount of computational resources to render the virtual objects and ensure a smooth transition throughout the player's interaction with the VR scene. Unfortunately, some players may find the delays in rendering the VR scene annoying and detract from the authentic VR experience. As a result, the player may not be provided with a fully immersive VR experience, which may result in the player not wanting to continue playing the game.
[0004] It is in this context that embodiments of the present disclosure arise. [Means for solving the problem]
[0005] Implementations of the present disclosure include methods, systems, and devices related to fetching graphics data for rendering a game scene presented on a display. In some embodiments, a method is disclosed that enables fetching and loading graphics data into a system memory based on game actions and gestures of a user playing a (VR) video game. For example, a user playing a VR video game may be immersed in a VR environment of the VR game. During the user's gameplay, when the user performs various game actions while interacting with the VR scene, the user's game actions may help infer and predict that the user focuses on a particular content item in the scene and is interested in interacting with the particular content item. In one example, game actions such as the user's gaze and the user's gestures (e.g., head movement, hand movement, body movement, position, gesture, etc.) may indicate that the user is interested in interacting with a particular content item in the game scene.
[0006] Thus, in one embodiment, the system is configured to process the user's gaze and gestures to generate a prediction of an interaction with the content item. Using the generated prediction of the interaction, the system may include a pre-fetching operation configured to pre-fetch graphics data associated with the content item and load the graphics data into the system memory in anticipation of the user interacting with the content item. Since the user's game actions are analyzed and tracked to identify content items that may be of interest to the user, the method disclosed herein outlines a method for fetching graphics data associated with a particular content item and loading the graphics data into the system memory in anticipation of the user interacting with the content item. Storing the graphics data in the system memory in this manner allows the graphics data to be quickly accessed by the system and may be used to render the content item or to render additional details related to the content item to improve image quality. In this manner, delays may be eliminated when the system renders a particular content item.
[0007] In one embodiment, a method is provided for fetching graphics data for rendering a scene presented on a display device. The method includes receiving gaze information of a user's eyes while the user interacts with the scene. The method includes tracking a user's gesture while the user interacts with the scene. The method includes identifying a content item within the scene as a potential focus of interactivity by the user. The method includes processing the user's gaze information and gestures to generate a prediction of the user's interaction with the content item. The method includes processing a pre-fetch operation to access graphics data and load the graphics data into a system memory in anticipation of the user's interaction with the content item. In this manner, once a prediction of an interaction with a content item within the scene is determined based on the user's gaming actions (e.g., gaze, gestures, etc.), graphics data associated with the content item can be loaded into the system memory and used for further rendering of the content item, which may help eliminate delays when a user interacts with the content item in a gaming scene.
[0008] In another embodiment, a system is provided for fetching graphics data for rendering a scene presented on a display. The system includes a server receiving eye gaze information of a user while the user interacts with the scene. The system includes the server tracking gestures of the user while the user interacts with the scene. The system includes the server identifying content items within the scene as potential foci of interactivity by the user. The system includes the server processing the user gaze information and gestures to generate a prediction of the user's interaction with the content item. The system includes the server accessing the graphics data in anticipation of the user interacting with the content item and processing a pre-fetch operation to load the graphics data into system memory.
[0009] Other aspects and advantages of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrating by way of example the principles of the disclosure.
[0010] The present disclosure may be better understood by reference to the following description taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0011] [Figure 1A] 1 illustrates an embodiment of a system for interaction with a virtual environment via a head mounted display (HMD), according to an embodiment of the present disclosure. [Figure 1B] 1 illustrates an embodiment of a system for tracking gaze information and gestures of a user 100 while the user 100 interacts with a game scene during game play, according to an embodiment of the present disclosure. [Figure 2A] 1 illustrates an embodiment of a user's view into a virtual environment, showing the user interacting with a virtual reality scene while wearing an HMD, according to an embodiment of the present disclosure. [Figure 2B]2B illustrates an embodiment of the user's virtual environment shown in FIG. 2A showing the user interacting with a virtual reality scene with the user's virtual arm reaching out towards a content item, according to an embodiment of the present disclosure. [Figure 3A] 1 illustrates another embodiment of a user's view into a virtual environment showing the user interacting with a virtual reality scene while wearing an HMD, according to an embodiment of the present disclosure. [Figure 3B] 1 illustrates another embodiment of a user's view into a virtual environment showing the user interacting with a virtual reality scene while wearing an HMD, according to an embodiment of the present disclosure. [Figure 3C] 1 illustrates another embodiment of a user's view into a virtual environment showing the user interacting with a virtual reality scene while wearing an HMD, according to an embodiment of the present disclosure. [Figure 4] 1 illustrates an embodiment of a system for fetching graphics data corresponding to an identified content item for rendering in a game scene, according to an embodiment of the present disclosure. [Diagram 5] 1 illustrates an embodiment of a table illustrating a user's gaze information and gestures tracked during the user's gameplay and generated predictions of the user's interactions with content items according to an embodiment of the present disclosure. [Figure 6] 1 illustrates components of an exemplary device that can be used to implement aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] The following embodiments of the present disclosure provide methods, systems, and devices for fetching graphics data used to render a game scene presented on a display. Specifically, the display may be a head-mounted display (HMD) of a user playing a virtual reality (VR) video game, or a display associated with the user's device. In one embodiment, the graphics data corresponds to one or more content items in the game scene and can be used to render additional details related to the content items. In some embodiments, the content items are identified based on a user's game actions while interacting with the game scene. For example, during the user's gameplay, the user's game actions, such as gaze and gestures (e.g., body movements), are tracked in real time while the user interacts with the game scene. In one example, the user's gaze and user gestures are processed to identify content items in the game scene that the user is potentially interested in interacting with. Accordingly, graphics data corresponding to the identified content items are pre-fetched and loaded into a system memory in anticipation of the user interacting with the content items. Storing the graphics data in system memory in this manner allows the graphics data to be quickly accessed by the system and used to render a particular content item or to render additional detail related to a content item. In this manner, delays associated with rendering various content items within a game scene can be eliminated, thereby enhancing the user's gaming experience by providing the user with an uninterrupted VR gaming experience.
[0013] By way of example, in one embodiment, a method is disclosed for facilitating fetching of graphics data used to render a scene presented on a display. The method includes receiving eye gaze information of a user while the user is interacting with the scene. In one embodiment, the method may further include tracking user gestures while the user is interacting with the scene. In another embodiment, the method may include identifying a content item within the scene as a potential focus of interactivity by the user. In some embodiments, the method includes processing the user gaze information and gestures to generate a prediction of the user's interaction with the content item. In other embodiments, the method includes processing a pre-fetch operation to access graphics data and load the graphics data into a system memory in anticipation of the user interacting with the content item. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without some or all of the specific details described herein. In other examples, well-known process operations have not been described in detail so as not to unnecessarily obscure the present disclosure.
[0014] With the above summary in mind, the following provides several illustrative figures to facilitate understanding of example embodiments.
[0015] FIG. 1A illustrates an embodiment of a system for interaction with a virtual environment via a head mounted display (HMD) according to an embodiment of the present disclosure. The HMD may also be referred to as a virtual reality (VR) headset. As used herein, the term "virtual reality" (VR) generally refers to user interaction with a virtual space / environment including viewing the virtual space through an HMD (or a VR headset) in a manner that responds in real time to the (user-controlled) movements of the HMD to provide the user with the sensation of being in the virtual space. For example, a user may see a three-dimensional (3D) representation of the virtual space when facing a given direction, and when the user turns to the side, thereby changing the orientation of the HMD in the same way, a representation of that side of the virtual space is rendered on the HMD.
[0016] As shown in FIG. 1A, a user 100 is shown physically located in a real-world space 120 wearing an HMD 102 and manipulating interface objects 104 to provide input to a video game. The HMD 102 is worn in a manner similar to glasses, goggles, or a helmet and is configured to display video games or other content to the user 100. The HMD 102 provides a highly immersive experience to the user by providing a display mechanism in close proximity to the user's eyes. Thus, the HMD 102 can provide a display area for each of the user's eyes that occupies a large portion or even the entirety of the user's field of view, and can also provide viewing with three-dimensional depth and perspective.
[0017] In some embodiments, the HMD 102 may provide the user with a gameplay point of view (POV) 108 into the VR scene. Thus, when the user 100 turns his / her head to look towards a different area in the virtual environment, the VR scene is updated to include any additional virtual objects that may be present within the gameplay POV 108 of the user 100. In one embodiment, the HMD 102 may include an eye-tracking camera configured to capture images of the user's 100 eyes while the user is interacting with the VR scene. The gaze information captured by the eye-tracking camera may include information related to the user's 100 gaze direction and specific virtual objects and content items within the VR scene that the user 100 is looking at or is interested in interacting with. Thus, based on the user's 100 gaze direction, the system may detect specific virtual objects and content items, e.g., game characters, game objects, game items, etc., that may be potential focal points for the user if the user is interested in interacting and engaging with them.
[0018] In some embodiments, the HMD 102 may include an outward-facing camera configured to capture images of the user's 100 real-world space 120, such as the user's body movements, and any real-world objects that may be located in the real-world space. In some embodiments, the images captured by the outward-facing camera may be analyzed to determine the position / orientation of the real-world objects relative to the HMD 102. With the known position / orientation of the HMD 102, the real-world objects, as well as inertial sensor data from the HMD, the user's gestures and movements, may be continuously monitored and tracked while the user interacts with the VR scene. For example, while interacting with an in-game scene, the user 100 may perform various gestures, such as pointing at or walking toward a particular content item in the scene. In one embodiment, the gestures may be tracked and processed by the system to generate a prediction of an interaction with a particular content item in the game scene. In other embodiments, the HMD 102 may include one or more lights that may be tracked to determine the position and orientation of the HMD 102.
[0019] As described above, the user 100 may manipulate the interface objects 104 to provide input to the video game. In various implementations, the interface objects 104 include lights and / or inertial sensor(s) that may be tracked to enable the determination of the interface object's position and orientation, as well as tracking of its movement. The manner in which the user 100 interfaces with the virtual reality scene displayed on the HMD 102 may vary, and other interface devices may be used in addition to the interface objects 104. For example, various types of one-handed and two-handed controllers may be used. In some implementations, the controller itself may be tracked by tracking lights included in the controller, or by tracking shape, sensor, and inertial data associated with the controller. These various controllers, or even simple hand gestures performed and captured by one or more cameras, may be used to interface, control, manipulate, interact with, and participate in the virtual reality environment presented on the HMD 102.
[0020] In the illustrated implementation, the HMD 102 is wirelessly connected to the cloud computing and gaming system 114 via the network 112. In one embodiment, the cloud computing and gaming system 114 maintains and executes the video game being played by the user 100. In some embodiments, the cloud computing and gaming system 114 is configured to receive input from the HMD 102 and the interface objects 104 via the network 112. The cloud computing and gaming system 114 is configured to process the input to affect the game state of the running video game. Outputs, such as video data, audio data, and haptic feedback data from the running video game are sent to the HMD 102 and the interface objects 104. As an example, video and audio streams are provided to the HMD 102, while haptic / vibration feedback commands are provided to the interface objects 104. In other implementations, the HMD 102 can communicate with the cloud computing and gaming system 114 wirelessly via alternative mechanisms or channels, such as a cellular network.
[0021] Additionally, although embodiments of the present disclosure may be described with respect to a head-mounted display, it will be understood that in other embodiments, a non-head-mounted display may be used instead, including, but not limited to, a portable device screen (e.g., a tablet, smartphone, laptop, etc.) or any other type of display that may be configured to render video and / or provide a display of an interactive scene or virtual environment in accordance with the present embodiments.
[0022] FIG. 1B illustrates an embodiment of a system for tracking gaze information and gestures of a user 100 while the user 100 interacts with a game scene during gameplay. As illustrated in the figure, the user 100 is shown standing in front of a display 105 and playing a game. The user 100 can play the game using interface objects 104 that provide input to the game. A computer 110 is connected to the display 105 via a wire. A camera 116 is positioned on top of the display 105 and configured to capture the user playing the game while the user is immersed in the gameplay. The camera 116 includes a camera point of view (POV) 118 that captures the user 100 and objects within the POV. According to the illustrated embodiment, the computer 110 can communicate with a cloud computing and gaming system 114 via a network 112.
[0023] The camera 116 may include eye tracking to enable tracking of the gaze of the user 100. The camera 116 is configured to capture and analyze images of the user's eyes to determine the gaze 106 of the user 100. In some embodiments, the camera 116 may be configured to capture and process gestures and body movements of the user 100 during gameplay. For example, during the user's 100 gameplay, the user may encounter various content items (e.g., game objects, game characters, etc.) with which the user may be interested in interacting. When the gaze 106 focuses on a particular content item while the user is moving in the direction of the content items, the noted actions (e.g., gaze, body movements) may be processed and the particular content item in the game scene may be identified as a potential focus of interactivity by the user. Thus, the system may be configured to track the gaze, gestures, and body movements of the user 100 during gameplay and use these to generate a prediction of interaction with a particular content item in the game scene.
[0024] In other embodiments, the camera 116 may be configured to track and capture facial expressions of the user 100 during gameplay, and these facial expressions are analyzed to determine the emotion associated with the expressions. In some embodiments, the camera 116 may be mounted on a three-axis gimbal that allows the camera to rotate freely about any axis to allow for the capture of various angles of the user. In one embodiment, the camera 116 may be a pan-tilt-zoom camera that can be configured to automatically zoom in and track the user's face and body as the user moves during gameplay.
[0025] In some embodiments, the interface object 104 may include one or more microphones to capture sounds from the real-world space 120 in which the game is being played. Sounds captured by the microphones may be processed to identify the location of the sound source. Sounds from the identified locations may be selectively utilized or processed to eliminate other sounds that are not from the identified locations. This information may be utilized in a variety of ways, including filtering out unwanted sound sources, associating sound sources with visual identifications, and the like. In some implementations, the interface object 104 may be tracked by tracking lights included in the interface object 104 or by tracking shape, sensor, and inertial data associated with the interface object 104. In various implementations, the interface object 104 includes lights and / or inertial sensor(s) that may be tracked to enable determination of the controller's position and orientation, as well as tracking of movement.
[0026] After the computer 110 captures data related to the user 100 during gameplay (e.g., gaze data, gesture information, body movement data, facial expression data, voice acting data, inertial sensor data, controller input data), the data may be transmitted over the network 112 to the cloud computing and gaming system 114. In some embodiments, the cloud computing and gaming system 114 may receive, process, and execute various data from the user 100 to generate a prediction of the user's interaction with a content item. In some embodiments, the cloud computing and gaming system 114 may utilize a pre-fetching operation to access and load into system memory graphics data corresponding to a content item in anticipation of the user's interaction with the content item. In some embodiments, the graphics data corresponds to a particular content item in a scene with which the user is interested in interacting, and the graphics data may be used to render the particular content item or to further enhance the image quality (e.g., roughness, curvature, geometry, vertices, depth, color, lighting, shading, texturing, motion, etc.) of the particular content item.
[0027] FIG. 2A illustrates an embodiment of a user's 100 view into a virtual environment, showing the user 100 interacting with a virtual reality scene 202a while wearing an HMD 102. As the user 100 holds an interface object 104 or uses his or her arms and hands to interact with the virtual reality scene 202, the system is configured to track the user's 100 gestures and the user's line of sight 106. For example, in the example shown in FIG. 2A, the virtual reality scene 202a includes multiple content items 204a-c that are rendered within the scene. As shown, content item 204a represents a sculpture of the Statue of Liberty, content item 204b represents a picture frame, and content item 204c represents a wardrobe closet. Notably, the Statue of Liberty sculpture is shown at a lower image quality (e.g., lower resolution) when a wireframe of the statue is rendered within the scene. During the user's interaction with the virtual reality scene 202, the user's gaze 106 is directed toward a content item 204a (e.g., the Statue of Liberty sculpture), such that the system can identify the content item 204a as a potential focus of interactivity by the user 100.
[0028] In some embodiments, the system is configured to process user gestures and gaze information, such as gaze 106, to generate a prediction of an interaction by the user 100 with a content item 204a (e.g., the Statue of Liberty sculpture). In one embodiment, the user gestures may include user actions, such as the user's head movements, hand movements, body movements, body language, and position. For example, while the user's gaze 106 is focused on the sculpture, the user's gestures may indicate that the user 100 is turning toward and walking toward the sculpture. Thus, using the gaze and gesture information, the system can generate a prediction of an interaction with the content item. In this example, the interaction prediction may include that the user wants to pick up and touch the sculpture.
[0029] FIG. 2B illustrates an embodiment of the virtual environment of the user 100 shown in FIG. 2A, showing the user 100 interacting with the virtual reality scene 202a with the user's virtual arm 100' reaching out towards a content item 204a (e.g., the Statue of Liberty sculpture). As illustrated in FIG. 2B, the gestures 204a-204b of the user 100 are tracked while the user interacts with the virtual reality scene 202a. Gesture 206a illustrates the user reaching out in the real world space 120 to touch the content item 204a while the user's line of sight 106 is focused and directed towards the content item 204a. As further illustrated in the figure, gesture 206b illustrates the user's 100 body position closer to the content item 204a as compared to the user's body position shown in FIG. 2A. Thus, using the user's 100 gaze 106 and the user's 100 gestures 204a-204b, the system can predict one or more interaction types in which the user may be interested, such as touching the sculpture, grabbing the sculpture, looking at the sculpture from a closer distance to inspect its details, etc.
[0030] In some embodiments, one or more actions of a user can be used to predict interactions with a content item in a scene, e.g., user input from any device can be used to pre-render content associated with expected user actions. In other embodiments, a sequence of gestures 204 and actions of a user 100 can predict interactions with a content item that can have predictable outcomes in a scene, allowing for more effective pre-rendering of the content item.
[0031] In some embodiments, once a content item 204 is identified as a potential focus of interactivity and a prediction of interaction with the content item 204 is generated, the system may utilize a prefetch operation to access and load the graphics data into system memory. In one embodiment, the graphics data corresponds to the content item 204 that the system identifies as a potential focus of interactivity with which a user is expected to interact.
[0032] In one embodiment, the system may include a central processing unit (CPU) and a graphics processing unit (GPU) configured to access graphics data from a system memory to render additional details corresponding to an identified content item 204. In one example, the graphics data may include data defining the geometry, vertices, depth, color, lighting, shading, texturing, motion, etc. of the content item 204. For example, referring simultaneously to FIGS. 2A and 2B, a user's gaze 106 is focused on a content item 204a (e.g., the Statue of Liberty sculpture). As the user 100 walks toward the sculpture and reaches out to touch it with his or her hands (e.g., gestures 204a-204b), the system is configured to access graphics data corresponding to the sculpture and use the graphics data to render additional details of the sculpture that enhance the image quality of the sculpture.
[0033] As shown in FIG. 2B, the user is located at a close distance to the content item 204a where the user can touch and interact with the content item 204a. As further shown, the content item 204a includes more detail and improved image quality (e.g., depth, shading, texturing, etc.) as compared to the content item 204a shown in FIG. 2A. In one embodiment, the amount of graphics data used to render the content item is based on the user's relative distance to the content item. For example, the system is configured to start rendering the details of the content item 204a when the user begins to move towards the content item 204a, and the amount of graphics data used to render the content item 204a increases as the user approaches the content item 204a. In another embodiment, the system is configured to start rendering the details of the content item 204a when the user's gaze 106 is focused and maintained on the content item 204a. In one embodiment, the amount of graphics data used to render the details of the content item is based on the length of time the user's gaze is maintained on the content item and the brightness of the user's eyes.
[0034] 3A-3C illustrate another embodiment of a user's 100 view into a virtual environment, showing the user 100 interacting with a virtual reality scene 202b while wearing an HMD 102. As illustrated in FIG. 3A, the user 100 is shown holding an interface object 104 while viewing the virtual reality scene 202b through the HMD 102. The virtual reality scene 202b includes a content item 204d representing a treasure chest placed on the floor. As described above, as the user 100 interacts with the virtual reality scene 202b, the system is configured to track the user's gaze and the user's gestures during the user's interaction with the game scene.
[0035] For example, referring to FIG. 3B, the user's gaze 106 is focused on the treasure chest. The user's gaze 106 data and other information related to the eyes (e.g., pupil light reflex, pupil size, eye movements, etc.) are tracked and processed by the system to identify content items that may be potential foci of interactivity by the user. As further illustrated, gesture 206c illustrates a body position of the user 100 closer to the treasure chest as compared to the body position of the user shown in FIG. 3A. Using the gaze information and gestures made by the user, the system is configured to generate a prediction of an interaction with the treasure chest, and graphics data corresponding to the treasure chest is retrieved from data storage and stored in system memory.
[0036] As illustrated in the example shown in FIG. 3B, the gaze 106 or gesture 206c can initiate a pre-fetch operation configured to access and load graphics data corresponding to the treasure chest into system memory. As the user 100 walks toward the treasure chest while maintaining the gaze 106 on the treasure chest, graphics data related to the treasure chest (e.g., content items, geometric shapes, vertices, depth, color, lighting, shading, texturing, motion, etc.) is loaded into system memory in anticipation of the user wanting to interact with the treasure chest. In the example shown in FIG. 2B, graphics data related to the contents of the treasure chest is loaded into system memory in anticipation of the user opening the treasure chest to discover and see what may be inside, e.g., gold, silver, gems, etc.
[0037] In other embodiments, if the user's gaze is no longer directed at the treasure chest or if the user's gestures 206 suggest that the user is no longer interested in the treasure chest, the system may be configured to pause loading of the graphics data into system memory and resume loading of the graphics data at a time after the user indicates an interest in interacting with the treasure chest. In some embodiments, if the user's gaze is no longer directed at an identified treasure chest or if the user's gestures 206 suggest that the user is no longer interested in the treasure chest, the graphics data corresponding to the content item 204d treasure chest is deleted from system memory.
[0038] Referring to FIG. 3C, while the user's gaze 106 is focused on the treasure chest, gesture 206d illustrates the user reaching out in a direction toward the treasure chest to explore its contents. As shown in virtual reality scene 202b, the user's virtual arm 100' is shown opening the lid of the treasure chest. In one embodiment, when the user's hand touches the lid of the treasure chest, the system is configured to fetch graphics data corresponding to the contents of the treasure chest from the system memory to render the contents of the treasure chest, e.g., gold coins. In this manner, detailed graphics (e.g., high-definition resolution) of the gold coins can be rendered quickly because the graphics data of the gold coins is stored in the system memory, and delays associated with rendering the gold coins or the contents of the treasure chest can be avoided. Thus, the system described with respect to FIGS. 3A-3C provides a method for tracking the user's gaze and gestures to identify content items in a scene that can be pre-fetched and loaded into the system memory and used to render content items in a scene that may be of interest to the user. Doing so may help eliminate delays in rendering the game scene while the user is interacting with the virtual reality scene 202.
[0039] 4 illustrates an embodiment of a system for fetching graphics data corresponding to identified content items for rendering in a game scene. The diagram shows how a behavior model 402 is used, and input data 406 such as user gaze information, user gestures, and interactive data (e.g., game play data) as inputs, to fetch graphics data for loading into system memory 412. In one embodiment, during the user's 100 gameplay, the user's 100 gaze 106 and gestures 206 are tracked and transmitted over network 112 to cloud computing and gaming system 114.
[0040] The method then proceeds to a behavioral model 402 configured to receive input data 406, such as user gaze information, user gestures, and interactive data. In some embodiments, other inputs that are not direct inputs may be obtained as inputs to the behavioral model 402. During the user's gameplay, the behavioral model 402 may also use machine learning models that are used to identify content items in the scene as potential foci of interactivity by the user, as well as generate predictions of the user's interactions 408 with the content items. The behavioral model 402 may also be used to identify patterns, similarities, and relationships between the user's gaze information, user gestures, and interactive data. Using the patterns, similarities, and relationships, the behavioral model 402 may be used to identify content items in the scene that may potentially be of interest to the user, and predictions of the user's interactions 408. In one embodiment, the interaction predictions 408 may include a wide range of interaction types that the user may perform in the game. Such predicted interactions may include the user reaching out to an identified content item to see more of the content item, touching the content item to feel it, exploring content stored within the content item, opening a door to explore what is on the other side, etc. Over time, the behavior model 402 is trained to predict the likelihood that a user will interact with a particular content item in a game scene, and the amount of graphics data used to render the content item can be adjusted based on the prediction.
[0041] After identifying a content item 204 in a scene as a potential focus of interactivity by a user and generating a prediction 408 of a user's interaction with the content item 204, the method flows to a cloud computing and gaming system 114, where the cloud computing and gaming system 114 is configured to process the identified content item 204 and the prediction 408 of the interaction. In some embodiments, a pre-fetching operation 404 may be utilized to access graphics data corresponding to the identified content item 204 from a game rendering data storage 410 using the prediction 408 of the interaction with the content item 204. In one embodiment, the pre-fetching operation 404 is configured to make adjustments based on updates to the prediction 408 of the interaction to increase or decrease the amount of graphics data accessed and loaded into the system memory 412. For example, if a user's gaze is centered on a series of different content items, the pre-fetching operation 404 may increase the amount of graphics data for the content items on which the user's gaze is primarily centered and decrease the amount of graphics data for content items on which the user's gaze is less centered.
[0042] In one embodiment, the graphics data may be used to render or further enhance the image quality (e.g., roughness, curvature, geometry, vertices, depth, color, lighting, shading, texturing, motion, etc.) of a particular content item 204. After the graphics data is accessed, it may be loaded and stored in system memory 412 and used by the CPU and GPU to render the content item 204 in a scene or at a higher resolution or to enhance the image quality of the content item 204.
[0043] For example, as shown in FIG. 4, a user 100 is shown opening a lid of a treasure chest (e.g., content item 204d) while directing a gaze 106 at the chest. A virtual arm 100′ is shown in the virtual reality scene 202 to represent the user's arm opening the lid of the treasure chest to explore the contents of the treasure chest. Using the user's gaze 106, the user's gestures 206, and the interaction data, the behavior model 402 can generate a prediction of an interaction 408 including that the user 100 may want to explore the contents of the treasure chest to see what is inside. Using the interaction prediction 408, the cloud computing and gaming system 114 can access and load graphics data related to the contents of the treasure chest since it predicts that the user may want to sift through all of the contents to see what is in the treasure chest. After storing the graphics data related to the contents of the treasure chest in the system memory 412, the graphics data can be used to render the contents of the treasure chest as the user 100 begins sift through the contents of the treasure chest. In this way, images of the treasure chest's contents can be rendered seamlessly in high resolution without any delay or lag time as the user explores the content.
[0044] FIG. 5 illustrates an embodiment of a table 502 illustrating a user's gaze information 504 and gestures 206 tracked during the user's gameplay and generated predictions of the user's interactions 408 with content items 204. As shown in the figure, the table 502 includes a user's gaze information 504 and gestures 206 tracked while the user interacts with a game scene. As further illustrated in the table, a content item 204 within the game scene has been identified as a potential focus of interactivity based on the corresponding user's gaze information 504 and gestures 206. In one embodiment, the gaze information 504 may include the user's gaze 106 and other information related to the user's eyes captured during the user's gameplay, such as eye movement, pupil light reflex, pupil size, etc. In some embodiments, the user's gestures 206 are tracked during gameplay, which may include various body movements and body language data that are processed to aid in determining predictions of the user's interactions 408 with content items within the scene.
[0045] To provide an illustration of table 502 of FIG. 5, in one example, the system may determine based on game scene context 506 that the user is in a scene involving a fight with an enemy character (e.g., a cyberdemon). During the user's interaction with the scene, the system may determine that the user's eyes are blinking rapidly while the user's gaze is focused on an exit door. The user's gestures may indicate that the user's hands are fidgeting or that the user is frightened by the idea of having to fight an enemy character without a suitable weapon. The recorded information may be processed and received by the behavior model 402 and used to generate a prediction of an interaction 408 that includes the user running toward the exit door to escape or to find a weapon. Using the prediction of the interaction, the system is configured to access and load graphics data related to content that may be on the other side of the exit door when the user opens the door to escape. Thus, as the graphics data is loaded into the system memory, the CPU and GPU may access the graphics data and render an image of the content of the scene that may be behind the exit door.
[0046] FIG. 6 illustrates components of an exemplary device 600 that can be used to implement aspects of various embodiments of the present disclosure. The block diagram illustrates device 600, which can incorporate or be a personal computer, video game console, personal digital assistant, server, or other digital device suitable for implementing embodiments of the present disclosure. Device 600 includes a central processing unit (CPU) 602 for executing software applications and optionally an operating system. CPU 602 may be comprised of one or more homogeneous or heterogeneous processing cores. For example, CPU 602 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture that is particularly adapted for highly parallel and computationally intensive applications, such as interpreting queries, identifying contextually relevant resources, and immediately implementing and rendering the contextually relevant resources within a video game. Device 600 may be one that is localized to a player playing a game segment (e.g., a game console) or one that is remote from the player (e.g., a back-end server processor), or one of many servers that use virtualization in a game cloud system for remote streaming of game play to clients.
[0047] Memory 604 stores applications and data used by CPU 602. Storage 606 provides non-volatile storage and other computer-readable media for applications and data and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input devices 608 communicate user input from one or more users to device 600, and examples of user input devices 608 may include a keyboard, mouse, joystick, touchpad, touch screen, still or video recorder / camera, gesture-recognizing tracking device, and / or microphone. Network interface 614 enables device 600 to communicate with other computer systems over an electronic communications network, which may include wired or wireless communications over a local area network or a wide area network such as the Internet. The audio processor 612 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 602, memory 604, and / or storage 606. The components of the device 600, including the CPU 602, memory 604, data storage 606, user input device 608, network interface 610, and audio processor 612, are connected via one or more data buses 622.
[0048] A graphics subsystem 620 is further coupled to the data bus 622 and the components of the device 600. The graphics subsystem 620 includes a graphics processing unit (GPU) 616 and a graphics memory 618. The graphics memory 618 includes a display memory (e.g., a frame buffer) used to store pixel data for each pixel of an output image. The graphics memory 618 may be integrated in the same device as the GPU 608, may be coupled as a separate device from the GPU 616, and / or may be incorporated within the memory 604. The pixel data may be provided to the graphics memory 618 directly from the CPU 602. Alternatively, the CPU 602 provides data and / or instructions defining a desired output image to the GPU 616, from which the GPU 616 generates pixel data for one or more output images. The data and / or instructions defining the desired output image may be stored in the memory 604 and / or the graphics memory 618. In an embodiment, GPU 616 includes 3D rendering capabilities for generating pixel data for output images from instructions and data defining the geometry, lighting, shading, texturing, motion, and / or camera parameters of a scene. GPU 616 may further include one or more programmable execution units capable of executing shader programs.
[0049] Graphics subsystem 614 periodically outputs pixel data of an image from graphics memory 618 for display on display device 610. Display device 610 may be any device capable of displaying visual information in response to signals from device 600, including CRT, LCD, plasma, and OLED displays. Device 600 may provide analog or digital signals to display device 610, for example.
[0050] It should be noted that access services distributed over a wide geography, such as providing access to games in the current embodiment, often use cloud computing. Cloud computing is a computing paradigm in which dynamically scalable, often virtualized resources are provided as a service over the Internet. Users do not need to be experts in the technical infrastructure of the "cloud" that supports them. Cloud computing can be divided into different services such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services often provide common applications, such as video games, online, accessed from a web browser, but the software and data are stored on servers in the cloud. The term cloud is used as a metaphor for the Internet based on how the Internet is depicted in computer network diagrams, and is an abstraction of the complex infrastructure it hides.
[0051] In some embodiments, a game server may be used to perform the operation of the duration information platform for video game players. Most video games played over the Internet operate through a connection to a game server. Typically, the game uses a dedicated server application that collects data from the player and distributes the collected data to other players. In other embodiments, the video game may be executed by a distributed game engine. In these embodiments, the distributed game engine may run on multiple processing entities (PEs), each PE executing a functional segment of the given game engine in which the video game is executed. Each processing entity is viewed by the game engine as simply a computational node. The game engine typically performs a functionally diverse set of operations to execute the video game application along with additional services experienced by the user. For example, the game engine implements game logic and performs game calculations, physics, geometry transformations, rendering, lighting, shading, audio, as well as additional in-game or game-related services. The additional services may include, for example, messaging, social utilities, audio communication, game play playback functions, help functions, and the like. The game engine may run on an operating system virtualized by a hypervisor on a particular server, but in other embodiments the game engine itself may be distributed across multiple processing entities, each of which may reside on a different server unit in a data center.
[0052] According to this embodiment, each processing entity for the execution of operations may be a server unit, a virtual machine, or a container depending on the needs of each game engine segment. For example, if a game engine segment is responsible for camera transformation, that particular game engine segment may be provisioned with a virtual machine associated with a graphics processing unit (GPU) since it will be performing a lot of relatively simple mathematical operations (e.g., matrix transformations). Other game engine segments, requiring fewer but more complex operations, may be provisioned with processing entities associated with one or more higher-powered central processing units (CPUs).
[0053] By distributing the game engine, the game engine has elastic computational characteristics that are not bound by the capabilities of a physical server unit. Instead, the game engine is provisioned with more or fewer computation nodes as needed to meet the demands of the video game. From the perspective of the video game and the video game player, a game engine that is distributed across multiple computation nodes is indistinguishable from a non-distributed game engine running on a single processing entity, since a game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output component to the end user.
[0054] Users access remote services through a client device that includes at least a CPU, a display, and I / O. The client device may be a PC, a cell phone, a netbook, a PDA, etc. In one embodiment, a network running on the game server recognizes the type of device used by the client and adjusts the communication method employed. In another case, the client device accesses the application on the game server over the Internet using a standard communication method such as HTML.
[0055] Of course, a given video game or game application may be developed for a particular platform and a particular associated controller device. However, when such games are made available through a game cloud system as described herein, users may access the video game using different controller devices. For example, a game may be developed for a game console and its associated controller, but a user may access a cloud-based version of the game from a personal computer utilizing a keyboard and mouse. In such a scenario, the input parameter settings may define a mapping from inputs that may be generated by the controller devices available to the user (keyboard and mouse in this case) to inputs that are acceptable for execution of the video game.
[0056] In another example, a user may access the cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen driven device. In this case, the client device and the controller device are integrated together in the same device, and input is provided by detected touchscreen input / gestures. In such a device, an input parameter setting may define a particular touchscreen input that corresponds to game input of a video game. For example, a button, a directional pad, or other type of input element may be displayed or overlaid during the execution of the video game to indicate a location on the touchscreen that the user can touch to generate game input. Gestures such as swipes in a particular orientation, or particular touch motions may also be detected as game input. In one embodiment, to familiarize the user with control operations on a touchscreen, a tutorial may be provided to the user showing how to input gameplay via the touchscreen, for example, before starting gameplay of the video game.
[0057] In some embodiments, the client device serves as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection and transmits inputs from the controller device to the client device. The client device may then process these inputs and then transmit the input data to the cloud gaming server over a network (e.g., a network accessed through a local network device such as a router). However, in other embodiments, the controller itself may be a networked device that has the ability to communicate inputs directly to the cloud gaming server over a network, without the need for such inputs to be communicated through the client device first. For example, the controller may connect to a local network device (such as the router mentioned above) to send and receive data to the cloud gaming server. Thus, while the client device is still required to receive video output from the cloud-based video game and render it on a local display, input latency may be reduced by allowing the controller to bypass the client device and transmit inputs directly over the network to the cloud gaming server.
[0058] In one embodiment, the networked controller and client devices can be configured to transmit certain types of inputs directly from the controller to the cloud gaming server and other types of inputs via the client device. For example, inputs whose detection does not rely on any additional hardware or processing, apart from the controller itself, can be transmitted directly from the controller to the cloud gaming server over the network, bypassing the client device. Such inputs can include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), and the like. However, inputs that utilize additional hardware or require processing by the client device can be transmitted by the client device to the cloud gaming server. These can include captured video or audio from the gaming environment that can be processed by the client device before transmission to the cloud gaming server. In addition, inputs from the controller's motion detection hardware can be processed by the client device in conjunction with the captured video to detect the position and movement of the controller, which is then communicated by the client device to the cloud gaming server. It should be understood that the controller device according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.
[0059] It should be understood that the various embodiments defined herein may be combined or assembled into specific implementations that use various features disclosed herein. Thus, the examples provided are only some of the possible examples and are not intended to limit the various implementations that are possible by combining various elements to define more implementations. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.
[0060] Embodiments of the present disclosure may be practiced with a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire-based or wireless network.
[0061] Although the method operations have been described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be coordinated to occur at slightly different times, or operations may be distributed within the system to allow processing operations to occur at various processing-related intervals, so long as the processing of telemetry and game state data to generate modified game state is performed in the desired manner.
[0062] One or more embodiments may also be fabricated as computer readable code on a computer readable medium. A computer readable medium is any data storage device that can store data which can then be read by a computer system. Examples of computer readable media include hard drives, network attached storage (NAS), read only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. A computer readable medium may include computer readable tangible media distributed across network connected computer systems such that the computer readable code is stored and executed in a distributed fashion.
[0063] In one embodiment, the video game is executed locally on a game console, personal computer, or server. In some cases, the video game is executed by one or more servers in a data center. When the video game is executed, some instances of the video game may be simulations of the video game. For example, the video game may be executed by an environment or server that generates a simulation of the video game. A simulation, in some embodiments, is an instance of the video game. In other embodiments, the simulation may be generated by an emulator. In either case, when the video game is represented as a simulation, the simulation may be executed to render interactive content that can be interactively streamed, executed, and / or controlled by user input.
[0064] Although the foregoing embodiments have been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims. Thus, the present embodiments should be considered as illustrative and not restrictive, and the present embodiments should not be limited to the details set forth herein, but may be modified within the scope of the appended claims and their equivalents.
Claims
1. 1. A method of fetching graphics data for rendering a scene presented on a display, comprising: receiving eye gaze information of the user while the user interacts with the scene; tracking gestures of the user while the user interacts with the scene; identifying content items within the scene as potential foci of interactivity by the user; processing the gaze information and the gestures of the user to generate a prediction of an interaction of the user with the content item; accessing the graphics data in anticipation of the user interacting with the content item and processing a prefetch operation to load the graphics data into a system memory; Including, The method, wherein the graphics data is used to render the content items in the scene to enhance the quality of the content items that were initially presented in the scene at a lower quality.
2. 2. The method of claim 1, wherein the prefetching operation is adjusted to increase or decrease an amount of the graphics data accessed and loaded into the system memory based on updating the prediction.
3. The method of claim 1 , wherein the interaction prediction is processed in part using a behavioral model trained over time to predict the likelihood that the user will interact with the content item or another content item.
4. The method of claim 1 , wherein the amount of graphics data used in the rendering of the content item is based on a distance of the user to the content item.
5. The method of claim 4 , wherein the amount of the graphics data used in the rendering of the content item is based on the amount of time the user's gaze is maintained on the content item.
6. The method of claim 5 , wherein the amount of the graphics data used for the rendering is dynamically adjusted based on learning that the user has previously interacted with the content item in the scene.
7. The method of claim 1 , wherein the graphics data is used to render another content item within the scene in anticipation of the user interacting with the other content item.
8. The method of claim 1 , wherein the gesture of the user is a head movement, a hand movement, a body movement, a position of the user relative to the content item, a body gesture, or a combination of two or more thereof.
9. The method of claim 1 , wherein the gaze information includes the user's gaze, pupil size, eye movement, or a combination of two or more thereof.
10. 2. The method of claim 1, wherein the graphics data includes information related to the identified content item, the information including roughness, curvature, geometry, vertices, depth, color, lighting, shading, texturing, or a combination of two or more thereof.
11. The method of claim 1 , further comprising using the graphics data to render additional detail associated with the content item to enhance the image quality of the content item.
12. The method of claim 1 , further comprising: using the graphics data to render other content items in the scene in anticipation of the user interacting with the other content items.
13. 2. The method of claim 1 , wherein the prediction of the interaction with the content item is based on processing the gaze information, the gestures, and the interactive data through a behavioral model, the behavioral model configured to identify relationships between the gaze information, the gestures, and the interactive data to generate the prediction of the interaction with the content item.
14. 1. A system for fetching graphics data for rendering a scene presented on a display, comprising: a server receiving eye gaze information of the user while the user interacts with the scene; the server tracking gestures of the user while the user interacts with the scene; the server identifying content items within the scene as potential foci of interactivity by the user; the server processing the gaze information and the gestures of the user to generate a prediction of the user's interaction with the content item; the server accessing the graphics data in anticipation of the user interacting with the content item and processing a pre-fetch operation to load the graphics data into system memory; Including, The system, wherein the graphics data is used to render the content items in the scene to enhance the quality of the content items that were initially presented in the scene at a lower quality.
15. 15. The system of claim 14, wherein the prefetching operation is adjusted to increase or decrease an amount of the graphics data accessed and loaded into the system memory based on updating the prediction.
16. 15. The system of claim 14, wherein the interaction prediction is processed in part using a behavioral model trained over time to predict the likelihood that the user will interact with the content item or another content item.
17. The system of claim 14 , further comprising the server using the graphics data to render additional detail associated with the content item to enhance the image quality of the content item.
18. 15. The system of claim 14, wherein the prediction of the interaction with the content item is based on processing the gaze information, the gestures, and the interactive data through a behavioral model, the behavioral model configured to identify relationships between the gaze information, the gestures, and the interactive data to generate the prediction of the interaction with the content item.
Citation Information
Patent Citations
Eye-tracking image viewer for digital pathology
CN112805670A
Attention-Based Rendering and Faithfulness
JP2016536699A
Virtual Scene Preload Method
JP2021513702A
Transmodal Input Fusion for Wearable Systems
JP2021524629A
Invoking automated assistant functions based on detected gestures and gaze
JP2021524975A