Hand Gesture Magnitude Analysis and Gearing for Contextually Correct Communication

The system addresses the challenge of miscommunication in virtual reality by tracking and adjusting gesture magnitude based on context, ensuring clear and contextually appropriate input in the metaverse.

JP2026508324APending Publication Date: 2026-03-10SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing systems fail to accurately interpret gesture magnitude and context in virtual or augmented reality environments, leading to potential miscommunication or distraction due to exaggerated or subdued inputs.

Method used

A system that tracks user gestures and facial features using multiple sensors, analyzes gesture attributes, and adjusts input magnitude based on context and interaction environment to ensure appropriate communication in the metaverse.

Benefits of technology

Accurately conveys the urgency or importance of user gestures, avoiding misinterpretation by adjusting inputs to match the context and environment, enhancing communication clarity in virtual or augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026508324000001_ABST
    Figure 2026508324000001_ABST
Patent Text Reader

Abstract

Provides hand gesture magnitude analysis and gearing for contextually correct communication. A method and system for interpreting a gesture provided by a user includes capturing an image of a gesture provided by a user during the user's interaction with a metaverse, analyzing the image to identify attributes of the gesture captured in the image, converting the attributes of the gesture into input based on the context of the user's interaction with the metaverse, and communicating the input to the metaverse for application to the interaction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1. Field of the Invention

[0002] The present disclosure relates to tracking gestures of users interacting with content in the metaverse and interpreting the gestures to generate inputs to apply to the content. [Background technology]

[0003] 2. Description of Related Technology

[0004] With the growth of the Internet, users can interact with a variety of content available in the digital space. Over time, more content becomes available in the digital space, and users can access and interact with the content from a variety of devices. For example, users can access content from mobile devices such as desktop computers, laptop computers, tablet computers, mobile phones, head-mounted displays, and smart glasses. Similarly, users can provide input through a variety of input devices, such as keyboards, mice, touchpads, touchscreens, scanners, wearable and / or handheld devices such as controllers, smart gloves, wearable clothing, smart pens / other smart writing implements, and microphones (for audio input). Users can provide different types of input, such as keystrokes, button presses, hover actions, taps, swipes, and touch input.

[0005] As virtual reality, augmented reality, and the metaverse grow in popularity, the ways in which users interact with each other and with the content presented in the metaverse continue to evolve. The term "metaverse" is a portmanteau of the prefix "meta" (meaning beyond) and "universe." In its simplest form, the metaverse can be understood as a virtual reality space in which users can interact with computer-generated environments and other users. Simply stated, the metaverse is defined using an augmented reality approach by integrating the physical environment with the digital world (i.e., the virtual environment), where the real world is overlaid with virtual reality objects. Users can interact with the metaverse in much the same way as they interact with video games or virtual or augmented reality content.

[0006] As the presentation and sharing of content continues to evolve, so too do the ways in which we interact with that content.

[0007] It is against this background that embodiments of the present invention arise. Summary of the Invention

[0008]

[0003] Embodiments of the present disclosure relate to systems and methods for tracking a user's gestures as the user interacts in the metaverse (or virtual or augmented reality space) and using the gestures to determine gesture magnitude and other attributes. The gesture magnitude is interpreted in the context of the user's communication in the metaverse to define an input. The gesture magnitude is further used to generate a communication signal that conveys the urgency or importance of the gesture. If the gesture magnitude of the user's gesture is large, the input defined from the gesture is amplified to convey the urgency or importance of the communication expressed through or accompanying the gesture. Similarly, if the gesture magnitude is small, the input is attenuated to avoid distraction and / or to convey a message appropriate for communication between the user and the metaverse. To avoid accidentally amplifying or attenuating communications, the user's gestures are tracked over time, and the tracked gestures are used to learn gesture attributes and gesture magnitudes typically expressed by users for various contexts of interaction.

[0009] Correctly interpreting attributes, including gesture magnitude and other gesture details, is useful because such details accurately convey the degree to which content should be adjusted or the user's mood when communicating with other users. Gesture magnitude captures the speed, velocity, physical interaction space the user occupies when performing the gesture, the extent of space to consider for capturing the gesture, etc. In addition to capturing gestures to define gesture attributes and gesture magnitude, the system also captures the user's facial features as the user delivers the gesture. The facial features are used to determine the user's facial expression. The user's facial expression can be used to further confirm the gesture magnitude. For example, the system can determine whether the amount of excitement conveyed by the user's facial expression correlates with the gesture and gesture magnitude. The system can also determine whether the intent of the user's expression corresponds to the context of the interaction, etc.

[0010] The interpreted gesture magnitude and other gesture attributes are used to define input that is communicated to the metaverse to apply to content, communicate with other users, or adjust the facial expressions and gestures of an avatar representing the user in the metaverse. The input is amplified or attenuated according to the urgency or importance conveyed by the gesture so that the input appropriately conveys urgency or importance when applied to the metaverse.

[0011] In one embodiment, a method for interpreting gestures provided by a user is disclosed. The method includes capturing images of a user's gestures as the user interacts with a metaverse. The images are captured using multiple sensors distributed within a physical environment in which the user is manipulating. The images of the gestures are analyzed to identify attributes of the gestures captured in the images. The gesture attributes are evaluated to identify at least a magnitude of the gesture. The gesture attributes, including the magnitude of the gesture, are transformed based on a context of the user's interaction with the metaverse to define an input from the user. The input is communicated to the metaverse for application to content associated with the interaction, and the input is interpreted in the context of the user's interaction with the metaverse.

[0012] In an alternative embodiment, a method for interpreting a gesture provided by a user is disclosed. The method includes tracking movements of a user's body parts. The tracking is performed by capturing images of the user's body parts using multiple sensors distributed within a physical environment in which the user is operating. The images are analyzed to identify attributes of the captured gesture. The attributes are evaluated to identify at least a magnitude of the gesture. The attributes of the gesture, including the magnitude of the gesture, are transformed based on a context of the user's interaction with the metaverse to define an input from the user. The magnitude of the gesture is evaluated in the context of the user's interaction with the metaverse to define a communication signal for conveying an importance of the input. The input is communicated in the communication signal for application to the metaverse. The input is interpreted in the context of the user's interaction and applied according to the importance conveyed by the communication signal.

[0013] Other aspects of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the embodiments described in the present disclosure. [Brief explanation of the drawings]

[0014] [Figure 1] 1 illustrates a simplified block diagram of a system including a wearable device worn by a user while the user interacts with computer-provided content, according to one embodiment. [Figure 2] FIG. 1 illustrates a simplified block diagram of a gesture interpretation engine used to interpret a user's gestures and determine gesture attributes and gesture magnitudes to apply as input to the metaverse, according to one embodiment. [Figure 2A] 3 shows a simplified block diagram of a gesture identification engine of the gesture interpretation engine of FIG. 2 according to an alternative embodiment. [Figure 3-1] 1A and 1B show an image of a gesture performed by a user and an image of a facial feature used to confirm the gesture as the user performs the gesture, according to one embodiment. [Figure 3-2] FIG. 3C shows an image according to an alternative embodiment that captures an alternative gesture performed by a user and uses images of facial features to identify the user expressing a different emotion than that captured in FIGS. 3A and 3B. [Figure 4] 10A and 10B show sample data flows following analysis of a user's gestures to identify inputs to apply in the metaverse based on the context of the communication, according to an alternative embodiment. [Figure 5] 1 illustrates an operational flow of a method for using images of a user's gestures captured by multiple sensors to determine inputs to apply in the metaverse, according to one implementation. [Figure 6] 1 illustrates components of an exemplary system that may be used to process requests from and provide content and assistance to users to carry out aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015]

[0010] Systems and methods are described for adjusting rendering attributes of content presented to a user at a client device. It should be noted that various embodiments of the present disclosure may be practiced without some or all of the specific details. In other instances, well-known process operations have not been described in detail so as not to unnecessarily obscure various embodiments of the present disclosure.

[0016] Various implementations described herein enable a gesture interpretation engine executing on a server computing system to monitor movements of a user's body parts as the user interacts with the metaverse using images of the movements captured by multiple sensors and to identify gestures performed by the user using the images. Gestures provided during the current interaction (i.e., the current gesture) are analyzed to identify gesture attributes. The gesture attributes are evaluated to determine the magnitude of the gesture. The attributes obtained from the analysis of the user's gestures are transformed based on the context of the user's interaction with the metaverse to generate input for application in the metaverse. The user's interaction with the metaverse may be with an interactive application such as a video game, and the input may be game input from the user or feedback regarding the player's gameplay or interaction with another user. Alternatively, a user's interaction in the metaverse may be part of a conversation the user has with another user, either through a social media application rendered in conjunction with a video game, or a messaging or chat application, and input converted from the user's gestures may be used to convey a sense of urgency or importance of the interaction. In some examples, the sense of urgency conveyed by a current gesture may be related to other input provided by the user, such that the sense of urgency conveyed by the current gesture may be applied to other input communicated to the metaverse. In an alternative example, a gesture provided by a user may be a command to a second user to perform a particular activity in the metaverse, possibly with a sense of urgency (e.g., to move in a particular direction, or head in a particular direction, or move quickly, or in a particular manner).

[0017] In some examples, input derived from gestures provided by a user may be used to adjust the corresponding gestures and facial expressions of an avatar used to represent the user in the metaverse. The user's input may be scaled so that the gestures and facial expressions applied to the user's avatar accurately reflect the user's mood, blend with the moods of other users with whom the user is interacting, or are adjusted to avoid being overly exaggerated or overly subdued. Additionally or alternatively, the user's input may be scaled so that exaggerated gestures and facial expressions of a first user can be interpreted in relation to subdued gestures and facial expressions of a second user. The amount of scaling may be determined based on the amount of exaggeration (or subduedness) the user expresses in their gestures. In some examples, the amount of scaling is based on the context of the interaction and / or the environment in which the user is interacting in the metaverse (e.g., excited interactions with friends versus subdued interactions with family, or vice versa) so that the adjustment may blend with the gestures and facial expressions of other users in the metaverse. To determine the amount of scaling that should be performed on a user's gesture, the current gesture for the context is compared to the user's gesture style for the same context. If a match occurs, no scaling is performed. Otherwise, the gesture is scaled, with the amount of scaling based on the amount of deviation of the current gesture from the gesture style recorded for the user. The scaling is performed so that when the input details are applied to the metaverse, the input is not exaggerated and blends easily with the rest of the content. For example, the input may be applied in the metaverse by adjusting the facial expressions and gestures of the user's avatar to convey the facial expressions and gestures provided by the user in the physical world. The gestures applied to the user's avatar are scaled versions (if the gestures were exaggerated) so that the gestures and facial expressions on the user's avatar blend with or are equivalent to the gestures of other users.In some implementations, scaling is performed according to the identities of other users present in the metaverse with which the user is interacting, such that the input, when applied to the user's avatar, is of the same scale of exaggeration as is expressed by the other users' avatars. Capturing and interpreting a user's gestures provides an additional way to provide input and verify interactions in the metaverse.

[0018] It should be noted that throughout this application, the term "metaverse" is used to implement various embodiments described herein. The metaverse can be defined as a virtual realm, where a network of virtual worlds is interconnected with each independently managed virtual world. The metaverse can be accessed through digital avatars using different technologies, such as augmented reality, virtual reality, or mixed reality. Virtual objects used in the metaverse can be tracked using blockchain technology. Interactions in the metaverse employ social media concepts. The metaverse is thus defined as combining the physical world with the virtual world to create an environment that enables multimodal interaction with virtual worlds and real people. Various embodiments and implementations described herein can be extended to user interactions in immersive virtual reality and augmented reality.

[0019] To aid the user in understanding the present disclosure in its overall implementation, certain implementations will now be described in more detail with reference to various figures. It should be noted that various implementations of the present disclosure may be practiced without some or all of the specific details. In other instances, well-known process operations have not been described in detail so as not to unnecessarily obscure various embodiments of the present disclosure.

[0020] FIG. 1 shows a simplified block diagram of an exemplary system in which different devices are used to capture images of gestures provided by a user while interacting with the Metaverse and to analyze the gestures and define inputs for application in the Metaverse. A user can interact with the Metaverse (or virtual or augmented reality environment) using a wearable computing device or computer, such as a desktop or laptop. In the exemplary system shown in FIG. 1 , a user 100 is shown wearing a wearable device (e.g., a head-mounted display (HMD)) 102. The HMD 102 can be worn in a manner similar to glasses, goggles, or a helmet and is configured to allow the user to interact within the Metaverse. In an alternative implementation, instead of the HMD 102, the user 100 may wear smart glasses having a display screen used to interact with the Metaverse, allowing the user to experience the real-world environment and augmented reality content. The user 100 may be interacting with an interactive application, such as a video game, running in the Metaverse, or may be interacting with another user or performing a specific action in the Metaverse. The HMD 102 provides a highly immersive experience to the user by providing a display mechanism in close proximity to the user's eyes. The HMD 102 and glasses may provide a display area for each of the user's eyes that occupies a large portion or even the entirety of the user's field of view. Optics provided within the HMD 102 allow the user to view rendered content in close proximity to the user's eyes. The optics take the user's visual characteristics into account when presenting content to the user.

[0021] In one embodiment, the HMD 102 is connected to a computer. The computer may be a local computer 106 (e.g., a gaming console) or a computer that is part of an application cloud 112 located remotely from the HMD 102. As a result, the connection between the HMD 102 and the computer may be wired or wireless. The computer (106 / or part of the cloud 112) may be any general-purpose or special-purpose computer known in the art, including, but not limited to, a gaming console, a personal computer, a laptop, a tablet computer, a mobile device, a mobile phone, a tablet, a thin client, part of a set-top box, a media streaming device, a virtual computer, a remote server computer, etc. With respect to a remote server computer (i.e., part of the cloud 112), the server computer may be a cloud server in a data center of an application cloud system. The data center includes multiple servers that provide the resources necessary to host one or more interactive applications that are part of the metaverse. These interactive applications provide content to the HMD 102 or other computers / computing devices for rendering. One or more of the interactive applications may be distributed applications that may be instantiated on one or more cloud servers in one data center or distributed across multiple data centers, and if instantiated on multiple cloud servers, the data for each of the interactive applications is synchronized across the multiple cloud servers.In one embodiment, the interactive application may be a video game application (i.e., a virtual reality (VR) application) or an augmented / mixed reality (AR) application, and the computer is configured to run an instance of the video game application or the AR application and to output video and audio data from the video game application or the AR application for rendering on a display screen associated with the HMD 102. In another implementation, the server may be configured to manage one or more virtual machines capable of running an instance of the interactive application (e.g., an AR application or a video game application) and may be configured to provide content for rendering in real time or with a delay.

[0022] Alternatively, a server may include multiple consoles, and an instance of an interactive application may be accessed from one or more consoles (e.g., game consoles). The consoles may be standalone consoles or may be rack-mounted or blade servers. A blade server may similarly include multiple server blades, each having the circuitry and resources necessary to instantiate a single instance of a video game application, for example, to generate a game content data stream. Other types of cloud servers, including other forms of blade servers, may also be involved to run instances of interactive applications (e.g., video game applications) that generate content (e.g., game content data streams, etc.) for users to interact with in the metaverse. Alternatively, the interactive application may be a chat application, or a social media application, or any other application that can receive interactions from users and share them with other users. In some implementations, a user's gestures may be used to adjust an avatar representing the user in the metaverse (e.g., the avatar's facial expressions and gestures), or to share content through the user's avatar with avatars of other users present in the metaverse, or may be in the form of reactions to content or comments from or about other users. Interactions between the user and other users may be about content presented in the metaverse, or about part of a conversation related to the content or other users, or to provide additional information related to the content or other users.

[0023] Gestures provided by a user while interacting with the Metaverse are captured using image capture devices such as a camera 108 located on the exterior of the HMD 102 or on the exterior of other devices such as the computer 106, controller 104b, one or more external cameras (109) located in the physical environment outside the HMD 102 and computer 106, or an internal camera (a camera located on the interior surface of the HMD 102) (not shown). Images capturing finger gestures are used to determine the location, position, orientation, and / or movement of the user 100, the HMD 102, the glove interface object 104a, the controller 104b, and / or other input devices.

[0024] The controller 104b may be a one-handed controller (i.e., a controller that can be operated using one of the user's hands) or a two-handed controller (i.e., a controller that can be operated using both of the user's hands). The glove interface object 104a, the controller 104b, and other wearable and input devices may be tracked by tracking visual indicators, such as lights, tracking shapes, etc., associated with or disposed on such devices to determine their respective locations and orientations. Additionally or alternatively, various input and wearable devices (e.g., HMD 102, smart glasses, glove interface object 104a, controller 104b, etc.) may be tracked using embedded sensors.

[0025] Embedded sensors located in various input and wearable devices worn and / or operated by the user may further be used to track and capture data regarding the user's finger gestures, facial features, and movements in the physical environment. Data captured by the sensors and image capture devices may be used to determine the user's location and orientation, including the location and orientation of the HMD 102, glove interface object 104a, controller 104b, etc., as well as the user's facial features and finger movements, when the user manipulates the glove interface object 104a and holds and manipulates the controller 104b. From details of the data captured by the sensors and image capture devices, finer details related to the movements of the user's body parts may be determined. For example, finer details that may be determined from the captured data may include the identity and finger positions of the user's specific fingers wearing the glove interface object 104a, the identity of the fingers used to hold the controller, the pressure applied by each finger holding the controller, the identity of the fingers used to provide input using the controls on the controller, the pressure applied when providing input, the user's location and orientation in the physical environment when interacting in the metaverse, etc. Details from finger gestures and facial features can be interpreted to define inputs to apply to the metaverse to control, navigate, interact with, participate in, and / or interact with other users present in the metaverse, a virtual reality space contained in the metaverse.

[0026] In addition to the camera 108 and other image capture devices, one or more microphones located on the HMD 102, the controller 104b, the computer 106, and elsewhere in the physical environment may be used to capture audio from the interactive environment. The audio captured by the one or more microphones may be processed to identify the location of the audio source (e.g., audio produced by the user, other users, interactive applications, or other sources). The audio captured by the microphones at the identified location may be selectively filtered, utilized, or processed to reject other audio not from the identified location. For example, audio produced by a user may be selectively filtered and used to confirm gestures. The filtered audio may also be provided as input to the metaverse.

[0027] In addition to capturing finger gestures, image capture devices and sensors distributed in the physical environment on various devices are also used to capture images of the user's facial features and details of changes in the user's facial features as the user interacts with the metaverse. The user's facial features are used to verify the finger gestures captured by the image capture devices. The images of the finger gestures and facial features are used to determine the location, position, orientation, and / or movement of the user 100, HMD 102, glove interface objects 104a, controllers 104b, and / or other input devices, as well as real-world objects located in the physical environment with which the user is manipulating. The finger gestures and facial feature details are forwarded to a client-side gesture interpretation engine 200a running on the computer 106 and / or a server-side gesture interpretation engine 200b running on a server in the application cloud 112.

[0028] In some embodiments, the computer 106 functions as a thin client in communication over the network 110 with a server computing device on the application cloud 112, and a cloud-side gesture interpretation engine 200b running on the server is used to process gesture-related data transferred by the computer 106. The server computing device (hereinafter simply referred to as the "server") may also be configured to run one or more interactive or augmented reality applications accessible within the metaverse. The interactive application may be a video game application selected for gameplay by the user 100. One or more servers of the application cloud 112 may be responsible for maintaining and running an instance of the video game using the respective server's processor or for instantiating the video game on the computer 106. For example, the computer 106 may be a local network device, such as a router, that typically facilitates the passage of network traffic rather than performing finger gesture and / or video game processing.

[0029] In one embodiment, the HMD 102, glove interface object 104a, controller 104b, and image capture devices such as camera 108, external camera 109, etc., may themselves be networked devices that connect independently and directly to a network 110 to communicate gesture-related data captured by the sensors and image capture devices to a server in an application cloud 112. The connections of the HMD 102, glove interface object 104a, controller 104b, and camera 108 to the network may be wired or wireless connections. Some of the sensors used to capture data regarding finger gestures include inertial measurement unit (IMU) sensors such as accelerometers, magnetometers, gyroscopes, and global positioning system trackers. Image capture devices used to capture images of finger gestures and facial features include one or more forward-facing cameras 108, multiple cameras or image capture devices (e.g., a stereo pair of cameras) positioned on the inner surface of the HMD 102 and pointed at the user's eyes and other facial features, an IR camera, a depth camera, an external camera 109 positioned in the physical environment and facing the user, or a combination of any two or more thereof.

[0030] In some implementations, gesture interpretation engine 200a in computer 106 performs preliminary processing of image and sensor data to identify the user's finger gestures and facial features captured as the user interacts in the metaverse, identify the specific one of the fingers used to provide the gesture and the specific one of the facial features used to provide an expression on the user's face, the extent of the finger and facial feature movement (including direction, amount of change, etc.), and forwards the identified details related to the finger gestures and facial features to gesture interpretation engine 200b running on a server in application cloud 112 for further processing. Gesture interpretation engine 200b processes the received finger gesture and facial feature details provided by gesture interpretation engine 200a to identify finger gesture attributes including gesture magnitude and uses details from the facial features to confirm the finger gesture and gesture magnitude. In an alternative embodiment, processing of finger gestures and facial features is performed by gesture interpretation engine 200a on computer 106 itself, and the gesture attributes, gesture magnitude, and facial feature images identified by gesture interpretation engine 200a are forwarded to gesture interpretation engine 200b on application cloud 112, where the finger gesture details are ascertained and interpreted to define inputs that can be applied to the user's interactions in the metaverse.

[0031] 2 illustrates various components of an exemplary gesture interpretation engine 200 that, in one implementation, are used to track a user's finger gestures and facial features and to interpret the captured finger gestures and facial features to define inputs for application to interactions within the metaverse. To assist in capturing and processing finger gestures and facial features, the gesture interpretation engine 200 includes multiple components. Some of the components include an image analysis engine 210, a gesture identification engine 220, a context analysis engine 230, a gesture verification engine 240, a gesture scaling engine 250, and a gesture interpretation engine 260. In addition to the various components, the gesture interpretation engine 200 also interacts with one or more databases to retrieve associated content stored internally for processing the gesture and facial feature data. For example, the gesture interpretation engine 200 may query and receive user profile data from a user profile data store (not shown) and metaverse-served content from a metaverse content data store (not shown), to name a few.

[0032] The gesture interpretation engine 200 receives images of a user as they interact with the metaverse and are captured by one or more sensors and image capture devices, and processes the images to identify and interpret gestures. An image analysis engine 210 within the gesture interpretation engine 200 analyzes the images to identify the captured user's body parts and determine whether the captured user's body parts are tracked, and if so, whether the tracked parts are being used to perform gestures. Once the image analysis engine 210 verifies and confirms that the user is performing gestures, details of the gestures captured in the images are forwarded to the gesture identification engine 220.

[0033] The gesture identification engine 220 receives gesture details from the image analysis engine 210 and performs further analysis of the image to identify various attributes (220a) of the gesture captured therein. The attributes are determined by tracking the movement of a particular part or parts of the user's body used to provide the gesture, such as a finger, arm, or facial feature, and the identity of the body part for tracking may be specified in the gesture interpretation engine 200. The gesture may be a finger gesture, an arm gesture, a facial gesture, an input gesture provided using one or more input devices or input surfaces, or the like. Depending on the type and amount of data the image capture device and sensor are configured to capture, the captured data may, in some implementations, include details of finger gestures only, finger gestures and arm gestures, or finger gestures, arm gestures, and facial gestures. In some implementations, data associated with finger and hand gestures is used to define gesture attributes, and facial gestures are used to further confirm the finger and hand gesture attributes. Each of the image capture and input devices used to provide the gestures is connected to the computer 106 and / or a server of the application cloud 112 via the network 110 using a wired or wireless connection so that details of the captured gestures can be sent to the gesture interpretation engine 200 running on the respective computer 106 or cloud server for further processing.

[0034] FIG. 2A illustrates some of the modules within the gesture identification engine 220 that are used to track the movement of a user's body parts in the physical environment while interacting within the Metaverse and to use details from the tracking to identify the type and extent (i.e., magnitude) of gestures provided by the user. Tracking a body part involves tracking and evaluating changes in the position, orientation, or location of the body part due to the body part's movement, where changes in the body part's position, orientation, or location may correspond to changes in the user's position, orientation, or location. The changes are tracked by first identifying an interactive space in which the user's body part is contained and then tracking the movement of the body part relative to the identified interactive space. The interactive space identification engine 222 of the gesture identification engine 220 is used to identify a space within the physical environment in which the user is present and then identify the portion of the space in which the tracked user's body part is contained. The interactive space identification engine 222 identifies the space occupied by the user by mapping the user's physical environment to identify the locations of various real-world objects contained therein and then mapping the user's location in relation to the real-world objects. The mapping is a three-dimensional mapping, and therefore the interactive space is defined as a three-dimensional space. The interactive space identification engine 222 may use the three-dimensional mapping of the interactive space to track different body parts of the user used to provide gestures. Details of defining an interactive space for tracking body parts of the user are described in more detail with reference to Figures 3A-3C.

[0035] The interactive space, in some embodiments, is defined using a virtual bounding box. The dimensions of the virtual bounding box are defined to be sufficient to encompass and completely contain the user's body part being tracked. Because the interactive space is defined to be a three-dimensional space, the virtual bounding box, in some embodiments, is defined to be three-dimensional and represented using (x, y, z) coordinates. In an alternative embodiment, the interactive space is defined as a two-dimensional space, and the virtual bounding box is defined to be two-dimensional and represented using (x, y) coordinates. The interactive space changes in response to movement of the user's body part, and the virtual bounding box is defined to move with the body part so that the virtual bounding box accurately represents the user's current interactive space.

[0036] Details of the interactive space are provided to the motion tracking engine 224 of the gesture identification engine 220 so that the interactive space can be monitored to detect any changes indicative of movement of the user's body parts. The motion tracking engine 224 analyzes images of the user captured by the image capture device and sensor data from the sensors to determine any change(s) in the position and / or orientation of the user and the user's body parts contained within a virtual bounding box. Changes in the position and / or orientation of the body parts may result in changes in the size and / or position of the virtual bounding box. When a change is detected in the position and / or orientation of one or more of the user's body parts, the motion tracking engine 224 calculates the extent of such change by (a) calculating the difference in coordinates of the virtual bounding box before and after the change, and (b) calculating the difference in coordinates of the body parts contained within the virtual bounding box. In some implementations, the user can make gestures that are well contained within the original (i.e., initial) virtual bounding box defined for the user. This may be true for users who typically make small gestures. In such cases, the change in the movement of the body part is determined by calculating the difference in the coordinates of the body part before and after providing the gesture. In other implementations, the user can perform a gesture that results in the body part moving out of the initial interactive space defined for the user (i.e., outside the initial virtual bounding box). This may be the case for users who perform large gestures (i.e., move their body part (e.g., hand, arm, etc.) in large amounts). In such cases, the change in the movement of the body part can be calculated as the difference in the respective coordinates of the virtual bounding box and the body part before and after the gesture.

[0037] The calculated changes in the coordinates of the virtual bounding box and the coordinates of each of the body parts are forwarded by the motion tracking engine 224 to the motion estimation engine 226. The motion estimation engine 226 of the gesture identification engine 220 evaluates the changes to determine various attributes of the gesture provided through the movement of the body parts. For example, the motion estimation engine 226 analyzes the changes caused by the movement of the body parts to identify attributes of the gesture. The identified attributes may include observed attributes and derived attributes. Some of the observed attributes identified by the movement evaluation engine 226 may include the identity of each body part used to provide the gesture, the type of change detected (e.g., a wave, swipe, tap, clap, fist bump, fist shake, etc. detected from hand and arm changes), the extent of such movement of the body part, the direction of the movement, the speed of the movement, the space occupied by the user when providing the gesture, the pressure applied while holding the input device (e.g., a controller), the pressure applied on a control key or button located on the surface or on the input device, the extent of movement of a facial feature (if the body part is a facial feature), etc. The observed attributes are used to define derived attributes. For example, the speed of finger movement is calculated from the direction and speed of finger movement across a touchscreen. In another example, the user's mood can be inferred from the pressure applied when providing input. For example, an excited player may grip the controller too tightly or may use a lot of pressure to provide input. Additionally, the user's mood can be inferred from the movement of one or more facial features. For example, observed player attributes may indicate that the player's eyes are wide and brows are furrowed to indicate that the player is fully engaged with the metaverse content. From the above example, an excited player may be seen holding the controller tightly or pressing the controller's buttons or controls too hard. The movement evaluation engine 226 may interpret the finger movements and the pressure applied to the controller's surface as the user holds it to calculate the grip force.Additionally, pressure applied while holding the controller, providing input to the controller's touch surface, or interacting with each control on the controller may be used to determine the user's mood. The user's mood may be verified by analyzing the user's observed facial features. Thus, the user's observed and derived attributes may be used by the movement evaluation engine 226 to determine the type of gesture presented, the gesture magnitude, the user's facial expression intent, the gesture intent, the user's mood, etc. The gesture attributes and gesture magnitude are used to verify the gesture.

[0038] Referring again to FIG. 2 , the gesture attributes and gesture magnitudes identified from the image and sensor data are used by gesture verification engine 240 to verify the user's gesture and determine whether the gesture is appropriate in the context of the interaction. The context analysis engine 230 is first used to determine the context of the user's interaction with the metaverse. For example, the user may be interacting with a video game application running in the metaverse, and the interaction may be as a player or a spectator. In this example, the context may be related to the video game, and more specifically, to a video game scene. Alternatively, the interaction may be in the form of the user engaging in a conversation with another user or group of users present in the metaverse, and the user may be communicating with the other user or group of users via their respective avatars. The input may communicate additional information related to particular content that is the subject of a conversation between users, or communicate a particular level of importance of a particular action that a user is performing via an avatar, or communicate instructions to a second user to perform a particular activity and the urgency level for performing the particular activity within the metaverse. In another example, a user may be interacting with shared content or providing content for sharing with other users. The context analysis engine 230 analyzes the content that the user is interacting with within the metaverse and identifies the context of the interaction. The interaction context and content are forwarded as input to the gesture verification engine 240.

[0039] The gesture verification engine 240 analyzes the attributes and gesture magnitude provided by the gesture identification engine 220 against the context and content provided by the context analysis engine to determine whether the gesture is in accordance with the user's interaction context. Additionally, the gesture verification engine 240 determines whether the user's gesture magnitude is exaggerated or subdued in relation to the interaction context. In some implementations, the user's gesture attributes and gesture magnitude are verified against the user's own previous gestures to determine whether the current gesture provided by the user is exaggerated, subdued, or consistent with the user's own previous gestures. To understand whether the user's gesture is exaggerated, subdued, or consistent, the gesture verification engine 240 examines the user's current gesture in the interaction context to determine whether the current gesture is appropriate for the interaction context. If appropriate, the current gesture is compared to previous gestures provided by the user for the same context. The user's previous gestures may be stored in a gesture data store (not shown) and provided to gesture verification engine 240 when the user's gestures need to be verified. In some implementations, the user's typical gestures for different contexts may be stored in a user profile data store (not shown) and retrieved for verification. If the user's current gestures are found to be appropriate for the interaction context and are detected as exaggerated or subdued, gesture verification engine 240 examines the user's interaction environment in the metaverse to determine whether the environment requires such exaggerated (or subdued) gestures from the user. For example, a user may exaggerate their gestures when with friends while watching a sporting event and subdue their gestures when with family or in a professional environment.Thus, by examining a user's gestures based on the environment, the gesture verification engine 240 can determine whether the user's exaggerated (or suppressed) gestures are appropriate in the context of the environment.

[0040] In some implementations, when a user's interaction is with one or more other users present in the metaverse, gesture verification engine 240 may examine the user's gesture attributes and gesture magnitude in the context of the interaction with the other users to determine whether the user's gesture is exaggerated, appropriate, or subdued. A gesture is deemed appropriate if the gesture attributes and gesture magnitude of the user's current gesture match the corresponding gesture attributes and gesture magnitude of other users present in the metaverse with whom the user is interacting. If the gesture attributes and gesture magnitude are appropriate or subdued, gesture verification engine 240 deems the user's response appropriate in the context of the interaction. However, if the gesture attributes and gesture magnitude are exaggerated, gesture verification engine 240 flags the gesture as exaggerated and should scale it down. Thus, the gesture validation engine 240 validates the gesture attributes and gesture magnitude of the user's current gesture in the context of the interaction, including the context of the content, the context of the user's previous gestures, the context of the environment the user is in, and the interaction context of other users who exist in the metaverse and with whom the user is interacting to determine whether the gesture is appropriate or if the gesture needs to be scaled down.

[0041] In some implementations, gesture attributes, including the interaction context, content, the environment in which the user is interacting in the metaverse, and the magnitude of the gesture, are provided to a machine learning algorithm within the gesture verification engine. The machine learning algorithm uses the context, content, environment, and gesture features identified for the media content to build and train an artificial intelligence (AI) model and generate an output that identifies the expressed intent of the user providing the gesture. Gestures provided by other users as they interact with different content, contexts, and environments in the metaverse are also used in training the generated AI model. The AI ​​model is improved as additional gestures are provided by users and as changes are detected in the content, context, and environment in which the user is interacting in the metaverse. The machine learning algorithm uses a classifier to classify various data and update the AI ​​model. The AI ​​model uses the gesture attributes of the user and other users, the context of the media content, the content, and the environment to generate an output corresponding to the current gesture provided by the user, where the output corresponds to the particular context of the interaction, the particular content, the particular events occurring within the content, and the particular environment in which the user is immersed in their interaction within the metaverse. The AI ​​model's determined output defines the user's expressed intent, which may include the user's mood, the urgency or importance the user wishes to communicate in the metaverse, etc. The AI ​​model's output becomes the input that needs to be applied to the metaverse.

[0042] Gesture details, including the user's and other users' gesture attributes, gesture magnitude, interaction context, interaction environment, and flags set to indicate exaggerated gestures, are provided as input to gesture scaling engine 250. Gesture scaling engine 250 uses the details provided by gesture validation engine 240 to scale the user's current gesture by adjusting the gesture attributes and gesture magnitude so that the gesture is deemed appropriate for the interaction context. The amount of scaling is determined according to the interaction context, the metaverse environment with which the user is interacting, and the amount by which the user's current gesture is exaggerated. The adjusted gesture attributes and gesture magnitude are forwarded by gesture scaling engine 250 to gesture interpretation engine 260. Gesture interpretation engine 260 interprets the gesture attributes and gesture magnitude in the interaction context to generate input for application to interactions in the metaverse. The input generated by gesture interpretation engine 260 is forwarded to input application engine 270. Input application engine 270 determines the context of the user's interaction with the metaverse, identifies the content or application or entity to which the input should be applied in the metaverse, and forwards the input accordingly.In some implementations, the input may be applied in the metaverse as additional information to content in the metaverse, as a comment or feedback to another user, as a response from the user to a conversation between the user and another user, as instructions to the user's avatar interacting with content or with avatars of other users present in the metaverse, as input to an interactive application available in the metaverse (e.g., a video game, or a chat, or a social media application), or to adjust the facial expression of the user's avatar to mimic the user's facial expression, or to adjust the user's avatar to behave in a manner similar to or appropriate within a group of other users' avatars, to name a few. If the input is to an interactive application, the input may be interpreted as defining an amount of adjustment to be made to the interactive application to affect a state. For example, the interactive application may be a video game application. In this example, the input may be interpreted as defining an amount of adjustment to be made to the user's input to generate a current state. In another example, the interactive application may be a sign language communication application. In this example, the input may be interpreted as defining an amount by which sign language must be adjusted to convey a level of urgency or importance defined using gestures. In yet another example, the interactive application may be a communication application, such as a social media application or a chat / messaging application. In this example, the input is interpreted to convey the importance of content shared by the user within the communication application. Thus, the input generated by interpreting the user's gestures and scaled according to the content, context, environment, and / or situation of the metaverse is applied according to the context of the user's interaction with the metaverse. The input from the interpretation of the gestures may be in addition to the input provided by the user through normal channels, such as an input device.

[0043] 3A-3C show some examples of interactive spaces identified within a user's physical environment in some implementations. The interactive spaces are defined to include various body parts of the user identified for tracking. In one implementation, FIGS. 3A and 3B show tracking of a user's hand gestures and facial features, where the tracked facial features indicate the user is expressing a particular emotion (e.g., Emotion 1), and FIG. 3C shows tracking of a user's hand gestures and facial features, where the tracked facial features indicate the user is expressing a different emotion (e.g., Emotion 2), in another implementation. The gesture interpretation engine 200 tracks a user's body parts as the user interacts with the metaverse by first identifying an interactive space within the physical environment in which the user is operating. In some implementations, because the physical environment is represented in three dimensions, the interactive space represented by the virtual bounding box is also defined in three dimensions. Thus, the virtual bounding box is represented using (x, y, z) coordinates. In an alternative implementation, the virtual bounding box representing the interactive space is represented in two dimensions using (x, y) coordinates. In such an implementation, the three-dimensional interactive space is transformed into two dimensions, and a virtual bounding box is defined in terms of (x, y) coordinates. The size of the bounding box is defined to include the user's tracked body parts, such as their arms, hands, and face. In FIG. 3A , bounding box “B1” represents the interactive space at time “t1” and is shown sized to include the user's upper body parts, including their arms, hands, and face. Furthermore, for simplicity, bounding box B1 is represented as a rectangle in two dimensions, but it can be represented to adopt any other shape besides a rectangle and can also be represented in three dimensions. As the user continues to interact with the metaverse, the user's face, arms, and / or hands continue to be tracked, and as movement is detected in one or more of the user's body parts, the gesture interpretation engine 200 detects the change(s) and dynamically adjusts the bounding box containing the user's tracked body parts from B1 to B2.

[0044] FIG. 3B illustrates changes in the interactive space due to movement of a user's hand and arm, resulting in a dynamic adjustment of the bounding box from B1 to B2 to fully encompass the arm. B2 represents the interactive space at time "t2." In the diagrams shown in FIGS. 3A and 3B, bounding boxes B1 and B2 are shown to have the same size, and the coordinates representing bounding box B2 have changed from the coordinates of bounding box B1 in terms of the horizontal coordinate. Depending on the degree of movement of the user's body parts, the horizontal and vertical coordinates representing the bounding box can change size, and the coordinates representing the bounding box are dynamically adjusted to reflect this. The coordinates of the virtual bounding boxes B1 and B2 are used to determine changes in the user's gesture due to movement of one or more body parts. The gesture changes are captured in an image by an image capture device, and movement data related to the body part movements is captured by a sensor. The gesture provided by the user is verified against changes in the user's facial features to determine whether the changes in the gesture correlate with corresponding changes in facial expressions expressed by the user.

[0045] FIG. 3C shows another example of an interactive space including a user's body parts as the user performs a gesture in an alternative embodiment. In this example, by tracking the user's facial features, it is possible to determine that the user is expressing a different emotion (i.e., anger) as evidenced by the user shaking a clenched fist in the air and the user's angry facial expression. As can be seen in FIG. 3C, the size of the interactive space represented by bounding box "B3" is much smaller than the bounding boxes B1 and B2 in FIGS. 3A and 3B. Thus, the interactive space can be defined based on the type of gesture provided by the user and the extent to which the user moves their arms and hands while performing the gesture, and the image capture device and sensor use the defined interactive space to track the movement of the user's body parts.

[0046] 4A illustrates the flow of data as, in one embodiment, image and sensor data capturing a user's gestures are processed by gesture interpretation engine 200 to generate input for application to interactions in the metaverse. Images of the user obtained from an image capture device and sensor data from multiple sensors are analyzed to determine whether the user is engaged in providing a gesture (410), and if so, gesture attributes are identified. The gesture attributes are evaluated to determine gesture magnitude. The gesture attributes and gesture magnitude are interpreted and verified (420) in the context of the interaction. Additionally, the gesture attributes and gesture magnitude are compared to the user's own previous gestures for the same or similar context of the interaction to determine whether they are substantially consistent or exaggerated. If the user's gesture is determined to be exaggerated relative to the context of the interaction (e.g., compared to gestures of other users or relative to the context of the content), the gesture attributes and magnitude are dynamically scaled 430 so that a gesture corresponding to the scaled-down gesture attributes and gesture magnitude is appropriate for application to an interactive application in which the user is interacting with that content in the metaverse. The scaled-down gesture attributes and gesture magnitude of the gesture are transferred to the metaverse for application to an interactive application, such as a video game application, where the gesture is provided as input to the video game, or to adjust the user's avatar, or to convey information (e.g., the gesture) to other avatars via the user's avatar.

[0047] FIG. 4B illustrates an alternative data flow when, in an alternative embodiment, image and sensor data capturing a user's gestures are processed by gesture interpretation engine 200 to generate input for application to interactions in the metaverse. The data flow in FIG. 4B differs from the flow illustrated in FIG. 4A at operations 425 and 430′. Accordingly, data flows common to FIGS. 4A and 4B are represented using the same reference numerals and will not be described in detail. As shown in FIG. 4B, after the gesture is interpreted and verified, gesture attributes associated with the gesture and gesture magnitude are compared with those of other users (425). The gesture magnitude and gesture attributes of the other users are determined relative to the context of the interaction and used in the comparison. If the comparison indicates that the gesture attributes and gesture magnitude of the current gesture are exaggerated relative to other users present in the video game (e.g., the user's gestures are large (i.e., waving their hands widely)), the gesture attributes and gesture magnitude associated with the current gesture are scaled down (430') so that the resulting magnitude and scale of the user's gesture matches the magnitude and scale of the other users' gestures. The scaled-down gesture is then interpreted to generate input to apply to the user's avatar to adjust the user's avatar's facial expressions and gestures in the metaverse.

[0048] 5 illustrates an operational flow of a method used in one implementation to process images of gestures made by a user and sensor data related to the gestures as the user interacts with content in the metaverse and interpret the user's gestures in the context of the metaverse interactions. The method begins at operation 510, where images of the user making gestures as the user interacts in the metaverse are captured using an image capture device and sensor data is collected from various sensors distributed in the user's physical environment. The captured images are of the user's body parts being tracked. As shown in operation 520, the images and sensor data are analyzed to identify gestures made by the user. The analysis occurs in the context of the interaction, where the user's interaction may be with an interactive application such as a video game, or a social media application, or a chat / messaging application, and the user may upload video content, gestures, text, GIFs (graphics interchange format files), graphics, audio content, or any other type of content that can be uploaded, shared with or by other users, or shared with other users' avatars in the metaverse or the user's own avatar. Consequently, the analysis occurs in the context of the content the user is interacting with and / or the context of the environment the user is interacting with and / or the context of other users with whom the user is interacting in the metaverse.

[0049] Once a gesture is identified, the user's gesture is evaluated to determine gesture attributes from which the gesture magnitude is calculated. The gesture attributes and gesture magnitude are compared to the user's previous gestures to determine whether the current gesture is exaggerated, and if so, whether the exaggeration is justified in the interaction context. If the gesture is justified in the interaction context, the gesture is considered as is without any adjustment. On the other hand, if the gesture is determined to be exaggerated in the interaction context, the gesture attributes and gesture magnitude are attenuated (i.e., scaled down) according to the interaction context. The amount of attenuation here is driven by the gesture magnitude and the amount of exaggeration determined from the interaction context. On the other hand, if the gesture is determined to be too restrained, the gesture attributes and gesture magnitude are amplified (i.e., scaled up) according to the interaction context. The amount of amplification here is driven by the magnitude of the gesture and the amount of suppression determined from the context of the interaction. In some implementations, the user's facial features may be analyzed to ascertain the gesture magnitude (i.e., the level of exaggeration or suppression expressed in the gesture) and the gesture magnitude may be attenuated or amplified based on the ascertained gesture magnitude. By attenuating or amplifying the gesture attributes and gesture magnitude, the gesture is within a level of expression appropriate for the interaction. In some implementations, when a gesture is exaggerated and the gesture magnitude indicates that the level or amount of gesture exaggeration exceeds a predefined threshold, the gesture interpretation engine 200 interprets the gesture magnitude and transmits a first signal indicating a level of urgency to be communicated in the input generated by interpreting the gesture. The level of urgency communicated in the input is amplified commensurate with the gesture magnitude of the gesture.Similarly, if the level or amount of gesture suppression is below a predefined threshold, the gesture interpretation engine 200 interprets the magnitude of the gesture and sends a second signal along with the input to avoid distracting the user during interactions in the metaverse.

[0050] As shown in operation 530, the gesture attributes and the magnitude of the gesture (either scaled down or original) are converted into input for application to the metaverse according to the context of the interaction. The input is then communicated to the metaverse for application to the content associated with the interaction, as shown in operation 540. The input applied to the content conforms to the content's context, the environment, and interaction standards defined by other users present in the metaverse, such that the user's input blends with other users and is not too loud, indiscreet, or flashy. For example, the input may be applied to adjust facial expressions and gestures.

[0051] In summary, various embodiments described herein are directed to interpreting gestures provided by a user interacting within the metaverse to generate input that can be applied to content in the metaverse. The user's movements are tracked and learned over time to avoid erroneously attenuating or exaggerating gesture communications based on the magnitude of the captured movements. The user's gestures may be finger gestures, and images can be used to detect which of the user's fingers are being used and which gestures are being provided by each finger and finger combination. The finger gestures may be interpreted as input into sign language or used to communicate information to another user in the metaverse. Communications may be translated and rendered through the user's avatar, to another avatar of another user, or to an avatar of the user's team. Finger gestures are interpreted to define a gesture magnitude, which may include attributes such as speed, direction, the space the user occupies, the amount of space considered to capture the gesture, the amount of pressure applied, and where the pressure is applied (e.g., on the controller handle when held, a button, an input surface, a joystick, etc.). From these attributes, other attributes such as the amount of excitement, the meaning of the user's facial expression, and the speed associated with the gesture may be calculated / derived. The gesture magnitude is interpreted to understand the urgency or importance or excitement of the user providing the gesture and other input, and to translate that excitement or importance into input to apply to the user's avatar or interactive application or communication between the user and other users via the respective avatar. Gestures are interpreted in the context of the content, type of environment, and type of user with which the user is interacting, to avoid erroneously attenuating or amplifying gesture attributes. Additionally, the gesture interpretation engine finds a way to interpret the gestures of a user who gestures in an exaggerated manner in a similar way to how it interprets the gestures of a user who gestures gently.The gestures are interpreted according to the content, context, environment, and other user(s) present in the environment, and the interpreted gestures are checked against other input cues provided by the user, such as facial features to provide expression, voice to provide speech (expressing emotion), and gesture speed to ensure that the user's facial expressions correlate with the observed and derived gesture attributes.

[0052] A user's gestures are dynamically adjusted as the user moves from one location to another (e.g., from a first game scene to a second game scene in a video game application, from a first interactive application to a second interactive application, etc.) such that the adjusted gestures allow the user or the user's avatar to blend with the environment when changes occur with respect to the presented content. Changes to the content may include changes to the scene, the content, the presence of other users, etc. In some implementations, the gesture interpretation engine 200 uses profiles of other users' avatars in the metaverse space with which the user is interacting to scale the user's own avatar so that the behavior of the user's avatar optimally blends with the behavior of the other avatars and / or the context occurring in the space within the metaverse. This type of dynamic adjustment of gestures not only allows the user's avatar to blend with the environment, but also ensures that the behavior (e.g., gestures, actions, etc.) of the user's avatar avoids obscuring the user.

[0053] The gesture interpretation engine analyzes a user's gestures over a period of time, recognizes the user's gesture style, and identifies specific patterns for different contexts of content and situations in which the user interacts in the metaverse. When the user provides a new gesture, the gesture interpretation engine performs a gesture recognition mapping with previous gestures to determine how close the user's current gesture is to the average gesture for the same or similar context defined by the previous gesture. Using the gesture recognition mapping, the gesture interpretation engine can quickly and efficiently recognize gestures and determine whether the user's gestures are exaggerated or suppressed. Based on this determination, the gesture interpretation engine can perform appropriate scaling to fit the environment and context associated with the user's interaction in the metaverse.

[0054] FIG. 6 is a block diagram of an exemplary gaming system 600 that may be used to provide content to an HMD for user consumption and interaction, according to various embodiments of the present disclosure. The gaming system 600 is configured to provide video streams to one or more clients 610 over a network 615, one or more of which may include an HMD (102), glasses, or other wearable devices. In one implementation, the gaming system 600 is depicted as a cloud gaming system, where an instance of a game runs on a cloud server and content is streamed to the clients 610. In an alternative implementation, the gaming system 600 may include a gaming console that runs an instance of a game and provides streaming content to the HMD for rendering. The gaming system 600 typically includes a video server system 620 and an optional game server 625. The video server system 620 is configured to provide video streams to one or more clients 610 with a minimum quality of service. For example, the video server system 620 may receive game commands that change the state of a video game or a viewpoint within a video game, and provide an updated video stream that reflects this change to the client 610 with minimal latency. The video server system 620 may be configured to provide video streams in a variety of alternative video formats, including formats that have not yet been defined. Furthermore, the video streams may include video frames configured for presentation to a user at a variety of frame rates. Typical frame rates are 30 frames per second, 60 frames per second, and 120 frames per second. However, alternative embodiments of the present disclosure include higher or lower frame rates.

[0055] The clients 610, individually referred to herein as 610A, 610B, etc., may include head-mounted displays, terminals, personal computers, game consoles, tablet computers, phones, set-top boxes, kiosks, wireless devices, digital pads, standalone devices, handheld gameplay devices, and / or the like. Typically, the clients 610 are configured to receive encoded video streams, decode the video streams, and present the resulting video to a user, e.g., a player of a game. The process of receiving the encoded video stream and / or decoding the video stream typically involves storing individual video frames in a receive buffer of the client. The video stream may be presented to the user on a display integrated into the client 610 or on a separate device such as a monitor or television. The clients 610 are optionally configured to support more than one game player. For example, a game console may be configured to support two, three, four, or more simultaneous players. Each of these players may receive a separate video stream, or a single video stream may include regions of frames generated specifically for each player, e.g., generated based on each player's point of view. The clients 610 are optionally geographically distributed. The number of clients included in the gaming system 600 can vary widely, from one or two to thousands, tens of thousands, or more. As used herein, the term "game player" is used to refer to a person who plays a game, and the term "gameplay device" is used to refer to a device used to play a game. In some embodiments, gameplay devices may refer to multiple computing devices that work together to deliver a gaming experience to a user. For example, a game console and an HMD may work with the video server system 620 to deliver a game for viewing through an HMD.In one embodiment, the game console receives the video stream from the video server system 620, and the game console forwards the video stream or updates to the video stream to the HMD for rendering.

[0056] Client 610 is configured to receive the video stream via network 615 (110 in FIG. 1). Network 615 can be any type of communications network, including a telephone network, the Internet, a wireless network, a powerline network, a local area network, a wide area network, a private network, and / or the like. In a typical embodiment, the video stream is communicated via a standard protocol such as TCP / IP or UDP / IP. Alternatively, the video stream is communicated via a proprietary standard.

[0057] A typical example of a client 610 is a personal computer that includes a processor, non-volatile memory, a display, decoding logic, network communication capabilities, and input devices. The decoding logic may include hardware, firmware, and / or software stored on a computer-readable medium. Systems for decoding (and encoding) video streams are well known in the art and vary depending on the particular encoding scheme used.

[0058] The clients 610 may, but need not, further include a system configured to modify the received video. For example, the clients may further be configured to render, overlay one video image on another, crop video images, and / or the like. For example, the clients 610 may be configured to receive various types of video frames, such as I-frames, P-frames, and B-frames, and process these frames into images for display to a user. In some embodiments, members of the clients 610 are configured to perform operations on the video stream, such as further rendering, shading, or 3D transformation. Members of the clients 610 are optionally configured to receive more than one audio or video stream. Input devices for the clients 610 may include, for example, a single-handed game controller, a two-handed game controller, a gesture recognition system, an eye-gaze recognition system, a voice recognition system, a keyboard, a joystick, a pointing device, a force-feedback device, a motion and / or position sensing device, a mouse, a touchscreen, a neural interface, a camera, input devices yet to be developed, and / or the like.

[0059] The video stream (and optionally, the audio stream) received by the client 610 is generated and provided by the video server system 620. As further described elsewhere herein, this video stream includes video frames (and the audio stream includes audio frames). The video frames are configured to meaningfully contribute to the image displayed to the user (e.g., the video frames include pixel information in a suitable data structure). As used herein, the term "video frame" is used to refer to a frame that primarily includes information configured to contribute to, e.g., act on, the image shown to the user. Most of the teachings herein regarding "video frames" are also applicable to "audio frames."

[0060] The client 610 is typically configured to receive input from a user. These inputs may include game commands configured to change the state of a video game or otherwise affect gameplay. The game commands may be received using an input device and / or may be generated automatically by computing instructions executing on the client 610. The received game commands are communicated from the client 610 to the video server system 620 and / or the game server 625 over the network 615. For example, in some embodiments, the game commands are communicated to the game server 625 via the video server system 620. In some embodiments, separate copies of the game commands are communicated from the client 610 to the game server 625 and the video server system 620. The communication of the game commands optionally depends on the identity of the command. The game commands are optionally communicated from the client 610A over a different route or communication channel than is used to provide the audio or video stream to the client 610A.

[0061] The game server 625 is optionally operated by a different entity than the video server system 620. For example, the game server 625 may be operated by a publisher of a multiplayer game. In this example, the video server system 620 is optionally viewed as a client by the game server 625 and is optionally configured to appear, from the perspective of the game server 625, to a conventional client running a conventional game engine. Communication between the video server system 620 and the game server 625 is optionally performed over the network 615. Thus, the game server 625 may be a conventional multiplayer game server that transmits game state information to multiple clients, one of which is the video server system 620. The video server system 620 may be configured to communicate with multiple instances of the game server 625 simultaneously. For example, the video server system 620 may be configured to provide multiple different video games to different users. Each of these different video games may be supported by a different game server 625 and / or published by a different entity. In some embodiments, several geographically distributed instances of video server system 620 are configured to provide gaming videos to multiple different users. Each of these instances of video server system 620 may communicate with the same instance of game server 625. Communication between video server system 620 and one or more game servers 625 optionally occurs over a dedicated communication channel. For example, video server system 620 may be connected to game server 625 over a high-bandwidth channel dedicated to communication between these two systems.

[0062] The video server system 620 comprises at least a video source 630, an I / O device 645, a processor 650, and non-transitory storage 655. The video server system 620 may include one computing device or may be distributed among multiple computing devices, optionally connected via a communication system such as a local area network.

[0063] The video source 630 is configured to provide a series of video frames that form a video stream, e.g., a streaming video or movie. In some implementations, the video source 630 includes a video game engine and rendering logic. The video game engine is configured to receive game commands from a player and maintain a copy of the video game state based on the received commands. This game state includes the positions of objects within the game environment and, typically, a viewpoint. The game state may also include object properties, images, colors, and / or textures. The game state is typically maintained based on game rules and movement, turning, attacking, focusing, interaction, use, and / or similar game commands. Portions of the game engine are optionally located within a game server 625. The game server 625 may maintain a copy of the game state based on game commands received from multiple players using geographically distributed clients. In these cases, the game state is provided by the game server 625 to the video source 630, where the copy of the game state is stored and rendered. The game server 625 may receive game commands directly from the client 610 via the network 615 and / or via the video server system 620.

[0064] The video source 630 typically includes rendering logic, e.g., hardware, firmware, and / or software stored on a computer-readable medium such as storage 655. This rendering logic is configured to create video frames of a video stream based on a game state. All or part of the rendering logic is optionally located within a graphics processing unit (GPU). The rendering logic typically includes processing stages configured to determine three-dimensional spatial relationships between objects based on the game state and viewpoint, and / or apply appropriate textures, etc. The rendering logic produces raw video, which is then typically encoded before communication to the client 610. For example, the raw video may be encoded according to Adobe Flash® standards, .wav, H.264, H.263, On2, VP6, VC-1, WMA, HuffyUV, Lagarith, MPG-x, Xvid, FFmpeg, x264, VP6-8, realvideo, mp3, or the like. The encoding process produces a video stream that is optionally packaged for delivery to a decoder on a remote device. The video stream is characterized by a frame size and a frame rate. Typical frame sizes include 800x600, 1280x620 (e.g., 620p), and 1024x768, although any other frame size may be used. The frame rate is the number of video frames per second. A video stream may contain various types of video frames. For example, the H.264 standard includes "P" frames and "I" frames. I-frames contain information that updates all macroblocks / pixels on a display device, while P-frames contain information that updates a subset of them. P-frames typically have a smaller data size than I-frames. As used herein, the term "frame size" is meant to refer to the number of pixels in a frame. The term "frame data size" is used to refer to the number of bytes required to store a frame.

[0065] In an alternative embodiment, video source 630 includes a video recording device, such as a camera. The camera may be used to generate delayed or live video that may be included in the video stream of a computer game. The resulting video stream optionally includes both rendered images and images recorded using a still or video camera. Video source 630 may also include a storage device configured to store pre-recorded video that is included in the video stream. Video source 630 may also include a motion or position sensing device configured to detect the movement or position of an object, such as a person, and logic configured to determine a game state or to produce video based on the detected movement and / or position.

[0066] Video source 630 is optionally configured to provide overlays configured to be placed over other video. For example, these overlays may include a command interface, login instructions, messages for the game player, images of other game players, and video feeds (e.g., webcam video) of other game players. In embodiments in which client 610A includes a touchscreen interface or an eye-gaze detection interface, the overlays may include a virtual keyboard, joystick, touchpad, and / or the like. In one example of an overlay, a player's voice is overlaid on an audio stream. Video source 630 optionally further includes one or more audio sources.

[0067] In embodiments in which video server system 620 is configured to maintain game state based on input from more than one player, the viewpoint, including the position and direction of the view, may be different for each player. Video source 630 is optionally configured to provide a separate video stream to each player based on the player's viewpoint. Furthermore, video source 630 may be configured to provide different frame sizes, frame data sizes, and / or encoding to each of clients 610. Video source 630 is optionally configured to provide 3D video.

[0068] The I / O devices 645 are configured to allow the video server system 620 to send and / or receive information such as video, commands, requests for information, game state, gaze information, device movement, device location, user movement, client identification information, player identification information, game commands, security information, audio, and / or the like. The I / O devices 645 typically include communications hardware such as a network card or modem. The I / O devices 645 are configured to communicate with the game server 625, the network 615, and / or the clients 610.

[0069] Processor 650 is configured to execute logic, e.g., software, included within various components of video server system 620 described herein. For example, processor 650 may be programmed with software instructions to perform the functions of video source 630, game server 625, and / or client qualifier 660. Video server system 620 optionally includes more than one instance of processor 650. Processor 650 may also be programmed with software instructions to execute commands received by video server system 620 or to coordinate the operation of various elements of game system 600 described herein. Processor 650 may include one or more hardware devices. Processor 650 is an electronic processor.

[0070] Storage 655 includes non-transitory analog and / or digital storage devices. For example, storage 655 may include an analog storage device configured to store video frames. Storage 655 may include computer-readable digital storage, such as a hard drive, optical drive, or solid-state storage. Storage 655 is configured (e.g., via a suitable data structure or file system) to store video frames, artificial frames, video streams including both video frames and artificial frames, audio frames, audio streams, and / or the like. Storage 655 is optionally distributed among multiple devices. In some embodiments, storage 655 is configured to store software components of video source 630, described elsewhere herein. These components may be stored in a format that allows them to be readily available as needed.

[0071] Video server system 620 optionally further includes a client qualifier 660. Client qualifier 660 is configured to remotely determine the capabilities of a client, such as client 610A or 610B. These capabilities may include both the capabilities of client 610A itself and the capabilities of one or more communication channels between client 610A and video server system 620. For example, client qualifier 660 may be configured to test the communication channel via network 615.

[0072] The Client Qualifier 660 may manually or automatically determine (e.g., discover) the capabilities of the Client 610A. Manual determination includes communicating with a user of the Client 610A and prompting the user to provide the capabilities. For example, in some embodiments, the Client Qualifier 660 is configured to display images, text, and / or the like within the browser of the Client 610A. In one embodiment, the Client 610A is an HMD that includes a browser. In another embodiment, the Client 610A is a game console with a browser that may be displayed on an HMD. The displayed objects request the user to enter information such as the Client 610A's operating system, processor, video decoder type, type of network connection, display resolution, etc. The information entered by the user is transmitted back to the Client Qualifier 660.

[0073] The automated determination may be made, for example, by running an agent on the client 610A and / or by sending a test video to the client 610A. The agent may include computing instructions, such as JavaScript, embedded in a web page or installed as an add-on. The agent is optionally provided by the client qualifier 660. In various embodiments, the agent may find out the processing capabilities of the client 610A, the decoding and display capabilities of the client 610A, the latency reliability and bandwidth of the communication channel between the client 610A and the video server system 620, the display type of the client 610A, any firewalls present on the client 610A, the hardware of the client 610A, software running on the client 610A, registry entries within the client 610A, and / or the like.

[0074] Client Qualifier 660 includes hardware, firmware, and / or software stored on a computer-readable medium. Client Qualifier 660 is optionally located on a computing device separate from one or more other elements of video server system 620. For example, in some embodiments, Client Qualifier 660 is configured to determine characteristics of a communication channel between client 610 and one or more instances of video server system 620. In these embodiments, the information discovered by the Client Qualifier can be used to determine an instance of video server system 620 that is best suited to deliver streaming video to one of clients 610.

[0075] It should be understood that the various embodiments defined herein may be combined or assembled into specific implementations that use various features disclosed herein. Thus, the examples provided are not limited to the various implementations possible by combining various elements to define even more implementations, but are merely some possible examples. In some instances, some implementations may include fewer elements without departing from the spirit of the disclosed implementations or equivalent implementations.

[0076] Embodiments of the present disclosure may be practiced with a variety of computer system configurations including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire-based or wireless network.

[0077] In some embodiments, communication may be facilitated using wireless technology. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. 5G networks are digital cellular networks in which a provider's coverage area is divided into small geographic areas known as cells. Analog signals representing sound and images are digitized by the telephone, converted by an analog-to-digital converter, and transmitted as a bit stream. All 5G wireless devices within a cell communicate over electromagnetic waves with a local antenna array and low-power automatic transceiver (transmitter and receiver) within the cell via frequency channels assigned by the transceiver from a frequency pool reused by other cells. The local antennas are connected to the telephone network and the Internet by high-bandwidth optical fiber or wireless backhaul connections. As with other cellular networks, mobile devices moving from one cell to another are automatically transferred to the new cell. Of course, 5G networks are merely example types of communication networks, and embodiments of the present disclosure may use previous generations of wireless or wired communication, as well as later generations of wired or wireless technologies that come after 5G.

[0078] In view of the above embodiments, it will be appreciated that the present disclosure may employ various computer-implemented operations involving data stored in computer systems. These operations are operations requiring physical manipulation of physical quantities. Any of the operations described herein that form part of the present disclosure are useful machine operations. The present disclosure also relates to devices or apparatus for performing these operations. An apparatus may be specially constructed for the required purposes, or the apparatus may be a general-purpose computer selectively enabled or configured by a computer program stored in the computer. In particular, various general-purpose machines may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus to perform the required operations.

[0079] Although the method operations have been described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be arranged to occur at slightly different times, or operations may be distributed within the system to allow processing operations to occur at various processing-related intervals, so long as the processing of telemetry and game state data to generate a modified game state is performed in the desired manner.

[0080] One or more embodiments may also be fabricated as computer-readable code on a computer-readable medium. The computer-readable medium is any data storage device that can store data which can thereafter be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium may include tangible computer-readable media distributed over network-coupled computer systems so that the computer-readable code is stored and executed in a distributed fashion.

[0081] Although the foregoing embodiments have been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims. Therefore, the present embodiments are to be considered as illustrative and not restrictive, and the embodiments are not to be limited to the details set forth herein, but may be modified within the scope of the appended claims and their equivalents.

[0082] It should be understood that the various embodiments defined herein may be combined or assembled into specific implementations that use various features disclosed herein. Thus, the examples provided are not limited to the various implementations possible by combining various elements to define even more implementations, but are merely some possible examples. In some instances, some implementations may include fewer elements without departing from the spirit of the disclosed implementations or equivalent implementations.

Claims

1. 1. A method for interpreting a gesture provided by a user, comprising: capturing images of the user's gestures interacting with a metaverse, the images being captured using multiple sensors distributed within a physical environment in which the user is operating; analyzing the image of the gesture to identify attributes of the gesture captured in the image, the attributes being evaluated to identify at least a magnitude of the gesture; Transforming the attributes of the gesture, including the magnitude of the gesture, based on a context of the user's interaction in the metaverse to define an input from the user; communicating the input to the metaverse to apply the input to the interaction, the input being interpreted in the context of the user's interaction with the metaverse; The method comprising:

2. the interaction is with an interactive application accessible within the metaverse, the gesture is a finger gesture detected on an input device, the input from the user is communicated to the interactive application, and the input is interpreted to define an amount of adjustment to be performed in the interactive application to define a current state; The method of claim 1 , wherein the interactive application is a video game application, or a sign language communication application, or a communication application used to communicate with other users.

3. the interaction is a conversation between the user and a second user in the metaverse, and the input converted from the gesture of the user is communicated to a second avatar representing the second user via an avatar representing the user in the metaverse; 10. The method of claim 1, wherein the input is used to communicate additional information related to particular content that is the target of the conversation, to communicate the importance level of a particular action to be performed by the user, or to communicate instructions to the second user to perform a particular activity and the urgency level of performing the particular activity within the metaverse.

4. 10. The method of claim 1, wherein applying the input to the interaction comprises adjusting gestures and facial expressions of an avatar representing the user in the metaverse according to the attributes of the gesture communicated in the input.

5. the metaverse includes avatars of other users who interact within the virtual space of the metaverse; 5. The method of claim 4, wherein the inputs communicated to the user's avatar are auto-scaled according to the context of the interaction the user is engaged in in the metaverse so as to blend the gestures and facial expressions of the user's avatar with the gestures and facial expressions of other avatars of other users present in the metaverse when the inputs are applied to the user's avatar.

6. the metaverse includes avatars of other users who interact within the virtual space of the metaverse; 5. The method of claim 4, wherein the input communicated to the user's avatar is auto-scaled based on identities of other users with whom the user is interacting within the virtual space of the metaverse.

7. 5. The method of claim 4, wherein the input communicated to the user's avatar is auto-scaled as the user moves from a first scene to a second scene within the metaverse, the auto-scaling being based on the context of content of the second scene relative to content of the first scene.

8. the gesture is identified from a movement of a body part of the user; Evaluating the gesture includes: tracking movement of the body part of the user to provide a current gesture during interaction in the metaverse; evaluating the movement of the body part to define attributes of the current gesture, the attributes being used to define a magnitude of the gesture based on a degree of change detected from the movement; The method of claim 1 , comprising:

9. Analyzing the image of a gesture includes: capturing images of facial features of the user as the user is providing a current gesture, the images of the facial features being captured using the plurality of sensors distributed within the physical environment, and changes in the facial features being evaluated to determine an expression shown by the user; verifying that the facial expression of the user correlates with the current gesture of the user; The method of claim 8 further comprising:

10. Transforming the attribute includes:

10. The method of claim 1, comprising interpreting the user's input in the context of content with which the user is interacting in the metaverse, wherein the interpreted input is applied to the content to adjust a current state of the content in the metaverse.

11. Interpreting the input comprises:

11. The method of claim 10, comprising attenuating or amplifying the input of the user applied in the metaverse based on a magnitude of the gesture, the level of attenuation or amplification corresponding to a magnitude and facial features of the gesture captured by the plurality of sensors distributed in the physical environment.

12. Interpreting the input comprises: capturing images of the user's facial features as the user provides a current gesture, wherein the images of the facial features captured using the plurality of sensors distributed within the physical environment are analyzed to identify changes, and the changes are evaluated to determine an expression exhibited by the user; interpreting the gesture magnitude of the current gesture and the facial expression of the user captured via image based on the context of the user's interaction in the metaverse and previous gestures of the user, the previous gestures identifying a gesture pattern of the user observed from previous interactions, and the current gesture and the facial expression being interpreted to define an urgency level conveyed in the input; The method of claim 10, comprising:

13. if the gesture magnitude of the current gesture is greater than a predefined threshold, the input is communicated in a first signal indicative of the urgency level of the communication, the urgency level communicated in the first signal being amplified commensurate with the magnitude of the gesture expressed by the user; 13. The method of claim 12, wherein if the gesture magnitude of the current gesture is less than the predefined threshold, the input is communicated in a second signal, the second signal designed to not distract the user during the interaction in the metaverse.

14. Capturing the gesture includes: identifying an interactive space in the physical environment with which the user is interacting, the interactive space being defined to encompass a portion of the space of the physical environment that includes particular body parts of the user used to provide the gesture being tracked, the interactive space being represented by a virtual bounding box using coordinates of the physical environment, the size of the virtual bounding box changing according to movement of one or more of the body parts as the user is providing the gesture; tracking the movement of the one or more of the user's body parts, wherein the movement is tracked by correlating the movement with changes in the coordinates of the virtual bounding box; assessing a degree of movement of the particular one of the body parts by calculating the change in the coordinates of the virtual bounding box, the degree of movement being used to determine attributes of the gesture and a magnitude of the gesture; The method of claim 1 , comprising:

15. The method of claim 14 , wherein the virtual bounding box is defined as a two-dimensional box or a three-dimensional box.

16. Capturing the gesture includes: Detecting a finger of the user on a controller used to provide the input to the metaverse, the finger being detected using one or more sensors of the plurality of sensors disposed on the controller; tracking the movement of the fingers on the controller as the user holds the controller and provides the input using controls disposed on the controller; interpreting the movements of the fingers to identify grip forces and input intensities using controls of the controller, wherein the grip forces and input intensities on each of the controls are used to determine gesture attributes including gesture magnitude, and the grip forces and input intensities are used to determine a mood of the user as the user interacts with the metaverse; The method of claim 1 , comprising:

17. 1. A method for interpreting a gesture provided by a user, comprising: tracking the movements of the user's body parts, said tracking being performed by capturing images of the body parts using a plurality of sensors distributed within a physical environment in which the user is operating; analyzing the image to identify attributes of a captured gesture, the attributes of the gesture being evaluated to identify at least a magnitude of the gesture; Transforming the attributes of the gesture, including the magnitude of the gesture, based on a context of the user's interaction in a metaverse to define an input from the user; evaluating a magnitude of the gesture in the context of the user's interaction with the metaverse to define a communication signal for conveying importance of the input; and communicating the input and the communication signal for application to the metaverse, wherein the input is interpreted in the metaverse according to the context of the user's interaction with the metaverse and applied according to the importance conveyed by the communication signal; The method comprising:

18. the interaction of the user in the metaverse is with a second user; Interpreting the input in the metaverse when the gesture of the user is exaggerated and the gesture of the second user is suppressed includes:

18. The method of claim 17, further comprising scaling down the input so that an exaggeration level of the gesture of the user represented in the input matches the exaggeration level of the gesture included in a second input of the second user, wherein the scaled down input of the user can be interpreted at the same level as the second input of the second user.

19. Comparing the magnitude of the gestures examining previous gestures of the user collected over time to identify the attributes and magnitudes of the gestures involved for different contexts of communication, the gestures being examined to understand the level of exaggeration typically expressed by the user for the different contexts; comparing attributes of the current gesture with corresponding attributes of the previous gesture determined for the interaction context to determine whether the attributes and gesture magnitude of the current gesture match the corresponding attributes and gesture magnitude recorded for the previous gesture for the interaction context; If the current gesture and the previous gesture match, interpreting the attributes of the current gesture according to a context of the interaction, the interpretation determining a level of urgency to be conveyed with the input for application to the user's interaction in the metaverse; 18. The method of claim 17, further comprising: if not matching, dynamically scaling the current gesture included in the input to match the previous gesture, the scaling being scaled up or down based on a difference in exaggeration levels expressed in the current gesture and the previous gesture.

20. 18. The method of claim 17, wherein transforming the attributes of the gesture including the magnitude of the gesture further comprises automatically scaling the gesture and the magnitude of the gesture based on a target identified in the user's interaction with the metaverse based on the context of the interaction with the metaverse and according to an environment in which the user is interacting within the metaverse.