Extending game controller functionality with virtual buttons using hand tracking.
A multimodal finger tracking system using IMU, radio signals, and image data improves input detection accuracy on handheld controllers by generating an ensemble model to verify finger gestures, addressing errors in single-mode tracking.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SONY INTERACTIVE ENTERTAINMENT LLC
- Filing Date
- 2025-09-16
- Publication Date
- 2026-04-10
AI Technical Summary
Existing methods for detecting input on handheld controllers using single-mode data, such as image tracking, are prone to errors and inaccuracies, particularly in interactive applications like video games.
A multimodal finger tracking system using multiple sensors and components, including IMU, radio signals, sound, and image data, generates an ensemble model to accurately detect and verify finger gestures, reducing errors by incorporating data from multiple sources.
The multimodal approach enhances the accuracy of detecting and interpreting finger gestures, ensuring correct input recognition for interactive applications by validating inputs through multiple data modalities.
Smart Images

Figure 2026062624000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to identifying input provided by finger gestures on a handheld controller, and more specifically, to using multimodal data collected from a plurality of sensors and components associated with the handheld controller to verify the input provided by finger gestures.
Background Art
[0002] As the number of interactive applications and video games available to users on various devices increases, accurate detection of input provided via various devices has become particularly important. For example, in order to accurately affect the game state of a video game, the video game input provided by a user using a handheld controller must be appropriately identified and correctly interpreted. Relying solely on single-mode data (e.g., image tracking of finger gestures) can result in incorrect results in video games.
[0003] Embodiments of the present disclosure have been made under such a background.
Summary of the Invention
[0004] Embodiments of the present disclosure relate to systems and methods for providing multimodal finger tracking to detect and verify finger gestures provided on an input device, such as a handheld controller. Multimodal finger tracking and verification ensure that finger gestures are appropriately identified and correctly interpreted, thereby reducing errors resulting from relying solely on single-mode tracking. A custom finger tracking model (e.g., an ensemble model) is generated and trained using multiple modalities of data captured by a plurality of sensors and components associated with the handheld controller (hereinafter simply referred to as the "controller"), thereby increasing the accuracy of detecting and interpreting finger gestures.
[0005] Traditional methods for detecting input relied on single-data-source models. For example, traditional methods relied on a common camera (i.e., a single data source) to detect and track a user's finger on a controller. The accuracy of tracking using a single source is unreliable and prone to errors, making it undesirable for interactive applications. To overcome the shortcomings of traditional methods, multimodal data is collected from multiple sensors and components associated with a controller, which are used to provide input and to validate the finger gestures detected by the controller. The collected multimodal data is used to generate and train a multimodal data model, which is then used to correctly interpret finger gestures. Because data from multiple modes is used to generate and train the model, the multimodal data model is also referred to herein as an “ensemble model.” The ensemble model is continuously trained using additional multimodal data collected over time, according to training rules defined for various finger gestures. The output is selected from the ensemble model and used to confirm / validate the finger gestures detected by the controller. Finger gestures can correspond to pressing physical buttons, pressing virtual buttons defined on the controller, or inputs provided on a touchscreen interface located on the controller, with outputs identified to correspond to the correct interpretation of the finger gesture. Virtual buttons can be identified on any surface of the controller where no physical buttons are located, and finger gestures can be defined on virtual buttons, such as single taps, double taps, presses, or swipes in a specific direction.
[0006] This model incorporates multimodal finger tracking techniques by considering several model components when generating and training an ensemble model, such as finger tracking using image feeds from an image capture device, IMU data from an inertial measurement unit (IMU) sensor located within the controller, radio signals from wireless devices located in the user's environment, and data from various sensors such as distance / proximity sensors and pressure sensors. The ensemble model helps accurately detect finger gestures provided to the controller by tracking and validating finger gestures using data from more than one mode.
[0007] In one embodiment, a method for validating inputs provided to a controller is disclosed. The method includes detecting finger gestures provided by a user on the surface of the controller. The finger gestures are used to define the inputs of an interactive application selected for user interaction. Multimodal data is collected by tracking finger gestures on the controller using multiple sensors and components associated with the controller. An ensemble model is generated using the multimodal data received from the multiple sensors and components. The ensemble model is continuously trained using additional multimodal data collected over time to produce different outputs, and the training follows training rules defined for different finger gestures. The ensemble model is generated and trained using machine learning algorithms to define various outputs. The outputs are identified from the ensemble model of finger gestures. The outputs identified from the ensemble model are interpreted to define the inputs of an interactive application.
[0008] In other embodiments, a method for defining inputs for an interactive application is disclosed. The method includes receiving a finger gesture provided by a user on the surface of a controller. The finger gesture is used to define the input for an interactive application selected for user interaction. Multimodal data capturing attributes of the finger gesture on the controller is received from multiple sensors and components associated with the controller. Weights are assigned to the modal data corresponding to each mode contained in the multimodal data captured by the multiple sensors and components. The weight assigned to each mode indicates that the modal data for each mode is used to accurately predict the finger gesture. The finger gesture and multimodal data are processed based on the weights assigned to each mode to identify the input for the interactive application corresponding to the finger gesture detected by the controller.
[0009] Other aspects and advantages of this disclosure will become apparent from the following detailed description, which illustrates the principles of this disclosure, in conjunction with the accompanying drawings.
[0010] This disclosure will be best understood by referring to the following description, which will be interpreted in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0011] [Figure 1A] An abstract pipeline used to construct an ensemble model for detecting user-provided gestures on a controller, according to one embodiment of this disclosure, is shown. [Figure 1B] A simplified block diagram of various components used to define an abstract pipeline for building an ensemble model to detect input provided via finger gestures in a controller, according to one embodiment of the present disclosure, is shown. [Figure 2]A and B illustrate an example, according to one embodiment of the present disclosure, in which the controller determines whether a real button press or a virtual button press is occurring using modal data collected from an inertial measurement unit sensor associated with the controller in response to a finger gesture detected by the controller. [Figure 3A] The present disclosure shows exemplary radio and reflected signals captured using a wireless communication device to detect the movement of different fingers of a user holding a controller, according to one embodiment of the present disclosure. [Figure 3B] The present disclosure shows several exemplary wireless signal graphs that capture amplitude fluctuations caused by tracking the movement of some users, according to several embodiments of this disclosure. [Figure 4] An exemplary microphone array used to capture the sound of activating a control on a controller (e.g., a button press) in order to predict when a virtual button is pressed is shown according to one embodiment of the present disclosure. [Figure 5] Figures A and B show various diagrams of an image capture device coupled to a wireless controller to capture finger gestures provided to the controller, according to one embodiment of the present disclosure. [Figure 6] This invention provides a simplified flow of operation of a method for validating inputs provided to a controller of an interactive application, according to one embodiment of this disclosure. [Figure 7] This document illustrates components of an exemplary computing device that can be used to perform various embodiments of the present disclosure. [Modes for carrying out the invention]
[0012] In the embodiments for carrying out the following inventions, several specific details are provided to give a complete understanding of the Disclosure. However, it will be apparent to those skilled in the art that the Disclosure can be practiced without some or all of these specific details. In other embodiments, well-known process steps are not described in detail in order to avoid obscuring the Disclosure.
[0013] As the number of interactive applications increases, the ability to accurately identify inputs provided using various devices becomes particularly important. Specifically, in interactive applications such as high-intensity video game applications, inputs provided using controllers must be properly detected and accurately interpreted so that the game state of the video game can be updated correctly and in a timely manner. For this purpose, data from multiple modalities to detect inputs provided by the user on an input device, such as a handheld controller, and to track finger gestures on the controller in response, is collected from multiple sensors and components associated with the controller. A custom finger tracking model is generated and trained using this multimodal data, and then this model is used to identify outputs corresponding to finger gestures detected on the controller. The trained custom finger tracking model improves the accuracy of predicting finger gestures better than a typical camera-only finger tracking model because its finger gesture prediction relies on multiple modal data sources to validate the finger gestures.
[0014] Modal data, and some of the sensors and components associated with the controller that captured the modal data, include (a) IMU data captured using inertial measurement unit (IMU) sensors, such as magnetometers, gyroscopes, and accelerometers; (b) radio communication signals, including forward and reflected signals, captured using radio communication devices such as Bluetooth®-enabled devices and Wi-Fi routers placed in the environment; (c) sound data from microphone arrays; (d) sensor data captured using distance and / or proximity sensors; and (e) image data captured using image capture devices(s). The modal data and sensors and components described above are presented as examples only and should not be considered exhaustive or limiting. Various sensors and components capture attributes of finger gestures in various modal forms used to generate an ensemble model. When training the ensemble model to improve the accuracy of identifying finger gestures detected by the controller, defined training rules are applied to various finger gestures. In some embodiments, a multimodal data acquisition engine running on a server collects various modal data captured by multiple sensors and components at the controller, and generates and trains an ensemble model. In other embodiments, the multimodal data acquisition engine can run on the controller itself, or on a processor co-located with and coupled to the controller to reduce latency.
[0015] Each modal data captures several attributes of the finger gesture provided by the controller, and by using these attributes to validate the finger gesture, the correct input corresponding to the finger gesture can be identified to influence the outcome of the interactive application selected by the user for the interaction. By using multiple modalities of data captured by tracking finger gestures on the controller, additional validation can be provided to correctly determine the finger gesture, thus identifying the correct input for the interactive application.
[0016] Figure 1A shows an example of an abstract pipeline followed to build an ensemble model for detecting finger gestures, according to several embodiments. The abstract pipeline uses a multimodal finger tracking approach, where finger gestures are tracked using multiple sensors and components associated with the controller, and data generated from such tracking is used to identify and / or validate finger gestures on the controller. Each of the multiple sensors and components captures modal data in a specific mode. For example, multiple modal data 120-1 captured by sensors and components from finger tracking include, to name a few, camera feeds 120a1, inertial measurement unit sensor (IMU) data 120b1, WiFi data (including WiFi forward and reflected signals) 120c1, sound data 120d1, sensor data from distance / proximity sensors 120e1, and sensor data from pressure sensors 120f1. Naturally, the aforementioned list of data for capturing finger gestures is provided as an example and should not be considered exhaustive or limited, as other forms of modal data captured from different sensors and components can also be considered when identifying and / or validating finger gestures. Modal data from each mode is analyzed to detect finger gestures. Considering modal data from a single mode may result in errors or reduced reliability in predicting finger gestures. Therefore, modal data from multiple modes is considered in the analysis to improve the accuracy of finger gesture detection. In some embodiments, modal data associated with different modes are voted on using a voting module, and weights are assigned to the modal data for each mode. The weights assigned to the modal data for each mode can be equal or unequal, and the decision to assign equal or unequal weights (144) to the modal data for each mode is, in some embodiments, based on the reliability of each mode in correctly predicting finger gestures.Next, finger gestures are correctly predicted by using weights assigned to data from different modes in the analysis. Correct prediction of finger gestures allows for the identification of inputs associated with those gestures, which are then used by the user as interactive inputs to the interactive application selected by the user for the interaction. Details of the analysis are explained with reference to Figure 1B.
[0017] Figure 1B shows the various components of system 100 for correctly detecting finger gestures in the controller so that input to an interactive application can be properly identified. The various components and sensors used to build the ensemble model represent an abstract pipeline that can be used to correctly detect finger gestures. The ensemble model is generated and trained using modal data from several modal components and sensors that capture different attributes of finger gestures provided by the controller. As described above, the different modalities of the data captured by tracking finger gestures include image feeds, IMU data, WiFi signals (radio signals including reflected signals), distance / proximity sensors, and data acquired from pressure sensors. Attribute information about the finger gestures captured by each of these components and sensors is transferred to a server computing device for further processing. The server computing device processes the multimodal data to properly identify the finger gestures in the controller.
[0018] For this purpose, the system 100 for determining the input of an interactive application includes a controller 110, such as a handheld controller used by the user to provide finger gestures; a plurality of sensors and components 120 associated with the controller 110 to capture various attributes of the finger gestures; and a server device 101 used to process multimodal data capturing the finger gestures and various attributes of the finger gestures, and to validate the finger gestures. The finger gestures are provided by the user as input to an interactive application selected for user interaction. The server computing device (or hereafter simply referred to as the “server”) 101 uses a modal data acquisition engine 130 for collecting various modalities of data transferred by the plurality of sensors and components 120, and a modal data processing engine 140 for processing the multimodal data to identify the finger gestures and define the input to the interactive application. The modal data acquisition engine 130 and the modal data processing engine 140 may be part of a multimodal processing engine running on the server 101.
[0019] In some embodiments, the server 101 may be a game console or any other computing device co-located within the user-operated environment. The game console or computing device may then be coupled to other game consoles via a network as part of a video game multiplayer setup. In some embodiments, the controller 110 is a networked device and is directly coupled to the remote server 101 via a network (not shown), such as the Internet. In the case of a networked device, the controller 110 is coupled to the network via a router embedded within or outside the controller 110. In other embodiments, the controller 110 is coupled to the remote server 101 via the Internet through a game console or another client device (not shown), and the game console or client device is co-located with the controller 110. Controller 110 is paired with Server 101 as part of the initial setup, or when it detects the presence of Controller 110 near a game console, another client device, or a router (if Server 101 is connected to Controller 110 via a game console, router, or other computing device co-located with Controller 110), or when it detects activation of Controller 110 by a user (i.e., if Server 101 is located remotely from Controller 110). Details of the finger gestures provided on the surface of Controller 110 are transmitted to the game console / Server 101 for processing.
[0020] In response to detecting a finger gesture provided by a user on the surface of the controller 110, various sensors and components 120 associated with the controller 110 automatically become active to capture different attributes of the user's finger gesture. Some of the sensors and components 120 that automatically become active to collect various attributes of the finger gesture include image capture device(s) 120a, IMU 120b, WiFi device(s) 120c, microphone array 120d, distance / proximity sensor 120e, and pressure sensor 120f. In addition to the aforementioned sensors and components, other sensors and / or components may also be used to collect attributes of the finger gesture at the controller.
[0021] In some embodiments, the image capture device 120a is a camera incorporated within a mobile computing device such as a mobile phone or a tablet computing device. Alternatively, the camera can be a webcam, or a console camera, or a camera that is part of an HMD. The camera, or the device in which the camera is incorporated, is paired with the game console / server 101 using a pairing engine 125a. This pairing enables the image capture device (i.e., the camera) to receive an activation signal from the game console / server 101 to capture an image of the user's finger gesture on the controller 110. When the game console / server 101 detects a finger gesture on the surface of the controller 110, it generates an activation signal to the image capture device. A mobile computing device with a built-in camera is supported on a holding structure located on the controller 110, allowing the camera within the mobile computing device to capture close-up views of various features of the finger gesture provided on the controller 110. More information about the holding structure is described with reference to Figures 5A and 5B. When activated, the image capture device captures images of various attributes of the finger gesture, including the finger used to provide the gesture, the position of the finger relative to the input controls on the controller 110 (e.g., buttons, touchscreen surfaces, other interactive surfaces, etc.), the movement of the finger on the controller, and the type of finger gesture provided (e.g., single tap, double tap, slide gesture, etc.). The images capturing the attributes of the finger gesture are transferred as an image data camera feed to the modal data acquisition engine 130 running on the game console / server 101.
[0022] In response to the activation of various sensors and components, an inertial measurement unit (IMU) sensor 120b integrated within the controller 110 is used to capture IMU signals related to finger gestures. In some embodiments, while the user holds the controller 110 in their hand, the IMU signals captured by the IMU sensor are used to distinguish between different finger gestures detected by the controller 110. For example, an IMU signal that captures a subtle tapping at a position defined within the back of the controller 110 can be interpreted as meaning a first input (i.e., virtual button 1), and an IMU signal that captures a subtle tapping at a position defined within the front of the controller 110 that includes no buttons or interactive interfaces can be interpreted as meaning a second input (i.e., virtual button 2), an IMU signal that captures a tapping at the upper right corner of the back of the controller 110 can be interpreted as meaning a third input (i.e., virtual button 3), tapping the upper left corner of the back of the controller 110 can be interpreted as meaning a fourth input (i.e., virtual button 4), an IMU signal that captures a tapping at the back of the controller using the middle finger can be interpreted as meaning a fifth input (i.e., virtual button 5), tapping on a physical button on the front of the controller 110 can be interpreted as meaning a sixth input (e.g., pressing the physical button), and so on.
[0023] In some embodiments, the functionality of the controller 110 can be extended by using virtual buttons defined by tracking finger gestures. The extended functionality enables the user to interact with more than one application simultaneously, and such interactions can be performed without the need to interrupt one application for another. Virtual buttons can be defined by identifying the position of the fingers when the user is holding the controller 110 and by finger gestures provided by the user with respect to the identified finger positions. In some cases, when the user is playing a game, for example when running on a game controller or a game server, the user may also be listening to music provided by a second application (e.g., a music application). Normally, when a user needs to interact with a music application, they must pause the game they are currently playing, access the menu to interact with the music application, and use one of the buttons or interactive surfaces on the controller 110 to advance to the next song on their playlist. To avoid the user having to pause gameplay and to provide an alternative way to interact with a music application while playing a game, virtual buttons can be defined to extend the functionality of the controller 110. Because virtual buttons can be defined and associated with pre-assigned commands, users can use the virtual buttons to interact with a music application without having to interrupt their current gameplay.
[0024] In another embodiment, finger gesture tracking can be used while the user is holding the controller, enabling users with certain disabilities to communicate in online games. For example, finger gesture tracking may be used to detect different finger positions and gestures while a user with a disability is holding the controller 110. These finger positions and gestures can be interpreted using a machine learning (ML) algorithm as Morse code input (dots and dashes for taps and swipes), and such interpretation can be performed by the ML algorithm by recognizing the user's disability provided in the user's user profile. Furthermore, the Morse code input can be converted to text characters or provided as game input. The text characters can be used to communicate with other players / spectators / users by converting text to speech, or to provide as text responses on a chat interface. The Morse code input can be interpreted to correlate with game input and used to influence the game state of the game the user is playing. The aforementioned uses of tracking and interpreting finger gestures to identify virtual buttons and / or inputs to interactive applications for users with disabilities are provided as examples and should not be considered exhaustive or limited; other uses may also be conceivable.
[0025] Figures 2A and 2B show some exemplary signal amplitude variations captured in the respective IMU signals for different finger gestures using the IMU sensor 120b. Figure 2A shows the position of the user's finger relative to different input controls (i.e., buttons and touchscreen interactive interfaces) defined on the controller 110 when the user is operating the controller 110. Figure 2B shows some exemplary amplitude variations captured for three different finger gestures detected at different positions on the controller in some embodiments. The amplitude variations shown in Figure 2B represent the amplitude variations along the X, Y, and Z axes captured in the IMU signals for different finger gestures. For example, the amplitude variations shown along the X, Y, and Z axes in box "VB1" are with respect to finger gesture 1, which can be interpreted as meaning that finger gesture 1 provided to the controller 110 by the user corresponds to an input related to virtual button 1. Similarly, the amplitude fluctuations shown along the X, Y, and Z axes in box "VB2" correspond to finger gesture 2, which can be interpreted as meaning that finger gesture 2 provided by the user to controller 110 corresponds to input for virtual button 2, and the amplitude fluctuations shown along the X, Y, and Z axes in box "RB1" correspond to finger gesture 3, which can be interpreted as meaning that finger gesture 3 provided by the user to controller 110 corresponds to input for physical button 1. Naturally, virtual buttons 1, 2, 3, etc., and physical buttons 1, 2, 3, etc., can be defined as relating to different inputs of different interactive applications. The IMU signals capturing minute signals are transferred to the modal data acquisition engine 130 as IMU sensor data.
[0026] Returning to Figure 1B, in response to the activation of various sensors and components, WiFi devices (i.e., wireless devices) 120c distributed in the environment where the user is located begin to capture WiFi signals. The data captured from the WiFi signals is used to detect the user's position, the relative positions of various body parts of the user including hands and fingers, and body movements including finger movements / finger gestures. Using the data from the WiFi signals and the reflected signals of those WiFi signals, different finger movements can be detected while the user is holding the game controller 110.
[0027] Figures 3A and 3B show various WiFi signals captured within the environment (i.e., geographical location) where a user (i.e., user 1) is located, using a WiFi device. The WiFi device includes a transmitter 301, such as a router, laptop, or other computing device including a personal digital assistant (e.g., a voice assistant), and a receiver 302, such as a second laptop or desktop computing device. The aforementioned list of devices representing transmitter 301 and receiver 302 is provided merely as an example and should not be considered exhaustive or limiting. Transmitter 301 continuously transmits WiFi signals, and receiver 302 receives various WiFi signals. Figure 3A shows various WiFi signals transmitted by transmitter 301 and received by receiver 302 within the room where the user is located. Some of the WiFi signals received by receiver 302 include WiFi signal 303 reflected by the room walls, WiFi signal 304 reflected by user 1, WiFi signal 305 representing user 1's line of sight, and WiFi signal 306 reflected by the room floor. Receiver 302 continuously monitors the received WiFi signals to determine the user's movement within the room's geographical location. Initially, channel state information (CSI), representing the channel characteristics of the communication link between transmitter 301 and receiver 302 regarding the geographical location where user 1 is operating, is determined using WiFi signals transmitted by transmitter 301 and received by receiver 302 when no objects or users are present between transmitter 301 and receiver 302. Using the channel characteristics, a baseline of the WiFi signals is established, taking into account the combined effects of scattering, fading, and signal intensity attenuation due to distance. Next, the CSI is determined when user 1 is present in a geographical location (e.g., a room) and when user 1 moves within that room. Variations in the CSI are due to the user's body blocking or reflecting one or more of the aforementioned WiFi signals. Using these variations in each of the one or more WiFi signals collected over a period of time, the user's movement within the geographical location (e.g., a room) is determined, including various finger movements of user 1 while user 1 is holding and operating the game controller 110. Using the channel characteristics of the WiFi signals, snapshots of body parts, including user 1's fingers, can be captured. These snapshots can then be used to reconstruct the body parts to determine which body parts (e.g., fingers) moved. These WiFi signals may include signals provided by the router and Bluetooth® signals provided by the controller.
[0028] Figure 3B shows the variation in the CSI signal amplitude of a Wi-Fi signal received from a Wi-Fi device (e.g., single subcarrier) caused by user movement in several embodiments. The signal amplitude variation is plotted against time. Signal amplitude variation 321 is shown for a Wi-Fi signal transmitted when there is no user or object between transmitter 301 and receiver 302 within the geographical location. In some embodiments, signal amplitude variation 321 establishes the baseline CSI signal amplitude variation. Signal amplitude variations 322-326 capture the variation caused by user movement. For example, signal amplitude variation 322 shows an example of WiFi signal variation when detecting that user 1 is sitting in a room (i.e., geographical location), signal amplitude variation 323 captures WiFi signal variation when detecting the opening or closing of the room door, signal amplitude variation 324 captures WiFi signal variation when detecting that user 1 is typing on an input device such as a keyboard, signal amplitude variation 325 captures WiFi signal variation when detecting that user 1 is waving their hand, and signal amplitude variation 326 captures WiFi signal variation when detecting that user 1 is walking around the room. Thus, it is possible to use the WiFi signal to detect different finger movements when the user is operating the controller 110. The WiFi signals capturing various user movements are transferred as WiFi signals to the modal data acquisition engine 130.
[0029] Returning to Figure 1B, in response to finger gesture detection on the controller 110, a microphone array 120d embedded in or attached to the controller 110 becomes active. The finger gesture may be a press of a physical button on the controller 110 or a press of a virtual button, and the activated microphone array 120 is configured to capture the sound attributes of the button press. Specifically, the microphones in the microphone array 120d work together to determine the direction from which the sound is coming and to pinpoint its location. The attributes of the finger gesture (i.e., direction and location) are used to determine whether the finger gesture corresponds to a press of a physical button or a press of a virtual button. In some embodiments, different locations on the controller other than where the physical buttons and touchscreen interface are located can represent different virtual buttons. For example, the left-hand corner on the back of controller 110 can be defined to represent virtual button 1, the right-hand corner on the back of controller 110 can be defined to represent virtual button 2, the central top position on the back of controller 110 can be defined to represent virtual button 3, and so on. The sound attributes caused by finger gestures can be interpreted to determine whether the finger gesture corresponds to a press of a physical button or a virtual button, and whether a physical or virtual button was pressed. The sound attributes captured by the microphone array 120d are transferred to the modal data acquisition engine 130.
[0030] Figure 4 shows an exemplary microphone array 120d embedded in or coupled to the controller 110 to capture sound generated from a user's finger gestures on the surface of the controller 110. The microphone array 120d is shown to include four microphones (401a, 401b, 401c, and 401d). The intensity of the sound captured by each microphone 401 varies based on the distance of the sound from each microphone 401. The sound signals captured by each of the microphones 401a-401d are then transferred to a digital signal processor (DSP) 402 for processing. In some embodiments, the DSP402 is configured to assign different weights to different sounds captured by the microphones. The sounds captured by different microphones in the microphone array 120d, and the relative weights assigned to each sound, are analyzed, for example, using triangulation techniques, to identify the attributes of the various captured sounds, such as direction, position, duration, volume, frequency, etc. These attributes are used together with the relative weights to determine a specific button press or swipe and the direction of the swipe, etc. In other embodiments, instead of assigning a separate weight to each sound, the DSP402 assigns separate weights to different attributes of each sound detected / captured by the microphones in the microphone array 120d, and these weights and detected attributes can be used to associate the sound with a specific button press or finger swipe. Using the weights assigned to different attributes, it is possible to determine which sounds to ignore as ambient noise and which sounds to focus on in order to determine a finger gesture. The sound attributes and analysis details are then transferred as electrical signals to the modal data acquisition engine 130.
[0031] Returning to Figure 1B, in some embodiments, in response to the detection of a finger gesture on the controller 110, one or more distance / proximity sensors 120e are activated to capture attributes of the finger gesture provided by the controller 110. Similar to the microphone array, the distance / proximity sensors 120e can be used to capture attributes of the finger gesture provided to the back of the controller 110. The attributes of the finger gesture provided to the back of the controller 110, captured by the distance / proximity sensors 120e, can be used to independently determine a virtual button press, or they can be used in combination with sound attributes captured by the microphone array 120d to further verify the virtual button press determined from the finger gesture. The additional verification provided by the data captured by the distance / proximity sensors 120e makes the detection of a virtual button press from the finger gesture more accurate. In some embodiments, the distance / proximity sensors 120e can include ultrasonic sensors, infrared sensors, LED time-of-flight sensors, capacitive sensors, and the like. The aforementioned list of distance / proximity sensors 120e is provided as an example and should not be considered exhaustive or limited. The attributes of the finger gesture captured by the distance / proximity sensor 120e are transferred to the modal data acquisition engine 130. In addition to the distance / proximity sensor 120e, the pressure sensor 120f is also activated and can capture the attributes of the pressure applied to different locations on the controller 110 by the finger gesture (e.g., position, pressure amount, pressure application time, finger used to apply pressure, etc.). For example, if the pressure application time is less than the threshold, the pressure attributes captured by the pressure sensor 120f can be ignored. If the pressure application time is greater than the threshold, the attributes of the applied pressure can be transferred as input to the modal data acquisition engine 130. Similarly, if the applied pressure amount is less than the threshold amount, the data captured by the pressure sensor 120f can be ignored. However, if the applied pressure amount is greater than the threshold amount, the data captured by the pressure sensor 120f can be considered as input to the modal data acquisition engine 130. The attributes of the pressure given via a finger gesture that meets or exceeds the threshold / amount are transferred to the modal data acquisition engine 130 as pressure sensor data.
[0032] As described above, in some embodiments, the image capture device may be a camera built into a mobile computing device, such as a mobile phone or tablet computing device. In these embodiments, for example, a mobile phone camera may be preferred over a webcam, console camera, or camera built into an HMD. In other embodiments, a camera built into a mobile phone (i.e., a mobile computing device) can be used to capture images of finger gesture attributes, in addition to a webcam / console camera / HMD camera. In embodiments where images of finger gesture attributes are captured using a mobile computing device camera, the mobile computing device (e.g., a mobile phone) can be coupled to the controller 110.
[0033] Figures 5A and 5B illustrate one such embodiment in which a mobile computing device (e.g., a mobile phone 502) is coupled to a controller 110. In some embodiments, the mobile phone 502 is coupled to the controller 110 using a holding structure 504. Figure 5A shows a front perspective view of the controller 110 in which the holding structure 504 receives and holds the mobile phone 502. Figure 5B shows a rear view of the holding structure 504 coupled to the controller 110 and configured to receive, hold, and operate the mobile phone 502. In some embodiments, the holding structure 504 is a three-dimensional (3D) printed structure that can be attached to the controller 110. In some embodiments, the 3D printed structure includes a motor (not shown) that moves the mobile phone 502 to different positions, allowing the camera of the mobile phone 502 to capture various attributes of finger gestures.
[0034] Referring simultaneously to Figures 1B, 5A, and 5B, in some embodiments, to accommodate different mobile phone models (such as mobile phone size, camera position, and number of cameras), the game console / server 101 first performs an automatic pairing operation to pair the mobile phone 502 with the game console / server 101. A signal is sent from the pairing engine 125a to the mobile phone to initiate the pairing operation. After the mobile phone 502 has successfully paired with the game console / server 101, the mobile phone 502 is mounted in the holding structure by initiating a calibration operation. The calibration engine 125b is used to determine the type and model of the mobile phone 502 and sends a signal to the holding structure to adjust the size of the holding structure to accommodate the mobile phone 502. In response to a signal from the calibration engine 125b, a motor that operates the holding structure 504 adjusts the size of the holding structure 504 so that it can securely receive the mobile phone 502. In addition to automatically calibrating the size of the holding structure to accommodate the mobile phone 502, the calibration engine 125b also calibrates the angle at which the mobile phone needs to be positioned to capture images of different finger positions. This angle is automatically calibrated in response to detecting the presence of the user's fingers and finger gestures provided on the controller 110, and is dynamically determined based on the position of the user's fingers if the user has provided finger gestures to the controller 110. As part of the angle calibration operation, the calibration engine 125b tracks the user's hand and fingers and sends a second signal to the controller 110 and / or the holding structure 504 to adjust the motors to move the mobile phone 502, thereby aligning the camera(s) of the mobile phone 502 to the calibrated angle. In response to the second signal from the calibration engine 125b, the motors of the holding structure 504 are engaged to move and rotate the movable parts of the holding structure 504 to achieve good hand and finger position tracking. Once the mobile phone is moved into place, the camera(s) of the mobile phone 502 become active and capture images of various attributes of the finger gestures. The captured images are streamed to the game console / server 101 as image data camera feeds.
[0035] The modal data acquisition engine 130 collects inputs from multiple sensors and components 120 to generate multimodal data. The multimodal data is processed to identify the modes and the amount of modal data captured for each mode included in the multimodal data collected from the sensors and components. The details of the modes, the amount of modal data for each mode, and the multimodal data captured by the sensors and components 120 are transferred by the modal data acquisition engine 130 to the modal data processing engine 140 for further processing.
[0036] The modal data processing engine 140 analyzes the modal data of each mode included in the multimodal data to identify and / or verify finger gestures in the controller. As previously described with reference to Figure 1A, modal data from multiple modes are considered in the analysis to improve the accuracy of finger gesture detection. As part of the analysis, weights are assigned to the modal data associated with each mode. The weights assigned to the modal data of each mode can be equal or unequal. The decision (144) to assign equal or unequal weights to the modal data of each mode is, in some embodiments, based on the accuracy of correctly predicting finger gestures using the modal data of each mode. In some implementations, multimodal data can be broadly categorized into a first set of modal data captured using multiple sensors and a second set of modal data captured using components. In some embodiments, the weights assigned to mode-specific modal data captured by sensors are greater than the weights assigned to mode-specific modal data captured by components. For example, IMU sensor data captured by an IMU sensor, or distance sensor data captured by a distance / proximity sensor, are assigned a greater weight than image data camera feeds captured by an image capture device. In another example, IMU sensor data and sound data are assigned heavier weights than WiFi signal data. The weight allocation engine 144a is used to identify the mode associated with each modal data included in the multimodal data, and the reliability of the modal data for each mode when predicting finger gestures. Based on the reliability of each mode, the weight allocation engine 144a assigns weights to the modal data. The weights assigned to the data of different modes are used together to correctly predict finger gestures. For example, the modal data processing engine 140 uses the weights assigned to the modal data for each mode included in the multimodal data to generate cumulative weights. The cumulative weights are used to correctly predict finger gestures. The predicted finger gestures are more accurate because the game console / server 101 relies on data from more than one mode to identify and / or validate finger gestures. Once a finger gesture is identified / validated, inputs related to the finger gesture are then identified and used as user input to interactive applications such as video games, influencing the game state.
[0037] In some embodiments, the modal data processing engine 140 uses a machine learning (ML) algorithm 146 to analyze multimodal data captured by multiple sensors and components to identify and / or validate finger gestures provided to the controller 110. The ML algorithm 146 uses the multimodal data to generate and train an ML model 150. By predicting and / or validating finger gestures using the ML model 150, the appropriate input corresponding to the predicted / validated finger gesture can be identified and used to influence the state of an interactive application (e.g., a video game) selected by the user for interaction. The ML algorithm 146 uses a classifier engine (i.e., a classifier) 148 to generate and train the ML model 150. The ML model 150 includes a network of interconnected nodes, where each consecutive node pair is connected by an edge. The classifier 148 is used to introduce various nodes into the network of interconnected nodes of the ML model 150, each node relating to modal data in one or more modes. Interrelationships between nodes are established, the complexity of modal data in different modes is understood, outputs used to identify or validate finger gestures are identified, and inputs corresponding to finger gestures are identified.
[0038] In some embodiments, the classifier 148 is predefined for different modes to understand the complexity of the modal data for each mode when correctly predicting and / or validating finger gestures provided to the controller. When and where finger gestures are provided to the controller 110, the classifier 148 further trains the ML model 150 using modal data captured in real time by sensors and components, and uses the ML model 150 to determine the amount of influence the modal data for each mode has on correctly predicting / validating finger gestures. The ML model 150 can be trained according to training rules 142 to improve the accuracy of finger gesture prediction. The training rules are defined for each finger gesture based on the anatomical form of the finger, how the controller is held, the position of the finger relative to the button, etc. The machine learning (ML) algorithm 146 uses modal data of different modes contained in multimodal data as input to the nodes of the ML model 150, and incrementally updates the nodes using additional multimodal data received over time, adjusting the outputs to meet predefined criteria for different finger gestures. The ML algorithm 146 uses reinforcement learning to strengthen the ML model 150 by using the initial set of multimodal data to build the ML model 150, learning the complexity of each mode and how the modal data of each mode affects the correct prediction / validation of finger gestures, and strengthens the model's training and reinforcement using additional modal data received over time. The adjusted outputs of the ML model 150 are used to correctly predict / validate different finger gestures. Outputs from the adjusted outputs are selected to correspond to finger gestures, and such selection may be based on the cumulative weights of the multimodal data that demonstrate accurate predictions of finger gestures.
[0039] In some embodiments, tracking finger gestures using modal data captured for different modes may be user-specific. For example, each user may handle the controller 110 in a different way. For instance, a first user may hold the controller 110 in a particular way and provide input to the controller in a particular way, or at a particular speed or with a particular pressure. If a second user uses the controller 110 to provide input via finger gestures, their method of holding the controller 110 or using it to provide input may differ from that of the first user. To accommodate the different user ways of handling the controller 110, a reset switch may be provided on the controller 110, allowing the finger gesture tracking to be reset or reprogrammed so that the finger gesture interpretation can be user-specific and user-independent. The reset switch may be defined as a specific button press or a specific sequence of button presses. In another example, resetting or reprogramming the finger gesture tracking may be done on demand by the user, and such a request may be based on a specific context in which the finger gesture needs to be tracked, for example.
[0040] In the various embodiments discussed herein, a modal data acquisition engine 130 and a modal data processing engine 140 with an ML algorithm 146 are described for identifying / validating finger gestures implemented on the server 101. However, to reduce latency, the modal data acquisition engine 130 and the modal data processing engine 140 with the ML algorithm 146 can be implemented locally on a disk in a computing device coupled to the controller 110 (e.g., a game console co-located with the controller) instead of on the remote server 101. In such embodiments, the graphics processing unit (GPU) of the game console can be used to improve the speed at which finger gestures are predicted.
[0041] The various embodiments discussed herein teach a multimodal finger tracking mechanism in which several modal components capture multimodal data. Each modal component provides a voting engine with information regarding the detection of finger gestures. The voting engine then uses the modal data from each mode to assign appropriate weights to the modal data from each mode that indicate an accurate prediction of the finger gesture. For example, modal data generated from a sensor may be given a greater weight because the sensor tends to detect gestures more accurately than a webcam feed. Because the system relies on modal data from more than one mode, the error in detecting finger gestures is reduced by the relative weighting of the modal data based on predictive accuracy. The modal data processing engine 140 uses actual button press state data, video features (i.e., images from an image capture device), audio features using a microphone array associated with the controller 110, sensor data, and WiFi signals to train a custom finger tracking model (i.e., ML model 150) to predict finger gestures with higher accuracy than when predictions rely on only a single data source, such as a common camera-only finger tracking model.
[0042] Figure 6 shows the operation flow of a method for verifying input provided to a controller in several embodiments. The method begins in operation 610, in which case a finger gesture is detected on the surface of a controller operated by a user to interact with an interactive application such as a video game. The finger gesture may be provided on any surface of the controller that includes input controls such as physical buttons and touchscreen interactive interfaces. In response to the detection of the finger gesture, as shown in operation 620, multiple sensors and components are activated to capture various attributes of the finger gesture. Each sensor or component captures modal data in a specific mode, and the modal data captured by multiple sensors and components are collected to define multimodal data. The ensemble model is generated using a machine learning algorithm, as shown in operation 630. The ensemble model is generated to include a network of interconnected nodes, each node being fed with modal data in one or more modes. The knowledge generated at each node based on the modal data it contains is exchanged between different nodes in the network via interconnects, and this knowledge is propagated to other nodes, thereby building the knowledge. To improve the accuracy of finger gesture prediction / validation, modal data for each mode may be weighted to indicate accurate predictions of finger gestures using the modal data for that mode. Outputs are defined in an ensemble model (also referred to herein as “ML Model 150”) based on the cumulative weights of various modal data collected for finger gestures, with each output meeting a certain level of prediction criteria for predicting finger gestures. Outputs are identified from the ensemble model for finger gestures as shown in Operation 640. Outputs are identified such that they at least meet the prediction criteria defined or required for finger gestures. The outputs identified from the ensemble model are used to define inputs for finger gestures.
[0043] Figure 7 shows components of an exemplary device 700 that can be used to carry out various embodiments of the present disclosure. This block diagram shows device 700 which may or may incorporate a personal computer, video game console, personal digital assistant, head-mounted display (HMD), wearable computing device, laptop or desktop computing device, server, or any other digital device suitable for carrying out embodiments of the present disclosure. For example, device 700 represents not only a first device but also a second device in the various embodiments discussed herein. Device 700 includes a central processing unit (CPU) 702 for running software applications and, optionally, an operating system. The CPU 702 may consist of one or more homogeneous or heterogeneous processing cores. For example, the CPU 702 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments may be implemented using one or more CPUs having a microprocessor architecture particularly suited to highly parallel and computationally intensive applications, such as interpreting queries, identifying contextually relevant resources, and immediately executing and rendering contextually relevant resources within a video game. The device 700 may be localized to a player playing a game segment (e.g., a game console), or remote from the player (e.g., a backend server processor), or one of many servers in a game cloud system that uses virtualization for remote streaming of gameplay to client devices.
[0044] Memory 704 stores applications and data used by the CPU 702. Storage 706 provides non-volatile storage and other computer-readable media for applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray®, HD-DVDs, or other optical storage devices, as well as signal transmission and storage media. User input device 708 communicates user input from one or more users to device 700, and examples of user input device 708 may include keyboards, mice, joysticks, touchpads, touchscreens, still recorders / cameras or video recorders / cameras, gesture-recognizing tracking devices, and / or microphones. The network interface 714 enables device 700 to communicate with other computer systems via an electronic communication network, which may include wired or wireless communication over a local area network and a wide area network such as the Internet. The audio processor 712 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 702, memory 704, and / or storage 706. The components of device 700, including the CPU 702, memory 704, data storage 706, user input device 708, network interface 714, and audio processor 712, are connected via one or more data buses 722.
[0045] The graphics subsystem 720 is further connected to the data bus 722 and the components of device 700. The graphics subsystem 720 includes a graphics processing unit (GPU) 716 and graphics memory 718. The graphics memory 718 includes display memory (e.g., a frame buffer) used to store pixel data for each pixel of the output image. The graphics memory 718 may be integrated into the same device as the GPU 716, connected as a separate device from the GPU 716, and / or implemented within memory 704. Pixel data can be provided directly from the CPU 702 to the graphics memory 718. Alternatively, the CPU 702 provides the GPU 716 with data and / or instructions defining a desired output image, from which the GPU 716 generates pixel data for one or more output images. The data and / or instructions defining a desired output image can be stored in memory 704 and / or graphics memory 718. In the embodiment, the GPU 716 includes a 3D rendering function for generating pixel data for output images from instructions and data defining the scene geometry, lighting, shading, texturing, motion, and / or camera parameters. The GPU 716 may further include one or more programmable execution units capable of executing shader programs.
[0046] The graphics subsystem 720 periodically outputs pixel data of an image from the graphics memory 718 and displays it on the display device 710. The display device 710 can be any device capable of displaying visual information in response to signals from device 700, including CRT, LCD, plasma, and OLED displays. In addition to the display device 710, the pixel data can also be projected onto a projection surface. Device 700 can, for example, provide analog or digital signals to the display device 710.
[0047] It should be noted that access services delivered across a wide geographical area, such as providing access to games in their current form, often utilize cloud computing. Cloud computing is a computing style in which dynamically scalable, often virtualized, resources are delivered as a service over the internet. Users do not need to be experts in the technical infrastructure of the "cloud" that supports them. Cloud computing can be categorized into different services such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services often deliver common applications, such as video games, online, accessible through a web browser, but the software and data are stored on servers in the cloud. The term "cloud" is used as a metaphor for the internet, based on how the internet is depicted in computer network diagrams, and is an abstract concept of concealing complex infrastructure.
[0048] In some embodiments, a game server may be used to run a persistent information platform for video game players. Most video games played over the internet operate via a connection to a game server. Typically, the game uses a dedicated server application to collect data from players and distribute it to other players. In other embodiments, the video game may run on a distributed game engine. In these embodiments, the distributed game engine may run on multiple processing entities (PEs), so that each PE runs a functional segment of a given game engine on which the video game runs. Each processing entity is considered simply a compute node by the game engine. A game engine typically performs a functionally diverse set of operations to run a video game application, along with additional services that the user experiences. For example, a game engine implements game logic and runs game calculations, physical processes, geometry transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. Additional services may include, for example, messaging, social utilities, audio communication, gameplay / playback functions, and help functions. The game engine may run on an operating system virtualized by a hypervisor on a specific server, but in other embodiments, the game engine itself may be distributed across multiple processing entities, each residing in a different server unit of a data center.
[0049] According to this embodiment, each processing entity for performing an operation may be a server unit, a virtual machine, or a container, depending on the needs of each game engine segment. For example, if a game engine segment is responsible for camera transformations, that particular game engine segment will perform a large number of relatively simple mathematical operations (e.g., matrix transformations), and may be provisioned with a virtual machine associated with a graphics processing unit (GPU). Other game engine segments that require less complex operations may be provisioned with processing entities associated with one or more higher-powered central processing units (CPUs).
[0050] By distributing the game engine, it gains resilient computing characteristics that are not constrained by the capabilities of physical server units. Instead, the game engine is provisioned with more or fewer computing nodes as needed to meet the demands of the video game. From the perspective of the video game and the video game player, a game engine distributed across multiple computing nodes is indistinguishable from a non-distributed game engine running on a single processing entity, as the game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the end user with the video game output components.
[0051] The user accesses the remote service via a client device that includes at least a CPU, display, and I / O. The client device may be a PC, mobile phone, netbook, PDA, etc. In one embodiment, a network running on the game server recognizes the type of device the client is using and adjusts the communication method employed. In another example, the client device accesses the application on the game server over the internet using a standard communication method such as HTML.
[0052] It should be understood that a given video game or game application may be developed for a specific platform and specific associated controller device. However, when such a game is made available through a game cloud system, as presented herein, users may access the video game using different controller devices. For example, a game may be developed for a game console and its associated controller, but a user may access a cloud-based version of the game from a personal computer using a keyboard and mouse. In such a scenario, input parameter configuration can define a mapping from inputs that can be generated by the user's available controller devices (in this case, a keyboard and mouse) to inputs that are acceptable for running the video game.
[0053] In another embodiment, a user may access a cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen-driven device. In this case, the client device and controller device are integrated together within the same device, and input is provided by detected touchscreen inputs / gestures. For such a device, the input parameter configuration can define specific touchscreen inputs corresponding to game inputs for a video game. For example, buttons, directional pads, or other types of input elements may be displayed or overlaid during the video game to indicate positions on the touchscreen that the user can touch to generate game inputs. Gestures such as swiping in a specific orientation, or specific touch motions, can also be detected as game inputs. In one embodiment, to familiarize the user with control operations on the touchscreen, a tutorial showing how to input gameplay via the touchscreen may be provided to the user, for example, before starting gameplay of a video game.
[0054] In some embodiments, the client device acts as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection and transmits inputs from the controller device to the client device. The client device then processes these inputs and can subsequently transmit the input data to the cloud game server via a network (a network accessed via a local network device such as a router). However, in other embodiments, the controller itself can be a networked device with the ability to communicate inputs directly to the cloud game server via a network, without the need to first communicate such inputs through the client device. For example, the controller can connect to a local network device (such as the aforementioned router) to send and receive data with the cloud game server. Thus, while the client device may still be required to receive video output from the cloud-based video game and render it to its local display, input latency can be reduced by allowing the controller to transmit inputs directly to the cloud game server over the network, bypassing the client device.
[0055] In one embodiment, a networked controller and client device can be configured to transmit certain types of input directly from the controller to the cloud game server, and other types of input via the client device. For example, detection inputs that do not depend on any additional hardware or processing other than the controller itself can be transmitted directly from the controller to the cloud game server via the network, bypassing the client device. Such inputs may include button inputs, joystick inputs, and embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes). However, inputs that utilize additional hardware or require processing by the client device can be transmitted to the cloud game server by the client device. These may include video or audio captured from the game environment that can be processed by the client device before being transmitted to the cloud game server. In addition, input from the controller's motion detection hardware can be processed by the client device in conjunction with captured video to detect the controller's position and movement, which are then communicated to the cloud game server by the client device. It should also be understood that controller devices in various embodiments may also receive data (e.g., feedback data) from the client device or directly from the cloud game server.
[0056] In one embodiment, various technical examples can be implemented using a virtual environment via a head-mounted display (HMD). HMDs are sometimes referred to as virtual reality (VR) headsets. As used herein, the term “virtual reality” (VR) generally refers to user interaction with a virtual space / virtual environment, including viewing a virtual space through an HMD (or VR headset) in real-time response to the HMD’s (user-controlled) movements, in order to provide the user with the sensation of being in a virtual space or metaverse. For example, a user might see a three-dimensional (3D) representation of a virtual space when facing a given direction, and when the user turns to the side, thereby changing the orientation of the HMD in the same way, a representation of that side of the virtual space is rendered on the HMD. HMDs can be worn in the same way as glasses, goggles, or helmets and are configured to display video games or other metaverse content to the user. By being close to the user’s eyes and providing a display mechanism, HMDs can provide a highly immersive experience. Therefore, HMDs can provide each of the user's eyes with a display area that occupies a large portion, or even the entire, of the user's field of view, and can also provide viewing with three-dimensional depth and perspective.
[0057] In one embodiment, the HMD may include an eye-tracking camera configured to capture images of the user's eyes while the user interacts with the VR scene. The gaze information captured by the eye-tracking camera(s) may include the user's gaze direction and information related to specific virtual objects and content items in the VR scene that the user is paying attention to or is interested in interacting with. Thus, based on the user's gaze direction, the system can detect specific virtual objects and content items, such as game characters, game objects, and game items, that may be potential focal points for the user when the user is interested in interacting with and engaging with them.
[0058] In some embodiments, the HMD may include outward-facing cameras (or more) configured to capture images of the user's real-world space, such as the user's body movements, and images of any real-world objects that may be placed in the real-world space. In some embodiments, the images captured by the outward-facing cameras can be analyzed to determine the position / orientation of real-world objects relative to the HMD. Using the known position / orientation of the HMD, real-world objects, and inertial sensor data from an inertial motion unit (IMU), the user's gestures and movements can be continuously monitored and tracked while the user interacts with the VR scene. For example, while interacting with a scene in a game, the user may perform various gestures, such as pointing to a specific content item in the scene or walking in its direction. In one embodiment, the gestures may be tracked and processed by the system to generate predictions of interactions with specific content items in the game scene. In some embodiments, machine learning may be used to facilitate or assist in the above predictions.
[0059] While using the HMD, various types of one-handed and two-handed controllers may be used. In one embodiment, the controller itself can be tracked by tracking the lights contained within the controller, or by tracking the shape, sensors, and inertial data associated with the controller. By using these various controllers, or even simple hand gestures performed and captured by one or more cameras, it becomes possible to interface with, control, manipulate, interact with, and participate in a virtual reality environment or metaverse rendered on the HMD. In some cases, the HMD can be wirelessly connected to cloud computing and gaming systems via a network. In some embodiments, the cloud computing and game system maintains and runs the video game being played by the user. In some embodiments, the cloud computing and game system is configured to receive input from the HMD and interface objects over a network. The cloud computing and game system is configured to process the input to affect the game state of the running video game. Outputs from the running video game, such as video data, audio data, and haptic feedback data, are sent to the HMD and interface objects. In other embodiments, the HMD can communicate wirelessly with the cloud computing and game system via an alternative mechanism or channel, such as a cellular network.
[0060] Furthermore, while embodiments of this disclosure may be described in relation to head-mounted displays, it will be understood that other embodiments may instead use non-head-mounted displays, including but not limited to portable device screens (e.g., tablets, smartphones, laptops, etc.) or any other type of display, which can be configured to render video and / or provide a display of an interactive scene or virtual environment according to this embodiment. Naturally, the various embodiments specified herein may be combined or assembled into a specific implementation using the various features disclosed herein. Thus, the examples provided are only a selection of possible examples and do not limit the various embodiments, which may be defined by combining various elements. In some examples, some embodiments may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
[0061] As described above, embodiments of the present disclosure for communication between computing devices can be implemented using a variety of computer device configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, head-mounted displays, and wearable computing devices. Embodiments of the present disclosure can also be implemented in distributed computing environments where tasks are performed by remote processing devices linked via wired or wireless networks.
[0062] In some embodiments, communication may be facilitated using wireless technology. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network where the service area covered by a provider is divided into small geographical areas called cells. Analog signals representing sound and images are digitized by a telephone, converted by an analog-to-digital converter, and transmitted as a bitstream. All 5G wireless devices within a cell communicate electromagnetically with local antenna arrays and low-power automatic transceivers (transmitters and receivers) within the cell via frequency channels allocated by transceivers from a frequency pool reused by other cells. The local antennas are connected to the telephone network and the internet by high-bandwidth optical fiber or wireless backhaul connections. As with other cell networks, mobile devices moving from one cell to another are automatically moved to the new cell. Naturally, a 5G network is merely an example of a type of communication network, and embodiments of this disclosure may also use previous generations of wireless or wired communication, as well as later generations of wired or wireless technologies that come after 5G.
[0063] Considering the embodiments described above, it is natural that the Disclosure can employ a variety of computer operations involving data stored in a computer system. These operations require the physical manipulation of physical quantities. Any of the operations described herein that form part of the Disclosure are useful machine operations. The Disclosure also relates to devices or apparatus for performing these operations. The apparatus may be configured specifically for a required purpose, or the apparatus may be a general-purpose computer selectively operated or configured by computer programs stored in the computer. In particular, various general-purpose machines can be used with computer programs written in accordance with the teachings herein. Alternatively, it may be more convenient to construct an apparatus more specialized to perform the required operations.
[0064] Although the method operations have been described in a specific order, it should be understood that other housekeeping operations may be performed between operations, or operations may be coordinated to occur at slightly different times, or operations may be distributed throughout the system to allow processing operations to occur at various processing-related intervals, as long as the processing of telemetry and game state data for generating the modified game state is performed in the desired manner.
[0065] One or more embodiments may also be made as computer-readable code on a computer-readable medium. The computer-readable medium can be any data storage device capable of storing data that can then be read by a computer system. Examples of computer-readable mediums include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. The computer-readable medium may include computer-readable tangible media distributed across network-connected computer systems so that the computer-readable code is stored and executed in a distributed manner.
[0066] While the embodiments described above have been described in some detail for clarity, it will be clear that certain changes and modifications can be made within the scope of the appended claims. Therefore, these embodiments should be considered illustrative rather than restrictive, and should not be limited to the details described herein, but may be modified within the scope of the appended claims and equivalents.
[0067] It should be understood that the various embodiments defined herein may be combined or assembled into specific embodiments that utilize the various features disclosed herein. Therefore, the examples provided are only a selection of possible embodiments and do not limit the various embodiments to which more embodiments can be defined by combining various elements. In some examples, some embodiments may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
Claims
1. A method for verifying the input provided to the controller, The controller detects a finger gesture provided by the user on its surface, and the finger gesture is used to define the input of an interactive application selected for the user's interaction. The controller receives multimodal data collected by tracking the finger gesture, and the multimodal data includes modal data corresponding to different modes captured for the finger gesture using a plurality of sensors and components associated with the controller. An ensemble model is generated using the multimodal data received from the multiple sensors and components, the ensemble model is continuously trained using additional multimodal data collected over time to produce different outputs, and the ensemble model is trained using a machine learning algorithm according to training rules defined for different finger gestures. The controller identifies the output from the ensemble model for the finger gesture detected on the surface of the controller, and the output is interpreted to define the input of the interactive application based on the finger gesture detected by the controller. A method executed by the server's processor.
2. In identifying the output, Weights are assigned to the modal data included in the multimodal data captured by each of the plurality of sensors and components, and the weights assigned to each modal data captured by a sensor or component among the plurality of sensors and components indicate an accurate prediction of the finger gesture using each modal data. The method according to claim 1, wherein the cumulative weight of the multimodal data is calculated using the weights assigned to each modal data of the modes included in the multimodal data, and the cumulative weight is used to identify the output for the finger gesture.
3. The weight assigned to the modal data captured by each of the multiple sensors and components is greater than the weight assigned to the modal data captured by each of the multiple sensors and components. The method according to claim 2, wherein the weights assigned to the modal data captured by the plurality of sensors are equal, and the weights assigned to the modal data captured by the plurality of components are equal.
4. The method according to claim 2, wherein the modal data captured for different modes included in the multimodal data are assigned equal weights.
5. The method according to claim 2, wherein weights are separately assigned to the modal data captured for each of the modes included in the multimodal data.
6. The method according to claim 1, wherein the ensemble model is trained according to training rules defined for the different finger gestures, the training rules are defined based on the anatomical form of the fingers, the position of the fingers related to input control on the controller, and the user's controller holding style.
7. The multimodal data includes video data, audio data, image data, sensor data, and wireless signals collected from the plurality of sensors and components. The method according to claim 1, wherein the identification of the output involves identifying, based on the finger gesture, a press of a physical button or a virtual button on the controller, or an input provided on a touchscreen interface, and the output is interpreted to define the input to the interactive application.
8. The plurality of sensors include one or a combination of inertial measurement unit (IMU) sensors, pressure sensors, proximity sensors, distance sensors, or capacitive sensors. The method according to claim 1, wherein the plurality of components include an image capture device, a wireless communication device, or a microphone array.
9. The method according to claim 1, wherein the multimodal data includes a Wi-Fi signal having forward and reflected signals captured by one or more wireless communication devices, the forward and reflected signals are interpreted to define a snapshot of the user's body parts, and the snapshot of the body parts is used to reconstruct the movement of one or more of the user's fingers when the user provides the finger gestures.
10. The aforementioned plurality of sensors and components include an image capture device, The method according to claim 1, wherein the multimodal data includes images of different positions held by the user's finger, captured by the image capture device when the user provides the finger gesture, and the angle of the image capture device dynamically adjusted to capture images of different positions of the finger, the dynamic adjustment being performed by automatically calibrating the angle of the image capture device in response to the presence of the finger provided by the user on the controller and the detection of the finger gesture.
11. The method according to claim 10, wherein the image capture device is a camera integrated into a mobile computing device, or a webcam, or an image capture device for a game console, or an image capture device for a computing device, or a camera for a head-mounted display, and the image capture device is communicably coupled to the controller.
12. If the image capture device is the camera of the mobile computing device, the mobile computing device is positioned on a holding structure coupled to the controller, the holding structure includes a motor, the motor is configured to receive and hold the mobile computing device, dynamically adjust the angle of the camera, and align with the angle calibrated to enable the capture of images of different positions held by the user's fingers when the user is performing a finger gesture. The method according to claim 11, wherein the retaining structure is a three-dimensional printed structure.
13. The plurality of sensors and components include one or more inertial measurement unit sensors (IMUs), the IMU being configured to detect the user's finger gestures on the surface of the controller and generate an IMU signal. The method according to claim 1, wherein the multimodal data includes the IMU signals received from one or more IMUs, the IMU signals are interpreted to identify the attributes of the finger gesture, and the attributes are used to identify user input in the controller.
14. The plurality of sensors and components include a microphone array embedded in or coupled to the controller. The aforementioned finger gesture is a button press on the controller. The method according to claim 1, wherein the multimodal data includes attributes of audio data captured by the microphone array, and the attributes of the audio data captured by a plurality of microphones in the microphone array are interpreted using triangulation techniques to identify the orientation and position of each microphone in the microphone array, and the orientation and position are used to determine the button press.
15. A method for defining inputs for an interactive application, The controller receives a finger gesture provided by the user on its surface, and the finger gesture is used to define the input of the interactive application selected for the user's interaction. The controller receives multimodal data capturing the attributes of the finger gesture, and the multimodal data includes modal data corresponding to different modes captured by multiple sensors and components associated with the controller. Weights are assigned to the modal data corresponding to each mode included in the multimodal data, and the weights assigned to each mode indicate an accurate prediction of the finger gesture using the modal data of each mode. Based on the weights assigned to each mode, the finger gestures and the multimodal data are processed to identify the inputs to the interactive application. A method executed by the server's processor.
16. In the processing of the finger gesture and the multimodal data, An ensemble model is generated using the multimodal data received from the aforementioned multiple sensors and components. The ensemble model is continuously trained using additional multimodal data collected over time to generate different outputs, and the ensemble model is trained using a machine learning algorithm according to defined training rules for different finger gestures. The method of claim 15, wherein an output from the ensemble model corresponding to the finger gesture detected on the surface of the controller is identified, and the output is interpreted to define the input of the interactive application.
17. The method according to claim 16, wherein the training rules are defined based on the anatomical form of the fingers, the position of the fingers related to input control on the controller, and the user's controller holding style.
18. The method of claim 15, wherein the processing of the finger gesture includes calculating an accumulated weight of the multimodal data using the weights assigned to the modal data of each mode included in the multimodal data, the accumulated weight being used to identify the input of the interactive application.
19. The plurality of sensors include one or a combination of inertial measurement unit (IMU) sensors, distance sensors, pressure sensors, proximity sensors, or capacitive sensors. The method according to claim 15, wherein the plurality of components include one or a combination of an image capture device, a wired communication device, a wireless communication device, or a microphone array.
20. The multimodal data includes a first set of the modal data captured by the plurality of sensors, and a second set of the modal data captured by the plurality of components. The method according to claim 15, wherein, in the weight assignment, a first weight is assigned to the modal data of each mode included in the first set of modal data, and a second weight is assigned to the modal data of each mode included in the second set of modal data, wherein the first weight is greater than the second weight.