Expansion of Game Controller Functionality Using Virtual Buttons with Hand Tracking
By employing multimodal data from various sensors on a handheld controller to train an ensemble model, the accuracy of detecting finger gestures is enhanced, addressing the unreliability of single-mode data in existing technologies.
Patent Information
- Application Number
- JP2024570970
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-02
- Filing Date
- 2023-04-21
- Publication Date
- 2025-06-19
- Estimated Expiration
- 2043-04-21
AI Technical Summary
Existing methods for detecting input provided by finger gestures on handheld controllers are unreliable due to reliance on single-mode data, leading to errors in interactive applications such as video games.
The use of multimodal data from multiple sensors and components associated with the handheld controller to generate and train a custom ensemble model for accurate detection and verification of finger gestures.
This approach significantly improves the accuracy of detecting and interpreting finger gestures, reducing errors and enhancing the reliability of input detection in interactive applications.
Smart Images

Figure 2025518793000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to identifying input provided by finger gestures on a handheld controller, and more specifically, to using multimodal data collected from a plurality of sensors and components associated with the handheld controller to verify the input provided by finger gestures.
Background Art
[0002] As the number of interactive applications and video games available to users on various devices increases, accurate detection of input provided via various devices has become particularly important. For example, in order to accurately affect the game state of a video game, the video game input provided by a user using a handheld controller must be appropriately identified and correctly interpreted. Relying solely on single-mode data (e.g., image tracking of finger gestures) can result in incorrect results in video games.
[0003] Embodiments of the present disclosure have been made under such a background.
Summary of the Invention
[0004] Embodiments of the present disclosure relate to systems and methods for providing multimodal finger tracking to detect and verify finger gestures provided on an input device, such as a handheld controller. Multimodal finger tracking and verification ensure that finger gestures are appropriately identified and correctly interpreted, thereby reducing errors resulting from relying solely on single-mode tracking. A custom finger tracking model (e.g., an ensemble model) is generated and trained using multiple modalities of data captured by a plurality of sensors and components associated with the handheld controller (hereinafter simply referred to as the "controller"), thereby increasing the accuracy of detecting and interpreting finger gestures.
[0005] Conventional methods for detecting input relied on a single data source model. For example, conventional methods relied on a general camera (i.e., a single data source) to detect and track a user's finger on a controller. The accuracy of tracking using a single source was unreliable and prone to errors, resulting in less desirable results in interactive applications. To overcome the drawbacks of conventional methods, multimodal data is used to provide input and is collected from multiple sensors and components associated with the controller that are used to verify finger gestures detected by the controller. The collected multimodal data is used to generate and train a multimodal data model, and it is used to correctly interpret finger gestures. Since data from multiple modes is used to generate and train the model, the multimodal data model is also referred to herein as an "ensemble model". The ensemble model is continuously trained according to training rules defined for various finger gestures using additional multimodal data collected over time. The output is selected from the ensemble model and is used to confirm / verify the finger gesture detected by the controller. The finger gesture can correspond to pressing a physical button, or pressing a virtual button defined on the controller, or an input provided on a touch screen interface disposed on the controller, and the output is identified to correspond to the correct interpretation of the finger gesture. The virtual button can be identified on any surface of the controller where no physical button is placed, and finger gestures on the virtual button, such as a single tap, or a double tap, or a press, or a swipe in a specific direction, can be defined.
[0006] This model incorporates multimodal finger tracking technology by considering several model components such as the finger tracking using the image feed from an image capture device, the IMU data from an inertial measurement unit (IMU) sensor placed within the controller, the wireless signals from wireless devices placed within the environment where the user is present, and the data from various sensors such as distance / proximity sensors, pressure sensors, etc., when generating and training an ensemble model. The ensemble model helps accurately detect the finger gestures provided to the controller by using data from more than one modality to track and verify the finger gestures.
[0007] In one embodiment, a method for validating an input provided to a controller is disclosed. The method includes detecting a finger gesture provided by a user on the surface of the controller. The finger gesture is used to define an input to an interactive application selected for user interaction. Multimodal data is collected by tracking the finger gesture on the controller using a plurality of sensors and components associated with the controller. An ensemble model is generated using the multimodal data received from the plurality of sensors and components. The ensemble model is continuously trained using additional multimodal data collected over time to generate different outputs, and the training follows training rules defined for different finger gestures. The ensemble model is generated and trained to define various outputs using a machine learning algorithm. The outputs are identified from the ensemble model of the finger gestures. The outputs identified from the ensemble model are interpreted to define the input to the interactive application.
[0008] In other embodiments, a method for defining inputs of an interactive application is disclosed. The method includes receiving a finger gesture provided by a user on a surface of a controller. The finger gesture is used to define an input of the interactive application selected for interaction by the user. Multimodal data capturing attributes of the finger gesture on the controller is received from a plurality of sensors and components associated with the controller. Weights are assigned to the modal data corresponding to each mode included in the multimodal data captured by the plurality of sensors and components. The weight assigned to each mode indicates that the modal data of each mode is used to accurately predict the finger gesture. The finger gesture and the multimodal data are processed based on the weights assigned to each mode to identify an input of the interactive application corresponding to the finger gesture detected by the controller.
[0009] Other aspects and advantages of the present disclosure will become apparent from the following detailed description, which illustrates the principles of the present disclosure by way of example in conjunction with the accompanying drawings.
[0010] The present disclosure may be best understood by reference to the following description, which is to be interpreted in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0011]
Figure 1A
Figure 1B
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0012] In the following embodiments for carrying out the invention, some specific details are described in order to provide a complete understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without some or all of these specific details. In other instances, well-known process steps are not described in detail so as not to obscure the present disclosure.
[0013] As the number of interactive applications increases, a method for accurately identifying inputs provided using various devices becomes particularly important. Specifically, in interactive applications such as high-intensity video game applications, the inputs provided using a controller need to be appropriately detected and accurately interpreted so that the game state of the video game can be correctly updated in a timely manner. For this purpose, multiple modalities of data that detect inputs provided by a user with an input device such as a handheld controller and track finger gestures on the controller in response thereto are collected from a plurality of sensors and components associated with the controller. The multimodal data is used to generate and train a custom finger tracking model, and then this model is used to identify an output corresponding to the finger gesture detected on the controller. The trained custom finger tracking model improves the accuracy of predicting finger gestures better than a finger tracking model using only a general camera because finger gesture predictions rely on multiple modal data sources to verify finger gestures.
[0014] Modal data, and some of the plurality of sensors and components associated with a controller that captured the modal data, include (a) inertial measurement unit (IMU) data captured using an IMU sensor, such as a magnetometer, gyroscope, accelerometer, etc., (b) wireless communication signals, including forward and reflected signals captured using wireless communication devices, such as Bluetooth (trademark)-compatible devices and Wi-Fi routers, placed in the environment, (c) audio data from a microphone array, (d) sensor data captured using distance and / or proximity sensors, and (e) image data captured using an image capture device(s). The foregoing modal data and sensors, components are presented by way of example only and should not be considered exhaustive or limiting. The various sensors and components capture attributes of finger gestures in various modal forms used to generate an ensemble model. When training an ensemble model to improve the accuracy of identifying finger gestures detected by the controller, apply training rules defined for the various finger gestures. In some embodiments, a multimodal data collection engine running on a server collects the various modal data captured by the plurality of sensors and components at the controller, and generates and trains an ensemble model. In other embodiments, the multimodal data collection engine can be executed on the controller itself or on a processor co-located with and coupled to the controller to reduce latency.
[0015] Each modal data captures some attributes of the finger gesture provided by the controller and uses these attributes to verify the finger gesture, so as to identify the correct input corresponding to the finger gesture in order to affect the result of the interactive application selected by the user for interaction. By using multiple modalities of data captured by tracking the finger gesture on the controller to provide additional verification to correctly determine the finger gesture, the correct input of the interactive application can be identified.
[0016] Figure 1A shows an example of an Abstract pipeline followed to build an ensemble model for detecting finger gestures according to some embodiments. The Abstract pipeline uses a multimodal finger tracking approach, and finger gestures are tracked using a plurality of sensors and components associated with the controller, and data generated from such tracking is used to identify and / or verify finger gestures on the controller. Each of the plurality of sensors and components captures modal data of a specific mode. For example, the plurality of modal data 120-1 captured by the sensors and components from finger tracking includes, by way of example, camera feed 120a1, inertial measurement unit sensor (IMU) data 120b1, WiFi data (including WiFi forward and reflected signals) 120c1, audio data 120d1, sensor data from distance / proximity sensor 120e1, and sensor data from pressure sensor 120f1. Of course, the foregoing list of data for capturing finger gestures is provided by way of example and should not be considered exhaustive or limiting, and other forms of modal data captured from different sensors and components can also be considered in identifying and / or verifying finger gestures. Modal data from each mode is analyzed to detect finger gestures. Considering modal data from a single mode, there may be an error in predicting finger gestures or a decrease in reliability. Therefore, to increase the accuracy of finger gesture detection, modal data from multiple modes is considered in the analysis. In some embodiments, modal data associated with different modes is voted on using a voting module, and weights are assigned to the modal data of each mode. The weights assigned to the modal data of each mode can be equal or unequal, and in some embodiments, the determination to assign equal or unequal weights (144) to the modal data of each mode is made based on the reliability of each mode in correctly predicting finger gestures.Next, by using the weights assigned to data in different modes for analysis, finger gestures are correctly predicted. When finger gestures are correctly predicted, the inputs related to the finger gestures are identified and used as interactive inputs to the interactive application selected by the user for interaction from the user. The details of the analysis are described with reference to FIG. 1B.
[0017] FIG. 1B shows various components of a system 100 for correctly detecting finger gestures at a controller so that inputs to an interactive application can be appropriately identified. The various components and sensors used to build the ensemble model represent an abstract pipeline that can be used to correctly detect finger gestures. The ensemble model is generated and trained using modal data from several modal components and sensors that capture different attributes of the finger gestures provided at the controller. As described above, the different modalities of data captured by tracking finger gestures include data obtained from an image feed, IMU data, WiFi signals (radio signals including reflected signals), distance / proximity sensors, pressure sensors, and the like. The attribute information regarding the finger gestures captured by each of these components and sensors is transferred to a server computing device for further processing. The server computing device processes the multimodal data to appropriately identify the finger gestures at the controller.
[0018] For this purpose, a system 100 for determining the input of an interactive application includes a controller 110, such as a handheld controller, used by a user to provide finger gestures, a plurality of sensors and components 120 associated with the controller 110 to capture various attributes of the finger gestures, and a server device 101 used to process multimodal data that captures the finger gestures and various attributes of the finger gestures and to verify the finger gestures. The finger gestures are provided by the user as the input of the interactive application selected for user interaction. The server computing device (or hereinafter simply referred to as "server") 101 uses a modal data collection engine 130 for collecting various modalities of data transferred by the plurality of sensors and components 120, and a modal data processing engine 140 for processing the multimodal data to identify the finger gestures and define the input of the interactive application. The modal data collection engine 130 and the modal data processing engine 140 may be part of a multimodal processing engine executed on the server 101.
[0019] In some embodiments, server 101 may be a game console or any other computing device co-located within the environment that the user operates. Next, the game console or computing device may be coupled to other game consoles via a network as part of a multiplayer setup for a video game. In some embodiments, controller 110 is a networked device, and controller 110 is directly coupled to remote server 101 via a network (not shown) such as the Internet. In the case of a networked device, controller 110 is coupled to the network via a router embedded within or external to controller 110. In other embodiments, controller 110 is coupled to remote server 101 via the Internet through a game console or another client device (not shown), and the game console or client device is co-located with controller 110. Controller 110 is paired with server 101 as part of an initial setup, or when the presence of controller 110 is detected near the game console, another client device, or router (when server 101 is connected to controller 110 via the game console, router, or other computing device co-located with controller 110), or when activation of controller 110 by the user is detected (i.e., when server 101 is remotely located from controller 110). Details of finger gestures provided on the surface of controller 110 are transferred to the game console / server 101 for processing.
[0020] In response to detecting a finger gesture provided by a user on the surface of the controller 110, various sensors and components 120 associated with the controller 110 automatically become active to capture different attributes of the user's finger gesture. Some of the sensors and components 120 that automatically become active to collect various attributes of the finger gesture include the image capture device(s) 120a, IMU 120b, WiFi device(s) 120c, microphone array 120d, distance / proximity sensor 120e, and pressure sensor 120f. In addition to the aforementioned sensors and components, other sensors and / or components can also be used to collect the attributes of the finger gesture at the controller.
[0021] In some embodiments, the image capture device 120a is a camera incorporated within a mobile computing device such as a cellular phone or a tablet computing device. Alternatively, the camera can be a webcam, or a console camera, or a camera that is part of an HMD. The camera, or the device in which the camera is incorporated, is paired with the game console / server 101 using the pairing engine 125a. This pairing enables the image capture device (i.e., the camera) to receive an activation signal from the game console / server 101 to capture an image of the user's finger gesture on the controller 110. When the game console / server 101 detects a finger gesture on the surface of the controller 110, it generates an activation signal for the image capture device. The mobile computing device with a built-in camera is supported on a holding structure disposed on the controller 110, so that the camera in the mobile computing device can capture a close-up view of various features of the finger gesture provided on the controller 110. More information regarding the holding structure will be described with reference to FIGS. 5A and 5B. When activated, the image capture device captures images of various attributes of the finger gesture, including the finger used to provide the gesture, the position of the finger relative to input controls on the controller 110 (e.g., buttons, touch screen surface, other interactive surfaces, etc.), the movement of the finger on the controller, the type of finger gesture provided (e.g., single tap, double tap, slide gesture, etc.). The images capturing the attributes of the finger gesture are transferred as image data camera feeds to a modal data collection engine 130 running on the game console / server 101.
[0022] In response to the activation of various sensors and components, the inertial measurement unit (IMU) sensor 120b integrated within the controller 110 is used to capture IMU signals related to the finger gesture. In some embodiments, while the user holds the controller 110 within their hand, the IMU signals captured by the IMU sensor are used to distinguish different finger gestures detected by the controller 110. For example, an IMU signal that captures a subtle tapping at a defined position inside the back of the controller 110 can be interpreted as meaning a first input (i.e., virtual button 1), and an IMU signal that captures a subtle tapping at a defined position inside the front of the controller 110 that does not include any button or interactive interface can be interpreted as meaning a second input (i.e., virtual button 2), an IMU signal that captures a tapping at the upper right corner of the back of the controller 110 can be interpreted as meaning a third input (i.e., virtual button 3), tapping the upper left corner of the back of the controller 110 can be interpreted as meaning a fourth input (i.e., virtual button 4), an IMU signal that captures a tapping at the back of the controller using the middle finger can be interpreted as meaning a fifth input (i.e., virtual button 5), tapping on a physical button on the front of the controller 110 can be interpreted as meaning a sixth input (e.g., pressing the physical button), and so on.
[0023] In some embodiments, the functionality of the controller 110 can be extended by using virtual buttons defined by tracking finger gestures. The extended functionality enables the user to interact with more than one application simultaneously, and such interactions can be performed without the need to interrupt one application for the sake of another. The virtual buttons can identify the position of the fingers when the user is holding the controller 110 and can be defined by finger gestures provided by the user with respect to the identified finger positions. In some cases, when the user is playing a game, for example while it is running on a game controller or game server, the user may also be listening to music provided by a second application (e.g., a music application). Typically, when a user needs to interact with a music application, the user has to pause the currently playing game, access the menu to interact with the music application, and use one of the buttons on the controller 110 or the interactive surface to advance to the next song on the user's playlist. To avoid the user from pausing the gameplay of the game and to provide other ways to interact with the music application while playing the game, virtual buttons can be defined to extend the functionality of the controller 110. Since the virtual buttons can be defined and associated with pre-assigned commands, the user can use the virtual buttons to interact with the music application without having to interrupt the user's current gameplay.
[0024] In another embodiment, tracking of finger gestures while the user holds the controller can be used, enabling users with certain disabilities to communicate in online games. For example, tracking of finger gestures can be used to detect the positions and gestures of different fingers while a disabled user holds the controller 110. These finger positions and gestures can be interpreted as Morse code inputs (dots and dashes for taps and swipes) using a machine learning (ML) algorithm, and such interpretation can be performed by the ML algorithm by recognizing the user's disability provided in the user's user profile. Further, the Morse code input can be converted to text characters or provided as game input. Using the text characters, voice communication with other players / audience / users can be achieved by converting the text to speech, or provided as a text response on a chat interface. The Morse code input can be interpreted to correlate with game input and used when affecting the game state of the game the user is playing. The foregoing uses of tracking and interpreting finger gestures to identify virtual buttons and / or inputs for interactive applications of disabled users are provided as examples and should not be considered exhaustive or limiting, and other uses can also be envisioned.
[0025] Figures 2A and 2B show some exemplary signal amplitude variations captured in the respective IMU signals of different finger gestures using the IMU sensor 120b. Figure 2A shows the position of the user's finger with respect to different input controls (i.e., buttons and touch screen interactive interfaces) defined on the controller 110 when the user is operating the controller 110. Figure 2B shows some exemplary amplitude variations captured for three different finger gestures detected at different positions on the controller in some embodiments. The amplitude variations shown in Figure 2B represent the amplitude variations along the X-axis, Y-axis, and Z-axis captured in the IMU signals of different finger gestures. For example, the amplitude variations shown along the X-axis, Y-axis, and Z-axis within the box "VB1" relate to finger gesture 1, which can be interpreted to mean that finger gesture 1 provided by the user to the controller 110 corresponds to an input regarding virtual button 1. Similarly, the amplitude variations shown along the X-axis, Y-axis, and Z-axis within the box "VB2" relate to finger gesture 2, which can be interpreted to mean that finger gesture 2 provided by the user to the controller 110 corresponds to an input regarding virtual button 2, and the amplitude variations shown along the X-axis, Y-axis, and Z-axis within the box "RB1" correspond to finger gesture 3, which can be interpreted to mean that finger gesture 3 provided by the user to the controller 110 corresponds to an input regarding real button 1. Of course, virtual buttons 1, 2, 3, etc., and real buttons 1, 2, 3, etc., can be defined to relate to different inputs of different interactive applications. The IMU signals capturing the fine signals are transferred to the modal data collection engine 130 as IMU sensor data.
[0026] Returning to FIG. 1B, in response to the activation of various sensors and components, the WiFi devices (i.e., wireless devices) 120c distributed in the environment where the user is located start to capture WiFi signals. Using the data captured by the WiFi signals, various physical positions of the user, such as the position of the user, various body parts including the hands and fingers, and the movement of the body including finger movement / finger gestures are detected. Using the WiFi signals and the data from the reflected signals of the WiFi signals, different finger movements while the user holds the game controller 110 can be detected.
[0027] FIGS. 3A and 3B show various WiFi signals captured within the environment (i.e., geographical location) where the user (i.e., user 1) is located using the WiFi device. The WiFi device includes a transmitter 301 such as other computing devices including routers, laptops, personal digital assistants (e.g., voice assistants), and a receiving device 302 such as a second laptop or desktop computing device. The list of the aforementioned devices representing the transmitter 301 and the receiver 302 is provided as an example only and should not be considered exhaustive or limiting. The transmitter 301 continuously transmits WiFi signals, and the receiver 302 receives various WiFi signals. Figure 3A shows various WiFi signals transmitted by the transmitter 301 and received by the receiver 302 within the room where the user is located. A part of the WiFi signals received by the receiver 302 includes the WiFi signal 303 reflected by the wall of the room, the WiFi signal 304 reflected by the user 1, the WiFi signal 305 representing the line of sight of the user 1, and the WiFi signal 306 reflected by the floor of the room. The receiver 302 continuously monitors the received WiFi signals to determine the movement of the user within the geographical location of the room. First, the channel state information (CSI) representing the channel characteristics of the communication link between the transmitter 301 and the receiver 302 for the geographical location where the user 1 is operating is first determined using the WiFi signals transmitted by the transmitter 301 and received by the receiver 302 when no object or user exists between the transmitter 301 and the receiver 302. Using the channel characteristics, a baseline of the WiFi signal is established, taking into account the combined effects of scattering, fading, and signal strength attenuation due to distance. Next, the CSI is determined when the user 1 is present at the geographical location (e.g., the room) and when the user 1 moves within that room. The variations in the CSI are due to the user's body blocking or reflecting one or more of the aforementioned WiFi signals. Using this variation in each one or more WiFi signals collected over a certain period, the movement of the user within the geographical location (e.g., the room), including the various finger movements of the user 1 while holding and operating the game controller 110, is determined. Using the channel characteristics of the WiFi signal, a snapshot of the body part including the finger of the user 1 can be captured. These snapshots can then be used to reconstruct the body part to determine which body part (e.g., finger) has moved. This WiFi signal can include the signal provided by the router and the Bluetooth (registered trademark) signal provided by the controller.
[0028] Figure 3B shows, in some embodiments, the variation in the CSI signal amplitude of the Wifi signal received from a WiFi device (e.g., a single subcarrier) caused by the movement of a user. The variation in the signal amplitude is plotted against time. The variation in the signal amplitude 321 is shown for the WiFi signal transmitted when there is no user or object between the transmitter 301 and the receiver 302 within the geographical location. In some embodiments, the variation in the signal amplitude 321 establishes the baseline variation in the CSI signal amplitude. The variations in the signal amplitude 322 - 326 capture the variations caused by the movement of the user. For example, the variation in the signal amplitude 322 shows an example of the variation of the WiFi signal when detecting that user 1 is sitting in a room (i.e., the geographical location), the variation in the signal amplitude 323 captures the variation of the WiFi signal when detecting the opening or closing of the door of the room, the variation in the signal amplitude 324 captures the variation of the WiFi signal when detecting that user 1 is typing on an input device such as a keyboard, the variation in the signal amplitude 325 captures the variation of the WiFi signal when detecting that user 1 is waving a hand, and the variation in the signal amplitude 326 captures the variation of the WiFi signal when detecting that user 1 is walking in the room. Therefore, it is possible to detect different finger movements when the user is operating the controller 110 using the WiFi signal. The WiFi signal capturing various user movements is transferred to the modal data collection engine 130 as a WiFi signal.
[0029] Returning to FIG. 1B, in response to finger gesture detection on the controller 110, the microphone array 120d embedded in or attached to the controller 110 becomes active. The finger gesture may be a physical button press on the controller 110 or a virtual button press, and the activated microphone array 120 is configured to capture the attributes of the button press sound. Specifically, the microphones within the microphone array 120d cooperate together to determine the direction in which the sound is generated and to identify the location. Using the attributes (i.e., direction and location) of the finger gesture, it is determined whether the finger gesture corresponds to a physical button press or a virtual button press. In some embodiments, different locations on the controller other than where the physical buttons and touch screen interface are located can represent different virtual buttons. For example, the left-hand side corner within the back of the controller 110 can be defined to represent virtual button 1, the right-hand side corner within the back of the controller 110 can be defined to represent virtual button 2, the central top position on the back of the controller 110 can be defined to represent virtual button 3, and so on. By interpreting the attributes of the sound caused by the finger gesture, it is possible to determine whether the finger gesture corresponds to a physical button press or a virtual button press, and which of the physical or virtual buttons was pressed. The attributes of the sound captured by the microphone array 120d are transferred to the modal data collection engine 130.
[0030] FIG. 4 shows an exemplary microphone array 120d embedded within or coupled to the controller 110 to capture sound generated from a user's finger gesture on the surface of the controller 110. The microphone array 120d is shown to include four microphones (401a, 401b, 401c, and 401d). The intensity of the sound captured by each of the microphones 401 varies based on the distance of the sound from each respective microphone 401. The sound signals captured by each of the microphones 401a - 401d are then transferred to a digital signal processor (DSP) 402 for processing. In some embodiments, the DSP 402 is configured to assign different weights to different sounds captured by the microphones. The sounds captured by different microphones within the microphone array 120d, and the relative weights assigned to each of the sounds, are analyzed using, for example, triangulation techniques to identify various attributes of the captured sounds such as direction, position, duration, volume, frequency, etc., and these attributes are used with the relative weights to determine a specific button press, or a swipe and the direction of the swipe, etc. In other embodiments, instead of assigning separate weights to each sound, the DSP 402 assigns separate weights to different attributes of each sound detected / captured by the microphones within the microphone array 120d, and uses these weights and the detected attributes to associate the sound with a specific button press or a finger swipe. The weights assigned to different attributes can be used to determine which sounds to ignore as ambient sound / noise and which sounds to focus on in order to determine the finger gesture. The details of the sound attributes and analysis are then transferred as electrical signals to the modal data collection engine 130.
[0031] Returning to FIG. 1B, in some embodiments, in response to the detection of a finger gesture on the controller 110, one or more distance / proximity sensors 120e become active and capture the attributes of the finger gesture provided by the controller 110. Similar to the microphone array, the distance / proximity sensors 120e can be used to capture the attributes of a finger gesture provided on the back side of the controller 110. The attributes of the finger gesture provided on the back side of the controller 110 captured by the distance / proximity sensors 120e can be used to independently determine the pressing of a virtual button or, in combination with the attributes of the sound captured by the microphone array 120d, to further verify the pressing of the virtual button determined from the finger gesture. The additional verification provided by the data captured by the distance / proximity sensors 120e makes the detection of the pressing of the virtual button from the finger gesture more accurate. In some embodiments, the distance / proximity sensors 120e can include ultrasonic sensors, infrared sensors, LED time-of-flight sensors, capacitive sensors, and the like. The foregoing list of distance / proximity sensors 120e is provided by way of example and should not be considered exhaustive or limiting. The attributes of the finger gesture captured by the distance / proximity sensor 120e are transferred to the modal data collection engine 130. In addition to the distance / proximity sensor 120e, the pressure sensor 120f is also activated to capture the attributes of the pressure (such as position, amount of pressure, pressure application time, finger used for pressure application, etc.) applied to different positions of the controller 110 by the finger gesture. For example, if the pressure application time is less than the threshold value, the attributes of the pressure captured by the pressure sensor 120f can be ignored. If the pressure application time is greater than the threshold value, the attributes of the applied pressure can be transferred as an input to the modal data collection engine 130. Similarly, if the amount of applied pressure is less than the threshold amount, the data captured by the pressure sensor 120f can be ignored. However, if the amount of applied pressure is greater than the threshold amount, the data captured by the pressure sensor 120f can be regarded as an input to the modal data collection engine 130. The attributes of the pressure given via a finger gesture that meets or exceeds the threshold / amount are transferred to the modal data collection engine 130 as pressure sensor data.
[0032] As described above, in some embodiments, the image capture device can be a camera incorporated within a mobile computing device such as a cellular phone or a tablet computing device. In these embodiments, for example, the camera of the cellular phone may be preferred over a web camera, or a console camera, or a camera incorporated within an HMD. In other embodiments, in addition to the web camera / console camera / HMD camera, a camera incorporated within a cellular phone (i.e., a mobile computing device) can be used to capture an image of the attributes of the finger gesture. In embodiments where the camera of the mobile computing device is used to capture an image of the attributes of the finger gesture, the mobile computing device (e.g., a cellular phone) can be coupled to the controller 110.
[0033] Figures 5A and 5B show such an embodiment in which a mobile computing device (e.g., mobile phone 502) is coupled to controller 110. In some embodiments, mobile phone 502 is coupled to controller 110 using a holding structure 504. FIG. 5A shows a front perspective view of controller 110 with holding structure 504 receiving and holding mobile phone 502. FIG. 5B shows a rear view of holding structure 504 coupled to controller 110 and configured to receive, hold, and operate mobile phone 502. In some embodiments, holding structure 504 is a three-dimensional (3D) printed structure that can be attached to controller 110. In some embodiments, the 3D printed structure includes a motor (not shown) that moves mobile phone 502 to different positions to enable the camera of mobile phone 502 to capture various attributes of finger gestures.
[0034] Referring simultaneously to FIGS. 1B, 5A, and 5B, in some embodiments, to accommodate different mobile phone models (e.g., mobile phone size, camera position, number of cameras, etc.), game console / server 101 first performs an automatic pairing operation to pair mobile phone 502 with game console / server 101. A signal to initiate the pairing operation is sent from pairing engine 125a to the mobile phone. After successful pairing of mobile phone 502 to game console / server 101, mobile phone 502 is attached to the holding structure by initiating a calibration operation. Calibration engine 125b is used to determine the type and model of mobile phone 502 and send a signal to the holding structure to adjust the size of the holding structure to accommodate mobile phone 502. In response to a signal from the calibration engine 125b, the motor that operates the holding structure 504 adjusts the size of the holding structure 504 so that the mobile phone 502 can be fixedly received. In addition to automatically calibrating the size of the holding structure to accommodate the mobile phone 502, the calibration engine 125b calibrates the angle at which the mobile phone needs to be adjusted to capture images of different positions of the finger. This angle is automatically calibrated in response to detecting the presence of the user's finger and the finger gestures provided on the controller 110, and is dynamically determined based on the position of the user's finger when the user is providing finger gestures to the controller 110. As part of the angle calibration operation, the calibration engine 125b tracks the user's hand and fingers and sends a second signal to the controller 110 and / or the holding structure 504 to adjust the motor to move the mobile phone 502, so that the camera(s) of the mobile phone 502 are aligned with the calibrated angle. In response to the second signal from the calibration engine 125b, the motor of the holding structure 504 moves and rotates the movable parts of the holding structure 504 to be engaged to achieve good tracking of the hand and finger positions. When the mobile phone is moved into place, the camera(s) of the mobile phone 502 become active and capture images of various attributes of the finger gesture. The captured images are streamed as image data camera feeds to the game console / server 101.
[0035] The modal data collection engine 130 collects inputs from a plurality of sensors and components 120 to generate multimodal data. The multimodal data is processed to identify the mode and the amount of modal data captured for each mode included in the multimodal data collected from the sensors and components. The details of the mode, the amount of modal data for each mode, and the multimodal data captured by the sensors and components 120 are transferred by the modal data collection engine 130 to the modal data processing engine 140 for further processing.
[0036] The modal data processing engine 140 analyzes the modal data of each mode included in the multimodal data to identify and / or verify finger gestures at the controller. As previously described with reference to FIG. 1A, modal data from multiple modes is considered in the analysis to increase the accuracy of finger gesture detection. As part of the analysis, weights are assigned to the modal data associated with each mode. The weights assigned to the modal data of each mode can be equal or unequal. In some embodiments, the determination (144) to assign equal or unequal weights to the modal data of each mode is made based on the accuracy of correctly predicting finger gestures using the modal data of each mode. In some implementations, the multimodal data can be roughly classified into a first set of modal data captured using multiple sensors and a second set of modal data captured using components. In some embodiments, the weights assigned to the modal data for each mode captured by the sensors are greater than the weights assigned to the modal data for each mode captured by the components. For example, greater weights are assigned to IMU sensor data captured by an IMU sensor or distance sensor data captured by a distance / proximity sensor than to image data camera feeds captured by an image capture device. In another example, a greater weight is assigned to the IMU sensor data and audio data than to the WiFi signal data. The weight assignment engine 144a is used to identify the mode associated with each modal data included in the multimodal data and the reliability of the modal data of each mode when predicting finger gestures. Based on the reliability of each mode, the weight assignment engine 144a assigns weights to the modal data. The weights assigned to the data of different modes are used in combination to correctly predict finger gestures. For example, the modal data processing engine 140 uses the weights assigned to the modal data of each mode included in the multimodal data to generate a cumulative weight. The cumulative weight is used to correctly predict finger gestures. Since the game console / server 101 depends on data of more than one mode to identify and / or verify finger gestures, the predicted finger gestures are more accurate. Once the finger gesture is identified / verified, the input regarding the finger gesture is then identified and used as user input to an interactive application such as a video game, affecting the game state.
[0037] In some embodiments, the modal data processing engine 140 analyzes multimodal data captured by a plurality of sensors and components using a machine learning (ML) algorithm 146 to identify and / or verify finger gestures provided to the controller 110. The ML algorithm 146 uses the multimodal data to generate and train an ML model 150. The ML model 150 is used to predict and / or verify finger gestures to identify appropriate inputs corresponding to the predicted / verified finger gestures and use them to affect the state of an interactive application (e.g., a video game) selected by the user for interaction. The ML algorithm 146 uses a classifier engine (i.e., a classifier) 148 to generate and train the ML model 150. The ML model 150 includes a network of interconnected nodes, and each successive pair of nodes is connected by an edge. The classifier 148 is used to input various nodes into the network of interconnected nodes of the ML model 150, and each node relates to modal data of one or more modes. Interrelationships between the nodes are established, the complexity of the modal data of different modes is understood, an output used to identify or verify finger gestures is identified, and an input corresponding to the finger gesture is identified.
[0038] In some embodiments, classifier 148 is pre - defined for different modes to understand the complexity of the modal data of each mode when correctly predicting and / or verifying the finger gestures provided to the controller. Classifier 148 uses the modal data captured in real - time by sensors and components when and if a finger gesture is provided at controller 110 to further train ML model 150, and uses ML model 150 to determine the amount of influence that the modal data of each mode has on the correct prediction / verification of the finger gesture. ML model 150 can be trained according to training rules 142 to increase the accuracy of finger gesture prediction. The training rules are defined for each finger gesture based on, for example, the anatomical form of the finger, the way the controller is held, the position of the finger relative to the button, etc. Machine learning (ML) algorithm 146 uses the modal data of different modes included in the multimodal data as inputs to the nodes of ML model 150, gradually updates the nodes using additional multimodal data received over time, and adjusts the output to meet pre - defined criteria for different finger gestures. ML algorithm 146 uses reinforcement learning to strengthen ML model 150 by using an initial set of multimodal data to build ML model 150, learns the complexity of each mode and how the modal data of each mode affects the correct prediction / verification of the finger gesture, and reinforces the learning and strengthening of the model using additional modal data received over time. The adjusted output of ML model 150 is used to correctly predict / verify different finger gestures. The output from the adjusted output is selected to correspond to the finger gesture, and such selection may be based on the cumulative weights of the multimodal data indicating the correct prediction of the finger gesture.
[0039] In some embodiments, the tracking of finger gestures using modal data captured for different modes may be user-specific. For example, each user may handle the controller 110 in a different way. For example, a first user may hold the controller 110 in a particular way and provide inputs on the controller in a particular way, or at a particular speed or pressure, etc. When a second user uses the controller 110 to provide inputs by finger gestures, those ways of holding the controller 110 or providing inputs using it can be different from those of the first user. To account for the different user ways of handling the controller 110, a reset switch may be provided on the controller 110, making it possible to reset or reprogram the tracking of finger gestures, so that finger gesture interpretation is user-specific and can be made user-independent. The reset switch can be defined as a particular button press or a particular button press sequence. In another example, the reset or reprogramming of finger gesture tracking can be done on demand from the user, and such a request can be based on, for example, a particular context in which the finger gesture needs to be tracked.
[0040] In various embodiments discussed herein, a modal data collection engine 130 and a modal data processing engine 140 with an ML algorithm 146 for identifying / verifying finger gestures implemented on the server 101 have been described. However, the modal data collection engine 130 and the modal data processing engine 140 with the ML algorithm 146 can be implemented locally on a disk within a computing device (e.g., a game console co-located with the controller) coupled to the controller 110 instead of on the remote server 101 to reduce latency. In such an embodiment, the graphics processing unit (GPU) of the game console can be used to improve the speed of predicting finger gestures.
[0041] The various embodiments discussed herein teach a multimodal finger tracking mechanism in which several modal components capture multimodal data. Each of the modal components provides information regarding the detection of finger gestures to a voting engine. The voting engine then uses the modal data of each mode to assign appropriate weights to the modal data of each mode that indicate an accurate prediction of the finger gesture. For example, since the sensor tends to detect gestures more accurately than the webcam feed, a greater weight can be given to the modal data generated from the sensor. Because the system relies on modal data from more than one mode, the relative weights to the modal data based on prediction accuracy reduce the error in detecting finger gestures. The modal data processing engine 140 uses the actual button press state data, video features (i.e., images from the image capture device), audio features using the microphone array associated with the controller 110, sensor data, WiFi signals, to train a custom finger tracking model (i.e., the ML model 150) to predict finger gestures with higher accuracy than when relying on only a single data source, such as a finger tracking model of a typical camera only.
[0042] FIG. 6 shows the operation flow of a method for verifying an input provided to a controller in some embodiments. The method starts at operation 610, where in this case, a finger gesture is detected on the surface of a controller operated by a user to interact with an interactive application such as a video game. The finger gesture can be provided on any surface of the controller including input controls such as physical buttons and touch screen interactive interfaces. In response to the detection of the finger gesture, as shown in operation 620, a plurality of sensors and components are activated to capture various attributes of the finger gesture. Each sensor or component captures modal data of a specific mode, and the modal data captured by the plurality of sensors and components is collected to define multimodal data. An ensemble model is generated using a machine learning algorithm as shown in operation 630. The ensemble model is generated to include a network of interconnected nodes, and each node is input with modal data of one or more modes. The knowledge generated at each node based on the modal data contained therein is exchanged between different nodes in the network via the interconnections, and the knowledge is propagated to other nodes, thereby building the knowledge. To improve the accuracy of finger gesture prediction / verification, the modal data of each mode may be given a weight indicating an accurate prediction of the finger gesture using the modal data of each respective mode. The output is defined in an ensemble model (also referred to herein as "ML model 150") based on the cumulative weights of the various modal data collected for the finger gesture, and each output meets a certain level of prediction criteria for predicting the finger gesture. As shown in operation 640, an output is identified from the ensemble model of the finger gesture. The output is identified to meet at least the prediction criteria defined or required for the finger gesture. The output identified from the ensemble model is used to define the input regarding the finger gesture.
[0043] Figure 7 shows the components of an exemplary device 700 that can be used to implement aspects of various embodiments of the present disclosure. This block diagram shows a device 700 that can incorporate, or alternatively be, a personal computer, a video game console, a personal digital assistant, a head-mounted display (HMD), a wearable computing device, a laptop or desktop computing device, a server, or any other digital device suitable for implementing embodiments of the present disclosure. For example, in various embodiments discussed herein, device 700 represents not only a first device but also a second device. Device 700 includes a central processing unit (CPU) 702 for executing software applications and optionally an operating system. CPU 702 may be composed of one or more homogeneous or heterogeneous processing cores. For example, CPU 702 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs having a microprocessor architecture particularly adapted for highly parallel and compute-intensive applications such as query interpretation, identification of contextually relevant resources, and immediate implementation and rendering of contextually relevant resources within a video game. Device 700 may be local to a player playing a game segment (e.g., a game console), or remote from the player (e.g., a backend server processor), or one of many servers using virtualization in a game cloud system for remote streaming of game play to a client device.
[0044] Memory 704 stores the applications and data used by CPU 702. Storage 706 provides non-volatile storage and other computer-readable media for applications and data, and may include a fixed disk drive, a removable disk drive, a flash memory device, and a CD-ROM, DVD-ROM, Blu-ray (registered trademark), HD-DVD, or other optical storage device, as well as signal transmission and storage media. User input device 708 communicates user input from one or more users to device 700, and examples of user input device 708 may include a keyboard, a mouse, a joystick, a touch pad, a touch screen, a still recorder / camera or a video recorder / camera, a tracking device that recognizes gestures, and / or a microphone. Network interface 714 enables device 700 to communicate with other computer systems via an electronic communication network, and may include wired or wireless communication over a local area network and a wide area network such as the Internet. Audio processor 712 is adapted to generate an analog or digital audio output from instructions and / or data provided by CPU 702, memory 704, and / or storage 706. The components of device 700 including CPU 702, memory 704, data storage 706, user input device 708, network interface 714, and audio processor 712 are connected via one or more data buses 722.
[0045] The graphics subsystem 720 is further connected to the data bus 722 and the components of the device 700. The graphics subsystem 720 includes a graphics processing unit (GPU) 716 and a graphics memory 718. The graphics memory 718 includes a display memory (e.g., frame buffer) used to store pixel data for each pixel of the output image. The graphics memory 718 may be integrated with the same device as the GPU 716, connected as a separate device from the GPU 716, and / or implemented within the memory 704. The pixel data can be provided directly from the CPU 702 to the graphics memory 718. Alternatively, the CPU 702 provides data and / or instructions defining the desired output image to the GPU 716, from which the GPU 716 generates pixel data for one or more output images. The data and / or instructions defining the desired output image can be stored in the memory 704 and / or the graphics memory 718. In an embodiment, the GPU 716 includes a 3D rendering function for generating pixel data for the output image from instructions and data defining the geometry, lighting, shading, texturing, motion, and / or camera parameters of the scene. The GPU 716 can further include one or more programmable execution units capable of executing shader programs.
[0046] The graphics subsystem 720 periodically outputs pixel data of the image from the graphics memory 718 for display on the display device 710. The display device 710 can be any device capable of displaying visual information in response to a signal from the device 700, including CRT, LCD, plasma, and OLED displays. In addition to the display device 710, the pixel data can also be projected onto a projection surface. The device 700 can provide, for example, an analog signal or a digital signal to the display device 710.
[0047] Note that access services distributed over a wide area, such as providing access to the games of the present embodiment, often use cloud computing. Cloud computing is a computing paradigm in which dynamically scalable and often virtualized resources are provided as services over the Internet. A user does not need to be an expert in the technical infrastructure of the "cloud" that supports the user. Cloud computing can be classified into different services such as infrastructure as a service (IaaS), platform as a service (PaaS), and software as a service (SaaS). Cloud computing services often provide common applications such as video games online for access from a web browser, but the software and data are stored on servers within the cloud. The term "cloud" is used as a metaphor for the Internet based on the way the Internet is depicted in a computer network diagram and is an abstract concept that hides complex infrastructure.
[0048] In some embodiments, the game server can be used to execute the operation of a persistent information platform for video game players. Most video games played over the Internet operate via a connection to a game server. Typically, a game uses a dedicated server application that collects data from players and distributes it to other players of the video game. In other embodiments, the video game may be executed by a distributed game engine. In these embodiments, the distributed game engine may be executed on a plurality of processing entities (PEs), and as a result, each PE executes a given functional segment of the game engine on which the video game is executed. Each processing entity is regarded as merely a computing node from the game engine. Game engines typically perform a functionally diverse set of operations to execute video game applications along with additional services experienced by users. For example, a game engine implements game logic, performs game calculations, physical processes, geometry transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. Additional services may include, for example, messaging, social utilities, audio communication, game play / replay functionality, help functionality, and the like. The game engine may be executed on an operating system virtualized by a hypervisor of a particular server, but in other embodiments, the game engine itself may be distributed across multiple processing entities, and each entity may reside in a different server unit of a data center.
[0049] According to this embodiment, for the execution of operations, each processing entity may be a server unit, a virtual machine, or a container, depending on the needs of each game engine segment. For example, if a game engine segment is responsible for camera transformation, that particular game engine segment will perform a large number of relatively simple mathematical operations (e.g., matrix transformation), so it may be provisioned with a virtual machine associated with a graphics processing unit (GPU). Other game engine segments that require fewer but more complex operations may be provisioned with processing entities associated with one or more higher-powered central processing units (CPUs).
[0050] By distributing the game engine, the game engine has elastic computing characteristics that are not restricted by the capabilities of physical server units. Instead, the game engine is provisioned with more or fewer compute nodes as needed to meet the requirements of the video game. From the perspective of the video game and the video game player, a game engine distributed across multiple compute nodes is indistinguishable from a non-distributed game engine executed by a single processing entity because a game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output components to the end user.
[0051] The user accesses the remote service through a client device that includes at least a CPU, a display, and I / O. The client device may be a PC, a mobile phone, a netbook, a PDA, etc. In one embodiment, the network running on the game server recognizes the type of device used by the client and adjusts the communication method to be employed. In another case, the client device uses a standard communication method such as HTML to access the application on the game server via the Internet.
[0052] It should be understood that a given video game or game application can be developed for a specific platform and a specific associated controller device. However, when making such a game available via a game cloud system as presented herein, a user can access the video game with different controller devices. For example, a game may be developed for a game console and its associated controller, but a user can access the cloud-based version of the game from a personal computer using a keyboard and mouse. In such a scenario, the input parameter configuration can define a mapping from the inputs that can be generated by the controller devices available to the user (in this case, the keyboard and mouse) to the inputs that are acceptable for the execution of the video game.
[0053] In another example, a user can access the cloud game system via a tablet computing device, a touch screen smartphone, or other touch screen-driven devices. In this case, the client device and the controller device are integrated together within the same device, and the input is provided by the detected touch screen input / gesture. For such devices, the input parameter configuration can define specific touch screen inputs that correspond to game inputs for the video game. For example, buttons, directional pads, or other types of input elements can be displayed or overlaid during the execution of the video game, indicating positions on the touch screen that the user can touch to generate game inputs. Gestures such as swipes in a specific direction, or specific touch motions can also be detected as game inputs. In one embodiment, a tutorial showing how to input into the game play via the touch screen can be provided to the user, for example, before starting the game play of the video game, to familiarize the user with the control operations on the touch screen.
[0054] In some embodiments, the client device functions as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection and sends inputs from the controller device to the client device. Next, the client device processes these inputs and then may send the input data via a network (e.g., a network accessible via a local network device such as a router) to a cloud gaming server. However, in other embodiments, the controller itself has the ability to communicate inputs directly to the cloud gaming server via the network, and it is possible for such devices to be networked devices without the need to first communicate these inputs through the client device. For example, the controller can connect to a local network device (such as the aforementioned router) to send and receive data with the cloud gaming server. Thus, while the client device may still be required to receive video output from a cloud-based video game and render it on a local display, the input latency can be reduced by enabling the controller to send inputs directly to the cloud gaming server via the network, bypassing the client device.
[0055] In one embodiment, the networked controller and client device can be configured to send certain types of input directly from the controller to the cloud game server and other types of input via the client device. For example, apart from the controller itself, inputs detected without relying on any additional hardware or processing can bypass the client device and be sent directly from the controller to the cloud game server via the network. Such inputs can include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), etc. However, inputs that utilize additional hardware or require processing by the client device can be sent to the cloud game server by the client device. These can include video or audio captured from the game environment that can be processed by the client device before being sent to the cloud game server. Additionally, inputs from the controller's motion detection hardware are processed by the client device in conjunction with the captured video to detect the position and movement of the controller, and then communicated to the cloud game server by the client device. It should be understood that controller devices according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud game server.
[0056] In one embodiment, various technical examples can be implemented using a virtual environment via a head-mounted display (HMD). The HMD may also be referred to as a virtual reality (VR) headset. As used herein, the term "virtual reality" (VR) generally refers to a user interaction with a virtual space / virtual environment that includes viewing a virtual space through an HMD (or VR headset) in real time to provide the user with a sense of being in a virtual space or metaverse, responsive to the movement (controlled by the user) of the HMD. For example, a user can view a three-dimensional (3D) display of a virtual space when facing a given direction, and when the user turns sideways and thereby changes the orientation of the HMD in the same way, the display of that side of the virtual space is rendered on the HMD. The HMD can be worn in the same way as glasses, goggles, or a helmet and is configured to display video games or other metaverse content to the user. By providing a display mechanism in close proximity to the user's eyes, the HMD can provide the user with a very immersive experience. Thus, the HMD can provide each of the user's eyes with a display area that occupies most, or even all, of the user's field of view, and can also provide viewing with three-dimensional depth and perspective.
[0057] In one embodiment, the HMD can include an eye-tracking camera configured to capture an image of the user's eyes while the user is interacting with a VR scene. The eye-tracking information captured by the eye-tracking camera(s) can include information related to the user's line of sight direction and specific virtual objects and content items within the VR scene that the user is looking at or interested in interacting with. Thus, based on the user's line of sight direction, the system can detect specific virtual objects and content items, such as game characters, game objects, game items, etc., that may potentially be a focus for the user if the user is interacting with and interested in engaging with them.
[0058] In some embodiments, the HMD may include outward-facing camera(s) configured to capture images of the user's real-world space, such as the user's body movements, and images of any real-world objects that may be placed in the real-world space. In some embodiments, the images captured by the outward-facing camera can be analyzed to determine the position / orientation of real-world objects relative to the HMD. Using the known position / orientation of the HMD, real-world objects, and inertial sensor data from an inertial motion unit (IMU), the user's gestures and movements can be continuously monitored and tracked while the user is interacting with a VR scene. For example, while interacting with a scene within a game, the user may perform various gestures such as pointing at a specific content item within the scene or walking in its direction. In one embodiment, the system may track and process the gestures to generate a prediction of an interaction with a specific content item within the game scene. In some embodiments, machine learning may be used to facilitate or assist in the above prediction.
[0059] During HMD use, various types of single-handed controllers and two-handed controllers may be used. In certain embodiments, the controller itself can be tracked by tracking the lights included in the controller or by tracking the shape, sensors, and inertial data associated with the controller. By using these various controllers or even simple hand gestures that are executed and captured by one or more cameras, it becomes possible to interface with, control, operate, interact with, and participate in a virtual reality environment or metaverse rendered on the HMD. In some cases, the HMD can be wirelessly connected via a network to cloud computing and game systems. In an embodiment, the cloud computing and game system maintains and executes a video game that a user is playing. In some embodiments, the cloud computing and game system is configured to receive inputs from an HMD and interface object via a network. The cloud computing and game system is configured to process the inputs to affect the game state of the running video game. Outputs such as video data, audio data, and tactile feedback data from the running video game are transmitted to the HMD and interface object. In other embodiments, the HMD can communicate wirelessly with the cloud computing and game system via an alternative mechanism or channel such as a cellular network.
[0060] Furthermore, embodiments of the present disclosure may be described with respect to a head-mounted display, but in other embodiments, a portable device screen (e.g., a tablet, smartphone, laptop, etc.) or any other type of display, including but not limited to, a non-head-mounted display, may be used instead that is configured to render video and / or provide a display of an interactive scene or virtual environment according to this embodiment. It will be understood that, of course, the various embodiments defined herein may be combined or assembled in a particular implementation using the various features disclosed herein. Thus, the examples provided are only a part of the possible examples and do not limit the various embodiments that can define more embodiments by combining various elements. In some examples, some embodiments may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
[0061] As described above, embodiments of the present disclosure for communication between computing devices can be practiced using a variety of computer device configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable household appliances, minicomputers, mainframe computers, head-mounted displays, wearable computing devices, and the like. Embodiments of the present disclosure can also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked via a wired or wireless network.
[0062] In some embodiments, communication may be facilitated using wireless technologies. Such technologies may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network where the service area covered by a provider is divided into small geographical areas called cells. Analog signals representing sound and images are digitized by a telephone, converted by an analog-to-digital converter, and transmitted as a bitstream. All 5G wireless devices within a cell communicate with electromagnetic waves via a frequency channel assigned by a transceiver from a frequency pool that is reused in other cells, to a local antenna array and a low-power automatic transceiver (transmitter and receiver) within the cell. The local antenna is connected to the telephone network and the Internet by a high-bandwidth optical fiber or wireless backhaul connection. Similar to other cellular networks, a mobile device moving from one cell to another is automatically transferred to the new cell. Of course, the 5G network is simply an example of a type of communication network, and in embodiments of the present disclosure, previous generation wireless or wired communication, as well as subsequent generation wired or wireless technologies coming after 5G, may be used.
[0063] Considering the above embodiments, it is natural that the present disclosure can adopt various computer-implemented operations involving data stored in a computer system. These operations are operations that require physical manipulation of physical quantities. Any of the operations described herein that form part of the present disclosure are useful machine operations. The present disclosure also relates to devices or apparatuses for performing these operations. The apparatus can be specially configured for the required purpose, or it can be a general-purpose computer selectively activated or configured by a computer program stored in a computer. Specifically, various general-purpose machines can be used together with a computer program written according to the teachings herein. Or, it may be more convenient to construct a more specialized apparatus for performing the required operations.
[0064] Although the method operations have been described in a specific order, other housekeeping operations may be performed between the operations, or the operations may be adjusted to occur at slightly different times, or the operations may be distributed within the system to allow the processing operations to occur at various processing-related intervals, as long as the telemetry and game state data processing for generating the modified game state is performed in the desired manner. It should be understood.
[0065] One or more embodiments can also be made as computer-readable code on a computer-readable medium. The computer-readable medium can be any data storage device capable of storing data. The data can then be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. The computer-readable medium can include computer-readable tangible media distributed across a network-connected computer system so that the computer-readable code is stored and executed in a distributed manner.
[0066] Although the foregoing embodiments have been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications can be made within the scope of the appended claims. Accordingly, the present embodiments should be regarded as illustrative rather than restrictive, and the present embodiments should not be limited to the details described herein, but may be modified within the scope of the appended claims and equivalents thereof.
[0067] It should be understood that the various embodiments defined herein may be combined or assembled into specific embodiments using the various features disclosed herein. Accordingly, the examples provided are only a part of the possible examples and do not limit the various embodiments that can define more embodiments by combining various elements. In some examples, some embodiments may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
Claims
1. A method for verifying an input provided to a controller, comprising: Detecting a finger gesture provided by a user on a surface of the controller, the finger gesture being used to define the input of an interactive application selected for interaction by the user; Receiving multimodal data collected by tracking the finger gesture on the controller, the multimodal data including modal data corresponding to different modes captured for the finger gesture using a plurality of sensors and components associated with the controller; Generating an ensemble model using the multimodal data received from the plurality of sensors and components, the ensemble model being continuously trained using additional multimodal data collected over time to generate different outputs, the ensemble model being trained according to training rules defined for different finger gestures using a machine learning algorithm; Identifying an output from the ensemble model for the finger gesture detected on the surface of the controller, the output being interpreted to define the input of the interactive application based on the finger gesture detected on the controller; A method executed by a processor of a server.
2. In the identification of the output, Assigning weights to the modal data included in the multimodal data captured by each of the plurality of sensors and components, the weights assigned to each modal data captured by a sensor or component among the plurality of sensors and components indicating an accurate prediction of the finger gesture using each modal data; Using the weights assigned to the respective modal data of the modes included in the multimodal data, calculate the cumulative weight of the multimodal data, the cumulative weight being used when identifying the output for the finger gesture, the method according to claim 1.
3. The weights assigned to the modal data captured by each of the plurality of sensors and components are greater than the weights assigned to the modal data captured by each of the plurality of sensors and components, The weights assigned to the modal data captured by the plurality of sensors are equal, and the weights assigned to the modal data captured by the plurality of components are equal, the method according to claim 2.
4. Equal weights are assigned to the modal data captured for different modes included in the multimodal data, the method according to claim 2.
5. Separate weights are assigned to the modal data captured for each mode included in the multimodal data, the method according to claim 2.
6. The ensemble model is trained according to training rules defined for the different finger gestures, the training rules being defined based on the anatomical form of the finger, the position of the finger related to the input control on the controller, and the controller holding style of the user, the method according to claim 1.
7. The multimodal data includes video data, audio data, image data, sensor data, and wireless signals collected from the plurality of sensors and components, In the identification of the output, based on the finger gesture, the pressing of an actual button or a virtual button on the controller or the identification of an input provided on a touch screen interface is performed, and the output is interpreted to define the input to the interactive application, the method according to claim 1.
8. The plurality of sensors includes any one or a combination of an inertial measurement unit (IMU) sensor, a pressure sensor, a proximity sensor, a distance sensor, or a capacitance sensor, The plurality of components includes an image capture device, a wireless communication device, or a microphone array, the method according to claim 1.
9. The multimodal data includes a WiFi signal having a forward signal and a reflected signal captured by the one or more wireless communication devices, and the forward signal and the reflected signal are interpreted to define a snapshot of the user's body part, and the snapshot of the body part is used when reconstructing the movement of one or more fingers of the user when the user provides the finger gesture, the method according to claim 1.
10. The plurality of sensors and components includes an image capture device, The multimodal data includes images of different positions held by the user's finger captured by the image capture device when the user provides the finger gesture, and the angle of the image capture device dynamically adjusted to capture images of different positions of the finger, and the dynamic adjustment is performed by automatically calibrating the angle of the image capture device in response to the detection of the presence of the finger and the finger gesture provided by the user on the controller, the method according to claim 1.
11. The image capture device is a camera integrated into a mobile computing device, or a web camera, or an image capture device of a game console, or an image capture device of a computing device or a camera of a head-mounted display, and the image capture device is communicably coupled to the controller, the method according to claim 10.
12. When the image capture device is the camera of the mobile computing device, the mobile computing device is disposed on a holding structure coupled to the controller, the holding structure includes a motor, the motor receives and holds the mobile computing device, dynamically adjusts the angle of the camera, and is calibrated to align with the angle such that images at different positions held by the user's finger when the user performs a finger gesture can be captured. The holding structure is a three-dimensional printed structure, the method according to claim 11.
13. The plurality of sensors and components includes one or more of inertial measurement unit sensors (IMUs), the IMUs are configured to detect a user's finger gesture on the surface of the controller and generate IMU signals. The multimodal data includes the IMU signals received from the one or more IMUs, the IMU signals are interpreted to identify attributes of the finger gesture, and the attributes are used to identify user input on the controller, the method according to claim 1.
14. The plurality of sensors and components includes a microphone array embedded in or coupled to the controller. The finger gesture is a button press on the controller. The multimodal data includes attributes of audio data captured by the microphone array, and the attributes of the audio data captured by a plurality of microphones in the microphone array are used to identify the direction and position of each microphone in the microphone array using triangulation technology, and the direction and the position are interpreted to determine the button press, the method according to claim 1.
15. A method for defining an input of an interactive application, receiving a finger gesture provided by a user on a surface of a controller, the finger gesture being used to define the input of the interactive application selected by the user for interaction, receiving multimodal data that captures attributes of the finger gesture on the controller, the multimodal data including modal data corresponding to different modes captured by a plurality of sensors and components associated with the controller, assigning weights to the modal data corresponding to each mode included in the multimodal data, the weights assigned to each mode indicating an accurate prediction of the finger gesture using the modal data of each mode, performing processing of the finger gesture and the multimodal data based on the weights assigned to each mode to identify the input of the interactive application, A method executed by a processor of a server.
16. In the processing of the finger gesture and the multimodal data, generating an ensemble model using the multimodal data received from the plurality of sensors and components, Continuously train the ensemble model using additional multimodal data collected over time to generate different outputs, the ensemble model being trained using a machine learning algorithm according to training rules defined for different finger gestures. Identify the output from the ensemble model corresponding to the finger gesture detected on the surface of the controller, the output being interpreted to define the input to the interactive application, the method of claim 15. **Claim 17** The training rules are defined based on the anatomical form of the finger, the position of the finger relative to the input controls on the controller, and the controller holding style of the user, the method of claim 16. **Claim 18** The processing of the finger gesture includes calculating a cumulative weight of the multimodal data using the weights assigned to the modal data of each mode included in the multimodal data, the cumulative weight being used when identifying the input to the interactive application, the method of claim 15. **Claim 19** The plurality of sensors includes any one or a combination of an inertial measurement unit (IMU) sensor, a distance sensor, a pressure sensor, a proximity sensor, or a capacitance sensor. The plurality of components includes any one or a combination of an image capture device, a wired communication device, a wireless communication device, or a microphone array, the method of claim 15. **Claim 20** The multimodal data includes a first set of modal data captured by the plurality of sensors and a second set of modal data captured by the plurality of components. The method according to claim 15, wherein in the weight assignment, a first weight is assigned to the modal data of each mode included in the first set of the modal data, and a second weight is assigned to the modal data of each mode included in the second set of the modal data, and the first weight is greater than the second weight.
Citation Information
Patent Citations
Gesture cataloguing and recognition
JP2009165826A
Multi-surface Controller
JP2017535002A
Handheld controller with touch-sensitive controls
JP2021518612A
Projecting and receiving input from one or more input interfaces attached to a display device
US10824239B1
Motion-Input Device For a Computing Terminal and Method of its Operation
US20080174550A1