Pinch state detection systems and methods

The wrist-wearable device addresses the limitations of traditional controllers by using sensors and neural network classifiers to directly digitize hand movements, enhancing the responsiveness and intuitiveness of human-machine interaction in XR applications.

JP2025087611APending Publication Date: 2025-06-10DOUBLEPOINT TECH OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024202057
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing wearable devices for controlling computing and extended reality user interface applications are limited by the need for separate physical controllers, which introduce an extra technological layer slowing down human-machine interaction and are often intrusive, limiting the user's hand usage.

Method used

A wrist-wearable device equipped with a processing core, memory, and sensors (optical and IMU) that directly digitizes fine hand movements and gestures, converting them into machine commands using a neural network classifier for accurate gesture detection.

Benefits of technology

Enables seamless and intuitive control of devices and applications by directly translating hand movements into machine commands, improving responsiveness and reducing the intrusiveness of traditional controllers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087611000001_ABST
    Figure 2025087611000001_ABST
Patent Text Reader

Abstract

To provide an input device and a corresponding method for digitizing and transforming minute hand movements and gestures into user interface commands without interfering with the normal use of user's hands.SOLUTION: The device and method may, for example, determine a gesture state by detecting and classifying gesture transitions based on transient events within sensor data of received data streams from a plurality of sensors, where the sensors are preferably of different types.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various exemplary embodiments of the present disclosure relate to wearable devices and methods that can be used to control a device, and in particular, can be used to control a device in the field of computing and extended reality user interface applications. Extended reality (XR) includes augmented reality (AR), virtual reality (VR), and mixed reality (MR).

Background Art

[0002] Digital devices have traditionally been controlled by dedicated physical controllers. For example, a computer can be operated with a keyboard and a mouse, a game console can be operated with a handheld controller, and a smartphone can be operated with a touch screen. Typically, such physical controllers include sensors and / or buttons that receive input from a user based on the user's actions. Although such separate controllers are widespread, they slow down human-machine interaction by adding an extra technological layer between the user's hand and the computing device. In addition, such dedicated devices are typically only suitable for controlling a specific device. Also, such devices can be intrusive to the user, for example, the user cannot use their hand for other purposes while using the controller.

Summary of the Invention

Problems to be Solved by the Invention

[0003] The present invention has been made to solve the problems in the above prior art.

Means for Solving the Problems

[0004] In the field of XR user interfaces, there is a need for improvements to the above-mentioned problems. A suitable input device, such as the invention disclosed herein, directly digitizes fine hand movements and gestures and converts them into machine commands without interfering with the normal use of a person's hand. Embodiments of the present disclosure are capable of detecting user actions, for example, by identifying gesture states based on sensor data streams received from at least one sensor, preferably a plurality of sensors, and it is preferred that there are various types of sensors.

[0005] The present invention is defined by the features recited in the independent claims. Some specific embodiments are defined in the dependent claims.

[0006] According to a first aspect of the present invention, there is provided a wrist wearable device including a processing core and at least one memory including computer program code, the at least one memory and the computer program code being configured to, together with the at least one processing core, at least receive data streams from at least one optical sensor and at least one IMU, the sensors being configured to perform measurements of a user, the receiving; pass the received data streams to a gesture classifier; use the gesture classifier to identify a gesture state by detecting and classifying gesture transitions based on transient events in the sensor data within the received data streams, the gesture classifier including a neural network classifier trained to detect the gesture transitions, the identifying; and the data streams may optionally be pre-processed before being passed to the gesture classifier.

[0007] According to a second aspect of the present invention, there is provided a method for identifying a selection gesture from acquired data, the method comprising: receiving a data stream from at least one optical sensor and at least one IMU, wherein these sensors are configured to perform measurements of a user; passing the received data stream to a gesture classifier; using the gesture classifier to identify a gesture state by detecting and classifying a gesture transition based on transient events in the sensor data within the received data stream, wherein the gesture classifier includes a neural network classifier trained to detect the gesture transition; and the data stream may optionally be preprocessed before being passed to the gesture classifier.

[0008] According to a third aspect of the present invention, there is provided a computer program which, when executed by at least one processor, causes the apparatus to at least: receive a data stream from at least one optical sensor and at least one IMU, wherein these sensors are configured to perform measurements of a user; pass the received data stream to a gesture classifier; use the gesture classifier to identify a gesture state by detecting and classifying a gesture transition based on transient events in the sensor data within the received data stream, wherein the gesture classifier includes a neural network classifier trained to detect the gesture transition; and the data stream may optionally be preprocessed before being passed to the gesture classifier.

[0009] According to a fourth aspect of the present invention, there is provided a non-transitory computer-readable medium storing a set of computer-readable instructions which, when executed on a processor, cause the second aspect to be implemented or cause an apparatus including the processor to be configured according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0010]

Fig. 1A

Fig. 1B

Fig. 2

Fig. 3

Fig. 4

Fig. 5

Fig. 6

Fig. 7

[0011] This specification describes a multimodal biometric measurement device and a corresponding method. Multimodal biometric measurement is understood to be the acquisition, processing, and / or analysis of multiple pieces of measurement information related to biometric measurement, in other words, measurements related to a user's body and body actions achieved using multiple modes (e.g., sensors). The device and / or method may be used for at least any one of the tasks of (e.g.) measurement, detection, signal acquisition, analysis, and user interface. Such a device is preferably suitable for a user to wear. The user can control one or more devices using this device. The control may be in the form of a user interface (UI) or a human-machine interface (HMI). The device may include one or more sensors. The device may be configured to generate user interface data based on data from its one or more sensors. The user interface data may be used to enable the user to control this device or at least one secondary device to some extent. The secondary device may be, for example, at least one of a personal computer (PC), a server, a mobile phone, a smartphone, a tablet device, a smartwatch, or any suitable type of electronic device. The device to be controlled or controllable may implement at least any one of an application, a game, and / or an operating system, all of which may be controlled by the multimodal device.

[0012] When interacting with AR / VR / MR / XR applications, for example, when using devices such as smartwatches or extended reality headsets, it is necessary for the user to be able to perform various actions (e.g., selection, drag and drop, rotate and drop, slider use, and zoom), also known as user actions. This embodiment improves the detection of user actions, which can lead to an improvement in responsiveness when implemented in a controller or system, or by a controller or system.

[0013] Statelessness is a well-known term in the fields of computing and user interfaces. It essentially means that the state or value of user input and user output is retained in the system memory. Past actions or events not only monitor transitions but also affect the system to provide various contexts and / or responses to current or future actions. Statelessness is a very important aspect of human-computer interaction devices as it enables dense interactions compared to instantaneous input. Interactions such as drag and drop, pinch to zoom, etc. require statelessness.

[0014] The user may perform at least one user action. Typically, the user performs an action to affect a controllable device and / or an AR / VR / MR / XR application. The user action may include, for example, at least one of movement, gesture, interaction with an object, interaction with a body part of the user, and null action. An example of a user action is the "pinch" gesture, which is a gesture where the user touches the tip of the index finger and the tip of the thumb. Another example is the "thumbs up" gesture, which is a gesture where the user extends the thumb, curls the other fingers, and rotates the hand so that the thumb points upward. This embodiment is configured to identify (specify) at least one user action by specifying the gesture state based at least to some extent on sensor input. By reliably identifying the user action, action-based control (e.g., using the embodiments disclosed herein) becomes possible.

[0015] The user action may include at least one state transition. For example, the user action may be pinching. That is, it may be a transition from an unpinned state to a pinned state. In this embodiment, the gesture transition action (i.e., the state transition action) is specified (e.g., by a neural network) based at least to some extent on sensor data.

[0016] In at least some embodiments, user interface (UI) commands are generated by a device (e.g., device 100) based on a specified gesture state. The UI commands may be referred to, for example, as human machine interface (HMI) commands, user interface interactions, and / or HMI interactions. Such commands can be used to convey user interface actions. For example, a user interface action may correspond to the "enter" key on a personal computer. In at least some embodiments, an output generator, such as a user interface command generator, maps a user action to a predetermined user interface action and / or generates a user interface command based on the predetermined user interface action. Thus, the device to be controlled does not need to understand the user action, and only reports a descriptor for a UI command that can be provided in a standard format such as a human interface device (HID). The UI commands may include, for example, data packets and / or frames including UI instructions.

[0017] In at least some embodiments, the device includes a mounting component configured to be worn by a user. Thus, the mounting component is a wearable component. Such a component may be any of a strap, band, wristband, bracelet, glove. The mounting component may be attached to another device such as a smartwatch and / or formed by such a device, or may form part of a larger device such as a gauntlet. In some embodiments, the strap, band, and / or wristband has a width of 2 to 5 cm, preferably 3 to 4 cm, and more preferably 3 cm.

[0018] In this embodiment, a signal is measured by at least one sensor. The sensor may include a processor and a memory, or may be connected to a processor and a memory, and the processor may be configured to perform the measurement. The measurement may be referred to as detection or sensing. The measurement includes, for example, detecting a change in at least one of a user, an environment, and the physical world. The measurement may further include, for example, attaching at least one timestamp to the sensed data and transmitting the sensed data (regardless of the presence or absence of the at least one timestamp).

[0019] The signal data may include one or more transient events. A transient event may be any change in the signal. In particular, the transient event may include either a rising edge of the signal or a falling edge of the signal.

[0020] At least some embodiments are configured to use an inertial measurement unit (IMU) to measure user characteristics such as gestures, position, and / or movement. The IMU may be configured to provide position information of the device. The IMU may be referred to as an inertial measurement unit sensor. The IMU may include at least any one of a multi-axis accelerometer, a multi-axis gyroscope, a multi-axis magnetometer, an altimeter, and a barometer. It is preferable that the IMU includes a magnetometer because the magnetometer provides an absolute reference for the IMU. The barometer can be used as an altimeter and can give the IMU additional degrees of freedom. The IMU signal data may include transient events.

[0021] At least some embodiments are configured to use an optical sensor to measure user characteristics such as gestures, position, and / or movement. The optical sensor may include, for example, a photoplethysmography (PPG) sensor. The optical sensor (e.g., a PPG sensor) may be referred to as a biometric sensor and is capable of measuring at least one of, for example, heart rate, oxygen saturation, or other biometric data. Typically, a light-emitting diode (LED) is used to illuminate the underlying tissue, and the luminance changes that occur during the cardiac cycle are measured. The optical sensor can use both visible light and infrared light to perform these functions. The optical signal data may include transient events.

[0022] When used for gesture classification, it is possible to minimize motion artifacts in the optical sensor data, thereby achieving a more accurate reading. This minimization can be achieved, for example, by using a low-pass filter or a median filter for high-speed motion artifacts and a high-pass filter for low-speed motion artifacts. The optical sensor data may be used for other purposes in addition to measuring heart rate and oxygen saturation. For example, a PPG sensor may be used to measure the amount of reflected light. The amount of reflected light is affected by the proximity of the sensor to the user's skin and the reflectivity of the underlying tissue. The proximity is affected, for example, by the movement of tendons.

[0023] At least some embodiments are configured to measure vibration signals and / or acoustic signals. Such signals may be measured, for example, from a user, particularly from the user's wrist area. A bioacoustic sensor may be configured to measure vibrations such as internal waves of the user and / or surface waves of the user's skin. Such vibrations may be, for example, the result of a contact event. The contact event may be the user contacting an object and / or the user contacting himself and / or another user. Examples of bioacoustic sensors may include contact microphone sensors (piezo contact microphones), MEMS (microelectromechanical systems) microphones, and / or vibration sensors. The vibration sensor may be, for example, an accelerometer that detects the acceleration generated by a vibrating structure (such as a user or a mounting component).

[0024] Sensors (such as inertial measurement unit sensors, bioacoustic sensors, and / or optical sensors) may be configured to output a sensor data stream. This output may be within the device, for example, to a controller, a digital signal processor (DSP), or a memory. Alternatively or additionally, this output may be to an external device. The sensor data stream may include one or more raw signals measured by the sensor. Additionally, the sensor data stream may include at least one of synchronization data, configuration data, and / or identification data. Such data may be used, for example, by a controller to compare data from various sensors.

[0025] In at least some of the embodiments, the apparatus includes means for providing kinesthetic feedback and / or tactile feedback. Such feedback may include, for example, vibration and may also be referred to as "haptic feedback". The means for providing haptic feedback may include, for example, at least any one of an actuator, a motor, a haptic output device. A preferred actuator that can be used in all embodiments is a linear resonant actuator (LRA). Haptic feedback can be useful for improving the usability of the user interface by providing feedback to the user from the user interface. Further, haptic feedback can alert the user to events and context changes in the user interface.

[0026] Figures 1A and 1B show an apparatus 100 according to at least some embodiments of the present invention. The apparatus 100 may be suitable for detecting a user's movement and / or action and for operating as a user interface device and / or an HMI device. The apparatus 100 may include at least one optical sensor, which may be arranged to face the user's hand. The apparatus 100 may include at least one inertial measurement unit 104 within a housing 102. Sensor data is acquired by an on-board controller. The sensor data may be preprocessed by the controller. The controller may be configured to detect a gesture state and thus a user action (e.g., whether the hand is in a pinch state) using, for example, a classifier or a classifier algorithm (e.g., a neural network classifier).

[0027] In at least some embodiments, device 100 may include a smartwatch. Device 100 preferably includes a screen. This screen may be used to present gesture state information to the user. In at least some embodiments, device 100 may include at least one of the following components, for example, a haptic device, a screen, a touch screen, a speaker, a heart rate sensor (e.g., an optical sensor), a Bluetooth communication device.

[0028] Device 100 includes a mounting component 101, which is shown in the form of a strap in FIG. 1A. Device 100 includes a housing 102 attached to the mounting component 101. User 20 may wear device 100, for example, on the wrist such that sensor 104 faces the user's wrist area (especially the skin of the wrist area).

[0029] Device 100 further includes a controller 103 within housing 102. Controller 103 includes at least one processor, at least one memory, and a communication interface. The memory may include instructions that, when executed by the processor, enable communication with other devices (e.g., a head device and / or a computing device) via the communication interface. The controller may be configured to receive at least parallel sensor data streams from sensors 104 and 105. The controller may be configured to perform preprocessing on at least any of those sensor data streams, which may include, for example, at least any of data cleansing, feature transformation, feature engineering, and feature selection. The controller may be configured to process the sensor data streams received from the sensors. The controller may be configured to generate at least one user interface (UI) event and / or command based on the characteristics of the processed sensor data streams. Further, the controller may include a model, and the model may include at least one neural network.

[0030] Device 100 further includes an inertial measurement unit (IMU) 104. The IMU 104 may be within the housing 102. The IMU 104 is configured to output a sensor data stream (e.g., to a controller). The sensor data stream includes, for example, at least one of multi-axis accelerometer data, gyroscope data, and / or magnetometer data. These three types of data represent acceleration, angular velocity, and magnetic flux density, respectively. One or more of the components of device 100 may be combined together. For example, the IMU 104 and the processor of the controller may be on the same printed circuit board (PCB). Changes in the IMU data reflect, for example, the movement and / or actions of the user, and thus such movement and / or actions can be detected using the IMU data.

[0031] In at least some embodiments, an inertial measurement unit (IMU) 104, or a sensor (such as a gyroscope) or multiple sensors, housed in a single device or multiple devices, is installed and / or fixed on the user's arm. The fixing may be performed, for example, by the mounting component 101.

[0032] Figure 1B shows the inside of the device 100 configured to be worn against the user's skin (e.g., the wrist). An optical sensor 105 is shown on one surface of the housing 102. The optical sensor 105 may include a PPG sensor. The optical sensor 105 may include at least one LED. The optical sensor 105 is configured to output a sensor data stream (e.g., to a controller). The sensor data stream includes, for example, at least either heart rate data or analog-to-digital (ADC) data corresponding to the amount of reflected light. One or more of the components of the device 100 may be combined together. For example, the optical sensor 105 and the processor of the controller may be connected on the same PCB. Changes in the optical sensor data reflect, for example, the movement and / or action by the user. Therefore, such movement and / or action can be detected using the optical sensor data.

[0033] Figure 2 shows exemplary sensor data corresponding to user actions. In the lower part of the figure, a sequence of the state of the user's hand is shown. First, the user's hand is in a non-pinch state (also called an unpinch state). Next, the user forms a pinch with the thumb and index finger. Generally, these two fingers are used for the pinch, but other fingers may be used as well. Finally, the user releases the pinch and the hand returns to the non-pinch state.

[0034] The upper part of the figure shows the corresponding sensor data that can be measured (e.g., using the device 100) when the user performs a sequence of hand poses corresponding to the gesture interface state. In the figure, the transition from the unpinch state to the pinch state and the transition from the pinch state to the unpinch state are marked. Such transitions may be called gesture transitions. As can be seen in the figure, several transient events occur during the gesture transition in which each signal changes.

[0035] The data stream may include one or more signals. In the figure, signals 201A and 201B correspond to sensor data measured and received by an optical sensor. The optical sensor can output signals corresponding to the detected visible light data and the detected infrared light data. Among these, infrared data is preferred. Signals 202A, 202B, and 202C correspond to signals received from a gyroscope that may be part of an IMU. Since the gyroscope is three-axis, three signals are received from the gyroscope. Signals 203A and 203B correspond to signals received from an accelerometer that may be part of an IMU. Since the accelerometer is three-axis, three signals are received from the accelerometer. These sensors can measure data related to the user's hand, for example, the user's wrist area. The illustrated sensor signals are related to the posture of the user's hand shown in the lower part of Figure 2.

[0036] In the exemplary situation of Figure 2, the user and the device are initially in a "pinchless state". When a transition that matches a pre-learned definition of a pinch-pinch transition is detected, the device 100 enters the "pinch state". In the gesture domain, pinch and unpinch are opposite actions to each other, but the signal patterns associated with each of these gestures are not similar and are asymmetric. For example, a pinch (tap) is characterized by a distinct vibratory mechanical pattern of the wearable device, while an unpinch is less distinct in the vibratory mechanical domain. Both pinch and unpinch are characterized by rising and falling edges in the optical domain, but due to the characteristics inherent in optical sensing, these edges are not distinct enough to be used for inference alone. Therefore, it is necessary to combine an inertial sensor and an optical sensor. Furthermore, IMU data provides only transition information, while optical data can provide at least some stateful information along with the transition information. The model according to the present disclosure may be configured to utilize both types of information (e.g., by machine learning).

[0037] For detecting transitions, relative values and / or relative thresholds may be used instead of absolute values and / or absolute thresholds. This is because the change in the signal of the sensor data can be very small (i.e., "local peaks and valleys"). Such local peaks and valleys are not so distinct in the overall signal, so absolute values are not as useful as relative values for detecting transitions. Device 100 can perform adaptive offset removal using the AC component of the input signal. Adaptive offset removal enables the detection of local rising edges and falling edges. Device 100 is configured such that the sensitivity of pinch transition detection is adjusted (e.g., when the device is in a pinch state). Sensitivity may mean adjusting a time window or relative values. Sensitivity adjustment may also include the analysis of the hand trajectory. Generally, hand movement in a point-and-select task is characterized by a ballistic phase and a correction phase. A pinch gesture or an unpinch gesture is more likely to occur in the correction phase than in the ballistic phase. The state of the gesture (pinch or unpinch) may be either state in the ballistic phase. Therefore, hysteresis information of the state prior to the trajectory phase may be incorporated into the adaptive sensitivity.

[0038] In an embodiment, the value of the optical sensor is higher at the start and end of the gesture compared to the midpoint of the gesture session. That is, the start and end of the gesture may be more emphasized so that local peaks and valleys are detected.

[0039] As shown in the figure, when a pinch-unpinch state transition is detected, the device enters the unpinch state. Further, device 100 may measure the transition time from the unpinch state to the pinch state for each individual user. The device may ignore all transitions that occur immediately after the previous transition or unreasonable transitions. For example, a user cannot transition from the unpinch state to the unpinch state. This can be considered by the gesture classifier and / or the interaction / state interpreter.

[0040] In an embodiment, an adaptation time window is used as part of gesture classification. The adaptation time window may be specifically adjusted according to a typical gesture duration (e.g., based on training data). The adaptation time window may be, for example, 1 millisecond when the adaptation range is 1 to 50%. The data used for classification and / or interpretation is stored in the controller of the present device. User-specific data acquired (e.g., in a single session) may be retrained or discarded after inference. Retraining has the advantage of further adapting to the user, while discarding saves memory resources.

[0041] Figure 3 shows an exemplary measurement range of the device disclosed in this specification. In the figure, a measurement range 210 related to the hand of user 20 is shown. Sensors (e.g., sensors 104, 105) can measure signals from the user or at least signals related to the user within range 210. An IMU may be disposed within or near range 210. Preferably, the sensors and the IMU are arranged such that housing 102 is attached to mounting component 101, and mounting component 101 is, for example, worn by the user so as to perform measurements of the user.

[0042] Figure 4 shows an exemplary user interface (UI) interactive. Such an interactive is displayed to the user and is, for example, displayed by head device 21 as part of a VR, AR, MR, or XR application. Head device 21 may include, for example, at least one of a head-mounted device (HMD), a headset, a virtual reality (VR) headset, an extended reality (XR) headset, and an augmented reality (AR) headset. Head device 21 may include a processor, a memory, and a communication interface, and the memory may include instructions that, when executed by the processor, enable communication with other devices (e.g., device 100) via the communication interface.

[0043] The devices within the present disclosure (e.g., device 100) may be configured to communicate with at least one computing device. The computing device may include at least one of a computer, a server, a mobile device (e.g., a smartphone or a tablet). The above device may be configured to at least participate (e.g., computationally) in providing a VR / XR / AR experience to a user (e.g., by causing a scene to be displayed on an HMD). The computing device may include a processor, a memory, and a communication interface. The memory may include instructions that, when executed by the processor, enable communication with other devices (e.g., device 100 or a head device) via the communication interface. The computing device may be the head device 21, as described above.

[0044] Interactables 31, 32, 33, and 34 represent various types of interactables with which user 20 can interact with a VR, AR, or XR application (also referred to as a scene). User interactions are detected by device 100, and the scene displayed to user 20 by device 21 is appropriately updated according to the situation (e.g., by a computer or mobile device connected to device 100 and device 21).

[0045] The interactive 31 is a button interactive, and the user can perform an action by pressing this button. This button requires a one-axis movement, similar to a physical button. The interactive 32 is a slidable interactive, and the user can adjust a value, for example, by moving a slider along a given axis. The interactive 33 is a scrollable interactive, and the user can move a view or scroll through content, for example, by moving a slider and / or the screen along at least one axis. The interactive 34 is a pan interactive, and the user can move a view by moving a slider or its view along a given axis. As a further type of interactive, there is a "grabbable" interactive, i.e., any interactive that can be grabbed by the user's hand. The interactive is not limited to a single type. For example, a slider can be grabbable.

[0046] The XR application is configured to provide a display of the interactive to, for example, the device 100. This display may include at least one of the position, type, size, orientation, and state, which are attributes of the interactive. The device 100 may be configured to receive the display and adjust internal processing accordingly (e.g., by affordance).

[0047] In the field of user interfaces, affordances are sensory cues (generally visible elements) that provide users with information about how to interact with specific elements. These cues include geometric shapes, symbols (e.g., arrows), etc. Common affordances include a scrollable scroll bar and a pressable button. If the application can predict what kind of actions the user is likely to perform next, (e.g., an interaction interpreter) can use an affordance profile to improve the accuracy of action detection. For example, usually, when there is a pinch, there will be an unpinch sooner or later. An affordance profile may include, for example, the user actions expected for a certain interactable. The prediction may be performed, for example, by a neural network executed by a controller of the device.

[0048] In an embodiment, the device 100 may be configured to obtain or receive data from at least one optical sensor 105 and at least one IMU 104 and pass the obtained / received data to a gesture classifier, and the gesture classifier may be configured to determine whether the hand is in a gesture state (e.g., a pinch state). This determination may be made by detecting and classifying gesture transitions based on transient events in the obtained / received data. The device 100 may be configured to perform preprocessing on one or more data (e.g., one or more data streams) and then pass those data to the classifier. The device 100 may be configured to perform the detection using a neural network trained to detect transitions between gesture states by detecting the rising edge and falling edge of the gesture.

[0049] The behavior of gesture prediction (e.g., threshold processing and time limits based on rules) may be modified according to known a priori constraints of interactable and scene logic. Such constraints may be provided to (e.g., by a scene such as scene 351) and stored within device 100. The constraints may be used as part of gesture detection. For example, a time limit may be used that a reverse gesture cannot occur for a predetermined time (e.g., 50 milliseconds) from the previous gesture. Such modifications may be applied based on the determined gesture state by an interaction interpreter and / or a state interpreter configured to apply the modifications based on the application state and / or application context.

[0050] The present device may be configured to output a transition as an event. In other words, the present device may output an event (e.g., "pinch on"). The present device may be configured to return the user state as "polarble". In other words, the present device may be queried, for example, "is the pinch true or false", and the appropriate state is returned in response.

[0051] In some embodiments, in combination with the embodiments disclosed herein, gaze tracking information is available and usable. The gaze tracking information (e.g., information 308) may include at least a display indicating where the user is looking. This information may include details (e.g., identifiers) regarding the interactable that the user is looking at. This information may be received, for example, from a headset such as an AR / MR / XR / VR headset. A device (e.g., device 100, device 300, and device 700) may be configured to receive gaze tracking data, whereby the device can use affordances with a matching interactable type in gesture detection.

[0052] FIG. 5 shows, in flowchart form, an exemplary process that can support at least some embodiments of the present invention. This process may be implemented, in whole or in part, by any one of apparatuses 100, 300, 700. At the top of FIG. 5, an optical sensor (e.g., similar to sensor 105) and an IMU (e.g., similar to IMU 104) are shown. Each action and each step of the flowchart may be implemented, at least in part, by an apparatus (e.g., an apparatus disclosed herein such as apparatus 100, etc.). At least some actions (e.g., the action of outputting a scene) may be implemented by another apparatus (e.g., head device 21, etc.). At least some actions (e.g., the action of providing communication) may be implemented by yet another apparatus (e.g., a mobile device including at least one processor, etc.).

[0053] The flowchart of FIG. 5 includes exemplary and non-limiting steps 901, 902, 903, 904, 905. Each action does not necessarily have to fit within one step. Further, the actions within steps 901, 902, 903, 904, and 905 may not be implemented simultaneously.

[0054] Step 901 includes data acquisition. In this step, a controller (e.g., controller 103) receives a data stream from one or both of IMU 104 and optical sensor 105.

[0055] As seen in the flowchart, after a sample (sensor data) is acquired, the acquired sensor data is preprocessed. Step 902 includes preprocessing. With respect to the IMU, the IMU may be adjusted before or after preprocessing. Such adjustment may be self-referential. For example, the sensor is adjusted or calibrated according to a stored reference value or a known good value (e.g., magnetic north in the case of an IMU).

[0056] As can be seen in the flowchart, after preprocessing, the preprocessed data is input into the gesture classifier at step 903. In the flowchart, the orientation of the IMU (and thus the orientation of the user's hand) may be directly calculated (e.g., without a neural network) using the preprocessed IMU data.

[0057] Furthermore, at step 903, the model outputs an inference result corresponding to the gesture state and thus the user action. This output, i.e., the inferred gesture state, may be provided, for example, as a confidence value.

[0058] At 904, the inferred gesture state is passed to the interaction interpreter and / or the state interpreter. This interpreter identifies the user input, and then the user input is passed to the controlled scene (application), for example, in the form of a boolean value and / or an HID (human interface device) frame. That application may be executed on another device or a network-connected device.

[0059] Optional context information reflecting the application state (scene state) may likewise be passed to the event interpreter. The interpreter can use an affordance profile according to this context information to interpret the received user action. For example, if the context information indicates that the user is interacting with the pan interactive 34, the interpreter can lower the threshold required for a pinch action in anticipation of a pinch action being performed.

[0060] In step 905, the scene is updated in response to receiving interaction data. The interpreter may be configured to output the interaction data as an event and / or as a probability, whereby the scene can query the event state. The scene may provide intention estimation data. The intention estimation may include signs of future actions and / or information related to a set of appropriate actions regarding context information. For example, the scene may provide information related to a previous action (e.g., "play a video") and an expected future action (e.g., "pause the video", "adjust the volume"). The scene may provide an affordance profile to the interpreter. As described above, this can increase the accuracy of detecting user actions when the device and / or scene can predict the next action.

[0061] FIG. 6 shows an apparatus 300 that can support at least some embodiments of the present invention. This apparatus is the same as apparatus 100 unless otherwise specified.

[0062] In FIG. 6, an exemplary schematic diagram of apparatus 300 is shown. Apparatus 300 includes an IMU 304, an optical sensor 305, and a controller 392. Apparatus 300 may be similar to apparatus 100 in that it may include mounting components so that it can be worn by a user. Controller 392 is shown in the figure represented by a dashed line. As shown, controller 392 may include at least model 380 (e.g., within the memory of the controller). Controller 392 preferably may include an interaction / state interpreter 390.

[0063] Sensors 305 and 304 are each configured to transmit a sensor data stream (e.g., sensor data streams 313 and 314 respectively). These sensor data streams may include multiple channels with respect to signals 201, 202, and 203 as described herein. These sensor data streams may be received by controller 392. Additional sensor data streams (e.g., bioacoustic sensor data streams) may also be received by controller 392.

[0064] Within controller 392, the received sensor data streams are preprocessed in blocks 330 and 350. The preprocessed data streams may be sent to at least one model 380. Model 380 may be a gesture classifier, for example, a neural network gesture classifier.

[0065] Models such as model 380 may include a neural network. The neural network may be, for example, a feedforward neural network, a convolutional neural network, or a recurrent neural network, or a graph neural network. The neural network may include a classifier and / or regression. The neural network may apply a supervised learning algorithm. In supervised learning, samples of inputs whose outputs are known are used, from which the network learns to generalize. Alternatively, the model may be constructed using an unsupervised learning algorithm or a reinforcement learning algorithm. In some embodiments, the neural network has been trained such that specific signal characteristics correspond to specific gesture transitions. For example, the neural network may be trained to associate a specific signal, a change in the signal, or a change in multiple signals with a gesture transition (e.g., the pinch action or the unpinch action illustrated in FIG. 2). The neural network may be trained to present a confidence level of the output user action.

[0066] The model may include at least one of an algorithm, a heuristic, and / or a mathematical model. A sensor fusion method may be employed to reduce the influence of the specific non-ideal behavior (e.g., drift and noise) of a particular sensor. Sensor fusion may include the use of a complementary filter, a Kalman filter, or a Mahony & Madgwick orientation filter. Such a model may further include a neural network disclosed herein.

[0067] The model may include, for example, at least one convolutional neural network (CNN) that performs inferences on input data (e.g., signals received from sensors and / or preprocessed data). The convolution may be performed in the spatial or temporal dimension. Features (calculated from the sensor data fed to the CNN) may be selected algorithmically or manually. The model may further include a further RNN (recurrent neural network), which may be used in combination with a neural network that supports user action characteristic identification based on sensor data reflecting user behavior.

[0068] According to the present disclosure, the training of the model may be performed, for example, using a labeled data set including multi-modal biometric measurement data from a plurality of subjects. This data set may be enhanced and augmented using synthetic data. The sequence of computational operations constituting the model may be derived by backpropagation, Markov decision process, Monte Carlo method, or other statistical methods according to the adopted model construction technique. Model construction may include dimensionality reduction techniques and clustering techniques.

[0069] The above-described model 380 is each configured to output a confidence level of a user action based on the received sensor data stream. The confidence level may include one or more probabilities indicated in the form of a percentage value associated with one or more individual user actions. The confidence level may include an output layer of a neural network. The confidence level may be represented in vector form. Such a vector may include a 1×n vector filled with the probability values of each user action known to the controller. For example, the model 380 that receives the sensor data stream may receive a data stream including a signal that the model 380 interprets as a "pinch" gesture transition with a 90% confidence level (interprets that the user has pinched).

[0070] At least one specified gesture state and / or gesture transition (or the confidence level associated with the specified state and / or transition) is output from the gesture classifier 380 and received by the interaction / state interpreter 390. The interaction / state interpreter 390 is configured to identify user input corresponding to at least one received specified gesture state and / or gesture transition and / or the associated confidence level, based at least to some extent on the at least one received gesture state and / or gesture transition. For example, if the model indicates that a "pinch" gesture transition has occurred with a high probability, the interaction / state interpreter may conclude that a "pinch" user input has occurred.

[0071] The interaction / state interpreter 390 may include a list and / or set of user actions, gesture states, and respective confidence levels. For example, when a confidence value that meets at least the required threshold level is received from the gesture classifier, user input corresponding to one gesture state and / or gesture transition is identified. Further, the apparatus may be configured to adjust the threshold level based on, for example, the received affordance profile 311 and / or intention 310.

[0072] The Interaction / State Interpreter 390 may receive optional gaze tracking information 308. The gaze tracking information may include at least a display indicating where the user is looking. This information may include details (e.g., an identifier) regarding the interactable that the user is looking at. This information may be received, for example, from a headset such as an AR / MR / XR / VR headset. This information may be received as a data stream (which may include coordinates).

[0073] If a conflict occurs, for example, the user action indicated by the highest probability may be selected by the Interaction / State Interpreter (argmax function). The Interaction / State Interpreter 390 may include, for example, at least any one of a rule-based system, an inference engine, a knowledge base, a lookup table, a prediction algorithm, a decision tree, and a heuristic. For example, the Interaction / State Interpreter may include an inference engine and a knowledge base. The inference engine outputs a set of probabilities. These sets are processed by the Interaction / State Interpreter, which may include IF-THEN statements.

[0074] The intent data 310 may include the application state, context, and / or the game state of the application and / or device being interacted with. The application state corresponds to the state of the application running on a controllable device and may include, for example, at least any one of variables, static variables, objects, registers, open file descriptors, open network sockets, and / or kernel buffers.

[0075] The above data may be used by the Interaction / State Interpreter to adjust the threshold level of user actions. For example, when an application prompts the user to "pinch" to confirm, the corresponding "pinch" threshold may be lowered so that a pinch action is detected even if the confidence level that a pinch action is occurring, as output by the model, is low.

[0076] An example of the operation of device 300 is shown below. The user is wearing device 300 on the hand. The user performs an action. The action includes pressing the index finger against the thumb. The sensors of device 300 output data as follows during the action. · The IMU sensor 304 outputs a data stream 314 that reflects the orientation of the IMU (and thus device 300) and acceleration data. · The optical sensor 305 outputs a data stream 313, which includes data from a PPG sensor configured to measure, for example, the amount of reflected light from the user's skin, and this data corresponds to the movement of tendons and / or muscles within the user's wrist area.

[0077] The data streams 314 and 313 are received by the model 380. The data streams 314 and 313 may correspond to signals 201, 202, and 203. In the model 380, IMU sensor data and optical sensor data are used by the model, and a gesture transition from unpinch to pinch is detected by the model. Based on the detected gesture transition, a gesture state of "pinch" with a confidence level of 68% is passed to the interaction / state interpreter 390.

[0078] Accordingly, the event interpreter 390 receives the gesture state "pinch" and the associated confidence level. The interaction / state interpreter determines that the threshold of the user action is met. This is optionally done by referring to the received affordance profile 311 and intent data 310 and adjusting the threshold, for example, during the process, based on at least one of the affordance profile 311 and / or intent data 310. Further, the interpreter 390 may receive optional gaze tracking information 308. In this way, the interaction / state interpreter 390 agrees on the "pinch" state.

[0079] Next, state 309 is sent to an application (e.g., a scene such as scene 351) being executed by device 360, which may be, for example, a personal computer or an HMD. In this example, device 360 receives the pinch state without any sensor data, thereby saving the resources of device 360.

[0080] FIG. 7 shows a device 700 that can support at least some embodiments of the present invention.

[0081] Device 700 includes a controller 702. The controller includes at least one processor and at least one memory including computer program code and optionally data. Device 700 may further include a communication unit or communication interface. Such a unit may include, for example, a wireless and / or wired transceiver. Device 700 may further include sensors (e.g., sensors 703, 704) operatively connected to the controller. These sensors may include, for example, at least any one of an IMU 703 and an optical sensor 704. Device 700 may also include other elements not shown in FIG. 7.

[0082] Device 700 is shown as including one processor, but may include two or more processors. In one embodiment, the memory is capable of storing instructions and may be capable of storing, for example, at least any one of an operating system, various applications, models, neural networks, and / or preprocessing sequences. Further, the memory may include storage (which may be used to store at least some of the information and data used in embodiments of the present disclosure).

[0083] Furthermore, the processor is capable of executing the stored instructions. In one embodiment, the processor may be implemented as a multi-core processor, or as a single-core processor, or as a combination of one or more multi-core processors and one or more single-core processors. For example, the processor may be implemented as one or more of various processing devices, and such processing devices include processing cores, coprocessors, microprocessors, controllers, digital signal processors (DSPs), processing circuits with or without DSPs, or other various processing devices, and the other various processing devices include, for example, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), microcontroller units (MCUs), hardware accelerators, dedicated computer chips, and other integrated circuits. In one embodiment, the processor may be configured to execute hard-coded functions. In one embodiment, the processor is implemented as an executor of software instructions, and those instructions, when executed, may specifically configure the processor such that the processor performs at least any one of the models, sequences, algorithms, and / or operations described herein.

[0084] The memory may be implemented as one or more volatile memory devices, and / or as one or more non-volatile memory devices, and / or as a combination of one or more volatile memory devices and one or more non-volatile memory devices. For example, the memory may be implemented as semiconductor memory (e.g., mask ROM, PROM (programmable ROM), EPROM (erasable PROM), flash ROM, RAM (random access memory), etc.).

[0085] At least one memory and computer program code, together with at least one processor, receive a data stream from at least one optical sensor 105 and at least one IMU 104, the sensors being configured to make measurements of a user and / or the device 700 including the sensors 104 and 105 and the device 700 being configured to be worn by a user, the receiving, passing the received data stream to a gesture classifier 380, and using the gesture classifier 380 to identify a gesture state by detecting and classifying gesture transitions based on transient events in the sensor data within the received data stream, the gesture classifier 380 including a neural network classifier trained to detect the gesture transitions, and the identifying may be configured to cause the device 700 to perform at least the above, and the data stream may optionally be pre-processed before being passed to the gesture classifier. This has been discussed in detail above with reference to FIGS. 1-5 already.

[0086] The device 700 may be configured to pass the identified gesture state to an interaction interpreter and / or a state interpreter configured to apply corrections based on an application state and / or context. The interaction / state interpreter may be provided by a controller of the device 700.

[0087] The device 700 may be configured to predict a state transition by detecting the statelessness of a gesture based on the detected transition and to infer a pinch state from the acquired data.

[0088] Device 700 may be configured such that an interaction interpreter and / or a state interpreter is configured to pass a detected gesture state to at least one of XR / VR / MR / AR applications (e.g., running on a computing device). The passed gesture state may be provided as an event. Alternatively or additionally, the passed gesture state may be provided as a polable, whereby the device can be queried regarding the event state. Device 700 may be configured to receive at least one of an intent estimate or an affordance profile from an XR / VR / MR / AR application or a computing device. Device 700 may be configured such that an adaptation time window is used as part of the detection and the adaptation time window is specifically adjusted to match a typical gesture duration (e.g., 1 millisecond).

[0089] Device 700 may be configured to receive context information (e.g., a modification and / or a context clue) from an XR / VR / MR / AR application or a computing device, and the device is configured to use the received context information to adjust statefulness detection. For example, the context information may indicate a "drop zone" within a scene. When the user moves something to the drop zone, a drop may be initiated (e.g., by an interpreter). This may be initiated even if the confidence of a pinch-unpinch transition (pinch-release) does not exceed a global threshold. Device 700 may be configured to apply dead reckoning corrections and context clues to improve statefulness detection. Device 700 may be configured to improve statefulness detection by using a recurrent model to incorporate a gesture history (e.g., a user-specific gesture history). What is incorporated in this way may include, for example, information regarding previous user actions and / or previously interacted elements.

[0090] Device 700 is configured to convert discrete temporal events (e.g., taps or releases) into states (e.g., pinch or unpinch states) in order to detect transitions (e.g., touch or release of the index finger and thumb).

[0091] Device 700 may be configured to perform on-board processing, which may include at least any one of pre-processing, gesture classification, or state interpretation.

[0092] Device 700 may be configured to implement a method for identifying a selected gesture from acquired data, the method including: step (901) of acquiring data from at least one optical sensor (105) and at least one IMU (104); and step (903) of passing the data to a gesture classifier, the gesture classifier being configured to detect whether the hand is in a gesture state (e.g., a pinch state), and the data may optionally be pre-processed before being passed to the classifier, the above step (903).

[0093] The above method may include the step of using a neural network configured to detect transitions between gesture states by detecting the rising and falling edges of the gesture. The above method may include the step of detecting transitions between a pinch state and an unpinch state. The above method may include the step of performing on-board processing (e.g., any one of the steps of pre-processing, gesture classification, transition detection).

[0094] The present disclosure may also be utilized by the following clauses.

[0095] Clause 1. A multimodal biometric measurement device (100, 300, 700), the device comprising a controller including at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to receive (901) a data stream from at least one optical sensor (105) and at least one IMU (104), the sensors being configured to perform measurements on a user, the receiving (901); passing (903) the received data stream to a gesture classifier (380); using the gesture classifier (380) to identify a gesture state by detecting and classifying a gesture transition based on transient events in the sensor data within the received data stream, the gesture classifier (380) including a neural network classifier trained to detect the gesture transition, the identifying; and causing the controller to perform at least the above, the data stream may optionally be preprocessed before being passed to the gesture classifier, device.

[0096] Clause 2. The device according to clause 1, further comprising a strap and a housing including an optical sensor and an IMU, the optical sensor being disposed within the housing such that it is capable of performing measurements on a user's wrist area when worn by the user.

[0097] Clause 3. The device according to clause 1 or 2, including a smartwatch.

[0098] Clause 4. The device according to any one of clauses 1 to 3, wherein the optical sensor includes a PPG sensor.

[0099] Clause 5. The device according to any one of clauses 1 to 4, wherein the controller is disposed within the housing.

[0100] Clause 6. A kit including at least one wrist wearable device (100), the kit Receiving a data stream from at least one optical sensor and at least one IMU, wherein the sensor is configured to perform measurements of a user, said receiving, Passing the received data stream to a gesture classifier, Using the gesture classifier to identify a gesture state by detecting and classifying a gesture transition based on a transient event in sensor data within the received data stream, wherein the gesture classifier includes a neural network classifier trained to detect said gesture transition, said identifying, being configured to perform, The data stream may optionally be pre-processed and then passed to the gesture classifier. Kit.

[0101] Clause 7. The kit includes a head-mounted device (HMD) including a display, the HMD being connected to the device (100) and being configured to display the gesture state and / or gesture transition using the display, the kit according to clause 6.

[0102] Clause 8. The wrist-wearable device includes a display, the display being connected to the controller (103) and being configured to display the gesture state and / or gesture transition using the display, the kit according to clause 6 or 7.

[0103] Clause 9. The HMD is configured to pass context information to the wrist-wearable device and / or the HMD is configured to interpret and apply a confidence value of the gesture state passed from the wrist-wearable device, the kit according to clause 7.

[0104] Clause 10. The wrist-wearable device includes a neural network classifier trained to detect a gesture state based on sensor data, the device or kit according to any one of clauses 1 to 9.

[0105] As a further advantage of the present disclosure, by using IMU sensor data and optical sensor data obtained from a wearable sensor (e.g., a wrist-mounted sensor), common problems in VR / XR / AR interactions (e.g., problems with the field of view or obstacles) are overcome.

[0106] Embodiments of the disclosure provide a technical solution to a technical problem. One of the technical problems to be solved is the detection of pinch gestures or tap gestures by a device that performs on-board processing based on IMU sensor data and optical sensor data. In practice, this is a problem because various information related to the movement of the user's hand can be included in the data.

[0107] The embodiments described herein overcome these limitations. This is done by using machine learning to detect gesture states based on transient events in the sensor signals. In this way, on-board gesture detection can be achieved more accurately and robustly. This brings several advantages. First, since the processing can be performed near the sensor, communication latency is minimized. Second, interpreter corrections (e.g., based on the application state) can also be performed near the sensor. Third, by using the selected sensor, the user can carefully give UI commands and the influence of hand position is minimal (e.g., compared to a camera-based system). Other technical improvements may also result from these embodiments and other technical problems may also be solved.

[0108] Naturally, the disclosed embodiments of the present invention are not limited to the specific structures, processing procedures, or materials disclosed herein, but extend to equivalents that would be understood by those skilled in the art. Further, naturally, the terms used herein are used only for the purpose of describing specific embodiments and are not intended to be limiting.

[0109] References to "one embodiment" or "an embodiment" throughout this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. For example, when a numerical value is referenced using a phrase such as "about" or "substantially", the exact numerical value is also disclosed.

[0110] A plurality of items, structural elements, components, and / or materials used in this specification may, for convenience, be presented in a common list. However, such lists should be construed as if each element of the list were individually identified as a distinct and unique element. Thus, unless the contrary is indicated, the individual elements of such lists should be construed as being merely based on their presence in a common group and not as de facto equivalents of any other element of the same list. Further, in this specification, various embodiments and examples of the invention may be referred to in conjunction with their various components. Of course, such embodiments, examples, and alternatives should not be construed as being de facto equivalents of each other, but rather as distinct and independent representations of the invention.

[0111] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this description, various specific details, such as examples of length, width, shape, etc., are shown so as to provide a thorough understanding of the embodiments of the invention. However, as will be understood by those skilled in the art, the invention may be practiced without one or more of these specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail, but this is for the purpose of preventing the aspects of the invention from being obscured.

[0112] Each of the above embodiments illustrates the principles of the present invention in one or more specific applications. However, as will be apparent to those skilled in the art, various changes in the form, usage, and details of the implementation may be made without exercising inventive faculty and without departing from the principles and concepts of the present invention. Accordingly, the present invention is not limited except as defined by the following claims.

[0113] In this document, the verbs "to comprise" and "to include" are used as open limitations that do not require the exclusion or necessitation of the presence of features not described. The features recited in the dependent claims may be freely combined with each other unless otherwise expressly stated. Further, of course, the use of "a" or "an", i.e., the singular, does not exclude plurality throughout this document.

Industrial Applicability

[0114] At least some embodiments of the present invention have industrial applications in providing a user interface (e.g., related to XR) to a controllable device (e.g., a personal computer or an HMD).

Explanation of Signs

[0115] 100, 300, 700, 360 Devices 101 Mounting Component 102 Housing 103, 392, 702 Controllers 104, 304, 703 Inertial Measurement Unit (IMU) 105, 305, 704 Optical Sensors 210 Wrist Measurement Range 20 User 201A~B, 202A~C, 203A~B Sensor Signals 308 Fixation Tracking Information 309 Status Information 310 Intention Data 311 Affordance Profile 313, 314 Sensor Data Streams 330, 350 Preprocessing Blocks 351 Scene 380 Model 390 Interpreter 901, 902, 903, 904, 905 Each Step of the Method

Claims

1. A wrist wearable device (100) comprising a processing core (701) and at least one memory (707) containing computer program code, the at least one memory (707) and the computer program code together with the at least one processing core (701) performing at least: receiving a data stream from at least one optical sensor (105) and at least one IMU (104), the sensor being configured to perform measurements of a user; passing the received data stream to a gesture classifier (380); identifying a gesture state by detecting and classifying gesture transitions based on transient events in the sensor data of the received data stream using the gesture classifier (380), the gesture classifier (380) comprising a neural network classifier trained to detect the gesture transitions; and configured to cause the device to The data stream may be optionally pre-processed before being passed to the gesture classifier. Device.

2. The apparatus of claim 1 , further configured to pass the identified gesture state to an interaction interpreter and / or a state interpreter configured to apply modifications based on an application state and / or context.

3. The device of any one of claims 1 to 2, further configured to infer a pinch state from the received data stream based on the identified gesture state.

4. The apparatus of any one of claims 1 to 3, wherein the interaction interpreter and / or state interpreter are further configured to pass the identified gesture state to an extended reality (XR) application.

5. The device of any one of claims 1 to 4, further configured to receive at least one of an intent estimation or an affordance profile from an XR application.

6. 6. The apparatus of claim 1, further configured such that an adaptive time window is used as part of the detection, the adaptive time window being specifically tuned to a typical gesture duration (e.g. 1 millisecond).

7. 7. The apparatus of claim 1, further configured to receive context information (e.g., modifications and / or context clues) from the XR application and configured to adjust statefulness detection using the received context information.

8. 8. The device of claim 1 , further configured to receive gaze tracking information (e.g., information about the location of the user's gaze and nearby interactables) and configured to adjust statefulness using the received gaze tracking information.

9. The apparatus of any one of claims 1 to 8, further configured to apply dead reckoning corrections and context cues to improve statefulness detection.

10. The apparatus of any one of claims 1 to 9, further configured to incorporate gesture history using a recurrent model to improve statefulness detection.

11. The apparatus of any preceding claim, further configured such that the optical sensor value is higher at the start and end of the gesture as opposed to midpoints of the gesture session.

12. 12. The device of claim 1, further configured to convert discrete temporal events (e.g., taps and releases) into states (e.g., pinch and unpinch states) to detect transitions (e.g., touch and release of index finger and thumb).

13. The apparatus of any one of claims 1 to 12, comprising the IMU (104) and the optical sensor (105).

14. The device of claim 1 , further configured to perform on-board processing, the on-board processing including the pre-processing, the gesture classification, and the state interpretation.

15. 1. A method for identifying a selection gesture from captured data, comprising: receiving (901) a data stream from at least one optical sensor (105) and at least one IMU (104), the sensor being configured to perform measurements of a user; passing (903) the received data stream to a gesture classifier (380); identifying a gesture state by detecting and classifying gesture transitions based on transient events in the sensor data of the received data stream using the gesture classifier (380), the gesture classifier (380) comprising a neural network classifier trained to detect the gesture transitions; Including, The data stream may be optionally pre-processed before being passed to the gesture classifier. method.

16. The method of claim 15 , wherein the identified gesture state is passed to an interaction interpreter and / or a state interpreter configured to apply modifications based on application state and / or context.

17. 17. The method of claim 15 or 16, further comprising detecting a transition between a pinched condition and an unpinch condition.

18. The method according to any one of claims 15 to 17, further comprising the step of performing on-board processing (e.g. of the steps of pre-processing (902), the gesture classification (903), the transition detection).

19. 19. The method of any one of claims 15 to 18, further comprising receiving context information (e.g., corrections and / or context clues) from an XR application and adjusting statefulness detection using the received context information.

20. A non-transitory computer readable medium storing a set of computer readable instructions, the set of instructions, when executed by at least one processor, to perform at least: receiving a data stream from at least one optical sensor (105) and at least one IMU (104), the sensor being configured to perform measurements of a user; passing the received data stream to a gesture classifier (380); identifying a gesture state by detecting and classifying gesture transitions based on transient events in the sensor data of the received data stream using the gesture classifier (380), the gesture classifier (380) comprising a neural network classifier trained to detect the gesture transitions; The device performs the following: The data stream may be optionally pre-processed before being passed to the gesture classifier. Non-transitory computer-readable medium.