Kneading state detection system and method

The wrist-wearing device receives optical sensors and IMU data, and uses a neural network classifier to detect gesture status, solving the problem of traditional controllers hindering human-computer interaction, and achieving efficient and flexible user interface control.

CN120066247APending Publication Date: 2025-05-30DOUBLEPOINT TECH OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411720573.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional physical controllers add a technical layer between the user and the computing device, hindering human-computer interaction, and dedicated devices are usually only suitable for specific devices and may hinder the user's hand freedom when used.

Method used

A wrist-wearing device is designed to directly digitize and convert the subtle movements and gestures of the hands into machine commands by receiving data flow from optical sensors and IMUs.

Benefits of technology

It realizes accurate detection and classification of user gestures without interfering with the normal use of the user's hands, and improves the efficiency and flexibility of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066247A_ABST
    Figure CN120066247A_ABST
Patent Text Reader

Abstract

According to an exemplary aspect of the present invention, there is provided an input device and a corresponding method that can digitize and convert subtle actions and gestures of a hand into directional light without interfering with normal use of the hand. For example, the apparatus and method may detect and classify gesture transients from transients of sensor data in a data stream received from a plurality of sensors, preferably of different types, to determine a gesture state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various exemplary embodiments of the present disclosure relate to a wearable device and a method, where the device and method can be used for controlling a device, particularly in the field of computing and extended reality user interface applications. Extended reality (XR) includes augmented reality (AR), virtual reality (VR), and mixed reality (MR). Background Art

[0002] Traditionally, digital devices have been controlled by dedicated physical controllers. For example, a keyboard and mouse can be used to operate a computer, a handheld controller can be used to operate a gaming console, and a smartphone can be operated through a touch screen. Generally, these physical controllers include sensors and / or buttons for receiving input from a user based on the user's actions. Such discrete controllers are ubiquitous, but they add an unnecessary technological layer between the user's hands and the computing device, thus hindering human-computer interaction. In addition, such dedicated devices are usually only suitable for controlling specific devices. Moreover, such devices may, for example, be obstructive to the user, preventing the user from using their hands for other purposes while using the control device. Summary of the Invention

[0003] In view of the above problems, there is a need for improvement in the field of XR user interfaces. According to the present invention, a suitable input device can directly digitize and convert subtle movements and gestures of the hand into machine commands without interfering with the normal use of the user's hands. For example, embodiments of the present disclosure can determine a gesture state based on sensor data streams received from at least one sensor, preferably a plurality of sensors, so as to determine the user's actions, where the sensors are preferably of different types.

[0004] The present invention is defined by the features of the independent claims. Some specific embodiments are defined in the dependent claims.

[0005] According to a first aspect of the present invention, there is provided a wrist-worn device, comprising a processing core, and at least one memory including computer program code, the at least one memory and the computer program code being configured to, together with at least one processing core, cause the device to at least: receive data streams from at least one optical sensor and at least one IMU, where the sensors are configured to measure a user; provide the received data streams to a gesture classifier; and use the gesture classifier to determine a gesture state, where gesture transitions are detected and classified based on transient data in the sensor data of the received data streams, and the gesture classifier includes a neural network classifier trained to detect the gesture transitions. Wherein, the data streams can be selectively preprocessed before being provided to the gesture classifier.

[0006] According to a second aspect of the present invention, a method for recognizing a selection gesture from acquired data is provided. The method includes: receiving data streams from at least one optical sensor and at least one IMU, where the sensors are configured to measure a user; providing the received data streams to a gesture classifier; and using the gesture classifier to determine a gesture state, where gesture transitions are detected and classified based on transient data in the sensor data of the received data streams, and the gesture classifier includes a neural network classifier trained to detect the gesture transitions. Optionally, the data streams are preprocessed before being provided to the gesture classifier.

[0007] According to a third aspect of the present invention, a computer program is provided, which when executed by at least one processor, can cause a device to at least: receive data streams from at least one optical sensor and at least one IMU, where the sensors are configured to measure a user; provide the received data streams to a gesture classifier; and use the gesture classifier to determine a gesture state, where gesture transitions are detected and classified based on transient data in the sensor data of the received data streams, and the gesture classifier includes a neural network classifier trained to detect the gesture transitions. Optionally, the data streams are preprocessed before being provided to the gesture classifier.

[0008] According to a fourth aspect of the present invention, a non-transitory computer-readable medium is provided, on which a set of computer-readable instructions is stored, which when executed on a processor, causes the execution of the method according to the second aspect above, or causes a device including a processor to be configured according to the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1A and 1B shows a schematic diagram of an exemplary device 100 capable of supporting at least some embodiments of the present invention;

[0010] Figure 2 shows exemplary sensor data corresponding to a user action;

[0011] Figure 3 shows an exemplary measurement area of a device capable of supporting at least some embodiments of the present invention;

[0012] Figure 4 shows interactive elements of an exemplary user interface of a device capable of supporting at least some embodiments of the present invention;

[0013] Figure 5 shows an exemplary process in the form of a flowchart capable of supporting at least some embodiments of the present invention;

[0014] Figure 6 shows an apparatus capable of supporting at least some embodiments of the present invention; and

[0015] Figure 7 shows an apparatus capable of supporting at least some embodiments of the present invention. Detailed Description

[0016] A multimodal biometric measurement apparatus and a corresponding method are described herein. Multimodal biometric measurement can be understood as using multiple modalities (such as sensors) to acquire, process, and / or analyze multiple biometric-related measurement information, that is, measurements related to the user's body and its movements. The apparatus and / or method can be used for at least one of the following tasks, such as: measurement, sensing, signal acquisition, analysis, user interface tasks. Such an apparatus is preferably suitable for the user to wear. The user can use the apparatus to control one or more devices. Such control can take the form of a user interface (UI) or a human-machine interface (HMI). The apparatus can include one or more sensors. The apparatus can be configured to generate user interface data based on data from the one or more sensors. The user interface data can be at least partially used to allow the user to control the apparatus or at least one second device. For example, the second device can be at least one of a personal computer (PC), a server, a mobile phone, a smartphone, a tablet device, a smartwatch, or any type of suitable electronic device. The controlled or controllable device can execute at least one of the following: an application program, a game, and / or an operating system, any of which can be controlled by the multimodal device.

[0017] When interacting with AR / VR / MR / XR application programs, such as when using devices such as smartwatches or extended reality headsets, users need to be able to perform various actions (also known as user actions), such as select, drag and drop, rotate and drop, use sliders, and zoom. This embodiment improves the detection of user actions, which can improve the response speed when implemented in a controller or system or by a controller or system.

[0018] Statelessness is a well-known term in the fields of computing and user interfaces. Essentially, it means that both the state or value of user input and output are stored in the system memory. Past actions or events have an impact on the system, so that current or future actions will obtain different contexts and / or responses, rather than just monitoring transitions. Statelessness is an important aspect of human-computer interaction devices because it enables richer interactions compared to instantaneous input. Interaction operations such as drag and drop, pinch to zoom, etc. all require statelessness.

[0019] A user may perform at least one user action. Generally, the user will perform an action to affect a controllable device and / or an AR / VR / MR / XR application. For example, the user action may include at least one of the following: movement, gesture, interaction with an object, interaction with a user body part, air action. An example of a user action is a "pinch" gesture, i.e., the user touches the tip of the index finger with the tip of the thumb. Another example is a "thumbs up" gesture, i.e., the user extends their thumb, bends the fingers and rotates the palm so that the thumb points upwards. This embodiment is configured to determine the gesture state at least in part based on sensor input, thereby identifying (determining) at least one user action. Reliably identifying user actions enables action-based control, such as using the embodiments disclosed herein.

[0020] The user operation may include at least one state transition. For example, the user action may be a pinch, i.e., a transition from a non-pinch state to a pinch state. In some embodiments, the gesture transition, i.e., the transition between states, is determined at least in part based on sensor data by a neural network or the like.

[0021] In at least some embodiments, a user interface command is generated by the device (e.g., device 100) according to the determined gesture state. The user interface command may be referred to as a human-machine interface command, a user interface interaction, and / or a human-machine interface interaction, for example. Such a command can be used to convey a user interface action. For example, the user interface action may correspond to the "enter" key on a personal computer. At least in some embodiments, an output generator (e.g., a user interface command generator) is used to map the user action to a predetermined user interface action and / or generate a user interface command according to the predetermined user interface action. Therefore, the controlled device does not need to understand the user action, but only needs to understand the user interface command, which can be provided in a standard format, such as a human-machine interface device (HID) report descriptor. The UM command may include, for example, a data packet and / or a frame containing UI instructions.

[0022] In at least some embodiments, the device includes a mounting component configured to be worn by the user. Therefore, the mounting component is a wearable component. Such a component may be any of the following: a strap, a band, a wristband, a bracelet, a glove. The mounting component may be connected to and / or formed by another device (e.g., a smartwatch), or form part of a larger device (e.g., a gauntlet). In some embodiments, the width of the strap, band, and / or wristband is 2 to 5 cm, preferably 3 to 4 cm, more preferably 3 cm.

[0023] In some embodiments, a signal is measured by at least one sensor. The sensor may include a processor and a memory or be connected to a processor and a memory, where the processor may be configured to perform the measurement. The measurement may be referred to as sensing or detection. The measurement includes, for example, detecting a change in at least one of a user, an environment, and the physical world. For example, the measurement may further include applying at least one timestamp to the sensed data and transmitting the sensed data (with or without the at least one timestamp).

[0024] The signal data may include one or more transients. A transient may be any change in the signal. In particular, a transient may include a rising edge and a falling edge of the signal.

[0025] At least some embodiments are configured to measure user characteristics, such as posture, position, and movement, by using an inertial measurement unit (IMU). The IMU may be configured to provide position information of the device. The IMU may be referred to as an inertial measurement unit sensor. The IMU may include at least one of the following: a multi-axis accelerometer, a multi-axis gyroscope, a multi-axis magnetometer, an altimeter, a barometer. The IMU preferably includes a magnetometer because the magnetometer may provide an absolute reference for the IMU. A barometer that can be used as an altimeter may provide an additional degree of freedom for the IMU. The IMU signal data may include transients.

[0026] At least some embodiments are configured to measure user characteristics, such as gestures, position, and / or movement, by using an optical sensor. For example, the optical sensor may include a photoplethysmogram (PPG) sensor. The optical sensor (e.g., a PPG sensor) may be referred to as a biometric sensor and may measure, for example, at least one of the following data: heart rate, blood oxygen saturation, or other biometric data. A light-emitting diode (LED) is typically used to illuminate the underlying tissue and measure the brightness changes that occur during the cardiac cycle. The optical sensor may use visible light and infrared light to implement these functions. The optical signal data may include transients.

[0027] For application in gesture classification, motion artifacts in the optical sensor data may be minimized to obtain a more accurate reading. For example, a low-pass filter or a median filter may be used to minimize fast motion artifacts, while a high-pass filter may be used to minimize slow motion artifacts. In addition to measuring heart rate and blood oxygen saturation, the optical sensor data may be used for other purposes. For example, a PPG sensor may be used to measure the amount of reflected light. The amount of reflected light is affected by the proximity of the sensor to the user's skin and the reflectivity of the underlying tissue. The proximity is affected by tendon movement, for example.

[0028] At least some embodiments are configured to measure vibration signals and / or sound signals. For example, such signals can be measured from a user, particularly from the user's wrist area. A bioacoustic sensor can be configured to measure vibrations, such as internal waves within the user's body and / or surface waves on the user's skin. For example, such vibrations may result from a touch event. A touch event can include a user touching an object and / or the user touching themselves and / or another user. The bioacoustic sensor can include a contact microphone sensor (piezoelectric contact microphone), a MEMS (microelectromechanical systems) microphone, and / or a vibration sensor, such as an accelerometer for detecting the acceleration generated by an acoustically vibrating structure (such as the user or a mounting component).

[0029] Sensors, such as inertial measurement unit sensors, bioacoustic sensors, and / or optical sensors, can be configured to provide a sensor data stream. The data stream can be provided within the device, for example, to a controller, a digital signal processor (DSP), and a memory. As an alternative or in addition, it can also be provided to an external device. The sensor data stream can include one or more raw signals measured by the sensors. In addition, the sensor data stream can also include at least one of synchronization data, configuration data, and / or identification data. For example, the controller can use this data to compare data from different sensors.

[0030] In at least some embodiments, the device includes means for providing kinesthetic feedback and / or tactile feedback. For example, such feedback can include vibrations and can be referred to as "haptic feedback". The means for providing haptic feedback can include, for example, at least one of the following: an actuator, a motor, a haptic output device. A preferred actuator is a linear resonant actuator (LRA), which can be used in all embodiments. Haptic feedback can provide the user with feedback from the user interface, thereby improving the usability of the user interface. In addition, haptic feedback can also alert the user to events or context changes in the user interface.

[0031] Figure 1A and 1B Fig. 12 shows a device 100 according to at least some embodiments of the present invention. The device 100 can be used to sense the movement and / or actions of a user and can be used as a user interface device and / or an HMI device. The device 100 can include at least one optical sensor that is arranged to face the user's hand. The device 100 can include at least one inertial measurement unit 104 within a housing 102. Sensor data is acquired by an on-board controller. The controller can preprocess the sensor data. The controller can be configured to detect a gesture state and a corresponding user action, such as whether the hand is in a pinching state, for example, using a classifier or a classifier algorithm (such as a neural network classifier).

[0032] In at least some embodiments, the device 100 may include a smartwatch. Preferably, the device 100 includes a screen. The screen can be used to provide gesture status information to the user. In at least some embodiments, the device 100 may include at least one of the following, for example: a haptic device, a screen, a touch screen, a speaker, a heart rate sensor (such as an optical sensor), a Bluetooth communication device.

[0033] The device 100 includes a mounting member 101, which is in the form of a band. Figure 1A The device includes a housing 102 connected to the mounting member 101. The user 20 can wear the device 100, for example, on the wrist, such that the sensor 104 faces the user's wrist area, particularly the skin of the wrist area.

[0034] The device 100 further includes a controller 103 within the housing 102. The controller 103 includes at least one processor, at least one memory, and a communication interface, where the memory may include instructions that, when executed by the processor, allow communication with other devices (such as a head device and / or a computing device) through the communication interface. The controller can be configured to cause the controller to receive at least concurrent sensor data streams from sensors 104 and 105. The controller can be configured to perform preprocessing on at least one of the sensor data streams, where the preprocessing can include, for example, at least one of data cleaning, feature transformation, feature engineering, and feature selection. The controller can be configured to process the sensor data streams received from the sensors. The controller can be configured to generate at least one user interface, event, and / or command based on the characteristics of the processed sensor data streams. Additionally, the controller may include a model, which may include at least one neural network.

[0035] The device 100 further includes an inertial measurement unit (IMU) 104. The IMU 104 can be located within the housing 102. The IMU 104 is configured to provide, for example, a sensor data stream to the controller. The sensor data stream includes, for example, at least one of the following: multi-axis accelerometer data, gyroscope data, and / or magnetometer data. These three types of data represent acceleration, angular velocity, and magnetic flux density, respectively. One or more components of the device 100 can be combined together. For example, the IMU 104 and the processor of the controller can be located on the same PCB (printed circuit board). Changes in the IMU data reflect the user's movement and / or actions, and thus such movement and / or actions can be detected through the IMU data.

[0036] In at least some embodiments, the inertial measurement unit (IMU) 104 or a sensor (such as a gyroscope) or multiple sensors mounted in a single device or multiple devices are placed and / or fixed on the user's arm. For example, the mounting member 101 can be used for fixation.

[0037] Figure 1BShows the inside of device 100, which is configured to be worn against the user's skin, for example, on the wrist. The optical sensor 105 is shown on one surface of the housing 102. The optical sensor 105 may include a PPG sensor. The optical sensor 105 may include at least one LED. The optical sensor 105 is configured to provide a sensor data stream, for example, to the controller. The sensor data stream may include, for example, at least one of the following: heart rate data, analog-to-digital conversion (ADC) data corresponding to the amount of reflected light. One or more components of device 100 may be combined together. For example, the processor of the optical sensor 105 and the controller may be connected to the same printed circuit board. Changes in the optical sensor data reflect, for example, the user's movement and / or action, so the user's movement and / or action can be detected by using the optical sensor data.

[0038] Figure 2 Shows exemplary sensor data corresponding to the user's actions. The lower part of the figure shows a series of states of the user's hand. First, the user's hand is in an unpinched state (also known as a non-pinch state). Then, the user forms a pinch with the thumb and index finger. These two fingers are commonly used for pinching, but other fingers can also be used. Finally, the user releases the pinch and the hand is in the unpinched state again.

[0039] Figure 2 The upper part shows the corresponding sensor data, which can be measured, for example, using device 100 when the user performs a series of hand gestures corresponding to the gesture interface states. The transitions from the unpinched state to the pinched state and from the pinched state to the unpinched state are marked in the figure. Such transitions can be called gesture transitions. As shown, several transients occur during the gesture transition, where the signal changes during the transition.

[0040] The data stream may include one or more signals. In the figure, signals 201A and 201B correspond to the sensor data received and measured by the optical sensor. The optical sensor may provide signals corresponding to the sensed visible light data and the sensed infrared light data. Among them, the infrared data is preferred. Signals 202A, 202B, and 202C correspond to the received signals of the gyroscope, and the gyroscope may be part of the IMU. The gyroscope is three-axis, so three signals are received from it. Signals 203A and 203B correspond to the signals received from the accelerometer, and the accelerometer may be part of the IMU. The accelerometer is three-axis, so three signals are received from it. The sensor can measure data related to the user's hand, such as the user's wrist area. The sensor signals shown in the figure are related to Figure 2 the user's hand position shown in the lower part.

[0041] In Figure 2In an exemplary case, the user and the device are initially in an "unpinched state". When a transition that meets the pre-learned unpinched-pinched transition definition is detected, the device 100 enters the "pinched state". Although pinching and unpinching are reverse actions in the field of gestures, the signal patterns associated with these two gestures are different and asymmetric. For example, pinching (tapping) is characterized by an obvious vibration mechanical pattern in the wearable device, while the vibration mechanical pattern during unpinching is less obvious. In the optical field, both pinching and unpinching have rising and falling edges, but due to the inherent characteristics of optical sensing, these edges are not obvious enough to be used for reasoning alone. Therefore, it is necessary to combine inertial sensors and optical sensors. In addition, although IMU data only provides transition information, optical data can at least provide some state information and transition information. The model established according to the present disclosure can be configured to utilize these two types of information, for example, through machine learning.

[0042] Relative values and / or thresholds can be used instead of absolute values to detect transitions because the influence of sensor data on the signal may be small (i.e., "local peaks and valleys"). Since these local peaks and valleys are not so obvious in the global signal, absolute values are not as helpful as relative values for detecting transitions. The device 100 can use the AC component of the input signal for adaptive offset cancellation. Adaptive offset cancellation allows the detection of local rising and falling edges. The device 100 is configured to adjust the sensitivity of unpinched transition detection, for example, when the device is in the pinched state. Sensitivity means adjusting the time window or relative value. Sensitivity adjustment may also involve hand trajectory analysis. Generally, hand movements in a tap task are characterized by ballistic and correction phases. Pinching or unpinching gestures are more likely to occur in the correction phase rather than the ballistic phase. In the ballistic phase, the state of the gesture (pinched or not pinched) may be one of them. Therefore, the state lag information before the trajectory phase can be incorporated into the adaptive sensitivity.

[0043] In some embodiments, the optical sensor values are higher relative to the start and end of the gesture compared to the gesture session average. In other words, more importance can be attached to the start and end of the gesture to detect local peaks and valleys.

[0044] As shown in the figure, when a pinched-unpinched state transition is detected, the device enters the unpinched state. In addition, the device 100 can measure the transition time of each user from the unpinched state to the pinched state. The device can ignore any transitions that occur quickly after the previous transition or illogical transitions. For example, a user cannot transition from the unpinched state to the unpinched state. This can be considered by the gesture classifier and / or the interaction / state interpreter.

[0045] In some embodiments, an adaptive time window is used as part of gesture classification. The adaptive time window can be specifically adjusted according to the duration of typical gestures, for example, based on training data. For example, the adaptive time window can be 1 millisecond with an adaptive range of 1 - 50%. The data used for classification and / or interpretation is stored in the controller of the device. For example, user - specific data that has been acquired in a single session can be retained or discarded after inference. The benefit of retaining the data is that it can be further adapted to the user, while discarding the data can save memory resources.

[0046] Figure 3 An exemplary measurement area of the device disclosed herein is shown. The measurement area 210 related to the user's hand 20 is shown in the figure. Sensors, such as sensors 104, 105, can measure signals from the user within area 210 or at least signals related thereto. The IMU can also be arranged within or near area 210. Since the housing 102 is connected to a mounting component 101, such as a mounting component worn by the user, the sensors and the IMU are preferably arranged to enable measurements of the user.

[0047] Figure 4 Exemplary interactive elements of a user interface are shown. For example, as part of a VR, AR, MR, or XR application, these interactive elements can be displayed to the user via a head - mounted device 21. The head - mounted device 21 can include at least one of a head - mounted display (HMD), a head - frame, a virtual - reality head - mounted device, an extended - reality head - mounted device, an augmented - reality head - mounted device, etc. The head - mounted device 21 can include a processor, a memory, and a communication interface, where the memory can include instructions that, when executed by the processor, allow communication with other devices (such as device 100) via the communication interface.

[0048] The device in the present disclosure, such as device 100, can be configured to communicate with at least one computing device. The computing device can include at least one of a computer, a server, and a mobile device such as a smartphone or a tablet. The device can be configured to at least participate (e.g., by computing) in providing a VR / XR / AR experience to the user, for example, by providing a scene for the HMD to display. The computing device can include a processor, a memory, and a communication interface, where the memory can include instructions that, when executed by the processor, allow communication with other devices (such as device 100 or the head - mounted device) via the communication interface. As described above, the computing device can be the head - mounted device 21.

[0049] The interactive elements 31, 32, 33, and 34 show different types of interactions that the user 20 can have with a VR, AR, or XR application (also referred to as a scene). The device 100 detects the user interaction and accordingly updates the scene displayed to the user 20 via the device 21, for example, updated by a computer or a mobile device connected to the device 100 and the device 21.

[0050] The interactive element 31 is an interactive button that the user can press to perform an operation. Like a physical button, the button needs to be moved along an axis. The interactive element 32 is a slidable interaction method where the user can move a slider along a provided axis, such as to adjust a value. The interactive element 33 is a scrollable interaction method where the user can move a slider and / or the screen along at least one axis, such as to move a view or scroll through content. The interactive element 34 is a panning interaction method where the user can move a slider or a view along a provided axis to move the view. Other types of interactive elements include "grabbable elements", i.e., things that the user can grab with their hand. Interactive elements are not limited to a single type; for example, a slider is grabbable.

[0051] The XR application is configured to provide an indication of the interactive element to the device 100, for example. The indication may include at least one of the following attributes of the interactive element: position, type, size, orientation, status. The device 100 may be configured to receive the indication and adjust internal processing accordingly, such as through function visibility.

[0052] In the field of user interfaces, "function visibility" is a sensory cue (usually visual) that tells the user how to interact with a particular element. These cues include geometric shapes, symbols (such as arrows), etc. Common function visibility includes a scrollable scroll bar and a pressable button. When an application can predict the type of action the user may perform next, an interaction interpreter, etc., can use a function visibility profile to improve the accuracy of detecting the action. For example, after a pinch, it is usually released (not pinched) sooner or later. For example, the function visibility profile may include the user's expected actions for the interactive element. For example, the prediction can be performed by a neural network executed by the controller of the device.

[0053] In some embodiments, the device 100 may be configured to obtain or receive data from at least one optical sensor 105 and at least one IMU 104 and provide the obtained / received data to a gesture classifier, where the gesture classifier is configured to determine whether the hand is in a gesture state (such as a pinch state) or not. The determination can be performed by detecting and classifying a gesture transition based on transients in the obtained / received data. The device 100 may be configured to perform preprocessing on the data before providing one or more data (such as one or more data streams) to the classifier. The device 100 may be configured to use a neural network to perform the detection, and the neural network is trained to detect a transition between gesture states by detecting the rising and falling edges of the gesture.

[0054] The behavior of gesture prediction (such as thresholds and rule - based time limits) can be modified according to prior - known constraints of the interactive element and the scene logic. These constraints can be provided to the device, for example, provided by a scene such as scene 351 and stored in device 100. The constraints can be part of gesture detection, such as a time limit that no reverse gesture can occur within a predetermined time (such as 50 milliseconds) after the previous gesture. Such corrections can be applied by an interaction and / or state interpreter configured to apply corrections based on the determined gesture state and based on the application state and / or application context.

[0055] The device can be configured to provide conversions in the form of events. In other words, the device can provide an event, such as "pinch down". The device can be configured to return the user state in a "pollable element" manner. In other words, the device can be queried, such as "Is the pinch true or false", and then return the state accordingly.

[0056] In some embodiments, gaze - tracking information can be used in combination with the embodiments disclosed herein. The gaze - tracking information, such as information 308, can at least include an indication of where the user is gazing. The information can include details related to the interactive element that the user is viewing, such as an identifier. For example, the information can be received from a head - mounted device such as an AR / MR / XR / VR headset. Devices, such as device 100, device 300, and device 700, can be configured to receive gaze - tracking data, based on which the device can use function visibility that matches the interactive type in gesture detection.

[0057] Figure 5 An exemplary process that can support at least some embodiments of the present invention is shown in the form of a flowchart. The process can be executed in whole or in part by any one of devices 100, 300, 700. At the Figure 5 top, an optical sensor (e.g., similar to sensor 105) and an IMU (e.g., similar to IMU 104) are shown. At least some of the operations and steps in the flowchart can be executed by a device, such as the devices disclosed herein, such as device 100. At least some operations, such as providing a scene, can be executed by a second device, such as a head - mounted device 21. At least some operations, such as providing communication, can be executed by a third device, such as a mobile device including at least one processor.

[0058] Figure 5 The flowchart of

[0059] Step 901 includes data acquisition. In this step, the controller (e.g., controller 103) receives a data stream from one or both of the IMU 104 and the optical sensor 105.

[0060] As shown in the flowchart, after obtaining the samples (sensor data), the acquired sensor data is preprocessed. Step 902 includes preprocessing. Regarding the IMU, the IMU can be adjusted before or after preprocessing. Such adjustment can be a self-referenced adjustment, for example, adjusting or calibrating the sensor according to the stored reference values or known parameters (such as the magnetic north pole of the IMU).

[0061] As can be seen from the flowchart, after preprocessing, the preprocessed data is input into the gesture classifier in step 903. In the flowchart, the preprocessed IMU data can be directly used, for example, to calculate the orientation of the IMU, and thus calculate the orientation of the user's hand, without the need for a neural network.

[0062] In addition, in step 903, the model outputs an inference output corresponding to the gesture state and thus to the user operation. This output, i.e., the inferred gesture state, can be provided, for example, as a confidence value.

[0063] In step 904, the inferred gesture state is provided to the interaction and / or state interpreter. The interpreter determines the user input and then provides it to the controlled scenario (application), for example, in the form of a boolean value and / or an HID frame, etc. The application can run on other devices or network devices.

[0064] Optional context information reflecting the application state (scenario state) can also be provided to the interpreter. Based on this context information, the interpreter can use a functional visibility profile when interpreting the received user actions. For example, if the context information indicates that the user is interacting with a flat interactive element 34, the interpreter may expect a pinch action to be performed and thus lower the threshold required for the pinch action.

[0065] In step 905, the scenario receives the interaction data. The scenario is updated accordingly. The interpreter can be configured to provide the interaction data in the form of events and / or pollable elements, and the scenario can query the event status based on this. The scenario can provide intention estimation data. The intention estimation can include an indication of future actions and / or information related to a set of appropriate actions for the context information. For example, the scenario can provide information related to previous operations (such as "play video") and expected future actions (such as "pause video", "adjust volume"). The scenario can provide a functional visibility profile to the interpreter. As mentioned above, when the device and / or scenario can predict the next action, this can improve the accuracy of detecting user actions.

[0066] Figure 6 FIG. 300 shows an apparatus 300 that can support at least some embodiments of the present invention. Unless otherwise specified, the apparatus is the same as apparatus 100.

[0067] In Figure 6 FIG. shows an exemplary schematic diagram of apparatus 300. Apparatus 300 includes a wrist-worn IMU 304, an optical sensor 305, and a controller 392. Apparatus 300 may be similar to apparatus 100 and may include mounting components for wearing by a user. Controller 392 is shown as a dashed line in the figure. As shown, controller 392 may at least include a model 380, for example, in the controller memory. Controller 392 preferably includes an interaction / status interpreter 390.

[0068] Sensors 305 and 304 are respectively configured to transmit sensor data streams, such as sensor data streams 313 and 314. The sensor data streams may include multiple channels, as described herein for signals 201, 202, and 203. The sensor data streams may be received by controller 392. Controller 392 may also receive other sensor data streams, such as bioacoustic sensor data streams.

[0069] In controller 392, the received sensor data streams are preprocessed in blocks 330 and 350. The preprocessed data streams may be directed to at least one model 380. Model 380 may be a gesture classifier, such as a neural network gesture classifier.

[0070] A model, such as model 380, may include a neural network. The neural network may be, for example, a feedforward neural network, a convolutional neural network, a recurrent neural network, or a graph neural network. The neural network may include a classifier and / or regression analysis. The neural network may apply a supervised learning algorithm. In supervised learning, input samples with known outputs are used, and the network learns from them and generalizes. Alternatively, unsupervised learning or reinforcement learning algorithms may also be used to build the model. In some embodiments, the neural network is trained such that certain signal features correspond to specific gesture transitions. For example, as Figure 2 shown, the neural network can be trained to associate specific signals, changes in signals, or changes in multiple signals with gesture transitions (such as a pinching or non-pinching action). The neural network can be trained to provide a confidence level for the output user actions.

[0071] The model may include at least one of an algorithm, a heuristic, and / or a mathematical model. A sensor fusion method may be employed to reduce the impact of the inherent non-ideal characteristics of a specific sensor, such as drift or noise. Sensor fusion may use a complementary filter, a Kalman filter, or a Mahony & Madgwick orientation filter. Such a model may also include the neural network disclosed herein.

[0072] For example, the model may include at least one convolutional neural network (CNN) that infers on input data (such as signals received from sensors and / or pre - processed data). The convolution can be performed in spatial or temporal dimensions. Features (calculated from sensor data fed into the CNN) can be selected algorithmically or manually. The model may also include an RNN (recurrent neural network), which can be used in combination with the neural network to support the identification of user action features based on sensor data reflecting user activities.

[0073] According to the present disclosure, the training of the model can be performed, for example, using a labeled data set containing multimodal biometric data from multiple subjects. Synthetic data can be used to augment and expand the data set. According to the model construction techniques employed, the sequence of computational operations constituting the model can be derived by backpropagation, Markov decision process, Monte Carlo methods, or other statistical methods. Model construction may involve dimensionality reduction and clustering techniques.

[0074] The above - mentioned models 380 are all configured to output a confidence level of a user action based on the received sensor data stream. The confidence level may include one or more indication probabilities in the form of percentage values associated with one or more respective user actions. The confidence level may include the output layer of the neural network. The confidence level can be represented in vector form. Such a vector can be a 1 x n vector, which contains the probability values of each user action known to the controller. For example, the model 380 receiving the sensor data stream may receive such a data stream that contains a signal, and the model 380 interprets the signal as a "pinch" gesture transition (the user has performed a pinch) with a confidence level of 90%.

[0075] The gesture classifier 380 outputs at least one determined gesture state and / or gesture transition, or a confidence level associated with the determined state and / or transition, which is received by the interaction / state interpreter 390. The interaction / state interpreter 390 is configured to determine a user input corresponding to the at least one received determined gesture state and / or transition and / or the associated confidence level based at least in part on the at least one received gesture state and / or transition. For example, if the model shows a high probability of a "pinch" gesture transition occurring, the interaction / state interpreter can infer that a "pinch" user input has occurred.

[0076] The interaction / state interpreter 390 may include a list and / or set of user actions, gesture states, and corresponding confidence levels. For example, when a confidence value at least matching the required threshold level has been received from the gesture classifier, a user input corresponding to a gesture state and / or transition is determined. Additionally, the device can be configured to adjust the threshold level, for example, based on the received function visibility profile 311 and / or intent 310.

[0077] The Interaction / State Interpreter 390 may receive optional gaze tracking information 308. The gaze tracking information includes at least an indication of where the user is gazing. The information may include details related to the interactive element the user is gazing at, such as an identifier. For example, the information may be received from a head-mounted device (such as an AR / MR / XR / VR head-mounted device). The information may be received as a data stream, which may include coordinates.

[0078] In case of a conflict, the Interaction / State Interpreter (ArgMax function) may, for example, select the user action indicating the highest probability. The Interaction / State Interpreter 390 may include, for example, at least one of the following: a rule-based system, an inference engine, a knowledge base, a lookup table, a prediction algorithm, a decision tree, a heuristic. For example, the Interaction / State Interpreter may include an inference engine and a knowledge base. The inference engine outputs a set of probabilities. These sets of probabilities are processed by the Interaction / State Interpreter, which may include IF-THEN statements.

[0079] The intent data 310 may include the application and / or the application state, context, and / or game state of the device being interacted with. The application state corresponds to the state of the application running on the controllable device and may include, for example, at least one of variables, static variables, objects, registers, open file descriptors, open network sockets, and / or kernel buffers.

[0080] The Interaction / State Interpreter may use the above data to adjust the threshold level of the user action. For example, if the application prompts the user to "pinch" to confirm, the corresponding "pinch" threshold may be lowered so that even if the confidence level of the pinch action output by the model is low, the pinch action can be detected.

[0081] An example of the operation of the device 300 is as follows. The user wears the device 300 on the hand. The user performs an action that includes pressing the index finger against the thumb. The sensors of the device 300 provide the following data during this action:

[0082] - The IMU sensor 304 provides a data stream 314 that reflects the orientation and acceleration data of the IMU (and the device 300).

[0083] - The optical sensor 305 provides a data stream 313 that includes, for example, data from a PPG sensor configured to measure the amount of reflected light from the user's skin, which corresponds to the movement of tendons and / or muscles in the user's wrist area.

[0084] Data streams 314 and 313 are received by model 380. Data streams 314 and 313 may correspond to signals 201, 202, and 203. In model 380, IMU sensor data and optical sensor data are used by the model, and the model detects a gesture transition from "no pinch" to "pinch". Based on the detected gesture transition, a "pinch" gesture state with a confidence level of 68% is provided to interaction / state interpreter 390.

[0085] Thus, event interpreter 390 receives the gesture state "pinch" and the associated confidence level. The interaction / state interpreter determines that a threshold for the user action has been reached, optionally with reference to the received function visibility profile 311 and intent data 310, e.g., adjusting the threshold based on at least one of the function visibility profile 311 and / or intent data 310 during this process. Additionally, interpreter 390 may receive optional gaze tracking information 308. In this way, the interaction / state interpreter 390 agrees to the "pinch" state.

[0086] State 309 is then transmitted to an application, such as a scenario executed by device 360, such as scenario 351, and the device may be, for example, a personal computer or an HMD, etc. In this example, device 360 receives the "pinch" state instead of any sensor data, thus saving the resources of device 360.

[0087] Figure 7 Device 700 is shown that can support at least some embodiments of the present invention.

[0088] Device 700 includes controller 702. The controller includes at least one processor, and at least one memory including computer program code and optional data. Device 700 may further include a communication unit or interface. For example, such a unit may include a wireless and / or wired transceiver. Device 700 may also include sensors, such as sensors 703, 704, which are operatively connected to the controller. The sensors may include, for example, any one of IMU 703, optical sensor 704. Device 700 may also include Figure 7 other elements not shown.

[0089] Although device 700 is described as including one processor, device 700 may include more processors. In one embodiment, the memory is capable of storing instructions, such as at least one of an operating system, various applications, models, neural networks, and / or preprocessing sequences. Additionally, the memory may further include storage space, such as for storing at least some of the information and data used in the disclosed embodiments.

[0090] In addition, the processor is capable of executing the stored instructions. In one embodiment, the processor may be embodied as a multi-core processor, a single-core processor, or a combination of one or more multi-core processors and one or more single-core processors. For example, the processor may be embodied as one or more of various processor devices, such as processor cores, coprocessors, microprocessors, controllers, digital signal processors (DSPs), processing circuits with or without DSPs, or various other processor devices, including integrated circuits, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontroller units (MCUs), hardware accelerators, dedicated computer chips, etc. In one embodiment, the processor may be configured to execute hard-coded functions. In one embodiment, the processor is embodied as an executor of software instructions, where the instructions may be specifically configured such that when executed, the processor can execute at least one of the models, sequences, algorithms, and / or operations described herein.

[0091] The memory may be embodied as one or more volatile storage devices, one or more non-volatile storage devices, and / or a combination of one or more volatile storage devices and non-volatile storage devices. For example, the memory may be embodied as semiconductor memory (such as masked ROM, PROM, EPROM, flash ROM, RAM, etc.).

[0092] At least one memory and computer program code may be configured to, together with at least one processor, cause the device 700 to at least perform the following operations: receive a data stream from at least one optical sensor 105 and at least one IMU 104, where the sensors are configured to measure a user, and / or the device 700 includes sensors 104 and 105, where the device 700 is configured to be worn by a user, provide the received data stream to the gesture classifier 380, and use the gesture classifier 380 to detect and classify gesture transitions and determine a gesture state based on transients within the sensor data in the received data stream, the gesture classifier 380 including a neural network classifier trained to detect the gesture transitions, where the data stream is optionally preprocessed before being provided to the gesture classifier. This has been discussed in detail above in connection with FIGS. 1 to Figure 5 This has been discussed in detail above.

[0093] The device 700 may be configured to provide the determined gesture state to an interaction and / or state interpreter, which is configured to perform corrections based on the application state and / or context. The interaction / state interpreter may be provided by the controller of the device 700.

[0094] The device 700 may be configured to predict a state transition based on the detected transition and by detecting the stateliness of the gesture, and infer a pinch state based on the obtained data.

[0095] Device 700 can be configured to cause an interaction and / or state interpreter to be configured to provide a detected gesture state to at least one of XR / VR / MR / AR applications, such as in the case of executing an application on a computing device. The provided gesture state can be provided as an event. As an alternative or addition, the provided gesture state can be provided as a pollable element, such that the event state of the device can be queried. Device 700 can be configured to receive at least one of the following from an XR / VR / MR / AR application or a computing device: an intent estimate or a functional visibility profile. Device 700 can be configured to use an adaptive time window as part of the detection, where the adaptive time window is tailored to a typical gesture duration, such as 1 millisecond.

[0096] Device 700 can be configured to receive context information, such as calibration and / or context cues, from an XR / VR / MR / AR application or a computing device, where the device is configured to use the received context information to adjust stateful detection. For example, the context information can indicate a "drop zone" in a scene. When a user moves something to the drop zone, a drop can be initiated, for example, by an interpreter even if the confidence of the pinch-no pinch transition (pinch-release) is not higher than a global threshold. Device 700 can be configured to apply dead reckoning calibration and context cues to improve stateful detection. Device 700 can be configured to use a recurrent model to capture gesture history, such as user-specific gesture history, to improve stateful detection. Such capture can include information about previous user actions and / or previous interaction elements, for example.

[0097] Device 700 can be configured to convert discrete time events, such as taps and releases, into states, such as pinch and no pinch, for detecting transitions, such as the touch and release of an index finger and a thumb.

[0098] Device 700 can be configured to perform on-board processing, which includes at least one of preprocessing, gesture classification, or state interpretation.

[0099] Device 700 can be configured to perform a method for identifying a selection gesture from acquired data, the method including: acquiring 901 data from at least one optical sensor 105 and at least one IMU 104, and providing the data to a 903 gesture classifier, where the gesture classifier is configured to detect whether a hand is in a gesture state, such as a pinch state, and where the data can optionally be preprocessed before being provided to the classifier.

[0100] The method can include using a neural network, which is configured to detect a transition between gesture states by detecting the rising and falling edges of a gesture. The method can include detecting a transition between a pinch state and a no pinch state. The method can include performing on-board processing for any of the steps, such as preprocessing, gesture classification, transition detection steps.

[0101] The present disclosure is also reflected in the following clauses.

[0102] Clause 1, A multimodal biometric measurement device, comprising a processor, the processor including at least one processor, and at least one memory including computer program code, the at least one memory and the computer program code configured to, together with the at least one processor, cause the processor to at least: receive data streams from at least one optical sensor and at least one IMU, wherein the sensors are configured to measure a user; provide the received data streams to a gesture classifier; and use the gesture classifier to determine a gesture state, wherein gesture transitions are detected and classified based on transients in the sensor data of the received data streams, the gesture classifier including a neural network classifier trained to detect the gesture transitions. Wherein, the data streams may be selectively preprocessed before being provided to the gesture classifier.

[0103] Clause 2, The device according to clause 1, further comprising a strap and a housing, the housing including the optical sensor and the IMU, the optical sensor arranged in the housing so as to measure the wrist area of the user when the device is worn by the user.

[0104] Clause 3, The device according to clause 1 or 2, including a smartwatch.

[0105] Clause 4, The device according to any one of clauses 1 to 3, wherein the optical sensor includes a PPG sensor.

[0106] Clause 5, The device according to any one of clauses 1 to 4, wherein the controller is arranged within the housing.

[0107] Clause 6, A kit, including at least one wrist-worn device, the kit configured to: receive data streams from at least one optical sensor and at least one IMU, wherein the sensors are configured to measure a user; provide the received data streams to a gesture classifier; and use the gesture classifier to determine a gesture state, wherein gesture transitions are detected and classified based on transients in the sensor data of the received data streams, the gesture classifier including a neural network classifier trained to detect the gesture transitions. Wherein, the data streams may be selectively preprocessed before being provided to the gesture classifier.

[0108] Clause 7, The kit according to clause 6, including a head-mounted device having a display, the head-mounted device connected to the device and configured to use the display to indicate the gesture state and / or gesture transitions.

[0109] Clause 8. The kit according to Clause 6 or 7, wherein the wrist-worn device includes a display screen, which is connected to a controller and configured to use the display screen to indicate a gesture state and / or a gesture transition.

[0110] Clause 9. The kit according to Clause 8, wherein the head-mounted device is configured to provide context information to the wrist-worn device, and / or the head-mounted device is configured to interpret and apply a confidence value of a gesture state provided by the wrist-worn device.

[0111] Clause 10. The device or kit according to the above clauses, wherein the wrist-worn device includes a neural network classifier, which is trained to detect a gesture state based on sensor data.

[0112] Advantages of the present disclosure include the following. By using IMU and optical sensor data obtained from wearable sensors (such as wrist-worn sensors), common problems in VR / XR / AR interactions, such as the field of view or occlusion problems, can be overcome.

[0113] The disclosed embodiments provide a technical solution to solve a technical problem. One technical problem to be solved is to detect a pinch or tap gesture using a device that processes IMU and optical sensor data on board. In fact, this is problematic because the data may contain different information related to the user's hand movements.

[0114] Embodiments of the present invention overcome these limitations by using machine learning to detect a gesture state based on transients in sensor signals. In this way, on-board gesture detection can be completed in a more accurate and robust manner. There are several advantages to doing so. First, the processing can be done near the sensor, so communication latency can be minimized. Second, interpreter correction based on the application state, etc., can also be done near the sensor. Third, compared with camera-based systems, using selected sensors allows the user to carefully provide user interface commands with minimal impact on the hand position. These embodiments can also bring other technical improvements and solve other technical problems.

[0115] It should be understood that the embodiments disclosed in the present invention are not limited to the specific structures, process steps or materials described herein, but can extend to equivalents recognized by those of ordinary skill in the art. It should also be understood that the terms used herein are only for describing specific embodiments and are not intended to be limiting.

[0116] As used herein, "an embodiment" or "embodiment" means that the specific features, structures, or characteristics described in connection with that embodiment are included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment" or "in an embodiment" that appear in various places herein are not necessarily all referring to the same embodiment. When a numerical value is referred to using, for example, "about" or "substantially", the exact numerical value is also disclosed.

[0117] In this document, a plurality of objects, structural elements, components, and / or materials may be presented in a common list for convenience. However, these lists should be regarded as each element in the list being able to be treated independently as a separate and distinct element. Thus, each element in such a list should not be regarded as a practical equivalent of any other element in the same list merely based on the fact that the elements appear in a common group without conflicting statements. Additionally, various embodiments and examples of the present invention may also relate to alternatives of the various components. It should be understood that these embodiments, examples, and alternatives are not to be regarded as practical equivalents of each other, but rather as separate and autonomous manifestations of the present invention.

[0118] Furthermore, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In this document, many specific details (such as examples of length, width, shape, etc.) are presented to facilitate an understanding of the embodiments of the present invention. However, those skilled in the art should understand that the present invention may be practiced without one or more of the specific details, or with other methods, components, materials, etc. Additionally, known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the present invention.

[0119] Although the above examples illustrate the principles of the present invention in one or more specific applications, those of ordinary skill in the art should understand that many modifications may be made to the form, use, and details of the implementation without departing from the principles and concepts of the present invention without creative effort. Therefore, the present invention is intended to be limited only by the appended claims.

[0120] The verbs "comprise" and "include" are used herein as open-ended limitations, neither excluding nor necessarily having the presence of other unlisted features. Unless otherwise specified, the features recited in the dependent claims can be freely combined with each other. Additionally, it should be understood that the use of "a" or "an" (i.e., the singular form) herein does not exclude the plural form.

[0121] Industrial Applicability

[0122] At least some embodiments of the present invention are industrially applicable to provide a user interface for a controllable device, such as a personal computer or an HMD, for example, a user interface related to XR.

[0123] List of reference numerals

[0124] 100,300,700,360 Device 101 Mounting component 102 Housing 103,392,702 Controller 104,304,703 Inertial measurement unit (IMU) 105,305,704 Optical sensor 210 Wrist measurement area 20 User 201A - B, 202A - C, 203A - B Sensor signal 308 Gaze tracking information 309 Status information 310 Intention data 311 Function visibility profile 313,314 Sensor data stream 330,350 Pre - processing block 351 Scene 380 Model 390 Interpreter 901,902,903,904,905 Steps of the method

Claims

1. A wrist-worn device (100), comprising at least one processing core (701), and at least one memory (707) comprising computer program code, wherein the at least one memory (707) and the computer program code are configured to, together with the at least one processing core (701), cause the device (100) to at least: receiving data streams from at least one optical sensor (105) and at least one IMU (104), wherein the sensors are configured to take measurements of a user, providing the received data stream to a gesture classifier (380), and The gesture classifier (380) is used to determine a gesture state, wherein: detecting and classifying gesture transitions based on transients in sensor data of the received data stream, the gesture classifier (380) comprising a neural network classifier trained to detect the gesture transitions, Before providing the data stream to the gesture classifier, the data stream may be selectively preprocessed.

2. The device according to claim 1, wherein: The apparatus is further configured to provide the determined gesture state to an interaction and / or state interpreter configured to make corrections based on the application state and / or context.

3. The device according to any one of the preceding claims, wherein: The apparatus is further configured to infer a pinch state from the received data stream based on the determined gesture state.

4. A device according to any one of the preceding claims, wherein: The apparatus is further configured such that the interaction and / or state interpreter is configured to provide the determined gesture state to an extended reality application.

5. A device according to any one of the preceding claims, wherein: The apparatus is further configured to receive at least one of the following from the extended reality application: an intent estimate or an affordance profile.

6. A device according to any one of the preceding claims, wherein: The apparatus is further configured to use an adaptive time window as part of the detection, wherein the adaptive time window is specifically tuned for a typical gesture duration, such as 1 millisecond.

7. A device according to any one of the preceding claims, wherein: The apparatus is further configured to receive contextual information, such as corrections and / or contextual clues, from the extended reality application, and the apparatus is configured to use the received contextual information to adjust stateful detection.

8. A device according to any one of the preceding claims, wherein: The apparatus is further configured to receive gaze tracking information, such as information about a user's gaze location and nearby interactable elements, wherein the apparatus is configured to use the received gaze tracking information to adjust stateful detection.

9. The device according to any one of the preceding claims, wherein: The apparatus is further configured to apply dead reckoning corrections and contextual clues to improve statefulness detection.

10. The device according to any one of the preceding claims, wherein: The apparatus is also configured to use a recurrent model to capture gesture history and improve stateful detection.

11. The device according to any one of the preceding claims, wherein: The apparatus is further configured such that the value of the optical sensor at the beginning and at the end of the gesture session is higher than an average value for the gesture session.

12. The device according to any one of the preceding claims, wherein: The device is also configured to convert discrete time events, such as taps and releases, into states, such as pinched and unpinched, so as to detect transitions, such as touch and release of the index finger and thumb.

13. A device according to any one of the preceding claims, wherein: The device includes an IMU (104) and an optical sensor (105).

14. The device according to claim 1, wherein: The device is also configured to perform onboard processing, which includes pre-processing, gesture classification, and state interpretation.

15. A method for identifying a selection gesture from acquired data, the method comprising: receiving (901) data streams from at least one optical sensor (105) and at least one IMU (104), wherein the sensors are configured to measure a user, providing (903) the received data stream to a gesture classifier (380), and determining a gesture state using the gesture classifier (380), wherein gesture transitions are detected and classified based on transients in sensor data of the received data stream, the gesture classifier (380) comprising a neural network classifier trained to detect the gesture transitions, Before providing the data stream to the gesture classifier, the data stream may be selectively preprocessed.

16. The method according to claim 15, wherein: The determined gesture state is provided to an interaction and / or state interpreter configured to make corrections based on the application state and / or context.

17. The method according to claim 15 or 16, wherein: The method also includes detecting a transition between the pinched state and the un-pinched state.

18. The method according to any one of claims 15 to 17, wherein: The method also includes performing onboard processing, such as the steps of pre-processing (902), gesture classification (903), and state interpretation.

19. The method according to any one of claims 15 to 18, wherein: The method also includes receiving contextual information, such as corrections and / or contextual clues, from the extended reality application and using the received contextual information to adjust stateful detection.

20. A non-transitory computer-readable medium having stored thereon a set of computer-readable instructions that, when executed by at least one processor, enable an apparatus to at least: receiving data streams from at least one optical sensor (105) and at least one IMU (104), wherein the sensors are configured to take measurements of a user, providing the received data stream to a gesture classifier (380), and The gesture classifier (380) is used to determine a gesture state, wherein: detecting and classifying gesture transitions based on transients in sensor data of the received data stream, the gesture classifier (380) comprising a neural network classifier trained to detect the gesture transitions, Before providing the data stream to the gesture classifier, the data stream may be selectively preprocessed.