Detection of User Input from a Hand Multimodal Biometric Measurement Device

The multimodal biometric measurement device addresses the limitations of traditional XR controllers by directly digitizing hand movements and gestures, providing efficient and hands-free control of devices through integrated sensors and processing, enhancing interaction speed and usability.

JP7713484B2Active Publication Date: 2025-07-25DOUBLEPOINT TECH OY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023036208
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-02-17
Filing Date
2023-03-09
Publication Date
2025-07-25
Estimated Expiration
2043-03-09

AI Technical Summary

Technical Problem

Existing XR user interfaces rely on dedicated physical controllers that introduce a redundant technical layer between the user's hand and the computing device, limiting interaction speed and usability, and are often device-specific, cluttering the user's hand and restricting its normal use.

Method used

A multimodal biometric measurement device that includes sensors to directly digitize hand movements and gestures, utilizing a wrist contour sensor, bioacoustic sensor, and inertial measurement unit to generate user interface commands without interfering with hand use, incorporating a processor and memory to process sensor data streams and identify user actions.

Benefits of technology

Enables seamless and efficient control of devices by directly converting hand movements into machine commands, enhancing interaction speed and usability without device-specific limitations, allowing hands-free operation and improved human-machine interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713484000001
    Figure 0007713484000001
  • Figure 0007713484000002
    Figure 0007713484000002
  • Figure 0007713484000003
    Figure 0007713484000003
Patent Text Reader

Abstract

To provide: a biometric device that digitizes and transforms minute hand movements and gestures into user interface commands without interfering with the normal use of user's hands; and method for generating commands.SOLUTION: A multi-modal biometric device 100 includes a mounting component 101, a controller 102, and at least one inertial measurement device 103, wrist contour sensor 104, and bio-acoustic sensor 105. The controller receives each sensor data stream from each sensor, identifies at least one characteristic of a user action based on at least one of the sensor data streams, identifies at least one user action based on the characteristic, and generates at least one user interface command based on the user action.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various exemplary embodiments of the present disclosure relate to wearable devices and methods that can be used to control devices (especially in the fields of computing applications and extended reality user interface applications). Extended reality (XR) encompasses augmented reality (AR), virtual reality (VR), and mixed reality (MR).

Background Art

[0002] The control of digital devices has traditionally been performed with dedicated physical controllers. For example, a computer can be operated with a keyboard and a mouse, a game console can be operated with a handheld controller, and a smartphone can be operated with a touch screen. These physical controllers typically include sensors and / or buttons that receive input from the user based on the user's actions. Such discrete controllers are widespread, but human-machine interaction is slowed down because a redundant technical layer is added between the user's hand and the computing device. Furthermore, such dedicated devices are typically only suitable for controlling a specific device. Also, such devices can, for example, clutter the user's hand, so the user cannot use their hand for other purposes while using the control device.

Summary of the Invention

Problems to be Solved by the Invention

[0003] The present invention has been made to solve the problems in the above prior art.

Means for Solving the Problems

[0004] In view of the above problems, improvements in the field of XR user interfaces are needed. Appropriate input devices, such as the invention disclosed herein, directly digitize slight movements and gestures of the hand and convert them into machine commands without interfering with the normal use of the user's hand. Embodiments of the present disclosure can detect user actions based on, for example, user action characteristics detected from a plurality of sensors, and these sensors are preferably various types of sensors.

[0005] The present invention is defined by the features of the independent claims. Some specific embodiments are defined in the dependent claims.

[0006] According to a first aspect of the present invention, there is provided a multimodal biometric measurement device, the device including a mounting component configured to be worn by a user, at least one wrist contour sensor, at least one bioacoustic sensor including a vibration sensor, at least one inertial measurement unit (IMU) including an accelerometer and a gyroscope, and a controller including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to cause the at least one processor to receive a first sensor data stream from the at least one wrist contour sensor, receive a second sensor data stream from the at least one bioacoustic sensor, receive a third sensor data stream from the at least one inertial measurement unit (wherein the first, second, and third sensor data streams are received in parallel), identify at least one feature of a user action based on at least any one of the first, second, and third sensor data streams, identify at least one user action based on the identified at least one feature of the user action, and generate at least one user interface (UI) command based at least to some extent on the identified at least one user action.

[0007] According to a second aspect of the present invention, a method for generating a UI command is provided. The method includes receiving a first sensor data stream from at least one wrist contour sensor, receiving a second sensor data stream from at least one bioacoustic sensor, receiving a third sensor data stream from at least one inertial measurement device (wherein the first, second, and third sensor data streams are received in parallel), identifying at least one characteristic of a user action based on at least any one of the first, second, and third sensor data streams, identifying at least one user action based on the identified at least one characteristic of the user action, and generating at least one user interface (UI) command based at least to some extent on the identified at least one user action.

[0008] According to a third aspect of the present invention, a non-transitory computer-readable medium storing a set of computer-readable instructions is provided. When executed by a processor, this set of instructions causes the second aspect to be implemented or causes a device including the processor to be configured according to the first aspect.

Brief Description of the Drawings

[0009]

Figure 1A

Figure 1B

Figure 2A

Figure 2D

Figure 2E

Figure 2G

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 9

Figure 10A

Figure 10B

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0010] This specification describes a multimodal biometric measurement device and a corresponding method. The device and / or the method may be used for at least any one of (for example) measurement, detection, signal acquisition, analysis, user interface tasks. Such a device is preferably suitable for the user to wear. The user can control one or more devices using this device. The control may be in the form of a user interface (UI) or a human machine interface (HMI). The device may include one or more sensors. The device may be configured to generate user interface data based on data from the one or more sensors. The user interface data may be used to enable the user to at least to some extent control this device or at least one second device. The second device may be, for example, at least one of a personal computer (PC), a server, a mobile phone, a smartphone, a tablet device, a smartwatch, or any suitable type of electronic device. The device to be controlled or controllable may implement at least any one of an application, a game, and / or an operating system, all of which may be controlled by the multimodal device.

[0011] The user may perform at least one user action. Typically, the user performs an action acting on a controllable device. The user action may be, for example, at least any one of movement, gesture, interaction with an object, interaction with a part of the user's body, and null action. An example of a user action is a "pinch" gesture, which is an action where the user touches the tip of the index finger with the tip of the thumb. Another example is a "thumbs up" gesture, which is an action where the user extends the thumb, bends the other fingers, and turns the hand so that the thumb points upward. The present embodiment is configured to identify (specify) at least one user action based at least to some extent on sensor input. If the reliability of the identification of the user action is high, control based on the action becomes possible, for example, control based on the action using the embodiments disclosed herein becomes possible.

[0012] Null action may be understood as being at least one of no user action and / or no user action having an effect on the application state of the device or application under control. In other words, null action corresponds to an entry of "none" in at least any one of a list of predetermined user actions, a list of user interface commands, and / or a set of classified gestures. The device may be configured such that the set of possibilities not only has a possibility / reliability value corresponding to each gesture in the set, but also includes a possibility assigned to "none". For example, when an application prompts the user to "pinch if accepting a selection" and the user performs normal hand movements that do not include a pinch gesture, the device may regard these movements as null actions. In such a situation, the device may be configured not to send anything to the application, for example. In another example, when the user is scrolling a list and performs an action not related to that list, the performed action may be identified as a null action. The event interpreter of the device is configured to identify and / or manage null actions, and this identification and / or management is optionally performed at least to some extent based on context information received from the application.

[0013] A user action may include at least one feature called a user action feature and / or an event. Such features may be, for example, features of any of time, position, space, physiology, and / or motion. The feature may include an indication of a part of the user's body (e.g., the user's finger). The feature may include an indication of the user's movement. For example, the feature may be the movement of the middle finger. In another example, the feature may be a circular movement. In yet another example, the feature may be the time to an action and / or the time from an action. In at least some embodiments, the features of a user action are determined based at least to some extent on sensor data, for example, by a neural network.

[0014] In at least some embodiments, user interface (UI) commands are generated based on user actions. The UI commands may be referred to as, for example, human machine interface (HMI) commands, user interface interactions, and / or HMI interactions. Such commands can be used to convey user interface actions. For example, a certain user interface action may correspond to the "enter" key on a personal computer. In at least some embodiments, an output generator, such as a user interface command generator, maps a user action to a predetermined user interface action and / or generates a user interface command based on a predetermined user interface action. Thus, the device to be controlled does not need to understand the user action, and only needs to report the descriptor for the UI command that can be provided in a standard format such as a human interface device (HID). The UI command may include, for example, a data packet and / or a frame including UI instructions.

[0015] In at least some embodiments, the device includes a mounting component configured to be worn by a user. Thus, the mounting component is a wearable component. Such a component may be any of a strap, a band, a wristband, a bracelet, or a glove. The mounting component may be attached to another device such as a smartwatch and / or may be formed by such a device, or may form part of a larger device such as a gauntlet. In some embodiments, the strap, band, and / or wristband has a width of 2 to 5 cm, preferably 3 to 4 cm, and more preferably 3 cm.

[0016] In this embodiment, a signal is measured by at least one sensor. This sensor may include a processor and a memory, or may be connected to a processor and a memory, and the processor may be configured such that the measurement is performed. The measurement may be referred to as detection or sensing. The measurement includes, for example, detecting a change in at least one of a user, an environment, and the physical world. The measurement may further include, for example, attaching at least one timestamp to the sensed data and transmitting the sensed data (with the at least one timestamp or without a timestamp).

[0017] At least some embodiments are configured to measure the contour of a user's wrist. The contour of the wrist is the three-dimensional surface contour around the wrist. This is defined by at least any one of a point cloud, mesh, or other surface representation of the topology of the outer tissue adjacent to and / or proximate to the radiocarpal joint, ulnar carpal joint, and distal radioulnar joint. In practice, for example, the contour of the wrist is the outer shape of the wrist. The contour of the wrist may be measured from the user's wrist region. The wrist region may include, for example, the region from the lower end of the user's palm to a few centimeters below the lower end of the user's palm. When the user performs an action with the hand connected to the wrist, the surface of the wrist deforms in response to the movement of the tendons and / or muscles required for that action. In other words, the contour of the wrist changes in a measurable manner, and the state of the contour of the wrist corresponds to the user action characteristics. Measurements may be continuously acquired (e.g., from the wrist region) to provide sensor data that can be used in this embodiment. The measurement of the contour of the wrist may be performed, for example, using a matrix of sensors, and the measurement may be, for example, optical, capacitive, or piezoelectric proximity sensing or contour measurement sensing. Such a matrix may be referred to as a wrist contour sensor. Using multiple sensors, particularly in matrix form, has the advantage that the exact state of the user's wrist can be known with sufficient high resolution. Each sensor can provide at least one data channel, and in a preferred embodiment, the wrist contour sensor provides at least 32 channels. The matrix may be placed on the user's wrist, particularly on the palm side of the wrist. Placing it in such a way may be implemented, for example, using a strap, wristband, or glove.

[0018] At least some embodiments are configured to measure vibration signals and / or acoustic signals. Such signals may be measured, for example, from a user, particularly from the user's wrist area. A bioacoustic sensor may be configured to measure vibrations such as internal waves of the user and / or surface waves of the user's skin. Such vibrations may be, for example, the result of a contact event. A contact event may be the user contacting an object and / or the user contacting themselves and / or another user. The bioacoustic sensor may include a contact microphone sensor (piezo contact microphone), a MEMS (Micro-Electro-Mechanical System) microphone, and / or a vibration sensor. The vibration sensor may be, for example, an accelerometer that detects the acceleration generated by a vibrating structure (such as a user or a mounting component).

[0019] At least some embodiments are configured to use an Inertial Measurement Unit (IMU) to measure the position and movement of a user. The IMU may be configured to provide position information of the device. The IMU may be referred to as an inertial measurement unit sensor. The IMU may be at least any one of a multi-axis accelerometer, a gyroscope, a magnetometer, an altimeter, and a barometer. The IMU preferably includes a magnetometer. This is because the magnetometer provides an absolute reference for the IMU. The barometer can be used as an altimeter and can add degrees of freedom to the IMU.

[0020] Sensors such as inertial measurement unit sensors, bioacoustic sensors, and / or wrist profile sensors may be configured to output a sensor data stream. This output may be performed, for example, within the device to a controller, a digital signal processor (DSP), and a memory. Alternatively or additionally, this output may be performed to an external device. The sensor data stream may include one or more raw signals measured by the sensor. Further, the sensor data stream may include at least any one of synchronization data, configuration data, and / or identification data. Such data may be used, for example, by a controller to compare data from various sensors.

[0021] In at least some embodiments, the apparatus includes means for providing kinesthetic feedback and / or tactile feedback. Such feedback may include, for example, vibration and may also be referred to as "haptic feedback". The means for providing haptic feedback may include, for example, at least any one of an actuator, a motor, and a haptic output device. A preferred actuator that can be used in all embodiments is a linear resonant actuator (LRA). Haptic feedback can be useful for improving the usability of the user interface by providing feedback to the user from the user interface. Further, haptic feedback can alert the user to events and situation changes in the user interface.

[0022] Figures 1A and 1B show an apparatus 100 according to at least some embodiments of the present invention. The apparatus 100 may be suitable for detecting a user's movement and / or action and operating as a user interface device and / or an HMI device. The apparatus 100 includes a mounting component 101, which is shown in the form of a strap in Figure 1A.

[0023] Device 100 further includes a controller 102. The controller 102 includes at least one processor and at least one memory including computer program code. The controller may be configured to receive at least parallel sensor data streams from sensors 103, 104, and 105. The controller may be configured to perform processing on at least any of those sensor data streams, and this processing may include preprocessing disclosed elsewhere in this document. The controller may be configured to process the sensor data streams received from the sensors. The controller may be configured to generate at least one user interface (UI) event and / or command based on the characteristics of the processed sensor data streams. Further, the controller may include a model, and the model may include at least one neural network.

[0024] Device 100 further includes an inertial measurement unit (IMU) 103. The IMU 103 is configured to transmit sensor data streams (e.g., to the controller 102). These sensor data streams include, for example, at least any of multi-axis accelerometer data, gyroscope data, and / or magnetometer data. One or more of the components of device 100 may be combined together, for example, the IMU 103 and the controller 102 may be on the same printed circuit board (PCB). Changes in the IMU data reflect, for example, movement and / or actions by the user, and thus such movement and / or actions can be detected using the IMU data.

[0025] Device 100 further includes a bioacoustic sensor 105, which is shown in the form of a contact microphone in FIG. 1A. The sensor 105 is configured to transmit a sensor data stream (e.g., to the controller 102). This sensor data stream includes, for example, a single-channel audio waveform or a multi-channel audio waveform. The arrangement of the sensor 105 with respect to the mounting component may be optimized with respect to signal quality (reduction of acoustic impedance) and user comfort. The sensor 105 is configured such that contact with the user is stable. This is because large artifacts may occur in the signal if the contact is not stable. Changes in the bioacoustic data reflect, for example, movement and / or actions by the user, and thus such movement and / or actions can be detected using the bioacoustic data.

[0026] At least one microphone may be mounted on the back of the mounting component 101. Other sensors (such as an IMU) may be mounted on the side of the mounting component 101 opposite the microphone. For example, the microphone 105 may be mounted on the inner circumference of the mounting component 101, and the IMU may be mounted on the outer circumference of the mounting component 101. Further, a wrist contour sensor may also be mounted on the inner circumference of the mounting component 101. At least one microphone may be configured to contact the skin when the user wears the device.

[0027] The device 100 further includes a wrist contour sensor 104, which includes a matrix of contour measurement sensors in FIG. 1A. These sensors may be, for example, proximity sensors, optical sensors. The contour measurement sensors may be configured to measure the distance from the sensor to the user's skin. In other words, each contour measurement sensor is configured to detect whether it is close to the surface of the user's wrist. The actions performed by the user affect the wrist contour in a predictable manner, which affects the distance to each contour measurement sensor. Therefore, it is possible to detect the user's wrist contour using data from multiple contour measurement sensors. Since the state of the wrist contour corresponds to an action (e.g., movement of a specific finger), such an action can be detected based on the wrist contour sensor data. Furthermore, changes in the wrist contour can also be detected. The sensor 104 is configured to transmit a sensor data stream (e.g., to the controller 102). This sensor data stream includes a set of contour measurement readings corresponding to a given contour at a given point in time. The set of readings may include, for example, a set of changes in the base current of a phototransistor, which are converted to voltage readings by a digital rheostat. This voltage may be sampled by an ADC (analog-to-digital converter) chip. A suitable sampling rate (interval) is 1 to 1000 ms, preferably 10 ms.

[0028] FIG. 1B shows an exemplary configuration of the device 100. The figure shows a mounting component 101 and a wrist contour sensor 104 including a matrix of sensors 106. In the figure, the sensor 106 is shown as an optical sensor. The matrix in FIG. 1B is an exemplary 3×8 matrix of 3 rows and 8 columns. For use in this device, 1 to 10 rows and 1 to 10 columns are suitable. The spacing between rows and columns of the matrix may be from about 2 millimeters to about 10 millimeters. Preferably, a matrix such as a 4×8 matrix may be used, particularly a 4×8 matrix with a 4-millimeter spacing on the palm side of the hand may be used. The sensors 106 may be distributed in a linear, radial, or asymmetric pattern with respect to each other.

[0029] For example, the sensor 106 may be one or more of an infrared-based proximity sensor, a laser-based proximity sensor, an electronic image sensor-based proximity sensor, an ultrasonic-based proximity sensor, a capacitive proximity sensor, a radar-based proximity sensor, an accelerometer-based proximity sensor, a piezoelectric-based proximity sensor, and / or a pressure-based proximity sensor. As an example, an infrared-based proximity sensor may include a pair of an infrared light-emitting diode (LED) operating as an infrared light emitter and a phototransistor operating as an infrared light receiver. Such an active infrared-based proximity sensor can indirectly detect the distance to an object in front at a short distance within several millimeters. This detection is performed by irradiating the front with infrared light radially by the LED and detecting by the phototransistor how much light is reflected back. In practice, the more light is reflected back, the more current the phototransistor allows to pass. This property of the phototransistor may be used in series with a resistor, thereby forming a voltage divider circuit that can be used, for example, to measure the incident light of the phototransistor with high precision. For example, the resolution of measuring the proximity of the skin of the wrist can be as fine as 0.1 millimeter. In at least some embodiments, the wrist contour sensor may include other types of sensors in addition to the infrared-based proximity sensor, for example, to enhance the robustness and / or accuracy of the infrared-based proximity sensor.

[0030] FIG. 2A shows an apparatus 200 that can support at least some embodiments of the present invention. The apparatus 200 is similar to the apparatus 100. The illustrated apparatus 200 is worn on the user's hand 299. The apparatus 200 is worn on the user's wrist by a mounting component 201. A protrusion 298 is shown on the band 201, and the protrusion 298 may house all the components of the apparatus 200 or may be attached to all its components.

[0031] FIG. 2B shows an exemplary measurement area of the device 200 and other devices disclosed herein. In FIG. 2B, the bioacoustic measurement area 295, the IMU 296, and the wrist contour measurement area 297 are shown in relation to the user's hand 299. The wrist contour measurement area 297 is shown in the form of a grid, and each point of the grid is the position of a contour measurement sensor (e.g., sensor 106), and the contour of the grid is the measured wrist contour.

[0032] FIGS. 2C and 2D show an exemplary wrist contour of a user. On the left side of the figure, the state where the user 299 is wearing the device 200 is shown. On the right side of the figure, the wrist contour 294 of the user 299 is shown. In FIG. 2C, the user's hand is in a first hand pose, which corresponds to the wrist contour shown in that figure. In FIG. 2D, the user's hand is in a second hand pose, which corresponds to the wrist contour shown in that figure. The difference between these poses corresponds to the difference seen in the wrist contours of FIGS. 2C and 2D.

[0033] FIG. 2E shows an exemplary topological sequence corresponding to the movement of the user's index finger based on the measured wrist contour data. This exemplary sequence includes five topologies based on measurement data acquired every second, which reflects the change in the contour corresponding to the movement of the user's finger. This exemplary sequence shows a sequential data set obtained from the wrist contour sensors and the change in the wrist contour determinable from those data sets. As seen in this figure, the data from each contour measurement sensor is represented in three-dimensional coordinates. These sensors together represent a set of coordinates, and connecting them forms a contour. By comparing the contour with the previous contour state, the deformation and change of the contour can be known. By delta encoding, it is possible to transmit only the changes from the previous data set. In this embodiment, the topology is not necessarily calculated. For example, the measured data points representing the contour (e.g., 32 points) may be input into a convolutional neural network as a sensor data set, and the convolutional neural network is trained to output user action features.

[0034] Figure 2F shows an exemplary data set of bioacoustic measurements. The transient characteristics of an exemplary signal are visualized. In this graph, the X-axis represents elapsed time and the Y-axis represents magnitude (e.g., voltage). The dotted box in the figure is shown only for visualization and highlights the region of interest. Transient characteristics are seen within this exemplary region of interest. Such characteristics may correspond to, for example, the user's pinch action. In some embodiments, the illustrated signal may be preprocessed, for example, by fast Fourier transform (FFT).

[0035] Figure 2G shows an exemplary data set of an IMU during typical hand movements. In this exemplary data set, the orientation data is encoded as quaternions. This data format consists of one scalar component and three vector components. The scalar component encodes the rotation angle, and the three vector components, called the fundamental quaternions, encode the direction of the rotation axis in three-dimensional space. Signal 801 corresponds to the coefficient of the scalar component. Signals 802, 803, 804 correspond to the coefficients of the three fundamental quaternions. In this graph, the X-axis is time and the Y-axis is the magnitude of the coefficient. In some embodiments, it is possible to use the wrist quaternion to calculate the three-dimensional rotation relative to the quaternion measured by another IMU (e.g., mounted on the user's head).

[0036] Figures 3A and 3B show apparatuses 300 and 400 that can support at least some embodiments of the present invention. These apparatuses are the same as apparatuses 100 and 200 unless otherwise specified.

[0037] FIG. 3A shows an exemplary schematic diagram of an apparatus 300. The apparatus 300 includes an IMU 303, a wrist contour sensor 304, a bioacoustic sensor 305, and a controller 302. The controller 302 is represented by a dashed line in the figure. As shown, the controller 302 may include at least models 350, 330, and 370 (e.g., within the memory of the controller). The controller 302 may include an event interpreter 380. The controller 302 may include an output generator 390.

[0038] Sensors 303, 304, and 305 are each configured to transmit a sensor data stream, e.g., configured to transmit sensor data streams 313, 314, and 315, respectively. The sensor data streams may be received by the controller 302.

[0039] Within the controller 302, the received sensor data stream is sent to at least one model (e.g., models 350, 330, 370). Models such as models 350, 330, 370 may include neural networks. The neural network may be, for example, a feedforward neural network, a convolutional neural network, a recurrent neural network, or a graph neural network. The neural network may include a classifier and / or regression. The neural network may apply a supervised learning algorithm. In supervised learning, samples of inputs with known outputs are used and the network learns and generalizes from them. Alternatively, the model may be constructed using an unsupervised learning algorithm or a reinforcement learning algorithm. In some embodiments, the neural network has been trained such that specific signal characteristics correspond to specific user action features. For example, the neural network may be trained to associate a specific wrist contour with user action features associated with a specific finger (e.g., the middle finger illustrated in FIGS. 2C-2E). The neural network may be trained to present a confidence level of the output user action features.

[0040] The model may include at least any one of an algorithm, a heuristic, and / or a mathematical model. For example, the IMU model 330 may include an algorithm for calculating orientation. The IMU may output data from a three-axis accelerometer, a gyroscope, and a magnetometer. A sensor fusion method may be employed to reduce the effects of the inherent non-ideal behavior (e.g., drift and noise) of a particular sensor. Sensor fusion may include the use of a complementary filter, a Kalman filter, or a Mahony & Madgwick orientation filter. Such a model may further include the neural networks disclosed herein.

[0041] The model may include, for example, at least one convolutional neural network (CNN) that performs inferences on input data (e.g., signals received from sensors and / or preprocessed data). The convolution may be performed in spatial or temporal dimensions. Features (calculated from sensor data fed into the CNN) may be selected algorithmically or manually. The model may include another RNN (recurrent neural network), which may be used in combination with the neural network to support the identification of user action features based on sensor data reflecting user activity.

[0042] According to the present disclosure, the training of the model may be performed, for example, using a labeled dataset that includes multimodal biometric measurement data from a plurality of objects. This dataset may be enhanced and augmented using synthetic data. The sequence of computer operations that make up the model may be derived by backpropagation, Markov decision process, Monte Carlo method, or other statistical methods, depending on the model construction method employed. Model construction may include dimensionality reduction techniques and clustering techniques.

[0043] The above-described models 330, 350, and 370 are each configured to output a confidence level of user action features based on the received sensor data stream. The confidence level may include one or more probabilities shown in the form of percentage values associated with one or more individual user action features. The confidence level may include the output layer of the neural network. The confidence level may be expressed in vector form. Such a vector may include a 1×n vector filled with the probability values of each user action feature known to the controller. For example, model 315 that receives a bioacoustic sensor data stream may receive a data stream that includes a signal having a short-duration large-amplitude portion. Model 315 may interpret that portion of the signal as a "tap" user action feature (the user tapped something) with a 90% confidence level.

[0044] At least one user action feature and / or an associated confidence level output from at least one of those models is received by the event interpreter 380. The event interpreter is configured to identify a user action (the identified user action) corresponding to the received at least one user action feature and / or the associated confidence level, based at least to some extent on the received at least one user action feature. For example, if one model indicates that a "tap" user action feature is likely to have occurred and another model indicates that the user's index finger was likely to have moved, the event interpreter concludes that a "tap with index finger" user action has occurred.

[0045] The event interpreter includes a list and / or set of user action characteristics and their respective confidence levels. User action characteristics are identified when it is received from the model that they are likely to meet at least their required confidence levels. If a conflict occurs, for example, the user action characteristic with the highest shown likelihood may be selected by the event interpreter (argmax function). The event interpreter 380 may include, for example, at least any one of a rule-based system, an inference engine, a knowledge base, a lookup table, a prediction algorithm, a decision tree, a heuristic. For example, the event interpreter may include an inference engine and a knowledge base. The inference engine outputs a set of possibilities. These sets are processed by the event interpreter, which may include IF-THEN statements. For example, (IF) when the user action characteristic relates to a particular finger, (THEN) that finger is part of the user action. The event interpreter may also incorporate and / or utilize context information, which is, for example, context information 308 (e.g., game or menu state) from another device or an application of another device. Using this type of interpreter has the advantage of higher specific accuracy. This is because the input given to the application (e.g., user interface command) is not based solely on wrist sensor inferences. Context information can be useful in XR technologies such as eye tracking, where it may be important to be able to access the user's state with respect to the controllable device.

[0046] The event interpreter 380 may be configured to receive optional context information 308. Such information may include application state, context, and / or the game state of the interacting applications and / or devices. The context information 308 may be used by the event interpreter to adjust the likelihood threshold of a user action or user action characteristic. For example, when an application prompts the user to perform a "pinch" for confirmation, the corresponding "pinch" threshold may be lowered so that a pinch action is detected even if the confidence level of the occurrence of the pinch action output by the model is low. Such adjustments may be associated with a specific application state, context, and / or the game state of the interacting applications and / or devices. The application state corresponds to the state of the application operating on the controllable device and may include, for example, at least any one of variables, static variables, objects, registers, open file descriptors, open network sockets, and / or kernel buffers.

[0047] At least one user action output from the event interpreter 380 is received by the output generator 390. The output generator 390 includes mapping information including at least any one of a list and / or set of predetermined user actions, a list and / or set of user interface commands, a set of classified gestures, and / or relationship information linking at least one predetermined user action to at least one user interface command. Such links may be one-to-one, many-to-one, one-to-many, or many-to-many. Thus, a predetermined action “summarize” may be linked, for example, to the user interface command “enter”. In another example, the predetermined actions “tap” and “double tap” may be linked to the user interface command “escape”. The output generator 390 is configured to map a user action to a predetermined user interface action and generate a user interface command 307 based at least in part on the received user action. This mapping is performed using the mapping information. The controller 302 may be configured to transmit the user interface command 307 to the device and / or store the user interface command 307 in the memory of the controller 302.

[0048] An example of the operation of the device 300 is shown below. The user is using the device 300 with it placed in their hand. The user performs an action, which includes abducting the wrist and tapping the thumb twice in succession with the middle finger. The sensors of the device 300 output data during the action as follows. · The IMU sensor 303 outputs a data stream 313 reflecting the orientation of the IMU (and thus the device 300) and acceleration data. · The wrist contour sensor 304 outputs a data stream 314 reflecting the movement of tendons and / or muscles within the user's wrist area. · The bioacoustic sensor 305 outputs a data stream 315 reflecting vibrations within the user's wrist area. Data streams 313, 314, and 315 are received by models 330, 350, and 370, respectively. In model 330, IMU sensor data is used by the model to pass a user action feature "wrist abduction" with a confidence level of 68% to event interpreter 380. Model 350 uses wrist contour data to pass two user action features "middle finger movement" with confidence levels of 61% and 82%, respectively, to event interpreter 380. Model 370 uses biometric acoustic data to pass two user action features "tap" with confidence levels of 91% and 77%, respectively, to event interpreter 380.

[0049] Accordingly, the event interpreter 380 receives the user action features "wrist abduction", "middle finger movement" (twice), and "tap" (twice), as well as their respective confidence levels. The event interpreter determines that the thresholds for those user action features are met. This determination is optionally made using context data 308 for processing (e.g., by adjusting the thresholds based on data 308). In this way, the event interpreter 380 generates user actions "abduct the wrist" and "double tap with the middle finger" based on IF-THEN rules within the event interpreter. These user actions are received by the output generator 390. The output generator uses mapping information to determine which user action links to which command and generates an appropriate user interface command. In this example, "wrist abduction" is mapped to "shift" and "double tap with the middle finger" is mapped to "double click". In this way, the output generator 390 generates a user interface command 307 that includes "shift" + "double click", and this command is sent to a device such as a personal computer. In this example, since the personal computer receives standard user interface commands that are neither sensor data nor user actions, device 300 can be used without pre-programming the personal computer.

[0050] FIG. 3B shows an exemplary schematic diagram of apparatus 400. Apparatus 400 is similar to apparatus 300, except that apparatus 400 sends data streams 413, 414, and 415 to one model 435. Model 435 may include at least one neural network, similar to models 330, 350, and 370. Such a model is configured to perform inferences on input data (e.g., signals received from sensors and / or preprocessed data). The model is configured to output at least one identified user action feature and / or respective confidence levels based on the input data. The advantage of having one model to use is that less storage device is required.

[0051] FIGS. 4A and 4B show apparatuses 500 and 600 that can support at least some embodiments of the present invention. These apparatuses are the same as apparatuses 100, 200, 300, and 400 unless otherwise specified.

[0052] FIG. 4A shows an exemplary schematic diagram of apparatus 500. As shown, controller 502 may include preprocessing sequences 520, 540, and 560 (e.g., within the memory of the controller). Preprocessing sequences 520, 540, and 560 are configured to receive signals from sensors (e.g., signals 513, 514, and 515). Preprocessing sequences 520, 540, and 560 are configured to perform preprocessing on the received signals and send the preprocessed signals to a model. The preprocessed signals may be referred to as prepared data or prepared signals. The preprocessing may include, for example, at least any one of data cleansing, feature transformation, feature engineering, and feature selection.

[0053] The preprocessing sequences 520, 540, and 560 may be configured to communicate with each other (e.g., to perform sensor fusion as already described herein). This optional communication is shown in FIG. 4 by the dotted arrows connecting the sequences, and may include adjusting the preprocessing parameters of one sequence based on information within another sequence. This information may be, for example, at least any one of parameter values, sampling values, and / or signal values. Further, the preprocessing parameters may be adjusted based on context information such as context information 508. For example, the sampling rate of sensor data may be adjusted according to the application state. In another example, when the IMU data indicates no movement, the sampling range of the wrist contour may be narrowed so that more precise movement is detected. At least one sequence may be configured to transmit data to another sequence, and at least one sequence may be configured to receive data from another sequence. For example, if one sequence includes a filter, data from another sequence may be used to adjust the filter parameters. In another example, when the IMU indicates that the user is wearing the device loosely, this information may be used to adjust the preprocessing sequence of the wrist contour sensor. The advantage of preprocessing sequence communication is that the input to the model becomes more accurate, which is because the filter characteristics are improved and this leads to an improvement in accuracy. Further, communication enables an improvement in adaptation to the user's ecology and movement, which is also due to an improvement in filter characteristics and an improvement in the cross-referenceability between sensors. Also, adaptation is improved with respect to anthropometric variations between users and movement artifacts.

[0054] FIG. 4B shows an exemplary schematic diagram of apparatus 600. Apparatus 600 includes a controller 602 and a controllable device 660. In apparatus 600, preprocessing of signals 613, 614, and 615 is performed by sequence 625 within controller 602. In apparatus 600, the preprocessed data output from sequence 625 is used by model 635 within controller 602. Sequence 625 is configured to function in a manner similar to sequences 520, 540, and 560. Further, in sequence 625, data may be combined after preprocessing (e.g., by subsampling, quantization, complementary filtering, and / or Kalman filtering). Model 635 is similar to model 435. Sequence 625 and / or event interpreter 680 may be configured to use optional context data 608. Data 608 is similar to data 308, 408, and 508. Context data 608 is also bidirectional. That is, this data may change based on at least any one of the application state of apparatus 660, the data of preprocessing state 625, and / or the data of event interpreter 680. Apparatus 600 is configured such that event interpretation and UI command generation are performed within apparatus 660.

[0055] In at least some embodiments, for example, devices similar to devices 300, 400, 500, 600 may be configured such that a single model similar to models 435, 635 is used with a plurality of preprocessing sequences similar to 520, 540, 560 (including inter-sequence communication). In at least some embodiments, for example, devices similar to devices 300, 400, 500, 600 may be configured such that a single preprocessing sequence similar to sequence 625 may be used with a plurality of models similar to models 330, 350, 370.

[0056] FIG. 5 shows, in flowchart form, an exemplary process that can support at least some embodiments of the present invention. At the top of FIG. 5 are shown a wrist contour sensor (such as, for example, sensors 104, 304, 404, 504, 604), a bioacoustic sensor (such as, for example, sensors 105, 305, 405, 505, 605), and an IMU (such as, for example, 103, 303, 403, 503, 603). Each action and each step of the flowchart may be performed at least to some extent by a device (such as, for example, the devices disclosed herein, such as a controller). At least some actions (such as, for example, the action of outputting context information) may be performed by another device (such as, for example, a controllable system, etc.).

[0057] The flowchart of FIG. 5 includes exemplary and non-limiting steps 901, 902, 903. These steps are provided only to show how a process similar to the processes shown in FIGS. 3A, 3B, and 4A, 4B would be represented in a flowchart. As shown in FIG. 5, each action need not necessarily fit within one step. Further, the actions within steps 901, 902, 903 need not be performed in parallel.

[0058] Step 901 includes adjusting the sensors. Such adjustment may be self-referential. For example, the sensors may be adjusted or calibrated according to stored reference values or known good values (such as, for example, magnetic north in the case of an IMU). Additionally or alternatively, the adjustment may be context-dependent, and the gain of the bioacoustic sensor may be adjusted according to the ambient noise level.

[0059] In the case of the bioacoustic sensor, as seen in the flowchart, optionally, a bandpass filter may be implemented. Such a filter may serve to narrow the bandwidth to only the region where the activity of interest occurs. By narrowing the bandwidth in this way, the signal-to-noise ratio is improved and the size of the data payload for subsequent processing steps is reduced.

[0060] As can be seen in the flowchart, after the acquisition of the sample (sensor data), the acquired sensor data is preprocessed. Step 902 includes preprocessing and corresponds to blocks 520, 540, and 560 in FIG. 4A and block 625 in FIG. 4B.

[0061] As can be seen in the flowchart, after preprocessing, the preprocessed data is input into each model at step 903. In the flowchart, each sensor has its own model (similar to FIGS. 3A and 4A), but a solution using a single model (similar to FIGS. 3B and 4B) can also be used. Those models in the flowchart are similar to models 330, 530, 350, 550, 370, 570. In the flowchart, the preprocessed IMU data is used to directly calculate the orientation of the IMU (and thus the orientation of the user's hand) (e.g., without a neural network). Such calculations correspond to, for example, the discussion of quaternions above.

[0062] Furthermore, at step 903, the model outputs a normalized inference output corresponding to the user action feature. "Normalization" in this context means making the sum of the confidence values of a given inference equal to 1. In the case of IMU data, the output is a quaternion reflecting the orientation of the IMU sensor. The output data corresponding to the user action feature is passed to an event interpreter (similar to interpreters 380, 480, 580, 680).

[0063] Optional context information reflecting the application state may similarly be passed to the event interpreter. This context information is similar to information 308, 408, 508, 608. As can be seen in the flowchart, this context information may alternatively or additionally be passed to at least one of steps 901 and / or 902, and this passing is optional and may include passing only to some of the adjustment block and / or the preprocessing block, for example, including passing to at least one of the adjustment block and / or the preprocessing block.

[0064] The event interpreter identifies user input, and the identified user input is passed to the application under control (e.g., in a format corresponding to the format of UI commands 307, 407, 507, 607). The application may be executed on another device or a network-connected device.

[0065] Figure 6 shows a simplified flowchart reflecting at least some of the embodiments disclosed herein. In this figure, artificial neural networks (ANNs) 861 and 862 are shown in a simplified graphical representation. For example, models 330, 530, 350, 550, 370, 570 may include such neural networks. Such models may be configured to output a confidence range in vector form (vectors 864, 865) as shown in the figure. Such vectors may include a 1×n vector filled with likelihood values for each user action feature known to the controller. For example, model 861 may include a finger classifier model. For example, model 862 may include a transient detector model.

[0066] As seen in the figure, for IMU data, quaternions may be calculated within the model (graphically represented by element 864), and quaternion orientation data (866) may be output from that model.

[0067] As seen in Figure 6, outputs 864, 865, and 866 are passed to an event interpreter 868 (shown here as a simplified block). Event interpreter 868 may be similar to interpreters 380, 480, 580, 680. Event interpreter 868 may optionally be configured to receive context information 869 from a controllable application 870. Context information 869 is similar to information 308, 408, 508, 608. Controllable application 870 may include a user space.

[0068] Figure 7 shows an apparatus 700 that can support at least some of the embodiments of the present invention.

[0069] The apparatus 700 includes a controller 702. The controller includes at least one processor and at least one memory including computer program code and optionally data. The controller 700 may further include a communication device or interface. Such a device may include, for example, a wireless and / or wired transceiver. The apparatus 700 may further include sensors (e.g., sensors 703, 704, 705) operatively connected to the controller. These sensors may include, for example, at least any one of an IMU, a bioacoustic sensor, and a wrist contour sensor. The apparatus 700 may also include other elements not shown in Figure 7.

[0070] The apparatus 700 is shown as including one processor, but may include two or more processors. In one embodiment, the memory is capable of storing instructions and may be capable of storing, for example, at least any one of an operating system, various applications, models, neural networks, and / or preprocessing sequences. Further, the memory may include a storage device (which may be used to store at least some of the information and data used in the embodiments of the present disclosure).

[0071] Furthermore, the processor is capable of executing the stored instructions. In one embodiment, the processor may be implemented as a multi-core processor, or as a single-core processor, or as a combination of one or more multi-core processors and one or more single-core processors. For example, the processor may be implemented as one or more of various processing devices, and such processing devices include coprocessors, microprocessors, controllers, digital signal processors (DSPs), processing circuits with or without DSPs, or other various processing devices. Other various processing devices include, for example, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontroller units (MCUs), hardware accelerators, dedicated computer chips, and other integrated circuits. In one embodiment, the processor may be configured to execute hard-coded functions. In one embodiment, the processor is implemented as an executor of software instructions, and when those instructions are executed, the processor may be specifically configured such that the processor implements at least any one of the models, sequences, algorithms, and / or operations described herein.

[0072] The memory may be implemented as one or more volatile memory devices, and / or as one or more non-volatile memory devices, and / or as a combination of one or more volatile memory devices and one or more non-volatile memory devices. For example, the memory may be implemented as semiconductor memory (e.g., mask ROM, PROM (programmable ROM), EPROM (erasable PROM), flash ROM, RAM (random access memory), etc.).

[0073] The at least one memory and the computer program code, together with the at least one processor, are configured to cause the apparatus 700 to at least perform steps of receiving a first sensor data stream from at least one wrist contour sensor, receiving a second sensor data stream from at least one bioacoustic sensor, receiving a third sensor data stream from at least one inertial measurement device (wherein the first, second, and third sensor data streams are received in parallel), identifying at least one characteristic of a user action based on at least any one of the first, second, and third sensor data streams, identifying at least one user action based on the identified at least one characteristic of the user action, and generating at least one user interface (UI) command based at least to some extent on the identified at least one user action. These actions have already been described in detail with respect to FIGS. 1 - 6.

[0074] In at least some embodiments, an apparatus capable of supporting the present invention is configured such that the sensors of the apparatus fit within a single housing. Examples of such single - housing apparatuses are shown below. Placing the sensors (e.g., IMU 853, wrist contour sensor 854, and bioacoustic sensor 855) within a single housing results in a more robust structure.

[0075] FIGS. 8A and 8B show an apparatus 850 capable of supporting at least some embodiments of the present invention. The apparatus 850 may be exactly the same as the apparatuses 100, 200, 300, 400, 500, 600, and / or 700.

[0076] Figure 8A shows a side view of device 850. Device 850 includes a controller 852, an IMU 853, a wrist contour sensor 854, and / or a bioacoustic sensor 855. The wrist contour sensor 854 may include one or more contour measurement sensors 856 (e.g., arranged as an array). In one embodiment, device 850 includes sensors 853, 854, and 855 within a single housing 858. As can be seen from Figure 8A, sensors 854 and 855 are positioned such that they can measure signals from outside the housing within the housing. For example, these sensors may be connected to ports of the housing or disposed near the surface of the housing. That is, the signals measured by the wrist contour sensor 854 and the bioacoustic sensor 855 are from outside the housing. In at least some embodiments, device 850 may include mounting components 851 which may be, for example, a strap. In at least some embodiments, device 850 may include a smartwatch. In at least some embodiments, device 850 may include at least any one of components such as, for example, a haptic device, a screen, a touch screen, a speaker, a heart rate sensor, a Bluetooth communication device, etc.

[0077] Figure 8B shows a bottom view of device 850. The wrist contour sensor 854 includes an array of contour measurement sensors 856, which are visible on the surface of housing 858. This surface may be concave in shape. The array of contour measurement sensors 856 may be configured to detect the user's wrist contour measurements when the user is wearing device 850. The bioacoustic sensor 855 is also visible on the surface of housing 858. Sensor 855 may be configured to measure bioacoustic signals related to the user when the user is wearing device 850. Device 850 may be configured to acquire signals from sensors 853, 854, and 855 in parallel.

[0078] Figure 9 shows an apparatus 920 that can support at least some embodiments of the present invention. The apparatus 920 may be exactly the same as apparatuses 100, 200, 300, 400, 500, 600, and / or 700.

[0079] Figure 9 shows an isometric bottom view of the apparatus 920. The apparatus 920 includes a controller, an IMU, a wrist contour sensor including a contour measurement sensor 926, and / or a bioacoustic sensor 925 within a housing 928. An array of the contour measurement sensors 926 is seen on the surface of the housing 928. This surface may be concave in shape. The array of the contour measurement sensors 926 may be configured to detect wrist contour measurements of a user when the user is wearing the device 920. The bioacoustic sensor 925 is also seen on the surface of the housing 928. The sensor 925 may be configured to measure bioacoustic signals related to the user when the user is wearing the device 920.

[0080] As can be seen from Figure 9, the array of sensors 926 and the bioacoustic sensor 925 may be arranged annularly on the surface of the housing 928. With such a configuration, it is possible to arrange other components (e.g., sensors) in the central region 929 of the housing 928. That is, by arranging the array of sensors 926 and the sensor 925 along the peripheral portion of the housing 928 (e.g., annularly and / or rectangularly), it becomes possible to arrange other sensors in the remaining space, thus facilitating the accommodation of other sensors on the same surface of the housing 928. Further, those other sensors may be configured to be spaced apart from them and the contour measurement sensors to minimize any potential interference. For example, a heart rate sensor may be conveniently arranged in the central region 929. With such an arrangement, possible interference (e.g., optical interference) from the contour measurement sensors is minimized.

[0081] Device 920 may be configured to acquire signals from the IMU and signals from sensors 925 and 926 in parallel. In at least some embodiments, device 920 may include a smartwatch. In at least some embodiments, device 920 may include at least any one of components such as, for example, a haptic device, a screen, a touch screen, a speaker, a heart rate sensor, a Bluetooth communication device, and the like.

[0082] Figures 10A and 10B show an apparatus 940 that can support at least some embodiments of the present invention. FIG. 10A shows a bottom view of the apparatus 940. FIG. 10B shows an isometric cutaway top view of the apparatus 940. The apparatus 940 includes a controller, an IMU 943, a wrist contour sensor 944, and / or a bioacoustic sensor 945. The wrist contour sensor 944 may include one or more contour measurement sensors 946 (arranged, for example, as a concentric array as seen in FIG. 10A), or may be connected to such contour measurement sensors 946. The number of rings included in this array may be from 1 to 24, preferably from 3 to 6, and particularly 3. The number of rings may depend on the form factor of the device. In this context, a ring means any geometric arrangement of at least one sensor spaced along, for example, an open or closed curve. For example, loops including rings or rectangular rings, elliptical rings, straight lines, or X - shapes are suitable. For example, the innermost loop may include 6 sensors spaced along that loop, the second - innermost loop may include 12 sensors spaced along that loop, and the outermost loop may include 16 sensors spaced along that loop. The loops may be circular, rectangular, or other shapes. The apparatus 940 includes the sensors 943, 944, and 945 within a single housing 948. The sensors 946, 945 are disposed on the surface of the housing to measure signals coming from outside the housing. This surface may be concave in shape. In at least some embodiments, the apparatus 940 may include a smartwatch. In at least some embodiments, the apparatus 940 may include at least any one of components such as, for example, a haptic actuator, a screen, a touch screen, a speaker, a heart rate sensor, a Bluetooth communication device, etc.

[0083] Figure 10B shows an isometric cutaway top view of apparatus 940. In this figure, components such as the screen and / or upper casing and / or bezel are omitted to show the interior of housing 948. A circuit board is mounted within housing 948. Components such as IMU 943, processor 949, and memory 949B may be mounted on this circuit board. This processor and this memory may together include a controller of apparatus 940. Sensors including sensors 943, 944, 946, and 945 may be directly or indirectly connected to processor 949A and / or memory 949B such that sensor data is passed to the controller.

[0084] Figure 11 shows an apparatus 970 that can support at least some embodiments of the present invention. Figure 11 shows a bottom view of apparatus 970. Apparatus 970 is attachable to a wearable device (e.g., a smartwatch). That is, apparatus 970 is an example of an add-on device. An add-on device is designed to be used alone or in combination with another device and / or in combination with mounting components. An add-on device may, in at least some embodiments, not include mounting components and / or a screen. An add-on device may communicate wirelessly or via a port with another device (e.g., a wearable device). An add-on device may be used with another wearable device in at least the following ways. · The add-on device may be directly attached to the body of another wearable device (e.g., as seen in Figure 12) so as to be able to measure signals related to the user or the user's gestures. · The add-on device may be attached to a mounting component of another wearable device, where the mounting component is a strap (e.g., the same strap as the strap of another wearable device) or a bracelet (e.g., in the case of an activity bracelet). · The add-on device may be part of a mounting component such as a strap or may be connected to the mounting component. For example, as shown in FIG. 11, one side of the add-on device may be connected to another wearable device and the other side may be connected to a strap.

[0085] Device 970 includes a controller, an IMU, a wrist contour sensor, and / or a bioacoustic sensor 975 within a single housing. The wrist contour sensor may include an array of contour measurement sensors, for example, may include sensors 976 arranged on the surface of the housing. This surface may be concave in shape. The array of contour measurement sensors may be configured to detect wrist contour measurements of the user when the user is wearing device 970. The bioacoustic sensor 975 is also arranged on the surface of the housing. The wrist contour sensor may be configured to measure bioacoustic signals related to the user when the user is wearing device 970. As shown in FIG. 11, the measurement area corresponding to the sensor may vary according to embodiments of the present disclosure. For example, in FIG. 9, the measurement area is typically in the central part of the user's wrist, while in the embodiment of FIG. 11, the measurement area is closer to the side of the user's wrist.

[0086] As shown in FIG. 11, for device 970, the left end of the device is connected to a mounting component (e.g., a strap), and the right end of the device is connected to a wearable device. The connection to the mounting component may be made by any suitable means, for example, by a mounting bracket and / or an adhesive. The connection to the wearable device may be made, for example, by a pin 979 (e.g., a watch pin) that mates with the wearable device. As can be seen from FIG. 11, sensors 975 and 976 are located on the surface of the housing so as to be able to measure signals from the outside of the housing.

[0087] In one exemplary embodiment, sensors 975 and 976 may be arranged in an annular pattern similar to the pattern of device 920 or the pattern of device 940. In another exemplary embodiment, as seen in FIG. 11, the sensors may be arranged such that the rows and columns are orthogonal, similar to the pattern of device 850. Device 970 may be configured to acquire signals from the IMU and signals from sensors 975 and 976 in parallel. Device 970 may be configured to wirelessly (e.g., using Bluetooth) transmit data (e.g., user actions and / or user interface frames) to another device. The receiving side of such data transmission may be at least one of the attached wearable device, any wearable device, or a smartphone.

[0088] FIG. 12 shows a device 950 that can support at least some embodiments of the present invention. FIG. 12 shows a side view of device 950. Device 950 is attachable to a wearable device (e.g., a smartwatch). That is, device 950 is an example of an add-on device. Device 950 may be attached to the wearable device by any suitable means, e.g., by at least one of mechanical friction such as snap fit (press fit) and / or screwing, adhesives, magnetic connection, fitting by thermal expansion.

[0089] Device 950 includes a controller, an IMU, a wrist contour sensor, and / or a bioacoustic sensor 955 within a housing 958. An array of contour measurement sensors 956 is seen on the surface of housing 958. This surface may be concave in shape. The array of contour measurement sensors 956 may be configured to detect the user's wrist contour measurements when the user is wearing device 950. The bioacoustic sensor 955 is also seen on the surface of housing 958. Sensor 955 may be configured to measure bioacoustic signals related to the user when the user is wearing device 950.

[0090] As can be seen from FIG. 12, sensors 955 and 956 are located on the surface of the housing so as to be able to measure signals from outside the housing. In one exemplary embodiment, sensors 955 and 956 may be arranged in an annular pattern similar to the pattern of device 920 or the pattern of device 940. In another exemplary embodiment, the sensors may be arranged such that the rows and columns are orthogonal, similar to the pattern of device 850. Device 950 may be configured to acquire signals from the IMU and signals from sensors 955 and 956 in parallel. Device 950 may be configured to wirelessly (e.g., using Bluetooth) transmit data (e.g., user actions and / or user interface frames) to another device. The receiving side of such data transmission may be at least any one of the attached wearable device, any wearable device, and a smartphone.

[0091] The present disclosure may also be utilized by the following clauses.

[0092] Clause 1. A non-transitory computer-readable medium storing a set of computer-readable instructions, which, when executed by at least one processor, at least measuring parallel signals corresponding to at least one movement of a user using at least one wrist contour sensor, at least one bioacoustic sensor including a vibration sensor, and at least one movement sensor such as an IMU; receiving, in a controller including at least one processor and at least one memory including computer program code, the measured parallel signals; combining, by the controller, the received parallel signals; generating, by the controller, a human interface data (HID) data frame based at least to some extent on the characteristics of the combined received signals; A non-transitory computer-readable medium that causes a device to perform.

[0093] Clause 2. A multimodal biological measurement device, a strap, at least one wrist contour sensor, at least one bioacoustic sensor including a vibration sensor, at least one motion sensor, a controller including at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code thereof, together with the at least one processor, at least receive parallel signals from the wrist contour sensor, the bioacoustic sensor, and the motion sensor; combine the received signals; generate a Human Interface Data (HID) data frame based at least to some extent on the characteristics of the combined received signals; and a controller configured to cause the controller to perform the above steps, and a device including the same.

[0094] Clause 3. The device according to Clause 2, wherein at least one wrist contour sensor, at least one bioacoustic sensor, and at least one motion sensor are arranged within a single housing.

[0095] Clause 4. The device according to Clause 2 or 3, including a smartwatch.

[0096] Clause 5. The wrist contour sensor includes an array of contour measurement sensors, and the array is annularly arranged on the surface of a single housing together with at least one bioacoustic sensor, and the surface is preferably concave. The device according to any one of Clauses 2 to 4.

[0097] Clause 6. The wrist contour sensor includes an array of contour measurement sensors that form concentric circles on the surface of a single housing. The number of annuli included in this array is from 1 to 25, preferably from 2 to 6, and particularly 3. The device according to any one of Clauses 2 to 5.

[0098] Clause 7. The device includes an add-on device and is configured to be attached to another wearable device by at least one of the following configurations, namely, The add-on device is configured to be directly attached to the body of another wearable device such that the sensors of the add-on device can measure the user (e.g., such that the sensors face the user), and / or, The add-on device is configured to be attached to a mounting component of another wearable device. For example, the mounting component is a strap, and in particular, the same strap as the strap of another wearable device, and / or, The add-on device is configured to be part of or connected to a mounting component of another wearable device (where the mounting component is a strap). For example, one side of the add-on device is connected to another wearable device and the other side is connected to the strap. The device according to any one of Clauses 2 to 6.

[0099] Clause 8. The controller is arranged within a single housing. The device according to any one of Clauses 2 to 7.

[0100] Clause 9. The generating includes identifying at least one user action feature based on the characteristics of the combined received signals. The identifying includes inputting the received signals into at least one neural network trained to identify the at least one user action feature and output a confidence level corresponding to the at least one user action feature. The method according to Clause 1, or the device according to any one of Clauses 2 to 8.

[0101] Clause 10. A computer program configured to implement the method according to Clause 1 or Clause 9, which is storable on a non-transitory computer-readable medium.

[0102] The present disclosure has the following advantages. That is, when using multimodal sensing, the specific accuracy in detecting user actions is increased. This is because the detected action is specified based on the outputs of a plurality of sensors. Further, this device and method are highly adaptable to individual users. This is because each sensor can be configured to consider more the user-specific ecology and habits.

[0103] Naturally, the embodiments of the present invention disclosed herein are not limited to the specific structures, processing procedures, or materials disclosed herein, and are extended to their equivalents, which will be understood by those skilled in the art. Further, naturally, the terms used herein are only used for the description of specific embodiments and are not intended to be limiting.

[0104] References to one embodiment or an embodiment throughout this specification mean that the specific features, structures, or characteristics described in connection with that embodiment are included in at least one embodiment of the present invention. Accordingly, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. For example, when a numerical value is referenced using terms such as about or substantially, the exact numerical value is also disclosed.

[0105] A plurality of items, structural elements, compositional elements, and / or materials used herein may, for convenience, be present in a general list. However, these lists should be interpreted as if each element of the list is separately and uniquely identified as a distinct element. Thus, the individual elements of such a list should be interpreted as de facto equivalents of any other element of the same list only based on the fact that they are present in a common group, unless the opposite meaning is indicated. Further, in this specification, various embodiments and examples of the present invention may be referred to in conjunction with alternative forms with respect to their various components. Of course, such embodiments, examples, and alternative forms should not be interpreted as de facto equivalents of each other, but rather should be regarded as distinct and independent representations of the present invention.

[0106] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this description, various specific details such as examples of length, width, shape, etc. are shown so as to obtain a sufficient understanding of the embodiments of the present invention. However, as will be understood by those skilled in the art, the present invention can be implemented even without one or more of these specific details, or with other methods, components, materials, etc. In other examples, well-known structures, materials, or operations are not illustrated or described in detail, which is to prevent the aspects of the present invention from being ambiguous.

[0107] Each of the above examples illustrates the principles of the present invention in one or more specific applications. However, as will be apparent to those skilled in the art, various changes in the form, usage, and details of the embodiments may be made without exercising inventive faculty and without departing from the principles and concepts of the present invention. Accordingly, the present invention is not to be limited except as defined by the claims set forth below.

[0108] In this document, the verbs "to comprise" and "to include" are used as open limitations that do not require the exclusion or the presence of features not described. Features recited in the dependent claims may be freely combined with each other unless otherwise expressly specified. Further, of course, the use of "a" or "an", i.e., the singular, does not exclude plurality throughout this document. 〔Appendix 1〕 A multimodal biometric measurement device, comprising: a mounting component configured to be worn by a user; at least one wrist contour sensor; at least one bioacoustic sensor including a vibration sensor; at least one inertial measurement unit (IMU) including an accelerometer and a gyroscope; a controller including at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code, together with the at least one processor, are configured to at least: receive a first sensor data stream from the at least one wrist contour sensor; receive a second sensor data stream from the at least one bioacoustic sensor; receive a third sensor data stream from the at least one inertial measurement unit (wherein the first, second, and third sensor data streams are received in parallel), identify at least one feature of a user action based on at least any one of the first, second, and third sensor data streams; identify at least one user action based on the at least one identified feature of the user action; generate at least one user interface (UI) command based on at least to some extent on the at least one identified user action; and the controller configured to cause the controller to perform the above steps, and a device including the same. 〔Appendix 2〕 The step of identifying the at least one feature of the user action includes supplying at least any one of the first, second, and third sensor data streams as an input to at least one neural network trained to identify and output at least one feature of the user action, and the output of the at least one neural network includes a confidence value of the at least one feature of the user action. The device according to Appendix 1. [Appendix 3] The step of identifying the at least one user action includes supplying the output of the at least one neural network to an event interpreter configured to identify the user action based on the identified user action features and received context information, the apparatus according to Appendix 2. [Appendix 4] The controller is configured to perform preprocessing on at least any one of the first, second, and third sensor data streams and then supply the at least one preprocessed data stream as an input to the at least one neural network, the apparatus according to any one of Appendices 2 to 3. [Appendix 5] The controller is configured such that at least any one of the first, second, and third sensor data streams is preprocessed in a preprocessing sequence separate from other data streams, the apparatus according to any one of Appendices 1 to 4. [Appendix 6] The controller is configured such that at least one preprocessing sequence communicates with at least one other preprocessing sequence, and the communication includes adjusting preprocessing parameters of the other preprocessing sequence based on data from the at least one preprocessing sequence, the apparatus according to any one of Appendices 4 to 5. [Appendix 7] The controller is further configured such that the event interpreter includes at least one of an inference engine and a decision tree, the apparatus according to any one of Appendices 1 to 6. [Appendix 8] The wrist contour sensor includes an array of contour measurement sensors, the apparatus according to any one of Appendices 1 to 7. [Appendix 9] The at least one wrist contour sensor, the at least one bioacoustic sensor, and the at least one inertial measurement device are arranged in a single housing. For example, the apparatus includes a smartwatch, the apparatus according to any one of Appendices 1 to 8. [Appendix 10] The wrist contour sensor includes an array of contour measurement sensors, and the array is arranged annularly on the surface of the single housing together with the at least one bioacoustic sensor, or the wrist contour sensor includes an array of contour measurement sensors that form a concentric ring or a rectangular ring on the surface of the single housing, and the number of rings included in the array is from 1 to 25, preferably from 2 to 6, and particularly 3. The device according to any one of Appendices 1 to 9. [Appendix 11] The device includes an add-on device, and the device is configured to be attached to another wearable device by at least any one of the following configurations, that is, the add-on device is configured to be directly attached to the body of another wearable device so that the sensor of the add-on device can measure the user (for example, so that the sensor faces the user), and / or, the add-on device is configured to be attached to a mounting component of another wearable device. For example, the mounting component is a strap, and particularly, it is the same strap as the strap of another wearable device, and / or, the add-on device is a part of the mounting component of another wearable device or is configured to be connected to the mounting component (for example, the mounting component is a strap). Specifically, one side of the add-on device is connected to the other wearable device, and the other side is configured to be connected to the strap. The device according to any one of Appendices 1 to 10. [Appendix 12] A method for generating a UI command, comprising: receiving a first sensor data stream from at least one wrist contour sensor; receiving a second sensor data stream from at least one bioacoustic sensor; receiving a third sensor data stream from at least one inertial measurement device (wherein the first, second, and third sensor data streams are received in parallel), identifying at least one characteristic of a user action based on at least any one of the first, second, and third sensor data streams. identifying at least one user action based on the at least one identified characteristic of the user action; generating at least one user interface (UI) command based at least in part on the at least one identified user action; A method comprising: [Appendix 13] The step of identifying the at least one characteristic of the user action includes supplying at least any one of the first, second, and third sensor data streams as an input to at least one neural network trained to identify at least one characteristic of the user action and output it, and the output of the at least one neural network includes a confidence value of the at least one characteristic of the user action. The method according to Appendix 12. [Appendix 14] The step of identifying the at least one user action includes supplying the output of the at least one neural network to an event interpreter configured to identify the user action based on the identified user action characteristics. The method according to any one of Appendices 12 to 13. [Appendix 15] After preprocessing is performed on at least any one of the first, second, and third sensor data streams, the at least one preprocessed data stream is supplied as an input to the at least one neural network, and at least any one of the first, second, and third sensor data streams is preprocessed in a preprocessing sequence separate from other data streams, and at least one preprocessing sequence communicates with at least one other preprocessing sequence, and the communication includes adjusting preprocessing parameters of the other preprocessing sequence. The method according to any one of Appendices 12 to 14. [Appendix 16] A computer program configured to implement the method according to any one of Appendices 12 to 15, the computer program being storable on a non-transitory computer-readable medium.

Industrial Applicability

[0109] At least some embodiments of the present invention have industrial applications in providing a user interface (e.g., related to XR) to a controllable device (e.g., a personal computer).

Description of Symbols

[0110] 100, 200, 300, 400, 500, 600, 700, 850, 920, 940, 950, 970 Devices Components for mounting 101, 201, 851, 921 Controllers 102, 302, 402, 502, 602, 702, 852 Inertial measurement units (IMUs) 103, 303, 403, 503, 603, 703, 853, 943 Wrist contour sensors 104, 304, 404, 504, 604, 704, 854, 944 Bioacoustic sensors 105, 305, 405, 505, 605, 705, 855, 925, 945, 975, 955 Contour measurement sensors 106, 856, 926, 946, 976, 956 Wrist contour Bioacoustic measurement area IMU Wrist contour measurement area Housing protrusion User's hand UI commands 307, 407, 507, 607 Sensor data streams 313, 413, 513, 613 Sensor data streams 314, 414, 514, 614 Sensor data streams 315, 415, 515, 615 Models 330, 530, 861 Models 350, 550, 862 Models 370, 570, 863 Event interpreters 380, 480, 580, 680, 868 Output generators 390, 490, 590, 690 Controllable device 660 Models 435, 635 Preprocessing sequences 520, 540, 560 Preprocessing sequence 625 Context information 308, 408, 508, 608, 869 Housings 858, 928, 948 Central area of the housing Confidence level vectors 864, 865 866 Azimuth data (quaternion) 870 Controllable device 801, 802, 803, 804 Signals 901, 902, 903 Steps of method

Claims

1. A multimodal biometric measurement device, comprising: a mounting component configured to be worn by a user; at least one wrist contour sensor; at least one bioacoustic sensor including a vibration sensor; at least one inertial measurement unit (IMU) including an accelerometer and a gyroscope; a controller including at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code, together with the at least one processor, are configured to perform at least: receiving a first sensor data stream from the at least one wrist contour sensor; receiving a second sensor data stream from the at least one bioacoustic sensor; receiving a third sensor data stream from the at least one inertial measurement unit (wherein the first, second, and third sensor data streams are received in parallel), performing preprocessing on at least one of the first, second, and third sensor data streams in at least one preprocessing sequence and then supplying the at least one preprocessed data stream as an input to at least one neural network, wherein at least one first preprocessing sequence communicates with at least one second preprocessing sequence, and the communication includes adjusting preprocessing parameters of the second preprocessing sequence based on data from the at least one first preprocessing sequence; identifying at least one feature of a user action based on at least one of the preprocessed first, second, and third sensor data streams; identifying at least one user action based on the identified at least one feature of the user action; generating at least one user interface (UI) command based at least to some extent on the identified at least one user action; and the controller configured to cause the controller to perform the above. The step of identifying the at least one feature of the user action includes supplying at least any one of the preprocessed first, second, and third sensor data streams as an input to the at least one neural network, and the at least one neural network is trained to identify at least one feature of the user action as an output of the at least one neural network, and the output of the at least one neural network includes a confidence value of the at least one feature of the user action. Device. **Claim 2** The step of identifying the at least one user action includes supplying the output of the at least one neural network to an event interpreter configured to identify the user action based on the identified user action feature and the received context information. The device according to claim 1. **Claim 3** The controller is configured such that at least any one of the first, second, and third sensor data streams is preprocessed in a preprocessing sequence separate from other data streams. The device according to claim 1 or 2. **Claim 4** The controller is further configured such that the event interpreter includes at least one of an inference engine and a decision tree. The device according to any one of claims 2. **Claim 5** The wrist contour sensor includes an array of contour measurement sensors. The device according to any one of claims 1 to 4. **Claim 6** The at least one wrist contour sensor, the at least one bioacoustic sensor, and the at least one inertial measurement device are arranged in a single housing. The device according to any one of claims 1 to 5. **Claim 7** The wrist contour sensor includes an array of contour measurement sensors, and the array is arranged annularly on the surface of the single housing together with the at least one bioacoustic sensor. The device according to claim 6. **Claim 8** The wrist contour sensor includes an array of contour measurement sensors that form a concentric ring or a rectangular ring on the surface of the single housing, and the number of rings included in the array is from 1 to 25. The device according to claim 6. **Claim 9** The device includes an add-on device, and the device is configured to be attached to another wearable device by at least one of the following configurations, that is, the add-on device is configured to be directly attached to the body of another wearable device so that the at least one wrist contour sensor, the at least one bioacoustic sensor including a vibration sensor, and the at least one inertial measurement unit (IMU) can measure the user; the add-on device is configured to be attached to a mounting component of the other wearable device, and / or the add-on device is configured to be part of a mounting component of the other wearable device or to be connected to the mounting component. The device according to any one of claims 1 to 8.

10. The mounting component is a strap. The device according to claim 9.

11. The preprocessing parameter is adjusted based on context information reflecting the application state. The device according to claim 1.

12. A method for generating a UI command, comprising: receiving a first sensor data stream from at least one wrist contour sensor; receiving a second sensor data stream from at least one bioacoustic sensor; receiving a third sensor data stream from at least one inertial measurement unit (wherein the first, second, and third sensor data streams are received in parallel), performing preprocessing on at least one of the first, second, and third sensor data streams in at least one preprocessing sequence and then supplying at least one preprocessed data stream as an input to at least one neural network, wherein at least one first preprocessing sequence communicates with at least one second preprocessing sequence, and the communication includes adjusting preprocessing parameters of the second preprocessing sequence based on data from the at least one first preprocessing sequence; identifying at least one feature of a user action based on at least one of the preprocessed first, second, and third sensor data streams. identifying at least one user action based on the at least one identified feature of the user action; generating at least one user interface (UI) command based at least in part on the at least one identified user action; the step of identifying the at least one feature of the user action includes supplying at least any one of the pre-processed first, second, and third sensor data streams as an input to the at least one neural network, the at least one neural network being trained to identify at least one feature of the user action as an output of the at least one neural network, the output of the at least one neural network including a confidence value of the at least one feature of the user action; Method. **Claim 13** The step of identifying the at least one user action includes supplying the output of the at least one neural network to an event interpreter configured to identify the user action based on the identified user action feature. The method according to claim 12. **Claim 14** A computer program configured to implement the method according to any one of claims 12 to 13, the computer program being storable on a non-transitory computer-readable medium.

Citation Information

Patent Citations

  • Wearable controller for wrist

    EP3203350A1

  • Operation signal output unit and operation system

    JP2008027252A

  • Method and apparatus for combining myoelectric sensor signals and inertial sensor signals for gesture-based control

    JP2016507851A

  • Camera-guided interpretation of neuromuscular signals

    JP2021535465A

  • Wrist worn computing device control systems and methods

    US20210124417A1